<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/"><channel><title>Ænix Blog</title><link>https://aenix.io/blog/</link><description>Articles, deep dives, news, and field notes from the Ænix team — Cozystack, Kubernetes, sovereign cloud, DORA / NIS2, AI infrastructure.</description><language>en</language><lastBuildDate>Thu, 08 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://aenix.io/blog/" rel="self" type="application/rss+xml"/><item><title>Paleocomputing, part 2: Kubernetes in Oberon, Wirth's radio instead of a network, a cluster in your browser, and what forgotten technology says about tomorrow's infrastructure</title><link>https://aenix.io/blog/2026/10/kubernetes-over-wirths-radio/</link><guid isPermaLink="true">https://aenix.io/blog/2026/10/kubernetes-over-wirths-radio/</guid><pubDate>Thu, 08 Oct 2026 00:00:00 +0000</pubDate><dc:creator>Timur Tukaev</dc:creator><category>Kubernetes</category><category>Cozystack</category><category>KubeVirt</category><category>Open Source</category><category>Retrocomputing</category><category>Distributed Systems</category><description>Kube: a Kubernetes control plane in Oberon, on Niklaus Wirth's own machines talking over radio. Six browser labs, the design lessons, and how to run it.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/kubernetes-over-wirths-radio.jpg" alt=""></p><p>Picture a Kubernetes cluster without a single network card. Its nodes know nothing of TCP/IP, or even of Ethernet; they call to each other over the radio in short 32-byte packets, like foremen on walkie-talkies at a building site. The control plane is written in a language from the late eighties and, network protocol included, takes about 1,250 lines. Starting a pod involves no download: the node simply loads a module and calls a procedure in it. And best of all, the whole thing opens in a browser tab, where you can pull a node&rsquo;s plug, jam the air, slip the cluster a forged command and watch what happens.</p>
<p>I built this cluster over the last few weeks and called it Kube. It is a Kubernetes control plane written in Oberon, Niklaus Wirth&rsquo;s language, and it runs on Oberon machines, that is, on the processor and operating system Wirth designed himself, from the circuit to the windows on the screen. Kube&rsquo;s nodes are Oberon machines too, and they talk over a radio network from the same book: Wirth wrote that network for his workstations back in the late eighties, and the new edition of the project moved it onto a cheap radio module.</p>
<p>Let me say straight away why I am doing this, since in stories like this one that question always comes first, and usually with a smirk. For me this is not a project for laughs, nor an attempt to drag a pretty old technology into the light for people to admire and move on. It is research first and foremost. I want to understand what the infrastructure of the past might have become if the ideas we now call Kubernetes had appeared in the years when a computer was small, wholly understandable by one person, and connected to its neighbours by whatever was at hand. What interesting principles the forgotten technologies carried, and what we threw out along with them. What those technologies lacked, and why they were abandoned. And if you honestly sort out both, you can try to look into tomorrow and imagine an infrastructure where complexity need not grow faster than usefulness.</p>
<p>Kubernetes is a convenient point of reference here. Its main idea fits in one sentence: you do not tell the system what to do, you describe what should exist, and a set of independent loops compares the wanted with the actual over and over and closes the gap step by step. The idea knows nothing about Go, containers, etcd or clouds, and it can be carried over to any machine. Wirth&rsquo;s machine suits it better than any, because it can be read in full, from the processor&rsquo;s registers to the pod scheduler, with nothing hidden under a framework. When you move a big system into such a small one, it immediately becomes clear what in it is essence and what is sediment, and which decisions were forced and which were mere habit.</p>
<p>This is the second part of my series on paleocomputing. In <a href="https://aenix.io/blog/2026/09/nine-days-of-paleocomputing/">the first part</a> I told how I ran Wirth&rsquo;s processor in a browser, in QEMU, in Kubernetes and in Cozystack, where an Oberon machine can now be installed from the catalog with one click. You do not need to read it: whatever concerns Oberon and Wirth&rsquo;s machine I will explain along the way, and Kubernetes you probably know as well as I do. The article came out long again. First I will tell how Kube is built and how it works, then about the principles the Oberon language and operating system imposed on it, how it differs from real Kubernetes, where it turned out better and where noticeably worse. After that come six labs you can do right in the browser, instructions for those who want to install it all themselves, and at the end my conclusions about what such experiments can teach the people who build infrastructure today.</p>
<p>If you would rather touch first and read later, open the <a href="https://tym83.github.io/paleocomputing/oberon/kube.html">cluster lab</a>. Nothing needs installing. Within half a minute three machines boot, the control plane starts Kube, the nodes join, and a cluster table appears on the right, which the page builds simply by listening to the air. Sources, documentation and every measurement are in the repository <a href="https://github.com/tym83/paleocomputing">github.com/tym83/paleocomputing</a>.</p>
<figure><img src="en-02-formed.png" alt="Three Oberon machines in a browser tab" width="1440" height="900" loading="lazy" decoding="async">
  <figcaption><p>Three Oberon machines in a browser tab. On the left their screens, on the right the cluster table built from the air and the labs, which check themselves</p></figcaption>
</figure>

<h2 id="how-kube-is-built">How Kube is built</h2>
<h3 id="oberon-and-wirths-machine-in-brief">Oberon and Wirth&rsquo;s machine in brief</h3>
<p>Oberon is both a programming language and an operating system. They were made in the mid-eighties at ETH Zurich by Niklaus Wirth, the author of Pascal and Modula-2, and Jürg Gutknecht. The language is tiny: modules, records, arrays, pointers, garbage collection and strict typing, while exceptions, generics and even unsigned integers are absent. The system matches the language. It has no processes, no threads and no memory protection. At the very bottom spins a single loop, <code>Oberon.Loop</code>: it polls the mouse and keyboard, calls commands one at a time, and in the pauses between input events calls background tasks, <code>Oberon.Task</code>. A task cannot be interrupted, so it must do its piece of work quickly and return. Multitasking here, in other words, is cooperative and rests on the good manners of everyone involved.</p>
<p>In 2013 Wirth published a new edition of the book Project Oberon and added his own processor to the system, RISC5, described in Verilog. Loaded into an FPGA, it ran on a board with a megabyte of memory at 25 MHz. That is Wirth&rsquo;s machine. With us it lives in three guises: in the browser runs Wirth&rsquo;s actual circuit, translated to C++ and compiled to WebAssembly; in QEMU runs our own model of the machine; and in Kubernetes the same QEMU model runs inside KubeVirt.</p>
<h3 id="where-it-all-started">Where it all started</h3>
<p>The first version of Kube was a pure toy, and I said so honestly in a separate episode of the series. It was a single Oberon module holding an object store and three controllers, for deployments, ReplicaSets and nodes. The nodes were three names, <code>node-a</code>, <code>node-b</code> and <code>node-c</code>, and a pod counted as running only because a controller had written a node&rsquo;s name into it. Nothing ran anywhere. But even that version showed very clearly that the heart of Kubernetes is not containers but reconcile loops, and that these loops fit Oberon beautifully.</p>
<p>They fit because Oberon already has everything a control plane needs, namely a central loop that calls background work. <code>Kube.Start</code> installs the three controllers as three <code>Oberon.Task</code>s with a period of 50 milliseconds, in the same ring where the system&rsquo;s garbage collector spins, and from then on <code>Oberon.Loop</code> calls them whenever the user is not typing or moving the mouse. The controllers do not call each other or send each other anything. Each looks only at objects of its own kind and changes only what it owns, and the result of its pass becomes the input of a neighbour on its next pass. That is exactly how the controllers of real Kubernetes interact: through shared state, not through messages.</p>
<p>To call all this a cluster, two things were missing. First, nodes, meaning several machines and a way for them to talk. Second, rollouts, where one version of an application replaces another without downtime. With nodes I thought at first that I had hit a dead end, since Wirth&rsquo;s machine has no network. Then I reread the book and found the radio in it.</p>
<h3 id="wirths-radio">Wirth&rsquo;s radio</h3>
<p>The Oberon workstations at ETH were always networked. Wirth wrote that network back in the late eighties for the wired Ceres workstation, and in the 2013 edition Paul Reed, who worked with him on the new version of the project, moved it onto radio. The board carries an nRF24L01+ transceiver, a cheap chip still found today in wireless keyboards and hobby gadgets, and the processor talks to it over SPI, a simple serial bus. The chip is very modest. It sends a frame of at most 32 bytes at a time, received frames wait their turn in a three-slot buffer, all stations on one channel hear each other, and there is no collision avoidance at all: if two stations speak at once, the receiver gets either garbage or nothing.</p>
<p>The chip&rsquo;s driver, the module <code>SCC</code>, is about two hundred lines. Each of its packets starts with an eight-byte header holding addresses, a type and a length, and a long packet is cut into several 32-byte frames. On top of <code>SCC</code> sits the module <code>Net</code>, a small protocol stations use to send each other messages and files. From here on, by Wirth&rsquo;s radio I will mean exactly this pair, and by the air everything the stations on one channel hear.</p>
<p>Why radio rather than an ordinary network? Because Wirth&rsquo;s machine simply has no other. It has neither Ethernet nor a TCP/IP stack, and writing them would mean building a little Linux on Oberon. The radio, on the other hand, is described in the very same book, and I was curious what kind of system would come out if I honestly accepted the machine&rsquo;s limits instead of dragging a modern network onto it.</p>
<p>A virtual machine has no real radio, of course. So our QEMU model emulates the whole nRF24L01+ chip, registers and queues included, and wraps every frame it sends in a UDP datagram. The datagrams flow to a relay, a small program that hands every frame to all the other machines. That is the air. The relay can lose a given share of frames, and if you stop it, the air is gone altogether. In Cozystack, our open platform built on Kubernetes, the relay became a catalog application, <code>OberonAir</code>, and in the browser the page itself plays the air.</p>
<h3 id="three-messages-each-in-one-frame">Three messages, each in one frame</h3>
<p>The protocol, which I called KubeNet, has just three messages. Each goes to everyone at once, since a radio cannot pick a receiver anyway, and each fits into one frame.</p>
<p>A node sends its heartbeat once a second, giving its name and the ids of the pods that are actually running on it right now. The control plane sends an assignment to every live node once a second, listing the pods bound to that node. The third message, a spec, also comes from the control plane and tells which image the pod with a given id has. Specs go out one per tick, pods not yet started first.</p>
<p>Why one frame and not a message of any length? <code>SCC</code> can send packets of up to half a kilobyte by cutting them into frames, but the air has no collision avoidance. If two stations start sending long packets at the same time, their frames get interleaved at every receiver, and a receiver can no longer tell whose frame is whose. One could invent air arbitration, queues and retransmissions, but that is exactly the road along which networks arrived at the complexity I wanted to get away from. So every message fits into one frame, which leaves 24 bytes of data after the <code>SCC</code> header.</p>
<p>In a heartbeat and an assignment these bytes are laid out as follows: the cluster tag, six bytes of node name, the pod count, up to ten pod ids at a byte each, two bytes of counter and four bytes of signature. Exactly twenty-four. Hence all of Kube&rsquo;s odd limits: a node name is at most six characters, a pod id takes one byte, and a node runs at most ten pods, because no more fit into a heartbeat and the scheduler will not bind more.</p>
<p>The main property of the protocol is that it is level-triggered, like Kubernetes controllers, only not inside the control plane but on the network itself. A node runs exactly what its last assignment says, not a sequence of &ldquo;start&rdquo; and &ldquo;stop&rdquo; commands. If an assignment is lost, another with the same content comes a second later, so there is nothing to repair. If a heartbeat is lost, the control plane waits for the next one. There are no acknowledgements, no retransmissions and no sequence numbers for reliability: all of the reliability comes from each message carrying the full state rather than a change. Even on an air that loses 30 percent of its packets, not a single pod moved in a minute, and I checked that.</p>
<figure><img src="en-03b-air-log.png" alt="The air in the lab" width="420" height="341" loading="lazy" decoding="async">
  <figcaption><p>The air in the lab. Hearts are node heartbeats, arrows are assignments from the control plane, the pencil marks pod specs. A rollout is under way, and the control plane is telling the nodes about new pods with the image Ticker2</p></figcaption>
</figure>

<h3 id="a-pod-is-a-module">A pod is a module</h3>
<p>In real Kubernetes a pod&rsquo;s image is a file-system archive with a program inside: the kubelet pulls it from a registry and runs it in an isolated container. Oberon has nothing like that, but it has a mechanism that works similarly and much faster. A pod&rsquo;s image in Kube is simply the name of an Oberon module on the node.</p>
<p>On learning of a new pod, the kubelet calls <code>Modules.Load</code> with the image name. If the module is not loaded yet, the system finds its compiled file on disk, loads it into memory, links it with every module it imports, checks the keys of their interfaces and runs the module&rsquo;s body. A key is a kind of checksum of the interface that the compiler writes into every module. If an interface has changed, the system refuses to load modules compiled against the old version. Then the kubelet finds the module&rsquo;s <code>Start</code> command and calls it, and when the pod is to go away, it calls the same module&rsquo;s <code>Stop</code> command.</p>
<p>A command in Oberon is any exported procedure without parameters. You run it by middle-clicking the text <code>Module.Procedure</code> in any window, and if it needs parameters, it reads them itself from the text after its name. So a pod&rsquo;s id cannot be passed directly, and that is what a tiny module, <code>Pods</code>, with two variables, the id and the image, is for. The kubelet fills them in before the call, and the workload module reads them. Not the most elegant solution, but very much in Oberon&rsquo;s spirit: a module&rsquo;s global variable is a legitimate way to pass context here, because only one command runs at any moment and races have nowhere to come from.</p>
<p>The teaching workload comes in two versions, <code>Ticker</code> and <code>Ticker2</code>. For each pod it starts it keeps a seconds counter, and the command <code>Ticker.Show</code> prints which pods run on this machine and how many seconds each has lived. They make rollouts easy to watch: the node&rsquo;s log shows the first version&rsquo;s pods stopping and the second version&rsquo;s starting.</p>
<p>If there is no module of that name on the node, <code>Start</code> cannot be called, and the kubelet simply leaves the pod out of its heartbeat. The control plane sees that the pod is assigned but not running, and it stays Pending. In Kubernetes, that is what a pod whose image could not be pulled looks like.</p>
<h3 id="the-store-on-disk">The store on disk</h3>
<p>In Kubernetes the whole state of the cluster lives in etcd, and a restarted control plane reads it from there. In Kube the role of etcd is played by an array of 256 records in the control plane machine&rsquo;s memory, and so that it survives a restart, a background task writes it to disk whenever it changes.</p>
<p>It is written to two files in turn, <code>Kube.Store0</code> and <code>Kube.Store1</code>, each holding a generation number and a checksum. If power fails in the middle of a write, only one file is damaged, and the other, of the previous generation, survives. At start <code>Kube.Start</code> reads both and takes the newer of the whole ones. The files are rewritten in place rather than created anew, and Oberon demands this too: its file system frees the space of replaced files only at the next boot, and if a new file were created for every change, a busy cluster would simply run the disk out.</p>
<p>After a restart the control plane gives the nodes one heartbeat timeout to report, and only then begins to count them NotReady. Kubernetes does the same after its node controller restarts. The start of that timeout, by the way, had to move from the moment the store has been read to the moment the control plane starts listening to the air. In the cloud, a person managed to type something over VNC between those two commands, the timeout ran out, and every pod moved, although none had stopped.</p>
<h3 id="rollouts-and-two-bugs-kubernetes-knows-well">Rollouts, and two bugs Kubernetes knows well</h3>
<p>The command <code>Kube.Apply web 6 Ticker2</code>, given to a running <code>web 6 Ticker</code>, changes the image. The deployment controller creates a new ReplicaSet and starts moving pods one at a time. First a pod with the new image is added. As soon as its kubelet reports it running, there is one pod more than wanted, and the old ReplicaSet removes one of its own. This repeats until the old ReplicaSet is empty, and then it is deleted.</p>
<p>It looks simple, but two bugs surfaced along the way, and both turned out to be old acquaintances of Kubernetes.</p>
<p>The first concerns stopping a pod. When the control plane deletes a pod from the store, the pod keeps running on its node until the kubelet hears the next assignment, which can take up to a second. By then the controller already sees one pod fewer and adds a new one. As a result, for a moment eight pods ran where seven were allowed. Real Kubernetes behaves in exactly the same way, and only recently has the Deployment gained the field <code>podReplacementPolicy</code>, which makes it wait until old pods have stopped completely, and even that is still alpha, behind a feature gate. In Kube I made this the only behaviour: pods the nodes still report but the store no longer holds count as stopping, and while there are any, no new pod is added.</p>
<p>The second bug concerns pod ids. A node knows a pod only by its one-byte id. At first a new pod got the lowest free id, and that often turned out to be the id of the very old pod it had just replaced. The kubelet saw a familiar id in the assignment and decided nothing had changed. The store claimed Ticker2 was running, while Ticker kept spinning on the node. This is exactly why Kubernetes never reuses a pod&rsquo;s UID. Kube&rsquo;s ids now go round, and the check after every rollout requires that no id of an old pod is running any more.</p>
<figure><img src="en-04-rolled.png" alt="The rollout is over" width="1440" height="900" loading="lazy" decoding="async">
  <figcaption><p>The rollout is over: all four pods run the new image, and during the rollout there were never fewer than four pods nor more than five. The page counted that from heartbeats, not from what the control plane says</p></figcaption>
</figure>

<h3 id="the-failures-clusters-are-built-for">The failures clusters are built for</h3>
<p>A cluster exists to survive failures, and I decided to test that the way I would test a real one. A separate test starts a control plane and two nodes and judges them solely by the air. For that a passive listener joins the relay: it sends nothing, only records every heartbeat and every assignment and checks their signatures. The test asks Kube itself almost nothing: only the outcome of a rollout is compared with what Kube wrote to disk, and everything else is judged by what actually happened on the air.</p>
<p>Switch a node off, and after five seconds of silence the control plane marks it NotReady, and two seconds later moves its pods to another node, so that the whole thing takes about eight seconds. Switch the control plane off, and the nodes keep running their last assignment, because there is nobody to tell them otherwise. When it boots again, the store comes back from disk and not one assignment changes. If the nodes lack the module, the pods stay Pending. And if 30 percent of packets are lost for a whole minute, not one pod moves.</p>
<p>At first the timeout was three seconds. On an air losing 30 percent, three heartbeats in a row went missing several times a minute, and the control plane, taking the node for dead, moved pods that had never gone anywhere. With five seconds, false alarms became dozens of times rarer, and the price was slower recovery from a real failure. Kubernetes makes the same bargain at another scale: the kubelet reports every ten seconds, and the control plane waits forty to fifty.</p>
<p>The most instructive case was losing the air. When the relay falls silent for twenty seconds, all the nodes fall silent at once. A control plane that simply trusts its timeouts decides the nodes have died and tries to move every pod, although there is nowhere to move them. And when the air comes back, the nodes come to life one by one: the first one back receives everyone else&rsquo;s pods, the second takes some of them back, and so on. In the first version twenty seconds of outage changed 74 assignments. The nodes kept working the whole time and noticed nothing, while the cluster staged a disaster for itself.</p>
<p>Kubernetes knows this trap: when all nodes of a zone go NotReady at once, the node controller puts the zone into the FullDisruption state and stops evicting pods, reasoning that all nodes dying at the same moment is far less likely than a lost link. Kube does the same when all nodes, or at least 55 percent of three or more, go NotReady. But it did not work at once: it took two details that only came to light in a failed check.</p>
<p>First, nodes do not go NotReady at the same moment; their heartbeats are up to a second apart. So the first node to go silent lost its pods before the others went silent and it became clear that this was an outage. Now a node&rsquo;s pods move only after it has been NotReady for two seconds, and by then all the others have gone silent too. Kubernetes waits five minutes for this by default. Second, nodes also come back one by one, and the very first to return took the cluster out of that state before the others had reported, so their pods moved at once. Now the silent nodes get one more timeout to report after leaving it, as Kubernetes does too, resetting the nodes&rsquo; timers when a zone leaves FullDisruption. With these two corrections, twenty seconds of outage change not a single assignment.</p>
<h3 id="strangers-on-the-air">Strangers on the air</h3>
<p>Once, in the sandbox, two clusters ended up on one air, and both had a node of the same name. A node of one cluster kept starting and stopping its pods, because it obeyed the assignments of both control planes in turn. So a cluster tag was added to the start of every message, a byte computed from the cluster&rsquo;s name, and a node began to ignore messages with a foreign tag.</p>
<p>But a tag proves nothing; anyone can send it, so a signature came next. Every message carries an authentication code, a MAC, computed from the message itself and the cluster&rsquo;s secret key, and without the key the right code cannot be guessed. It is computed with HalfSipHash-2-4, a reduced variant of SipHash that works on 32-bit words and gives a 32-bit result. HMAC-SHA256 will not do here: Wirth&rsquo;s processor is 32-bit and runs at 25 MHz, and of the message&rsquo;s 24 bytes the signature can take four at most. HalfSipHash was made precisely for such small devices, and in Oberon it takes about thirty lines. Thirty-two bits is on the small side for serious cryptography, but for a teaching cluster it is an honest compromise.</p>
<p>A signature does not prevent a recorded genuine message from being replayed, so every message also carries a 16-bit counter, and the receiver drops anything not newer than the last one accepted from the same sender. The failure test checks this too. Right after a genuine assignment it sends the node, on behalf of a stranger, &ldquo;run nothing&rdquo; in three ways: tagged as another cluster&rsquo;s, signed with the wrong key, and as a genuine assignment recorded earlier. The node ignores all three. The same message, honestly signed with the key and with a fresh counter, it obeys, and that is the control case without which the check would prove nothing.</p>
<h3 id="how-much-it-takes">How much it takes</h3>
<p>A separate load test starts a control plane and two to eight nodes, three pods per node, and measures from the air how fast the cluster converges and how fast it recovers from losing a node. At any size it converges in about three seconds, and losing a node costs about eight: five seconds of timeout, two seconds of pause before eviction, and up to a second until the next assignment. There was never a false NotReady.</p>
<p>The cluster does have a ceiling, and it lies not in the radio itself but in Wirth&rsquo;s driver. Before every send, <code>SCC</code> waits 50 milliseconds to let someone else&rsquo;s acknowledgement finish, and the machine does nothing else meanwhile. So one station can send at most twenty messages a second. The control plane needs one assignment per node every second, so with eight nodes it spends almost half its time waiting, while received frames pile up in the three-slot buffer and the extra ones are lost. During a rollout specs join the assignments, one per tick, and the waiting takes up almost all the time. By my reckoning the protocol can carry about twenty nodes when idle and about ten during a rollout. That is a calculation from the code, not a measurement. Besides, every Oberon machine occupies a whole host core, since its loop never idles, and eight nodes on one machine load that machine rather than the air.</p>
<h3 id="two-bugs-in-our-qemu">Two bugs in our QEMU</h3>
<p>To make all of this work, two bugs had to be fixed not in Kube but in our machine model for QEMU, and without the cluster I would have found neither. The <code>MOD</code> operation after a multiplication sometimes returned the wrong half of the product, so the store&rsquo;s checksum always came out as zero. The background task saw no changes in it, and the store was never written once. And a machine that had once lost its relay stopped hearing the air for good, because QEMU&rsquo;s UDP channel quietly dropped its reader after a failed read. Both findings are described in the repository, and they are a good example of why I so like running real programs on an emulator: booting the system triggered neither bug.</p>
<h3 id="commands-at-start-and-oberonkube">Commands at start, and OberonKube</h3>
<p>The last step was about convenience, but without it everything else would have lost its point. Nobody will start a cluster a second time if commands have to be typed over VNC on every machine. I needed each machine to know who it is at start, the way a cloud VM learns its role from cloud-init.</p>
<p>QEMU now takes a string of commands for the machine to run at start and feeds it to the serial port, as if it had been typed at a console. A small module, <code>Boot</code>, reads the string and runs the commands one after another, each with its parameters. It is called at the very end of the body of <code>System</code>, the module loaded when the system starts. <code>System</code> itself is rebuilt from the image&rsquo;s own sources with that single line added, so its interface, and the key every other module checks, stayed the same.</p>
<p>Here the cluster taught one more lesson. The commands at start run at every boot, not just the first. If <code>Kube.Apply web 4 Ticker</code> is among them, then after a restart of the control plane the deployment returns to what the commands say, and a rollout done by hand since then is quietly undone. Whether that is good or bad depends on what you take as the source of truth. In the cloud it is the order form, and returning to it is exactly the right behaviour there. But in the browser lab, where a person rolls out a new version by hand, it was a bug, and it was the lab itself, by the way, not the QEMU tests, that found it. For such cases there is now the command <code>Kube.Ensure</code>, which creates a deployment only if there is none yet.</p>
<p>In Cozystack all of this came together in one catalog application, <code>OberonKube</code>. A user says how many nodes they need, the cluster key and a list of deployments, and the catalog creates the air, the control plane machine and the nodes, and gives each machine the commands for its role. The cluster forms by itself. The form here is the source of truth: a changed list of deployments takes effect at the next restart of the control plane. The machines of one air also try to land on different hosts, so that losing a host costs as few nodes as possible.</p>
<h2 id="what-oberon-imposed">What Oberon imposed</h2>
<p>When you write Kubernetes in Go for Linux, almost any decision can go any way. Need a queue, and channels are at hand; need a database, and there is etcd; need a network, and there is gRPC over TCP; need isolation, and there are kernel namespaces. There is always a choice, and so decisions are often made out of habit. On Oberon there is almost nothing to choose from, and that is the most interesting part: every limit of the language and the system forced a quite specific decision, and those decisions show clearly which properties of Kubernetes follow from its very idea and which from what it was built on.</p>
<p>I will split the limits into those that came from the language and those that came from the operating system and the machine, although with Wirth the boundary between them is a convention: it was all done by one hand.</p>
<h3 id="what-the-language-imposed">What the language imposed</h3>
<p><strong>Sizes are known in advance.</strong> In Oberon-07 an array&rsquo;s size is nearly always known at compile time, and dynamic memory is allocated only for records reached through pointers. Writing a program in such a language where everything grows as needed is awkward, and I did not try. Kube&rsquo;s store is an array of 256 objects, a node has at most ten pods, a node name is at most six characters, an image name at most fifteen, a pod id takes one byte. Each of these numbers has a reason, and all of them sit in two blocks of constants at the top of the modules <code>Kube</code> and <code>KubeNet</code>. Thanks to this, Kube allocates no memory in its working loop, cannot exhaust it under a flood of objects, and behaves predictably at its limits: an extra pod simply will not be created. Real Kubernetes lives with limits too, such as 110 pods per node by default or a megabyte and a half per object in etcd, but there they are scattered across the documentation, while here they cannot be hidden.</p>
<p><strong>Integers are 32-bit, and only bytes are unsigned.</strong> HalfSipHash works on exactly 32-bit words and produces a 32-bit signature that fits in the frame. Hence also the counter arithmetic modulo 65,536, with a careful &ldquo;is this number newer&rdquo; check that stays correct after wraparound. Hence too the store&rsquo;s checksum, computed so as not to depend on sign. A small thing, but it is exactly where the bug with <code>MOD</code> in our QEMU turned up.</p>
<p><strong>Commands without parameters.</strong> An exported procedure without parameters counts as a command, and one with parameters does not. So the kubelet cannot call <code>Start(pod)</code>: it puts the pod&rsquo;s id into the <code>Pods</code> module and calls <code>Start</code> with no arguments. In any other language this would be considered bad form, but here there is no other way, and it is safe, because only one command runs at a time.</p>
<p><strong>Modules with keys.</strong> Every compiled module carries the key of its interface, and at load time the system checks it against what the importing modules were compiled with. So a pod&rsquo;s image in Kube is not just a name but a name with a built-in compatibility check. If a workload module was compiled against an old version of <code>Pods</code>, the system refuses to load it, and the pod stays Pending instead of crashing in the middle of its work. In the container world the closest thing is a pinned image digest, but that only guarantees that you got exactly those bytes, not at all that they are compatible with their environment.</p>
<p><strong>The language is small.</strong> It sounds like a drawback, but in practice it disciplines more than anything else. Oberon has no generic collections, exceptions, interfaces, goroutines or reflection, so all of Kube is written as plain loops over arrays and procedures that return a boolean instead of throwing. It reads almost like pseudocode, and that is just how I wanted it to read.</p>
<h3 id="what-the-system-and-the-machine-imposed">What the system and the machine imposed</h3>
<p><strong>One loop and cooperative tasks.</strong> This is the main one. Oberon has no threads, so Kube&rsquo;s controllers run as tasks in the central loop, and the kubelet on a node is a task too. Each task does a short piece of work and returns, and three things follow at once. First, Kube has not a single lock, mutex or channel, because races have nowhere to come from: while one controller runs, the rest stand still. Second, behaviour is deterministic: given the same input, the controllers do the same things in the same order, and such a system is far easier to debug. Third, and this is a minus, any task that stops to think for long freezes the whole machine, kubelet and radio included. A pod whose <code>Start</code> command goes into an endless loop hangs the entire node.</p>
<p><strong>No memory protection.</strong> All modules live in one address space, and the only thing that keeps one module from corrupting another&rsquo;s memory is a strictly typed language with array bounds checks. For Kube this means that pods are in no way isolated, either from each other or from the kubelet. This is the most serious difference from real Kubernetes, and I will come back to it when we get to what Oberon lacked.</p>
<p><strong>The interface is text.</strong> In Oberon any line of the form <code>Module.Command</code> in any window is a command. So Kube has no API server, no YAML and no kubectl. Its API consists of commands such as <code>Kube.Apply web 6 Ticker2</code>, <code>Kube.Get</code> and <code>Kube.DeletePod</code>, which a person writes in any window and runs with a middle click, reading the result in the system log. The same principle made the commands at start possible: a machine&rsquo;s role is set by an ordinary line of commands that QEMU feeds to the serial port, and the module <code>Boot</code> runs it exactly as a person would. No special configuration format was needed; the configuration of a machine became the text of its commands.</p>
<p><strong>An old-school file system.</strong> Oberon frees the space of replaced files only at the next boot. So Kube&rsquo;s store is rewritten in place, in two files in turn, rather than created anew at every change, and that also protected it from a power failure in the middle of a write. A solution databases arrived at decades ago appeared here by itself, out of a file-system limitation.</p>
<p><strong>Radio instead of a network.</strong> I have already covered this in detail: a 32-byte frame, broadcast to everyone, no collision avoidance. Hence a protocol of one-frame messages, full state instead of acknowledgements, a cluster tag and a signature in every message. And hence the most unexpected property of the whole system, which the next section is about: everything the cluster knows can be heard on the air.</p>
<p><strong>Twenty-five megahertz.</strong> The machine&rsquo;s speed set the periods: controllers fire every 50 milliseconds, the kubelet every 100, a heartbeat goes out once a second. The control plane needs very little actual computation, and nearly all its time goes into waiting. But alas, the waiting is not free either: before every send the radio driver spins in a loop for fifty milliseconds, and it is this waiting, not computation, that limits the size of the cluster.</p>
<h3 id="how-kube-differs-from-real-kubernetes">How Kube differs from real Kubernetes</h3>
<p>To create no illusions, here is a brief comparison. Kube embodies the idea of Kubernetes, not Kubernetes itself, and the difference between them is enormous.</p>
<table>
  <thead>
      <tr>
          <th></th>
          <th>Kubernetes</th>
          <th>Kube</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>size</td>
          <td>millions of lines of Go</td>
          <td>about 1,400 lines of Oberon, of which the control plane is about 800</td>
      </tr>
      <tr>
          <td>store</td>
          <td>etcd, distributed and kept consistent by Raft</td>
          <td>an array of 256 objects in one machine&rsquo;s memory and two files on its disk</td>
      </tr>
      <tr>
          <td>API</td>
          <td>API server, REST, YAML, kubectl, RBAC</td>
          <td>Oberon commands a person writes in a window</td>
      </tr>
      <tr>
          <td>kinds of objects</td>
          <td>dozens, plus your own through CRDs</td>
          <td>four: Deployment, ReplicaSet, Pod and Node</td>
      </tr>
      <tr>
          <td>network between nodes</td>
          <td>TCP/IP, usually with a separate pod network</td>
          <td>radio broadcast, 32-byte frames</td>
      </tr>
      <tr>
          <td>pod image</td>
          <td>a container image from a registry</td>
          <td>an Oberon module on the node&rsquo;s disk</td>
      </tr>
      <tr>
          <td>pod isolation</td>
          <td>Linux kernel namespaces and cgroups</td>
          <td>none, all pods share one address space with the kubelet</td>
      </tr>
      <tr>
          <td>resources</td>
          <td>CPU and memory requests and limits</td>
          <td>only the number of pods per node, ten at most</td>
      </tr>
      <tr>
          <td>scheduler</td>
          <td>filters, scores, affinity, priorities</td>
          <td>the least loaded node</td>
      </tr>
      <tr>
          <td>networking for applications</td>
          <td>Service, DNS, Ingress</td>
          <td>none</td>
      </tr>
      <tr>
          <td>storage for applications</td>
          <td>PersistentVolume</td>
          <td>none</td>
      </tr>
      <tr>
          <td>control plane availability</td>
          <td>several replicas of the API server and etcd</td>
          <td>one machine; after a restart the state is read from disk</td>
      </tr>
      <tr>
          <td>readiness</td>
          <td>readiness and liveness probes</td>
          <td>a pod is ready as soon as its <code>Start</code> command returns</td>
      </tr>
      <tr>
          <td>protocol security</td>
          <td>TLS with mutual certificate checks</td>
          <td>a 32-bit signature with a shared key and a counter against replays</td>
      </tr>
      <tr>
          <td>scale</td>
          <td>thousands of nodes</td>
          <td>tested up to eight nodes, all on one air</td>
      </tr>
  </tbody>
</table>
<p>This table shows that Kube is good for nothing Kubernetes is good for. But it shows something else too: all the logic Kubernetes exists for, that is, the wanted state, independent reconcile loops, rollouts without downtime, recovery from a lost node and caution when the link is lost, fit into a small system with almost nothing listed in the middle column. Everything else in Kubernetes is there so that this logic works on thousands of machines, for thousands of users and for applications nobody trusts. Very important things, but not the foundation.</p>
<h2 id="where-kube-turned-out-better-and-where-oberon-fell-down">Where Kube turned out better, and where Oberon fell down</h2>
<p>Comparing a toy cluster on the radio with the Kubernetes that half the internet runs on seems unfair, and I am not going to claim Kube is better. What interests me is something else: which qualities emerged in Kube by themselves, with no effort at all, simply because it grew up on Oberon. I mean the properties that ordinary Kubernetes either lacks or pays far too much for. I counted six of them.</p>
<h3 id="the-whole-cluster-can-be-read-in-an-evening">The whole cluster can be read in an evening</h3>
<p>The control plane is about eight hundred lines, the network module with the kubelet and signatures about four hundred, and with the teaching workload and the commands-at-start module it comes to about fourteen hundred. Below lies only the Oberon system, which can also be read in full, and Wirth&rsquo;s processor in Verilog, for which a couple of evenings will do. So the whole path from &ldquo;I want six replicas of Ticker2&rdquo; to the register of the radio chip through which an assignment flies off to a node can be followed by eye, without once hitting a library that nobody has ever read.</p>
<p>With Kubernetes that stopped being possible long ago, and not because it is badly written: it solves a huge number of problems Kube simply does not have. But in the end even an experienced engineer who has run it for years usually knows how it behaves but not why, and finds out only during an outage. In Kube the answer to any &ldquo;why&rdquo; is in one procedure.</p>
<h3 id="the-air-is-the-observability">The air is the observability</h3>
<p>I did not plan this property; it arose by itself, and I like it best of all. Since every message goes to everyone and carries the full state rather than a change, at any moment the air carries everything the cluster knows about itself: which nodes are alive, what runs on each, what is bound to each, and which image each pod has. A passive listener that sends nothing reconstructs the whole picture of the cluster without asking it a single question.</p>
<p>All my checks work this way. Neither the failure test, nor the load test, nor the browser labs ever ask the control plane what is going on. They listen to the air and compare what actually ran on the nodes with what was assigned to them. A bug in Kube itself cannot fool such a check, since it looks not at Kube&rsquo;s reports but at the behaviour of the nodes. In ordinary Kubernetes the same takes metrics, logs, audit, exporters and a separate system to collect them all, and still you see only what the components chose to tell about themselves.</p>
<h3 id="a-pod-starts-in-the-blink-of-an-eye">A pod starts in the blink of an eye</h3>
<p>In Kubernetes starting a pod means pulling an image, unpacking layers, creating namespaces and starting a process, and that takes seconds, or even minutes with a cold cache and a big image. In Kube starting a pod means loading an Oberon module, if it is not loaded yet, and calling its command, and if the module is already in memory, starting a pod comes down to a procedure call. The whole cluster converges in three seconds, and almost all of that time goes into waiting for the next heartbeat, not into work.</p>
<p>This speed was paid for with isolation, of course, so the comparison is unfair. But it shows where the time really goes. When people talk today about cold starts, about WebAssembly on the server or about V8 isolates, this is exactly the question: can a unit of deployment be made so small and so module-like that starting it costs as much as a function call, without giving up isolation?</p>
<h3 id="compatibility-is-checked-at-load-time">Compatibility is checked at load time</h3>
<p>I wrote about this in the section on principles, but it bears repeating, because it is a strong idea. A pod&rsquo;s image in Kube is a module with an interface key, and the system refuses to load it if it was compiled against an incompatible version of what it imports. A pod with an incompatible image will not crash after an hour of work on an unexpected call: it simply will not start, it stays Pending, and that is visible at once. In the container world nobody checks an image&rsquo;s compatibility with its environment except your tests, and the compatibility of images that call each other is checked by an API schema at best, and only if you have one.</p>
<h3 id="no-locks-and-therefore-no-race-conditions">No locks, and therefore no race conditions</h3>
<p>Kube&rsquo;s code has no mutexes, no channels and no atomic operations. Controllers, the kubelet, writing the store to disk and receiving radio packets all run as tasks of one loop and execute strictly in turn. So data races, deadlocks and bugs that show up once in a thousand runs have nowhere to come from in Kube. Behaviour is deterministic: run the same scenario twice, and the controllers do the same things in the same order.</p>
<p>There is a flip side, covered a little further on, but the very thought that a control plane need not be multithreaded seems underrated to me. The control plane of eight nodes needs very little computation. The control plane of a thousand nodes needs more, of course, but there too most of the time goes not into computing but into waiting for the network and the disk, and single-threaded event loops proved long ago that such waiting can be served without threads.</p>
<h3 id="security-from-the-first-message">Security from the first message</h3>
<p>Early Kubernetes left a lot open by default; the kubelet, for one, accepted anonymous requests for a long time, and closing such holes took years. In Kube the signature and replay protection appeared before the cluster learned to do anything useful, simply because the radio left no choice: on the air anyone can say anything, and that is obvious from day one. The signature is 32-bit, which is too little for the real world, but architecturally the protocol is right: every message is authentic and belongs to its cluster, and the check costs about thirty lines.</p>
<h3 id="the-whole-cluster-in-a-browser-tab">The whole cluster in a browser tab</h3>
<p>And the last one, not about architecture but about what follows from it. An Oberon machine is so small that three of them, together with the air, fit into a browser tab. So everything I describe can not only be read but repeated without installing anything: switch off a node, jam the air, slip in a forged assignment. Kubernetes has good learning sandboxes, but they are always someone&rsquo;s cluster somewhere in a cloud, while here the cluster lives on your computer and can be broken any way you like without bothering anyone.</p>
<h3 id="what-oberon-lacked">What Oberon lacked</h3>
<p>Now about where Oberon proved too tight, and why such systems most likely lost. For the research this part matters no less.</p>
<p><strong>Isolation.</strong> This is the main one. In Oberon all modules live in one address space and trust each other. A pod that corrupts memory through <code>SYSTEM.PUT</code> corrupts the kubelet, the radio and everything else along with it. A pod that goes into an endless loop freezes the whole node, because a cooperative task that does not return stops the entire system. As long as the machine runs code by one author who trusts that code, all is well, and Wirth designed his system just that way: for one person at one computer. But a cluster exists precisely to run other people&rsquo;s code, and without isolation that is impossible. Kubernetes on Linux gets isolation from the kernel, while to get it Oberon would need memory protection in the processor, preemptive multitasking and resource accounting, that is, it would have to become a different system entirely.</p>
<p><strong>Preemption.</strong> Besides isolation, the control plane cannot interrupt a pod that computes for too long, nor share processor time among pods. No quotas, no priorities, no limits, only good manners. A teaching workload that adds one to a counter once a second does not care, but any real workload will stumble on this at once.</p>
<p><strong>A network.</strong> Radio with 32-byte frames is excellent for control, but there is no room in it for application data. Kube&rsquo;s pods have no addresses, no services and no way to talk to each other. Files can be sent over Wirth&rsquo;s radio, but that would be more like a USB stick than a network.</p>
<p><strong>Flexible sizes.</strong> The limits I praised for predictability turn into a ceiling: ten pods per node, 256 objects in the store, six characters per name and twenty messages a second per station. Each of these ceilings can be raised, but not without end, because they all follow from the 24-byte frame, static arrays and a simple driver. Kubernetes paid for having no such ceilings with enormous complexity, and looking at Kube you understand that this price was deliberate.</p>
<p><strong>A highly available control plane.</strong> Kube has one control plane machine. If it dies for good along with its disk, the cluster is left without a master. The nodes keep running their last assignment, which is a good property, but the control plane cannot be replaced by another machine, because there is no consistent store across several machines. Raft can be written in Oberon, but over a radio without delivery guarantees and with 32-byte frames that would be a big piece of research of its own.</p>
<p><strong>Multiple users.</strong> Oberon has one user, Kube one key per cluster. No namespaces, no roles, no separation of rights, and whoever knows the key can do anything. Kubernetes spends most of its complexity precisely on letting many people and teams safely share one cluster, while Kube has no such layer at all.</p>
<p>Put all these points together, and it becomes clear that Oberon lost not because it was badly designed, but because it was designed for another world, where a computer belongs to one person, all the code on it is written by that person or by people they trust, and a network is a way to pass a file to a neighbour. As soon as computers became shared and code became foreign, isolation, preemption and rights were needed, and simple systems gave way to complex ones.</p>
<h2 id="six-labs-in-your-browser">Six labs in your browser</h2>
<p>At first I doubted the cluster could be run in a browser at all. The Oberon machine in the labs of the previous article already ran in a tab: it is Wirth&rsquo;s actual circuit, translated to C++ and compiled to WebAssembly. But it had no radio, and a cluster needs three machines that hear each other. So the browser machine got the same nRF24L01+ transceiver model as the QEMU one, and learned to hand the frames it sends to the outside and to take in other machines&rsquo; frames. Each machine runs in a background thread of its own, and the page takes their frames and hands them to the others, that is, the page itself is the air. It can also lose a given share of frames, cut a single machine off the air and read Kube&rsquo;s messages, checking their signatures, which is why a cluster table built from the air appears on its right.</p>
<p>A serial port was needed too, so that machines learn their roles at power-on, as in the cloud. The control plane receives the commands <code>Kube.Start</code>, <code>KubeNet.Serve kube 00c0ffee00c0ffee</code> and <code>Kube.Ensure web 4 Ticker</code>, and the nodes <code>KubeNet.Join</code> with their names, so nothing has to be typed. <code>Kube.Ensure</code> is there for a reason, but more on that in the sixth lab.</p>
<p>Speed turned out to be a task of its own. In WebAssembly Wirth&rsquo;s circuit executes under a million instructions a second per machine, that is, dozens of times slower than the 25 MHz board. Kube, meanwhile, counts time in machine seconds: a heartbeat every second, a timeout of five. Had I kept the machines&rsquo; clocks honest, the cluster would take several minutes to form, and the lab with a node switched off would take about five minutes. So the page runs the machines&rsquo; clocks ten times faster than their instructions, and the machines believe they run at 2.5 MHz. Their seconds pass at a pace the eye can follow, and the protocol does not care: it has no idea how many instructions fit into a second.</p>
<p>Background tabs held another unexpected trap: the browser slows timers in them sharply, and the cluster nearly froze. So the machines now spin their own loop and do not depend on the page&rsquo;s timers. The last trap was input. The page types a command into a machine with keys and mouse clicks, and at first it did so in one go. That is about eleven million instructions, which with the fast clock is over four seconds of machine time during which the machine does not hear the air. After every button press the control plane declared both nodes NotReady, and all that saved it was that eviction stops in FullDisruption. Now input goes in small portions interleaved with the machine&rsquo;s work, and the air is delivered in between.</p>
<p>Each lab checks itself and shows a tick when what it describes has really happened on the air. You cannot press a button and get a tick: the check looks at heartbeats and assignments, not at what you did.</p>
<p><a href="https://tym83.github.io/paleocomputing/oberon/kube.html">Open the lab</a></p>
<figure><img src="en-kube-lab.gif" alt="All six labs in a row, sped up eight times" width="860" height="538" loading="lazy" decoding="async">
  <figcaption><p>All six labs in a row, sped up eight times. The cluster forms, new code rolls out, a node is switched off, the air is jammed, an intruder tries four ways, the control plane is restarted</p></figcaption>
</figure>

<h3 id="1-the-cluster-forms-by-itself">1. The cluster forms by itself</h3>
<p>Nothing to do but wait. The machines boot, and the control plane&rsquo;s screen shows the module <code>Boot</code> running the commands that came over the serial port. Kube installs three controllers, starts listening to the air and creates the deployment <code>web</code> with four pods, and a couple of seconds later the lines &ldquo;node node1 Ready&rdquo; and &ldquo;node node2 Ready&rdquo; appear in the log.</p>
<figure><img src="en-screen-plane.png" alt="The control plane&#39;s screen" width="1024" height="768" loading="lazy" decoding="async">
  <figcaption><p>The control plane&rsquo;s screen. Everything in the log below the system&rsquo;s version line was done by the commands at start; nobody pressed a single key</p></figcaption>
</figure>

<p>Meanwhile the nodes show the kubelet receiving its assignment and starting the pods of the Ticker module.</p>
<figure><img src="en-screen-node1.png" alt="The screen of node1" width="1024" height="768" loading="lazy" decoding="async">
  <figcaption><p>The screen of node1. The kubelet joined the air, received an assignment with two pods, loaded the Ticker module and called its Start for each. The last lines are the output of Ticker.Show: how many seconds each pod has lived</p></figcaption>
</figure>

<p>On the right of the page the cluster table fills in. For each node it shows when it was last heard, which pods it says it runs, and which are bound to it. If those two columns differ, the cluster has not converged yet, or something has gone wrong.</p>
<figure><img src="en-panel-cluster.png" alt="The cluster table built from the air" width="840" height="336" loading="lazy" decoding="async">
  <figcaption><p>The cluster table built from the air. The page asks Kube nothing; it just listens to heartbeats and assignments</p></figcaption>
</figure>

<h3 id="2-roll-out-new-code">2. Roll out new code</h3>
<p>Press the button <strong>web 4 Ticker2</strong>. The page types <code>Kube.Apply web 4 Ticker2 ~</code> on the control plane and runs it with a middle click, as a person would. Kube creates a new ReplicaSet and starts moving pods one at a time. The check passes the rollout when every pod runs the new image, and only if during all that time the number of pods never fell below four or rose above five. It counts from heartbeats, that is, from what the nodes actually ran.</p>
<figure><img src="en-panel-air-rollout.png" alt="The air during the rollout" width="840" height="682" loading="lazy" decoding="async">
  <figcaption><p>The air during the rollout. The control plane sends the specs of new pods with the image Ticker2, and the nodes report in their heartbeats that they started them</p></figcaption>
</figure>

<p>When the rollout is over, press <strong>Ticker2.Show</strong> on a node, and it shows the new version counting. The whole story is in the node&rsquo;s log: the Ticker v1 pods stopped, the Ticker v2 pods started, and the new pods have ids different from the old ones.</p>
<figure><img src="en-screen-node2-ticker2.png" alt="A node&#39;s screen after the rollout" width="1024" height="768" loading="lazy" decoding="async">
  <figcaption><p>A node&rsquo;s screen after the rollout. Every old pod got Stop, every new one Start, and Ticker2.Show lists the new pods</p></figcaption>
</figure>

<h3 id="3-a-node-dies">3. A node dies</h3>
<p>Press <strong>Power off</strong> above the screen of a node that runs pods. Its screen goes dark, its heartbeats stop, and after five machine seconds the control plane marks the node NotReady, and two seconds later moves its pods to the remaining one. Switch the node back on, and it boots, joins the air and becomes Ready, but nobody gives its pods back: like the real scheduler, Kube does not move running pods to a node that came later. New pods from scaling up, however, will go to it as the least loaded node.</p>
<figure><img src="en-06-moved.png" alt="The node is off" width="1440" height="900" loading="lazy" decoding="async">
  <figcaption><p>The node is off. It is red in the table and has not been heard for a while, and all pods already run on the second node</p></figcaption>
</figure>

<h3 id="4-the-air-dies">4. The air dies</h3>
<p>Press <strong>Switch the air off</strong> and wait half a minute. All nodes go NotReady at once, but Kube decides that the link broke, not the nodes, and moves nothing. Switch the air back on, and a couple of seconds later the same pods run on the same nodes. The lab passes only if not one assignment changed during the outage or after it.</p>
<figure><img src="en-07-air-off.png" alt="The air is off" width="1440" height="900" loading="lazy" decoding="async">
  <figcaption><p>The air is off. Both nodes are NotReady, but the assignments are unchanged: Kube understood this is a lost link</p></figcaption>
</figure>

<p>The most interesting part here is to open &ldquo;what is going on&rdquo; under the lab and read about the 74 reassignments the first version made in twenty seconds of outage. And if you like, switch the air off for a couple of seconds, shorter than the timeout, and see that nothing at all happens.</p>
<h3 id="5-an-intruder">5. An intruder</h3>
<p>The section &ldquo;An intruder on the air&rdquo; has four buttons. Each sends, on behalf of a stranger, the assignment &ldquo;run nothing&rdquo; to the node that runs pods, and does so right after a genuine assignment from the control plane, so that the forgery is the last thing the node heard. The first button tags it as another cluster&rsquo;s, the second signs it with the wrong key, the third replays a genuine assignment recorded earlier. The node ignores all three and keeps running its pods.</p>
<p>The fourth button signs the forgery with the real cluster key and a fresh counter. That is the control case, and the node obeys, stopping its pods. Without it the lab would prove nothing: what if the node simply listens to nobody but the control plane? A second later the control plane sends its next assignment, and the pods come back, because the protocol carries the full state, not commands.</p>
<figure><img src="en-panel-intruder.png" alt="The intruder panel after the fourth attempt" width="840" height="470" loading="lazy" decoding="async">
  <figcaption><p>The intruder panel after the fourth attempt: the node obeyed the assignment signed with the key. It ignored the three forgeries before it, and each time the page said so right here</p></figcaption>
</figure>

<p>The forgeries are visible in the air log too, by the way: the page checks signatures itself and marks messages with a foreign tag or a bad signature.</p>
<h3 id="6-the-control-plane-restarts">6. The control plane restarts</h3>
<p>Switch the control plane off and, a few seconds later, back on. It boots, runs its commands once more, reads its store from disk, and not one pod moves. The nodes kept running their last assignment all along, and the control plane, once back, gave them one timeout to report, and they all did.</p>
<p>It was this lab that found the last bug before publication. The commands at start first had <code>Kube.Apply web 4 Ticker</code>, and after a restart the deployment went back to Ticker, undoing the rollout from the second lab. The QEMU tests did not notice, because they restarted the control plane before any rollout. Here, where a person rolls out a new version by hand, the right thing is to keep what they did, so the commands now hold <code>Kube.Ensure</code>, which creates a deployment only if there is none yet. In the cloud, on the contrary, the order form is the source of truth, and <code>Kube.Apply</code> stayed there.</p>
<figure><img src="en-11-labs-done.png" alt="All six labs done" width="420" height="1075" loading="lazy" decoding="async">
  <figcaption><p>All six labs done</p></figcaption>
</figure>

<h3 id="what-else-to-try">What else to try</h3>
<p>The page does more than the labs. With the loss slider you can spoil the air, say, lose 30 percent of frames, and see that pods do not move anywhere. Push the losses much higher, and sooner or later five heartbeats in a row will go missing and Kube will take a live node for a dead one; that is the price of a timeout. <strong>Cut off the air</strong> isolates a machine without switching it off, and you can watch the cut-off node keep running its pods although the control plane has already handed them to others. That is the very case of a pod running in two places, and Kubernetes lives with it in just the same way. In the command field you can type any Oberon command, for example <code>Kube.Apply api 2 Ticker ~</code> to create a second deployment, or you can click right into a machine&rsquo;s screen and work in it by hand. The middle button there is a click with Alt.</p>
<h2 id="how-to-try-it-yourself">How to try it yourself</h2>
<p>There are four ways, from very simple, where nothing needs installing, to your own cloud. Everything below is open: the sources are in the repository <a href="https://github.com/tym83/paleocomputing">tym83/paleocomputing</a>, Kube&rsquo;s code in the <code>impl/kube</code> folder, and a detailed description with measurements in <code>impl/kube/README.md</code>.</p>
<h3 id="in-a-browser">In a browser</h3>
<p>Open the <a href="https://tym83.github.io/paleocomputing/oberon/kube.html">cluster lab</a> and wait half a minute. Nothing needs installing: everything runs in the tab, and the server only serves files. The page is happiest in a recent Chrome or Firefox on a computer with several cores, since each of the three machines takes a core of its own. Keep the tab in view: in the background the browser slows it down, and the cluster, though it does not stop, lives noticeably slower.</p>
<p>To run the same page locally, clone the repository, start any static server in the <code>impl/web</code> folder, for example <code>python3 -m http.server 8765</code>, and open <code>http://127.0.0.1:8765/kube.html</code>. And if you need no browser at all, the same three-machine cluster runs in Node.js with <code>node impl/web/kube-test.mjs</code>. It boots three machines on Wirth&rsquo;s actual circuit, waits for the cluster to form, and checks from the air that all pods run and all signatures are valid. The machines&rsquo; clocks are honest here, so it takes about a minute.</p>
<h3 id="in-qemu-on-your-computer">In QEMU on your computer</h3>
<p>You need Docker, git and make. First build our machine model for QEMU. The build runs in a container, so QEMU&rsquo;s dependencies stay out of your system, but it takes about ten minutes:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-sh" data-lang="sh"><span class="line"><span class="cl">git clone https://github.com/tym83/paleocomputing
</span></span><span class="line"><span class="cl"><span class="nb">cd</span> paleocomputing
</span></span><span class="line"><span class="cl">make -C qemu build
</span></span></code></pre></div><p>The system disk, with Kube, KubeNet, the teaching workload and the commands-at-start module already compiled onto it, is easiest to take from the published image. Each machine needs its own copy of the disk, grown to eight megabytes so the system has room to write its store:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-sh" data-lang="sh"><span class="line"><span class="cl">docker create --name payload ghcr.io/tym83/paleocomputing/oberon-run:v0.1.21
</span></span><span class="line"><span class="cl">docker cp payload:/opt/oberon/payload/prom.bin .
</span></span><span class="line"><span class="cl">docker cp payload:/opt/oberon/payload/oberon.dsk .
</span></span><span class="line"><span class="cl">docker rm payload
</span></span><span class="line"><span class="cl"><span class="k">for</span> m in plane node1 node2<span class="p">;</span> <span class="k">do</span> cp oberon.dsk <span class="nv">$m</span>.dsk<span class="p">;</span> truncate -s 8M <span class="nv">$m</span>.dsk<span class="p">;</span> <span class="k">done</span>
</span></span></code></pre></div><p>Next you need the air, that is, a Docker network with a relay in it:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-sh" data-lang="sh"><span class="line"><span class="cl">docker network create kube-air
</span></span><span class="line"><span class="cl">docker run -d --name relay --network kube-air -v <span class="s2">&#34;</span><span class="nv">$PWD</span><span class="s2">/qemu/radio:/r&#34;</span> <span class="se">\
</span></span></span><span class="line"><span class="cl">  qemu-build:risc5 <span class="s1">&#39;python3 -u /r/relay.py&#39;</span>
</span></span></code></pre></div><p>And finally three machines, each with its commands at start. The string after <code>commands=</code> is exactly what the machine runs after booting, the commands separated by semicolons:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-sh" data-lang="sh"><span class="line"><span class="cl">run<span class="o">()</span> <span class="o">{</span>
</span></span><span class="line"><span class="cl">  docker run -d --name <span class="nv">$1</span> --network kube-air -p 127.0.0.1:<span class="nv">$2</span>:5900 <span class="se">\
</span></span></span><span class="line"><span class="cl">    -v <span class="s2">&#34;</span><span class="nv">$PWD</span><span class="s2">/.qemu-work:/src:ro&#34;</span> -v <span class="s2">&#34;</span><span class="nv">$PWD</span><span class="s2">:/w&#34;</span> -w /w qemu-build:risc5 <span class="se">\
</span></span></span><span class="line"><span class="cl">    <span class="s2">&#34;/src/build/qemu-system-risc5 -machine &#39;oberon,radio=air,commands=</span><span class="nv">$3</span><span class="s2">&#39; \
</span></span></span><span class="line"><span class="cl"><span class="s2">     -bios prom.bin -drive if=none,id=sd0,file=</span><span class="nv">$1</span><span class="s2">.dsk,format=raw -vnc :0 \
</span></span></span><span class="line"><span class="cl"><span class="s2">     -chardev udp,id=air,host=relay,port=7524,localaddr=0.0.0.0,localport=7524&#34;</span>
</span></span><span class="line"><span class="cl"><span class="o">}</span>
</span></span><span class="line"><span class="cl"><span class="nv">KEY</span><span class="o">=</span>00c0ffee00c0ffee
</span></span><span class="line"><span class="cl">run plane <span class="m">5900</span> <span class="s2">&#34;Kube.Start;KubeNet.Serve kube </span><span class="nv">$KEY</span><span class="s2">;Kube.Apply web 4 Ticker&#34;</span>
</span></span><span class="line"><span class="cl">run node1 <span class="m">5901</span> <span class="s2">&#34;KubeNet.Join node1 kube </span><span class="nv">$KEY</span><span class="s2">&#34;</span>
</span></span><span class="line"><span class="cl">run node2 <span class="m">5902</span> <span class="s2">&#34;KubeNet.Join node2 kube </span><span class="nv">$KEY</span><span class="s2">&#34;</span>
</span></span></code></pre></div><p>This uses <code>Kube.Apply</code>, not <code>Kube.Ensure</code> as in the browser lab: release v0.1.21 does not have <code>Kube.Ensure</code> yet. The difference shows only if you roll out a new version by hand and restart the control plane: with <code>Kube.Apply</code> the deployment goes back to what the commands say.</p>
<p>The machines&rsquo; screens are on VNC ports 5900, 5901 and 5902. Booting in software emulation takes up to a minute, after which the control plane&rsquo;s log shows lines about the nodes being ready, and the nodes&rsquo; logs about the pods started. Oberon&rsquo;s mouse has three buttons, and the middle one runs the command it points at. So <code>Kube.Get</code> in any window of the control plane shows all objects, and <code>Ticker.Show</code> on a node shows its pods. To roll out a new version, write <code>Kube.Apply web 4 Ticker2 ~</code> on the control plane and middle-click that line.</p>
<figure><img src="qemu-plane.png" alt="The control plane in QEMU after a restart" width="1024" height="768" loading="lazy" decoding="async">
  <figcaption><p>The control plane in QEMU after a restart. The store came back from disk, and <code>Kube.Get</code> shows the same pods on the same node. The shot was taken by an automated check whose commands at start hold <code>Kube.Ensure</code>, which is why the log shows it left the existing deployment alone</p></figcaption>
</figure>

<figure><img src="qemu-node-b.png" alt="Node node-b in QEMU" width="1024" height="768" loading="lazy" decoding="async">
  <figcaption><p>Node node-b in QEMU. It joined first and got all four pods: like the real scheduler, Kube does not move running pods to a node that came later</p></figcaption>
</figure>

<p>You can listen to the air from the side, as the tests do. The listener joins the relay as one more machine, sends nothing, and with <code>--json</code> prints every heartbeat and every assignment together with whether its signature is valid:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-sh" data-lang="sh"><span class="line"><span class="cl">docker run --rm -it --network kube-air -v <span class="s2">&#34;</span><span class="nv">$PWD</span><span class="s2">/qemu/radio:/r&#34;</span> <span class="se">\
</span></span></span><span class="line"><span class="cl">  qemu-build:risc5 <span class="s1">&#39;python3 -u /r/listen.py relay --json --key 00c0ffee00c0ffee&#39;</span>
</span></span></code></pre></div><p>And then you can break things. <code>docker stop node2</code> switches a node off, <code>docker stop relay</code> jams the air, and <code>docker start</code> brings either back. If the relay gets a different address after a restart, the machines will find it themselves: after a failed send our model looks the address up by name again. The same folder holds <code>inject.py</code>, which can play the intruder.</p>
<p>The checks in the repository start the cluster in just this way, only automatically. After building QEMU and the tools with <code>make -C impl tools</code>, you can run them yourself: <code>python3 qemu/test/kube_dr_check.py</code> checks every failure discussed above (their table is in <code>impl/kube/README.md</code>), <code>python3 qemu/test/kube_boot_check.py</code> checks that the cluster forms from commands at start alone, and <code>python3 qemu/test/kube_load_check.py --nodes 4</code> measures load. These checks build the disk from source, so they already have <code>Kube.Ensure</code>. Keep in mind that each Oberon machine takes a whole core, since its loop never idles, and eight nodes on a laptop will measure the laptop rather than the cluster.</p>
<p>When you are done, remove the containers with <code>docker rm -f plane node1 node2 relay</code> and the network with <code>docker network rm kube-air</code>.</p>
<h3 id="in-your-own-kubevirt">In your own KubeVirt</h3>
<p>An Oberon machine also runs in plain KubeVirt, without Cozystack. It needs our virt-launcher image, which knows the RISC5 architecture, and KubeVirt&rsquo;s ability to attach hooks to virtual machines switched on. How to do that is described in detail in the <a href="https://github.com/tym83/paleocomputing/blob/main/kubevirt/GUIDE.md">guide</a>, which also has an example VirtualMachine. A cluster needs three such machines, a relay as an ordinary pod with a UDP service, and the commands at start in the machine&rsquo;s annotation, which the hook passes on to QEMU. That is exactly what the Cozystack catalog does for you, so without Cozystack the easiest way is to see what its charts in the <code>marketplace</code> folder create.</p>
<h3 id="in-cozystack">In Cozystack</h3>
<p>If you have Cozystack, plug in our catalog once with <code>cozypkg</code>:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-sh" data-lang="sh"><span class="line"><span class="cl">cozypkg tap oci://ghcr.io/tym83/paleocomputing/machines:v0.1.21
</span></span><span class="line"><span class="cl">cozypkg add paleocomputing.machines
</span></span></code></pre></div><p>The cluster administrator has to do two things for this: switch on the Sidecar feature gate in KubeVirt and install our virt-launcher image for your KubeVirt version. The details are on the <a href="https://tym83.github.io/paleocomputing/cozystack/">project page</a>.</p>
<p>After that, users see a Paleocomputing section in the dashboard, with OberonVM, OberonAir and OberonKube among other things. A Kube cluster is ordered with one form or one resource:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">apiVersion</span><span class="p">:</span><span class="w"> </span><span class="l">apps.cozystack.io/v1alpha1</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">kind</span><span class="p">:</span><span class="w"> </span><span class="l">OberonKube</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">metadata</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">farm</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">spec</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">nodes</span><span class="p">:</span><span class="w"> </span><span class="m">3</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="l">00c0ffee00c0ffee</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">deployments</span><span class="p">:</span><span class="w"> </span><span class="l">web 4 Ticker; api 2 Ticker2</span><span class="w">
</span></span></span></code></pre></div><p>From this order the catalog creates an air, a control plane machine and three nodes, and the cluster forms by itself. Any machine&rsquo;s screen opens with <code>virtctl vnc</code> under the tenant&rsquo;s own rights, and the machine names are shown in the dashboard. The source of truth here is the form: its deployments are applied at every start of the control plane, so a changed form takes effect after a restart, and a change made over VNC lasts only until then. Deleting the order deletes all its parts too.</p>
<p>You can also build a cluster by hand from separate machines. Then create an <code>OberonAir</code>, and in each <code>OberonVM</code> name that air in the field <code>air</code>, the role in <code>kubeRole</code> (<code>plane</code> for the control plane and <code>node</code> for the nodes), the node name in <code>kubeNode</code> and the same key in <code>kubeKey</code>. The control plane&rsquo;s deployments go in the field <code>commands</code>, and the cluster name, if it should not be <code>kube</code>, in <code>kubeCluster</code>. This is more fun if you want, for example, to put two clusters with different keys on one air and watch them not get in each other&rsquo;s way.</p>
<h2 id="what-all-this-says-about-tomorrows-infrastructure">What all this says about tomorrow&rsquo;s infrastructure</h2>
<p>At the start I mentioned that this is research as well, not just a joke, so let me try to put into words what the cluster on Wirth&rsquo;s radio taught me. Not in the sense that everyone should rewrite Kubernetes in Oberon, but in the sense of which ideas are worth carrying from this small system into a large one.</p>
<p><strong>State, not events.</strong> Inside its control plane Kubernetes has long worked on the level-triggered principle, but between components it still has events, watches, change streams and long-lived connections. Kube went further simply because the radio left it no choice: every message carries the full state, and the protocol needs no acknowledgements, no retries and no recovery after a broken connection. I think this approach applies far more widely than is usually assumed. Wherever the state can be described briefly enough and the network is unreliable, be it edge networks, satellite links, industrial networks or clusters of thousands of small devices, a protocol in which every message stands on its own turns out both simpler and more reliable.</p>
<p><strong>Observability as a property of the protocol.</strong> Since the air carries everything the cluster knows about itself, I did not have to build observability. In a large system you cannot broadcast everything to everyone, of course, but the very thought that what should be observed is the messages between components, not the components&rsquo; reports about themselves, seems interesting. A check that looks at the behaviour of the nodes, not at what the control plane thinks of them, catches the control plane&rsquo;s own bugs, and that is how nearly all of Kube&rsquo;s bugs were found.</p>
<p><strong>A unit of deployment the size of a module.</strong> An Oberon module with an interface key is very close to where WebAssembly is now heading with its component model: a small piece of code with an explicitly described interface, which can be loaded quickly and checked for compatibility before it runs. Oberon lacked isolation, and WebAssembly provides it without a separate process or kernel. Put the two together, and you get a pod that starts like a function call, is isolated like a container, and has its compatibility checked at load time, as in Oberon. I think tomorrow&rsquo;s infrastructure may well be built on such principles.</p>
<p><strong>Limits written in the code.</strong> All of Kube&rsquo;s ceilings sit in two blocks of constants, and each has a reason. Kubernetes has limits too, but they are scattered across flags, documentation and operating experience, and many of them you learn about only when you hit them. I would like large systems to declare their limits and the reasons for them more honestly, rather than pretend there are none.</p>
<p><strong>Simplicity that fits in a head.</strong> Kube can be read in an evening, and that is not decoration but a working property: when something went wrong, I found the cause by reading one procedure, not by going through issues in a tracker. Kubernetes will never be like that again, but new systems can be designed so that their core, the part responsible for the main idea, stays surveyable, with everything else built around it in layers you need not read.</p>
<p><strong>Security in the very first version.</strong> The radio forced messages to be signed from the very start, and it cost about thirty lines. Systems that began with a trusted network and added protection later paid more for it. The lesson is trite, but Oberon illustrates it well: if the environment is hostile from day one, the right architecture comes about by itself.</p>
<p>And separately, about what not to do. Kube shows clearly that isolation, preemption and separation of rights are not superfluous complexity to be thrown out for simplicity&rsquo;s sake, but what a shared computer is impossible without. This is exactly where Oberon lost, and any system that wants to be both simple and shared will have to find a way to get isolation without losing surveyability. It seems to me that today, for the first time, suitable building blocks for this exist, from WebAssembly to small verifiable kernels, and the only question is whether anyone will want to put together from them a system that can once again be read in full.</p>
<h2 id="instead-of-a-conclusion">Instead of a conclusion</h2>
<p>I set out wanting to check whether the idea of Kubernetes fits into a machine designed from start to finish by one person. It fits, together with rollouts, disaster recovery, a signed protocol and a store on disk, in about fourteen hundred lines of a language from the late eighties. Along the way it turned out that the machine&rsquo;s limits did not only get in the way but also suggested solutions, and some of them proved better than the ones we are used to. It also became clear exactly where such systems hit their ceiling, and that is perhaps the most valuable part, because it explains why the world took another road.</p>
<p>Next on the list are a readiness check that the workload itself answers, so that a rollout waits until new code is really ready rather than merely started, and a way to apply deployments from outside the control plane without VNC. And perhaps isolation, at least partial, to see how much simplicity it would cost.</p>
<p>The project is open. The sources are in the <a href="https://github.com/tym83/paleocomputing">repository</a>; our code is under the Apache-2.0 licence, and the QEMU machine model, like QEMU itself, under the GPL. The cluster lab lives at <a href="https://tym83.github.io/paleocomputing/oberon/kube.html">tym83.github.io/paleocomputing/oberon/kube.html</a>, and the other labs of the series and the page on the Cozystack catalog are on the <a href="https://tym83.github.io/paleocomputing/">project site</a>. You can read about Cozystack itself at <a href="https://cozystack.io/">cozystack.io</a>. I will be glad if someone switches the air off in the lab not for half a minute but in some cleverer way, and finds where Kube breaks. There have been many such findings in this story already, and each one taught something.</p>
<p>In the next part of the paleocomputing series we will build, from the specifications and the tender documents, Ada&rsquo;s rivals: the programming languages that took part in the US Department of Defense competition but failed to win.</p>
]]></content:encoded></item><item><title>Paleocomputing, part 1: Wirth's processor, Project Oberon, a new architecture in QEMU and KubeVirt, and running it in Cozystack and K8s</title><link>https://aenix.io/blog/2026/09/nine-days-of-paleocomputing/</link><guid isPermaLink="true">https://aenix.io/blog/2026/09/nine-days-of-paleocomputing/</guid><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><dc:creator>Timur Tukaev</dc:creator><category>Cozystack</category><category>KubeVirt</category><category>Kubernetes</category><category>Open Source</category><category>CHERI</category><category>Retrocomputing</category><description>Running Niklaus Wirth's RISC5 processor in a browser, in QEMU and in Kubernetes, and measuring what array-bounds checking really costs on a fully open system.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/nine-days-of-paleocomputing.jpg" alt=""></p><p>On 1 January 2024, Niklaus Wirth died — the man who gave us Pascal, Modula-2 and Oberon, won the Turing Award, and spent his whole life waging a stubborn war on bloated software. What fewer people know is that, already well into his seventies, he sat down and designed a processor of his own: small and simple, so that he could show students a whole computer at once, from the logic gates up to the windows on the screen.</p>
<p>In September 2026 I took that processor and tried to run it everywhere I could reach. First right inside a browser tab — and not as an emulator, but as the very circuit Wirth drew. Then in QEMU, as an ordinary virtual machine. Then in Kubernetes, which is normally home to a very different kind of VM, the sort that runs Ubuntu and databases. And finally in Cozystack, our cloud platform, where Wirth&rsquo;s machine can now be installed with a single click from the catalog, the way you&rsquo;d install PostgreSQL. Along the way I finally measured something programmers have argued about for decades — what array-bounds checking actually costs. I also ran a small language model on Wirth&rsquo;s processor. And, since we&rsquo;re being honest, with one careless move I sent every virtual machine in a live production cluster off on a migration (nobody was hurt, but it was an unpleasant few seconds).</p>
<p>This turned into a very long article, because a great deal happened over nine days, and a good part of it was my own mistakes, which I then had to find and fix. I&rsquo;ve tried to write so that someone who has never heard of Oberon can follow along, and I&rsquo;ve tucked everything only a specialist would care about under spoilers. You can read straight through, or jump to whatever part interests you. First comes the story of Oberon itself, then how my original plan fell apart, then the processor and bounds-checking, then the browser and the lab exercises, then QEMU, Kubernetes and Cozystack, and right at the end the language model and how to install all of this yourself.</p>
<p>All the source code, notes and instructions live in the repository <a href="https://github.com/tym83/paleocomputing">github.com/tym83/paleocomputing</a>, and the project site is <a href="https://tym83.github.io/paleocomputing/">tym83.github.io/paleocomputing</a>. If you&rsquo;d rather get your hands dirty first and read afterwards, open the <a href="https://tym83.github.io/paleocomputing/oberon/lab.html">lab</a>: nothing to install, the machine boots right there in your browser.</p>
<figure><img src="oberon-boot-screen.png" alt="Oberon booted on Wirth&#39;s actual circuit" width="1024" height="768" loading="lazy" decoding="async">
  <figcaption><p>Oberon on Wirth&rsquo;s actual circuit. I&rsquo;ve seen this picture five hundred times, and almost every time it matched the reference down to the last dot.</p></figcaption>
</figure>

<h2 id="what-oberon-is">What Oberon is</h2>
<p>Oberon is at once a programming language and an operating system, built in the mid-1980s at ETH Zürich by Niklaus Wirth and Jürg Gutknecht. They started in the autumn of 1985, and by 1988 the system was genuinely up and running. By Wirth&rsquo;s own account, two people wrote it in the hours left over from their day jobs — which, you have to admit, sounds fairly wild when you think about how many people it takes to write any operating system today. It ran on the Ceres workstation, also built at ETH, and students were taught on it until roughly the early 2000s.</p>
<p>The name, incidentally, came from Voyager. In January 1986 the probe sent back images of Uranus and its moons, and Wirth — who regarded Voyager as a model engineering project, a craft that kept working far beyond its design life — named the system after the moon Oberon. In the book he calls Oberon the largest moon of Uranus, though in fact Titania is bigger; even Wirth has his typos. And, for good measure, Oberon is also the king of the elves, which isn&rsquo;t a bad namesake either.</p>
<p>The whole enterprise has its guiding idea set out in the preface to the project&rsquo;s book: a system built from scratch should be something you can describe, explain and understand in full — one person should be able to read and understand the whole of it, from the processor to the windowing interface. By today&rsquo;s standards that sounds almost like fantasy, because nobody alive understands a modern OS, together with its compiler and its processor, in full.</p>
<p>In 2013 Wirth released a new edition of the project, and in it he carried the idea all the way. The processor Ceres ran on had long since gone out of production, and Wirth decided that since everything else in the system was his own and comprehensible, the processor should be his own too. He designed a simple RISC processor — called RISC5 in the sources — and described it in Verilog, the language used to describe digital circuits. A circuit like that can be flashed onto an FPGA, a programmable chip, and out comes a real, working computer. Wirth&rsquo;s ran on an inexpensive board with one megabyte of memory at 25 MHz. That&rsquo;s roughly a hundred times slower than a single core of your phone, but it is more than enough for Oberon, because the compiler, as Wirth writes, compiles itself in about three seconds. The book and the sources for the system, the compiler and the processor are all in the open, and that is exactly what makes Oberon unique. One more detail that will matter to us: the system itself dates from the late 1980s, while the processor we&rsquo;ll be running it on appeared only in the 2010s.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">What makes Oberon so unusual, if you&#39;ve never seen it</span>
  </summary>
  <div class="spoiler__body">
    <p><strong>Any text can be a command.</strong> Oberon has no command line in the usual sense. If somewhere on the screen, in any window, there&rsquo;s something of the form <code>Module.Procedure</code>, that is already a command. You point at it with the mouse, press the middle button, and it runs. There are menus too, but a menu here is just a line of text in a window&rsquo;s title bar listing commands. Want your own menu? Write the commands you need into a text file and open it. This idea went on to strongly influence the Acme editor in Plan 9, and Rob Pike said as much himself.</p>
<p><strong>You need a three-button mouse.</strong> The left button places the cursor, the middle one executes, the right one selects — and there are also chords, where you hold one button down and, without releasing it, press a second. On a laptop with no middle button we stand in for it with clicks plus modifier keys; more on that below.</p>
<p><strong>There is always exactly one program.</strong> The multitasking you&rsquo;re used to isn&rsquo;t here. Underneath, a single loop spins, polling the keyboard and mouse and calling commands one at a time, and until a command finishes, nothing else happens. Wirth himself admits this sounds very limiting, and explains at length why, for one person at one computer, it is enough.</p>
<p><strong>No header files, and no dependency hell.</strong> Each module describes its own interface, and that description carries something like a checksum. If the interface changes, the system simply refuses to run modules built against the old version — so a program built against a stale library and crashing somewhere inscrutable is, here, impossible in principle. And this is 1988.</p>
<p><strong>There is no memory protection; the language stands in for it.</strong> The processor has no mechanism to stop one program from reaching into another&rsquo;s memory. All the safety rests on the language being strictly typed and collecting its own garbage, with the compiler checking every array access. We&rsquo;ll break exactly this in the lab exercises.</p>
<p><strong>The language is small.</strong> The full report on Oberon-07 runs to 17 pages, and its epigraph is Einstein&rsquo;s line about making everything as simple as possible, but not simpler.</p>

  </div>
</details>

<p>Why bother with a museum piece in 2026? I have three reasons. The first is how deeply Oberon shaped what we use today. Go — the language half of today&rsquo;s cloud infrastructure is written in, Kubernetes included — names Oberon outright as one of its ancestors; it&rsquo;s right there in the <a href="https://go.dev/doc/faq">Go FAQ</a>. And Robert Griesemer, one of Go&rsquo;s three authors, did his doctorate under Wirth at ETH, and in his talk <a href="https://www.youtube.com/watch?v=0ReKdcpNyQg">&ldquo;The Evolution of Go&rdquo;</a> he draws a family tree with a straight line running from Oberon to Go. The second reason is that Oberon really does fit in your head. If you want to understand how a computer works from the gate to the window, I know of no better teaching aid. And the third reason is that in 1995 Wirth wrote <a href="https://people.inf.ethz.ch/wirth/Articles/LeanSoftware.pdf">&ldquo;A Plea for Lean Software&rdquo;</a>, a manifesto against bloated software — the source of Wirth&rsquo;s law, that software slows down faster than hardware speeds up, though Wirth honestly credited the observation to his colleague Martin Reiser — and Oberon was its proof that things could be otherwise. Thirty years on, matters have, to put it mildly, only got worse.</p>
<p>And Oberon, pleasingly, is still alive. It has active forks that were updated just days ago, and a Russian-speaking community, OberonCore, briskly discussing all of it.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">What to read and watch about Oberon (the best I found)</span>
  </summary>
  <div class="spoiler__body">
    <ol>
<li><a href="https://people.inf.ethz.ch/wirth/ProjectOberon/index.html">Project Oberon, 2013 edition</a> — the primary source: the book on the system, the compiler and the processor, with all of its sources.</li>
<li><a href="https://www.projectoberon.net/">projectoberon.net</a> — Paul Reed&rsquo;s site, the most convenient way in: mirrors, archives, a ready-made disk image.</li>
<li><a href="https://people.inf.ethz.ch/wirth/Oberon/Oberon07.Report.pdf">The Programming Language Oberon (Oberon-07)</a> — the whole language in 17 pages, an evening&rsquo;s read.</li>
<li><a href="https://people.inf.ethz.ch/wirth/Articles/Modula-Oberon-June.pdf">Wirth, &ldquo;Modula-2 and Oberon&rdquo;</a> — the history first-hand: what Wirth took from Xerox PARC and what he deliberately threw out. My favorite line from it: &ldquo;We were inspired by what could be done, and shown how not to do it.&rdquo;</li>
<li><a href="https://people.inf.ethz.ch/wirth/Articles/LeanSoftware.pdf">Wirth, &ldquo;A Plea for Lean Software&rdquo;</a> — the 1995 manifesto itself.</li>
<li><a href="https://people.inf.ethz.ch/wirth/CompilerConstruction/CompilerConstruction1.pdf">Wirth, &ldquo;Compiler Construction&rdquo;</a> — a short, very readable textbook on compilers, built around a pared-down Oberon for the same processor.</li>
<li><a href="https://people.inf.ethz.ch/wirth/FPGA-relatedWork/RISC-Arch.pdf">The RISC Architecture</a> — the processor described in a few pages.</li>
<li><a href="https://www.youtube.com/watch?v=5niGplCza7s">Video: Wirth demonstrates Oberon on Ceres, 2011</a>.</li>
<li><a href="https://inf.ethz.ch/news-and-events/spotlights/infk-news-channel/2021/11/niklaus-wirth-video-interview.html">ETH video interview with Wirth, 2021</a> — the link opens the first of three parts.</li>
<li><a href="https://www.youtube.com/watch?v=0ReKdcpNyQg">Griesemer, &ldquo;The Evolution of Go&rdquo;</a> and the <a href="https://go.dev/talks/2015/gophercon-goevolution.slide">slides</a> — the bridge from Oberon to Go.</li>
<li>Emulators: <a href="https://github.com/pdewacht/oberon-risc-emu">oberon-risc-emu</a> by Peter De Wachter, <a href="https://schierlm.github.io/OberonEmulator/">OberonEmulator</a> by Michael Schierl (runs in the browser), <a href="https://github.com/pdewacht/project-norebo">Norebo</a> — an Oberon compiler you can run from an ordinary command line, without which half of this article wouldn&rsquo;t exist, and <a href="https://github.com/andreaspirklbauer/Oberon-extended">Extended Oberon</a> by Andreas Pirklbauer.</li>
<li>In Russian: the <a href="https://oberoncore.ru/library/start">OberonCore library</a> with translations of Wirth, the <a href="https://forum.oberoncore.ru/">forum</a>, <a href="https://www.inr.ac.ru/~info21/">Informatika-21</a>, and the book <em>Project Oberon: The Design of an Operating System and Compiler</em>, published by DMK Press in 2012. Judging by the year, that&rsquo;s a translation of the pre-2013 edition — the one without Wirth&rsquo;s processor — though I haven&rsquo;t checked myself; correct me if I&rsquo;m wrong.</li>
</ol>

  </div>
</details>

<h2 id="where-this-whole-idea-came-from">Where this whole idea came from</h2>
<p>I&rsquo;d been keeping a list of forgotten systems for a long time. Operating systems, languages and whole machines that once worked, sometimes very well, and then vanished — often taking with them ideas nobody has properly reproduced since. The Burroughs B5000, which back in 1961 checked data types right there in the hardware. KeyKOS, for which a power cord yanked from the wall was a normal event, not a disaster. Transputers, Lilith, the iAPX 432. Articles and books have been written about all of them, but almost nobody actually runs them, and running them was exactly what I wanted — for real, with all the guts, to see what they can do.</p>
<p>That&rsquo;s how the series I called Paleocomputing came about. The idea behind it is simple. Many of these systems lost not because they were bad but because they were expensive. Custom silicon for a custom language cost insane money, while ordinary Intel processors got cheaper faster than anyone could build something clever. The economics are different now. Programmable chips cost pennies, the open RISC-V architecture officially permits adding your own instructions to the processor, and the CHERI project — much more on it below — essentially brings back into modern processors the very hardware memory protection Burroughs was doing before I was born. Which means the old ideas can be tested by hand again, and that is what I set about doing.</p>
<p>For the first episode I chose Oberon, though I wanted Burroughs more. The reason is straightforward: Oberon has absolutely everything you need to work. The processor sources, the compiler, the operating system, the book that explains all of it, living emulators and ready-made disk images. You can dive straight to full depth instead of spending a month excavating documentation — and to start a series, scale was exactly what I needed.</p>
<h2 id="how-it-was-done">How it was done</h2>
<p>Let me admit up front that most of the code in this project was written not by me but by a neural network. More precisely, by several AI agents to whom I handed out different roles. Some wrote code, others reviewed it, and still others were given the job of finding out why the whole thing didn&rsquo;t work — and took it very seriously. They&rsquo;ll turn up throughout the text as reviewers and auditors, and I want it clear from the start that these are not living people but the same neural networks, which I set to play the part of a captious reader, a hardware specialist, or a measurement methodologist. They worked independently of one another and of the agent writing the code, and that, I think, turned out to be the single most useful invention of the whole project. What stayed with me were the design, the decisions, the endless questions about whether everything had really been checked, and a live cluster that, as it turns out, is very easy to upset.</p>
<p>I mention this not to ride a fashionable topic but because the rest of the story makes no sense without it. First, everything described here was done in nine days, from the 21st to the 29th of September, and that pace would have been impossible had I written it all by hand. Second, it&rsquo;s the reason some of the code can&rsquo;t be handed to the big open-source projects, which I&rsquo;ll come to in the section on QEMU. And third, a neural network has one very characteristic habit: it loves to report that everything is done and all the checks are green. In practice that often meant the check simply wasn&rsquo;t checking anything. So most of those nine days went not into writing code but into teaching the checks to blush honestly when something was broken — a thread that runs through the whole article.</p>
<h2 id="the-plan-that-fell-apart-in-a-day">The plan that fell apart in a day</h2>
<p>On paper the plan sounded splendid. Wirth&rsquo;s real processor was to run right inside a browser tab, cycle by cycle — and not as an emulator written from the documentation, but as his own circuit. A reader could change the instruction set of that processor right there on the page, press a button, and a few seconds later watch the operating system carry on running on the altered hardware. And on top of that I meant to lay three loud claims. That the whole computer, from gates to windows, runs in a browser. That if you add to the processor a special instruction which multiplies and adds in one go, a small neural network on it will run at least twice as fast, and you can do it in about a minute. And that if you add hardware array-bounds checking to the processor, you can for the first time honestly measure what such a check costs — because, as I confidently wrote in the draft, nobody has that number.</p>
<p>I gave myself two to four weeks. Right.</p>
<p>Before sitting down to write code, I sent the plan out to five reviewers, each with an area of their own: the schedule, the hardware, the compiler, the measurement methodology, and the browser. The brief was the same for all of them: find why this won&rsquo;t work. The comments came to more than eighty kilobytes of text, and after them not much of the plan was left standing.</p>
<p>The most galling part was that the room for new instructions in the processor, which I thought I&rsquo;d found, wasn&rsquo;t free at all. I&rsquo;d seen a bit in the instruction encoding that is always zero, and I was going to seat my new instructions there. The hardware reviewer opened Wirth&rsquo;s circuit and showed that the processor simply doesn&rsquo;t look at that bit — so on a real processor all my new instructions would silently execute as the most common instruction of all, a data move, and the program would quietly do something entirely unlike what was written. There was no free room in the instruction set at all. Later, it&rsquo;s true, one small spot did turn up, where Wirth&rsquo;s compiler always writes zeros, and there&rsquo;s a separate story about how I tried to squeeze in there.</p>
<p>The two-times speed-up for the neural network didn&rsquo;t survive the review either. The reviewers counted straight off the circuit how many cycles each operation takes, and even if the new instruction were entirely free, it could not in principle speed the program up by more than a factor of two — and the one actually built would be lucky to manage a few percent. And right there was the key point, later borne out: the bottleneck isn&rsquo;t the missing instruction at all, it&rsquo;s that Wirth&rsquo;s multiplier is very slow, computing the product one bit per cycle.</p>
<p>The number nobody supposedly has proved quite an embarrassment. The methodologist brought a list of papers where the cost of hardware memory protection had been measured many times over, including on modern Arm processors and in the CHERI project. About the only thing that was relatively new was that I wanted to use, as the workload, a system that rebuilds itself — but even that is a detail rather than a discovery.</p>
<p>It also emerged that in Oberon you can&rsquo;t turn off array-bounds checking in ordinary code at all; it&rsquo;s welded into the compiler. There was simply nothing to compare Wirth&rsquo;s system against without checks, and I had to plan from the very start for my own compiler patch that produces a check-free build.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">What else the reviewers tore apart</span>
  </summary>
  <div class="spoiler__body">
    <p>I was going to report results with a confidence interval, as proper statistics demands. The methodologist pointed out that on a simulator this is meaningless, because the same run always yields the same number, down to the cycle — there&rsquo;s no spread there for an interval to come from. Spread appears only if you take many different programs, and that is what you should be measuring.</p>
<p>The reviewers also noted that on a real machine the video controller takes about 7% of the processor&rsquo;s time, because it constantly reads memory for the screen, and that ignoring it makes every number a little rosier than reality. That two multiplications in a row on Wirth&rsquo;s machine cost more than two multiplications apart. That the standard open tool for estimating a circuit&rsquo;s area loses more than half of the memory elements by default. And that the whole plan had not a single intermediate point where you could stop and show people something — which, for a project done in your spare time, is fatal.</p>
<p>The revised schedule estimate after the review: nine to thirteen weeks working evenings. We managed it in nine days, but only because the code was written by agents — which I&rsquo;ve already mentioned.</p>

  </div>
</details>

<p>After the review I rewrote the plan. Two claims remained in the first episode. The first, that the whole computer runs in a browser. The second, that you can honestly measure the cost of bounds-checking on a fully open system where everything is visible, from the processor circuit to the compiler. The neural network and its new instruction I moved to a separate episode. And a line I&rsquo;m very fond of appeared in the risk table: that the number will most likely come out boring, around six percent — and that this is the result.</p>
<p>Getting ahead of myself: I did eventually run a neural network on Wirth&rsquo;s processor, and the reviewers were right about the main thing — it all came down to the multiplier. But that&rsquo;s nearer the end.</p>
<h2 id="running-wirths-real-processor">Running Wirth&rsquo;s real processor</h2>
<p>It all starts simply. Wirth described his processor in Verilog, and that description isn&rsquo;t a program but a circuit — registers, wires, and the logic between them. To run a circuit like that without an actual chip, there&rsquo;s a tool called Verilator, which turns it into a C++ program that faithfully recomputes every wire of the processor on every cycle. It runs slower than real hardware, but exactly like it, with no approximations. The processor still needs a screen, a disk, a keyboard and a mouse bolted on. I built that harness myself, reproducing precisely the interface the system expects, and I didn&rsquo;t touch the processor itself by a single line. That was a point of principle for me, because everything I&rsquo;m going to measure later has to be measured on his circuit, not on mine.</p>
<p>The first working day began at one in the morning on 22 September, and by lunchtime Wirth&rsquo;s real circuit had booted the real Project Oberon. Booting takes about twelve million instructions, and on my laptop the simulation gets through it in four seconds. The picture at the top of the article is from this very run.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">What the boot hung on until lunchtime</span>
  </summary>
  <div class="spoiler__body">
    <p>First I picked the wrong bootloader. Wirth&rsquo;s repository holds a bootloader that receives the system over a serial port and doesn&rsquo;t read from disk at all, and since its beginning is identical to the one I needed, I spent quite a while staring at code that looked right.</p>
<p>Then the processor executed nonsense as its first instruction. While the processor is resetting, it still reads an instruction off the data bus, and if there are zeros there at that moment, the first thing executed is a meaningless data move instead of a jump to the bootloader, and the machine just sits. I stepped on this same bug three more times on different test rigs, and once it hid extremely well — more on that below.</p>
<p>And the third was the screen upside down, because Wirth&rsquo;s video controller reads the image out of memory from the bottom up. The hardware reviewer predicted it in advance, and the warning saved me a heap of time.</p>

  </div>
</details>

<p>Getting it to run is only half the job; you have to be sure it ran correctly. For that there&rsquo;s a technique called lockstep comparison. You take a reference you trust and drive it alongside your own machine through the same program, instruction by instruction, comparing the contents of every register after each one. If even a single bit diverges anywhere, you find out immediately, and you know exactly which instruction it was on. For the reference I took the emulator from the Norebo project, and across the entire system boot — nearly fifteen million instructions — my circuit never once diverged from it. There was a divergence at first, mind you, and it was my fault: on a reviewer&rsquo;s advice I&rsquo;d written a housekeeping value into memory just in case, one the emulator writes only in a particular situation that didn&rsquo;t arise in this run. The right thing was to touch nothing.</p>
<p>Here I should say why this matters at all. Every cycle count in this article is computed by a separate fast model, because running the real circuit under heavy loads takes far too long. And that model can be trusted exactly as far as it matches the circuit. I checked the two against each other on the same boot and got a perfect match, down to the cycle — but at first the comparison claimed the model was off by nearly sixteen percent, and I very nearly believed it. The comparison itself was lying, because it was peeking at the instruction address one cycle earlier than it should. Had I published back then that the model was 16% off, it would have devalued every number in the project. Since then I have a rule: a negative result must be checked just as carefully as a positive one, because you can go wrong in both directions.</p>
<h2 id="what-array-bounds-checking-costs">What array-bounds checking costs</h2>
<p>Now for the central question of the first episode. If you write <code>a[i]</code> in C, and <code>i</code> happens to be larger than the array, the program will silently read or write someone else&rsquo;s memory. From that simple mistake grew a huge share of the vulnerabilities of the last fifty years, from the buffer overflows of the nineties to today&rsquo;s browser security bulletins. Guarding against it is very simple: before every array access, compare the index against the length and, if it&rsquo;s out of bounds, stop the program. Many languages do exactly this; Oberon always does; and in C and C++ the check traditionally isn&rsquo;t written, because it&rsquo;s held to be slow. So I wanted to find out just how slow it actually is, on a system where absolutely everything is visible.</p>
<p>For that you need three variants of the same system. The first has no checks at all, and no such Oberon exists, so I made it myself by changing a single line in the compiler. The second has the software check, as Wirth&rsquo;s does, where the compiler inserts a comparison and a jump to the fault handler before every access. And the third has a hardware check, where the processor gains a new instruction that does all of this itself in one cycle. That instruction gets a section of its own. For the workload I took the Oberon compiler compiling several modules of the system itself, because it&rsquo;s a real, large program rather than a synthetic test.</p>
<p>The first number I got was under half a percent, and I was delighted that checks cost almost nothing. Then I realized I&rsquo;d compared the wrong things. I was running two different compilers, one with checks and one without, and the check-free compiler simply did less work, because it didn&rsquo;t have to insert checks into anyone else&rsquo;s code. What you must compare is the same work — that is, build one and the same compiler in two variants and give them the same task. Once I fixed that, the number grew several-fold.</p>
<p>There were a few more passes after that, and the honest answer was this. As long as the checks are removed only from the compiler itself, they cost about two percent of the time. And if you remove them from the whole system the compiler runs on, including its handling of text, files and memory, it comes to almost five percent. So for every hundred cycles of the compiler&rsquo;s work, between two and five go towards making sure the program never runs off the end of an array. Whether that&rsquo;s a lot or a little, each reader can decide; to my taste it&rsquo;s very cheap for the absence of an entire class of vulnerabilities. With one caveat the reviewers made on the very first day: I have only one workload, so this is a number about the Oberon compiler, not about every program in the world.</p>
<p>Here the first surprise was waiting for me. When I recounted which checks actually sit in the code, I found there are very few array-bounds checks in Oberon, and the overwhelming majority are checks that a pointer isn&rsquo;t empty — that it isn&rsquo;t NIL. In the compiler alone there are more than three hundred of these against a handful of index checks. So the cost of safety in Oberon comes mostly from the program being unable to dereference a null pointer, and arrays as such are a secondary matter here.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">How I got caught retelling other people&#39;s papers</span>
  </summary>
  <div class="spoiler__body">
    <p>First I compared my percentages against published work on memory protection on other processors, and got a very pretty little table. The auditor withdrew it in its entirety, because I&rsquo;d misrepresented three of the four numbers. In one place I&rsquo;d taken the figure for a lightweight protection mode, while the full one cost five times more. In another the bulk of the loss came from cache misses, and Wirth&rsquo;s processor has no cache at all, so comparing that with our number is meaningless. In another I&rsquo;d mistaken two different operating modes for a range of values. And in one paper the result was given as a multiplier, and I&rsquo;d read it as a percentage, so a slowdown of one-and-a-half to two-and-a-half times turned, in my hands, into a slowdown of one-and-a-half to two-and-a-half percent. That last one I&rsquo;m most ashamed of.</p>
<p>The closest thing to my own setup turned out to be a dissertation where several CHERI processors came in at 10–16%, but what&rsquo;s measured there is full pointer protection, a noticeably broader thing, so it&rsquo;s a bearing rather than a comparison. And the claim that nobody has this number didn&rsquo;t survive even this, because as far back as 1981 there was a paper measuring the cost of checks in Pascal. Measuring it, of course. The correct thing to say is only that we didn&rsquo;t find something in such-and-such sources, and that&rsquo;s the only way I put it now.</p>

  </div>
</details>

<h2 id="how-i-taught-the-processor-to-check-for-itself">How I taught the processor to check for itself</h2>
<p>Now the hardware check. The idea is simple. Instead of two instructions — a comparison and a conditional jump — the processor gains one new instruction, which I named <code>CHK</code>. It takes an index and an array length and, if the index is out of range, jumps to the fault handler itself. All of it in one cycle instead of two.</p>
<p>The difficulty turned out to be where in the instruction to put the array length. An instruction on Wirth&rsquo;s processor is 32 bits, and almost all of them are already taken. The only place the processor genuinely doesn&rsquo;t use — which I proved by enumerating every possible value — is twelve bits in the middle of the instruction. Twelve bits means arrays of up to four thousand elements, which in Oberon is the overwhelming majority, so I was pleased and put the length there.</p>
<p>And I got a very nasty result. When a check fires in Oberon, the system reports where and exactly what error occurred, and it takes the error number from the very same instruction bits where I&rsquo;d put the length. I ran an experiment with a real error, a program that reaches past the end of a hundred-element array. The ordinary system honestly reported that the index was out of the array&rsquo;s bounds. Mine, with the new instruction, reported just as confidently that there had been a null-pointer dereference, because a chunk of the number 100 had landed in the error-number field. A programmer getting that message would go off hunting a non-existent pointer bug and lose half a day to it. That&rsquo;s worse than the system saying nothing at all.</p>
<p>After that I tried to wriggle out of it a couple of times by shrinking the length to eight bits, and both times it went wrong. First the statistics showed eight bits were enough for most arrays, but then it emerged that the majority in those statistics were identical buffers for file names, hardly ever used, while the hot loops that run millions of times work with precisely the large arrays and don&rsquo;t fit in eight bits. The solution turned up elsewhere. The new instruction, as it happens, doesn&rsquo;t need a destination register, because it computes nothing and only checks, and the four bits that usually name that register are free. The length can be sliced in two: the four high bits go there, and the eight low bits into the middle of the instruction, positioned so they don&rsquo;t touch the error number. That gave both the twelve bits of length and the correct error messages.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">What the new instruction looks like in the processor circuit, and the traps it held</span>
  </summary>
  <div class="spoiler__body">
    <p>The entire processor edit sits under one switch, so you can build exactly the same circuit with and without the new instruction and compare them (simplified here, without the two-piece reassembly of the length):</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-verilog" data-lang="verilog"><span class="line"><span class="cl"><span class="no">`ifdef</span> <span class="n">WITH_CHK</span>
</span></span><span class="line"><span class="cl"><span class="k">assign</span> <span class="n">CHK</span>     <span class="o">=</span> <span class="o">~</span><span class="n">p</span> <span class="o">&amp;</span> <span class="o">~</span><span class="n">q</span> <span class="o">&amp;</span> <span class="o">~</span><span class="n">u</span> <span class="o">&amp;</span> <span class="n">v</span> <span class="o">&amp;</span> <span class="p">(</span><span class="n">op</span> <span class="o">==</span> <span class="mh">1</span><span class="p">);</span>
</span></span><span class="line"><span class="cl"><span class="k">assign</span> <span class="n">chkFail</span> <span class="o">=</span> <span class="n">CHK</span> <span class="o">&amp;</span> <span class="p">(</span><span class="n">B</span> <span class="o">&gt;=</span> <span class="n">chkLim</span><span class="p">);</span>
</span></span><span class="line"><span class="cl"><span class="no">`else</span>
</span></span><span class="line"><span class="cl"><span class="k">assign</span> <span class="n">CHK</span>     <span class="o">=</span> <span class="mh">1</span><span class="mb">&#39;b0</span><span class="p">;</span>
</span></span><span class="line"><span class="cl"><span class="k">assign</span> <span class="n">chkFail</span> <span class="o">=</span> <span class="mh">1</span><span class="mb">&#39;b0</span><span class="p">;</span>
</span></span><span class="line"><span class="cl"><span class="no">`endif</span>
</span></span></code></pre></div><p>At first the first line had no <code>~u</code>, and because of that the new instruction occupied two encodings instead of one, hijacking another, unrelated one along the way. This was found by a check that enumerates every possible instruction on the ordinary processor and on the processor with the new instruction and looks for where they behave differently. The right answer is in exactly one place; it found two.</p>
<p>The auditor also worked out what happens if a module carrying the new instruction is accidentally run on an ordinary processor. The ordinary processor takes it for a shift and silently corrupts one of the registers — and which one depends on the array length, and at a length around four thousand it&rsquo;s the register holding a procedure&rsquo;s return address. From the outside this would look like inexplicable stack corruption. Now such modules are tagged with a new format version, and the ordinary system simply refuses to load them. The funny thing is that the requirement to tag the version was in the very first review; I accepted it and then lost it.</p>

  </div>
</details>

<p>So how much does the hardware check save in the end? On the compiler it removes about a sixth of the cost of checks; on a mixed load with arrays of various sizes, a tenth; and on purely computational code with small arrays, a half. A half, by the way, is the ceiling, and it will never be more, because two instructions became one. The difference between the loads has a very simple explanation. The instruction helps only where the array length fits in twelve bits. Where the array is larger, the compiler falls back on the old software check, and there&rsquo;s no gain.</p>
<p>And one more lovely story from this section. For the second workload I wrote a small program with a sort and a matrix multiply. In the checked variant it fell over immediately with an out-of-bounds message, and the bug was in my program itself: I&rsquo;d computed an index wrong and was overrunning the array by a factor of three and a half. The check-free variant ran silently, as though all were well, and quietly wrote two and a half thousand numbers into someone else&rsquo;s memory. The bug was found only because the checks were on, and honestly I can&rsquo;t think of a better advertisement for bounds-checking.</p>
<p>How much the new instruction costs in hardware — how much room it takes on the die, and whether it slows the processor down — I measured too, and the answer came out boring. It adds less than a percent of logic and doesn&rsquo;t touch the clock frequency at all. The interesting part here is how I arrived at that boring answer, because along the way the noise in such measurements proved larger than the effect itself.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">How the noise turned out to be larger than the effect</span>
  </summary>
  <div class="spoiler__body">
    <p>A circuit&rsquo;s area is estimated with open synthesis tools, which turn a Verilog description into a set of logic elements. The first trap was that the tool, in its default mode, silently left out three quarters of the memory elements, and the only hint of it was a plus sign at the very end of one line of the report. The second was that, without explicitly stated constraints, it simply ignored the speed target, and the circuit for 200 picoseconds and for 50 nanoseconds came out identical down to the last digit.</p>
<p>But the worst was something else. I took Wirth&rsquo;s circuit and rewrote it four times so that the logic didn&rsquo;t change at all — adding redundant parentheses, say, or a pointless OR with zero. The area after that wandered by about as much as my new instruction adds. That is, the effect I wanted to measure was the same size as the difference from how a completely untouched expression happens to be written. So all one can honestly say is that it&rsquo;s under a percent; and as for frequency, one way of building it showed slightly faster, another slightly slower, and picking either would be choosing a convenient answer rather than measuring one.</p>
<p>Then I placed and routed the circuit on a real FPGA — that is, I asked the tools to lay it out across the actual cells of a specific chip and run the wires between them. Only after that can you see the frequency the circuit will really run at. The native 25 MHz held with a large margin, the new instruction didn&rsquo;t affect the frequency, and most of the delay in the circuit is in the wires, not the logic.</p>

  </div>
</details>

<h2 id="the-auditors-didnt-accept-the-work">The auditors didn&rsquo;t accept the work</h2>
<p>By the evening of the first day I had a lot of pretty numbers, and I handed them to five auditors for acceptance. At 17:40 they reported back, and all five wrote the same word: NOT ACCEPTED, in capitals.</p>
<p>The most unpleasant and, at the same time, the most useful was the mutation audit. Its idea is simple and very cruel. The auditor deliberately introduces plausible errors into the processor — changing a &ldquo;greater than or equal&rdquo; to a &ldquo;greater than&rdquo; in a comparison, say — and watches whether my tests notice. If the tests stay green on a broken processor, they aren&rsquo;t checking anything. Of thirty errors introduced, my tests caught ten. Two thirds of the broken processors passed every check as sound.</p>
<p>The reasons were very instructive. In one place a check that had failed to run counted as passed, and the report proudly announced that three of five cases had been checked and no errors found. In another, the test build was configured so that its failure didn&rsquo;t stop the run, and success was determined by a smiley on the last line of output, so zero tests executed gave a fully green report. And an instruction that supposedly checked the system boot in fact checked nothing at all: with a processor that had one of its instructions broken, the machine hung right at the start of the boot, and the check still reported success. On top of that, the system boot contains not a single operation on fractional numbers, so my claim that the comparison had checked all of the processor&rsquo;s arithmetic was simply untrue.</p>
<p>The auditors&rsquo; main conclusion was that almost all the errors happened in carrying the results over into the text. A number obtained under one set of conditions migrated into the summary as a general one. We rewrote everything we found, added the same kind of deliberate breakage to the automatic check on every change, and introduced a rule that any check must be able to fail, and that this must be demonstrated. And the next day one more gem surfaced. In all two hundred-plus processor tests, the very first instruction of the program wasn&rsquo;t actually executing, because of that same bus-during-reset bug I wrote about above. It was impossible to notice, because every test began with an instruction whose result didn&rsquo;t matter anyway. The most galling part is that I&rsquo;d already found and fixed this bug on another rig, and had even left a comment there about tripping on exactly this.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">Small joys of the early days</span>
  </summary>
  <div class="spoiler__body">
    <p>I wrote a script that waits for a long run to finish and checks for it by process name. The script waited an hour and a half, because it found itself — the process name was written in its own command line. Very Zen, and the whole time I thought a long compilation was under way.</p>
<p>Four billion instructions went to waste because the emulator and the processor write device addresses differently, and the check for whether the program touches a device never once fired.</p>
<p>One and the same field in a jump instruction is treated by Wirth&rsquo;s compiler as 24-bit, by the disassembler as 20-bit, and by the processor itself as 22-bit. Three parts of one system, written by one man, and each right in its own way.</p>
<p>One of the tools for preparing tests was silently corrupting the reference disk image right there in the repository, and the screen check on the corrupted image still matched. Now everything works on copies, and the reference has checksums beside it.</p>
<p>Two multiplications in a row on Wirth&rsquo;s machine cost almost one and a half times more than two multiplications apart, because the multiplier doesn&rsquo;t manage to reset its counter in time. I was pleased with the find at first, and then learned that the multiplier&rsquo;s slowness had been discussed on the Oberon forums as far back as 2016. This particular mechanism I didn&rsquo;t find there, but my not finding something doesn&rsquo;t mean nobody knew it.</p>

  </div>
</details>

<p>And at 19:30 on the first day the circle closed. The Oberon compiler, running on Wirth&rsquo;s real circuit, compiled itself, and the result matched byte for byte what the emulator produces. The next day the same thing came off inside the system itself, with windows and a mouse. A script of clicks and keystrokes made the system rebuild itself completely, and every file came out exactly as in the original disk image. Along the way it emerged that the official 2016 system image isn&rsquo;t entirely consistent with itself: a couple of its modules were out of date, and one was missing altogether. The ordinary life of any living project — and honestly, it rather moved me.</p>
<h2 id="oberon-in-a-browser-tab">Oberon in a browser tab</h2>
<p>Now the first claim of this episode, the one about the browser. Wirth&rsquo;s circuit, which Verilator turned into a C++ program, can then be compiled further into WebAssembly — the format in which a browser can run ordinary compiled code at nearly native speed. And then what&rsquo;s spinning in the tab isn&rsquo;t an emulator someone wrote from the documentation, but the very same circuit Wirth flashed onto his FPGA, with all its wires and cycles. I&rsquo;ll note again that the screen, disk, keyboard and mouse around the processor are mine, built to the same interface as Wirth&rsquo;s, and that the video controller, which on a real machine takes a little of the processor&rsquo;s time, doesn&rsquo;t figure in the measurements.</p>
<p>At first I was told this was a bad idea, because recomputing every wire is too slow for a browser. In practice the browser version runs almost as fast as the same simulation launched directly on the computer — a couple of percent apart. That&rsquo;s about six times slower than Wirth&rsquo;s real machine, so the system takes a few seconds to boot, and after that you can quite happily work with it. And the whole machine together with its disk image weighs about three hundred kilobytes — less than the average picture on any news site.</p>
<figure><img src="oberon-in-browser.png" alt="The same system running in a browser" width="778" height="1156" loading="lazy" decoding="async">
  <figcaption><p>The same system in a browser. The log with the Oberon splash, and a System.Tool window whose text you click with the middle button.</p></figcaption>
</figure>

<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">What it took to make this work</span>
  </summary>
  <div class="spoiler__body">
    <p>The WebAssembly build needed a few stubs for functions the browser lacks, and a couple of workarounds inside Verilator itself, but that&rsquo;s the dull part. What came next was more interesting.</p>
<p>If you open the page in a background tab, the browser draws nothing on it, and the screen stayed black even after you switched to that tab. I had to separately catch the moment the tab becomes visible.</p>
<p>The mouse was the most interesting of all. Oberon needs to know which of the three buttons are pressed at the same time, and the browser reports presses individually, so the state of all the buttons has to be assembled by hand. On a laptop the middle button is stood in for by a click with Alt held — not Ctrl, as one would want, because on a Mac Ctrl-click is the right button. And a Shift-click stands in for a tricky chord, where you press the left button and, without releasing it, add the right. The page first presses the left itself and then adds the right, because to Oberon it matters which button things started with.</p>
<p>The machine itself runs in a separate thread so as not to hang the page, and passes finished screen frames to the main thread without copying. It also computes only while someone is looking at it, and if the tab is hidden or you&rsquo;ve scrolled the machine off the edge of the screen, it stops. Otherwise it would honestly burn a whole core of your computer&rsquo;s processor, because the circuit doesn&rsquo;t know how to idle and simply counts cycles.</p>

  </div>
</details>

<p>Thanks to this, the machine can be dropped into any page with two lines:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-html" data-lang="html"><span class="line"><span class="cl"><span class="p">&lt;</span><span class="nt">script</span> <span class="na">type</span><span class="o">=</span><span class="s">&#34;module&#34;</span> <span class="na">src</span><span class="o">=</span><span class="s">&#34;https://tym83.github.io/paleocomputing/oberon/embed.js&#34;</span><span class="p">&gt;&lt;/</span><span class="nt">script</span><span class="p">&gt;</span>
</span></span><span class="line"><span class="cl"><span class="p">&lt;</span><span class="nt">oberon-machine</span> <span class="na">base</span><span class="o">=</span><span class="s">&#34;https://tym83.github.io/paleocomputing/oberon/&#34;</span><span class="p">&gt;&lt;/</span><span class="nt">oberon-machine</span><span class="p">&gt;</span>
</span></span></code></pre></div><p>Details and settings are on the <a href="https://tym83.github.io/paleocomputing/oberon/embed.html">Embed it</a> page. There you can also switch the machine to the processor with the new check instruction and place your own files on its disk before boot.</p>
<h2 id="the-page-where-you-change-the-processor">The page where you change the processor</h2>
<p>Originally I wanted the reader to be able to change the processor&rsquo;s instruction set themselves, right there on the page. That, I&rsquo;ll freely admit, I didn&rsquo;t do, because rebuilding the circuit from Verilog directly in the browser proved too heavy. The fallback idea from the same plan worked instead: build several processor variants in advance and let you switch between them. On the <a href="https://tym83.github.io/paleocomputing/oberon/checks.html">Change the processor</a> page there are two cores — the ordinary one, like Wirth&rsquo;s, and one with the new <code>CHK</code> instruction — and on each you can run the same loop that walks an array.</p>
<p>On the ordinary processor, one array access in this loop takes eleven cycles, two of which go on the software check. On the processor with the new instruction it&rsquo;s ten, because the check takes one cycle instead of two. One cycle out of eleven, exactly as intended. That boring line from the risk table came true.</p>
<figure><img src="checks-page.png" alt="The page where the processor is switched" width="940" height="760" loading="lazy" decoding="async">
  <figcaption><p>The page where you change the processor. A core switch and listings of the two variants of the loop.</p></figcaption>
</figure>

<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">How this measurement lied tenfold at first</span>
  </summary>
  <div class="spoiler__body">
    In the first version the page ran the program to the end, and tracked the moment it finished by peeking into memory once every two hundred thousand instructions. As a result the measurement also caught the idle tail after the loop ended, and the check appeared to cost ten instructions, not one. It looked quite plausible, and I nearly believed it. Now the loop is deliberately longer than what we measure, both processor versions execute exactly the same number of instructions, and how many times the loop managed to run is written by the program itself into a register.
  </div>
</details>

<p>And the main thing this page checks isn&rsquo;t even the cost of the check. The ordinary system, with no new instructions at all, must boot on both processors in exactly the same way, with the same image on the screen and the same number of instructions executed. If, after the new instruction is added, the old programs behave even slightly differently, then I&rsquo;ve broken compatibility and ended up with some other machine.</p>
<figure><img src="checks-measure.gif" alt="A measurement on the ordinary core, a switch to the CHK core, and another measurement" width="720" height="429" loading="lazy" decoding="async">
  <figcaption><p>A measurement on the ordinary core, a switch to the core with CHK, and one more. The difference, as promised, is a single cycle.</p></figcaption>
</figure>

<h2 id="thirteen-lab-exercises">Thirteen lab exercises</h2>
<p>Since the machine runs in the browser, you can learn on it. The <a href="https://tym83.github.io/paleocomputing/oberon/lab.html">lab</a> has thirteen exercises, and each is checked by the machine rather than taken on trust. The check looks into the memory, registers or disk of the emulated computer itself and sees whether you did what was asked. The exercises are graded by level: first just look, then change, break, measure and, finally, build your own.</p>
<figure><img src="lab-page.png" alt="The lab: machine on the left, exercise and check button on the right" width="1440" height="1000" loading="lazy" decoding="async">
  <figcaption><p>The lab. On the left the machine, on the right the exercise and a check button that looks straight into the emulated computer&rsquo;s memory.</p></figcaption>
</figure>

<p>And this is what the Oberon interface looks like in action. A middle-click on the text <code>System.ShowModules</code> opens the list of loaded modules, and another, on <code>Hilbert.Draw</code>, draws a Hilbert curve. No buttons, only text:</p>
<figure><img src="lab-showmodules-hilbert.gif" alt="A middle-click on text is how you run a program" width="720" height="540" loading="lazy" decoding="async">
  <figcaption><p>A middle-click on text — that is how you run a program.</p></figcaption>
</figure>

<table>
  <thead>
      <tr>
          <th>#</th>
          <th>Exercise</th>
          <th>Level</th>
          <th>What you learn</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>1</td>
          <td>The system on real hardware</td>
          <td>look</td>
          <td>that Wirth&rsquo;s circuit runs under the picture, and how an interface where any text can be a command is built</td>
      </tr>
      <tr>
          <td>2</td>
          <td>Your first module</td>
          <td>look</td>
          <td>how to type, save and compile a program in the built-in editor</td>
      </tr>
      <tr>
          <td>3</td>
          <td>The interface key</td>
          <td>change</td>
          <td>why Oberon has no header files, and how the system guards against incompatible modules</td>
      </tr>
      <tr>
          <td>4</td>
          <td>There is no memory protection here</td>
          <td>break</td>
          <td>what happens if you write garbage straight into screen memory, and how one instruction kills the machine outright</td>
      </tr>
      <tr>
          <td>5</td>
          <td>How many cycles per instruction</td>
          <td>measure</td>
          <td>why, on a machine with no cache, you can count the cycles in your head</td>
      </tr>
      <tr>
          <td>6</td>
          <td>Memory runs out mid-instruction</td>
          <td>break</td>
          <td>why garbage collection works only between instructions</td>
      </tr>
      <tr>
          <td>7</td>
          <td>The system rebuilds itself</td>
          <td>build</td>
          <td>how to rebuild a module inside the system and find the image&rsquo;s inconsistency with your own hands</td>
      </tr>
      <tr>
          <td>8</td>
          <td>Two generations of the compiler</td>
          <td>look</td>
          <td>why a compiler that builds itself still proves nothing</td>
      </tr>
      <tr>
          <td>9</td>
          <td>Inside the compiler</td>
          <td>change</td>
          <td>where in the compiler the processor&rsquo;s instructions are born</td>
      </tr>
      <tr>
          <td>10</td>
          <td>The garbage collector from the inside</td>
          <td>look</td>
          <td>that garbage collection is an ordinary task the system calls about once a second</td>
      </tr>
      <tr>
          <td>11</td>
          <td>One task at a time</td>
          <td>break</td>
          <td>why one hung task stops the whole system</td>
      </tr>
      <tr>
          <td>12</td>
          <td>The cost of a check, by hand</td>
          <td>measure</td>
          <td>how to measure three variants on your own loop — no check, software, and hardware</td>
      </tr>
      <tr>
          <td>13</td>
          <td>Your own built-in procedure</td>
          <td>build</td>
          <td>how to add a new built-in procedure to the language by rebuilding the compiler right inside the system</td>
      </tr>
  </tbody>
</table>
<p>By default the lab opens in English, and switches to Russian with a button or a link carrying <a href="https://tym83.github.io/paleocomputing/oberon/lab.html?lang=ru"><code>?lang=ru</code></a>, and the choice is remembered. Beside it sits an eight-chapter <a href="https://tym83.github.io/paleocomputing/oberon/book/">handbook</a> that begins with what&rsquo;s real here and ends with what we measured, and it has an <a href="https://tym83.github.io/paleocomputing/oberon/book/en/">English version</a> too. If you teach computer architecture and want to take these exercises for your own use, write to me and I&rsquo;ll help you set them up. They&rsquo;re open, and the machine can be embedded in your own page with the same tag.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">What the lab exercises showed me</span>
  </summary>
  <div class="spoiler__body">
    <p>The garbage collector in Oberon is very lazy. It runs only if the user has managed twenty actions or memory is nearly out, so a few hundred kilobytes of garbage can sit around indefinitely. And it works only between instructions because it can&rsquo;t scan the stack, and finds all live objects solely through modules&rsquo; global variables.</p>
<p>The automatic typing in one of the exercises was silently corrupting the program. Because of an error in the key table, the closing bracket wasn&rsquo;t being typed, yet the file was saved just fine, so the check, which looked only for the file&rsquo;s presence, noticed nothing.</p>
<p>And the button that rolls the machine back to the start didn&rsquo;t work for a long time, because in Wirth&rsquo;s circuit the processor&rsquo;s registers aren&rsquo;t zeroed on reset. On the first run Verilator zeroes them; on a subsequent one whatever was there remains. On a real FPGA it would be exactly the same, so this is a crutch specific to my rig.</p>

  </div>
</details>

<p>A separate embarrassment of this section is that the lab on the site was dead for a while. The exercise list was empty and an eternal spinner hung on the screen. First I found one error in the page&rsquo;s code and fixed it, but that didn&rsquo;t help. The real cause was that the page&rsquo;s file had no closing tag on its script. I&rsquo;d considered this harmless — the browser will forgive it — and had even taught the site&rsquo;s automatic check to forgive it too. But by the HTML standard, if a file ends inside an unclosed script, the browser marks that script as already executed and simply doesn&rsquo;t run it, issuing neither an error nor a warning. Now the automatic check opens the site in a real browser and requires the page to have as many exercises as the sources do. In short, I no longer say the browser will forgive anything.</p>
<h2 id="and-what-does-the-check-cost-on-modern-processors">And what does the check cost on modern processors?</h2>
<p>All right — on Wirth&rsquo;s processor the software check costs two cycles out of eleven, and the hardware one costs a single cycle. But Wirth&rsquo;s processor is a very simple machine that executes one instruction after another and guesses at nothing. How do things stand on the processors in our laptops and servers?</p>
<p>I took the same array loop, wrote it in C and in Rust in three variants. In the first there&rsquo;s no check at all; in the second it&rsquo;s written the way the language itself writes it; and in the third the compiler is forbidden to throw it away. All three variants I ran everywhere I could reach. Let me note straight away, because without it the numbers can&rsquo;t be read: the loop is deliberately chosen to be as favorable to the check as possible. The array is small and always sits in the processor&rsquo;s fastest memory, and the check always passes, so it&rsquo;s easy for the processor to predict its outcome. On top of that, I forbade the compiler to vectorize — to process several array elements with one instruction — even though in real code checks are costly primarily because they get in the way of exactly that. I wanted to see the check on its own, without everything else.</p>
<table>
  <thead>
      <tr>
          <th>Processor</th>
          <th>How much the check slows this loop</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>RISC5, software check</td>
          <td>by 22%</td>
      </tr>
      <tr>
          <td>RISC5, new <code>CHK</code> instruction</td>
          <td>by 11%</td>
      </tr>
      <tr>
          <td>Apple M4</td>
          <td>not visible, lost in the noise</td>
      </tr>
      <tr>
          <td>AMD EPYC</td>
          <td>not visible, lost in the noise</td>
      </tr>
      <tr>
          <td>Server-class Arm Neoverse N2</td>
          <td>by 2–6%</td>
      </tr>
      <tr>
          <td>CHERIoT, hardware check</td>
          <td>not one extra instruction</td>
      </tr>
  </tbody>
</table>
<p>First, the check itself hasn&rsquo;t changed at all in forty years — everywhere it&rsquo;s the same comparison and conditional jump as on Wirth&rsquo;s machine. Second, the check as the language writes it cost nothing at all anywhere, because modern C and Rust compilers worked out for themselves that the index in this loop never goes out of bounds and threw the check away. Wirth&rsquo;s compiler can&rsquo;t do that; it always leaves the check in.</p>
<p>And why is the check invisible on AMD and Apple but visible on the server Arm? This strikes me as the most interesting result of the first half of the project. A modern processor executes several instructions per cycle and reorders them itself to keep all its execution units busy. If a program has little work, the extra check instructions simply drop into the free slots and cost nothing, like a passenger who boards a half-empty bus and crowds no one. I checked this by gradually adding work and checks to the loop and watching for when the checks start to cost time. On AMD one or two checks really do cost nothing, and from the fourth they begin to. On the server Arm each check adds a little time from the very start, because that core has fewer free slots. Wirth&rsquo;s processor has no free slots at all — it executes one instruction at a time, and every check instruction is a cycle that always gets paid for.</p>
<h2 id="what-cheri-does">What CHERI does</h2>
<p>CHERI is an architecture in which a pointer knows its own bounds — that is, along with the address it stores where the region of memory it&rsquo;s allowed to reach begins and ends, and the processor checks those bounds itself on every memory access. Such a pointer is twice as wide as an ordinary address, and a program can&rsquo;t forge it. I wanted to put CHERI on the same ladder, and not on a big, complex processor where an extra instruction can cost anything at all, but on the closest relative to Wirth&rsquo;s processor. I took CHERIoT-Ibex, a simple 32-bit core that also executes instructions one at a time and has no cache, and ran the same loop on its circuit.</p>
<p>The result was very telling. On CHERI the check takes not a single instruction in the loop, because it&rsquo;s built into the memory read itself. The loop with protection is exactly the same length as the loop without it. To make sure the protection really works, I narrowed the pointer&rsquo;s bounds to 32 elements, and the program fell over on precisely the 33rd. And there is no check-free variant on CHERI, and never will be, because there&rsquo;s simply nothing to turn it off with. The published work confirms this. On CHERI the check itself is nearly free, and what you pay for is the pointer width, because wide pointers take up twice the room in memory and in cache.</p>
<p>And so a ladder emerged. On Wirth&rsquo;s processor the software check costs two instructions and two cycles, the new <code>CHK</code> instruction costs one instruction and one cycle, and CHERI costs zero instructions. The wide modern processors stand off to the side: they have as many instructions as Wirth&rsquo;s, but as long as the core has free slots those instructions cost almost nothing. And between <code>CHK</code> and CHERI a gap yawned on this ladder for a long time — one I closed only at the very end of these nine days.</p>
<h2 id="wirths-machine-in-qemu">Wirth&rsquo;s machine in QEMU</h2>
<p>The browser is wonderful, but I wanted Wirth&rsquo;s machine to live where ordinary virtual machines live, with its own console, disks, restarts and all the trappings of a proper hypervisor. Almost every VM on Linux is launched, one way or another, by QEMU, a program that can impersonate computers of the most varied architectures. Wirth&rsquo;s processor, of course, isn&rsquo;t among them, so it had to be written from scratch — that is, QEMU had to be taught to understand all of RISC5&rsquo;s instructions and its idiosyncratic fractional arithmetic, and then a board had to be assembled around the processor, with memory, a disk, a keyboard, a mouse and a screen.</p>
<p>We had one advantage here that almost nobody writing a new architecture for QEMU has. Usually the correctness of such work is checked against documentation and test suites, whereas we had the processor&rsquo;s own circuit and a reference emulator already checked against it over fifteen million instructions. So we checked QEMU not against paper but against the circuit, instruction by instruction, by that same lockstep technique. Across one and a half million instructions of the boot, no divergence was found; the image on the screen matched the circuit down to the last dot; and the list of modules that opens on a click in <code>System.ShowModules</code> was exactly the same, with the same addresses in memory. The number of dark dots on the screen after boot — 18,607 — became my reference from then on, and wherever the machine ran, that is what I compared the screen against.</p>
<figure><img src="qemu-oberon-showmodules.png" alt="A middle-click on System.ShowModules in QEMU" width="1024" height="768" loading="lazy" decoding="async">
  <figcaption><p>A middle-click on System.ShowModules in QEMU. The same modules at the same addresses as on Wirth&rsquo;s circuit.</p></figcaption>
</figure>

<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">A rule that&#39;s in no description of the processor</span>
  </summary>
  <div class="spoiler__body">
    <p>The first divergence in the lockstep comparison happened very early, on the two hundred and twenty-fourth instruction, and the bootloader hadn&rsquo;t even reached its first disk read. Wirth&rsquo;s processor, it emerged, updates the flags showing whether a result is negative and whether it&rsquo;s zero on any write to a register, including when it merely loads a number from memory. This isn&rsquo;t in the processor&rsquo;s description, but it is in the circuit, and Wirth&rsquo;s bootloader relies on it: right after loading a number from memory, it checks whether it&rsquo;s zero without doing a separate comparison. It&rsquo;s for things like this that you check against the circuit and not the documentation.</p>
<p>The fractional arithmetic I did myself too, rather than taking QEMU&rsquo;s ready-made version, because Wirth&rsquo;s arithmetic isn&rsquo;t standard. It rounds differently, handles very small numbers differently, and so on, and a standard implementation would give plausible but different results. The first comparison showed a couple of dozen divergences, and I was already about to fix my code when the comparison itself proved wrong, and the arithmetic had matched on the first try. Had I trusted it, I&rsquo;d have broken working code.</p>

  </div>
</details>

<p>You can build and run it like this (in detail, in the <a href="https://github.com/tym83/paleocomputing/blob/main/qemu/GUIDE.md">QEMU guide</a>):</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-sh" data-lang="sh"><span class="line"><span class="cl">git clone https://github.com/tym83/paleocomputing <span class="o">&amp;&amp;</span> <span class="nb">cd</span> paleocomputing
</span></span><span class="line"><span class="cl">make -C qemu build        <span class="c1"># builds QEMU in a Docker container</span>
</span></span><span class="line"><span class="cl">docker create --name oberon-payload ghcr.io/tym83/paleocomputing/oberon-run:v0.1.17
</span></span><span class="line"><span class="cl">docker cp oberon-payload:/opt/oberon/payload/prom.bin .
</span></span><span class="line"><span class="cl">docker cp oberon-payload:/opt/oberon/payload/oberon.dsk .
</span></span><span class="line"><span class="cl">docker rm oberon-payload
</span></span><span class="line"><span class="cl">.qemu-work/build/qemu-system-risc5 -machine oberon -bios prom.bin <span class="se">\
</span></span></span><span class="line"><span class="cl">  -drive <span class="k">if</span><span class="o">=</span>none,id<span class="o">=</span>sd0,file<span class="o">=</span>oberon.dsk,format<span class="o">=</span>raw -vnc :0
</span></span></code></pre></div><p>The first command builds QEMU with our processor; the next three pull the bootloader and the system disk out of a ready-made image; and the last one starts the machine. Watch the screen with any VNC client at <code>127.0.0.1:5900</code>, and if you add <code>chk=on</code> to <code>-machine oberon</code>, you get the processor with the hardware check. Booting without hardware acceleration takes up to a minute, so don&rsquo;t be alarmed if at first you see only the bootloader screen.</p>
<p>Why this is a separate build rather than a patch to mainline QEMU is also worth explaining. The QEMU and libvirt projects don&rsquo;t accept code that a language model had a hand in writing — even where it&rsquo;s only suspected. I have a public repository that says honestly how it was made, so this can go upstream only if someone rewrites the code by hand. The GPL, meanwhile, expressly permits your own build, and the hardest part of the work — an exact description of the processor&rsquo;s behavior and a way to check any implementation of it — remains useful regardless. One further nuisance is that our processor is written against the very newest development version of QEMU, so already-released versions won&rsquo;t build it.</p>
<h2 id="libvirt-which-wont-take-your-word-for-it">libvirt, which won&rsquo;t take your word for it</h2>
<p>The next layer is libvirt. It&rsquo;s the library through which almost everything that manages VMs on Linux talks to QEMU, from the virsh command-line tool to KubeVirt, which is coming up next. The question was whether a new architecture could be plugged in with no edits at all. It couldn&rsquo;t — but not for the reason I expected.</p>
<p>First I simply put the architecture <code>risc5</code> in the machine description, and libvirt refused right away, saying it didn&rsquo;t know such an architecture. So I got clever, called myself an architecture it did know, and slipped it my own program. libvirt refused again, but this time more deeply. It doesn&rsquo;t trust what&rsquo;s written in the machine description; it asks QEMU itself what it can impersonate, QEMU honestly calls itself risc5, and there&rsquo;s no such name in libvirt&rsquo;s list. Deceiving it through the machine description is impossible.</p>
<p>So there&rsquo;s no getting around a libvirt edit, and the edit was tiny — about ten lines in five places. I made it so that the list of architectures is taken from a separate file, and the next old machine would be a new line in that file rather than new code. The build checks each edit separately, because an edit that silently failed to apply is worse than none: libvirt would build without errors but not know the architecture. This came in handy at once, when in a new version of libvirt the relevant piece of code moved to another file, and the build announced it loudly instead of quietly building a broken library.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">How libvirt crashed on an honest answer</span>
  </summary>
  <div class="spoiler__body">
    After the edit libvirt recognized the architecture, but on its very first query to our QEMU it crashed on a null-pointer dereference. It asks the emulator for a list of supported CPU models, and our QEMU honestly answered that it has no models, because Wirth&rsquo;s machine has exactly one processor. libvirt simply isn&rsquo;t built for such an answer. One option was to teach libvirt to survive such a refusal, which is more correct in substance; another was to declare a single lone CPU model in QEMU. I chose the second, because it&rsquo;s fewer edits to someone else&rsquo;s code.
  </div>
</details>

<p>And here at once is a fine story about why you can&rsquo;t trust your eyes. The first screenshot from under libvirt looked entirely correct, but a byte-by-byte comparison with the reference found several thousand differences. I managed to suspect the mouse, the moment the screenshot was taken, and libvirt itself — and I&rsquo;d simply miscomputed the address of video memory from hexadecimal into decimal and missed by 512 bytes. That&rsquo;s exactly four rows of the screen, and the difference is completely invisible to the eye. After the fix, the screen matched the reference byte for byte.</p>
<figure><img src="libvirt-oberon-screen.png" alt="Oberon under libvirt" width="1024" height="768" loading="lazy" decoding="async">
  <figcaption><p>Oberon under libvirt. It looks exactly the same as it did with the 512-byte shift, which is why I no longer trust my eyes.</p></figcaption>
</figure>

<h2 id="kubevirt-without-a-fork">KubeVirt without a fork</h2>
<p>One layer up sits KubeVirt. It lets you run virtual machines in Kubernetes the same way you run ordinary containers, and manage them with the same tools. Inside, each such VM has a housekeeping pod, and it&rsquo;s there that libvirt and QEMU live. I needed to slip a machine of an entirely different architecture in there, and to do it without making my own copy of either KubeVirt or Cozystack. A private copy of a large project stays with you forever, and you have to drag it by hand through every update, so I very much wanted to avoid that.</p>
<p>A standard extension point that few people know about came to the rescue. Before starting a VM, KubeVirt can hand its description to an external handler and then run whatever the handler returns. The handler can be an ordinary script sitting in the Kubernetes configuration, so you don&rsquo;t even need to build a separate image for it. Our handler receives the description of a perfectly ordinary VM — with an Intel processor, disks and networking — and reshapes it into Wirth&rsquo;s machine. It changes the architecture, turns off hardware acceleration (which, for a foreign architecture, doesn&rsquo;t exist anyway), substitutes our emulator, bootloader and disk image, and throws out everything Wirth&rsquo;s machine doesn&rsquo;t have — which is almost everything KubeVirt adds by default.</p>
<p>The emulator itself and the patched libvirt do have to be placed into the housekeeping image that KubeVirt starts all its VMs from. Everything else KubeVirt and Cozystack handle themselves. So for Wirth&rsquo;s machine the cluster needs exactly two things an ordinary user can&rsquo;t do: allow external handlers, and install our housekeeping image. The administrator agrees to this once, and after that any user installs the machine from the catalog, much as you&rsquo;d install a driver package.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">The refusals that only show up on real KubeVirt</span>
  </summary>
  <div class="spoiler__body">
    <p>First I ran the description KubeVirt gives an ordinary VM through the handler and fed the result to libvirt on my own desk. It all started up, the screen matched the reference, and I decided it was done. On real KubeVirt inside our Cozystack the machine wouldn&rsquo;t start, and there were four refusals, not one of which had appeared on the desk.</p>
<p>First QEMU refused, because KubeVirt requires the machine to support power management, which Wirth&rsquo;s machine doesn&rsquo;t have. Then it refused because KubeVirt asks to be allowed to add processors on the fly, and Wirth&rsquo;s machine has one processor and never any more. The third refusal was the funniest. I removed the section with system information from the description, and now KubeVirt itself fell over, because it reads that section after startup. I put the section back, and QEMU refused again, because it doesn&rsquo;t support such information for this architecture. In the end the section had to stay, and one small setting had to be removed — the one that makes libvirt pass it to QEMU. Since then I remove exactly what&rsquo;s refused and not a line more, because a broad sweep of the broom breaks what was working. And the fourth refusal came when I updated KubeVirt while our housekeeping image had been built for the previous version. Now the image is built separately for each supported version of KubeVirt.</p>

  </div>
</details>

<p>There was one more quiet trap with the housekeeping image. We take KubeVirt&rsquo;s standard image and replace only libvirt in it, and the version of our libvirt has to match, exactly, the one already in the image. Get it wrong and the image builds without a single error, but our library ends up beside the standard one rather than in place of it, and it&rsquo;s the standard one that runs. So now the build checks that there&rsquo;s exactly one libvirt in the image, and the correspondence between KubeVirt and libvirt versions is recorded in a single file and can&rsquo;t be set by hand.</p>
<h2 id="sixteen-seconds-to-a-cluster-wide-migration">Sixteen seconds to a cluster-wide migration</h2>
<p>This is probably the most instructive story of the project, and it&rsquo;s about my own inattention; the code has nothing to do with it.</p>
<p>The housekeeping image KubeVirt starts VMs from is set for the entire cluster at once. To swap it out, I edited the KubeVirt configuration right on the live cluster — our working rig, on which, among other things, other people&rsquo;s VMs were running. Sixteen seconds later KubeVirt began migrating every virtual machine in the cluster onto the new image, without stopping them. The thing is, Cozystack has automatic updates of running machines enabled by default, and KubeVirt did exactly what it was told: since the housekeeping image had changed, all the machines needed moving to the new one. Over five hours it attempted to migrate the machines more than a hundred times, and more than half the attempts failed. Nothing crashed and no data was lost, and the successful migrations incidentally showed that our image can migrate machines — but the picture was not a pretty one. Here&rsquo;s how I stopped it:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-sh" data-lang="sh"><span class="line"><span class="cl">kubectl -n cozy-kubevirt patch kubevirt kubevirt --type<span class="o">=</span>merge <span class="se">\
</span></span></span><span class="line"><span class="cl">  -p <span class="s1">&#39;{&#34;spec&#34;:{&#34;workloadUpdateStrategy&#34;:{&#34;workloadUpdateMethods&#34;:[]}}}&#39;</span>
</span></span></code></pre></div><p>This command turns off automatic updates of machines, but it&rsquo;s an emergency brake, not a solution. It doesn&rsquo;t cancel migrations already under way, the setting will come back at the next Cozystack update, and going entirely without auto-updates is also bad, because after a KubeVirt update the machines would stay on the old housekeeping image. The most galling part of this story is that I knew where to look. I simply checked that the new image existed, and didn&rsquo;t think to check what it would do to the cluster.</p>
<p>Why half the migrations failed became clear from the logs, and the image had nothing to do with it. By default KubeVirt migrates no more than two machines at a time off a single server; the rest wait their turn and often don&rsquo;t get one. Quotas got in the way too. During a migration a machine exists in two copies at once and takes up twice the memory, so a machine that has eaten its user&rsquo;s entire quota will never migrate live. The quota needs headroom, at least enough for the largest machine.</p>
<h2 id="a-catalog-for-cozystack">A catalog for Cozystack</h2>
<p>Cozystack is an open platform from which you assemble your own cloud on top of Kubernetes. We at Ænix build it together with the community, and it&rsquo;s part of the CNCF, the foundation where Kubernetes itself lives. The platform has a web interface, users — here called tenants (like separate accounts in a cloud) — virtual machines, and an application catalog holding databases, Kubernetes clusters, caches and the rest. And recently a mechanism for pluggable catalogs appeared in the community. Anyone can publish their own set of applications and plug it into their platform, and then those applications show up for users alongside the built-in ones. For paleocomputing I made exactly such a catalog, inventing nothing beyond what&rsquo;s already in the platform.</p>
<p>The catalog is split into four parts, because they have to be trusted differently. The first holds the machines and environments for users — the real VM with Wirth&rsquo;s processor, the browser lab, the handbook, and a bundle that installs all of it at once. The second holds the Oberon language environment, in which you can run Oberon programs as ordinary jobs in the cluster. The third places boot images into the platform&rsquo;s shared storage, and the fourth swaps out that KubeVirt housekeeping image. These last two touch the whole cluster, and so are installed only with the administrator&rsquo;s explicit consent.</p>
<p>The catalog is plugged in like this:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-sh" data-lang="sh"><span class="line"><span class="cl">cozypkg tap oci://ghcr.io/tym83/paleocomputing/machines:v0.1.17
</span></span><span class="line"><span class="cl">cozypkg tap oci://ghcr.io/tym83/paleocomputing/languages:v0.1.17
</span></span><span class="line"><span class="cl">cozypkg add paleocomputing.machines
</span></span><span class="line"><span class="cl">cozypkg add paleocomputing.languages
</span></span><span class="line"><span class="cl"><span class="c1"># what touches the whole cluster is installed only with the administrator&#39;s explicit consent</span>
</span></span><span class="line"><span class="cl">cozypkg tap oci://ghcr.io/tym83/paleocomputing/images:v0.1.17
</span></span><span class="line"><span class="cl">cozypkg tap oci://ghcr.io/tym83/paleocomputing/platform:v0.1.17
</span></span><span class="line"><span class="cl">cozypkg add paleocomputing.platform --allow-privileged
</span></span></code></pre></div><p>The <code>tap</code> command only plugs the catalog into the platform; <code>add</code> installs its contents. That distinction once confused me too, and at first it wasn&rsquo;t in the documentation at all.</p>
<p>Before installing the last part, be sure to look at the automatic-machine-update setting in KubeVirt. With Cozystack&rsquo;s default settings, changing the housekeeping image will send every VM in the cluster off on a migration, and this happens on install, on uninstall, and at every KubeVirt update. Exactly what happened to me. This hole was found by one of the reviewers while checking this very article, so now, if auto-update is on, the component changes nothing and waits until the administrator explicitly permits the migration. You can permit it with the setting <code>allowWorkloadUpdate: true</code> at install time, or with the annotation <code>paleocomputing.io/allow-workload-update=true</code> on the KubeVirt resource. And one more thing. The catalog-plugging utility doesn&rsquo;t yet verify the digital signature, so if you&rsquo;re installing this somewhere more serious than a home rig, check the signature yourself; the instructions say how.</p>
<p>After it&rsquo;s plugged in, a Paleocomputing section of its own appears in users&rsquo; catalogs. There&rsquo;s a wrinkle, though, which I hit while taking the screenshots for this article. The current Cozystack web interface shows only three sections in its sidebar — the ones hard-wired into its code — and hides all the rest at the very end of the full list of applications. At first I couldn&rsquo;t find my own section either. And yet the whole point of a pluggable catalog is to bring your own sections, not to dissolve into someone else&rsquo;s. So we&rsquo;ll be fixing this in Cozystack itself, and the interface will have to show any sections a catalog brings.</p>
<figure><img src="cozystack-catalog.jpg" alt="The Paleocomputing section in the full application list" width="1456" height="827" loading="lazy" decoding="async">
  <figcaption><p>The Paleocomputing section in the full application list. It isn&rsquo;t in the sidebar yet; we&rsquo;re fixing that in Cozystack.</p></figcaption>
</figure>

<p>The machine is installed with a form in the web interface or with one short description:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">apiVersion</span><span class="p">:</span><span class="w"> </span><span class="l">apps.cozystack.io/v1alpha1</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">kind</span><span class="p">:</span><span class="w"> </span><span class="l">OberonVM</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">metadata</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">wirth</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">spec</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">memory</span><span class="p">:</span><span class="w"> </span><span class="l">128Mi</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">hardware</span><span class="p">:</span><span class="w"> </span><span class="l">chk    </span><span class="w"> </span><span class="c"># or base, the ordinary Wirth processor</span><span class="w">
</span></span></span></code></pre></div><figure><img src="cozystack-oberonvm-form.jpg" alt="The OberonVM form: memory, processor variant, disk size" width="1456" height="827" loading="lazy" decoding="async">
  <figcaption><p>The OberonVM form: memory, processor variant and disk size. That&rsquo;s the whole machine.</p></figcaption>
</figure>

<figure><img src="cozystack-oberonvm-card.jpg" alt="The running machine in the interface, with status and what runs under it" width="1456" height="827" loading="lazy" decoding="async">
  <figcaption><p>The finished machine in the interface, with its status and a list of what&rsquo;s running under it.</p></figcaption>
</figure>

<p>There&rsquo;s no screen for the machine in the web interface yet; in the current versions only ordinary VMs have one. You can look at it with the <code>virtctl</code> utility, with the user&rsquo;s own privileges and any VNC client:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-sh" data-lang="sh"><span class="line"><span class="cl">virtctl -n tenant-sandbox vnc oberon-vm-oberon-vm-habr
</span></span></code></pre></div><figure><img src="cozystack-habr-vnc.png" alt="Wirth&#39;s machine in the cloud, captured with an ordinary user&#39;s privileges" width="1024" height="768" loading="lazy" decoding="async">
  <figcaption><p>Wirth&rsquo;s machine in the cloud, captured with an ordinary user&rsquo;s privileges. The very same 18,607 dark dots.</p></figcaption>
</figure>

<p>It&rsquo;s arranged so that the next old machine won&rsquo;t require new code. Everything that sets Wirth&rsquo;s machine apart from the others is recorded in one small file, which I call the machine&rsquo;s passport: the architecture, the emulator, the bootloader, the disk, the processor variants and the memory limits. Everything else is shared, and the handler for KubeVirt is one and the same for all machines — it simply reads the passport. The next machine, Lilith for instance, is a new passport, not a copy of the code. To prove this isn&rsquo;t empty talk, the tests include a second, fictional machine of a different architecture, and it builds without a single edit to the code.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">What a machine&#39;s passport looks like</span>
  </summary>
  <div class="spoiler__body">
    <div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">kind</span><span class="p">:</span><span class="w"> </span><span class="l">Machine</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">oberon</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">image</span><span class="p">:</span><span class="w"> </span><span class="l">ghcr.io/tym83/paleocomputing/oberon-run:v0.1.17</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">domain</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">arch</span><span class="p">:</span><span class="w"> </span><span class="l">risc5</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">machine</span><span class="p">:</span><span class="w"> </span><span class="l">oberon</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">emulator</span><span class="p">:</span><span class="w"> </span><span class="l">/usr/local/bin/qemu-system-risc5</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">vcpus</span><span class="p">:</span><span class="w"> </span><span class="m">1</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">graphics</span><span class="p">:</span><span class="w"> </span><span class="l">vnc</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">terminationGracePeriodSeconds</span><span class="p">:</span><span class="w"> </span><span class="m">0</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">payload</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">path</span><span class="p">:</span><span class="w"> </span><span class="l">/payload</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">files</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span>- {<span class="nt">name: prom.bin,   role: firmware, qemu</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="s2">&#34;-bios&#34;</span><span class="p">,</span><span class="w"> </span><span class="s2">&#34;{path}&#34;</span><span class="p">]</span>}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span>- {<span class="nt">name: oberon.dsk, role: disk,     qemu</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="s2">&#34;-drive&#34;</span><span class="p">,</span><span class="w"> </span><span class="s2">&#34;if=none,id=sd0,file={path},format=raw&#34;</span><span class="p">]</span>}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">variants</span><span class="p">:</span><span class="w"> </span>{<span class="nt">base</span><span class="p">:</span><span class="w"> </span>{<span class="nt">}, chk</span><span class="p">:</span><span class="w"> </span>{<span class="nt">chk</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;on&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">memory</span><span class="p">:</span><span class="w"> </span>{<span class="nt">default: 128Mi, min: 128Mi, max</span><span class="p">:</span><span class="w"> </span><span class="l">1Gi}</span><span class="w">
</span></span></span></code></pre></div><p>The time allowed for a graceful shutdown is set to zero here, because Wirth&rsquo;s machine can&rsquo;t hear a request to shut down, and KubeVirt would wait half a minute on it in vain at every restart. The bootloader is replaced with a new one at every catalog update, while the disk is placed once and thereafter belongs to the user, so its files survive updates.</p>
  </div>
</details>

<h2 id="everything-green-and-nothing-works">Everything green, and nothing works</h2>
<p>While I was fussing with the cluster, the same trap recurred six times in one evening, and I started a note under exactly that title. Each time a check reported success while in fact nothing worked. The catalog is built, but the main file isn&rsquo;t inside it. The install succeeded, but the machine isn&rsquo;t running. The server answers that all is well, but the image doesn&rsquo;t contain the machine itself. The build is green, but the file being copied doesn&rsquo;t exist. Then there was a seventh time — those very sixteen seconds.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">What only shows up on a live cluster</span>
  </summary>
  <div class="spoiler__body">
    <p>The browser-lab image was published for a long time without the machine itself. The page was there, the processor and disk weren&rsquo;t, and the check made do with the page being served. Along the way I discovered that, because of one line in the list of ignored files and the fact that the Mac&rsquo;s filesystem doesn&rsquo;t distinguish upper- and lower-case, an entire forty-one-file language library never made it into the repository.</p>
<p>The machine with the hardware-check processor started up, but was actually running on the ordinary one. The package for Cozystack, it emerged, contained its own copy of the handler, and I&rsquo;d been editing another. Now the automatic check compares the copies.</p>
<p>The machine&rsquo;s screen in Kubernetes was blank at first, because the handler had thrown out the image output along with everything else superfluous. When I put it back, QEMU stopped starting, because the keyboard layouts hadn&rsquo;t been placed in the image.</p>
<p>One housekeeping task hung with no errors at all. In Cozystack, only pods with a special label may reach the Kubernetes control-plane server, and the network layer silently doesn&rsquo;t answer the rest. Had it been a lack of permissions, the refusal would have come immediately, so the hang itself was the clue.</p>
<p>And with the digital signature, none of my releases matched what the community catalog expects, because the catalog expects a signature from a build off the main branch, and I was releasing by tag. This was found by reading the sources of the utility, which, as I discovered, doesn&rsquo;t verify the signature at all.</p>

  </div>
</details>

<p>On the evening of 27 September there was an episode I&rsquo;m still a little ashamed of. We were reworking how the machine is described in the cluster, and the releases came in a queue. The first stopped at its own checks. The second hung dead, because the task that prepares the machine&rsquo;s disk was waiting for the machine to start, and the machine was waiting for the disk. The third found two more bugs. At that point I lost my patience and asked the agent why we were shipping release after release, and couldn&rsquo;t it just check things properly the first time (in the original this was put more forcefully).</p>
<p>It could, and after that the release process changed. Now every change is first built as a test version, a separate sandbox in the cluster switches to it, and a scenario, acting as an ordinary user, installs the machine, checks that it started on the right processor, compares the screen against the reference, restarts it, deletes it, and confirms that nothing was left behind. Only if all of that passes does the change reach the main branch and become a release. A second such scenario checks the component that changes the housekeeping image, and, for good measure, runs an ordinary Ubuntu on our image to confirm we haven&rsquo;t broken the neighbors. The first release to pass all of this before publication was v0.1.14, and the stuck machines came up by themselves afterwards.</p>
<p>Amusingly, the scenario itself had bugs of its own, even while the system was working. For example, Ubuntu&rsquo;s boot was at first verified via a helper program inside the guest, which simply isn&rsquo;t in a clean Ubuntu image, so the check would have waited forever. Then it was verified via the login prompt on the console, and a console without a real terminal stays silent. And the wait function counted only the pauses, not the total elapsed time, so twenty minutes of waiting turned into two hours. The script waited very patiently by a long-since-booted Ubuntu.</p>
<h2 id="the-component-that-watches-the-housekeeping-image">The component that watches the housekeeping image</h2>
<p>Swapping out KubeVirt&rsquo;s housekeeping image by hand is a bad idea, for three reasons. It&rsquo;s shared by all the cluster&rsquo;s VMs, so a mistake breaks everyone. It has to match the KubeVirt version, and after a Cozystack update it will stay old and break all the machines. And my first manual method incidentally froze several other KubeVirt settings I had no intention of touching.</p>
<p>So the catalog has a separate component that looks at the KubeVirt configuration every half a minute and puts its own part in order. If we have an image for the current KubeVirt version, it installs it. If not, it removes its edit, and the cluster returns to the standard image. Wirth&rsquo;s machines then won&rsquo;t start, but everything else will work. If anything at all looks questionable, it likewise removes its edit, and if it couldn&rsquo;t read the configuration, it touches nothing. It makes its edit in such a way that if KubeVirt ever changes the shape of its configuration, the edit won&rsquo;t slot the image into the wrong place but will loudly refuse to apply.</p>
<p>And this component the sandbox caught too. On the live cluster it crashed, because it passed one of the utilities information about every server in the cluster at once, and on a real cluster that information is enormous. In the tests the servers were tiny and the test cluster consisted of a single one, so nothing crashed there. Now a test with a cluster of three thousand large servers turns red on the old code.</p>
<h2 id="servers-on-intel-and-on-arm">Servers on Intel and on Arm</h2>
<p>When I started building all this for Arm servers too, I was asked why builds for different processors are needed at all if we&rsquo;ve already added the Oberon architecture to KubeVirt. It&rsquo;s a good question, and the confusion is natural, because there are two architectures here. One is the architecture of Wirth&rsquo;s machine, which QEMU impersonates, and it doesn&rsquo;t depend on the server. The other is the server&rsquo;s real processor, Intel or Arm. The emulator and libvirt are ordinary programs, and they have to be built separately for each server processor. It&rsquo;s like an emulator for a games console, where the console is one thing but the emulator&rsquo;s builds for Windows and for Mac are different.</p>
<p>Our housekeeping image was built only for Intel. And since it replaces the standard image for every VM in the cluster, on a cluster of Arm servers not a single machine would start — not just Oberon. The Arm build went through at once, but the actual launch on Arm stopped three times. First, the image with the bootloader and disk existed only for Intel. Then, that one of the build tools is also released only for Intel. And the third problem was the most interesting. On Arm, KubeVirt always gives a VM modern firmware with its own flash memory, and Wirth&rsquo;s machine has no flash memory at all, so QEMU exited immediately after starting. On Intel this isn&rsquo;t visible, because the default firmware there is different. Not one of these three problems was visible in the code or in the built image — only on a live server with the right processor.</p>
<p>Now all of this works on two versions of KubeVirt and on both processor types, and the screen matches the reference in all four combinations. In fairness I&rsquo;ll note that on Arm I tested bare KubeVirt in a throwaway test cluster, and I haven&rsquo;t yet run Cozystack on Arm servers.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">How to install the machine without Cozystack, on your own KubeVirt</span>
  </summary>
  <div class="spoiler__body">
    <p>All of this works on ordinary KubeVirt too, without Cozystack, and the automatic check proves it on every change. It brings up a throwaway cluster, installs KubeVirt, swaps out the housekeeping image, starts the machine and compares the screen. Detailed instructions are in <a href="https://github.com/tym83/paleocomputing/blob/main/kubevirt/GUIDE.md">kubevirt/GUIDE.md</a>; in short, there are three steps:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-sh" data-lang="sh"><span class="line"><span class="cl"><span class="c1"># 1. allow external handlers (on servers without /dev/kvm you also need useEmulation: true)</span>
</span></span><span class="line"><span class="cl">kubectl -n kubevirt get kubevirt kubevirt -o json <span class="se">\
</span></span></span><span class="line"><span class="cl">  <span class="p">|</span> jq <span class="s1">&#39;.spec.configuration.developerConfiguration.featureGates |= ((. // []) + [&#34;Sidecar&#34;] | unique)&#39;</span> <span class="se">\
</span></span></span><span class="line"><span class="cl">  <span class="p">|</span> kubectl replace -f -
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># 2. install the housekeeping image for your KubeVirt version</span>
</span></span><span class="line"><span class="cl">git clone --depth <span class="m">1</span> -b v0.1.17 https://github.com/tym83/paleocomputing <span class="o">&amp;&amp;</span> <span class="nb">cd</span> paleocomputing
</span></span><span class="line"><span class="cl"><span class="nv">KV</span><span class="o">=</span><span class="k">$(</span>kubectl -n kubevirt get kubevirt kubevirt -o <span class="nv">jsonpath</span><span class="o">=</span><span class="s1">&#39;{.status.observedKubeVirtVersion}&#39;</span><span class="k">)</span>
</span></span><span class="line"><span class="cl"><span class="nb">echo</span> <span class="s2">&#34;</span><span class="nv">$KV</span><span class="s2"> ghcr.io/tym83/paleocomputing/virt-launcher:</span><span class="nv">$KV</span><span class="s2">-paleo-v0.1.17&#34;</span> &gt; /tmp/launchers.txt
</span></span><span class="line"><span class="cl">kubectl -n kubevirt create configmap kubevirt-paleo-launcher-status
</span></span><span class="line"><span class="cl"><span class="nb">export</span> <span class="nv">KUBECTL</span><span class="o">=</span>kubectl <span class="nv">KUBEVIRT_NAMESPACE</span><span class="o">=</span>kubevirt <span class="nv">LAUNCHER_TABLE</span><span class="o">=</span>/tmp/launchers.txt
</span></span><span class="line"><span class="cl"><span class="nv">R</span><span class="o">=</span>marketplace/repos/platform/packages/system/kubevirt-paleo-launcher/files/reconcile.sh
</span></span><span class="line"><span class="cl"><span class="k">until</span> sh <span class="nv">$R</span> once <span class="o">&amp;&amp;</span> <span class="o">[</span> <span class="s2">&#34;</span><span class="k">$(</span>kubectl -n kubevirt get cm kubevirt-paleo-launcher-status <span class="se">\
</span></span></span><span class="line"><span class="cl">  -o <span class="nv">jsonpath</span><span class="o">=</span><span class="s1">&#39;{.data.state}&#39;</span><span class="k">)</span><span class="s2">&#34;</span> <span class="o">=</span> Applied <span class="o">]</span><span class="p">;</span> <span class="k">do</span> sleep 10<span class="p">;</span> <span class="k">done</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># 3. install the machine with ordinary Helm and open its screen</span>
</span></span><span class="line"><span class="cl">python3 marketplace/tools/pin-images.py --release v0.1.17
</span></span><span class="line"><span class="cl">helm install wirth marketplace/repos/machines/packages/apps/oberon-vm <span class="se">\
</span></span></span><span class="line"><span class="cl">  -n oberon --create-namespace --set <span class="nv">storageClass</span><span class="o">=</span>&lt;your StorageClass&gt;
</span></span><span class="line"><span class="cl">virtctl -n oberon vnc oberon-vm-wirth
</span></span></code></pre></div><p>Bear in mind that allowing external handlers takes effect across the whole cluster, and anyone who can create VMs directly will be able to slip their own handler into them. On a shared cluster this is worth weighing. And the automatic machine updates from the story above are worth remembering here too.</p>

  </div>
</details>

<h2 id="the-second-episode-a-language-model-on-wirths-processor">The second episode: a language model on Wirth&rsquo;s processor</h2>
<p>Remember the claim from that very first plan, that a new instruction which multiplies and adds in one go would double the speed of a neural network on Wirth&rsquo;s processor? The reviewers demolished it, but a line with their estimates stayed in my list of deferred tasks afterwards. The new instruction, by their reckoning, would give a few percent; a fast multiplier, more than one and a half times. It sounded like a result, but behind those numbers there was neither a program nor a way to reproduce them — only arithmetic off the cycle table. I&rsquo;d wanted to check it by hand from the start.</p>
<p>The task boiled down to this. A small language model was to run inside the Oberon system itself. A program in Oberon, compiled by the system&rsquo;s own compiler, reads the model&rsquo;s weights from a file and prints text, and all of it executes on Wirth&rsquo;s real circuit, cycle by cycle. And then I&rsquo;d need to work out where the time goes and check both of the reviewers&rsquo; estimates.</p>
<p>Wirth&rsquo;s machine hasn&rsquo;t much memory — one megabyte for everything, screen included — so the model came out truly tiny. It predicts the next letter from the previous eight, and has about forty-three thousand parameters. For comparison, the models we chat with have millions of times more. Yet even a model like this takes up almost half of the machine&rsquo;s free memory. I trained it on an ordinary laptop in a few dozen seconds, on the text of <em>Alice&rsquo;s Adventures in Wonderland</em>, long since in the public domain.</p>
<details class="spoiler spoiler--default">
  <summary class="spoiler__summary">
    <span class="spoiler__icon" aria-hidden="true">▸</span>
    <span class="spoiler__title">Why not a transformer, and why not integers</span>
  </summary>
  <div class="spoiler__body">
    <p>Modern models are built as transformers, and one could write such a model too, but it does the same underlying arithmetic and, on top of that, demands several hundred more lines of intricate maths on Wirth&rsquo;s non-standard fractional numbers. For the question of what a multiplication costs, that would add nothing.</p>
<p>And converting the model to integers, which usually speeds up computation on weak hardware, would give nothing on Wirth&rsquo;s processor. On Wirth&rsquo;s machine, integer multiplication is even slower than fractional multiplication, so integers would only add work. This was settled by arithmetic, not fashion.</p>

  </div>
</details>

<p>Here&rsquo;s what the model writes if you start it with <code>alice was</code>:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">alice was one thought all the tell you spo
</span></span></code></pre></div><p>Shakespeare can rest easy. But what interests us here isn&rsquo;t the literature, it&rsquo;s the stopwatch.</p>
<p>The main requirement was that the text the model prints on Wirth&rsquo;s circuit match, byte for byte, a reference computed on an ordinary computer. For that, the reference had to be computed not with the usual Python facilities but in exactly the same arithmetic as Wirth&rsquo;s processor, repeating every operation in the same order. Otherwise nothing would have matched. I separately worked out how far Wirth&rsquo;s arithmetic diverges from the standard, and found that almost all the intermediate numbers differ in their last digits. The generated text never once diverged, mind you, but that&rsquo;s simple luck. A comparison against ordinary Python would have been right almost always, and one day inexplicably wrong — the worst kind of bug, because it can be neither reproduced nor explained.</p>
<p>The text matched everywhere. On the emulator, where this check now runs automatically on every change; on Wirth&rsquo;s circuit; on the circuit with the fast multiplier, of which more below; and inside the real system with windows. The program and the weights are placed on the disk, two middle-clicks first compile the program and then start the generation, and the finished text is read off the disk and matches the reference. In QEMU, the same.</p>
<h2 id="where-the-time-goes">Where the time goes</h2>
<p>On Wirth&rsquo;s processor with its native multiplier, the model prints about nine letters a second. On a live machine, where the video controller takes a little of the processor&rsquo;s time, slightly fewer. In short, nine letters a second, with luck.</p>
<p>When I looked at what the processor was busy with all that time, the picture was very clear. Almost forty percent of all cycles go on fractional multiplication, and most of that time is simply waiting for the slow multiplier to finish counting. Wirth&rsquo;s is serial and computes the product one bit per cycle, so a single multiplication takes twenty-six cycles. Another quarter of the time goes on reading from memory. And here another interesting thing emerged. The simplest line of the program, where a product is added to a sum, Wirth&rsquo;s compiler turns into twenty-seven instructions, of which only four are useful. All the rest is reading and writing the sum to memory, the loop counter and — yes — bounds checks and NIL checks again. The thing is, Wirth&rsquo;s compiler can&rsquo;t keep variables in the processor&rsquo;s registers and goes to memory for them every time. One of the reviewers predicted this from the sources on the first day and guessed the cycle count almost exactly. This, by the way, isn&rsquo;t an oversight but deliberate simplicity: the whole of Wirth&rsquo;s compiler is under three thousand lines and builds itself in seconds, and complex optimization simply didn&rsquo;t fit that budget.</p>
<h2 id="the-fast-multiplier">The fast multiplier</h2>
<p>Since it all comes down to the multiplier, I wrote a fast one. It does the same thing as Wirth&rsquo;s, but not one bit per cycle — all at once, in one or two cycles — with the rounding and other subtleties rewritten exactly after Wirth. It&rsquo;s switched on the same way as the <code>CHK</code> instruction, with a single toggle when building the processor.</p>
<p>The main requirement of it was as strict as everything else: the results must match Wirth&rsquo;s multiplier to the last bit. I ran tens of millions of pairs of numbers through it, including all the awkward edge cases, and found not a single divergence. And to make sure the check can notice errors at all, I deliberately spoiled the rounding in one place, and it found tens of thousands of divergences.</p>
<p>With the fast multiplier the model ran 1.6 times faster — about fourteen letters a second instead of nine. The estimate from the profile predicted exactly that, but on a simple machine with no caches that&rsquo;s more a confirmation that the profile was computed correctly than a real prediction. Even if multiplication became entirely free, the program couldn&rsquo;t be sped up by more than 1.64 times, because the other cycles don&rsquo;t go anywhere. So the reviewers&rsquo; number didn&rsquo;t quite come true, but on the main point they were right: the multiplier was indeed the bottleneck.</p>
<p>For the system&rsquo;s ordinary work, though, the fast multiplier gives nothing at all. The compiler building itself does, across forty million instructions, a mere thirty-one fractional multiplications, and the system boot does none. The slow multiplier was a perfectly sensible choice for a machine on which you write text and build programs. The language model was the first task for which it became the bottleneck.</p>
<p>And what of the new instruction that multiplies and adds, the one this all began with? I didn&rsquo;t implement it; I estimated it from the counters. On Wirth&rsquo;s ordinary multiplier it would give one and a half to six percent, on the fast one up to ten. So the main lever here isn&rsquo;t the new instruction but a compiler that learns to keep the sum in a register instead of running to memory for it. But that I no longer measured.</p>
<p>In hardware the fast multiplier was even smaller than the native one, because an FPGA has ready-made hardware multiply blocks (its DSP slices), and all the serial-counting logic simply vanished. The single-cycle variant, mind you, noticeably lowers the processor&rsquo;s top frequency, while the two-cycle one is almost as fast and doesn&rsquo;t touch the frequency. The native 25 MHz holds with a large margin in both cases, so you&rsquo;d only have to choose between them if someone wanted to overclock Wirth&rsquo;s machine.</p>
<h2 id="two-finds-along-the-way">Two finds along the way</h2>
<p>The first find concerns the comparison of fractional numbers. Wirth&rsquo;s compiler compares them the same way as integers, through subtraction and a check of the processor&rsquo;s flags, and for some comparisons it uses the overflow flag. Only that flag is changed exclusively by integer operations. As a result, if an integer overflowed somewhere in the program before a comparison of fractional numbers, the comparison can give a wrong answer. I checked this with a very simple program. First it honestly says that one is less than two; then it overflows an integer; after which it reports, just as confidently, that one is not less than two. This reproduces both on the emulator and on the circuit. My model has no overflows, and the reference checks for that. I claim no novelty here: surely someone has seen this already; correct me if you know where.</p>
<p>The second find was funnier. The model needed to save a text file to disk, and here it emerged that in the whole project nobody had ever written anything to our QEMU, because the system boot only reads from disk. And it couldn&rsquo;t write, because the disk was attached without write permission, and the very first attempt to save anything brought QEMU down entirely. On top of that, Oberon&rsquo;s filesystem places new data past the end of the disk, and the disk image was exactly the size the system occupied, so the file silently wasn&rsquo;t saved either.</p>
<p>When I fixed that and ran the machine in Kubernetes, writing still didn&rsquo;t work — now because of file permissions. The task that places the machine&rsquo;s disk onto its volume was supposed to permit writing, but did it in a way that actually permitted nothing. All that time the disk in the cluster was read-only, and had anyone saved a file inside Oberon, the machine would have crashed. Nobody noticed, because nobody was saving anything. Now the permissions are set explicitly, including on already-created disks, and the sandbox check scenario asks QEMU itself whether it opened the disk for writing.</p>
<p>All the details, tables and commands for reproducing this are in the <a href="https://github.com/tym83/paleocomputing/blob/main/13-episode-lm-on-risc5.md">episode write-up</a>, and the essentials can be reproduced in a couple of minutes with <code>cd impl &amp;&amp; make lm-check &amp;&amp; make lm-profile</code>.</p>
<h2 id="descriptors-between-chk-and-cheri">Descriptors, between CHK and CHERI</h2>
<p>The <code>CHK</code> instruction checks an index against a length that the compiler baked into the instruction itself, which means the compiler has to know the array length in advance. CHERI keeps the bounds in the pointer, and the check can be neither forgotten nor bypassed. Between these two extremes there historically sat descriptors as well, as in that very Burroughs B5000 of the early sixties. A descriptor is a pointer that carries the array length with it, and the processor checks it on every access, so the compiler no longer needs to know anything in advance. I wanted to put this rung on the ladder and measure it on the same fully open system.</p>
<p>The very first observation narrowed the task sharply. In Oberon the compiler knows the length of almost any array in advance, because the language simply has no pointers to arrays and no variable-length arrays. There&rsquo;s one exception: when an array is passed to a procedure prepared to accept an array of any length. Inside such a procedure the length isn&rsquo;t known in advance, and that&rsquo;s exactly where a descriptor has something to do. Everywhere else <code>CHK</code> already copes.</p>
<p>I fit the descriptor into an ordinary 32-bit word. Wirth&rsquo;s processor has a 24-bit address, but the machine has only a megabyte of memory, and twenty bits are enough for a megabyte, so I gave the remaining twelve bits to the array length — up to four thousand elements and a bit. And with it one new instruction, which takes a descriptor and an index, computes the element&rsquo;s address and, if the index runs past the length, jumps to the fault handler. An ordinary address, fed in place of a descriptor, looks like a zero-length array, so any access through it is immediately treated as an error. This is the closest analogue of CHERI&rsquo;s central rule — that without permission there is no access — that you can build without a special hardware tag.</p>
<p>The new instruction was checked as strictly as <code>CHK</code>: with the same deliberate breakages, lockstep comparison against the emulator, and a check that the ordinary system boots the same on a processor with the new instruction as on the ordinary one. And once again the first attempt lied. The run with deliberate breakages first reported that it had caught them all. The script that introduced the breakages was itself crashing, the circuit wasn&rsquo;t building, and the fact that nothing had built was being counted as a caught breakage. Now a circuit that fails to build counts as a failed check, not a passed one.</p>
<p>On the same loop as on the two-core page, the descriptor turned out to be the fastest of all — faster even than the variant with no checks at all, eight cycles per access against nine. There&rsquo;s no miracle in this. The new instruction also computes the array element&rsquo;s address, which usually takes two more instructions, and had I made an identical instruction without the check, it would run just as fast. The lesson here is a different one, and I like it: if the bound travels with the pointer, the check can be hidden inside an instruction you need anyway. That is exactly how it&rsquo;s done in CHERI, where the check sits inside the memory read.</p>
<p>Next I taught the compiler to use descriptors where the length isn&rsquo;t known in advance, and asked the system to rebuild itself with the new compiler. It rebuilt, and the next generation of the compiler matched the previous one byte for byte. This is an important check, because if the compiler had anywhere forgotten to turn a descriptor back into an ordinary address, the address would have been wrong and there&rsquo;d have been no match.</p>
<p>And then the real system found two boundaries I&rsquo;d never have thought of myself. The first concerns array length. In the whole of Project Oberon there isn&rsquo;t a single array longer than four thousand elements, but in the tools I built it with there turned up a sixteen-thousand-element buffer, and the compiler honestly refused to build it rather than silently truncating the length. The second boundary was more interesting. When I built many modules in one run, memory in the emulator crept past a megabyte, and my descriptor can address only a megabyte, so the build simply hung, silently. Hence a very telling conclusion. A pointer with bounds can&rsquo;t be squeezed into the width of an ordinary address, because as soon as there&rsquo;s more memory, there&rsquo;s no room left for the length. That&rsquo;s exactly why CHERI&rsquo;s pointers are twice as wide as an address and, on top of that, compress the bounds cleverly.</p>
<p>What does all this construction cost on large loads? On a synthetic program that works heavily with arrays of unknown length, descriptors give a speed-up of about a quarter, for the same reason as in the loop: they save on computing the address. On the compiler, though — a real program — they instead give a tiny slowdown, under a percent, and I spent quite a while hunting for where it comes from, because by calculation it should have been a speed-up. It had nothing to do with descriptors: every time the program passes a string to a procedure, the compiler now assembles a descriptor for it, which made the code a little longer, and the system spends slightly more time loading modules. Had the compiler prepared the string descriptors in advance, this wouldn&rsquo;t happen, but that I no longer did.</p>
<p>On the FPGA the processor with descriptors still holds its 25 MHz, has about five percent more logic, and the new instruction landed, for the first time, on the circuit&rsquo;s critical path — its longest signal path — meaning it could limit the frequency in future. <code>CHK</code> never once did that.</p>
<p>In the end the ladder looks like this:</p>
<table>
  <thead>
      <tr>
          <th>Rung</th>
          <th>Where the bound is stored</th>
          <th style="text-align: right">Check instructions in the loop</th>
          <th>Main cost</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>RISC5, software check</td>
          <td>in the code</td>
          <td style="text-align: right">2</td>
          <td>slowest of all</td>
      </tr>
      <tr>
          <td>RISC5, <code>CHK</code></td>
          <td>in the instruction itself</td>
          <td style="text-align: right">1</td>
          <td>doesn&rsquo;t work if the length isn&rsquo;t known in advance</td>
      </tr>
      <tr>
          <td>RISC5, descriptor</td>
          <td>in the pointer</td>
          <td style="text-align: right">0</td>
          <td>memory up to a megabyte, arrays up to four thousand elements</td>
      </tr>
      <tr>
          <td>CHERI</td>
          <td>in a wide, tagged pointer</td>
          <td style="text-align: right">0</td>
          <td>pointers twice as wide</td>
      </tr>
  </tbody>
</table>
<p>A descriptor removes the check from the loop just as CHERI does, but in Oberon it has almost nothing to protect beyond what <code>CHK</code> already does. And the key thing that sets it apart from CHERI: a descriptor stays an ordinary number, and a program permitted low-level operations can assemble any descriptor from any length and any address. It&rsquo;s precisely here that CHERI puts a special tag that can&rsquo;t be forged. A system built entirely with descriptors I haven&rsquo;t yet booted on the real circuit; this processor isn&rsquo;t in the browser or the catalog yet either; and all the details and commands for reproducing it are in the <a href="https://github.com/tym83/paleocomputing/blob/main/14-episode-descriptors.md">episode write-up</a>.</p>
<h2 id="how-to-install-all-of-this-yourself">How to install all of this yourself</h2>
<p>There are four ways, from the very simplest, where you need install nothing, to your own cloud.</p>
<p>The easiest is to open the <a href="https://tym83.github.io/paleocomputing/oberon/lab.html">lab</a> in your browser and do at least the first four exercises. You&rsquo;ll see how an interface where any text can be a command works, write your first module, and break the machine by writing garbage straight into screen memory. If you don&rsquo;t want the exercises, the <a href="https://tym83.github.io/paleocomputing/oberon/run.html">Just run the system</a> page holds the bare system, and the <a href="https://tym83.github.io/paleocomputing/oberon/checks.html">Change the processor</a> page lets you see the cost of a check with your own eyes. The middle mouse button is stood in for by a click with Alt, and the two-button chord by a click with Shift. Of the things worth trying yourself, I&rsquo;d suggest a middle-click on <code>System.ShowModules</code>, to see how few modules the system needs after boot, and on <code>Hilbert.Draw</code>, because it&rsquo;s simply pretty. And then lab exercises twelve and thirteen, the most hardware-facing: in one you measure the three processor variants yourself, in the other you add a new built-in procedure to the language and rebuild the compiler right inside the system.</p>
<p>If you want Wirth&rsquo;s machine as an ordinary VM, build your own QEMU from the commands in the QEMU section or from the <a href="https://github.com/tym83/paleocomputing/blob/main/qemu/GUIDE.md">guide</a>. It&rsquo;s worth trying the processor variants with the hardware check and with descriptors there.</p>
<p>If you have your own KubeVirt, there&rsquo;s a <a href="https://github.com/tym83/paleocomputing/blob/main/kubevirt/GUIDE.md">guide</a> and the short version in the spoiler above. The two latest KubeVirt versions are supported, and servers on both Intel and Arm. You can start with a throwaway test cluster; the same scenario the automatic check runs can be launched by hand. And I&rsquo;ll repeat the warning once more, because I got burned by it myself: the housekeeping image is changed for all of the cluster&rsquo;s VMs at once, so on a working cluster, first check whether automatic machine updates are on.</p>
<p>And if you have Cozystack, it&rsquo;s enough to plug in the catalog with the commands from the Cozystack section, and your users will get Wirth&rsquo;s machine, the lab, the handbook and a bundle that installs it all at once. The details are on the <a href="https://tym83.github.io/paleocomputing/cozystack/">project page</a>. If you need to show students or colleagues how a whole computer is built, this is probably the quickest way. As for the automatic machine updates, the component now asks about them itself and won&rsquo;t touch anything without explicit consent.</p>
<p>And if you&rsquo;d like to check up on us, the command <code>cd impl &amp;&amp; make deps &amp;&amp; make check</code> will, in about eight minutes, run all the processor tests, the system boot, the lockstep comparison against the reference, the compiler self-build, and a rebuild of the whole system with a byte-for-byte comparison. For this you&rsquo;ll need Verilator and a C++ compiler.</p>
<h2 id="what-i-took-away-from-this">What I took away from this</h2>
<p><strong>First, on bounds-checking.</strong> It&rsquo;s cheaper than commonly thought, but I wouldn&rsquo;t call it entirely free. On Wirth&rsquo;s processor — very simple, no caches, no branch prediction — it cost the compiler between two and five percent of the time, depending on which part of the system you remove it from. And most of that cost comes from null-pointer checks, not array-bounds checks. On large modern processors, in my deliberately favorable loop, it&rsquo;s invisible, because the extra instructions drop into free slots; on the server Arm it&rsquo;s visible, if only a little. In real programs checks are costly primarily because they stop the compiler from processing arrays in batches, and that I didn&rsquo;t measure. CHERI hides the check right inside the memory read, and then there&rsquo;s no check in the program at all.</p>
<p><strong>Second, the bottleneck often turns out not to be where you look for it.</strong> For the neural network everyone, myself included, wanted a clever new instruction, while the bottleneck was the old slow multiplier. Replacing it sped the model up 1.6 times, whereas the new instruction, by estimate, would have given a few percent. The reviewers said this on the very first day, and a measurement a week later confirmed it. First measure, and only then add hardware.</p>
<p><strong>Third, a pointer that knows its own bounds has to be made wider than an ordinary address.</strong> My descriptors, fit into an ordinary word, ran up against a megabyte of memory and arrays of up to four thousand elements, and that is exactly why CHERI&rsquo;s pointers are twice as wide.</p>
<p><strong>Fourth, the nastiest bugs are found only by a live environment.</strong> My tests were good, but the most interesting things were found by real KubeVirt, real Cozystack and a live Arm server — from a machine that needs no power management, to a network layer that silently doesn&rsquo;t answer, to sixteen seconds before the whole cluster migrated. Since then I release a new version only if it has first passed a full check in the sandbox.</p>
<p><strong>And fifth, a neural network writes code fast and proves it works slowly.</strong> Nine days for all of this happened only because the code was written by agents. But more than half of those days went into teaching the checks to blush honestly, and for every pretty number there was an auditor — also an agent — who took it apart. If a check has never once turned red, it most likely isn&rsquo;t checking anything, especially when you badly want everything to work out. And a separate open question is what to do with code like this in open-source projects. QEMU and libvirt won&rsquo;t accept it at all; the CNCF, by contrast, takes it calmly; and how open source is to live with this isn&rsquo;t yet very clear.</p>
<h2 id="in-place-of-a-conclusion">In place of a conclusion</h2>
<p>All of this is open. The sources are in the <a href="https://github.com/tym83/paleocomputing">repository</a>, our code under the Apache-2.0 licence, and the processor for QEMU under the GPL, like QEMU itself. The site with the lab is <a href="https://tym83.github.io/paleocomputing/">tym83.github.io/paleocomputing</a>. Every find, with the commands to reproduce it, is in the repository in the <code>impl/docs</code> folder, and about Cozystack itself you can read at <a href="https://cozystack.io/">cozystack.io</a>.</p>
<p>Coming up next in the series: the Burroughs B5000 with its hardware memory protection; Lilith, Wirth&rsquo;s very first machine, which will need only a new passport; and, patience permitting, my own board with Oberon on an FPGA, to finally measure the cycles not in simulation but on real hardware.</p>
<p>If you use Oberon in teaching, or simply wrote in it once, tell me in the comments — I&rsquo;m very curious where it lives now. And do correct me if I&rsquo;ve gone wrong somewhere; judging by this article, that happens to me regularly.</p>
<p><strong>A quick poll: what will you do after this article?</strong></p>
<ul>
<li>Open the lab and break the machine by writing garbage into screen memory</li>
<li>Go check auto-updates for the VMs on my own cluster. Right now</li>
<li>Rewrite everything in Oberon (Go is basically Oberon, only with goroutines)</li>
<li>Keep turning off bounds-checking for the sake of a couple of percent</li>
<li>I wrote in it back in the nineties, and I have something to say in the comments</li>
<li>Made it as far as the poll, which is an achievement in itself</li>
</ul>
<p>P.S. A special thank-you to Niklaus Wirth. I never met him, but for nine days I talked to his circuit, and it lied to me almost never. Unlike my own checks.</p>
]]></content:encoded></item><item><title>What dozens of AI agents taught me: how I wrote the Blockstor storage system as an experiment</title><link>https://aenix.io/blog/2026/08/what-dozens-of-ai-agents-taught-me-how-i-wrote-the-blockstor-storage-system-as-an-experiment/</link><guid isPermaLink="true">https://aenix.io/blog/2026/08/what-dozens-of-ai-agents-taught-me-how-i-wrote-the-blockstor-storage-system-as-an-experiment/</guid><pubDate>Mon, 10 Aug 2026 00:00:00 +0000</pubDate><dc:creator>Andrei Kvapil</dc:creator><category>Kubernetes</category><category>LINSTOR</category><category>Storage</category><category>AI and ML</category><category>Cozystack</category><category>Open Source</category><description>Andrei Kvapil on building Blockstor as a clean-room, Kubernetes-native block storage orchestrator by driving up to 60 AI agents with TDD and hard exit gates.</description><content:encoded><![CDATA[<p>A couple of months ago I decided to run an experiment: build a clean-room implementation of LINSTOR from scratch, working only from its references and public API types. It started as a Friday joke. I wanted to spend as little time on it as possible, leave it running in the background, and see where it went. The point was to find out how far a modern model can get on its own, with no human in the loop.</p>
<p><img src="https://aenix.io/img/blog/medium/what-dozens-of-ai-agents-taught-me-how-i-wrote-the-blockstor-storage-system-as-an-experiment/cover.jpg" alt="Blockstor storage system built with AI agents" width="1200" height="630" loading="lazy" decoding="async"></p>
<p>Spoiler: full autonomy didn’t happen, and I ended up wrestling with the project quite a bit. But the process pulled me in completely, and the end result beat every expectation I had.</p>
<blockquote>
<p>People in the community kept asking how it actually went. Fair question — building this taught me a great deal about driving models effectively, and it turned up a pile of working methods and patterns that have since made me much more productive in everyday work too.</p>
</blockquote>
<h2 id="what-blockstor-is">What Blockstor is</h2>
<p>I called the project Blockstor. It’s a block device orchestrator. Roughly: you request a replicated volume of the size you need, and that volume gets created on several nodes in ZFS or LVM and set up for replication with DRBD — a smart, network-aware take on RAID 1. The system supports snapshots, replica reallocation, resize, automatic failover, and more.</p>
<p>Blockstor is not a storage system designed and written on a blank page. I built the experiment on LINSTOR — a mature distributed block storage manager that I’ve run in production myself for years.</p>
<p>LINSTOR was close to an ideal storage system for me, because its API types are Kubernetes-like. Its backend logic, though, is organized in a rather peculiar way. The main problem, as I see it, is the request-based model: for most API calls it goes out to the nodes in real time and polls their current state to build a response. To my mind that’s an unacceptable pattern for a distributed system, because it doesn’t hold up at scale. The absence of a reconciliation loop also makes automatic recovery hard. I’m not going to second-guess the developers’ decisions — I’m sure they had their reasons. But given how committed I am to Kubernetes patterns, I asked myself how I — an experienced architect and developer — would write a system like this in Kubernetes-native logic.</p>
<p>Rewrites from one language to another aren’t rare, by the way. That, for example, is how the rusternetes project came about: Kubernetes being rewritten from Go to Rust. But how is that even possible? Kubernetes has an enormous codebase with an enormous number of person-hours in it. How do you convince yourself that the resulting “AI slop” works correctly? And can you rely on it at all, let alone run it in production?</p>
<h2 id="tdd">TDD</h2>
<p>The answer lies in how Kubernetes itself is developed in the open. Free, community-driven projects have a body of established practice, and contributors do their best to follow it.</p>
<p>All kinds of people contribute to projects like these, at every skill level. That makes tests critical: tests are what keep functionality someone contributed from breaking later. Say you land a new feature, it passes every test, and it goes into the project. Later someone lands another feature. If a test for your feature fails while theirs is being checked, that means the new functionality has a problem that breaks yours — and the system automatically refuses to let the change through. Over time, tests in public projects have become so important that they’re now valued more highly than the code itself. Tests like these are exactly what makes rewriting X in Y possible.</p>
<p>The rusternetes project claims:</p>
<blockquote>
<p><em>Actively conformance-tested against the official Kubernetes e2e test suite — currently passing 94% of conformance tests (415/441) across 160 rounds of testing.</em></p>
</blockquote>
<p>So passing the tests is the primary yardstick for whether a project counts as a Kubernetes replacement.</p>
<blockquote>
<p><em>When I see a bird that walks like a duck and swims like a duck and quacks like a duck, I call that bird a duck.</em></p>
<p><em>— The famous <a href="https://en.wikipedia.org/wiki/Duck_test">duck test</a></em></p>
</blockquote>
<h2 id="hooks-and-setup">Hooks and setup</h2>
<p>For the same reason, the first thing I did was force the model into test-driven development (TDD), where tests are written before the implementation. Tests let you pin down the conformance requirements you need. They also give me confidence that new functionality arriving during development won’t break what already works.</p>
<p>On a colleague’s advice I wired in golangci-lint right away and made it mandatory for the model to run all code through it. <a href="https://github.com/lexfrei">@lexfrei</a> — the colleague in question — argues that this step saves a substantial number of tokens.</p>
<h2 id="the-exit-gate-and-a-deterministic-result">The exit gate and a deterministic result</h2>
<p>This is probably the first and most important lesson: you can tell a model “keep writing until the tests pass”, and that’s a good condition on the artifact it produces. That’s exactly the method I used to build this project. With one catch: I had no tests, because LINSTOR doesn’t publish a test suite for its own functionality.</p>
<p>All I had to go on was years of running LINSTOR, dozens of my own talks and articles, and the Apache 2.0 code of projects in the LINSTOR ecosystem — golinstor, the Container Storage Interface (CSI) driver, Piraeus-operator, and their API contracts. LINSTOR itself is licensed under GPL, which rules out using its code to build an Apache 2.0 project on top of it. That was the central challenge, and it’s precisely why I went the clean-room route.</p>
<p>So my main job became defining those exit gates. Get them right and I don’t have to police the whole resulting codebase: if the code satisfies the conditions I set myself, that is the evidence the job is done.</p>
<h2 id="the-api-as-a-contract">The API as a contract</h2>
<p>I like the type definitions LINSTOR uses in its API. On top of that, both the official CSI driver and Piraeus-operator work with those types, and I had no wish to rewrite either of them.</p>
<p>From experience, designing APIs is one of the most painful jobs in this profession. You can always change backend logic. Ship one breaking change to the API and everything falls apart on the client side. So an API should change rarely, in exceptional cases only — and the risk of getting it wrong at the very start is high. That’s why I decided to use exactly the types the original project uses, and map them onto Kubernetes-first CustomResourceDefinition (CRD) types.</p>
<p>While building Cozystack and selling solutions based on it, we’ve run into the same question from prospects more than once: what’s the story with your API? What if we put you into production, build a solution around you, and then your API changes and we have to rewrite our half of the project?</p>
<p>As a reminder: LINSTOR’s code is published under GPLv3, and I planned to publish Blockstor under Apache 2.0. Reusing the original code was off the table. So the first task was to find compatible contracts in the Kubernetes tooling around it: the <a href="https://github.com/LINBIT/golinstor">golinstor</a> library, linstor-csi, and Piraeus-operator, all distributed under Apache 2.0, plus the official LINSTOR documentation, licensed under CC BY-SA.</p>
<p>That gave me my first contract: the API must be compatible with the LINSTOR Go library, linstor-csi, and Piraeus-operator.</p>
<h2 id="the-test-environment">The test environment</h2>
<p>The model needed somewhere to work and somewhere to test what it produced. For the test environment I picked a beefy bare-metal node and a generated test suite that brought up a Kubernetes cluster on Talos inside virtual machines (VMs) and deployed Blockstor into the cluster. That choice was deliberate: Blockstor needs the DRBD module, and DRBD has a habit of hanging when it’s configured wrong, so I needed a fast way to stand up and recreate a broken environment. Blockstor’s architecture also stores configuration as Kubernetes CRDs, so using Talos closed the question of bootstrapping Kubernetes itself.</p>
<h2 id="the-first-result">The first result</h2>
<p>When the first proof of concept (PoC) was ready, the model had built me a working prototype, guided only by the sources above and its own dataset. For the interaction model, though, it had implemented the very same request-based model as the original LINSTOR — presumably picked up from the project’s documentation. And it already worked! I could talk to the API using the official CLI, though there were hundreds of bugs and gaps.</p>
<p>To fix that, I asked the model to freeze the API contracts as tests and rewrite the logic on controller-runtime, the way I needed it. With a clear, deterministic goal, the architecture that came back began to resemble what I actually wanted: fully asynchronous, with the API translator as a separate pluggable module.</p>
<p>Later I had the model implement a set of end-to-end (e2e) tests that requested volumes through the official CSI plugin and the Kubernetes kubernetes-csi/csi-test framework. In other words, the model kept working until Blockstor provisioned volumes and snapshots through standard Kubernetes abstractions. After a while, I had a working prototype. The system was still a long way from stable, though, which left me with a question: how do you reach the stability I needed when there are no tests to start from?</p>
<p>This is where all my material went in — my articles on debugging LINSTOR, my talks, my Claude Code debugging skills, our plunger scripts integrated into Cozystack, and bugs reported on GitHub. I made the model study all of it and compile a set of issues that had to be tested and that our system had to satisfy. Out came a huge Markdown file. I reviewed it and sent the model off to test and fix every problem it had caught, on the test rig. There were many such iterations, and each one took a colossal amount of time. Most of the time the model would bring up the environment, run tests, fix errors, re-provision the environment — and all of that took time.</p>
<p>Eventually I started asking how to speed the process up.</p>
<h2 id="speeding-up-development">Speeding up development</h2>
<p>That’s where the idea of parallelizing agents came from. Some problems could be closed out at the unit-test level. Others could only be verified on a real environment. Running the same set of bugs through again and again, I arrived at this method:</p>
<ol>
<li>I ask one agent to gather the references and put together a plan of sorts for the bugs.</li>
<li>Once I have the huge document, I ask the model to dispatch a batch of agents, each working a specific bug.</li>
<li>In the code, every agent gets its own isolated environment and its own worktree, where it submits its fixes.</li>
<li>Each agent’s job is to study the problem, bring up an environment, prepare a fix and tests, and hand all of it back to the main agent.</li>
<li>The main agent pulls the work from every agent into the main tree — and that repeats for several iterations.</li>
<li>In the end, the number of agents running at once reached 60: 30 at the unit-test level, 30 on e2e in a concrete environment.</li>
</ol>
<p>After a while the project started to look like something you could actually use. Plenty of things still weren’t stable.</p>
<h2 id="the-marathon-continues">The marathon continues</h2>
<p>While nursing this whole zoo along, I was already doing integration testing myself and checking the results by hand. I’d ask Claude for access to an environment, drive the linstor CLI manually, and try to reproduce the bugs, of which there were still plenty.</p>
<p>At that point I had to stop building large blocks and start digging into far more meticulous testing of the user journey. The bulk of the problems went away once I made the agents implement tests using the official CLI and laid out the user journey for Day-2 operations from the official LINSTOR documentation. In some places, though, the model started spinning its wheels: over all that time it never got every test passing reliably, and I had to step in personally. Drawing on what I know, I asked it to walk me through each problem, then put a set of leading questions to it about the architecture.</p>
<h2 id="the-fine-tuning-stage">The fine-tuning stage</h2>
<p>A lot of my questions came down to the asynchronous nature of the controllers. DRBD is by nature fussy about when, how, and at what point configuration gets applied, so it was critical to build a stable state machine — one that would let the reconciliation loop through to a specific action only when the conditions were met.</p>
<p>The main problem here was snapshot logic. To create a snapshot, LINSTOR first freezes I/O at the DRBD level, then sends the command simultaneously to the several nodes holding the backing device in ZFS. Once the operation succeeds, LINSTOR unfreezes DRBD. All of it has to happen instantly, and on error the state has to roll back automatically so that the container or VM never blocks while working with its data.</p>
<p>The second problem was how to skip the initial sync of a new replica. I knew the observable behavior from years of operating LINSTOR: the first replica records its current generation identifier, and later replicas take that as their starting value and sync from there.</p>
<p>First I tried to reconstruct the exact command sequence from observed behavior alone, running operations against a live controller and capturing the commands. But the mechanics aren’t documented publicly, and observation on its own didn’t add up. So I fell back on the clean-room technique. One agent — the “dirty room” — reconstructed a functional specification from the original sources: which commands run, and under what conditions. It recorded those functional facts and nothing else, carrying over no code and no expression of it. The second agent — the “clean room” — had no access to anyone else’s sources in the first place, and implemented the approach from scratch, strictly from the specification.</p>
<p>One more problem was consistent node-id allocation — every DRBD replica needs a unique number in the cluster, from 1 to 8 — along with allocating TCP ports for replication, which in the current implementation is no longer tied to DRBD and is handed out from a per-node pool instead. This is exactly where the state machine above paid off.</p>
<h2 id="building-the-ci-system">Building the CI system</h2>
<p>By now it was clear the experiment was reaching its final stage, and I started thinking about the project’s future. To finish it and get to production, we needed a serious continuous integration (CI) system that would guarantee no untested code in the codebase. Given the volume of tests, we couldn’t let a run stretch over several hours, so they had to be parallelized the same way.</p>
<p>And we built one: for every pull request it ran a big pile of tests across six or seven runners in parallel and returned “ok” or “not ok”.</p>
<p>I worked in that mode for several more rounds, until CI was genuinely green and stable. Then we switched from local worktrees to pull requests on GitHub.</p>
<h2 id="the-final-stage">The final stage</h2>
<p>Data loss is not acceptable, so before declaring the project finished I had to check thoroughly that Blockstor behaves correctly and reliably on that front. Reading and reviewing all the generated code was beyond both my energy and my capacity. On the other hand, Blockstor — like LINSTOR — is essentially an orchestrator. The data itself is stored by ZFS and DRBD, and I had no doubts about their reliability. What mattered was confirming that the controller really does configure resources correctly and survives dropouts.</p>
<p>But how do you establish whether the current code can be trusted with production? It’s an ambiguous and difficult question. And how do you get a deterministic answer to it? Here I used one more pattern I worked out over the course of the project.</p>
<p>My colleague @lexfrei had mentioned earlier that models are terrified of being told their actions could lead to serious financial losses. I decided to use that as an exit gate. I launched an agent responsible for releasing the code to production and had it work under conditions where its life and financial well-being “depended” on the final result. It drew up an acceptance plan with a long list of items, including a 24-hour burn-run on real infrastructure. A second agent tried to satisfy those requirements. That went on for several more days and several releases, until the agent in charge of shipping to production finally gave its blessing.</p>
<p>A system nobody uses can’t be called stable, though. So the next step was integrating Blockstor with Cozystack. Our test suite covers many functions at once: requesting volumes with and without DRBD, RWX, snapshots, and other details. The new agent’s job was to prepare a draft pull request, turn the tests green, and get through the final stage.</p>
<p>As of July 15, 2026, that pull request isn’t merged yet — but there’s a good chance you’ll soon have a new Kubernetes-native storage backend in Cozystack. Watch this space.</p>
<blockquote>
<p><strong>Editor’s note (September 2026):</strong> Blockstor is still a separate experimental project — <a href="https://github.com/cozystack/blockstor">github.com/cozystack/blockstor</a>, Apache 2.0 — and it does not ship in Cozystack. LINSTOR/DRBD via Piraeus remains the storage Cozystack ships and its default backend. Blockstor sits on the roadmap as an opt-in backend for 2027.</p>
</blockquote>
<h2 id="the-main-takeaway">The main takeaway</h2>
<p>Blockstor still has experimental status. But the experience let me speed up and automate work on other projects considerably, and I now apply it every day.</p>
<p>Plenty of people still think of AI as an autocomplete tool. For me that stopped being true a long time ago. AI-first development isn’t occasionally asking a model to write a function. It’s rebuilding your entire engineering process around agents, context, tests, skills, plans, review, and automation.</p>
<p>Working that way, you can take on tasks that used to look too big for a small team — experimenting with a Kubernetes-native implementation of a LINSTOR-class storage system, for instance. But it only works on one condition: you need discipline. Without a plan, AI creates chaos far faster than a human can. With a plan, tests, agents, and proper review, AI becomes a force multiplier for engineering.</p>
<p>And this, I think, is what building complex infrastructure looks like from here on: not one engineer against an enormous codebase, but an engineer as the architect of a process, with a swarm of specialized agents working around them. The engineer’s job isn’t to watch each agent. It’s to build a system where the agents watch themselves and come to a human only with what genuinely needs one.</p>
<h2 id="join-the-community">Join the community</h2>
<ul>
<li><a href="https://github.com/cozystack/blockstor">Blockstor on GitHub</a></li>
<li><a href="https://github.com/cozystack/cozystack">Cozystack on GitHub</a></li>
<li>Telegram <a href="https://t.me/cozystack">group</a></li>
<li>Slack <a href="https://kubernetes.slack.com/archives/C06L3CPRVN1">group</a> (Get invite at <a href="https://slack.kubernetes.io/">https://slack.kubernetes.io</a>)</li>
<li><a href="https://calendar.google.com/calendar?cid=ZTQzZDIxZTVjOWI0NWE5NWYyOGM1ZDY0OWMyY2IxZTFmNDMzZTJlNjUzYjU2ZGJiZGE3NGNhMzA2ZjBkMGY2OEBncm91cC5jYWxlbmRhci5nb29nbGUuY29t">Community Meeting Calendar</a></li>
</ul>
<hr>
<p><a href="https://blog.aenix.io/what-dozens-of-ai-agents-taught-me-how-i-wrote-the-blockstor-storage-system-as-an-experiment-921f7d3a1137">What dozens of AI agents taught me: how I wrote the Blockstor storage system as an experiment</a> was originally published in <a href="https://blog.aenix.io">Ænix</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>
]]></content:encoded></item><item><title>Focus on today: how we built aeman, a daily planning board for engineers on top of GitHub Projects</title><link>https://aenix.io/blog/2026/07/focus-on-today-how-we-built-aeman-a-daily-planning-board-for-engineers-on-top-of-github-projects/</link><guid isPermaLink="true">https://aenix.io/blog/2026/07/focus-on-today-how-we-built-aeman-a-daily-planning-board-for-engineers-on-top-of-github-projects/</guid><pubDate>Fri, 24 Jul 2026 00:00:00 +0000</pubDate><dc:creator>Andrei Kvapil</dc:creator><category>Platform Engineering</category><category>Developer Tools</category><category>Open Source</category><category>Kubernetes</category><category>Cozystack</category><description>How Ænix built aeman, an open-source daily planning board for engineers that uses GitHub Projects v2 as its only storage and a Kubernetes-style watch API.</description><content:encoded><![CDATA[<p>My name is Andrei Kvapil, and I’m a co-founder of Ænix — we created Cozystack, an open source cloud platform, and co-maintain it, and we help companies build infrastructure. We’re a fully remote company: 15 people at the time of writing (July 2026), several teams (two reliability teams, a development team, marketing, back office, and so on), spread across several time zones. We started out living in GitHub Projects, but the moment we began to grow we ran straight into the limits of our own process: tasks scattered across boards and chats, half of the morning sync spent figuring out what was even in flight, and unplanned work eating entire days without leaving a trace anywhere.</p>
<p>This article is the story of how we fixed that with aeman — a tool we built ourselves and recently <a href="https://github.com/aenix-io/aeman">open sourced</a>. But I have to start further back.</p>
<p><img src="https://aenix.io/img/blog/medium/focus-on-today-how-we-built-aeman-a-daily-planning-board-for-engineers-on-top-of-github-projects/cover.png" alt="aeman board" width="1397" height="650" loading="lazy" decoding="async"></p>
<h2 id="where-this-came-from">Where this came from</h2>
<p>Years ago I worked at another company, and it had two internal systems with wonderful names: Ford and Nixon. Colleagues there wrote about them at length, so here’s the short version. They’re boards for tracking daily tasks, where an engineer sees exactly what they’re working on today. Not a three-month backlog, not a hundred kanban columns — your day, and nothing else. For years that system kept a lot of remote teams working effectively and kept many customers’ infrastructure running.</p>
<p>Let me be honest: the process behind aeman isn’t entirely my idea. I took a working process-management system I’d seen at my previous company and mapped it onto our own processes and tasks. Our board now forces us to focus on what matters most — the business tasks.</p>
<p>Unlike Trello, this model starts from the assumption that one person can work on tasks in several teams during a single day. And the daily plan must not turn into a wall of tasks. If an engineer has 6 to 10 tasks planned but can realistically finish only 3 or 4, they end up with a permanent sense of debt, while management gets a false picture of heavy utilization. People also start switching between tasks constantly, and the brain naturally reaches for the ones that are easier to close. Hard tasks can then sit untouched for a long time. So the daily plan has to be short and honest.</p>
<h2 id="why-not-use-an-off-the-shelf-tool">Why not use an off-the-shelf tool?</h2>
<p>I tried to build a process like this on existing tools: Trello, Asana, Notion, GitHub Projects — with varying success. Notion came closest, but I wasn’t willing to drag the whole team into yet another system and pay for it for the sake of the boards alone. And over time the board inevitably grew until it no longer fit on a screen.</p>
<p>GitHub Projects turned out to be the most viable option. It’s free, it offers a convenient API, and it already has everything you need: sprints, priorities, statuses, and other entities. It also extends well. That’s why we worked on it for a long time.</p>
<p>But it became clear fairly quickly that the problem wasn’t the tool’s capabilities — it was the UX. It takes far too many extra actions to create a task, assign an owner, and fill in the required fields. And we never managed to build a comfortable daily process around GitHub Projects. Sprints exist, for example, but they proved too awkward to use: you can’t close the current sprint and start the next one with the unfinished tasks carried over in a single action, or quickly push a task to “later” so it drops out of sight and stops competing for an engineer’s attention.</p>
<p>Another problem showed up over time. At every sync we were effectively looking at one column — <strong>In Progress</strong>. That’s where all the work happened, while cards in the other columns slowly turned into a task graveyard: still on the board, but no longer part of the daily process and going nowhere. GitHub Projects turned out to be an excellent place to store tasks and a weak one for managing a team’s daily focus.</p>
<p><strong>That’s exactly why, when I started building aeman, I chose GitHub Projects as the backend of the new system and built the user interface and the workflow on top of it.</strong></p>
<h2 id="the-core-idea-exactly-the-cards-you-need-today">The core idea: exactly the cards you need today</h2>
<p>The goal of aeman is to surface the tasks that matter today and focus the engineer on them. In an ideal world, an engineer leaves the morning sync with a list of tasks they will <strong>definitely</strong> work on that day. Everything else moves to tomorrow or to next week — and physically disappears from the board so it stops catching your eye.</p>
<p>The second idea: unplanned work has to be visible. Everyone has days that get eaten by a sudden incident or an urgent request from a neighboring team. Usually that work is recorded nowhere, and at the end of the sprint nobody can explain where the time went. In aeman it has a zone of its own: you log it after the fact as a single card, and the day stops vanishing without a trace (screenshots in the sections below). At the daily sync you can then talk it through and move the task into the planned block.</p>
<p>The third idea: multi-team work is a property of the model, not a filter. In a small company one engineer can work on tasks in several teams — this is especially true of founders, who still juggle everything. In aeman, teams are a dimension of the board: a sprint is tracked per team, each team has its own weekly plan, and the Me board collects one day’s cards from all of the engineer’s teams at once. You don’t have to hop between boards to see your day. A task also moves freely between people and teams — it can change assignees from one day to the next, and that’s a normal mode of work the tool shouldn’t get in the way of.</p>
<p>And the fourth idea: a team lead needs a tool that shows the status of every task within a minute, without pulling people away with extra questions.</p>
<p>Philosophically, aeman is closer to Todoist than to a classic kanban board. Cards are deliberately short: in normal mode you see only the title, and the description opens on a double click. The description takes free-form links — pull requests, GitHub issues, Telegram chats, anything else — and aeman recognizes them automatically and shows a button for jumping straight there. That keeps the main screen clean, and closing another task delivers the satisfying feeling familiar to anyone who uses a good task manager.</p>
<h2 id="what-it-looks-like">What it looks like</h2>
<h2 id="the-me-board-your-day">The Me board: your day</h2>
<p>An engineer opens aeman in the morning and sees their day: cards grouped into four colored zones. This isn’t the Eisenhower matrix; the breakdown is different:</p>
<ul>
<li><strong>Red</strong> — <em>urgent</em>, has to be done today</li>
<li><strong>Gray</strong> — <em>planned</em>, ordinary planned work, which may take more than one day</li>
<li><strong>Yellow</strong> — <em>unplanned</em>, whatever landed during the day</li>
<li><strong>Green</strong> — <em>nice to have</em>, do it only once everything else is finished</li>
</ul>
<p>The board pulls in cards from all of the engineer’s teams, meaning from the current sprint of each of them. The team selector at the top narrows the list to a single team, or hides every card that’s done, blocked, or in review, leaving only what you can work on right here and right now, when it’s time to get into deep work. Each card has a completion slider from 0 to 100% in steps of 10% (dragging the progress bar to done is surprisingly satisfying), a stage (Review / Locked / Recurrent / Done) that recolors the bar, a counter of days in progress, and links to an issue or a PR. On the right there’s a notes panel: a personal log of the day where you can write things down as you work. That way you don’t have to dig through your memory at the next standup — you skim the notes and report what you did. Notes live in the context of the day; tomorrow the board is clean, but you can always step back a day and reread them.</p>
<p><img src="https://aenix.io/img/blog/medium/focus-on-today-how-we-built-aeman-a-daily-planning-board-for-engineers-on-top-of-github-projects/02.png" alt="The aeman Me board" width="1444" height="966" loading="lazy" decoding="async"></p>
<h2 id="the-team-board-the-whole-team-at-a-glance">The Team board: the whole team at a glance</h2>
<p>The team lead’s view: a grid of people by zones for a chosen day. Columns are engineers with their avatars, rows are the same colored zones. You see immediately who’s working on what, whose work is on fire, and who’s overloaded. The lead creates cards from here too: a couple of clicks and a task is created, assigned, and prioritized.</p>
<p><img src="https://aenix.io/img/blog/medium/focus-on-today-how-we-built-aeman-a-daily-planning-board-for-engineers-on-top-of-github-projects/03.png" alt="The aeman Team board" width="1444" height="966" loading="lazy" decoding="async"></p>
<p>In Me mode the team lead can also impersonate an engineer (the <strong>View as</strong> button) — look at the board through their eyes and tidy it up if needed.</p>
<h2 id="daily-sprints-and-carry-over">Daily sprints and carry-over</h2>
<p>Our sprints are short — one day — and they count <strong>forward</strong>, not backward. At the morning sync the team first discusses what got done yesterday, and then the team lead hits <strong>Carry over</strong>: every unfinished card moves into the new sprint (today), and finished ones stay in yesterday’s history. Recurring tasks are recreated automatically. That’s what day planning is here, and it takes minutes.</p>
<p>If planning decides a task definitely isn’t happening today, there are <strong>+1 day</strong> and <strong>+1 week</strong> buttons: the card disappears from the board until its day, and no history is lost. When that day comes, it shows up again.</p>
<p><img src="https://aenix.io/img/blog/medium/focus-on-today-how-we-built-aeman-a-daily-planning-board-for-engineers-on-top-of-github-projects/04.png" alt="Carry-over controls in aeman" width="529" height="287" loading="lazy" decoding="async"></p>
<h2 id="the-weekly-plan">The weekly plan</h2>
<p>Here we departed from the original idea and started fitting our own workflow onto the tool we’d ended up with.</p>
<p>The founders and I used to plan the team’s week by writing a new list of weekly tasks into Google Sheets and sending each team a “focus for the week” message on Slack. That process has now moved into aeman as well.</p>
<p>Under the team grid lives the weekly plan: the team’s business tasks for the week, laid out in two lanes, “by Wednesday” and “by Friday”. Once a week the founders put tasks there, and the tasks wait their turn. The team lead drags a plan card onto an engineer; it appears on that engineer’s daily board while staying in the plan with a colored marker. A shared progress bar shows how the team is doing against the week, and the weekly <strong>Carry over</strong> moves whatever is still open into the next week.</p>
<p><img src="https://aenix.io/img/blog/medium/focus-on-today-how-we-built-aeman-a-daily-planning-board-for-engineers-on-top-of-github-projects/05.png" alt="The aeman weekly plan" width="1444" height="966" loading="lazy" decoding="async"></p>
<h2 id="reviews-subtasks-and-the-activity-log">Reviews, subtasks, and the activity log</h2>
<p>Sending a task to review creates a linked card for the reviewer, with a feedback loop: while the review is open, the original sits at the Review stage, and once the reviewer takes their own card to 100%, the original is released automatically. Big tasks break down into subtasks whose progress rolls up into the parent. Every action on the board is written to the card’s activity log, so you can reconstruct its whole history after the fact: who moved it, when the progress changed, and why it ended up in a different sprint.</p>
<h2 id="a-day-with-aeman">A day with aeman</h2>
<p>Here’s what the whole flow looks like:</p>
<ol>
<li><strong>Morning sync.</strong> We open the Team board for yesterday, the team lead shares their screen, and we go through the cards: what’s done, what’s stuck, person by person, with the lead fixing statuses as we go. Then the lead hits <strong>Carry over</strong> and everything unfinished moves into today. Then we discuss what’s new: the lead creates cards on the fly and distributes them across people and zones.</li>
<li><strong>The day.</strong> Everyone works from their own Me board. Something urgent lands — a card goes into the yellow zone. Something gets blocked — the Locked stage, and the lead sees it. Progress moves along the way; thoughts and findings go into the day’s notes.</li>
<li><strong>Done</strong> — the card is at 100%, stage Done. If it needs review, <strong>Send to review</strong>, and it shows up for the reviewer.</li>
<li><strong>Tomorrow</strong> it all repeats.</li>
</ol>
<p>Every board is <strong>live</strong>: edits by colleagues and by AI agents appear on everyone’s screen in about a second, with no reload. More on how that works below.</p>
<h2 id="under-the-hood">Under the hood</h2>
<p>Now for the part the engineering crowd came for.</p>
<h2 id="github-projects-v2-as-the-only-storage">GitHub Projects v2 as the only storage</h2>
<p>aeman has no database of its own at all. Every card is an item on a GitHub Projects board, and every field — zone, progress, sprint, weekly plan — is an ordinary project field. aeman even provisions the fields it needs lazily: point it at any empty project and the first change creates whatever is missing. You could say aeman is a specialized view over a GitHub board: the same board opens in GitHub’s native interface, the data is always yours, and you can walk away from aeman at any moment without migrating anything.</p>
<p><img src="https://aenix.io/img/blog/medium/focus-on-today-how-we-built-aeman-a-daily-planning-board-for-engineers-on-top-of-github-projects/06.png" alt="The same board in the GitHub Projects interface" width="1444" height="966" loading="lazy" decoding="async"></p>
<h2 id="a-kubernetes-style-api">A Kubernetes-style API</h2>
<p>The aeman API is deliberately built on the same principles as the Kubernetes API — a pattern that has proven itself for synchronizing distributed state. Each entity — Card, Sprint, Ordering, Presence — is a separate resource with the familiar kind, metadata, and spec.</p>
<p>The client follows the familiar <strong>list + watch</strong> scheme. It first fetches a full snapshot of the board, then opens a WebSocket and starts receiving a stream of ADDED, MODIFIED, and DELETED events. After a full resync the server sends a special Sync frame that signals the local state is completely up to date.</p>
<p>In effect, every open browser tab works like a Kubernetes informer with its own local cache. Any change made by another user or by an AI agent shows up on every open screen in about a second, and your own changes aren’t echoed back to you. GET /api/v1 returns a machine-readable catalog of all available resources and endpoints, so clients and AI agents don’t need to know the shape of the API in advance.</p>
<h2 id="the-github-api-is-slow-what-to-do-about-it">The GitHub API is slow. What to do about it</h2>
<p>The project’s main engineering problem: GitHub’s GraphQL mutations take hundreds of milliseconds, sometimes whole seconds; bursts run into secondary rate limits; and — surprise — GitHub’s read replicas lag behind writes by several seconds. Write straight through and the UI turns into a slide show: move a slider, wait; drag a card, wait.</p>
<p>So writes in aeman use a <strong>write-behind</strong> scheme, and the user never waits for GitHub. Every change is applied to the server cache instantly, becomes visible on all open boards at once, and only then goes to GitHub asynchronously. Writes run through a background queue with rate limiting and automatic retries on transient errors, to stay clear of GitHub API rate limits.</p>
<p>The queue can coalesce consecutive changes, much like the Kubernetes DeltaFIFO. If a user changes the same card field several times in a row, only the final value is sent to GitHub. Card content follows a different rule: text is never lost. If a user edits a description or adds notes in quick succession, the queue merges them into one final version and sends it in a single request. Until the write is confirmed, new notes exist under temporary identifiers, which are then swapped for real ones automatically.</p>
<p>If a write still fails after all the retries, the server notifies every connected client and rereads the board state from GitHub. GitHub remains the single source of truth. On shutdown, aeman first waits for the queue to drain, so that changes already confirmed to the user aren’t lost.</p>
<p>GitHub doesn’t guarantee that a change is available for reading the moment it’s written. Sometimes, right after a successful write, the API keeps returning the old state for a while. Because of that a user could see a freshly moved card jump back, a deleted card reappear, or a new one land in the wrong place. To avoid this, aeman treats its own recent writes as more trustworthy than fresh data from GitHub for a short window. Once GitHub starts returning the current state, the system switches back to it automatically. As a result, the user never sees the interface roll back.</p>
<h2 id="comments-and-the-activity-log">Comments and the activity log</h2>
<p>Notes and change history also live in GitHub, with no separate database. For draft cards the log sits right in the issue body, after a special <code>&lt;!-- aeman:log --&gt;</code> marker. If a card is linked to an existing issue or pull request, notes become ordinary GitHub comments and are visible to everyone in the discussion.</p>
<p>Every action on the board — creating a card, changing progress, moving between sprints, sending to review, and the rest — is recorded by the server as a machine-readable event line. To avoid burying subscribers under dozens of notifications, all those events are collected into a single log comment that gets appended to over time. The log is capped at the last 200 events, and user notes are never deleted.</p>
<p>The result is a complete change history for every card: who moved it and when, how the progress changed, and why it ended up in one sprint or another. That log is useful not only for audit, but also as a data source for statistics and metrics.</p>
<h2 id="a-pluggable-backend">A pluggable backend</h2>
<p>aeman was designed from the start for a pluggable backend. Today cards live in GitHub Projects, but the system itself isn’t tied to GitHub. All interaction with storage is hidden behind a Backend interface, so supporting a new backend comes down to implementing that one interface.</p>
<p>In future iterations I want to add a git backend that stores cards as plain files in the repository. Neither the user interface nor the application’s business logic will have to change.</p>
<p>The packages under pkg/ can also be used as an ordinary Go library, which lets you embed the board engine into your own tool. There’s more on that in docs/embedding.md.</p>
<h2 id="mcp-for-ai-agents">MCP for AI agents</h2>
<p>The same binary can also run as an MCP server. That makes Claude and other AI agents full participants in the process: they create and move cards, leave notes, run carry-over between sprints, and perform other operations.</p>
<p>There’s no separate API for AI and no special privileges. Every action goes through the same domain layer as a human action, with the same contract and the same checks, and the changes show up for everyone immediately.</p>
<p>Besides the local mode, aeman offers a publicly available MCP server. Engineers can connect AI clients without installing anything locally and start working with their tasks directly from Claude Code and other AI agents.</p>
<p><img src="https://aenix.io/img/blog/medium/focus-on-today-how-we-built-aeman-a-daily-planning-board-for-engineers-on-top-of-github-projects/07.png" alt="aeman driven from an AI agent over MCP" width="1444" height="966" loading="lazy" decoding="async"></p>
<h2 id="how-to-try-it">How to try it</h2>
<p>The simplest way to try aeman is to run it locally with your own GitHub token.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">gh auth login <span class="c1"># the project and repo scopes are required</span>
</span></span><span class="line"><span class="cl">git clone https://github.com/aenix-io/aeman
</span></span><span class="line"><span class="cl"><span class="nb">cd</span> aeman
</span></span><span class="line"><span class="cl">make build
</span></span><span class="line"><span class="cl">./aeman serve
</span></span></code></pre></div><p>The browser opens automatically at http://127.0.0.1:8765. You can use any GitHub Project as the board — preferably a new, empty one. aeman creates all the fields it needs on the first change.</p>
<p>For teamwork there’s a multi-user mode. The repository ships a ready-made docker-compose.yml, login through a GitHub OAuth App is supported, and every user works with their own GitHub token. Detailed instructions are in docs/deploy.md.</p>
<h2 id="where-we-landed">Where we landed</h2>
<p>The whole company has been living on aeman for a few weeks now. Product managers and team leads quickly took to the new approach to planning, though — as with any new tool — there was some skepticism. But the main effect turned out to be a different one: in the morning every engineer knows what they’ll be working on today, unplanned work has stopped being invisible, and the daily syncs have become far more constructive and focused.</p>
<p>The project is open source under the Apache-2.0 licence and available on GitHub: <a href="https://github.com/aenix-io/aeman">github.com/aenix-io/aeman</a>.</p>
<p>If you decide to try aeman, I’d be glad to hear any feedback, any problems you run into, and your stories about how planning works where you are.</p>
<hr>
<p><a href="https://blog.aenix.io/focus-on-today-how-we-built-aeman-a-daily-planning-board-for-engineers-on-top-of-github-projects-c59da4451b8b">Focus on today: how we built aeman, a daily planning board for engineers on top of GitHub Projects</a> was originally published in <a href="https://blog.aenix.io">Ænix</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>
]]></content:encoded></item><item><title>Cozystack 1.6: Talos tenant workers, tenant SSO, SecurityGroups, and hierarchical quotas</title><link>https://aenix.io/blog/2026/07/cozystack-1-6-talos-workers-tenant-sso-and-hierarchical-quotas/</link><guid isPermaLink="true">https://aenix.io/blog/2026/07/cozystack-1-6-talos-workers-tenant-sso-and-hierarchical-quotas/</guid><pubDate>Wed, 22 Jul 2026 00:00:00 +0000</pubDate><dc:creator>Timur Tukaev</dc:creator><category>Cozystack</category><category>Kubernetes</category><category>Talos</category><category>Multi-tenancy</category><category>KubeVirt</category><category>Platform Engineering</category><description>Cozystack v1.6.0 moves tenant workers to Talos via Cluster API and adds tenant OIDC, a SecurityGroup firewall API, hierarchical quotas and etcd v1alpha2.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/cozystack-1-6-talos-workers-tenant-sso-and-hierarchical-quotas.jpg" alt="Cozystack 1.6: Talos tenant workers, tenant SSO, SecurityGroups, and hierarchical quotas" width="1200" height="630" loading="lazy" decoding="async"></p>
<p>Cozystack v1.6.0 was published on 22 July 2026. It replaces the Ubuntu and kubeadm bootstrap of tenant Kubernetes workers with Talos Linux driven by Cluster API, completes the etcd-operator <code>v1alpha2</code> migration with in-place adoption of live clusters, adds OIDC single sign-on for tenant kube-apiservers and per-instance Grafana, introduces a tenant-facing <code>SecurityGroup</code> firewall API, and makes tenant resource quotas hierarchical. It rolls up every fix from v1.5.1 and v1.5.2.</p>
<p>It also carries the largest upgrade surface since v1.0. The platform migration <code>targetVersion</code> moves from 45 to 54, so migrations 45 through 53 run as pre-upgrade hooks — and three of them can block or wedge the upgrade if their preconditions are not met. Plan a window and read the checks below.</p>
<h2 id="read-this-before-upgrading">Read this before upgrading</h2>
<p><strong>Upgrade to v1.6.4 (or the latest 1.6.x), not v1.6.0.</strong> Three of the fixes that landed after v1.6.0 are the kind that fail silently:</p>
<ul>
<li>v1.6.0 shipped a <strong>fail-open</strong> copy of <code>hack/seaweedfs-naming-audit.sh</code> — the script whose output gates a runbook step that deletes PVCs. Every <code>kubectl</code> call in it was silenced with <code>2&gt;/dev/null</code>, so a timeout or an RBAC denial produced an empty &ldquo;all clean&rdquo; table byte-identical to an honestly clean fleet. Fixed in v1.6.1. Run the audit from a v1.6.1-or-later checkout, and read the <strong>exit code</strong>, not the table. A clean result from the v1.6.0 copy is not evidence of anything.</li>
<li>v1.6.2 fixes Velero CRDs staying frozen at whatever version was first installed, because Helm never touches a chart&rsquo;s <code>crds/</code> directory on upgrade. Once the Velero image moved to a version with new backup phases, the apiserver rejected phase transitions against the stale CRDs and <strong>backups stopped while the HelmRelease stayed green</strong>.</li>
<li>v1.6.2 also repairs the default backup <code>Strategy</code> CRs and the Velero <code>BackupStorageLocation</code>, which were gated on a Helm <code>lookup</code> performed while the referenced object was still being created. When the lookup came back empty the objects were skipped permanently, since helm-controller does not re-render a release whose chart and values are unchanged.</li>
</ul>
<p>If you are coming from the 1.5 line, <strong>go through v1.5.4 first</strong>. v1.5.4 is the first 1.5.x release stamped <code>targetVersion: 46</code>, and its slot 45 stamps <code>helm.sh/resource-policy: keep</code> onto the CAPI <code>KubeadmConfigTemplate</code> objects. 1.6 drops <code>KubeadmConfigTemplate</code> from the tenant <code>kubernetes</code> chart entirely — workers move to <code>TalosConfigTemplate</code> — so without that pin Helm deletes the template while the kubeadm-backed MachineSet is still mid-rollover with its <code>bootstrap.configRef</code> pointing at it. Verify before upgrading:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">kubectl get kubeadmconfigtemplates.bootstrap.cluster.x-k8s.io -A <span class="se">\
</span></span></span><span class="line"><span class="cl">  -o custom-columns<span class="o">=</span><span class="s1">&#39;NS:.metadata.namespace,NAME:.metadata.name,KEEP:.metadata.annotations.helm\.sh/resource-policy&#39;</span>
</span></span></code></pre></div><p>A row showing <code>&lt;none&gt;</code> means the pin did not land. Fix it before going to 1.6.</p>
<h3 id="pre-upgrade-checks">Pre-upgrade checks</h3>
<p>Run these against the management cluster before applying the v1.6 Platform Package.</p>
<p><strong>1. etcd adoption needs a reachable backup target.</strong> Migration 50 adopts every legacy <code>etcd.aenix.io/v1alpha1</code> cluster onto the new operator and takes a mandatory snapshot first. If it cannot resolve the snapshot target it exits 1 and halts the upgrade.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">kubectl get etcdclusters.etcd.aenix.io -A
</span></span><span class="line"><span class="cl">kubectl get etcds.strategy.backups.cozystack.io cozy-default-etcd
</span></span><span class="line"><span class="cl">kubectl get secret cozy-backups-creds -n cozy-velero
</span></span><span class="line"><span class="cl">kubectl get buckets.apps.cozystack.io cozy-backups -n tenant-root
</span></span></code></pre></div><p>If the first command prints nothing, migration 50 is a no-op.</p>
<p><strong>2. SeaweedFS naming audit.</strong> v1.6.0 pins <code>fullnameOverride: seaweedfs</code> and adopts the running workload set in place, but two states cannot be adopted automatically and the chart fails the render rather than guess: class <code>S</code> (installed fresh on 1.5.x, data on <code>data1-seaweedfs-system-volume-*</code> PVCs) and class <code>MIXED</code> (both naming generations present). Both must be recovered before upgrading, following the <code>seaweedfs-431-rename-recovery</code> runbook. Class <code>L</code> needs nothing. A cluster that went 1.4.x straight to 1.6 never renamed and is unaffected.</p>
<p><strong>3. Tenant clusters still on Kubernetes v1.30.</strong> v1.30 leaves the Talos-to-Kubernetes support matrix and the chart refuses to render it. Migration 46 patches live CRs to v1.31, but a GitOps-managed CR is overwritten again by the next source reconcile — so bump <code>spec.version</code> in Git.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">kubectl get kuberneteses.apps.cozystack.io -A <span class="se">\
</span></span></span><span class="line"><span class="cl">  -o custom-columns<span class="o">=</span><span class="s1">&#39;NS:.metadata.namespace,NAME:.metadata.name,VERSION:.spec.version&#39;</span>
</span></span></code></pre></div><p><strong>4. Hand-made tenant StorageClasses that collide.</strong> Remote-accessible LINSTOR StorageClasses are now created inside each tenant cluster under the same name. A manually created tenant StorageClass with a colliding name — typically <code>replicated</code> — blocks the propagated class and stalls the tenant CSI release. Delete any such class that is not Helm-managed. Infra classes that must stay node-local need an explicit <code>linstor.csi.linbit.com/allowRemoteVolumeAccess: &quot;false&quot;</code>; an absent annotation is treated as remote-accessible.</p>
<p><strong>5. Deprecated etcd module backup values.</strong> The <code>backup.*</code> block on the <code>etcd</code> tenant module is removed in this release. Move to a <code>BackupClass</code> bound to the <code>Etcd</code> strategy first.</p>
<p><strong>6. Coming from v1.4.x?</strong> The v1.5.0 requirement still applies: Kubernetes 1.33 or newer on the management cluster and on any tenant cluster enabling the Flux addon.</p>
<h2 id="talos-linux-for-tenant-kubernetes-workers">Talos Linux for tenant Kubernetes workers</h2>
<p>Tenant worker nodes no longer boot Ubuntu and bootstrap through kubeadm. Phase 1 of the Kubernetes-app split replaces that path with Talos Linux, driven by <code>cluster-api-bootstrap-provider-talos</code> (CABPT v0.6.12) registered as a second bootstrap provider alongside kubeadm, plus a <code>clastix/talos-csr-signer</code> sidecar embedded in the Kamaji control-plane pod.</p>
<p>Talos PKI — an Ed25519 CA plus trustd TLS — is generated through cert-manager with stable, lookup-and-reuse Talos secrets. The Kamaji control-plane provider carries an upstream-bound patch exposing <code>KamajiControlPlane.spec.network.additionalServicePorts</code>, so trustd (50001/TCP) can be published on the apiserver Service. Workers boot the Talos openstack raw image via a CDI <code>DataVolume</code> streamed from the image factory, with the system disk exposed as virtio-blk using <code>blockSize.custom: logical=512, physical=4096</code> — so 4Ki-native backends such as LINSTOR/DRBD behave under QEMU&rsquo;s <code>O_DIRECT</code> writes while SeaBIOS still boots.</p>
<p>The separate <code>disk-kubelet</code> PVC is gone. Talos lays out <code>EPHEMERAL</code> itself on the single system disk that <code>nodeGroups[*].diskSize</code> now sizes.</p>
<p>What this means operationally:</p>
<ul>
<li><strong>Existing tenants roll over automatically.</strong> Old machines are replaced by Talos workers without manual intervention, and migration 45 pins the outgoing <code>KubeadmConfigTemplate</code> objects with <code>helm.sh/resource-policy: keep</code> so Helm cannot prune them mid-rollover. But plan for a full worker-pool replacement per tenant cluster: worker disks are reprovisioned and container images re-pulled.</li>
<li><strong>Worker MachineHealthCheck remediation is now ON by default.</strong> <code>maxUnhealthy</code> moved from a hard-coded <code>0</code> to <code>nodeHealthCheck.maxUnhealthy</code>, defaulting to <code>&quot;50%&quot;</code>. CAPI now deletes and replaces unhealthy worker Machines. Set <code>nodeHealthCheck.maxUnhealthy: &quot;0%&quot;</code> to keep the old behaviour until your fleet is stable on Talos. Per-node-group <code>maxUnhealthy</code> and <code>nodeStartupTimeout</code> overrides also ship, so a stateful group can sit at <code>0%</code> while a stateless one tolerates <code>50%</code>.</li>
<li><strong>The default <code>md0</code> node group is no longer merged into every cluster.</strong> <code>nodeGroups</code> defaults to <code>{}</code> and the built-in <code>md0</code> applies only when no node groups are configured. Migration 47 pins <code>md0</code> explicitly on existing CRs to preserve live topology; it is fail-closed, aborting the upgrade rather than letting Helm prune a live <code>md0</code> MachineDeployment.</li>
<li><strong>Fresh clusters with <code>nodeGroups: {}</code> come up with zero workers.</strong> The chart no longer manages <code>MachineDeployment.spec.replicas</code> — the cluster-autoscaler owns it alone, seeded from <code>minReplicas: 0</code>. Supply a node group with <code>roles: [ingress-nginx]</code> and <code>minReplicas &gt;= 1</code>, or let the autoscaler bring up <code>md0</code> once ingress-nginx pods go <code>Pending</code>. The upside for existing clusters: <code>helm upgrade</code> no longer drains workers back to a hardcoded <code>replicas: 2</code> on every platform bump.</li>
<li><strong>Air-gapped installations lose the <code>registries.mirrors</code> passthrough</strong> to tenant workers — the Helm-rendered <code>*-patch-containerd</code> Secret has no consumer in the Talos machine config and was removed. The Talos OS image and installer are overridable via <code>talos.imageFactoryURL</code> and <code>talos.installerRepository</code>, but in-guest registry mirroring is a Phase 2 follow-up.</li>
<li><strong>Worker node disks now default to the application-level <code>replicated</code> StorageClass.</strong> A node group leaving <code>storageClass</code> empty previously took the management-cluster default; it now falls back to the application <code>storageClass</code>, because linstor-csi v1.11.2 rejects ReadWriteMany volumes on a non-DRBD class and worker VMs need RWX to live-migrate. The rendered worker template changes, so each tenant&rsquo;s worker MachineDeployment rolls once.</li>
</ul>
<p>The tenant <code>kubernetes</code> HelmRelease now reports Ready as soon as Helm completes (<code>DisableWait</code>) and no longer blocks on worker or addon readiness; worker-rollout health is tracked by <code>WorkloadMonitor</code> instead.</p>
<h2 id="etcd-operator-v1alpha2-with-in-place-adoption">etcd-operator v1alpha2 with in-place adoption</h2>
<p>The etcd stack moves to the donated <code>etcd-operator.cozystack.io/v1alpha2</code> API — a Membership-API lifecycle replacing the StatefulSet model — served by a Cozystack-authored chart at appVersion v0.5.2, with CRDs split into their own <code>etcd-operator-crds</code> package so they can be installed ahead of the controller.</p>
<p>The interesting part is the upgrade. Migration 50 runs as a pre-upgrade hook, while the legacy operator is still the running one, and uses the <code>etcd-migrate</code> tool baked into the migrations image to rewrite ownership, labels and CRs so the new operator takes over the live data plane on its first reconcile — no data move, no pod restart. Before touching anything it takes a mandatory snapshot of every adopted cluster to the <code>cozy-backups</code> bucket under a system key prefix with tenant-invisible credentials, and it fails loudly rather than adopt without one. It also re-issues each cluster&rsquo;s server and peer certificates with the new operator&rsquo;s native member wildcard SAN, because in <code>secretRef</code> TLS mode the operator never mints certs — an adopted cluster whose certs lack the native domain would silently fail TLS the first time a member is replaced, while still reporting <code>Available=True</code>.</p>
<p>If migration 50 halts on a cluster that genuinely has no backup storage, the documented escape hatch is:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">kubectl edit package.cozystack.io cozystack.cozystack-platform
</span></span></code></pre></div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">spec</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">components</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">platform</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">values</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span><span class="nt">migrations</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">          </span><span class="nt">etcdAdoptSkipBackup</span><span class="p">:</span><span class="w"> </span><span class="kc">true</span><span class="w">
</span></span></span></code></pre></div><p>Flux re-renders the platform chart and re-runs the hook, adopting every legacy etcd without a snapshot. <strong>A botched adoption on a tenant control-plane etcd is unrecoverable without a snapshot.</strong> Set the flag back to <code>false</code> afterwards, or the next etcd migration skips its snapshot too.</p>
<h2 id="securitygroup-tenant-managed-network-policy">SecurityGroup: tenant-managed network policy</h2>
<p>A new tenant-facing, namespace-scoped firewall resource — <code>sdn.cozystack.io/v1alpha1 SecurityGroup</code> — lets tenants manage their applications&rsquo; network policy without any access to the <code>cilium.io</code> API group.</p>
<p>A <code>SecurityGroup</code> is a membership group. It owns a membership label (<code>securitygroup.sdn.cozystack.io/&lt;name&gt;</code>) that a new <code>securitygroup-controller</code> stamps onto the pods of every managed application listed in <code>spec.attachments</code>, and removes when an attachment is dropped. The group projects 1:1 onto a <code>CiliumNetworkPolicy</code> in the same namespace whose <code>endpointSelector</code> is exactly that membership label — so one <code>SecurityGroup</code> can cover several applications at once, and <code>fromSG</code> / <code>toSG</code> peers resolve group-to-group references live in the Cilium dataplane. A finalizer guards the backing policy so member labels are always stripped before it is removed, and the aggregated-API REST storage re-asserts them on every write so a full-replace <code>PUT</code> cannot orphan them.</p>
<p>The important limitation: ingress and egress rules only ever <strong>add</strong> allowed traffic. An empty rule list does not isolate the member pods, because effective connectivity remains the union of every policy selecting a pod — including the platform&rsquo;s blanket-allow baseline. Default-deny enforcement is tracked separately as future work. Do not treat a <code>SecurityGroup</code> as a segmentation boundary in this release.</p>
<h2 id="hierarchical-tenant-resource-quotas">Hierarchical tenant resource quotas</h2>
<p><code>tenant.spec.resourceQuotas</code> previously became a plain per-namespace <code>ResourceQuota</code> that ignored the tenant tree entirely: a super-admin inside a quota&rsquo;d tenant could create a sub-tenant with a larger quota, or with no quota at all, and consume more than was allocated. A tenant&rsquo;s declared quota is now the budget for its whole sub-tree.</p>
<p>Enforcement is two cooperating layers, mirroring OpenShift&rsquo;s ClusterResourceQuota split. A declaration-time gate in the aggregated apiserver deterministically rejects a sub-tenant whose declared quota exceeds the parent&rsquo;s remaining budget. A runtime enforcer in <code>cozystack-controller</code> maintains a per-namespace <code>tenant-quota-allocated</code> ResourceQuota that clamps each pool member to its share, aggregating sub-tree usage so members collectively stay within it — Kubernetes enforces the most restrictive ResourceQuota in a namespace, so this binds without fighting Flux over the chart-rendered <code>tenant-quota</code>.</p>
<p>Review sub-tenant quotas before upgrading. A parent quota lowered below the sum of its children emits a <code>QuotaOvercommitted</code> event, and <code>--tenant-quota-buffer-percent</code> on <code>cozystack-controller</code> temporarily inflates pool budgets during rollout so workloads already over a freshly-binding quota keep running. There are no API type changes and no tenant chart changes — the feature works with the existing field.</p>
<h2 id="oidc-single-sign-on-for-tenant-kubernetes-and-grafana">OIDC single sign-on for tenant Kubernetes and Grafana</h2>
<p>Two symmetric Phase 1 features let a tenant opt an individual workload into the platform&rsquo;s Keycloak <code>cozy</code> realm through one flat selector, with no platform-level configuration.</p>
<p><strong>Tenant kube-apiserver.</strong> Each <code>Kubernetes</code> CR gains <code>spec.oidc.mode: System | CustomConfig | None</code>, defaulting to <code>None</code>. <code>System</code> trusts the <code>cozy</code> realm via a per-cluster public client and audience binding; <code>CustomConfig</code> accepts a tenant-supplied structured <code>AuthenticationConfiguration</code>. <code>spec.oidc.users[]</code> drives one ClusterRoleBinding per user inside the tenant cluster (<code>admin</code> maps to <code>cluster-admin</code>, <code>view</code> to <code>view</code>). In <code>System</code> mode a <code>&lt;release&gt;-oidc-kubeconfig</code> Secret carrying a ready-to-use <code>kubectl oidc-login</code> exec block is surfaced through the Cozystack Dashboard, so a tenant user gets from &ldquo;cluster exists&rdquo; to &ldquo;kubectl works with my SSO identity&rdquo; without an operator in the loop. The underlying <code>controlPlane.apiServer.extraArgs</code>, <code>extraVolumes</code> and <code>extraVolumeMounts</code> passthrough on <code>KamajiControlPlane</code> is also exposed for anyone hand-rolling authentication.</p>
<p><strong>Grafana.</strong> Each <code>Monitoring</code> CR gains the same selector. <code>System</code> wires a per-instance confidential Keycloak client with audience binding; authorization is app-side, with <code>spec.oidc.users: [{email, role: Admin|Editor|Viewer}]</code> reconciled into Grafana&rsquo;s Main Org by a chart-owned post-install and post-upgrade Job that pre-provisions users, adds them to the org, patches their role and prunes stale members. <code>CustomConfig</code> accepts a tenant-supplied <code>[auth.generic_oauth]</code> payload, inline as a map or as a Secret with an <code>auth.ini</code> key — in the Secret case <code>spec.oidc.users</code> is unsupported and the chart fails the render on that combination rather than silently ignoring it. The <code>admin_user</code> / <code>admin_password</code> Secret remains a documented break-glass path in every mode.</p>
<p>Reference: <a href="https://cozystack.io/docs/v1.6/operations/oidc/enable_oidc/">enabling OIDC</a>.</p>
<h2 id="application-deletion-now-reclaims-storage">Application deletion now reclaims storage</h2>
<p>Deleting a managed application used to leave its volumes and operator-generated Secrets behind. A ten-PR sweep across the catalog adds post-delete cleanup hooks and PVC retention policies so a deleted application actually releases what it held: ClickHouse keeper, data and log PVCs; Qdrant&rsquo;s data PVC; OpenBAO&rsquo;s data PVC; the tenant monitoring module&rsquo;s VictoriaMetrics and VictoriaLogs storage and TLS Secrets; the tenant seaweedfs module&rsquo;s volume PVCs; the tenant etcd module&rsquo;s <code>data-etcd-*</code> PVCs; Harbor&rsquo;s jobservice and trivy PVCs; MariaDB&rsquo;s operator-generated password Secrets; Bucket and Gateway ACME resources.</p>
<p><strong>This makes deletion irreversible.</strong> That is the point of the change, but it is a real behaviour change from &ldquo;delete leaks the volume, and you could recover from it&rdquo;. Harbor&rsquo;s PVCs are force-deleted, overriding <code>persistence.resourcePolicy=keep</code>; Harbor&rsquo;s CNPG database PVCs are not touched. Migrations 48 and 51 backfill the release label onto pre-existing ClickHouse keeper and monitoring storage PVCs so already-deployed clusters are covered; both are best-effort, and the worst case if a relabel is skipped is the pre-existing leak.</p>
<p>Back up before deleting. Tell your tenants.</p>
<h2 id="also-in-v160">Also in v1.6.0</h2>
<ul>
<li><strong>Wildcard certificates, end to end.</strong> v1.5&rsquo;s <code>publishing.certificates.wildcardSecretName</code> covered the root tenant only. The platform controller now replicates that certificate into every tenant namespace that terminates TLS, so per-tenant ingress controllers and Gateways serve it automatically, with no cross-namespace Secret read. A new opt-in <code>publishing.certificates.wildcard</code> (default <code>false</code>) has the platform mint a single <code>*.&lt;root-host&gt;</code> certificate via DNS-01 on the default ingress-nginx path — the practical fix for hitting Let&rsquo;s Encrypt rate limits at scale.</li>
<li><strong>Keycloak</strong> gains four independent opt-ins: a KMS-encrypting database proxy for column-level PII encryption at rest, backed by a static KEK or Vault Transit (with Vault Kubernetes and AppRole auth); an <code>ingress.adminHost</code> that serves the admin console and Administration REST API on a separate hostname and route, attachable to a private Gateway or ingressClass via <code>publishing.ingressNameAdmin</code>; Barman-based S3 backups of the CNPG database; and a selectable realm login theme shipped through platform <code>branding</code> values. All are off by default.</li>
<li><strong>Immutable tags and rc-to-stable promotion.</strong> A stable release is now the byte-identical promotion of the release candidate that was tested: no tag is ever force-moved, stable is never rebuilt, and images are retagged by digest. The cron-driven <code>auto-release.yaml</code> patch-tag workflow is deleted outright.</li>
<li><strong>Kamaji defaults to 2 replicas</strong> with soft pod anti-affinity, and its telemetry handler and webhook are removed — cutting apiserver admission p99 for <code>TenantControlPlane</code> webhooks by roughly 70% on multi-tenant clusters.</li>
<li><strong>Remote-accessible LINSTOR StorageClasses propagate into tenant clusters</strong> under the same name, with the class named by <code>storageClass</code> (default <code>replicated</code>) becoming the tenant default and the legacy <code>kubevirt</code> class kept as an alias.</li>
<li>The <strong>Cozystack Dashboard</strong> now shows the LoadBalancer external IP on the application Services tab, and renders explicit error and unknown-type states instead of an infinite spinner. The console SPA is now built from in-tree source; the shipped image and digest are unchanged.</li>
<li>After upgrading, <strong>re-check metrics scrape targets on port 10250</strong> — cert-manager&rsquo;s webhook, external-secrets&rsquo; webhook and metrics-server&rsquo;s listener all moved off it.</li>
</ul>
<h2 id="platform-components">Platform components</h2>
<table>
  <thead>
      <tr>
          <th>Component</th>
          <th>Change</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Talos Linux</td>
          <td>v1.13.0 → v1.13.6, kernel 6.18.29 → 6.18.38; closes CVE-2026-53359 and CVE-2026-46113 KVM guest-to-host escapes</td>
      </tr>
      <tr>
          <td>etcd-operator</td>
          <td>v0.4.5 → v0.5.2, new <code>v1alpha2</code> API</td>
      </tr>
      <tr>
          <td>Cilium</td>
          <td>1.19.3 → 1.19.5</td>
      </tr>
      <tr>
          <td>KubeVirt</td>
          <td>v1.8.4 (fixes VMIs stuck <code>Scheduled</code> on Kubernetes 1.36; closes CVE-2026-35469)</td>
      </tr>
      <tr>
          <td>Velero</td>
          <td>1.17.0 → 1.18.1, concurrent backup processing, cache volumes for data movers</td>
      </tr>
      <tr>
          <td>Vertical Pod Autoscaler</td>
          <td>1.3.0 → 1.5.0, in-place pod resizing promoted to Beta</td>
      </tr>
      <tr>
          <td>Harbor</td>
          <td>2.14.2 → 2.15.1</td>
      </tr>
      <tr>
          <td>Keycloak</td>
          <td>26.5.2 → 26.6.3</td>
      </tr>
      <tr>
          <td>LINSTOR</td>
          <td>1.33.2 → 1.33.3, linstor-csi v1.11.2</td>
      </tr>
      <tr>
          <td>FoundationDB operator</td>
          <td>v2.13.0 → v2.30.0</td>
      </tr>
      <tr>
          <td>HAMi</td>
          <td>2.8.1 → 2.9.0</td>
      </tr>
      <tr>
          <td>Percona MongoDB operator</td>
          <td>1.21.1 → 1.22.0</td>
      </tr>
      <tr>
          <td>OpenBAO</td>
          <td>v2.5.0 → v2.5.1</td>
      </tr>
      <tr>
          <td>csi-driver-nfs</td>
          <td>4.11.0 → 4.13.3</td>
      </tr>
      <tr>
          <td>CoreDNS chart</td>
          <td>1.43.2 → 1.46.0</td>
      </tr>
      <tr>
          <td>OpenCost</td>
          <td>1.111.0 → 1.120.3</td>
      </tr>
  </tbody>
</table>
<p>Talos system extensions are refreshed to the 20260622 set, including DRBD 9.3.2 and ZFS 2.4.3. Applying that alone rolls each tenant worker pool once, since the worker template is renamed to carry the new image. Managed Kubernetes patch versions move to v1.32.13, v1.33.13, v1.34.9 and v1.35.6; v1.30 is out of the tenant support matrix.</p>
<p>Two API groups are worth knowing about as non-events. <code>SecurityGroup</code> had <code>spec.targetRef</code> replaced by the <code>spec.attachments[]</code> membership model before release, so only people tracking <code>main</code> between 30 June and 16 July 2026 need to rewrite objects. The <code>network.cozystack.io</code> group (<code>ExposureClass</code> / <code>ServiceExposure</code>) was introduced and removed within the cycle and is <strong>not</strong> part of v1.6.0 — native <code>Service</code> <code>type: LoadBalancer</code> plus <code>loadBalancerClass</code> covers the same ground, exposed as <code>publishing.loadBalancerClass</code>.</p>
<h2 id="the-16-patch-line">The 1.6 patch line</h2>
<ul>
<li><strong>v1.6.1</strong> (5 August 2026): CNPG operator and CRDs aligned to 1.28.2, fixing a PVC resize deadlock that could leave a single-instance PostgreSQL cluster with zero instances; the <code>talos-reconcile</code> Job now renders for the implicit <code>md0</code> group, so the first autoscaler-driven scale-up no longer leaves Machines permanently blocked with no matching <code>TalosConfigTemplate</code>; the <code>keycloak-configure</code> pre-delete Job patches the HelmRelease in the right namespace, unblocking Keycloak uninstall and reinstall; the SeaweedFS naming audit fails closed; etcd-operator to v0.5.4.</li>
<li><strong>v1.6.2</strong> (19 August 2026): the backup-strategy lookup gate, the Velero CRD upgrade, and the kube-ovn webhook certificate reload — that last one mattering because <code>kube-ovn-webhook</code> loaded its serving certificate once at startup and never re-read it, so after a cert-manager renewal every pod creation in tenant namespaces was rejected under <code>failurePolicy: Fail</code>, including <code>virt-launcher</code> pods, blocking VMI startup. Also: barman-cloud requests an S3 checksum only when required, so backups to Ceph RGW and some MinIO and R2 builds stop failing outright; sharded helm-controllers no longer crashloop behind an HTTP proxy; <code>kubectl apply --validate</code> works again against the <code>core</code> and <code>sdn</code> API groups.</li>
<li>Later patch releases v1.6.3 and v1.6.4 followed. Use the latest 1.6.x; see the <a href="https://github.com/cozystack/cozystack/releases">Cozystack releases</a>.</li>
</ul>
<h2 id="upgrading">Upgrading</h2>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">kubectl annotate namespace cozy-system helm.sh/resource-policy<span class="o">=</span>keep --overwrite
</span></span><span class="line"><span class="cl">kubectl annotate configmap -n cozy-system cozystack-version helm.sh/resource-policy<span class="o">=</span>keep --overwrite
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">helm upgrade cozystack oci://ghcr.io/cozystack/cozystack/cozy-installer <span class="se">\
</span></span></span><span class="line"><span class="cl">  --version 1.6.4 <span class="se">\
</span></span></span><span class="line"><span class="cl">  --namespace cozy-system
</span></span></code></pre></div><p>The annotations are required — without them, removing or upgrading the installer release can delete the <code>cozy-system</code> namespace and everything in it. Reuse your existing values, watch the migration hooks, and expect the rollouts listed above: tenant worker pools, the dashboard and linstor-gui gatekeepers, the VPA admission controller, and the linstor-scheduler admission Deployment. Full procedure: <a href="https://cozystack.io/docs/v1.6/operations/cluster/upgrade/">upgrade guide</a>.</p>
<h2 id="where-ænix-fits">Where Ænix fits</h2>
<p>Cozystack is a CNCF Sandbox project under Apache 2.0. v1.6 is a good release and a demanding upgrade — the Talos worker rollover, the etcd adoption and the deletion-semantics change all deserve a rehearsal on a non-production cluster first. Ænix created Cozystack and is one of its maintainers; it sells <a href="https://aenix.io/products/cozystack-enterprise-support/">enterprise support for Cozystack</a>, including upgrade planning and hands-on migration for teams running it at scale.</p>
<h2 id="release-links">Release links</h2>
<ul>
<li><a href="https://github.com/cozystack/cozystack/releases/tag/v1.6.0">Cozystack v1.6.0 on GitHub</a></li>
<li><a href="https://github.com/cozystack/cozystack/releases/tag/v1.6.2">Cozystack v1.6.2 on GitHub</a></li>
<li><a href="https://github.com/cozystack/cozystack/releases">All Cozystack releases on GitHub</a></li>
<li><a href="https://cozystack.io/docs/v1.6/">Cozystack v1.6 documentation</a></li>
<li><a href="https://t.me/cozystack">Telegram</a> and <a href="https://kubernetes.slack.com/archives/C06L3CPRVN1">Slack</a> (invite at <a href="https://slack.kubernetes.io/">slack.kubernetes.io</a>)</li>
</ul>
]]></content:encoded></item><item><title>Cozystack 1.5: Gateway API, default backups, Flux sharding, and TLS for managed services</title><link>https://aenix.io/blog/2026/06/cozystack-1-5-gateway-api-default-backups-and-tls-for-managed-services/</link><guid isPermaLink="true">https://aenix.io/blog/2026/06/cozystack-1-5-gateway-api-default-backups-and-tls-for-managed-services/</guid><pubDate>Mon, 22 Jun 2026 00:00:00 +0000</pubDate><dc:creator>Timur Tukaev</dc:creator><category>Cozystack</category><category>Kubernetes</category><category>Cilium</category><category>KubeVirt</category><category>GPU</category><category>Platform Engineering</category><description>Cozystack v1.5.0 adds opt-in Gateway API via Cilium, a default BackupClass, Flux v2.8 with sharding, TLS for managed databases, and GPU passthrough.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/cozystack-1-5-gateway-api-default-backups-and-tls-for-managed-services.jpg" alt="Cozystack 1.5: Gateway API, default backups, Flux sharding, and TLS for managed services" width="1200" height="630" loading="lazy" decoding="async"></p>
<p>Cozystack v1.5.0 was published on 22 June 2026. It rolls up every fix from the v1.4.1 to v1.4.4 patch line and pushes the platform in five directions: a second ingress path via Gateway API, backups that work without per-app S3 configuration, stricter and shardable Flux reconciliation, TLS on externally published managed services, and GPU passthrough that no longer needs a manual <code>kubectl patch</code>.</p>
<p>This release also has a real upgrade surface. Read the next section before you plan the window.</p>
<h2 id="read-this-before-upgrading">Read this before upgrading</h2>
<p><strong>Do not install v1.5.0, v1.5.1 or v1.5.2 today. Go to v1.5.4.</strong> The v1.5.0 SeaweedFS chart bump (4.05 to 4.31) renamed SeaweedFS workloads from chart-based names to release-based names. StatefulSet names are immutable, so Helm could not rename in place — it stood up a second, duplicate set beside the running one. On a cluster with more nodes than master replicas the empty duplicate schedules successfully, both generations carry identical pod labels, and the <code>seaweedfs-s3</code> Service load-balances across two filers writing into the same metadata store. That is a data-integrity incident, not a cosmetic duplicate. A related defect in migration 43 could prune the <code>seaweedfs-db</code> CNPG cluster — the index for every object in a tenant&rsquo;s S3 — for any instance not literally named <code>seaweedfs</code>.</p>
<p>Both are fixed in v1.5.4 (19 August 2026), which pins <code>fullnameOverride: seaweedfs</code>, adopts the running workloads in place, and ships a fail-closed audit script. v1.5.4 is the final release of the 1.5 line and the only one worth deploying.</p>
<p>Five further changes need attention on any 1.5 upgrade:</p>
<ul>
<li><strong>Kubernetes 1.33 or newer is required on the management cluster</strong>, and on any tenant cluster that enables the Flux addon. Flux v2.8&rsquo;s helm-controller v1.5 requires it. Upgrade Kubernetes first.</li>
<li><strong><code>upgrade.force: true</code> is gone.</strong> Immutable-field changes no longer self-heal. If a chart upgrade changes a StatefulSet <code>volumeClaimTemplates</code> or <code>serviceName</code>, the apply fails; recreate the object by hand with <code>kubectl delete sts &lt;name&gt; --cascade=orphan</code> and let Flux reconcile.</li>
<li><strong>GPU VM operators must move custom host devices first.</strong> With <code>cozystack.gpu-operator</code> enabled, the bundle now owns <code>KubeVirt.spec.configuration.permittedHostDevices</code> and overwrites it on the first reconcile. Move hand-edited entries into <code>.gpu.permittedHostDevices</code> before upgrading and confirm every <code>resourceName</code> matches what the nodes advertise.</li>
<li><strong>MetalLB moves to the FRR-K8s BGP backend and HTTPS-only metrics.</strong> MetalLB v0.15.2 to v0.16.1 adopts the upstream-default FRR-K8s backend; classic FRR mode is deprecated. <code>kube-rbac-proxy</code> is replaced by native TLS plus RBAC, so any scrape config pointing at the old plain-HTTP metrics endpoints must be updated. The host-network port denylist is rotated to match, including the new <code>/healthz</code> and <code>/readyz</code> port 17472.</li>
<li><strong>Externally published databases and messaging gain TLS automatically.</strong> Instances with <code>external: true</code> flip to TLS-on after upgrade. The trust anchor is a self-signed CA, so external clients must retrieve and pin it. Cluster-internal instances are unaffected.</li>
</ul>
<p>One more, if you plan to go on to 1.6 later: v1.5.4 is the first 1.5.x release stamped migration <code>targetVersion: 46</code>, and the pin it applies to the CAPI <code>KubeadmConfigTemplate</code> objects is what keeps a later 1.6 upgrade from pruning them out from under a live, mid-rollover MachineSet. Reaching 1.6 from anything older on the 1.5 line is the path that hurts.</p>
<h2 id="gateway-api-via-cilium">Gateway API via Cilium</h2>
<p>Cozystack-native services can now be published through the Gateway API backed by Cilium, as an opt-in alternative to the per-tenant ingress-nginx controllers. It is materialised per tenant by a new <code>gateway.cozystack.io/v1alpha1 TenantGateway</code> CRD reconciled by <code>cozystack-controller</code>.</p>
<p>Enable it platform-wide with <code>publishing.gateway.enabled=true</code>. A tenant then either gets its own Gateway, LoadBalancer IP and certificate with <code>tenant.spec.gateway=true</code>, or inherits the nearest ancestor&rsquo;s Gateway through the same label-based selector model that already drives ingress inheritance. Two certificate solver modes ship: HTTP-01 (the default — a per-app certificate, no platform configuration for new apps) and DNS-01 (opt-in — one wildcard certificate covering an apex, with Cloudflare, Route 53, DigitalOcean and RFC 2136 providers).</p>
<p>Defaults stay on ingress-nginx, so existing clusters do not change behaviour. Two side effects do land on everyone:</p>
<ul>
<li>Cilium Envoy and Gateway API support are now always enabled. That is an extra <code>cilium-envoy</code> DaemonSet, roughly 100 MB RAM per node at idle. Budget for it on dense nodes.</li>
<li><code>cozystack-api</code> now invokes admission (<code>createValidation</code> / <code>deleteValidation</code>) on Create <strong>and</strong> Delete for <code>apps.cozystack.io/*</code>. Any custom ValidatingAdmissionPolicy or webhook on those kinds now fires on all three verbs.</li>
</ul>
<p>Reference: <a href="https://cozystack.io/docs/v1.5/networking/gateway-api/">Gateway API guide</a>.</p>
<h2 id="backups-that-work-out-of-the-box">Backups that work out of the box</h2>
<p>Previous releases installed the backup machinery. v1.5 closes the gap between &ldquo;installed&rdquo; and &ldquo;working without per-app S3 configuration&rdquo;.</p>
<p>A platform-managed default <code>BackupClass</code> named <code>cozy-default</code> now ships, backed by a system bucket <code>cozy-backups</code>. Applications opt in with <code>useSystemBucket</code>, after which the platform projects shared backup credentials into the tenant namespace with RBAC isolation and projection metrics, and skips per-release credential Secrets entirely. Default strategies are provided for every backup-capable app — Velero for VMDisk and VMInstance, CNPG for PostgreSQL, Altinity for ClickHouse, plus MariaDB, FoundationDB and etcd — and a Velero <code>BackupStorageLocation</code> is wired to the system bucket. The legacy per-tenant S3 fields on Postgres and ClickHouse are deprecated in favour of this flow.</p>
<p><strong>Velero is now a default system package</strong>, not an optional one. This fixes a deterministic failure: the default <code>backupstrategy-controller</code> hard-depends on Velero, so on clusters without it the controller sat in <code>DependenciesNotReady</code> and kept the platform HelmRelease from ever reaching Ready. Existing clusters get Velero in the <code>cozy-velero</code> namespace on upgrade. If you do not back up VMs, opt out via <code>bundles.disabledPackages</code>.</p>
<p>Two new strategies join the catalog. An <strong>etcd</strong> strategy (cluster-scoped <code>strategy.backups.cozystack.io Etcd</code>, S3-only) with a snapshot BackupJob and a destructive in-place RestoreJob. And a generic, application-agnostic <strong>Job</strong> strategy: the operator supplies a Kubernetes Job template, Cozystack renders and runs it as a one-shot backup, then re-renders it with <code>.Mode == &quot;restore&quot;</code> for recovery. Cross-namespace restore is not supported by the Job strategy.</p>
<p>References: <a href="https://cozystack.io/docs/v1.5/applications/backup-and-recovery/">application backup and recovery</a>, <a href="https://cozystack.io/docs/v1.5/operations/services/managed-app-backup-configuration/">managed app backup configuration</a>.</p>
<h2 id="flux-v28-and-helm-controller-sharding">Flux v2.8 and helm-controller sharding</h2>
<p>Flux moves from v2.7.3 to v2.8.0 across both the embedded management-cluster Flux and the optional tenant Flux addon; the flux-operator and flux-instance charts go v0.33.0 to v0.50.0.</p>
<p>helm-controller v1.5 ships Server-Side Apply with <code>--force-conflicts</code> and kstatus-based health checking by default. Two practical consequences: chart fields that v2.7 silently dropped are now hard errors (fixed in this release for foundationdb, kafka, kubevirt-instancetypes, vm-instance and the platform chart), and parent HelmReleases now wait for every child resource to be Ready before reporting Ready themselves. Correctness improves; first-install timings get longer and stricter, which is why several packages gained explicit <code>spec.timeout</code> and <code>dependsOn</code> entries in the same release.</p>
<p>Alongside it, a new <strong>flux-shard-operator</strong> spreads tenant HelmReleases across multiple helm-controller shards, so one noisy tenant — the canonical case being a HelmRelease stuck in infinite remediation — can no longer degrade reconciliation for everyone else. Placement is per tenant: all of a tenant&rsquo;s HelmReleases share one shard, assigned greedily by least load, with a CREATE-time mutating webhook stamping the shard label. It defaults to <code>shardCount: auto</code>, which sizes shards from the tenant HelmRelease count — small clusters stay on a single shard, large fleets shard out automatically — and an integer pins the count. The hand-rolled <code>flux-tenants</code> deployment is drained and retired by migration 44.</p>
<h2 id="tls-for-managed-databases-and-messaging">TLS for managed databases and messaging</h2>
<p>Four managed-app charts gain TLS through a single <code>tls.enabled</code> value with consistent tri-state semantics. Unset, it inherits <code>external</code>: TLS is on when the service is published externally and off when it is cluster-internal. An explicit <code>true</code> or <code>false</code> always wins. In every case the trust anchor is a chart- or operator-managed self-signed CA that clients retrieve and pin — there is no publicly trusted CA in this path.</p>
<ul>
<li><strong>Kafka</strong> serves TLS on its external LoadBalancer listener (port 9094), certificates managed end to end by the Strimzi operator. Clients trust via the operator-published <code>&lt;release&gt;-cluster-ca-cert</code> and <code>&lt;release&gt;-clients-ca-cert</code> Secrets. The external listener is now gated only on <code>external: true</code>, decoupled from <code>tls.enabled</code>.</li>
<li><strong>NATS</strong> and <strong>Qdrant</strong> use a self-contained cert-manager chain (self-signed Issuer, CA, leaf) rendered in the tenant namespace. NATS covers client connections and cluster routes; Qdrant covers REST and gRPC. Clients trust the <code>&lt;release&gt;-ca</code> Secret.</li>
<li><strong>PostgreSQL</strong> already served TLS unconditionally via CNPG, so <code>tls.enabled</code> here injects the external hostname into the operator-managed server certificate&rsquo;s SANs when <code>external: true</code> — which is what makes <code>sslmode=verify-full</code> work against the external endpoint. Clients read <code>ca.crt</code> from the <code>&lt;release&gt;-credentials</code> Secret.</li>
</ul>
<p>The TLS work in this release stops at these four charts. The other managed services — MariaDB, Valkey, ClickHouse, OpenSearch, RabbitMQ, MongoDB — are unchanged.</p>
<h2 id="gpu-passthrough-without-manual-patching">GPU passthrough without manual patching</h2>
<p>GPU enablement is wired across all three paths a GPU can reach a workload, each of which previously needed manual reconciliation.</p>
<p>For <strong>tenant Kubernetes</strong>, node groups declaring <code>gpus</code> automatically get the <code>gpu=on</code> kubelet label, so HAMi&rsquo;s device plugin schedules and advertises <code>nvidia.com/gpu</code>. The tenant GPU operator loads the driver with <code>NVreg_NvLinkDisable=1</code>, which fixes single-SXM-GPU passthrough that previously hung at &ldquo;Fabric State: In Progress&rdquo; with CUDA reporting &ldquo;system not yet initialized&rdquo;. Both defaults are overridable via <code>addons.gpuOperator.valuesOverride</code>.</p>
<p>For <strong>KubeVirt VMs</strong>, enabling <code>cozystack.gpu-operator</code> auto-populates the KubeVirt CR: it injects the <code>HostDevices</code> feature gate and fills <code>permittedHostDevices</code> (plus <code>mediatedDevicesConfiguration</code> for vGPU) from shipped NVIDIA default tables. GPU VMs now schedule without a manual patch — at the cost of the ownership change listed in the upgrade notes above.</p>
<p>A third <strong><code>container</code> variant</strong> of the gpu-operator is added for hosts where the NVIDIA driver and container toolkit are already installed by the OS; it exposes GPUs to ordinary containerized pods through the device plugin only.</p>
<p>GPU sharing in Cozystack remains NVIDIA GPU Operator plus HAMi. Reference: <a href="https://cozystack.io/docs/v1.5/kubernetes/gpu-sharing/">GPU sharing and operator variants</a>.</p>
<h2 id="also-in-v150">Also in v1.5.0</h2>
<ul>
<li><strong>Deletion protection.</strong> Objects labelled <code>platform.cozystack.io/no-delete=true</code> cannot be deleted. The check runs in-process in the kube-apiserver through a ValidatingAdmissionPolicy — no webhook, no DaemonSet, no TLS, no extra image. Protected in this release: the <code>cozy-system</code> and <code>tenant-root</code> namespaces, the <code>tenant-root</code> HelmRelease, the <code>cozystack-version</code> ConfigMap, the <code>cozystack-packages</code> OCIRepository, the cert-manager ClusterIssuers, the LinstorCluster, and the packages CRDs. To delete one, remove the label first: <code>kubectl label &lt;kind&gt; &lt;name&gt; platform.cozystack.io/no-delete-</code>. Requires Kubernetes 1.30+.</li>
<li><strong>Runtime-populated dashboard dropdowns.</strong> A namespaced, read-only <code>Option</code> resource (<code>core.cozystack.io/v1alpha1</code>), computed on read by a privileged in-process provider registry, plus an <code>x-cozystack-options</code> schema keyword that charts declare. GPU devices, KubeVirt instancetypes and preferences, Multus networks, VM images, storage pools, storage classes, backup classes and plans become real dropdowns in the Cozystack Dashboard instead of free text or stale static enums. Tenants get read-only access to options in their own namespace.</li>
<li><strong>Tenants can start, stop and restart their own VMs.</strong> The missing <code>update</code> on the <code>virtualmachines/start</code>, <code>/stop</code> and <code>/restart</code> KubeVirt subresources is granted at <code>cozy:tenant:use:base</code>; the dashboard power buttons previously returned 403 for every tenant role.</li>
<li><strong>Tenant Overview Grafana dashboard</strong> for platform admins, deployed only to the root Grafana in <code>cozy-monitoring</code> and never to per-tenant Grafanas, plus a <strong>Cluster Usage</strong> admin page backed by a dedicated <code>cozystack-dashboard-cluster-usage</code> ClusterRole (cluster-wide and per-node utilization, GPUs included). The sidebar entry is fail-closed without the binding.</li>
<li><strong><code>storageClass</code> marked immutable on 16 stateful apps</strong> — changing it never migrates data, because PVCs pin <code>storageClassName</code> at creation. Enforcement in v1.5.0 is UI-only: the aggregated apiserver does not yet evaluate the CEL rule on Update, so a direct <code>kubectl patch</code> is still accepted.</li>
<li>Notable fixes: <code>cozystack-api</code> now publishes free-form <code>.spec</code> as <code>x-kubernetes-preserve-unknown-fields</code> rather than <code>additionalProperties: true</code>, which was crashing kube-controller-manager cluster-wide with a nil-pointer panic in the VAP type-checker; tenant kubeconfigs use the root host for the Keycloak OIDC issuer, so <code>kubectl oidc-login</code> no longer fails TLS verification for non-root tenants; <code>config_path</code> registry configuration works on containerd 2.x tenant nodes; RWX Block volumes are routed to the upstream hotplug detach path instead of the NFS-cleanup branch; OpenSearch is finally referenced by the PaaS bundle.</li>
</ul>
<h2 id="platform-components">Platform components</h2>
<table>
  <thead>
      <tr>
          <th>Component</th>
          <th>Change</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Flux</td>
          <td>v2.7.3 → v2.8.0 (flux-operator/flux-instance charts v0.33.0 → v0.50.0)</td>
      </tr>
      <tr>
          <td>MetalLB</td>
          <td>v0.15.2 → v0.16.1, FRR-K8s default backend, HTTPS-only metrics</td>
      </tr>
      <tr>
          <td>SeaweedFS</td>
          <td>4.05 → 4.31 (see the upgrade warning above)</td>
      </tr>
      <tr>
          <td>etcd-operator</td>
          <td>v0.4.3 → v0.4.5 (v0.4.5 fixes a restore-datadir path that made restore non-functional)</td>
      </tr>
      <tr>
          <td>ouroboros</td>
          <td>v0.7.2 → v0.8.0</td>
      </tr>
      <tr>
          <td>seaweedfs-cosi-driver</td>
          <td>v0.3.1, with stale-socket self-heal</td>
      </tr>
      <tr>
          <td>kuberture</td>
          <td>new optional system package v0.1.1</td>
      </tr>
      <tr>
          <td>Go toolchain</td>
          <td>1.26.4, <code>golang.org/x/net</code> v0.55.0</td>
      </tr>
  </tbody>
</table>
<p><code>kuberture</code> is off by default. It bridges an external-dns gap — external-dns cannot read EndpointSlices — by watching the <code>default/kubernetes</code> API-server EndpointSlice and emitting annotated headless Services that external-dns consumes to publish the Kubernetes API endpoint to DNS. Enable via <code>bundles.enabledPackages</code> with at least one <code>config.outputs</code> entry.</p>
<h2 id="the-15-patch-line">The 1.5 patch line</h2>
<ul>
<li><strong>v1.5.1</strong> (24 June 2026) fixes a v1.5.0 regression: persistent EFI/TPM state was restored for the <code>windows.11</code>, <code>windows.2k22</code> and <code>windows.2k25</code> KubeVirt preferences, which made KubeVirt provision a ReadWriteOnce <code>persistent-state-for-&lt;vm&gt;</code> PVC on the <code>replicated</code> StorageClass. That pins the VM to its node and blocks live migration and node drains — on clusters using <code>evictionStrategy: LiveMigrate</code> it can stall a cluster upgrade outright. Secure Boot and the vTPM still work; only the state persistence across reboots is dropped.</li>
<li><strong>v1.5.2</strong> (3 July 2026) covers ten fixes, including a Kamaji DataStore deletion deadlock that wedged tenant namespaces, a victoria-metrics-operator dependency that could never resolve when <code>certManager.enabled: false</code>, and single-replica MariaDB instances being rejected by the operator webhook.</li>
<li><strong>v1.5.3</strong> was tagged but its GitHub release was left a draft and never published. No operator received it.</li>
<li><strong>v1.5.4</strong> (19 August 2026) is the final 1.5 release and the one to deploy.</li>
</ul>
<h2 id="upgrading">Upgrading</h2>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">kubectl annotate namespace cozy-system helm.sh/resource-policy<span class="o">=</span>keep --overwrite
</span></span><span class="line"><span class="cl">kubectl annotate configmap -n cozy-system cozystack-version helm.sh/resource-policy<span class="o">=</span>keep --overwrite
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">helm upgrade cozystack oci://ghcr.io/cozystack/cozystack/cozy-installer <span class="se">\
</span></span></span><span class="line"><span class="cl">  --version 1.5.4 <span class="se">\
</span></span></span><span class="line"><span class="cl">  --namespace cozy-system
</span></span></code></pre></div><p>The annotations are required, not optional — without them, removing or upgrading the installer release can delete the <code>cozy-system</code> namespace and everything in it. Reuse your existing values. Full procedure and post-upgrade checks: <a href="https://cozystack.io/docs/v1.5/operations/cluster/upgrade/">upgrade guide</a>.</p>
<h2 id="where-ænix-fits">Where Ænix fits</h2>
<p>Cozystack is a CNCF Sandbox project under Apache 2.0, and the upgrade paths above are the same for everyone running it. Ænix created Cozystack and is one of its maintainers; it sells <a href="https://aenix.io/products/cozystack-enterprise-support/">enterprise support for Cozystack</a> — including upgrade planning for the sharp edges in this release — for teams that would rather not carry the platform alone.</p>
<h2 id="release-links">Release links</h2>
<ul>
<li><a href="https://github.com/cozystack/cozystack/releases/tag/v1.5.0">Cozystack v1.5.0 on GitHub</a></li>
<li><a href="https://github.com/cozystack/cozystack/releases/tag/v1.5.4">Cozystack v1.5.4 on GitHub</a></li>
<li><a href="https://cozystack.io/docs/v1.5/">Cozystack v1.5 documentation</a></li>
<li><a href="https://t.me/cozystack">Telegram</a> and <a href="https://kubernetes.slack.com/archives/C06L3CPRVN1">Slack</a> (invite at <a href="https://slack.kubernetes.io/">slack.kubernetes.io</a>)</li>
</ul>
]]></content:encoded></item><item><title>White-label cloud playbook — for MSPs and resellers in 2026</title><link>https://aenix.io/blog/2026/05/white-label-cloud-msp-reseller-playbook/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/white-label-cloud-msp-reseller-playbook/</guid><pubDate>Sun, 31 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Cozystack</category><category>Multi-tenancy</category><category>Hosting</category><category>Observability</category><description>Architecture and reseller economics for launching a white-label cloud under your own brand, and how the engagement is structured.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/white-label-cloud-msp-reseller-playbook.jpg" alt=""></p><h2 id="why-white-label-cloud-matters-for-msps">Why white-label cloud matters for MSPs</h2>
<p>MSPs have customer relationships hyperscalers can&rsquo;t easily replicate. They lack the cloud product to monetize those relationships at scale. White-label cloud — branded as MSP&rsquo;s product, running on shared or dedicated infrastructure — bridges this.</p>
<p>Pattern in 2026: MSP gets branded multi-tenant cloud product on open-source platform; customers consume MSP-branded cloud; MSP collects margin between platform cost and customer pricing.</p>
<h2 id="architecture">Architecture</h2>
<ul>
<li><strong>Multi-tier Tenant CRD</strong> — root tenant → MSP tenant → MSP customer tenant. Per-tier isolation.</li>
<li><strong>Branded Cozystack Dashboard</strong> — MSP can customize colors, logo, domain, service catalog options (white-labelling is part of open-source Cozystack; Ænix support covers it from the Standard tier)</li>
<li><strong>WHMCS integration</strong> — billing flows through MSP&rsquo;s existing customer-management system via the Ænix <a href="https://aenix.io/products/whmcs-integration/">WHMCS integration</a>, a proprietary Ænix module</li>
<li><strong>Service catalog</strong> — MSP can curate which services to expose to customers (e.g., hide Kafka if MSP doesn&rsquo;t support it)</li>
<li><strong>SLA management</strong> — per-customer SLA tracking through Cozystack observability</li>
</ul>
<h2 id="reseller-economics">Reseller economics</h2>
<p>Typical economics for an MSP running white-label cloud:</p>
<ul>
<li><strong>Platform cost</strong> — Ænix engagement + hardware + colocation</li>
<li><strong>Per-customer cost</strong> — incremental hardware/storage/bandwidth</li>
<li><strong>Customer pricing</strong> — typically 30-50% above raw platform cost</li>
<li><strong>Margin</strong> — covers MSP support, sales, operations</li>
</ul>
<p>Break even at 30-50 paying customers when you are covering the platform and tooling; 50-100 when a dedicated on-call rota is funded alongside it. Positive economics after that, depending on customer mix.</p>
<h2 id="engagement-structure">Engagement structure</h2>
<ul>
<li><strong>Free 30-minute discovery call</strong>, then a fixed-price <a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a> (14 days focused or 28 days full)</li>
<li><strong>Platform live in weeks</strong> once hardware is ready, using the productized installer; branding, catalogue curation and billing integration follow at your pace</li>
<li><strong>Support subscription</strong> — published tiers per 10 nodes per month; for a white-label product, plan on Standard ($3,000) or higher (see <a href="https://aenix.io/pricing/">/pricing/</a>)</li>
<li><strong>Optional managed services</strong></li>
</ul>
]]></content:encoded></item><item><title>When Cozystack fits SMB and mid-market — and when it doesn't</title><link>https://aenix.io/blog/2026/05/when-cozystack-fits-smb-and-mid-market/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/when-cozystack-fits-smb-and-mid-market/</guid><pubDate>Sat, 30 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>DORA</category><category>VMware</category><category>Proxmox</category><category>Kubernetes</category><category>Cozystack</category><category>Sovereignty</category><description>Most SMB organizations do not need Cozystack. An honest test for when they do, and what to run instead when they do not.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/when-cozystack-fits-smb-and-mid-market.jpg" alt=""></p><h2 id="the-honest-test">The honest test</h2>
<p>Cozystack fits when at least three of:</p>
<ol>
<li><strong>Regulated data</strong> — banking, healthcare, public sector data with sovereignty / residency requirements</li>
<li><strong>Multi-tenant model</strong> — SaaS / customer-facing / cross-BU isolation needed</li>
<li><strong>Sustained workloads</strong> — 24/7 utilization where dedicated economics beat hyperscaler</li>
<li><strong>Internal platform team</strong> — capacity to operate (or building it)</li>
<li><strong>AI/GPU workloads at scale</strong> — sustained inference / training</li>
<li><strong>Specific exit trigger</strong> — VMware migration, hyperscaler repatriation</li>
</ol>
<p>If you have 0-1 of these, Cozystack is over-engineering. If 2, marginal. If 3+, fits.</p>
<h2 id="when-it-doesnt-fit--what-to-do-instead">When it doesn&rsquo;t fit — what to do instead</h2>
<h3 id="for-smb-without-regulated-data">For SMB without regulated data</h3>
<ul>
<li><strong>Hyperscaler-managed deployments</strong> (AWS, Azure, GCP) — operationally simple</li>
<li><strong>DigitalOcean / Hetzner / OVHcloud</strong> — managed cloud-adjacent</li>
<li><strong>Proxmox VE</strong> for on-prem virtualization</li>
</ul>
<h3 id="for-mid-market-with-simple-needs">For mid-market with simple needs</h3>
<ul>
<li><strong>Vanilla Kubernetes</strong> if container-only — lighter than Cozystack</li>
<li><strong>Existing managed cloud</strong> if it works — don&rsquo;t fix what isn&rsquo;t broken</li>
<li><strong>Hetzner cloud + VPS</strong> if team is small — operationally simple</li>
</ul>
<h2 id="when-it-does-fit--examples">When it does fit — examples</h2>
<h3 id="mid-market-with-regulated-data">Mid-market with regulated data</h3>
<ul>
<li>Regional bank with multi-jurisdictional operations under DORA</li>
<li>Mid-size insurance carrier with claims-processing AI on regulated data</li>
<li>Public-sector or quasi-public with procurement-mandated sovereignty</li>
</ul>
<h3 id="mid-market-becoming-multi-tenant">Mid-market becoming multi-tenant</h3>
<ul>
<li>SaaS company with 100+ customers needing hard isolation</li>
<li>B2B platform serving regulated industry customers</li>
<li>Specialty cloud product for vertical market</li>
</ul>
<h3 id="mid-market-with-strong-platform-team">Mid-market with strong platform team</h3>
<ul>
<li>Mid-market with intentional in-house platform engineering investment</li>
<li>Fast-growing tech-mid-market scaling beyond hyperscaler simple model</li>
</ul>
<h2 id="ænix-engagement-model-for-mid-market">Ænix engagement model for mid-market</h2>
<ul>
<li><strong><a href="https://aenix.io/contact/">30-minute discovery call</a></strong> — free, no sales pressure</li>
<li><strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong> (fixed price, 14 days focused) — if you want a structured assessment</li>
<li><strong>Implementation</strong> — only if it actually fits; self-run Cozystack with <a href="https://aenix.io/products/cozystack-enterprise-support/">Ænix enterprise support</a> is often enough at mid-market scale</li>
</ul>
<p>For most SMB outreach, the honest answer is &ldquo;stay where you are.&rdquo; We&rsquo;re explicit about this.</p>
]]></content:encoded></item><item><title>VMware replacement after Broadcom: a guide for service providers, banks, and sovereign clouds in 2026</title><link>https://aenix.io/blog/2026/05/vmware-replacement-after-broadcom/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/vmware-replacement-after-broadcom/</guid><pubDate>Sat, 30 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>VMware</category><category>Kubernetes</category><category>Cozystack</category><category>Sovereignty</category><category>AI and ML</category><category>GPU</category><description>What changed under Broadcom, a component-by-component VMware-to-Cozystack mapping, how the migration actually runs, and the FAQ engineers ask first.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/vmware-replacement-after-broadcom.jpg" alt=""></p><p>After Broadcom, the VMware bill stopped being predictable. Subscription-only licensing, mandatory VCF bundling, two-to-five-times price increases on renewal, and the end of perpetual licences changed the math for every infrastructure team running VMware at scale. The result has been a documented wave of VMware replacement projects across service providers, banks, government, telecom, and AI/GPU operators evaluating how to exit VMware safely.</p>
<hr>
<h2 id="vmware-competitors-at-a-glance">VMware competitors at a glance</h2>
<p>If you&rsquo;re scoping the market, here is how Cozystack compares to the most-cited alternatives to VMware:</p>
<ul>
<li><strong>Cozystack (this page)</strong> — open-source, Kubernetes-native, multi-tenant. VMs + containers + databases + object storage + GPU on one platform. Best for service providers and regulated enterprises that need sovereignty and want one stack.</li>
<li><strong>Nutanix AHV</strong> — proprietary KVM-based hypervisor inside Nutanix HCI. Strong VM-only story, weak on Kubernetes-native multi-tenancy, expensive at scale.</li>
<li><strong>Proxmox VE</strong> — open-source KVM + LXC. Excellent for SMB and labs; thin on managed databases, multi-tenancy, and enterprise support model.</li>
<li><strong>Scale Computing HC3</strong> — appliance-based hyperconverged stack. Good for ROBO/edge; closed ecosystem.</li>
<li><strong>Red Hat OpenShift Virtualization</strong> — KubeVirt-based, similar foundations to Cozystack. Heavier OpenShift footprint and Red Hat licensing model.</li>
<li><strong>OpenStack</strong> — broad, mature, large operational surface area. Best when you have a dedicated team running it.</li>
<li><strong>Microsoft Azure Stack HCI</strong> — Hyper-V on validated hardware. Locks into Microsoft licensing.</li>
<li><strong>Verge.io / Spectro Cloud / Platform9</strong> — vendor-led KubeVirt or hyperconverged stacks; comparable to Cozystack on architecture, different commercial models.</li>
</ul>
<p>The rest of this page goes deep on Cozystack as a Cozystack-specific VMware replacement. For a head-to-head comparison of VMware alternatives, see <a href="https://aenix.io/alternatives/vmware-alternatives/">VMware alternatives</a>.</p>
<hr>
<h2 id="why-teams-replace-vmware-in-2026">Why teams replace VMware in 2026</h2>
<p>The technical case for moving off VMware existed before Broadcom. Broadcom turned it into a board-level decision.</p>
<h3 id="1-unpredictable-subscription-only-economics">1. Unpredictable, subscription-only economics</h3>
<p>VCF subscription pricing replaced perpetual licensing. Renewal quotes have come back at 2× to 5× prior spend across our pipeline. ELAs have been broken or restructured mid-term. Standalone vSphere SKUs were retired in favour of bundled VCF tiers that include components most customers do not need.</p>
<p>For service providers, this collapses margin: end-customer prices are sticky, licence costs are not. For banks and regulated enterprises, it breaks multi-year capex planning that was built around perpetual entitlements.</p>
<h3 id="2-vendor-lock-in-across-the-stack">2. Vendor lock-in across the stack</h3>
<p>VMware&rsquo;s strength was integration: vSphere, vSAN, NSX, vRealize, Horizon, vCD, all assumed each other. That same integration is now the lock-in surface. Replacing one component meant rebuilding adjacent ones, so most teams stayed.</p>
<p>Cozystack inverts this. Each layer is an independent open-source project (KubeVirt, LINSTOR/DRBD, Cilium, Kube-OVN, KubeVirt CDI, SeaweedFS, Velero, Flux, etc.), composed by a Kubernetes operator. You can replace any layer without rewriting the rest, and you can audit every line.</p>
<h3 id="3-sovereignty-regulator-pressure-and-us-vendor-risk">3. Sovereignty, regulator pressure, and US-vendor risk</h3>
<p>DORA (in force across the EU since January 2025) and NIS2 require demonstrable control over critical ICT third parties. For European banks, telcos, and government workloads, depending on a US-headquartered closed-source hypervisor stack is a documented operational risk.</p>
<p>Cozystack is open source under Apache 2.0. Your binaries, your hardware, your data plane. Ænix ships air-gap install workflows and an advisory support model that does not require direct customer-environment access.</p>
<h3 id="4-pricing-isnt-the-only-cliff--capability-is">4. Pricing isn&rsquo;t the only cliff — capability is</h3>
<p>VMware&rsquo;s roadmap is now Broadcom&rsquo;s roadmap. KubeVirt and the Kubernetes-native virtualization stack have a community of hundreds of contributors and a release cadence VMware can no longer match in the open. GPU support, multi-tenant tenancy, modern storage replication, GitOps-native ops — all moved to Kubernetes-native projects years ago. VMware tries to retro-fit them through Tanzu and VCF.</p>
<hr>
<h2 id="what-cozystack-gives-you-instead">What Cozystack gives you instead</h2>
<p>Cozystack is a single platform you install on bare metal. Once it&rsquo;s up, you have:</p>
<ul>
<li><strong>Virtual machines</strong> through KubeVirt — full KVM-based VMs with live migration (CPU only), block-storage attach, snapshots, and templates.</li>
<li><strong>Tenant Kubernetes clusters</strong> for customers who want containers — every tenant gets their own K8s, isolated.</li>
<li><strong>Managed databases</strong> — PostgreSQL, MariaDB, MongoDB, Redis, Valkey, RabbitMQ, Kafka, NATS, ClickHouse, OpenSearch, Qdrant, FoundationDB — exposed as cloud services your tenants self-provision.</li>
<li><strong>S3-compatible object storage</strong> — for backups, application data, and AI training sets.</li>
<li><strong>GPU as a service</strong> — for VMs (passthrough of whole GPUs, or NVIDIA vGPU, which requires your NVIDIA vGPU licence) and for Kubernetes pods (whole-GPU through the NVIDIA GPU Operator, fractional sharing through HAMi). NVIDIA data-centre GPUs are supported through the NVIDIA GPU Operator; MIG and time-slicing are on the roadmap.</li>
<li><strong>Multi-tenant control plane</strong> — <code>Tenant</code> Kubernetes CRD, nested tenants, per-tenant quotas and presets.</li>
<li><strong>Observability built in</strong> — VictoriaMetrics + VictoriaLogs, open source and included.</li>
<li><strong>Backup and DR</strong> — Velero + S3 + per-database point-in-time recovery for managed services.</li>
<li><strong>Self-service portal</strong> — Cozystack Dashboard. Billing runs in your own system or through the Ænix <a href="https://aenix.io/products/whmcs-integration/">WHMCS integration</a>, a proprietary Ænix module (two integration modes) that is not part of open-source Cozystack.</li>
</ul>
<p>It runs on your bare metal. No public cloud dependency. No phone-home telemetry by default.</p>
<hr>
<h2 id="vmware-to-cozystack-architecture-mapping-vsphere-alternative-esxi-alternative-vcenter-alternative--all-in-one">VMware-to-Cozystack architecture mapping (vSphere alternative, ESXi alternative, vCenter alternative — all in one)</h2>
<p>The questions every VMware admin asks first: <em>&ldquo;What replaces vSphere? What&rsquo;s a real ESXi alternative? What replaces vCenter, NSX, vSAN, vCloud Director?&rdquo;</em> Here is the direct one-to-one mapping.</p>
<table>
  <thead>
      <tr>
          <th>VMware / VCF component</th>
          <th>Cozystack equivalent</th>
          <th>Acts as</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><strong>vSphere / ESXi</strong></td>
          <td>KubeVirt on Talos</td>
          <td>vSphere alternative, ESXi alternative — KVM-based VMs with live migration, snapshots, templates</td>
      </tr>
      <tr>
          <td><strong>vCenter</strong></td>
          <td>Cozystack control plane (Kubernetes API + Cozystack Dashboard)</td>
          <td>vCenter alternative — multi-cluster federation through Cluster API</td>
      </tr>
      <tr>
          <td><strong>vSAN</strong></td>
          <td>LINSTOR (DRBD, via the Piraeus operator)</td>
          <td>vSAN alternative — hyperconverged synchronous replicated block storage</td>
      </tr>
      <tr>
          <td><strong>NSX</strong></td>
          <td>Cilium (eBPF)</td>
          <td>NSX alternative — native L4/L7, network policies, observability, no NSX licensing</td>
      </tr>
      <tr>
          <td><strong>vCloud Director (vCD)</strong></td>
          <td>Tenant CRD + Cozystack Dashboard</td>
          <td>vCloud Director alternative — multi-tenancy, self-service, quotas, RBAC</td>
      </tr>
      <tr>
          <td><strong>vRealize / Aria Operations</strong></td>
          <td>VictoriaMetrics + VictoriaLogs + Grafana</td>
          <td>Aria Operations alternative — open-source observability stack</td>
      </tr>
      <tr>
          <td><strong>Site Recovery Manager (SRM)</strong></td>
          <td>Velero + S3 + PostgreSQL PITR</td>
          <td>Backup and restore with rehearsed runbooks; replication for stateful services. No SRM-style orchestrated cross-site failover</td>
      </tr>
      <tr>
          <td><strong>Horizon (VDI)</strong></td>
          <td>Not in scope of Cozystack</td>
          <td>Pair with KasmWorkspaces or similar; talk to us about reference designs</td>
      </tr>
      <tr>
          <td><strong>Tanzu Kubernetes Grid</strong></td>
          <td>Tenant Kubernetes (native)</td>
          <td>Tanzu alternative — each tenant gets a real K8s control plane</td>
      </tr>
      <tr>
          <td><strong>vRealize Automation / vRA</strong></td>
          <td>ApplicationDefinition + portal catalog</td>
          <td>vRA alternative — self-service service catalog, GitOps-native</td>
      </tr>
      <tr>
          <td><strong>VMware Cloud Foundation (VCF)</strong></td>
          <td>Cozystack</td>
          <td>VCF alternative — whole-stack equivalent, single open-source distribution</td>
      </tr>
  </tbody>
</table>
<p>Two areas need redesign rather than one-for-one mapping: <strong>networking</strong> (Cilium is fundamentally different from NSX) and <strong>multi-tenancy model</strong> (Cozystack tenants are Kubernetes-native; vCD orgs are vSphere-native). For both, we run an architecture review before any migration commits.</p>
<hr>
<h2 id="how-vmware-migration-actually-works">How VMware migration actually works</h2>
<p>Migration paths are workload-dependent. For most teams, this is the realistic sequence.</p>
<h3 id="1-discovery-and-assessment">1. Discovery and assessment</h3>
<p>We run a structured assessment of the current vSphere / VCF / vCD inventory: workload count, OS mix, vSAN / NSX dependencies, integrations (backup, identity, monitoring), and tenancy model. Output is a migration plan with workload buckets, risk flags, and timing.</p>
<h3 id="2-cozystack-deployed-in-parallel">2. Cozystack deployed in parallel</h3>
<p>Cozystack is installed on new or repurposed hardware alongside the existing VMware estate. No big-bang cutover. Tenants migrate cohort by cohort.</p>
<h3 id="3-vm-by-vm-image-migration-to-kubevirt">3. VM-by-VM image migration to KubeVirt</h3>
<p>For most VMs, migration is a disk-image copy. We use KubeVirt CDI plus a set of dedicated migration scripts we&rsquo;ve built and reused across customer deployments. For Windows VMs and VMs with VMware Tools dependencies, we run an automated cleanup pass before boot on KubeVirt.</p>
<h3 id="4-network-and-storage-cutover">4. Network and storage cutover</h3>
<p>Networking: VLAN mapping into Cilium, with policy parity checked against the source NSX rules. Storage: import disks into LINSTOR, validate IOPS and replication.</p>
<h3 id="5-validation-and-dr-cutover">5. Validation and DR cutover</h3>
<p>Each migrated workload runs in parallel on Cozystack until validated by the application owner. Backup-and-restore runbooks (Velero, PostgreSQL PITR) are written and rehearsed before final cutover; there is no SRM-style orchestrated cross-site failover.</p>
<h3 id="6-vmware-decommission">6. VMware decommission</h3>
<p>Licences end on their own terms. Hardware repurposed into the Cozystack cluster (we run on commodity x86 — your existing servers usually qualify).</p>
<p>For OpenStack, CloudStack, and Proxmox sources the same playbook applies, with different image-import and network-mapping steps.</p>
<blockquote>
<p><strong>Looking for a checklist?</strong> Use the <a href="https://aenix.io/resources/vmware-migration-checklist/">VMware migration checklist</a>. A structured assessment is part of the engagement — <a href="https://aenix.io/contact/">book a call</a> to scope it.</p>
</blockquote>
<hr>
<h2 id="multi-tenancy-and-sovereignty">Multi-tenancy and sovereignty</h2>
<p>Cozystack was built for service providers first. The same model works for any organization that needs hard separation between business units.</p>
<ul>
<li><strong>Tenant CRD</strong> — each tenant is a Kubernetes object. Namespaces are derived (<code>&lt;parent&gt;-&lt;name&gt;</code>), nested tenants are supported, quotas are enforced via Kubernetes ResourceQuotas.</li>
<li><strong>RBAC by default</strong> — tenants cannot see each other. Tenant operators get a scoped Kubernetes API.</li>
<li><strong>Air-gap install</strong> — supported and documented; works behind Harbor / Nexus / proxy patterns.</li>
<li><strong>No phone-home</strong> — telemetry is opt-in and disabled by default.</li>
<li><strong>DORA / NIS2-aligned controls</strong> — operational resilience, incident reporting, supplier-risk documentation. Architecture maps to the controls; certification, where required, is the customer&rsquo;s audit process to run.</li>
<li><strong>Ænix support model</strong> — access is your choice. Advisory, runbooks and GitOps PR review need no access to your production cluster; remote access to your clusters, with your approval, and external monitoring are available on the higher <a href="https://aenix.io/pricing/">support tiers</a>. Critical for banks, telcos, and any regulated buyer who cannot expose infrastructure to a vendor.</li>
</ul>
<hr>
<h2 id="cozystack-vs-vmware--feature-comparison">Cozystack vs VMware — feature comparison</h2>
<table>
  <thead>
      <tr>
          <th>Capability</th>
          <th>VMware (VCF, post-Broadcom)</th>
          <th>Cozystack + Ænix</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><strong>Licence model</strong></td>
          <td>Subscription only (VCF bundles)</td>
          <td>Apache 2.0 open source + optional Ænix support tier</td>
      </tr>
      <tr>
          <td><strong>Renewal risk</strong></td>
          <td>2–5× increases observed; bundling forced</td>
          <td>Predictable support pricing; OSS code remains usable regardless</td>
      </tr>
      <tr>
          <td><strong>Compute</strong></td>
          <td>vSphere / ESXi</td>
          <td>KubeVirt (KVM-based)</td>
      </tr>
      <tr>
          <td><strong>Live migration</strong></td>
          <td>Yes (incl. with GPU under vGPU)</td>
          <td>Yes (CPU); GPU live migration not supported (industry-wide limitation, not Cozystack-specific)</td>
      </tr>
      <tr>
          <td><strong>Storage</strong></td>
          <td>vSAN</td>
          <td>LINSTOR (DRBD); synchronous replication, production-grade</td>
      </tr>
      <tr>
          <td><strong>Network</strong></td>
          <td>NSX (proprietary)</td>
          <td>Cilium (CNCF Graduated, eBPF)</td>
      </tr>
      <tr>
          <td><strong>Multi-tenancy</strong></td>
          <td>vCloud Director</td>
          <td>Tenant CRD; native to Kubernetes</td>
      </tr>
      <tr>
          <td><strong>Service catalog</strong></td>
          <td>vRealize Automation / Aria</td>
          <td>ApplicationDefinition + Cozystack Dashboard</td>
      </tr>
      <tr>
          <td><strong>Backup / DR</strong></td>
          <td>Site Recovery Manager (SRM)</td>
          <td>Velero + S3 + PostgreSQL PITR with rehearsed runbooks; no SRM-style orchestrated failover</td>
      </tr>
      <tr>
          <td><strong>GPU for VMs</strong></td>
          <td>NVIDIA vGPU under Horizon</td>
          <td>NVIDIA vGPU + KubeVirt (requires your NVIDIA vGPU licence); passthrough of whole GPUs</td>
      </tr>
      <tr>
          <td><strong>GPU for containers</strong></td>
          <td>Tanzu (limited)</td>
          <td>Whole-GPU via NVIDIA GPU Operator; fractional sharing (GPU memory + cores) via HAMi — Kubernetes-native</td>
      </tr>
      <tr>
          <td><strong>Observability</strong></td>
          <td>vRealize / Aria (licensed separately)</td>
          <td>VictoriaMetrics + VictoriaLogs (OSS, included)</td>
      </tr>
      <tr>
          <td><strong>Ops model</strong></td>
          <td>Vendor support requires environment access</td>
          <td>Your choice: advisory + runbooks + GitOps PR review with no cluster access, or remote access with your approval</td>
      </tr>
      <tr>
          <td><strong>Sovereignty</strong></td>
          <td>Closed source, US vendor</td>
          <td>Open source, hosted on your hardware, EU contracting entity (AENIX s.r.o.)</td>
      </tr>
      <tr>
          <td><strong>Air-gap install</strong></td>
          <td>Supported (additional licensing)</td>
          <td>Supported (no additional cost)</td>
      </tr>
      <tr>
          <td><strong>Compliance posture</strong></td>
          <td>Customer responsibility on top of VCF</td>
          <td>Architecture aligned with DORA / NIS2 controls; CNCF Certified Kubernetes distribution, CNCF Kubernetes AI Conformance, OpenSSF Best Practices; AENIX s.r.o. holds <a href="https://aenix.io/compliance/iso-27001/">ISO/IEC 27001</a></td>
      </tr>
      <tr>
          <td><strong>Pricing transparency</strong></td>
          <td>Quote-driven; non-public</td>
          <td>Public pricing on aenix.io/pricing; OSS is free</td>
      </tr>
  </tbody>
</table>
<hr>
<h2 id="faq">FAQ</h2>
<h3 id="can-we-keep-our-existing-hardware">Can we keep our existing hardware?</h3>
<p>Yes — in most cases. Cozystack runs on commodity x86. The standard scenario is to deploy Cozystack on a new pod of servers, migrate workloads off VMware, then repurpose the freed VMware hardware into Cozystack as licences lapse.</p>
<h3 id="how-long-is-a-typical-migration">How long is a typical migration?</h3>
<p>Two numbers matter, and conflating them is how migration plans go wrong. The <strong>first production cohort</strong> typically runs in a live environment 6-12 weeks after kickoff — that is when the platform stops being a proof of concept. The <strong>whole estate</strong> takes about 8-12 months for ~100 VMs and 18-24 months for ~1,000 VMs, including planning and migration waves, after a 14- or 28-day Platform Readiness Assessment; mid-size estates fall in between, depending on dependencies. Small, flat estates land at the short end; vCD- or NSX-heavy regulated ones at the long end, run in cohorts. The driver is rarely raw migration speed — it&rsquo;s regression testing and the parallel-run windows application owners will agree to.</p>
<h3 id="do-you-support-windows-vms">Do you support Windows VMs?</h3>
<p>Yes. KubeVirt runs Windows. We have an automated cleanup step that removes VMware Tools and installs the right virtio drivers before first boot on KubeVirt. Activation handling depends on your Windows licensing model — discussed during assessment.</p>
<h3 id="what-happens-to-our-existing-vmware-operational-skills">What happens to our existing VMware operational skills?</h3>
<p>The hypervisor concepts (VMs, snapshots, templates, networking) carry over. The control-plane shifts from vCenter to Kubernetes — for most teams this is a 4–8 week ramp with the right training and runbooks. Ænix runs the training as part of professional services.</p>
<h3 id="what-about-vmware-cloud-foundation-specifically--is-the-migration-different">What about VMware Cloud Foundation specifically — is the migration different?</h3>
<p>VCF migrations are larger (more components, more integrations). The discovery phase covers SDDC Manager, Workspace ONE, Aria, NSX-T overlays, and any custom service definitions, and the assessment sizes the extra work.</p>
<h3 id="what-if-we-use-vcloud-director-for-our-customers">What if we use vCloud Director for our customers?</h3>
<p>vCD migration is the most common path for service providers. Tenant model maps to Cozystack Tenant CRD, service catalog maps to ApplicationDefinition, billing maps to the Ænix WHMCS integration or your own billing system. We are happy to walk through the architecture in a call.</p>
<h3 id="is-gpu-live-migration-supported">Is GPU live migration supported?</h3>
<p>No — industry-wide limitation, not Cozystack-specific. VMware vGPU live migration has known caveats too. For inference workloads we recommend designing for stateless restart rather than relying on GPU live migration.</p>
]]></content:encoded></item><item><title>VMware migration tools and strategy in 2026 — what works, what fails</title><link>https://aenix.io/blog/2026/05/vmware-migration-tools-and-strategy/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/vmware-migration-tools-and-strategy/</guid><pubDate>Fri, 29 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>VMware</category><category>Nutanix</category><category>OpenShift</category><category>Kubernetes</category><category>Cozystack</category><category>KubeVirt</category><description>Three VMware migration paths, the tooling for KubeVirt-based migration, where migrations stumble, and realistic cost ranges.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/vmware-migration-tools-and-strategy.jpg" alt=""></p><p>The VMware migration market in 2026 is a different conversation than in 2022. Broadcom-induced exits have produced enough customer experience that the patterns that work and the patterns that fail are documented. This article covers the working version.</p>
<h2 id="migration-paths--three-options">Migration paths — three options</h2>
<p>There are three honest VMware migration paths in 2026:</p>
<h3 id="path-1-vmware-managed-migration-vendor-tools-led">Path 1: VMware-managed migration (vendor-tools-led)</h3>
<p>Use Broadcom-supplied or Red Hat-supplied tooling to migrate within or out of the VMware ecosystem. Tools: VMware HCX (intra-VMware), Red Hat Migration Toolkit for Virtualization (MTV) for VMware → OpenShift, vendor-specific (Nutanix Move, Scale Computing tools).</p>
<p><strong>Pros:</strong> Vendor support during migration; established tooling.</p>
<p><strong>Cons:</strong> Locks into vendor relationship at destination. Limited choice of destination architecture.</p>
<h3 id="path-2-kubevirt-based-migration-open-source-destination">Path 2: KubeVirt-based migration (open-source destination)</h3>
<p>Convert VMware VMs to KubeVirt format on the destination platform (Cozystack, OpenShift Virtualization, vendor KubeVirt platforms). Tools: virtv2v (RedHat), forklift (RedHat), KubeVirt CDI (community), custom scripts.</p>
<p><strong>Pros:</strong> Open destination architecture, no vendor lock-in at destination, modern Kubernetes-native foundation.</p>
<p><strong>Cons:</strong> More integration work upfront; team learning curve.</p>
<h3 id="path-3-lift-and-shift-to-public-cloud">Path 3: Lift-and-shift to public cloud</h3>
<p>Move VMware workloads to VMware-on-cloud (VMware on AWS, Azure VMware Solution, Oracle VMware Solution). Defers the architectural question.</p>
<p><strong>Pros:</strong> Fastest migration path; smallest re-architecture.</p>
<p><strong>Cons:</strong> Doesn&rsquo;t address Broadcom pricing pressure (often makes it worse); doesn&rsquo;t address sovereignty; usually a stop-gap rather than a strategy.</p>
<p>For most 2026 migrations driven by Broadcom pricing or sovereignty, Path 2 (KubeVirt-based) wins. This article focuses there.</p>
<h2 id="tooling-for-kubevirt-based-migration">Tooling for KubeVirt-based migration</h2>
<h3 id="kubevirt-cdi-containerized-data-importer">KubeVirt CDI (Containerized Data Importer)</h3>
<p>Native KubeVirt tool. Imports disk images (qcow2, VMDK, VHD, ISO) into KubeVirt PersistentVolumeClaims. Free, open-source, works.</p>
<p><strong>Strengths:</strong> Native to KubeVirt; supported by KubeVirt operators; works with various source formats.
<strong>Limits:</strong> Disk-level migration; doesn&rsquo;t handle VM metadata, networking, or post-migration cleanup automatically.</p>
<h3 id="forklift--migration-toolkit-for-virtualization-red-hat">Forklift / Migration Toolkit for Virtualization (Red Hat)</h3>
<p>Red Hat&rsquo;s tool for vSphere → KubeVirt (OpenShift Virtualization) migration. Open-source; works on any KubeVirt platform with manual configuration.</p>
<p><strong>Strengths:</strong> Bulk VM migration; metadata preservation; production-tested.
<strong>Limits:</strong> Originally OpenShift-focused; some integration work for non-OpenShift destinations.</p>
<h3 id="virtv2v-red-hat">virtv2v (Red Hat)</h3>
<p>Lower-level conversion tool. Operates on individual VM images; converts vSphere VMs to KVM-compatible format with VMware Tools cleanup.</p>
<p><strong>Strengths:</strong> Mature, used by Forklift internally; handles Windows VM driver cleanup.
<strong>Limits:</strong> Single-VM tool; orchestration is your problem.</p>
<h3 id="cozystack-specific-migration-tooling">Cozystack-specific migration tooling</h3>
<p>KubeVirt CDI + dedicated migration scripts that Ænix has built and reused across customer deployments. Covers VM image conversion, multi-tenant placement, network mapping into Cilium policies.</p>
<p><strong>Strengths:</strong> used by Ænix in production migrations; Cozystack-tenant-aware.
<strong>Limits:</strong> delivered as part of an Ænix engagement, not as a standalone open-source tool.</p>
<h3 id="vendor--commercial-tools">Vendor / commercial tools</h3>
<ul>
<li><strong>Nutanix Move</strong> — for VMware → Nutanix AHV migrations</li>
<li><strong>Scale Computing tools</strong> — for VMware → Scale HC3</li>
<li><strong>Various commercial migration tools</strong> — Carbonite, Veeam, etc.</li>
</ul>
<p>These work for their specific destination; not relevant for KubeVirt-based migration to open destinations.</p>
<h2 id="strategy-for-vmware-migration">Strategy for VMware migration</h2>
<h3 id="step-1-discovery-and-assessment">Step 1: Discovery and assessment</h3>
<p>Before tools, agree on:</p>
<ul>
<li><strong>Trigger</strong> — what&rsquo;s pushing the migration? (Pricing, sovereignty, scale, AI.)</li>
<li><strong>Workload portfolio</strong> — count, OS mix, criticality, dependencies.</li>
<li><strong>Destination architecture</strong> — what you&rsquo;re moving to.</li>
<li><strong>Timeline constraints</strong> — VCF subscription expirations.</li>
<li><strong>Team capacity</strong> — to run the destination after migration.</li>
</ul>
<p>Output: structured assessment with workload classification (migrate-now / later / stay / re-platform).</p>
<h3 id="step-2-destination-platform-engineered">Step 2: Destination platform engineered</h3>
<p>The destination platform must be production-ready before migration starts. This is where many migrations fail: workloads move to a destination that&rsquo;s been engineered as a PoC, not as a production platform.</p>
<p>For Cozystack-based destinations: 1-3 months destination-build before migration cohort 1.</p>
<h3 id="step-3-cohort-based-migration">Step 3: Cohort-based migration</h3>
<p>Cohorts of 10-50 workloads at a time. Each cohort:</p>
<ol>
<li>Pre-migration validation (workload state, dependencies)</li>
<li>Image conversion</li>
<li>Network and storage cutover</li>
<li>Post-migration validation</li>
<li>Parallel-run (VMware + destination both active)</li>
<li>Customer signoff</li>
<li>VMware-side decommission</li>
</ol>
<p>Cohort cadence: 1-3 cohorts per month at steady state.</p>
<h3 id="step-4-vcf-expiration-aligned-sequencing">Step 4: VCF expiration-aligned sequencing</h3>
<p>Workloads move when VCF commitments lapse. Moving early into a paid-for commitment costs money. Moving late risks expiration without prepared destination.</p>
<h3 id="step-5-decommission">Step 5: Decommission</h3>
<p>VMware-side decommission as cohorts complete. Hardware repurposed where applicable.</p>
<h2 id="where-migrations-stumble">Where migrations stumble</h2>
<h3 id="stumble-1-destination-not-ready">Stumble 1: destination not ready</h3>
<p>Workloads moved to destination that&rsquo;s not actually production-ready. Operational issues blamed on migration.</p>
<h3 id="stumble-2-big-bang-cutover-attempted">Stumble 2: big-bang cutover attempted</h3>
<p>Weekend &ldquo;we&rsquo;ll move it all&rdquo; — rarely works at enterprise scale. Cohort approach is the proven pattern.</p>
<h3 id="stumble-3-data-gravity-ignored">Stumble 3: data gravity ignored</h3>
<p>50TB production database moves are not a weekend job. Cross-network movement, cutover windows, dual-write periods all need engineering.</p>
<h3 id="stumble-4-networking-redesign-skipped">Stumble 4: networking redesign skipped</h3>
<p>Cilium ≠ NSX. Network policies, ingress patterns, service-to-service auth all need redesign.</p>
<h3 id="stumble-5-post-migration-capacity-underestimated">Stumble 5: post-migration capacity underestimated</h3>
<p>Migration completes; platform team thinks they&rsquo;re done. Actually, post-migration operations is when the workload starts. Capacity for ongoing ops is non-negotiable.</p>
<h2 id="cost-ranges">Cost ranges</h2>
<p>For planning:</p>
<ul>
<li><strong>Assessment:</strong> <a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a>, fixed price, 14 days focused or 28 days full.</li>
<li><strong>Destination platform foundation:</strong> live in weeks once hardware is ready; up to 3 months with integrations, depending on scale.</li>
<li><strong>Migration cohort labor:</strong> 8-15 person-days per cohort of 10-50 VMs.</li>
<li><strong>Total elapsed:</strong> about 8-12 months for ~100 VMs and 18-24 months for ~1,000 VMs, including planning and migration waves; mid-size estates fall in between, depending on dependencies.</li>
</ul>
<p>Compared to ongoing VCF subscription: most customer engagements show net positive after Year 2 even accounting for migration cost.</p>
]]></content:encoded></item><item><title>Transport and logistics cloud architecture — NIS2, AI, edge in 2026</title><link>https://aenix.io/blog/2026/05/transport-logistics-cloud-architecture-nis2/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/transport-logistics-cloud-architecture-nis2/</guid><pubDate>Fri, 29 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>NIS2</category><category>Cozystack</category><category>Sovereignty</category><category>AI and ML</category><category>GPU</category><description>A three-tier architecture for transport and logistics, the NIS2 controls that apply to the sector, and where AI workloads fit.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/transport-logistics-cloud-architecture-nis2.jpg" alt=""></p><h2 id="three-pressures">Three pressures</h2>
<ol>
<li><strong>NIS2 scope</strong> — transport is an Annex I sector (essential or important entity depending on size); Article 21/23 apply to ICT</li>
<li><strong>AI optimization</strong> — routing, demand forecasting, predictive maintenance</li>
<li><strong>Edge compute density</strong> — depots, ports, terminals, vehicles</li>
</ol>
<h2 id="three-tier-architecture-pattern">Three-tier architecture pattern</h2>
<ul>
<li><strong>HQ Cloud</strong> — TMS, fleet management, AI training, customer-facing</li>
<li><strong>Regional sites</strong> — operational centres, regional dispatch, AI inference</li>
<li><strong>Edge</strong> — depots, ports, terminals, on-vehicle compute</li>
</ul>
<p>Cozystack at all three tiers, with OT boundary at edge for safety-critical systems (rail signalling, automated terminal operations).</p>
<h2 id="nis2-architecture-controls-for-transport">NIS2 architecture controls for transport</h2>
<p>Standard Article 21 + 23 mapping; transport-specific:</p>
<ul>
<li>Multi-modal data sovereignty (cross-border freight data)</li>
<li>Sub-contractor visibility (logistics chains often have 5+ levels)</li>
<li>BCP for kinetic disruption scenarios</li>
<li>Air-gap for safety-critical OT</li>
</ul>
<h2 id="ai-use-cases-in-transport">AI use cases in transport</h2>
<ul>
<li>Route optimization (real-time + planning)</li>
<li>Demand forecasting</li>
<li>Predictive maintenance for fleet</li>
<li>Last-mile optimization</li>
<li>Customer-facing AI (delivery ETA, customer service)</li>
</ul>
<p>Most workloads are sustained 24/7 inference where dedicated GPU economics fit.</p>
<h2 id="how-ænix-engages">How Ænix engages</h2>
<p>Standard <strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong> with transport workstream emphasis.</p>
]]></content:encoded></item><item><title>Telco cloud modernization in 2026 — from legacy NFV to Kubernetes-native edge</title><link>https://aenix.io/blog/2026/05/telco-cloud-edge-nfv-modernization/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/telco-cloud-edge-nfv-modernization/</guid><pubDate>Thu, 28 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Telco</category><category>Sovereignty</category><category>Multi-tenancy</category><category>Cozystack</category><category>Cloud</category><category>AI and ML</category><description>How tier-1 and tier-2 telecom operators modernize legacy NFV environments into Kubernetes-native sovereign cloud platforms — a guide for architects.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/telco-cloud-edge-nfv-modernization.jpg" alt=""></p><p>The telco cloud conversation in 2026 sits at an unusual intersection
of pressures that no other vertical faces simultaneously. NFV
deployments from 2015-2020 are aging out. NIS2 obligations apply
across the operator&rsquo;s infrastructure. Sovereign-cloud branded
products are commercially attractive (regional sovereignty
positioning is hard for hyperscalers to replicate). Edge-compute
demand for 5G/6G workloads is growing. AI for network operations
(traffic prediction, anomaly detection, customer-facing assistants)
is emerging.</p>
<p>The architectural answer is rarely &ldquo;one platform for everything&rdquo; —
telcos have multiple operational domains, each with its own
constraints. But the patterns that work across all of them have
converged on Kubernetes-native multi-site platforms with strong
multi-tenant isolation, sovereign operation, and edge-aware
scheduling.</p>
<h2 id="what-modernization-actually-replaces">What modernization actually replaces</h2>
<p>Most tier-1 European telcos in 2026 operate three or four parallel
infrastructure environments:</p>
<h3 id="1-legacy-nfv-environment-2015-2020-vintage">1. Legacy NFV environment (2015-2020 vintage)</h3>
<p>Typically built on VMware Cloud Foundation, OpenStack-based vendor
distros (Red Hat OSP, Mirantis Cloud Platform, Wind River, etc.), or
vendor-specific NFVI (Ericsson CEE, Nokia CloudBand). Hosts VNFs from
network-equipment vendors with vendor-specific certification
requirements.</p>
<p>Modernization driver: vendor distribution lifecycles (Red Hat OSP
EOL, Mirantis transition), VMware Broadcom pricing, operational
complexity that compounds with team turnover. Most tier-1 operators
have parallel modernization programmes running.</p>
<h3 id="2-it-cloud-separate-from-nfv">2. IT cloud (separate from NFV)</h3>
<p>For business workloads, customer-facing portals, billing, OSS/BSS.
Typically a separate VMware estate or hyperscaler-region deployment.
Same Broadcom pressure on VMware; sovereignty pressure on hyperscaler
deployment.</p>
<h3 id="3-edge-compute-5g6g-era">3. Edge compute (5G/6G era)</h3>
<p>Distributed edge sites at substations, central offices, MEC nodes.
Smaller-footprint deployments with intermittent connectivity to core.
Often built on different stack from core platform (legacy small-
footprint orchestration).</p>
<h3 id="4-ai--data-lake--network-analytics">4. AI / data lake / network analytics</h3>
<p>Newer environment for traffic prediction, anomaly detection, customer-
facing AI. Often hyperscaler-based today; under sovereignty pressure
to move on-prem for sensitive workload patterns.</p>
<h2 id="the-single-platform-vision-and-where-it-works">The single-platform vision (and where it works)</h2>
<p>The architectural attraction of Cozystack-based modernization is a
unified platform across all four environments. One operational model,
one upgrade lifecycle, one observability stack, federated identity,
GitOps-driven changes.</p>
<p>Where this works:</p>
<ul>
<li><strong>IT cloud workloads</strong> — straightforward Cozystack-based deployment.
Most tier-1 telco IT cloud modernizations follow this pattern.</li>
<li><strong>AI / data lake / analytics</strong> — Ænix AI Platform workload
patterns; sovereign by default; multi-tenant for cross-BU access.</li>
<li><strong>Edge compute</strong> — Cozystack supports small-footprint edge
deployments with federation to core. Standardised across sites.</li>
</ul>
<p>Where it requires nuance:</p>
<ul>
<li><strong>NFV environment specifically</strong> — VNF certification with vendor
equipment is still vendor-managed. The Cozystack platform underneath
can host KubeVirt-based VNFs, but the certification stack remains
vendor-bound. Some VNFs are certified on specific OpenStack
distributions and need a parallel modernization track.</li>
<li><strong>Critical SCADA-style network OAM</strong> — air-gapped network management
systems with their own operational model. Cozystack can host them,
but the OT-style operational discipline differs from IT cloud.</li>
</ul>
<p>The practical pattern: Cozystack as the IT cloud platform + AI
platform + edge platform; NFV gets its own modernization track with
Cozystack-equivalent architecture where vendor certification allows.</p>
<h2 id="multi-tenancy-for-telco">Multi-tenancy for telco</h2>
<p>Telco operators have layered multi-tenancy needs:</p>
<ul>
<li><strong>Operational BU separation</strong> — fixed broadband, mobile, enterprise,
consumer; each BU has different operational ownership.</li>
<li><strong>Customer-facing services</strong> — telco offers cloud capacity to its
business customers (sovereign cloud product line); each customer
is its own tenant.</li>
<li><strong>Internal vs external workloads</strong> — internal-only workloads (OAM,
observability backends) versus customer-facing workloads (cloud
product, customer portal).</li>
<li><strong>Sectoral / regulated tenants</strong> — telcos increasingly host
regulated workloads (financial sector adjacencies, public-sector
contracts) under sectoral tenant isolation.</li>
</ul>
<p>Cozystack&rsquo;s Tenant CRD model with nested tenants supports all four
layers natively. A tier-1 telco typically operates 5-50 top-level
tenants with hundreds to thousands of nested tenants.</p>
<h2 id="sovereignty-positioning">Sovereignty positioning</h2>
<p>For telcos, sovereignty is not just compliance — it&rsquo;s a commercial
differentiator. European customers increasingly view US-headquartered
hyperscaler dependencies as structural risk. Telcos can offer sovereign
cloud products in their region that hyperscaler-managed offerings
cannot match on the substantive sovereignty criteria.</p>
<p>Cozystack-based architecture supports this commercially:</p>
<ul>
<li><strong>Encryption with a customer-held passphrase</strong> — opt-in volume
encryption at rest (LINSTOR and LUKS); the key-management process
is designed with the telco, which provides operational support</li>
<li><strong>Air-gap option</strong> — for classified or restricted-egress use cases</li>
<li><strong>Open-source substrate</strong> — exit-readiness built in; telco doesn&rsquo;t
lock customers into vendor relationship</li>
<li><strong>EU jurisdiction</strong> — telco&rsquo;s EU presence + Ænix EU contracting
entity (AENIX s.r.o.)</li>
</ul>
<p>For sovereign cloud commercial product lines, the Ænix Public Cloud
Platform is the typical pairing — multi-region, multi-DC,
service-catalog depth, brand-engineered customer portal.</p>
<h2 id="edge-compute-realities">Edge compute realities</h2>
<p>5G/6G drove distributed compute requirements that NFV-era architecture
didn&rsquo;t anticipate. Modern edge compute for telco operations spans:</p>
<ul>
<li><strong>MEC (Multi-access Edge Computing) nodes</strong> at base-station or
central-office locations for low-latency customer workloads</li>
<li><strong>Distributed RAN intelligence</strong> — vRAN / O-RAN workloads with
real-time constraints</li>
<li><strong>Customer-edge compute</strong> — telco-managed edge instances on customer
premises (factories, smart-grid sites, transport hubs)</li>
<li><strong>Sectoral edge</strong> — telco-operated edge for sectoral customers
(financial branch offices, healthcare facilities, retail)</li>
</ul>
<p>Cozystack supports edge deployment with reduced footprint (~3-node
clusters at edge sites typical) with federation to regional and core
platforms. Same operational model; same observability; differentiated
service catalog per tier.</p>
<h2 id="ai-for-telco-operations">AI for telco operations</h2>
<p>The AI workload patterns at tier-1 telcos:</p>
<ul>
<li><strong>Traffic prediction and capacity planning</strong> — 24/7 inference on
network telemetry</li>
<li><strong>Anomaly detection</strong> — security and operational anomalies in real-
time</li>
<li><strong>Customer-facing AI</strong> — chatbot, billing assistance, customer
service routing</li>
<li><strong>Network optimization</strong> — routing, traffic engineering, energy
efficiency</li>
<li><strong>Fraud detection</strong> — transaction anomaly detection, identity
validation</li>
<li><strong>Sectoral AI services</strong> — telco-hosted AI capacity for sectoral
customers (banking AI, healthcare AI, public-sector AI)</li>
</ul>
<p>Sustained-utilisation workload profiles dominate; this is exactly the
case where dedicated GPU economics beat hyperscaler. The Ænix AI
Platform fits.</p>
<h2 id="phasing-a-tier-1-telco-modernization">Phasing a tier-1 telco modernization</h2>
<p>The cloud platform follows the operator-scale pattern: a 3-6 month
pilot, then 9-18 months to full multi-region operation. Edge expansion
and NFV modernization run as longer parallel tracks, paced by site
roll-outs and vendor lifecycles. Typical phasing:</p>
<h3 id="phase-0--strategic-engagement-start-of-the-pilot">Phase 0 — Strategic engagement (start of the pilot)</h3>
<p>Architecture review across all four environments. Commercial alignment
on sovereign cloud product line, sectoral positioning, modernization
sequencing. Sponsor and workstream-lead identification.</p>
<h3 id="phase-1--it-cloud-modernization">Phase 1 — IT cloud modernization</h3>
<p>Ænix Private Cloud Platform / Public Cloud Platform deployed. Internal
workloads migrated. Customer-facing portal launched for sovereign
cloud product.</p>
<h3 id="phase-2--ai--data-lake-parallel">Phase 2 — AI / data lake (parallel)</h3>
<p>Ænix AI Platform deployed. Network analytics workloads moved
on-prem. Customer-facing AI services launched.</p>
<h3 id="phase-3--edge-expansion-ongoing-paced-by-site-roll-out">Phase 3 — Edge expansion (ongoing, paced by site roll-out)</h3>
<p>Edge sites stood up at MEC / central-office / customer-edge locations.
Federated identity and observability across.</p>
<h3 id="phase-4--nfv-modernization-parallel-paced-by-vendor-lifecycles">Phase 4 — NFV modernization (parallel, paced by vendor lifecycles)</h3>
<p>Vendor-certified VNF replatforming where allowed. Greenfield deployments
of new VNFs on Cozystack-based architecture. Legacy NFV maintained
in parallel until vendor lifecycle forces upgrade.</p>
<p>Phase 5+ depends on customer-facing product growth: regional expansion,
sectoral SKUs, new service families.</p>
<h2 id="when-this-fits-a-telco">When this fits a telco</h2>
<p>Strong fit:</p>
<ul>
<li>Tier-1 or tier-2 European telecom operator</li>
<li>Active legacy NFV modernization programme</li>
<li>Strategic intent to offer sovereign cloud product line</li>
<li>Multi-million-euro budget envelope across multi-year programme</li>
<li>Senior executive sponsorship (CIO / CTO / Cloud BU lead)</li>
</ul>
<p>Marginal fit:</p>
<ul>
<li>Smaller operators (regional, MVNO-style) that sell cloud — Ænix
Public Cloud Platform at provider scale, from the <a href="https://aenix.io/pricing/">published price
list</a>, rather than a full operator-scale programme</li>
<li>Operators with deep OpenStack-based NFV investment that still
works — modernization can wait for vendor lifecycle to force it</li>
</ul>
<h2 id="where-to-dig-deeper">Where to dig deeper</h2>
<ul>
<li><strong><a href="https://aenix.io/industries/telco/">Telco industry page</a></strong> — commercial landing</li>
<li><strong><a href="https://aenix.io/products/public-cloud-platform/">Public Cloud Platform product page</a></strong> —
the product for telcos that sell cloud services</li>
<li><strong><a href="https://aenix.io/services/sovereign-cloud-builder/">Sovereign cloud builder services</a></strong> —
for sovereign cloud product line builds</li>
<li><strong><a href="https://aenix.io/solutions/sovereign-ai/">Sovereign AI services</a></strong> — for AI
workload patterns</li>
<li><strong><a href="https://aenix.io/blog/2026/05/public-cloud-edition-multi-tenant-cloud-builder/">Public Cloud Platform build phasing</a></strong> —
operator-scale build phasing</li>
<li><strong><a href="https://aenix.io/blog/2026/05/sovereign-ai-architecture-decisions/">Sovereign AI architecture decisions</a></strong> —
seven decisions for sovereign AI</li>
</ul>
]]></content:encoded></item><item><title>SRE as a product discipline — what an SRE engagement actually changes</title><link>https://aenix.io/blog/2026/05/sre-engagement-reliability-as-product-discipline/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/sre-engagement-reliability-as-product-discipline/</guid><pubDate>Thu, 28 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>DevOps</category><category>Platform Engineering</category><category>Observability</category><category>Cozystack</category><description>Embed SRE in product teams, centralize it as a function, or buy an engagement — what each delivers and how to measure it.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/sre-engagement-reliability-as-product-discipline.jpg" alt=""></p><p>SRE — Site Reliability Engineering — is one of the most-adopted and
most-misunderstood engineering disciplines of the past decade. Most
mid-to-large engineering organisations claim to &ldquo;do SRE.&rdquo; Far fewer
actually run the discipline as Google&rsquo;s original SRE book describes
it: software engineering applied to operations, with explicit error
budgets, SLOs that affect prioritisation, and a hard ceiling on
operational toil.</p>
<h2 id="what-sre-means-precisely">What SRE means, precisely</h2>
<p>The original Google formulation has three load-bearing properties:</p>
<h3 id="1-sre-is-software-engineering-applied-to-operations">1. SRE is software engineering applied to operations</h3>
<p>SREs are software engineers — they write code, they ship systems,
they automate operations. Not a relabelled ops team with the same
job and a fancier title. The &ldquo;is your SRE actually writing code?&rdquo;
test is a real test; teams that fail it have an ops function, not
an SRE function.</p>
<h3 id="2-error-budgets-affect-prioritisation">2. Error budgets affect prioritisation</h3>
<p>The SLO defines the bar. The error budget is the difference between
100% and the SLO target. When the error budget burns down, the product
team&rsquo;s prioritisation shifts — feature work pauses, reliability work
becomes the priority. When the error budget is healthy, feature work
can proceed.</p>
<p>Crucially, this is a <em>contract</em>, not a suggestion. SLO breach has
real consequences for product roadmap. Organisations that have SLOs
but don&rsquo;t let them affect roadmap have observability, not SRE
practice.</p>
<h3 id="3-operational-toil-is-capped-often-at-50">3. Operational toil is capped (often at 50%)</h3>
<p>The Google formulation: SREs cannot spend more than 50% of their
time on operational toil (repetitive, manual, automatable work).
The remainder must go to software engineering that reduces toil.
This makes the function self-improving over time — toil decreases,
engineering capacity increases.</p>
<p>Teams that exceed the toil cap structurally have an ops function
masquerading as SRE. The discipline degrades over months until the
function is indistinguishable from traditional ops.</p>
<h2 id="three-engagement-models">Three engagement models</h2>
<p>In 2026, three structurally different SRE models exist in production:</p>
<h3 id="model-1-embedded-sre">Model 1: embedded SRE</h3>
<p>SREs sit inside product teams. The product team includes 2-5
engineers, one of whom holds the SRE function. Same team owns
features and reliability; the SRE shapes the architecture and
on-call from inside.</p>
<p>Strengths: alignment between SRE and product priorities, fast
feedback, no cross-team friction.</p>
<p>Weaknesses: SRE-quality varies per team, depending on the embedded
person. Centralised SRE expertise doesn&rsquo;t compound. Hard to scale
across many product teams without losing consistency.</p>
<p>Fits: 100-300-engineer organisations where each product team has
clear ownership and ~1 SRE-skilled engineer is available per team.</p>
<h3 id="model-2-centralised-sre-function">Model 2: centralised SRE function</h3>
<p>SREs sit in a separate function with their own reporting line.
They consult with product teams, set reliability standards, run
shared infrastructure components, and operate the critical-path
services.</p>
<p>Strengths: SRE expertise compounds. Standards are consistent across
teams. Critical-path services have dedicated reliability owners.</p>
<p>Weaknesses: cross-team friction when SRE recommendations conflict
with product team priorities. Risk of SRE-as-blocker rather than
SRE-as-enabler.</p>
<p>Fits: 300+-engineer organisations with multiple critical-path
services and a clear platform-engineering function. Often paired
with platform engineering as a sister function.</p>
<h3 id="model-3-hybrid-embedded--centralised">Model 3: hybrid (embedded + centralised)</h3>
<p>Embedded SREs in product teams handle team-specific reliability.
Centralised SRE function operates shared infrastructure, sets
standards, escalates for cross-team incidents.</p>
<p>Strengths: gets benefits of both. Most common pattern at large
organisations (Google, Netflix, large fintech).</p>
<p>Weaknesses: requires substantial scale to justify the dual
investment. Below 500 engineers, the overhead is hard to amortise.</p>
<p>Fits: 500+-engineer organisations with multiple business units and
substantial reliability-sensitive workload portfolio.</p>
<h2 id="where-sre-shows-up-as-theatre">Where SRE shows up as theatre</h2>
<p>Patterns we see in assessments:</p>
<h3 id="theatre-1-slos-exist-but-dont-affect-roadmap">Theatre 1: SLOs exist, but don&rsquo;t affect roadmap</h3>
<p>SLOs are documented. Dashboards show SLO compliance. Quarterly
reviews mention SLO trends. But product team prioritisation
proceeds independently of error-budget state. When SLO breach
happens, the response is incident-specific firefighting, not a
roadmap re-prioritisation.</p>
<p>This is observability with SLO labels, not SRE practice. The
distinction matters for buying decisions: organisations in this
state don&rsquo;t need more dashboards, they need a different governance
model.</p>
<h3 id="theatre-2-sre-title-applied-to-traditional-ops">Theatre 2: SRE title applied to traditional ops</h3>
<p>Operations engineers get retitled &ldquo;SRE&rdquo; without a change in job
content. They still spend 90% of their time on ticket queue. They
write no code beyond shell scripts. They have no error-budget
authority.</p>
<p>This is a label change, not a discipline shift. Often happens
during DevOps-to-SRE transitions that didn&rsquo;t get executive sponsor
investment.</p>
<h3 id="theatre-3-error-budgets-defined-but-never-invoked">Theatre 3: error budgets defined but never invoked</h3>
<p>Error budgets are calculated. Some dashboards show them. But no
process exists for what happens when they burn down. Effectively
the same as no error budget.</p>
<p>Fix: write down the explicit feature-pause protocol that activates
on error budget burn. Get product leadership to sign off on it.
Test it once with an artificial burn-down before assuming it works
in production.</p>
<h3 id="theatre-4-incident-post-mortems-without-action-items">Theatre 4: incident post-mortems without action items</h3>
<p>Post-mortems happen after incidents. They get written. They go
into a folder. No action items are tracked to completion. The same
incident class recurs within 6 months.</p>
<p>Fix: post-mortem action items go into the same backlog as feature
work, with named owners, due dates, and explicit prioritisation.
The post-mortem hasn&rsquo;t worked if its action items don&rsquo;t ship.</p>
<h2 id="what-ænix-sre-engagement-delivers">What Ænix SRE engagement delivers</h2>
<p>We typically engage with organisations where SRE practice is at
one of three states:</p>
<ul>
<li><strong>Pre-SRE</strong> — no formal SRE function; reliability is incident-
driven firefighting. Engagement covers function definition,
hiring plan, initial SLO definition for critical services.</li>
<li><strong>Theatre-state SRE</strong> — SRE titles and dashboards exist, but
discipline hasn&rsquo;t landed. Engagement diagnoses which theatre
patterns are operating, designs corrections, often involves
cross-functional governance work.</li>
<li><strong>Mature SRE needing expansion</strong> — discipline works in one
business unit, needs to scale across the organisation or pick up
new service families (AI/GPU workloads, edge compute, sovereign
cloud). Engagement focuses on consistency and expertise transfer.</li>
</ul>
<h3 id="workstream-1--slo-definition">Workstream 1 — SLO definition</h3>
<p>For each critical service, define:</p>
<ul>
<li>SLI (Service Level Indicator) — what we measure</li>
<li>SLO (Service Level Objective) — the target</li>
<li>Error budget — the difference between 100% and the SLO</li>
<li>Burn-down policy — what happens when the budget is being consumed</li>
<li>Recovery threshold — what restores feature-work prioritisation</li>
</ul>
<p>Ænix doesn&rsquo;t define SLOs for you in isolation — we facilitate the
workshop where engineering and product leadership co-author them.
SLOs without joint ownership don&rsquo;t stick.</p>
<h3 id="workstream-2--incident-response-process">Workstream 2 — Incident response process</h3>
<p>Roles (incident commander, scribe, communicator). Severity
classification. Runbook structure. Escalation paths. Blameless
post-mortem template. Action-item tracking.</p>
<p>This is often the highest-leverage workstream — the framework
multiplies the effectiveness of every future incident.</p>
<h3 id="workstream-3--toil-measurement-and-reduction">Workstream 3 — Toil measurement and reduction</h3>
<p>Inventory current SRE / ops team work. Categorise: toil (repetitive,
manual, automatable) versus engineering (durable, automation-
generating). Measure toil ratio.</p>
<p>Recommend the 3-5 automations that reduce the most toil. Stage
them by ROI. Hand off to engineering team for implementation; we
support if needed.</p>
<h3 id="workstream-4--observability-stack">Workstream 4 — Observability stack</h3>
<p>Ænix&rsquo;s default observability recommendation: VictoriaMetrics for
metrics, VictoriaLogs for logs, OpenTelemetry for tracing where
applicable. Self-hosted (sovereignty-friendly, lower-overhead than
Prometheus + Loki at scale, no SaaS vendor data-residency leak).</p>
<p>For organisations already on a different stack, we work with what&rsquo;s
in place rather than push replacement. Observability stack matters;
SRE discipline matters more.</p>
<h3 id="workstream-5--function-design">Workstream 5 — Function design</h3>
<p>Embedded / centralised / hybrid model recommendation per the
engineering organisation&rsquo;s profile. Headcount planning. Hiring
priorities. Reporting line. Interface with platform engineering
(if separate function) and product teams.</p>
<h2 id="the-cozystack-reliability-defaults">The Cozystack reliability defaults</h2>
<p>For organisations running an Ænix platform product, SRE practice
gets a head start because the platform ships with SRE-aligned
defaults:</p>
<ul>
<li><strong>Observability built in</strong> — VictoriaMetrics + VictoriaLogs
pre-deployed, with alert rules for the platform components</li>
<li><strong>Declarative change history</strong> — tenants and services are
Kubernetes resources managed through GitOps, so every change has a
commit, an author and a review</li>
<li><strong>Backup-restore patterns</strong> — Velero + per-app PITR, with RPO / RTO
documented per service during the engagement</li>
</ul>
<p>SLO definitions and failure-injection (chaos) tooling are not shipped
as platform features; they are designed with you in the engagement.</p>
<p>This lets the SRE engagement focus on the organisation-specific
work (function design, SLOs aligned to business priorities,
governance) rather than rebuilding the technical substrate.</p>
<h2 id="when-this-engagement-fits">When this engagement fits</h2>
<p>Strong fit:</p>
<ul>
<li>Engineering organisation 200+ engineers with reliability
becoming a board-level concern</li>
<li>Recent incident pattern that exposed reliability gaps</li>
<li>Regulator-driven RTO/RPO obligations (DORA Articles 11-12, NIS2
Article 21(2)(c))</li>
<li>Existing observability investment but no clear SRE discipline</li>
<li>Platform engineering function exists or is being built (SRE pairs
naturally with platform engineering)</li>
</ul>
<p>Marginal fit:</p>
<ul>
<li>Smaller organisations (&lt;100 engineers) — usually embedded SRE
rather than separate function; Ænix engagement can be lighter
(workshop + advisory rather than full multi-month engagement)</li>
<li>Organisations with mature SRE in one BU needing expansion —
scope can be narrower</li>
</ul>
<p>Poor fit:</p>
<ul>
<li>Pure firefighting culture without engineering leadership
sponsorship for the discipline shift — SRE engagement without
executive backing degrades to incident response training, which
is helpful but not what we sell</li>
</ul>
<h2 id="where-to-dig-deeper">Where to dig deeper</h2>
<ul>
<li><strong><a href="https://aenix.io/services/sre-consulting/">SRE consulting services</a></strong> — the
commercial landing</li>
<li><strong><a href="https://aenix.io/blog/2026/05/devops-best-practices-2026/">DevOps best practices for 2026</a></strong> —
the eight DevOps practices including SRE</li>
<li><strong><a href="https://aenix.io/blog/2026/05/platform-engineering-vs-devops-vs-sre/">Platform engineering vs DevOps vs SRE</a></strong> —
terminology and function design</li>
<li><strong><a href="https://aenix.io/services/cloud-engineering/">Cloud engineering disciplines</a></strong> —
the seven disciplines that compound</li>
</ul>
]]></content:encoded></item><item><title>Cozystack 1.4: New Dashboard UI, Persistent Tenant Workers, Backup Strategies, and Fractional GPU Sharing</title><link>https://aenix.io/blog/2026/05/cozystack-1-4-new-dashboard-ui-persistent-tenant-workers-backup-strategies-and-fractional-gpu-sharing/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/cozystack-1-4-new-dashboard-ui-persistent-tenant-workers-backup-strategies-and-fractional-gpu-sharing/</guid><pubDate>Thu, 28 May 2026 00:00:00 +0000</pubDate><dc:creator>Timur Tukaev</dc:creator><category>Cozystack</category><category>Kubernetes</category><category>KubeVirt</category><category>GPU</category><category>Multi-tenancy</category><category>Talos</category><description>Cozystack v1.4.0 is now available. The release was published on May 19, 2026, and rolls up every fix shipped in the v1.3.1 to v1.3.3 patch line.</description><content:encoded><![CDATA[<blockquote>
<p><strong>Editor’s note (September 2026):</strong> this release announcement is preserved as it was published on May 28, 2026. Cozystack has moved on since — the current release line is 1.6. For the latest patch release, see the <a href="https://github.com/cozystack/cozystack/releases">Cozystack releases</a> and the <a href="https://cozystack.io/docs/">current documentation</a>. Documentation links below deliberately point at the v1.4 docs, which describe the release as shipped.</p>
</blockquote>
<p>Cozystack v1.4.0 is now available. The release was published on May 19, 2026, and rolls up every fix shipped in the v1.3.1 to v1.3.3 patch line.</p>
<p>This cycle is focused on the operational experience of running Cozystack as a production platform: a faster dashboard architecture, more durable tenant Kubernetes workers, clearer resource sizing, backup workflows for managed applications, better GPU utilization, safer ingress publishing, and fewer race conditions during first installs and upgrades.</p>
<p><img src="https://aenix.io/img/blog/medium/cozystack-1-4-new-dashboard-ui-persistent-tenant-workers-backup-strategies-and-fractional-gpu-sharing/cover.jpg" alt="Cozystack 1.4 release" width="1200" height="630" loading="lazy" decoding="async"></p>
<h2 id="main-highlights">Main highlights</h2>
<h3 id="new-schema-driven-dashboard-ui">New schema-driven dashboard UI</h3>
<p>Cozystack 1.4 ships a rewritten dashboard from the <code>cozystack/cozystack-ui</code> project. The old <code>openapi-ui</code> plus BFF stack has been replaced by a React 19 and TypeScript frontend that talks directly to the Kubernetes API.</p>
<p><img src="https://aenix.io/img/blog/medium/cozystack-1-4-new-dashboard-ui-persistent-tenant-workers-backup-strategies-and-fractional-gpu-sharing/02.png" alt="The new Cozystack dashboard" width="1915" height="672" loading="lazy" decoding="async"></p>
<p>The new architecture removes an extra process and proxy layer while keeping the dashboard schema-driven. It also improves several day-to-day workflows:</p>
<ul>
<li>VNC access for virtual machines now uses dynamic WebSocket URLs instead of deployment-specific <code>localhost</code> assumptions.</li>
<li>The dashboard can read <code>ApplicationDefinition</code> resources for the application catalog and marketplace.</li>
<li>Operators can inject runtime branding through a ConfigMap, including logos, names, and brand colors, without rebuilding the image.</li>
<li>Existing <code>/openapi-ui/*</code> bookmarks are redirected to the new console.</li>
<li>The package has been renamed consistently to <code>cozy-dashboard</code>.</li>
</ul>
<p>The new dashboard exposes the IaaS marketplace driven by ApplicationDefinition resources.</p>
<p><img src="https://aenix.io/img/blog/medium/cozystack-1-4-new-dashboard-ui-persistent-tenant-workers-backup-strategies-and-fractional-gpu-sharing/03.png" alt="The IaaS marketplace in the new Cozystack dashboard" width="1902" height="874" loading="lazy" decoding="async"></p>
<p>The PaaS catalog covers managed databases, messaging, object storage, secrets, search, and inference services.</p>
<p>Deploying a managed Kubernetes cluster uses the same schema-driven form, with cluster addons exposed declaratively.</p>
<p>A managed HTTP cache deployment, with PVC sizing, storage class, endpoints, and resource parameters all generated from the application schema.</p>
<p>Documentation:</p>
<ul>
<li><a href="https://cozystack.io/docs/v1.4/getting-started/deploy-app/">Deploying applications through the new dashboard</a></li>
<li><a href="https://cozystack.io/docs/v1.4/cozystack-api/application-definitions/">ApplicationDefinition reference</a></li>
<li><a href="https://cozystack.io/docs/v1.4/operations/configuration/white-labeling/">White-labeling and runtime branding</a></li>
</ul>
<h3 id="persistent-worker-node-storage-for-tenant-kubernetes">Persistent worker-node storage for tenant Kubernetes</h3>
<p>Tenant Kubernetes worker VMs now use PVC-backed persistent disks through KubeVirt <code>dataVolumeTemplates</code>. Previously, workers used ephemeral <code>emptyDisk</code> storage, which meant kubelet certificates, kubeconfig, and containerd state were lost after a VM reboot. A rebooted worker could lose its identity and require manual recovery.</p>
<p>In v1.4, worker state survives VM restarts. The NodeGroup field <code>ephemeralStorage</code> is renamed to <code>diskSize</code>, and a new per-nodeGroup <code>storageClass</code> option lets operators control where worker disks are provisioned. Migration 39 rewrites legacy values automatically during upgrade.</p>
<p>Existing tenant clusters will roll worker nodes once because the KubeVirt machine template changes. Operators should plan capacity for the rollout and choose the storage class intentionally. For many worker-node disk scenarios, the <code>local</code> StorageClass is recommended because worker disks now survive restarts and do not need DRBD replication semantics.</p>
<p>Documentation: <a href="https://cozystack.io/docs/v1.4/kubernetes/">Tenant Kubernetes configuration</a>.</p>
<h3 id="instance-type-resource-presets">Instance-type resource presets</h3>
<p>Resource presets now follow a cloud-style <code>&lt;series&gt;.&lt;size&gt;</code> taxonomy. The new model covers five CPU-to-memory series:</p>
<ul>
<li><code>t1</code> for tiny and low-memory workloads.</li>
<li><code>c1</code> for compute-balanced workloads.</li>
<li><code>s1</code> for standard services such as proxies and caches.</li>
<li><code>u1</code> for universal workloads such as databases and messaging.</li>
<li><code>m1</code> for memory-heavy workloads such as search and analytics.</li>
</ul>
<p>Each series includes eight sizes from <code>nano</code> to <code>4xlarge</code>, giving operators and tenants 40 presets in total.</p>
<p>The previous flat names such as <code>small</code>, <code>medium</code>, and <code>large</code> are still accepted as deprecated aliases. Existing deployments keep the same CPU and memory values, while Migration 39 rewrites stored values to the new names. The Cozystack API now emits deprecation warnings when app CRs still use legacy preset names.</p>
<p>Documentation: <a href="https://cozystack.io/docs/v1.4/guides/resource-management/">Resource presets</a>.</p>
<h3 id="declarative-backup-strategies-for-managed-applications">Declarative backup strategies for managed applications</h3>
<p>The backup strategy controller now supports PostgreSQL, MariaDB, ClickHouse, and FoundationDB. Tenants can define a strategy together with <code>BackupClass</code>, <code>Plan</code>, <code>BackupJob</code>, and <code>RestoreJob</code> resources, while the controller composes the backend-specific objects for each managed service.</p>
<p>The new strategies support scheduled backups, ad-hoc snapshots, in-place restores, and restore-to-copy workflows against S3-compatible object storage. Credentials are referenced through Kubernetes Secrets instead of being stored inline, and controller RBAC is constrained so it can only access explicitly referenced secrets.</p>
<p>This extends the existing backup flows for VMInstance and VMDisk and moves Cozystack closer to full backup coverage across the managed application catalog.</p>
<p>Documentation:</p>
<ul>
<li><a href="https://cozystack.io/docs/v1.4/operations/services/managed-app-backup-configuration/">Managed app backup configuration</a></li>
<li><a href="https://cozystack.io/docs/v1.4/applications/backup-and-recovery/">Application backup and recovery</a></li>
</ul>
<h3 id="hami-based-fractional-gpu-sharing">HAMi-based fractional GPU sharing</h3>
<p>Cozystack 1.4 adds <code>hami</code> as an optional system package. HAMi v2.8.1, a CNCF Sandbox project, provides fractional GPU sharing for tenant Kubernetes clusters.</p>
<p>With HAMi enabled, tenant workloads can request resources such as <code>nvidia.com/gpu</code>, <code>nvidia.com/gpumem</code>, and <code>nvidia.com/gpucores</code>, allowing multiple pods to share one physical NVIDIA GPU with explicit memory and compute slicing. The integration includes the device plugin, scheduler extender, mutating webhook, and RuntimeClass. It is exposed through an opt-in <code>hami.enabled</code> toggle and depends on the NVIDIA GPU Operator.</p>
<p>There is one important compatibility note: HAMi compute isolation depends on container images with glibc older than 2.34. Memory enforcement works broadly, but Alpine and musl-based images are not supported for HAMi-core compute isolation.</p>
<p>Documentation: <a href="https://cozystack.io/docs/v1.4/kubernetes/gpu-sharing/">GPU sharing with HAMi</a>.</p>
<h3 id="one-switch-for-proxy-protocol-and-hairpin-nat">One switch for PROXY protocol and hairpin NAT</h3>
<p>The new <code>publishing.proxyProtocol: true</code> option enables PROXY protocol on the host ingress-nginx and deploys Ouroboros to solve the related hairpin-NAT problem.</p>
<p>When PROXY protocol is enabled, in-cluster traffic to the cluster’s own public hostnames can otherwise reach ingress-nginx without the required PROXY header. Ouroboros fixes that path through CoreDNS rewrite snippets. Cozystack exposes it as both a host-level system package and a per-tenant addon through <code>addons.ouroboros.enabled</code>.</p>
<p>The default behavior is unchanged. Clusters that do not enable PROXY protocol do not receive new resources.</p>
<p>Documentation: <a href="https://cozystack.io/docs/v1.4/networking/hairpin-proxy-protocol/">PROXY protocol and hairpin NAT</a>.</p>
<h3 id="better-helmrelease-behavior-and-tenant-bootstrap-reliability">Better HelmRelease behavior and tenant bootstrap reliability</h3>
<p>The Cozystack operator now exposes HelmRelease generation knobs as operator flags and chart values, including interval, retry interval, install timeout, upgrade timeout, and max history.</p>
<p>The retry strategy now uses <code>RetryOnFailure</code>, which avoids uninstall-and-reinstall loops when a first install is slow. Applications can also set a per-Application install and upgrade timeout through the <code>release.cozystack.io/helm-install-timeout</code> annotation. Tenant Kubernetes uses this to give Kamaji enough time during cold bootstrap, fixing the recurring <code>wait hr/tenant-kubernetes timeout</code> failure mode.</p>
<p>Documentation:</p>
<ul>
<li><a href="https://cozystack.io/docs/v1.4/kubernetes/">Tenant Kubernetes operations</a></li>
<li><a href="https://cozystack.io/docs/v1.4/operations/troubleshooting/flux-cd/">Troubleshooting Flux CD</a></li>
</ul>
<h3 id="kubelet-reservations-for-worker-nodes">Kubelet reservations for worker nodes</h3>
<p>Tenant Kubernetes worker nodes now receive automatically computed kubelet reservations for CPU and memory. This keeps kubelet itself from being targeted under memory pressure and makes scheduler and autoscaler decisions more accurate.</p>
<p>Cluster-autoscaler annotations now report allocatable CPU and memory instead of raw totals, so autoscaling decisions match what Kubernetes can actually schedule.</p>
<p>Documentation: <a href="https://cozystack.io/docs/v1.4/kubernetes/">Tenant Kubernetes operations</a>.</p>
<h2 id="also-in-v140">Also in v1.4.0</h2>
<ul>
<li>PostgreSQL parameters are now typed and protected by a denylist for dangerous values such as <code>archive_command</code>, <code>restore_command</code>, <code>ssl_passphrase_command</code>, <code>dynamic_library_path</code>, and <code>*_preload_libraries</code>.</li>
<li>Keycloak gains <code>extraEnv</code> and user profile customization support.</li>
<li>The etcd application exposes S3 backup schedules through the updated etcd-operator.</li>
<li>Per-package <code>upgradeCRDs</code> policy is now configurable.</li>
<li><code>cozyreport</code> now collects Flux, cert-manager, host context, application resources, and a top-level <code>summary.txt</code>.</li>
<li>SeaweedFS tenant ingress limits single PUT requests to 5 GB.</li>
<li>GPU observability dashboards and recording rules were added for Grafana and VictoriaMetrics.</li>
<li>VMInstance port filtering is fixed under the new cozy-proxy v0.3.0 mode.</li>
<li>LINSTOR CSI is updated with fixes for dual-attach and transient demotion errors.</li>
</ul>
<p>Documentation:</p>
<ul>
<li><a href="https://cozystack.io/docs/v1.4/applications/postgres/">PostgreSQL configuration</a></li>
<li><a href="https://cozystack.io/docs/v1.4/operations/oidc/">Keycloak and OIDC</a></li>
<li><a href="https://cozystack.io/docs/v1.4/operations/services/etcd/">etcd service configuration</a></li>
<li><a href="https://cozystack.io/docs/v1.4/operations/services/seaweedfs/">SeaweedFS service configuration</a></li>
<li><a href="https://cozystack.io/docs/v1.4/operations/services/monitoring/dashboards/">Monitoring dashboards</a></li>
<li><a href="https://cozystack.io/docs/v1.4/operations/troubleshooting/">Troubleshooting and diagnostics</a></li>
</ul>
<h2 id="platform-components">Platform components</h2>
<p>Cozystack 1.4 updates the platform base and several core packages:</p>
<ul>
<li>Talos: v1.12.7 to v1.13.0</li>
<li>cert-manager: v1.19.3 to v1.20.2</li>
<li>Cilium: v1.19.1 to v1.19.3</li>
<li>NVIDIA GPU Operator: v25.3.0 to v26.3.1</li>
<li>etcd-operator: v0.4.2 to v0.4.3</li>
<li>KubeVirt: v1.6.3 to v1.8.2</li>
<li>cozy-proxy: v0.2.0 to v0.3.0</li>
<li>linstor-csi: v1.10.6</li>
<li>HAMi: v2.8.1</li>
<li>Ouroboros: v0.7.2</li>
</ul>
<p>Documentation:</p>
<ul>
<li><a href="https://cozystack.io/docs/v1.4/guides/platform-stack/">Platform stack overview</a></li>
<li><a href="https://cozystack.io/docs/v1.4/operations/cluster/upgrade/">Upgrade guide</a></li>
</ul>
<h2 id="upgrade-notes">Upgrade notes</h2>
<p>Most operators can upgrade to v1.4.0 without manual configuration changes. Cozystack keeps the same API surface for existing workloads, and in-platform migrations handle the main value rewrites.</p>
<p>There are several operational details to plan for:</p>
<ul>
<li>Tenant Kubernetes workers will roll once. The <code>ephemeralStorage</code> to <code>diskSize</code> migration is automatic, but existing worker VMs are replaced one by one because the KubeVirt machine template changes.</li>
<li>KubeVirt VMs that were already running before the platform upgrade need a cold restart after upgrading. The KubeVirt jump from v1.6.3 to v1.8.2 crosses an upstream QEMU change, and live migration of pre-upgrade VMs can fail. New VMs created after the upgrade are unaffected.</li>
<li>Legacy resource preset names still work as deprecated aliases, but new deployments should use the <code>&lt;series&gt;.&lt;size&gt;</code> names.</li>
<li>PostgreSQL deployments that use denylisted parameters will fail to render until those parameters are removed.</li>
<li>cert-manager v1.20 changes the default container UID/GID to 65532. Operators with custom PodSecurityPolicy, imagePullSecrets, or filesystem-mounted certificates pinned to the previous UID should review their configuration.</li>
</ul>
<p>Documentation:</p>
<ul>
<li><a href="https://cozystack.io/docs/v1.4/operations/cluster/upgrade/">Upgrade guide</a></li>
<li><a href="https://cozystack.io/docs/v1.4/kubernetes/">Tenant Kubernetes operations</a></li>
<li><a href="https://cozystack.io/docs/v1.4/virtualization/">Virtualization operations</a></li>
<li><a href="https://cozystack.io/docs/v1.4/guides/resource-management/">Resource management</a></li>
<li><a href="https://cozystack.io/docs/v1.4/applications/postgres/">PostgreSQL configuration</a></li>
</ul>
<h2 id="documentation-worth-knowing-about">Documentation worth knowing about</h2>
<ul>
<li><a href="https://cozystack.io/docs/v1.4/getting-started/deploy-app/">New dashboard and application catalog</a></li>
<li><a href="https://cozystack.io/docs/v1.4/cozystack-api/application-definitions/">ApplicationDefinition reference</a></li>
<li><a href="https://cozystack.io/docs/v1.4/operations/configuration/white-labeling/">White-labeling and runtime branding</a></li>
<li><a href="https://cozystack.io/docs/v1.4/kubernetes/">Tenant Kubernetes configuration</a></li>
<li><a href="https://cozystack.io/docs/v1.4/guides/resource-management/">Resource presets</a></li>
<li><a href="https://cozystack.io/docs/v1.4/operations/services/managed-app-backup-configuration/">Managed app backup configuration</a></li>
<li><a href="https://cozystack.io/docs/v1.4/applications/backup-and-recovery/">Application backup and recovery</a></li>
<li><a href="https://cozystack.io/docs/v1.4/kubernetes/gpu-sharing/">GPU sharing with HAMi</a></li>
<li><a href="https://cozystack.io/docs/v1.4/networking/hairpin-proxy-protocol/">PROXY protocol and hairpin NAT</a></li>
<li><a href="https://cozystack.io/docs/v1.4/operations/cluster/upgrade/">Upgrade guide</a></li>
<li><a href="https://cozystack.io/docs/v1.4/">Cozystack v1.4 documentation</a></li>
</ul>
<h2 id="thank-you-to-all-contributors">Thank you to all contributors</h2>
<p>This release was shaped by the work of <a href="https://github.com/androndo">@androndo</a>, <a href="https://github.com/Arsolitt">@Arsolitt</a>, <a href="https://github.com/dislogical">@dislogical</a>, <a href="https://github.com/dvc">@dvc</a>, <a href="https://github.com/IvanHunters">@IvanHunters</a>, <a href="https://github.com/kvaps">@kvaps</a>, <a href="https://github.com/lexfrei">@lexfrei</a>, <a href="https://github.com/matthieu-robin">@matthieu-robin</a>, <a href="https://github.com/mattia-eleuteri">@mattia-eleuteri</a>, <a href="https://github.com/myasnikovdaniil">@myasnikovdaniil</a>, <a href="https://github.com/sircthulhu">@sircthulhu</a>, and <a href="https://github.com/tym83">@tym83</a>.</p>
<p>A special welcome to first-time contributors <a href="https://github.com/dvc">@dvc</a> and <a href="https://github.com/dislogical">@dislogical</a>. Thank you all.</p>
<h2 id="release-links">Release links</h2>
<ul>
<li><a href="https://github.com/cozystack/cozystack/releases/tag/v1.4.0">Cozystack v1.4.0 on GitHub</a></li>
<li><a href="https://github.com/cozystack/cozystack/compare/v1.3.0...v1.4.0">Full changelog v1.3.0 to v1.4.0</a></li>
<li><a href="https://github.com/cozystack/cozystack-ui">Cozystack UI</a></li>
<li><a href="https://github.com/Project-HAMi/HAMi">HAMi</a></li>
<li><a href="https://github.com/lexfrei/ouroboros">Ouroboros</a></li>
</ul>
<h2 id="join-the-community">Join the community</h2>
<ul>
<li>GitHub: <a href="https://github.com/cozystack/cozystack">cozystack/cozystack</a></li>
<li>Telegram: <a href="https://t.me/cozystack">@cozystack</a></li>
<li>Slack: <a href="https://kubernetes.slack.com/archives/C06L3CPRVN1">#cozystack</a> on the Kubernetes workspace (<a href="https://slack.kubernetes.io/">invite</a>)</li>
<li><a href="https://zoom-lfx.platform.linuxfoundation.org/meetings/cozystack">Subscribe to our community meetings calendar</a></li>
<li><a href="https://webcal.prod.itx.linuxfoundation.org/lfx/lfsixxnFWxbvsyEuC2">Add meetings to your calendar</a></li>
</ul>
<hr>
<p><a href="https://blog.aenix.io/cozystack-1-4-0c5e399a7308">Cozystack 1.4: New Dashboard UI, Persistent Tenant Workers, Backup Strategies, and Fractional GPU Sharing</a> was originally published in <a href="https://blog.aenix.io">Ænix</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>
]]></content:encoded></item><item><title>Seven decisions when designing sovereign AI architecture</title><link>https://aenix.io/blog/2026/05/sovereign-ai-architecture-decisions/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/sovereign-ai-architecture-decisions/</guid><pubDate>Wed, 27 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>DORA</category><category>NIS2</category><category>Sovereignty</category><category>AI and ML</category><category>Multi-tenancy</category><category>Financial Services</category><description>Seven architecture decisions behind a sovereign AI stack, how they interlock, and the combinations that recur in real deployments.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/sovereign-ai-architecture-decisions.jpg" alt=""></p><h2 id="the-seven-decisions">The seven decisions</h2>
<h3 id="1-trigger-profile">1. Trigger profile</h3>
<p>What&rsquo;s pushing toward sovereign? (Regulated data class / inference economics / auditability / air-gap requirement)</p>
<h3 id="2-regulatory-scope">2. Regulatory scope</h3>
<p>Which regulators bind you? (DORA, NIS2, sectoral, sovereign-cloud mandate, GDPR cross-border)</p>
<h3 id="3-model-selection">3. Model selection</h3>
<p>Open-weight vs proprietary. Common 2026 open-weight: Llama, Mistral, Qwen, DeepSeek, Phi, Gemma. Choice depends on:</p>
<ul>
<li>Language requirement (multilingual vs English)</li>
<li>Workload type (chat / RAG / code / vision / embedding)</li>
<li>Licence (commercial use, attribution, redistribution)</li>
<li>Capability target</li>
</ul>
<h3 id="4-hardware-sizing">4. Hardware sizing</h3>
<ul>
<li>NVIDIA H100/H200 — workhorse for fine-tuning and high-throughput inference</li>
<li>NVIDIA Blackwell — newer, highest memory bandwidth</li>
<li>NVIDIA L40S — 48GB, multi-tenant inference fleet</li>
<li>NVIDIA A100 — cost-effective, second-hand market</li>
<li>AMD MI300/MI325 — alternative when ecosystem fits</li>
</ul>
<h3 id="5-multi-tenancy-model">5. Multi-tenancy model</h3>
<ul>
<li>Single-tenant: lab / single-team / PoC</li>
<li>Multi-tenant via Tenant CRD: enterprise platform / customer-facing</li>
<li>Cluster-per-tenant: highest isolation, operationally expensive</li>
</ul>
<h3 id="6-sovereignty-controls">6. Sovereignty controls</h3>
<ul>
<li>Encryption keys customer-controlled (HSM)</li>
<li>Supplier transparency to second hop</li>
<li>Audit-trail completeness in regulator-consumable formats</li>
<li>Air-gap option for most sensitive workloads</li>
</ul>
<h3 id="7-operational-model">7. Operational model</h3>
<ul>
<li>Customer-operated (you run it)</li>
<li>Vendor-operated (Ænix or similar runs it)</li>
<li>Hybrid (you operate; vendor 2nd-line)</li>
</ul>
<h2 id="how-decisions-interlock">How decisions interlock</h2>
<p>The seven aren&rsquo;t independent. Trigger profile shapes regulatory scope; regulatory scope shapes sovereignty controls; sovereignty controls shape operational model; operational model affects model selection feasibility.</p>
<h2 id="common-combinations">Common combinations</h2>
<p><strong>Pattern 1: Regulated finance + sustained inference + multi-tenant</strong>
DORA + Article 28 controls + multi-tenant Tenant CRD + opt-in volume encryption with a customer-held passphrase + Ænix-managed operations + open-weight (Llama 70B class) on H100/L40S fleet.</p>
<p><strong>Pattern 2: Public sector + air-gapped + classified data</strong>
Sovereign-cloud mandate + air-gap + customer-operated + open-weight (Llama / Phi) on customer hardware.</p>
<p><strong>Pattern 3: AI startup + sustained 24/7 inference + customer-facing</strong>
No specific regulator + cost economics trigger + multi-tenant + customer-operated + open-weight + GPU mix matched to workload.</p>
<h2 id="how-to-use-the-decision-guide">How to use the decision guide</h2>
<p>Answer the questions above in order and note your answers; the architecture options narrow naturally. The <a href="https://aenix.io/resources/sovereign-ai-decision-guide/">sovereign AI decision guide</a> walks through the same decisions in more depth.</p>
<p>For specific engagement see <strong><a href="https://aenix.io/solutions/sovereign-ai/">Sovereign AI services</a></strong>.</p>
]]></content:encoded></item><item><title>Smart grid platform architecture — IT/OT convergence, edge, and AI on customer-controlled infrastructure</title><link>https://aenix.io/blog/2026/05/smart-grid-platform-architecture-it-ot/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/smart-grid-platform-architecture-it-ot/</guid><pubDate>Tue, 26 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>NIS2</category><category>AI and ML</category><category>GPU</category><category>Compliance</category><description>A smart-grid architectural reference for energy operators: IT/OT boundaries that work, NIS2 controls, AI on grid-operational data, and legacy migration.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/smart-grid-platform-architecture-it-ot.jpg" alt=""></p><p>The energy sector&rsquo;s infrastructure-modernization conversation in 2026 sits at an unusual intersection: NIS2 compliance pressure, AI-driven grid-optimization demand, IT/OT convergence reality, edge compute requirements at substation density, and the irreducible operational fact that grid hardware refresh cycles are measured in decades. Few other sectors face this combination simultaneously.</p>
<h2 id="three-pressures-converging-on-energy-infrastructure">Three pressures converging on energy infrastructure</h2>
<h3 id="pressure-1-nis2-compliance-with-operational-reality">Pressure 1: NIS2 compliance with operational reality</h3>
<p>NIS2 Article 21 risk-management measures apply to energy operators (Annex I sector; essential or important entity depending on size). Article 23 incident reporting (24-hour / 72-hour / 1-month timelines) requires telemetry tuned for security, not just performance. Article 21(2)(d) supply-chain security requires mapping ICT third parties to the second hop at minimum.</p>
<p>For energy operators with legacy SCADA + DCS + GIS + energy-management systems integrated through years of one-off engineering, Article 21&rsquo;s &ldquo;documented risk register per critical workload&rdquo; is non-trivial.</p>
<h3 id="pressure-2-ai-for-grid-operations">Pressure 2: AI for grid operations</h3>
<p>Grid forecasting (load, generation, weather impact), demand response automation, predictive maintenance (transformer / line / substation health), market-price optimization — all increasingly use ML.</p>
<p>The economics drive sustained workloads (24/7 model serving) where dedicated GPU infrastructure usually wins over hyperscaler GPU. Combined with grid-operational data sensitivity (often legally protected), sovereign AI is the natural answer.</p>
<h3 id="pressure-3-edge-compute-at-substation-density">Pressure 3: Edge compute at substation density</h3>
<p>Modern grid operations distribute compute toward generation, distribution, and storage edges. Substations (10s-100s per operator), distributed generation sites, microgrids, EV charging clusters — each may host workload that needs local compute with intermittent central connectivity.</p>
<p>Traditional centralized SCADA architecture doesn&rsquo;t scale to this density. Modern smart-grid architecture distributes the platform.</p>
<h2 id="smart-grid-architectural-reference">Smart-grid architectural reference</h2>
<p>Bring the three pressures together and the architecture looks roughly like this:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-gdscript3" data-lang="gdscript3"><span class="line"><span class="cl"><span class="n">HQ</span> <span class="n">Cloud</span> <span class="p">(</span><span class="n">Cozystack</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">   <span class="err">├──</span> <span class="n">Energy</span> <span class="n">management</span> <span class="n">system</span> <span class="p">(</span><span class="n">EMS</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">   <span class="err">├──</span> <span class="n">Grid</span> <span class="n">analytics</span> <span class="o">+</span> <span class="n">AI</span> <span class="n">training</span>
</span></span><span class="line"><span class="cl">   <span class="err">├──</span> <span class="n">Forecasting</span> <span class="n">models</span> <span class="p">(</span><span class="nb">load</span><span class="p">,</span> <span class="n">generation</span><span class="p">,</span> <span class="n">weather</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">   <span class="err">├──</span> <span class="n">Market</span> <span class="n">integration</span> <span class="p">(</span><span class="n">ENTSO</span><span class="o">-</span><span class="n">E</span><span class="p">,</span> <span class="n">market</span> <span class="n">operators</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">   <span class="err">├──</span> <span class="n">Customer</span> <span class="n">billing</span> <span class="o">+</span> <span class="n">customer</span><span class="o">-</span><span class="n">facing</span> <span class="n">portals</span>
</span></span><span class="line"><span class="cl">   <span class="err">└──</span> <span class="n">Compliance</span> <span class="o">+</span> <span class="n">audit</span> <span class="nb">log</span> <span class="n">aggregation</span>
</span></span><span class="line"><span class="cl">        <span class="err">↓</span>
</span></span><span class="line"><span class="cl"><span class="n">Regional</span> <span class="n">control</span> <span class="n">sites</span> <span class="p">(</span><span class="n">Cozystack</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">   <span class="err">├──</span> <span class="n">SCADA</span> <span class="n">aggregation</span>
</span></span><span class="line"><span class="cl">   <span class="err">├──</span> <span class="n">Regional</span> <span class="n">grid</span> <span class="n">management</span>
</span></span><span class="line"><span class="cl">   <span class="err">├──</span> <span class="n">AI</span> <span class="n">inference</span> <span class="p">(</span><span class="n">forecasting</span> <span class="n">models</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">   <span class="err">└──</span> <span class="n">Disaster</span> <span class="n">recovery</span>
</span></span><span class="line"><span class="cl">        <span class="err">↓</span>
</span></span><span class="line"><span class="cl"><span class="n">Substation</span> <span class="n">edge</span> <span class="p">(</span><span class="n">Cozystack</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">   <span class="err">├──</span> <span class="n">Local</span> <span class="n">SCADA</span> <span class="o">/</span> <span class="n">RTU</span> <span class="n">integration</span>
</span></span><span class="line"><span class="cl">   <span class="err">├──</span> <span class="n">Real</span><span class="o">-</span><span class="n">time</span> <span class="n">grid</span><span class="o">-</span><span class="n">edge</span> <span class="n">compute</span>
</span></span><span class="line"><span class="cl">   <span class="err">├──</span> <span class="n">IoT</span> <span class="n">data</span> <span class="n">ingestion</span> <span class="p">(</span><span class="n">sensors</span><span class="p">,</span> <span class="n">smart</span> <span class="n">meters</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">   <span class="err">└──</span> <span class="n">Air</span><span class="o">-</span><span class="n">gapped</span> <span class="n">OT</span> <span class="n">boundary</span>
</span></span></code></pre></div><p>Cozystack runs at all three layers under one operational model, with policy and identity federated, but the OT/IT boundary at substation tier is intentionally restricted.</p>
<h2 id="itot-convergence--boundaries-that-work">IT/OT convergence — boundaries that work</h2>
<p>The naive &ldquo;merge IT and OT&rdquo; pattern is dangerous. The working pattern in 2026 is <strong>convergence with boundaries</strong>:</p>
<h3 id="strict-boundary-ot-zone">Strict boundary: OT zone</h3>
<ul>
<li>Substation OT systems (SCADA, RTU, IEDs) live in OT zone</li>
<li>Air-gapped or strictly restricted-egress to IT zone</li>
<li>Updates via controlled channels (Harbor mirror, manual approval)</li>
<li>Cybersecurity tooling is OT-aware (not generic IT EDR)</li>
<li>Identity model separate from workforce IT</li>
</ul>
<h3 id="permeable-boundary-it-zone">Permeable boundary: IT zone</h3>
<ul>
<li>Grid analytics, AI inference, forecasting, customer-facing apps live in IT zone</li>
<li>Standard cloud-native operational model</li>
<li>IT cybersecurity tools and identity federation</li>
<li>Cloud-native observability + compliance tooling</li>
</ul>
<h3 id="bridge-data-fabric">Bridge: data fabric</h3>
<ul>
<li>Carefully-controlled data flows from OT to IT (typically read-only or one-way)</li>
<li>Data validation + sanitization at boundary</li>
<li>Audit-trail completeness at boundary</li>
</ul>
<p>The architectural pattern: Cozystack runs the IT zone and the bridge layer. OT zone is its own dedicated infrastructure with controlled egress. Both can be deployed on Cozystack but with strict policy boundaries.</p>
<h2 id="cozystack-architectural-advantages-for-energy">Cozystack architectural advantages for energy</h2>
<h3 id="1-multi-site-under-one-operational-model">1. Multi-site under one operational model</h3>
<p>Cozystack platforms federate across central + regional + substation tiers. Single operational model, single platform team, consistent observability, GitOps-driven changes.</p>
<h3 id="2-air-gap-support-for-ot">2. Air-gap support for OT</h3>
<p>Documented air-gap install workflow. Suitable for OT zones that cannot have internet egress. Updates via Harbor mirror or controlled channels.</p>
<h3 id="3-ai-infrastructure-native">3. AI infrastructure native</h3>
<p>KubeVirt for legacy AI workloads, native Kubernetes for modern ML pipelines. Passthrough of whole GPUs or NVIDIA vGPU (requires your NVIDIA vGPU licence) for VM-bound workloads, and HAMi fractional sharing (GPU memory and compute cores) for containers sharing GPUs across forecasting models. NVIDIA data-centre GPUs are supported through the NVIDIA GPU Operator; MIG and time-slicing are on the roadmap.</p>
<h3 id="4-multi-tenant-for-cross-bu">4. Multi-tenant for cross-BU</h3>
<p>Tenant CRD model accommodates generation / transmission / distribution / retail BUs with separate isolation. For unbundled markets, this is non-optional.</p>
<h3 id="5-sovereign-by-architecture">5. Sovereign by architecture</h3>
<p>Open-source platform on customer-controlled hardware. Opt-in volume encryption at rest (LINSTOR and LUKS) with a passphrase you hold. Audit-trail completeness in regulator-consumable formats. NIS2-aligned without bolt-on workarounds.</p>
<h3 id="6-long-operational-horizon">6. Long operational horizon</h3>
<p>Apache 2.0 licence + CNCF Project community governance fits decade-plus grid operational planning. Vendor roadmap risk is minimized.</p>
<h2 id="nis2-specific-architecture-controls">NIS2-specific architecture controls</h2>
<p>For each Article 21 sub-requirement that touches infrastructure:</p>
<ul>
<li><strong>Risk register per critical function</strong> — workload-level mapping, including OT-zone boundaries</li>
<li><strong>Incident detection at 24-hour timeline</strong> — telemetry tuned for security, not just performance</li>
<li><strong>Business continuity tested</strong> — annual exercises with documented results, including substation-tier failover</li>
<li><strong>Supply-chain transparency to second hop</strong> — every ICT vendor + sub-contractor mapped</li>
<li><strong>MFA for privileged accounts</strong> — universal in IT zone; appropriate substitutes in OT zone (smart-card, dedicated KVM)</li>
<li><strong>Cryptography with customer-controlled keys</strong> — HSM-based for sensitive grid-operational data</li>
<li><strong>Vulnerability management with CVE response SLA</strong> — challenging in OT zone (long maintenance windows); architecture must support staged remediation</li>
</ul>
<h2 id="ai-workloads-on-grid-operational-data">AI workloads on grid-operational data</h2>
<p>Common AI use cases at energy operators:</p>
<ul>
<li><strong>Load forecasting</strong> — short-term (1-24 hours), medium-term (1-30 days), long-term (1-5 years)</li>
<li><strong>Generation forecasting</strong> — especially for renewables (wind, solar) where weather is decisive</li>
<li><strong>Predictive maintenance</strong> — transformer health, line condition, substation equipment</li>
<li><strong>Demand response automation</strong> — for flexible loads (industrial customers, EV charging, storage)</li>
<li><strong>Market price optimization</strong> — for participating in wholesale markets</li>
<li><strong>Grid topology optimization</strong> — power flow analysis, congestion management</li>
<li><strong>Customer-facing AI</strong> — chatbot, billing analysis, energy-advice AI</li>
</ul>
<p>The data feeding these models is often legally protected (customer data, grid-operational data, market-sensitive data). Sovereign AI infrastructure on customer hardware is the natural deployment pattern.</p>
<p>Typical hardware sizing for mid-size energy operator (5-10 GW generation portfolio): 16-64 GPUs across H100/L40S, with elastic burst capacity for re-training cycles.</p>
<h2 id="migration-patterns-from-legacy">Migration patterns from legacy</h2>
<p>Most energy-sector smart-grid platforms in 2026 are migrations from legacy mixes of:</p>
<ul>
<li>Vendor-specific SCADA platforms (often heavily customized)</li>
<li>Legacy virtualization (VMware-heavy)</li>
<li>Standalone forecasting systems (commercial software)</li>
<li>Spreadsheet-based grid analytics (still surprisingly common)</li>
</ul>
<p>Migration sequencing:</p>
<ol>
<li><strong>Discovery</strong> — workload inventory, OT/IT zone mapping, regulatory scope</li>
<li><strong>Cozystack foundation</strong> — central tier first</li>
<li><strong>Regional tier rollout</strong> — cohort by cohort</li>
<li><strong>AI workload migration</strong> — from legacy forecasting to Kubernetes-native</li>
<li><strong>Edge tier rollout</strong> — substation-by-substation, slowest tier due to operational risk</li>
<li><strong>Legacy decommission</strong> — staged as cohorts complete</li>
</ol>
<p>For a mid-size operator, the multi-site platform follows the operator-scale pattern: a 3-6 month pilot on the central tier, then 9-18 months to full multi-site operation. The substation roll-out can run longer, paced by maintenance windows.</p>
<h2 id="common-pitfalls">Common pitfalls</h2>
<h3 id="pitfall-1-weak-otit-boundary">Pitfall 1: weak OT/IT boundary</h3>
<p>&ldquo;We&rsquo;ll converge OT and IT&rdquo; without strict boundary controls produces operational risk. The cybersecurity model in OT is different; the boundary must be deliberate.</p>
<h3 id="pitfall-2-hyperscaler-only-ai">Pitfall 2: hyperscaler-only AI</h3>
<p>Putting grid-operational AI in hyperscaler region. Often incompatible with NIS2 supplier-concentration requirements; sometimes legally problematic for customer-data-driven workloads.</p>
<h3 id="pitfall-3-substation-tier-underspecified">Pitfall 3: substation-tier underspecified</h3>
<p>Treating substation tier as &ldquo;just SCADA&rdquo; misses that modern smart-grid architecture distributes compute to substations. Needs platform-engineering treatment.</p>
<h3 id="pitfall-4-skipping-nis2-architecture-during-platform-build">Pitfall 4: skipping NIS2 architecture during platform build</h3>
<p>&ldquo;We&rsquo;ll add NIS2 controls in v2&rdquo; produces retrofit cost that exceeds doing it from start.</p>
<h3 id="pitfall-5-under-engineered-dr-for-grid-criticality">Pitfall 5: under-engineered DR for grid criticality</h3>
<p>Grid operations have public-safety implications. DR posture must be tested and proven, not documented only.</p>
<h2 id="where-ænix-engages">Where Ænix engages</h2>
<p>The standard <strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong> with energy-specific workstream emphasis covers the full picture. For details and engagement structure see <strong><a href="https://aenix.io/industries/energy/">energy industry page</a></strong>.</p>
<p>For specific triggers see <strong><a href="https://aenix.io/solutions/nis2-compliance/">NIS2 compliance</a></strong>, <strong><a href="https://aenix.io/solutions/sovereign-ai/">Sovereign AI</a></strong>, <strong><a href="https://aenix.io/solutions/data-sovereignty/">Data sovereignty</a></strong>.</p>
<hr>
<h2 id="want-to-dig-deeper">Want to dig deeper?</h2>
<ul>
<li><strong><a href="https://aenix.io/industries/energy/">Energy industry page</a></strong> — engagement details</li>
<li><strong><a href="https://aenix.io/solutions/nis2-compliance/">NIS2 compliance</a></strong> — essential-entity regulatory</li>
<li><strong><a href="https://aenix.io/solutions/sovereign-ai/">Sovereign AI</a></strong> — AI on grid data</li>
<li><strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong> — methodology</li>
<li><strong><a href="https://aenix.io/products/cozystack/">Cozystack</a></strong> — open-source platform foundation</li>
</ul>
]]></content:encoded></item><item><title>Reverse cloud migration — a practical playbook for leaving public cloud in 2026</title><link>https://aenix.io/blog/2026/05/reverse-cloud-migration-playbook/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/reverse-cloud-migration-playbook/</guid><pubDate>Tue, 26 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>DORA</category><category>NIS2</category><category>Sovereignty</category><category>Cloud Repatriation</category><category>AI and ML</category><category>GPU</category><description>A five-step cloud repatriation playbook, the pitfalls that recur, when not to repatriate, and how long a realistic move actually takes.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/reverse-cloud-migration-playbook.jpg" alt=""></p><p>Most coverage of cloud repatriation is either ideological (&ldquo;public cloud was always too expensive&rdquo;) or vendor-led (&ldquo;buy our private-cloud appliance&rdquo;). Neither helps the platform engineer or infrastructure lead who has to translate a board-level decision into running systems. The work below is what we actually do during an Ænix repatriation engagement.</p>
<h2 id="why-repatriation-is-happening-now">Why repatriation is happening now</h2>
<p>Three independent pressures hit the same architectures at the same time:</p>
<p><strong>Cost cliff after the renewal cycle.</strong> Public-cloud spend that looked acceptable in 2020-2022 has compounded for several years. Renewal cycles are now meeting boards that did not approve the trajectory. Reservation under-utilization, egress costs, and &ldquo;we&rsquo;ll right-size later&rdquo; technical debt are all visible.</p>
<p><strong>Regulatory pressure.</strong> DORA in force since January 2025; NIS2 transposition across EU member states; sectoral rules sharpening in financial services, public sector, healthcare; and explicit sovereign-cloud mandates in Kazakhstan, France, Germany, and other markets. Many critical-function workloads can no longer comfortably run in a single hyperscaler region under a US-headquartered provider.</p>
<p><strong>AI and inference economics.</strong> GenAI training and inference at scale have egress, GPU-pricing, and data-residency profiles that hyperscalers were not designed to optimize. Some workloads make sense in hyperscaler GPU; many do not.</p>
<p>The combination has shifted repatriation from &ldquo;what some hipsters do&rdquo; to a normal part of cloud strategy. Done well, it produces 30-60% cost reduction on the workloads that move, plus regulator-aligned architecture. Done badly, it produces an under-engineered on-prem environment that combines the worst of both worlds.</p>
<h2 id="repatriation-isnt-all-or-nothing">Repatriation isn&rsquo;t all-or-nothing</h2>
<p>The single most common mistake at the strategy level is treating repatriation as a binary decision. It almost never is.</p>
<p>A typical repatriated estate, after a year of work, looks like:</p>
<ul>
<li><strong>30-50% on-prem or private cloud</strong> — steady-state workloads, regulated workloads, expensive workloads, latency-critical workloads</li>
<li><strong>30-50% remaining in public cloud</strong> — elastic / spike workloads, hyperscaler-proprietary services with no realistic alternative, latency-sensitive customer-facing workloads where hyperscaler edge is decisive, very small workloads where the operational cost of repatriation exceeds the savings</li>
<li><strong>10-20% in transition</strong> — workloads being moved, in PoC, or under reassessment</li>
</ul>
<p>The workstream that classifies workloads — repatriate now / repatriate later / keep in cloud — is the most consequential single deliverable of a repatriation engagement.</p>
<h2 id="five-step-playbook">Five-step playbook</h2>
<p>Here is the sequence of work in a structured reverse cloud migration. Each step has a defined deliverable and a defined precondition for moving to the next step.</p>
<h3 id="step-1--honest-tco-modelling">Step 1 — honest TCO modelling</h3>
<p>Most organizations do not actually know their real public-cloud TCO. The bill is one number; the real cost includes:</p>
<ul>
<li>Egress charges, especially for backup, observability, and cross-region traffic</li>
<li>Reservation / commitment under-utilization</li>
<li>Idle and over-sized resources never reclaimed</li>
<li>Hyperscaler-managed services priced at a premium over self-managed equivalents</li>
<li>Hidden cost of vendor-lock-in: switching cost when something fails</li>
<li>Cost of platform-engineering capacity dedicated to managing public-cloud-specific complexity</li>
</ul>
<p>A honest TCO model captures all of these and compares them to a realistic destination-cost model that includes:</p>
<ul>
<li>Hardware acquisition and refresh cost over 5 years</li>
<li>Datacenter or colocation cost</li>
<li>Network bandwidth, including egress between sites</li>
<li>Storage tiering and growth</li>
<li>Backup and DR infrastructure</li>
<li>Identity, observability, and platform tooling</li>
<li>Platform-engineering capacity needed to operate the destination</li>
<li>Software licences where applicable</li>
</ul>
<p>The honest model usually shows on-prem economics 30-60% better for steady-state workloads, and 0-20% worse for highly elastic workloads. The interesting question is which workloads are which.</p>
<p><strong>Deliverable:</strong> TCO model in a spreadsheet your CFO can audit. Public-cloud-current-state vs. destination-target-state, sensitive to occupancy assumptions.</p>
<h3 id="step-2--workload-classification">Step 2 — workload classification</h3>
<p>Every workload gets one of four labels:</p>
<ul>
<li><strong>Repatriate now</strong> — clear cost or regulatory case, low migration friction, no hyperscaler-only dependencies.</li>
<li><strong>Repatriate later</strong> — case is clear but commitments not yet expired, or destination architecture not ready.</li>
<li><strong>Stay in cloud</strong> — workload is genuinely better-suited to hyperscaler economics or hyperscaler-only services.</li>
<li><strong>Reassess</strong> — not enough data to decide; needs PoC or instrumentation.</li>
</ul>
<p>The classification considers:</p>
<ul>
<li>Compute pattern (steady vs. spiky)</li>
<li>Data gravity (size, growth rate, regulatory class)</li>
<li>Hyperscaler-proprietary dependencies</li>
<li>Network adjacency (does it need to be next to other workloads?)</li>
<li>Latency sensitivity</li>
<li>Available commitment / reservation expiration</li>
<li>Migration effort estimate</li>
</ul>
<p><strong>Deliverable:</strong> workload table with classification, ranked by net repatriation ROI. The top-10 list usually accounts for 60-80% of the cost case.</p>
<h3 id="step-3--destination-architecture">Step 3 — destination architecture</h3>
<p>This is where most repatriations go wrong: treating &ldquo;the destination&rdquo; as &ldquo;an on-prem cluster&rdquo; without engineering the platform underneath.</p>
<p>A real destination architecture has answers for all of:</p>
<ul>
<li><strong>Compute:</strong> virtualization platform (KubeVirt, KVM/libvirt, Hyper-V, VMware), hardware specs, lifecycle management, live migration, snapshots</li>
<li><strong>Storage:</strong> primary storage (LINSTOR / Ceph / SAN / vendor HCI), backup, DR, snapshots, replication topology</li>
<li><strong>Network:</strong> datacenter fabric, software-defined networking (Cilium, NSX equivalent), load balancing, BGP / routing, edge connectivity to remaining cloud workloads</li>
<li><strong>Identity:</strong> how IAM works, including federation with workforce identity</li>
<li><strong>Observability:</strong> metrics (VictoriaMetrics, Prometheus), logs (VictoriaLogs, Loki), tracing</li>
<li><strong>Backup and DR:</strong> Velero, restic, cross-site replication, RPO/RTO targets</li>
<li><strong>Platform-engineering function:</strong> the team that operates this; size, structure, on-call model</li>
<li><strong>Self-service surface:</strong> how product teams provision environments without filing tickets</li>
</ul>
<p>A Kubernetes-native architecture (KubeVirt + Cilium + LINSTOR + Velero, or equivalent) is increasingly the default for repatriation projects in 2026 because it gives a coherent answer to all of the above without integrating ten unrelated vendors. <strong><a href="https://aenix.io/products/cozystack/">Cozystack</a></strong> is one such platform; OpenStack, OpenShift Virtualization, and several vendor-led options are alternatives.</p>
<p><strong>Deliverable:</strong> target architecture document. Diagram + named components + sizing + cost.</p>
<h3 id="step-4--cutover-sequencing">Step 4 — cutover sequencing</h3>
<p>The cutover plan respects three constraints:</p>
<ol>
<li><strong>Commitment expirations.</strong> Workloads move when AWS RI / Azure RI / Savings Plan expirations open the economic case. Moving early into the lockup costs money.</li>
<li><strong>Data gravity.</strong> Workloads that share data move together, or with explicit cross-cloud data flow during transition.</li>
<li><strong>Risk concentration.</strong> No single cohort moves more than your operational capacity to manage rollback.</li>
</ol>
<p>A typical 100-VM repatriation runs as 4-8 cohorts within a programme of about 8-12 months. Each cohort is parallel-run in cloud and on-prem until validated by the application owner.</p>
<p><strong>Deliverable:</strong> cutover plan with cohort definitions, dates, dependencies, rollback paths, and signoff criteria.</p>
<h3 id="step-5--operate-the-platform">Step 5 — operate the platform</h3>
<p>The post-repatriation steady state is where the strategic gain materializes — <em>if</em> the platform was actually engineered. The post-repatriation team needs:</p>
<ul>
<li>Runbooks for the platform stack</li>
<li>Monitoring and alerting that catches issues before users do</li>
<li>Self-service paths for product teams (golden paths, IaC-based environment provisioning, GitOps-based deployment)</li>
<li>A platform-engineering function with the right headcount to maintain pace</li>
</ul>
<p>The most common failure mode at this stage is a successful migration followed by under-staffing of the platform team. Repatriation with a 5-person platform team responsible for everything that AWS used to handle is not sustainable.</p>
<p><strong>Deliverable:</strong> operating model, headcount plan, runbook library, and a 12-month platform-engineering roadmap.</p>
<h2 id="common-pitfalls">Common pitfalls</h2>
<p>Beyond &ldquo;TCO is wishful&rdquo; and &ldquo;destination is left for later,&rdquo; four more pitfalls recur:</p>
<h3 id="pitfall-1--treating-repatriation-as-a-cost-project-rather-than-a-platform-project">Pitfall 1 — treating repatriation as a cost project rather than a platform project</h3>
<p>Repatriation that&rsquo;s measured purely on cost reduction tends to deprioritize the platform work that makes the cost reduction sustainable. Two years in, the team has saved money but lost velocity, and the result is a partial reverse-repatriation back into hyperscalers.</p>
<p>The fix: repatriation goals include <em>platform engineering maturity</em> alongside cost. A successful repatriation produces a self-service platform whose time-to-environment is at least as good as what you had on hyperscalers — usually better.</p>
<h3 id="pitfall-2--underestimating-data-gravity">Pitfall 2 — underestimating data gravity</h3>
<p>Moving 50 TB of production data is not a weekend job. Cross-network movement, cutover windows, dual-write periods, rollback paths, and backup-during-migration all need explicit engineering. Teams that don&rsquo;t plan for this end up in a multi-week emergency in the middle of the migration.</p>
<h3 id="pitfall-3--buying-a-vendor-led-repatriation-appliance">Pitfall 3 — buying a vendor-led &ldquo;repatriation appliance&rdquo;</h3>
<p>Several vendors sell &ldquo;private cloud in a box&rdquo; as a repatriation answer. These work for a narrow class of customers but rebuild the lock-in problem with a different vendor. The vendor&rsquo;s roadmap becomes your roadmap; the vendor&rsquo;s support availability becomes your operational ceiling. Open-source platforms (Cozystack, OpenStack, OpenShift) avoid this trap.</p>
<h3 id="pitfall-4--ignoring-the-human-side">Pitfall 4 — ignoring the human side</h3>
<p>Platform-team morale, product-team relationships, and internal communications all shift during a repatriation. Engineers who built deep AWS skills may not see private cloud as a career step. Product teams accustomed to AWS Console may push back against new self-service paths. A repatriation program that doesn&rsquo;t account for the human transition stalls in month 6.</p>
<h2 id="when-not-to-repatriate">When NOT to repatriate</h2>
<p>Repatriation is the wrong answer when:</p>
<ul>
<li>You have a small IT team running a handful of services. Public cloud&rsquo;s operational simplicity is genuinely better than what you&rsquo;d build.</li>
<li>Your workload portfolio is dominated by hyperscaler-proprietary services (e.g., heavy use of AWS Lambda + DynamoDB + Kinesis with no realistic alternatives).</li>
<li>Your business is fundamentally elastic — traffic patterns where 10× spike capacity is needed for hours per day. Hyperscaler economics are good for this.</li>
<li>Your platform-engineering capacity is already stretched beyond capacity. Adding repatriation to an under-resourced team makes both worse.</li>
<li>Your renewal cycle is at year 1 of a 5-year commitment with steep penalties. Wait for commitment expirations.</li>
</ul>
<p>A good repatriation engagement is honest about these cases. The Ænix engagement specifically does not push repatriation when staying in cloud is the right answer.</p>
<h2 id="what-about-hybrid">What about hybrid?</h2>
<p>Most repatriated estates end up hybrid — a useful term, often vague in practice. There are three coherent patterns:</p>
<ul>
<li><strong>Steady-state on-prem, elastic in cloud.</strong> Predictable workloads run on private platform; spike capacity overflows to cloud. Operationally complex; usually worth it for workloads with strong elastic patterns.</li>
<li><strong>Critical on-prem, non-critical in cloud.</strong> Regulated and core workloads on private cloud; auxiliary workloads (analytics, internal tooling, dev/test) in public cloud. Operationally simpler; common for financial-services repatriations.</li>
<li><strong>Geographic split.</strong> EU workloads on-prem in EU; non-EU workloads in regional public cloud. Driven by sovereignty rather than cost.</li>
</ul>
<p>The hybrid pattern that fits depends on the trigger that drove repatriation — cost, regulator, or sovereignty.</p>
<h2 id="how-long-does-this-take">How long does this take?</h2>
<p>A typical repatriation, end-to-end:</p>
<ul>
<li><strong>14 or 28 days:</strong> assessment phase (<a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a>, fixed price, with cost-and-cloud-spend workstream emphasis).</li>
<li><strong>3-12 months, overlapping with the first cohorts:</strong> destination platform build — greenfield infrastructure, base services, observability, identity, IaC and GitOps tooling, runbooks. The platform is usable for the first cohort well before the build is complete.</li>
<li><strong>The bulk of the programme:</strong> workload migration in cohorts. Earliest cohorts move quickly; later cohorts respect commitment ladders. For ~1,000 VMs, expect 18-24 months in total.</li>
<li><strong>Ongoing:</strong> platform operation and continuous optimization. The post-repatriation platform is a long-term asset that compounds value.</li>
</ul>
<p>For an organization with 100 VMs and a moderate cloud bill, total elapsed time from &ldquo;we should look at this&rdquo; to &ldquo;we are running on the destination platform&rdquo; is about 8-12 months.</p>
<h2 id="where-to-start">Where to start</h2>
<p>If repatriation is on the table for your organization, the structured next step is a focused assessment. The output is honest enough to support a board-level decision either way: repatriate (with a plan), don&rsquo;t repatriate (with the reasons), or selective repatriation (with the workload list).</p>
<p>Ænix runs this as a 14- or 28-day <strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong> with the cost-and-cloud-spend workstream emphasized. See the <strong><a href="https://aenix.io/solutions/cloud-repatriation/">cloud repatriation services page</a></strong> for the engagement details.</p>
<hr>
<h2 id="want-to-dig-deeper">Want to dig deeper?</h2>
<ul>
<li><strong><a href="https://aenix.io/solutions/cloud-repatriation/">Cloud repatriation services page</a></strong> — the engagement details and pricing</li>
<li><strong><a href="https://aenix.io/solutions/cloud-cost-optimization/">Cloud cost optimization</a></strong> — adjacent FinOps trigger</li>
<li><strong><a href="https://aenix.io/solutions/data-sovereignty/">Data sovereignty</a></strong> — regulatory side of the same shift</li>
<li><strong><a href="https://aenix.io/products/cozystack/">Cozystack</a></strong> — the platform we typically recommend as repatriation destination</li>
</ul>
]]></content:encoded></item><item><title>Public-sector sovereign cloud — from procurement framework to running platform</title><link>https://aenix.io/blog/2026/05/public-sector-sovereign-cloud-procurement/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/public-sector-sovereign-cloud-procurement/</guid><pubDate>Mon, 25 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Public Sector</category><category>Sovereignty</category><category>Compliance</category><category>NIS2</category><category>Cozystack</category><description>How public-sector procurement leads and IT directors turn sovereignty mandates (EUCS, SecNumCloud, BSI C5, NIS2) into a running, auditable cloud platform.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/public-sector-sovereign-cloud-procurement.jpg" alt=""></p><p>The public-sector sovereign-cloud conversation in 2026 is more
fragmented than in financial services. There&rsquo;s no single regulation
like DORA driving alignment. Instead, each jurisdiction has its own
framework, often layered on top of GDPR, NIS2, and sectoral overlays.
A multinational public-sector engagement — or even a national one
crossing regions — usually maps against three or more frameworks
simultaneously.</p>
<h2 id="the-framework-landscape">The framework landscape</h2>
<h3 id="eu-level">EU level</h3>
<ul>
<li><strong>EUCS (EU Cybersecurity Certification Scheme for Cloud Services)</strong> —
proposed EU-wide scheme (adoption still pending). Three assurance levels
(Basic, Substantial, High). High level requires substantive
sovereignty controls.</li>
<li><strong>NIS2</strong> — public administration is an Annex I sector; central-government
entities are essential entities (Article 3). Article 21 + Article 23 obligations.</li>
<li><strong>GDPR</strong> — personal-data baseline, cross-border-transfer rules under
Articles 44-50.</li>
</ul>
<h3 id="member-state-level">Member-state level</h3>
<ul>
<li><strong>France: SecNumCloud</strong> — strict French national framework. The most
demanding EU-member-state sovereign-cloud scheme. Reference standard
for several other national initiatives.</li>
<li><strong>Germany: BSI C5</strong> — German cloud security catalogue. Now widely
referenced beyond Germany for DACH operations.</li>
<li><strong>Italy: ACN</strong> — National Cybersecurity Agency frameworks; Polo
Strategico Nazionale infrastructure rules.</li>
<li><strong>Spain: ENS High</strong> — Esquema Nacional de Seguridad highest level.</li>
</ul>
<p>Other member states have variations.</p>
<h3 id="central-asia-and-apac">Central Asia and APAC</h3>
<ul>
<li><strong>Kazakhstan</strong> — procurement-mandated sovereignty for public-sector
workloads via goszakup.gov.kz / mitwork.kz / zakup.sk.kz.</li>
<li><strong>Singapore: IM8</strong> — Government IT security standards.</li>
<li><strong>India: MeitY</strong> — Ministry of Electronics IT, including STQC
Empanelled CSP framework.</li>
<li><strong>Australia: IRAP</strong> — Information Security Registered Assessors
Program; Protected / Secret levels for government workloads.</li>
</ul>
<h3 id="sectoral-overlays">Sectoral overlays</h3>
<ul>
<li>Classified workloads (most jurisdictions): national
classification overlay</li>
<li>Healthcare: national health-data sovereignty rules</li>
<li>Critical-infrastructure: sectoral cybersecurity overlays</li>
</ul>
<h2 id="what-substantively-sovereign-means">What &ldquo;substantively sovereign&rdquo; means</h2>
<p>A &ldquo;sovereign cloud&rdquo; product that doesn&rsquo;t satisfy all of the following
substantively will fail under audit at the High-assurance level of
most frameworks:</p>
<h3 id="1-data-residency-at-every-layer">1. Data residency at every layer</h3>
<p>Not just production storage. Backup, observability, CI/CD artefacts,
managed-service telemetry, cross-border replication, sub-contractor
processing — every layer must respect the residency requirement.</p>
<h3 id="2-customer-controlled-encryption-keys">2. Customer-controlled encryption keys</h3>
<p>HSM-backed for sensitive data classes. Documented rotation. Emergency-
access procedures. Provider personnel cannot extract or copy keys
under any circumstances. This is the most common point where
hyperscaler &ldquo;sovereign&rdquo; offerings fall short — the provider retains
operational access to keys, which fails the substantive condition.</p>
<h3 id="3-open-source-platform-foundation">3. Open-source platform foundation</h3>
<p>For transparency, exit-readiness, and auditability. Closed-source
platforms tied to a single vendor&rsquo;s roadmap fail several frameworks'
substantive requirements (even where they pass procurement-policy
checks).</p>
<h3 id="4-supplier-chain-transparency">4. Supplier-chain transparency</h3>
<p>Supply-chain provisions across the frameworks (DORA Art. 28, NIS2
Art. 21(2)(d), national schemes) expect
documentation of the supplier chain at least to the second hop. Most
hyperscaler-based sovereign-cloud arrangements stop at the first hop
(the hyperscaler itself).</p>
<h3 id="5-air-gap-deployment-option">5. Air-gap deployment option</h3>
<p>For the most sensitive workloads — classified,
healthcare-with-strict-residency. Updates flow through controlled
channels (customer-side artefact registry, manual approval). Most
sovereign-cloud frameworks at High level require air-gap support as
an architectural option even if not used by every workload.</p>
<h3 id="6-audit-trail-completeness-in-standard-formats">6. Audit-trail completeness in standard formats</h3>
<p>Logs in standard formats (Syslog, CEF, OpenTelemetry) that the
customer&rsquo;s audit team can consume independently of the platform
vendor. Tamper-evident. Retention per the longest applicable
regulatory requirement.</p>
<h3 id="7-no-phone-home-telemetry">7. No phone-home telemetry</h3>
<p>Telemetry that leaves the customer perimeter must be opt-in and
explicitly documented. Many hyperscaler-managed cloud products have
non-optional telemetry channels that fail this criterion.</p>
<h3 id="8-operational-independence-under-sovereign-jurisdiction">8. Operational independence under sovereign jurisdiction</h3>
<p>Provider personnel access logged and time-limited. Sovereign
jurisdiction for the support entity (Ænix has AENIX s.r.o. in
Czechia for EU contracts and AENIX INC in Delaware for US contracts).
No cross-jurisdictional support routing for sovereignty-sensitive
workloads.</p>
<h2 id="what-cozystack-based-architecture-delivers-across-frameworks">What Cozystack-based architecture delivers across frameworks</h2>
<p>The architectural pattern that satisfies all major frameworks
simultaneously:</p>
<ul>
<li><strong>Open-source platform</strong> — Cozystack under Apache 2.0, CNCF Project,
vendor-neutral substrate. Customer can audit, modify, or replace
the platform vendor.</li>
<li><strong>Customer-held keys</strong> — volume encryption is opt-in per storage
class and key management is designed with you; Ænix never holds keys.</li>
<li><strong>Air-gap support</strong> — documented air-gapped install workflow for
classified-data use cases.</li>
<li><strong>Self-hosted observability</strong> — VictoriaMetrics + VictoriaLogs in
jurisdiction; no SaaS-observability residency leak.</li>
<li><strong>Customer-controlled identity</strong> — Keycloak / Active Directory /
national IdP integration; Ænix never holds production credentials.</li>
<li><strong>Multi-tenant Tenant CRD</strong> — strong isolation per data class /
business unit / sectoral overlay.</li>
<li><strong>Audit-isolated environments</strong> — separate clusters for production,
audit, and forensic copy.</li>
<li><strong>EU jurisdiction support entity</strong> — AENIX s.r.o. (Czechia).</li>
</ul>
<p>The architectural pattern is the same; the certification work is
framework-specific. For SecNumCloud-tier engagements, the customer
typically engages a certified auditor; Ænix provides the architecture
and documentation deliverables, the customer runs the audit cycle.</p>
<h2 id="procurement-realities">Procurement realities</h2>
<p>Public-sector engagements are procurement-framework-driven in a way
that private-sector engagements are not. A few practical realities:</p>
<h3 id="tender-response">Tender response</h3>
<p>Public-sector RFPs typically specify which frameworks must be
satisfied (SecNumCloud High, BSI C5, EUCS Substantial, etc.). The
response must demonstrate substantive compliance, not just intent.
Ænix engagement model includes tender-response support; references
can be shared under NDA where the customer allows it.</p>
<h3 id="multi-year-framework-agreements">Multi-year framework agreements</h3>
<p>Many public-sector engagements run through framework agreements with
specific compliance and exit clauses. Ænix&rsquo;s commercial entity
(AENIX s.r.o. in Czechia for EU; AENIX INC in Delaware for US) is
the contracting entity; engagement structure adapts to framework-
agreement requirements.</p>
<h3 id="ænix-is-not-a-hyperscaler--thats-the-point">Ænix is not a hyperscaler — that&rsquo;s the point</h3>
<p>Several public-sector mandates explicitly require non-hyperscaler
sovereign provision. Ænix&rsquo;s open-core model — customer hardware,
customer-held keys where encryption is enabled, customer operational
control, optional Ænix support
— fits these mandates structurally rather than via contractual
workarounds.</p>
<h2 id="engagement-phases-for-public-sector">Engagement phases for public-sector</h2>
<h3 id="phase-0--framework-scoping">Phase 0 — Framework scoping</h3>
<p>Confirm applicable frameworks. Identify the highest-bar one (usually
SecNumCloud High for French, BSI C5 for German, EUCS High where
EU-wide requirements apply, national procurement rules elsewhere). Design the
architecture against the highest bar; map down to the others.</p>
<h3 id="phase-1--architecture-and-procurement-response-work">Phase 1 — Architecture and procurement-response work</h3>
<p>Produce procurement-response artefacts: technical proposal, framework
compliance mapping, reference architecture, sample evidence catalogue.
Typical duration: 2-4 months.</p>
<h3 id="phase-2--phase-1-platform-build-per-public-cloud-platform-">Phase 2 — Phase-1 platform build (per Public Cloud Platform /</h3>
<p>Private Cloud Platform models)</p>
<p>Multi-DC deployment, air-gap option enabled if applicable, sovereign
identity integration, audit-isolated environments. A national
multi-region programme runs a 3-6 month pilot, then 9-18 months to full
multi-region operation; a single-agency private cloud is a 3-12 month
build after a 14- or 28-day assessment.</p>
<h3 id="phase-3--certification-cycle">Phase 3 — Certification cycle</h3>
<p>Customer engages accredited auditor; Ænix provides architecture
documentation, control mapping, evidence catalogue. Ænix engineers
participate in technical interviews with the auditor where allowed.
Typical certification cycle: 6-12 months parallel to Phase 2.</p>
<h3 id="phase-4--production-operations">Phase 4 — Production operations</h3>
<p>Customer team operates the platform with Ænix advisory and a Plus or
Enterprise support tier (see <a href="https://aenix.io/pricing/">/pricing/</a>). Annual
recertification cycle (most frameworks).</p>
<p>Certified production comes later than in private-sector engagements
because of the certification cycle, but the certification value compounds —
once certified, the platform retains certification with annual
recertification rather than per-engagement.</p>
<h2 id="ænixs-existing-public-sector-posture">Ænix&rsquo;s existing public-sector posture</h2>
<p>Ænix contracts through AENIX s.r.o. (EU) and AENIX INC (US) and works
within public-sector procurement frameworks. Specific engagements are
confidential; references are available under NDA in the discovery call.</p>
<h2 id="when-this-engagement-model-fits">When this engagement model fits</h2>
<p>Strong fit:</p>
<ul>
<li>National sovereign cloud initiatives (public, public-private
partnership, sovereign cloud operator)</li>
<li>EU member-state regional / sectoral cloud programmes</li>
<li>Classified-data hosting with an air-gap requirement</li>
<li>Healthcare sovereign cloud at national or regional level</li>
<li>Education / research consortia with multi-decade planning horizon</li>
</ul>
<p>Marginal fit:</p>
<ul>
<li>Single ministry / agency procurement with smaller scope — may fit
Private Cloud Platform rather than Public Cloud Platform</li>
</ul>
<p>Poor fit:</p>
<ul>
<li>Workloads where hyperscaler-managed cloud is already framework-
compliant (some specific procurement frameworks)</li>
<li>Organisations without sovereignty pressure (use the private-sector
product matching the workload profile)</li>
</ul>
<h2 id="where-to-dig-deeper">Where to dig deeper</h2>
<ul>
<li><strong><a href="https://aenix.io/industries/public-sector/">Public-sector industry page</a></strong> — the
commercial landing</li>
<li><strong><a href="https://aenix.io/solutions/data-sovereignty/">Data sovereignty services</a></strong> —
buyer-trigger sovereignty landing</li>
<li><strong><a href="https://aenix.io/solutions/nis2-compliance/">NIS2 compliance services</a></strong> —
for the NIS2 overlay (public administration is Annex I)</li>
<li><strong><a href="https://aenix.io/services/sovereign-cloud-builder/">Sovereign Cloud Builder services</a></strong> —
the engagement type</li>
<li><strong><a href="https://aenix.io/blog/2026/05/build-sovereign-cloud-eu-and-central-asia/">Build sovereign cloud — playbook for EU and Central Asia</a></strong> —
EU and Central Asia sovereign cloud playbook</li>
<li><strong><a href="https://aenix.io/blog/2026/05/data-residency-requirements-2026/">Data residency requirements in 2026</a></strong> —
per-layer residency walkthrough</li>
</ul>
]]></content:encoded></item><item><title>Public Cloud Platform at operator scale — what it takes to launch a national sovereign cloud</title><link>https://aenix.io/blog/2026/05/public-cloud-edition-multi-tenant-cloud-builder/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/public-cloud-edition-multi-tenant-cloud-builder/</guid><pubDate>Mon, 25 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Cozystack</category><category>Multi-tenancy</category><category>Sovereignty</category><category>Cloud</category><category>Platform Engineering</category><description>What an operator-scale sovereign cloud build on Ænix Public Cloud Platform covers for telcos, banks and national operators: phases, timeline, team, pitfalls.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/public-cloud-edition-multi-tenant-cloud-builder.jpg" alt=""></p><p>Most hosting providers run Ænix Public Cloud Platform at provider
scale: the productized installer puts the platform live in weeks once
hardware is ready, priced from the published support tiers. This post
is about the other end — the operator-scale programme. The question
there is not &ldquo;should we use Cozystack?&rdquo; — that&rsquo;s already decided. It&rsquo;s
&ldquo;we are launching a cloud product at national or tier-1-customer scale;
what does the partnership with Ænix look like across a 3-6 month pilot
and the 9-18 months to full multi-region operation?&rdquo;</p>
<h2 id="who-runs-public-cloud-platform-at-operator-scale">Who runs Public Cloud Platform at operator scale</h2>
<p>Five buyer profiles dominate operator-scale Public Cloud Platform engagements:</p>
<ol>
<li><strong>Tier-1 telcos / national operators</strong> — incumbent telecom
operators launching or scaling a public cloud as part of their
product portfolio. Often paired with sovereignty positioning
(&ldquo;our sovereign cloud&rdquo;, &ldquo;national cloud&rdquo;).</li>
<li><strong>Big banks operating their own cloud</strong> — the bank consumes its
own cloud product for internal workloads and, sometimes, sells
capacity to its customer base.</li>
<li><strong>Sovereign cloud initiatives</strong> — government-mandated cloud
products, sometimes with public-private-partnership structure,
with explicit sovereignty requirements and regulator alignment.</li>
<li><strong>Hosting providers at large scale</strong> — providers above ~5,000
customers where the Public Cloud Platform operational model needs scaling
into multi-region with multi-DC active/active.</li>
<li><strong>National AI/GPU operators</strong> — sustained inference + training
capacity for sectoral customers (banks, healthcare, public
sector) where AI sovereignty is a national-level
requirement.</li>
</ol>
<p>All five share the same operational reality: multi-region or
multi-DC active/active; multi-million-euro infrastructure investment;
customer-facing SLAs that map to national regulator expectations; and
a partnership model with Ænix that lasts years, not months.</p>
<h2 id="what-an-operator-scale-build-adds">What an operator-scale build adds</h2>
<h3 id="multi-region--multi-dc-activeactive">Multi-region / multi-DC active/active</h3>
<p>Single-DC deployments are served by Public Cloud Platform at provider
scale (or Private Cloud Platform for internal use). Operator-scale builds
assume from day one that the customer needs
active/active across regions or datacentres with cross-DC replication
tuned for RTO/RPO targets. The platform&rsquo;s control plane, observability,
identity, and storage layers all design for multi-region from the
foundation rather than retrofitting.</p>
<h3 id="service-catalog-depth">Service-catalog depth</h3>
<p>A provider-scale build exposes ~20 managed services. An operator-scale build
typically targets 30-50+ services across compute, storage, networking,
managed databases, observability, AI/GPU, message queues, search,
content delivery, security tooling. Cozystack&rsquo;s package architecture
(Package + PackageSource + ApplicationDefinition resources, as of
v1.x) supports the catalog expansion.</p>
<h3 id="operations-team-at-scale">Operations team at scale</h3>
<p>10-30+ engineers running the platform, depending on customer count
and SLA. An operator-scale engagement includes operations team
hiring and training as a substantial workstream — not &ldquo;you find
people, we&rsquo;ll train them&rdquo; but &ldquo;we design the org structure with you,
participate in interviews, do hands-on training, and provide escalation
support (Plus or Enterprise tier) for the first 12-18 months while your
team builds confidence.&rdquo;</p>
<h3 id="regulator-and-sovereignty-alignment">Regulator and sovereignty alignment</h3>
<p>Whatever the sovereignty framework is in the customer&rsquo;s market —
SecNumCloud, BSI C5, EUCS, sectoral overlays, national procurement
mandates — the architecture is designed to satisfy it substantively,
not just contractually. Compliance evidence catalogue is a deliverable.</p>
<h3 id="customer-facing-brand-engineering">Customer-facing brand engineering</h3>
<p>Beyond Cozystack Dashboard customisation, an operator-scale engagement includes
brand-engineering work: customer portal that looks like a top-tier
cloud product, not a customised Cozystack instance. UX flows tuned
to how customer&rsquo;s customers think about ordering, configuring,
paying. Designer-led, not engineering-led.</p>
<h2 id="how-an-operator-scale-engagement-phases">How an operator-scale engagement phases</h2>
<p>Phases 0 and 1 form the pilot and take 3-6 months together. Phases 2-4
take 9-18 months to full multi-region operation, overlapping where the
teams allow.</p>
<h3 id="phase-0--discovery-and-partnership-formation-start-of-the-pilot">Phase 0 — Discovery and partnership formation (start of the pilot)</h3>
<p>Before engineering, agreement on:</p>
<ul>
<li>Strategic objectives (what cloud product, what customer base, what
competitive positioning)</li>
<li>Regulatory scope (which frameworks bind the platform)</li>
<li>Org structure (who owns what; how Ænix and customer teams interact)</li>
<li>Commercial structure (engagement model, IP, support model post-go-live)</li>
<li>Roadmap (phasing of services, geographic expansion, SLA tiers)</li>
</ul>
<p>Output: signed engagement plan with named workstream leads on both
sides.</p>
<h3 id="phase-1--foundation-rest-of-the-pilot">Phase 1 — Foundation (rest of the pilot)</h3>
<p>Hardware procurement and racking. Talos / Cozystack platform deployment
in the first datacentre. Storage layer (LINSTOR/DRBD at scale).
Networking foundation. Identity integration with customer&rsquo;s existing
workforce identity (Keycloak / Okta / Active Directory / sovereign IdP).
Initial observability stack.</p>
<p>End state: working platform, single region, internal access only.
Not yet customer-ready.</p>
<h3 id="phase-2--multi-region-foundation">Phase 2 — Multi-region foundation</h3>
<p>Second datacentre stood up. Cross-DC replication validated. Federated
identity. Multi-region storage replication (LINSTOR async or Ceph
cross-region). Disaster recovery patterns tested. Compliance
documentation foundation built.</p>
<p>End state: multi-DC platform, internal access, RTO/RPO validated
against targets.</p>
<h3 id="phase-3--service-catalog-buildout">Phase 3 — Service catalog buildout</h3>
<p>Service-by-service rollout. Start with foundational services (compute,
storage, basic networking, managed PostgreSQL). Layer in managed
service families (databases, queues, caches, search, observability).
Add product-specific services (GPU, AI inference, sectoral compliance
tooling).</p>
<p>Each service goes through: deployment → internal testing → friendly-
customer pilot → production GA. Cohort-based rollout, not big-bang.</p>
<h3 id="phase-4--customer-onboarding-and-limited-ga">Phase 4 — Customer onboarding and limited GA</h3>
<p>Customer-facing portal launched (brand-engineered). Billing integration
validated end-to-end. Support runbooks documented. First 10-50 friendly
customers onboarded. SLA monitoring operationalised.</p>
<p>End state: cloud product live with first customer cohort, billing
and support workflows proven.</p>
<h3 id="phase-5--general-availability-and-scale-ongoing">Phase 5 — General availability and scale (ongoing)</h3>
<p>Open market launch. Marketing and sales activated. Operations team
scales to support customer growth. Ænix escalation support (Plus or
Enterprise tier) continues until the customer team is ready to absorb
it (typically 12-24 months post-GA).</p>
<p>Subsequent phases are roadmap-driven: new services, new regions, new
sectoral SKUs.</p>
<h2 id="where-multi-million-euro-cloud-projects-fail">Where multi-million-euro cloud projects fail</h2>
<p>Three failure patterns we&rsquo;ve seen across the industry:</p>
<h3 id="1-under-investing-in-brand-engineering">1. Under-investing in brand engineering</h3>
<p>Engineering-led platform with engineering-grade UX. Customers click
around, find it functional but unappealing, sign up for hyperscaler
instead. An operator-scale engagement includes design partnership
explicitly to avoid this.</p>
<h3 id="2-operations-team-sized-for-go-live-not-18-month-out-volume">2. Operations team sized for go-live, not 18-month-out volume</h3>
<p>Cloud products grow exponentially during the first year of GA if
positioning is right. Operations teams sized for go-live customer
count get overwhelmed at month 6-12. Plan operations capacity for
18-month-out volume; hire ahead.</p>
<h3 id="3-regulator-dialog-deferred">3. Regulator dialog deferred</h3>
<p>Sovereignty positioning depends on regulator endorsement (explicit or
implicit). Projects that defer the regulator conversation until
late-phase find themselves rebuilding architecture to satisfy
expectations they could have designed for at the start. Engage
regulators in Phase 0-1, not Phase 4.</p>
<h2 id="when-an-operator-scale-programme-is-the-right-answer">When an operator-scale programme is the right answer</h2>
<p>Strong fit:</p>
<ul>
<li>Tier-1 telco / national operator / large bank / sovereign cloud
initiative</li>
<li>Multi-region or multi-DC operational reality</li>
<li>5,000+ target customer count or strategic customer base</li>
<li>Multi-million-euro budget envelope across a multi-year programme</li>
<li>Sovereignty / regulator positioning is core to value proposition</li>
<li>Senior executive sponsorship (CIO or CTO level minimum)</li>
</ul>
<p>Marginal fit:</p>
<ul>
<li>Large hosting providers below tier-1 telco scale — Public Cloud
Platform at provider scale, extended region by region, often fits
better than a full operator programme, depending on growth profile</li>
<li>AI/GPU-focused operators where the AI workload dominates — the Ænix AI
Platform may fit better, with selective Public Cloud Platform
components</li>
</ul>
<p>Poor fit:</p>
<ul>
<li>Smaller hosting providers — Public Cloud Platform at provider scale
(from the <a href="https://aenix.io/pricing/">published price list</a>) fits substantially better
on economics and operational model</li>
<li>Regulated enterprises consuming cloud rather than producing it —
Private Cloud Platform is the right answer</li>
</ul>
<h2 id="engagement-structure">Engagement structure</h2>
<ul>
<li><strong>Discovery call</strong> (executive level, 60-90 min) — strategic fit
assessment</li>
<li><strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong>
(fixed price, 28 days full) — input to partnership formation;
output is a signed engagement plan</li>
<li><strong>Pilot</strong> (Phases 0-1, 3-6 months), then <strong>Phases 2-4</strong> (9-18 months
to full multi-region operation)</li>
<li><strong>Support subscription</strong> (ongoing) — Plus or Enterprise tier (see
<a href="https://aenix.io/pricing/">/pricing/</a>) until the customer team is ready to absorb
escalation</li>
</ul>
<p>Engagement size: multi-year programme, quoted per RFP.</p>
<h2 id="where-to-dig-deeper">Where to dig deeper</h2>
<ul>
<li><strong><a href="https://aenix.io/products/public-cloud-platform/">Public Cloud Platform landing</a></strong> —
feature list, product-specific FAQ</li>
<li><strong><a href="https://aenix.io/services/public-cloud-builder/">Public Cloud Builder services</a></strong> —
engagement details</li>
<li><strong><a href="https://aenix.io/services/sovereign-cloud-builder/">Sovereign Cloud Builder services</a></strong> —
for the sovereignty-specific variant</li>
<li><strong><a href="https://aenix.io/blog/2026/05/build-sovereign-cloud-eu-and-central-asia/">Build sovereign cloud — playbook for EU and Central Asia</a></strong> —
sovereign cloud architectural patterns</li>
</ul>
]]></content:encoded></item><item><title>Proxmox vs VMware vs Cozystack — a 2026 comparison for the post-Broadcom era</title><link>https://aenix.io/blog/2026/05/proxmox-vs-vmware-vs-cozystack-comparison/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/proxmox-vs-vmware-vs-cozystack-comparison/</guid><pubDate>Sun, 24 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>VMware</category><category>Proxmox</category><category>Kubernetes</category><category>Cozystack</category><category>Sovereignty</category><category>Multi-tenancy</category><description>Proxmox VE, VMware after Broadcom, and Cozystack compared by architecture and use case, with a feature matrix and the realistic migration paths.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/proxmox-vs-vmware-vs-cozystack-comparison.jpg" alt=""></p><p>The post-Broadcom virtualization market has three main options this article compares: Proxmox VE and Cozystack (both open source) and VMware itself (XCP-ng is a less common fourth). Each has a different architectural target. Picking the right one is mostly a function of scale and use case.</p>
<h2 id="proxmox-ve--smb-friendly-vm-focused">Proxmox VE — SMB-friendly, VM-focused</h2>
<p><strong>Architecture:</strong> KVM + LXC + ZFS + Ceph (community), single-cluster Proxmox VE, no native multi-tenancy.</p>
<p><strong>Strengths:</strong></p>
<ul>
<li>Mature, stable, easy to install.</li>
<li>Strong community.</li>
<li>AGPLv3 licence, commercial subscription available.</li>
<li>Excellent for single-team or single-tenant deployments.</li>
<li>Proxmox Backup Server is good.</li>
</ul>
<p><strong>Limits:</strong></p>
<ul>
<li>Multi-tenancy through namespaces and permissions; not designed for hard isolation.</li>
<li>Service catalog beyond VMs (managed databases, S3, etc.) requires manual integration.</li>
<li>Service-provider use cases (billing per tenant, self-service portal) require external software.</li>
<li>Federation across clusters is heavier than Kubernetes.</li>
</ul>
<p><strong>Best for:</strong> SMB IT departments, labs, dev environments, single-tenant private virtualization, teams under ~50 hosts.</p>
<h2 id="vmware-post-broadcom--enterprise-legacy">VMware (post-Broadcom) — enterprise legacy</h2>
<p><strong>Architecture:</strong> vSphere + vSAN + NSX + vCloud Director + vRealize/Aria. Closed source, subscription-licensed.</p>
<p><strong>Strengths:</strong></p>
<ul>
<li>Mature, well-known, extensive ecosystem.</li>
<li>Strong enterprise tooling integration.</li>
<li>Broad operational expertise in market.</li>
</ul>
<p><strong>Limits:</strong></p>
<ul>
<li>Subscription-only licensing (post-Broadcom 2023).</li>
<li>2-5× price increases on renewal observed across our pipeline.</li>
<li>Vendor lock-in across the stack.</li>
<li>Sovereignty concerns (US-headquartered vendor).</li>
</ul>
<p><strong>Best for:</strong> Existing VMware estates that haven&rsquo;t yet been triggered out by economics. New deployments rarely choose VMware in 2026.</p>
<p>(See <strong><a href="https://aenix.io/alternatives/vmware-alternative/">VMware alternative</a></strong> for migration guidance.)</p>
<h2 id="cozystack--open-source-kubernetes-native-multi-tenant">Cozystack — open-source, Kubernetes-native, multi-tenant</h2>
<p><strong>Architecture:</strong> KubeVirt + Cilium + Kube-OVN + LINSTOR (DRBD) + Tenant CRD + Cozystack Dashboard. Open-source CNCF Project.</p>
<p><strong>Strengths:</strong></p>
<ul>
<li>Kubernetes-native virtualization — same platform for VMs, containers, databases.</li>
<li>Multi-tenancy structural (Tenant CRD) — production-grade for service providers and regulated multi-tenant.</li>
<li>First-class managed database, S3 object storage, GPU services.</li>
<li>Apache 2.0 licence, no per-CPU pricing.</li>
<li>Air-gapped deployment supported.</li>
</ul>
<p><strong>Limits:</strong></p>
<ul>
<li>Newer than Proxmox or VMware; smaller community.</li>
<li>Kubernetes operational expertise required (mitigated by Ænix support tier).</li>
<li>Not optimized for single-tenant SMB use case (Proxmox better here).</li>
</ul>
<p><strong>Best for:</strong> Service providers, regulated enterprises, multi-team platforms, AI/GPU operators, sovereign-cloud builders.</p>
<h2 id="comparison-matrix">Comparison matrix</h2>
<table>
  <thead>
      <tr>
          <th></th>
          <th>Proxmox VE</th>
          <th>VMware (VCF)</th>
          <th>Cozystack</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><strong>Licence</strong></td>
          <td>AGPLv3 + commercial subscription</td>
          <td>Subscription-only</td>
          <td>Apache 2.0</td>
      </tr>
      <tr>
          <td><strong>Compute</strong></td>
          <td>KVM + LXC</td>
          <td>vSphere</td>
          <td>KubeVirt (KVM) + K8s</td>
      </tr>
      <tr>
          <td><strong>Storage</strong></td>
          <td>ZFS, Ceph</td>
          <td>vSAN</td>
          <td>LINSTOR (DRBD)</td>
      </tr>
      <tr>
          <td><strong>Network</strong></td>
          <td>Linux SDN</td>
          <td>NSX</td>
          <td>Cilium</td>
      </tr>
      <tr>
          <td><strong>Multi-tenancy</strong></td>
          <td>Namespace + permissions</td>
          <td>vCloud Director</td>
          <td>Tenant CRD</td>
      </tr>
      <tr>
          <td><strong>Managed databases</strong></td>
          <td>Manual / community</td>
          <td>Limited</td>
          <td>First-class (PostgreSQL, MariaDB, MongoDB, Redis, Valkey, Kafka, ClickHouse, OpenSearch, etc.)</td>
      </tr>
      <tr>
          <td><strong>S3 object storage</strong></td>
          <td>Manual</td>
          <td>Limited</td>
          <td>First-class</td>
      </tr>
      <tr>
          <td><strong>GPU</strong></td>
          <td>Passthrough</td>
          <td>vGPU under Horizon</td>
          <td>Passthrough of whole GPUs or NVIDIA vGPU (requires your NVIDIA vGPU licence) for VMs; HAMi fractional sharing for containers; MIG and time-slicing on the roadmap</td>
      </tr>
      <tr>
          <td><strong>Self-service</strong></td>
          <td>Web UI for ops</td>
          <td>vCD</td>
          <td>Cozystack Dashboard</td>
      </tr>
      <tr>
          <td><strong>Backup/DR</strong></td>
          <td>PBS</td>
          <td>SRM</td>
          <td>Velero + PG PITR</td>
      </tr>
      <tr>
          <td><strong>Best scale</strong></td>
          <td>&lt;50 hosts</td>
          <td>Enterprise</td>
          <td>Multi-tenant scale</td>
      </tr>
      <tr>
          <td><strong>Best for</strong></td>
          <td>SMB, labs</td>
          <td>Existing VMware</td>
          <td>Cloud builders, regulated multi-tenant</td>
      </tr>
  </tbody>
</table>
<h2 id="migration-paths">Migration paths</h2>
<h3 id="proxmox--cozystack">Proxmox → Cozystack</h3>
<p>VM images (qcow2) import directly into KubeVirt CDI. Multi-tenant model designed during migration (Proxmox didn&rsquo;t have one to migrate). Storage and network re-architecture. Typical: a 14- or 28-day assessment + 3-9 months implementation.</p>
<h3 id="vmware--cozystack">VMware → Cozystack</h3>
<p>KubeVirt-based migration with image conversion. Windows VMs supported; specific tooling for VMware Tools cleanup. Multi-tenant model maps from vCloud Director to Tenant CRD. (Full guidance: <strong><a href="https://aenix.io/alternatives/vmware-alternative/">VMware alternative landing</a></strong>.)</p>
<h3 id="proxmox--vmware">Proxmox → VMware</h3>
<p>Rare in 2026; reverse migration usually doesn&rsquo;t make economic sense post-Broadcom.</p>
<h2 id="how-to-choose">How to choose</h2>
<ol>
<li><strong>You&rsquo;re under 50 hosts, single-tenant, mostly VMs:</strong> Proxmox VE.</li>
<li><strong>You&rsquo;re a service provider, multi-tenant cloud, regulated multi-tenant:</strong> Cozystack.</li>
<li><strong>You&rsquo;re already on VMware and the budget supports staying:</strong> stay (but plan an exit). If renewal pressures bite: see VMware alternative.</li>
<li><strong>AI/GPU workloads at scale:</strong> Cozystack (KubeVirt + GPU operators).</li>
<li><strong>Pure container workloads, no VMs:</strong> vanilla Kubernetes (Cozystack still works but is over-spec).</li>
</ol>
]]></content:encoded></item><item><title>Proxmox to Cozystack — when single-tenant outgrows itself</title><link>https://aenix.io/blog/2026/05/proxmox-migration-when-cozystack-fits/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/proxmox-migration-when-cozystack-fits/</guid><pubDate>Sat, 23 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Proxmox</category><category>Cozystack</category><category>Migration</category><category>Multi-tenancy</category><category>Hosting</category><description>When does Proxmox VE outgrow single-tenant? A Proxmox-to-Cozystack migration guide for MSPs and growing teams hitting multi-tenancy and scale limits.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/proxmox-migration-when-cozystack-fits.jpg" alt=""></p><p>Proxmox VE is one of the most successful open-source virtualisation
platforms of the last decade. Mature, easy to install, strong
community, AGPLv3 with commercial subscription. We talk to a lot of
operators who started on Proxmox, grew, and are evaluating what comes
next.</p>
<p>Crucially: Proxmox is the right answer for many of them. This article
covers when migration is warranted and when it&rsquo;s premature.</p>
<h2 id="where-proxmox-keeps-winning">Where Proxmox keeps winning</h2>
<p>Proxmox VE remains the right answer for:</p>
<ul>
<li><strong>SMB IT departments</strong> — small-to-mid businesses running 10-50
virtualised workloads on premises, no customer-facing multi-
tenancy needs</li>
<li><strong>Single-tenant labs and dev environments</strong> — Proxmox&rsquo;s
operational simplicity beats any heavier alternative</li>
<li><strong>Mature operators with stable customer base under ~200 customers</strong> —
Proxmox&rsquo;s commercial economics still work; migration cost would
exceed the value</li>
<li><strong>Mostly-VM workloads</strong> — Proxmox&rsquo;s KVM + LXC scope fits cleanly</li>
<li><strong>Existing operators with deep Proxmox expertise and stable team</strong> —
switching cost includes team retraining</li>
</ul>
<p>If your situation matches these, <em>don&rsquo;t migrate</em>. The Ænix Public Cloud
Platform is over-engineered for SMB single-tenant operation. We say
this in discovery calls rather than push the engagement.</p>
<h2 id="when-proxmox-is-being-outgrown">When Proxmox is being outgrown</h2>
<p>Migration warrants serious evaluation when at least three of these
hold:</p>
<h3 id="1-customer-count-growing-past-300">1. Customer count growing past ~300</h3>
<p>Proxmox&rsquo;s multi-tenancy model (pools, realms and permissions, not hard
isolation) starts to feel thin above ~300 customer-facing tenants.
Per-customer audit trails, isolation guarantees, and quota enforcement
become operational pain.</p>
<h3 id="2-customers-asking-for-services-beyond-vms">2. Customers asking for services beyond VMs</h3>
<p>Managed PostgreSQL, MariaDB, MongoDB, Redis, Valkey, Kafka, S3-compatible object
storage, tenant Kubernetes clusters, GPU services. Proxmox&rsquo;s scope is
VMs + LXC; everything else is bolted on with manual integration or
external systems.</p>
<h3 id="3-whmcs-or-similar-customer-management-integration">3. WHMCS or similar customer-management integration</h3>
<p>Proxmox has WHMCS integration, but the service catalog beyond VMs is
manual integration work. Ænix Public Cloud Platform adds a WHMCS
integration (a proprietary Ænix module, not part of open-source
Cozystack) that covers the full service catalog.</p>
<h3 id="4-multi-dc-activeactive">4. Multi-DC active/active</h3>
<p>Proxmox clustering is single-DC. Geographic distribution requires
manual cross-cluster replication patterns. Cozystack handles multi-
DC active/active as a first-class deployment mode.</p>
<h3 id="5-container-native-customer-demand">5. Container-native customer demand</h3>
<p>Customers want tenant Kubernetes clusters or container-native
service catalogs. Proxmox can host containers via LXC but isn&rsquo;t the
right operational model for tenant-facing Kubernetes-as-a-service.</p>
<h3 id="6-recurring-licence--subscription-pressure-on-commercial-proxmox">6. Recurring licence / subscription pressure on commercial Proxmox</h3>
<p>Proxmox&rsquo;s commercial subscription is competitive but real cost.
Operators with growing infrastructure footprint sometimes find the
total subscription cost approaching what Ænix charges for Public Cloud
Platform support — at which point the service-catalog and operational
upside of Cozystack tips the decision.</p>
<h2 id="architectural-mapping-proxmox--cozystack">Architectural mapping: Proxmox → Cozystack</h2>
<table>
  <thead>
      <tr>
          <th>Proxmox VE</th>
          <th>Cozystack equivalent</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><strong>KVM hypervisor</strong></td>
          <td>KubeVirt (KVM-based)</td>
      </tr>
      <tr>
          <td><strong>LXC containers</strong></td>
          <td>Native Kubernetes containers (different model — LXC system-style vs Kubernetes application-style)</td>
      </tr>
      <tr>
          <td><strong>ZFS storage</strong></td>
          <td>LINSTOR (DRBD)</td>
      </tr>
      <tr>
          <td><strong>Ceph (Proxmox-managed)</strong></td>
          <td>LINSTOR (DRBD); Cozystack does not ship Ceph</td>
      </tr>
      <tr>
          <td><strong>Linux SDN / bridges</strong></td>
          <td>Cilium (eBPF)</td>
      </tr>
      <tr>
          <td><strong>Proxmox web UI</strong></td>
          <td>Cozystack Dashboard</td>
      </tr>
      <tr>
          <td><strong>Proxmox Backup Server (PBS)</strong></td>
          <td>Velero + S3-compatible target + per-app PITR</td>
      </tr>
      <tr>
          <td><strong>PVE-Storage replication</strong></td>
          <td>LINSTOR DRBD replication</td>
      </tr>
      <tr>
          <td><strong>Proxmox API / pvesh, qm, pct</strong></td>
          <td>Kubernetes API</td>
      </tr>
      <tr>
          <td><strong>Datacenter / Pool / VM</strong></td>
          <td>Tenant CRD + namespace + KubeVirt VM</td>
      </tr>
      <tr>
          <td><strong>Permission model (roles)</strong></td>
          <td>Kubernetes RBAC + Tenant CRD scope</td>
      </tr>
  </tbody>
</table>
<p>Two areas need redesign rather than 1:1 mapping:</p>
<ul>
<li><strong>LXC vs Kubernetes containers</strong> — Proxmox LXC is system-container
(full OS image), Kubernetes container is application-container
(single process or small set). Workloads using LXC for system-
container patterns either migrate to KubeVirt VMs or get
refactored.</li>
<li><strong>Multi-tenancy model</strong> — Proxmox tenant model (pools, realms
and permissions) versus Cozystack Tenant CRD (Kubernetes-native).
Customer-facing isolation is stronger in Cozystack; operational
abstraction is different.</li>
</ul>
<h2 id="migration-phases">Migration phases</h2>
<h3 id="phase-0--assessment-14-or-28-days">Phase 0 — Assessment (14 or 28 days)</h3>
<p>Inventory: customer count, customer-facing services consumed, VM
count, OS mix, LXC usage, storage tiers, network topology, backup
patterns, WHMCS / customer-management integration.</p>
<p>Honest TCO comparison: current Proxmox + commercial subscription +
operational team versus Ænix Public Cloud Platform + hardware refresh +
Ænix support tier. For operators under ~300 customers, this often
shows Proxmox staying competitive; above ~500, Cozystack typically
wins on service-catalog and operational depth.</p>
<p>Output: go/no-go decision with quantified justification.</p>
<h3 id="phase-1--cozystack-foundation-live-in-weeks">Phase 1 — Cozystack foundation (live in weeks)</h3>
<p>The platform goes live in weeks once hardware is ready, using the
productized installer; catalogue and brand work continue alongside
the pilot. Cozystack platform deployed on new hardware or repurposed Proxmox
hardware (commodity x86 servers move easily). Cilium networking
configured. LINSTOR storage operationalised. Identity integration
(typically Keycloak + customer IdP). Cozystack Dashboard brand customisation
matching the operator&rsquo;s existing brand.</p>
<p>WHMCS integration validated end-to-end. Service catalog populated
with the operator&rsquo;s chosen services (VMs first, managed databases
next, S3 then, expanding from there).</p>
<h3 id="phase-2--pilot-customer-migration">Phase 2 — Pilot customer migration</h3>
<p>5-20 friendly customers migrated to Cozystack as the first cohort.
Pattern per customer:</p>
<ol>
<li>Customer VMs converted from Proxmox qcow2 to KubeVirt-compatible
format</li>
<li>Network configuration translated (Proxmox bridges → Cilium
ClusterPool + NetworkPolicies)</li>
<li>Storage migrated (ZFS / Ceph volumes → LINSTOR in
Cozystack)</li>
<li>Customer-side validation window (7-14 days)</li>
<li>DNS / load balancer cutover</li>
</ol>
<p>During the pilot, customer support team builds operational
familiarity with Cozystack. Documentation patterns shake out.</p>
<h3 id="phase-3--production-migration-cohorts">Phase 3 — Production migration cohorts</h3>
<p>Cohorts of 30-100 customers at a time. Same per-customer pattern as
pilot, with operational efficiency improvements as the team
internalises the workflow.</p>
<p>LXC-using customers receive special handling: either system-style
KubeVirt VM (1:1 replacement) or refactor to Kubernetes-native
application container (depending on customer&rsquo;s preference and
support).</p>
<h3 id="phase-4--proxmox-decommission">Phase 4 — Proxmox decommission</h3>
<p>As migration cohorts complete, Proxmox hardware moves into the
Cozystack cluster. Proxmox subscription wound down per renewal
cycle. Proxmox Backup Server data archived per customer agreements.</p>
<h2 id="timeline-realities">Timeline realities</h2>
<p>For typical mid-size hosting provider (300-1,000 customers):</p>
<ul>
<li>Phase 0: 14 or 28 days</li>
<li>Phase 1: platform live in weeks once hardware is ready</li>
<li>Phases 2-3 (pilot and production cohorts): typically 3-9 months</li>
<li>Phase 4: alongside the last cohorts, timed to Proxmox renewals</li>
</ul>
<p><strong>Total: the platform is live in weeks once hardware is ready;
moving workloads and customers typically takes 3-9 months</strong></p>
<p>For larger operators (1,000-5,000 customers), Phase 3 runs longer
for sustainable cohort pacing; the assessment sets the schedule.</p>
<h2 id="where-proxmox-to-cozystack-migrations-stumble">Where Proxmox-to-Cozystack migrations stumble</h2>
<h3 id="1-lxc-workloads">1. LXC workloads</h3>
<p>If a substantial fraction of customer workloads use LXC for
system-container patterns (e.g., per-customer LAMP stack as a single
LXC), the migration to Kubernetes-native containers requires
refactoring. The alternative is running them as KubeVirt VMs (1:1
mapping but heavier resource footprint). Plan time for this in
Phase 0.</p>
<h3 id="2-customer-facing-api-divergence">2. Customer-facing API divergence</h3>
<p>Some customers built tooling against the Proxmox API. Cozystack
exposes Kubernetes API + Cozystack Dashboard API; the contracts differ.
Customer-facing migration support (documentation, sometimes API
compatibility shim) is engagement work.</p>
<h3 id="3-operations-team-training">3. Operations team training</h3>
<p>Proxmox operators are comfortable with the Proxmox web UI and the
imperative <code>qm</code> / <code>pct</code> / <code>pvesh</code> CLI tools. Cozystack expects GitOps for production
changes. Operations team needs 4-8 weeks of focused training plus
3-6 months of practice. Ænix engagement includes training; customer
investment in the transition is also required.</p>
<h3 id="4-zfs-specific-workloads">4. ZFS-specific workloads</h3>
<p>Some customers chose Proxmox specifically for ZFS-on-host features
(advanced snapshots, ZFS-replicated backups). Cozystack ships LINSTOR
(DRBD); ZFS-specific operational patterns don&rsquo;t translate. Customer
dialogue about feature equivalence is part of Phase 0.</p>
<h2 id="versus-other-alternatives">Versus other alternatives</h2>
<p><strong>Versus building it yourself on raw KVM + libvirt + Kubernetes:</strong>
Same trade-offs as for any open-source-build option. Cozystack
gets a multi-tenant platform to production in weeks to a few months;
raw builds take 12-24 months to reach the same level. For operators with
strong platform engineering capacity, the raw-build is a credible
alternative.</p>
<p><strong>Versus VMware (post-Broadcom):</strong> Proxmox-to-VMware migration is
rare in 2026 — reverse migration usually doesn&rsquo;t make economic sense
post-Broadcom.</p>
<p><strong>Versus Nutanix:</strong> Nutanix AHV is closed-source proprietary KVM.
For operators valuing open-source substrate, Cozystack wins on that
property alone. For operators valuing integrated commercial support
without open-source overhead, Nutanix wins.</p>
<p><strong>Versus OpenShift Virtualization:</strong> Both are KubeVirt-based. OpenShift
fits existing Red Hat / OpenShift customers; Cozystack fits operators
preferring open-source-first procurement and lighter operational
footprint.</p>
<h2 id="when-this-engagement-model-fits">When this engagement model fits</h2>
<p>Strong fit:</p>
<ul>
<li>Hosting provider or MSP with 300+ customers</li>
<li>Growth trajectory toward 1,000+ customers</li>
<li>Customer demand for services beyond VMs</li>
<li>Multi-DC operational reality</li>
<li>Budget for a migration programme (typically 3-9 months of customer moves)</li>
</ul>
<p>Marginal fit:</p>
<ul>
<li>200-300 customer providers — borderline; depends on growth
trajectory and service-catalog ambition</li>
</ul>
<p>Poor fit:</p>
<ul>
<li>SMB IT (&lt;100 internal VMs) — Proxmox is still better</li>
<li>Lab / dev environments — Proxmox simplicity wins</li>
<li>Sub-200-customer hosting providers — poor fit for a full migration
programme; consider a greenfield service line on Ænix Public Cloud
Platform at provider scale instead</li>
</ul>
<h2 id="engagement-structure">Engagement structure</h2>
<ul>
<li><strong>Discovery call</strong> (30 min, free)</li>
<li><strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong>
(fixed price, 14 days focused or 28 days full) — go/no-go with TCO
comparison</li>
<li><strong>Pilot deployment</strong> — Cozystack stood up (live in weeks once
hardware is ready), 5-20 friendly customers migrated</li>
<li><strong>Cohort migration</strong> — customer migration in cohorts; pilot and
cohorts together typically take 3-9 months</li>
<li><strong>Proxmox decommission</strong> (parallel) — as cohorts
complete</li>
<li><strong>Support subscription</strong> (ongoing) — Plus or Enterprise support tier
for 24×7 coverage (see <a href="https://aenix.io/pricing/">/pricing/</a>)</li>
</ul>
<h2 id="where-to-dig-deeper">Where to dig deeper</h2>
<ul>
<li><strong><a href="https://aenix.io/migration/proxmox/">Proxmox migration hub</a></strong> — commercial landing</li>
<li><strong><a href="https://aenix.io/blog/2026/05/proxmox-vs-vmware-vs-cozystack-comparison/">Proxmox vs VMware vs Cozystack comparison</a></strong> —
decision matrix</li>
<li><strong><a href="https://aenix.io/alternatives/proxmox-alternative/">Proxmox alternative</a></strong> —
alternative-focused commercial landing</li>
<li><strong><a href="https://aenix.io/industries/hosting-providers/">Hosting providers industry page</a></strong> —
industry-specific positioning</li>
<li><strong><a href="https://aenix.io/blog/2026/05/isp-edition-economics-hosting-providers/">Public Cloud Platform economics for hosting providers</a></strong> —
unit-economics walkthrough</li>
<li><strong><a href="https://aenix.io/blog/2026/05/hosting-provider-platform-modernization/">Hosting provider platform modernization</a></strong> —
modernisation pattern</li>
</ul>
]]></content:encoded></item><item><title>Private LLM deployment — a practical guide to on-premise AI infrastructure in 2026</title><link>https://aenix.io/blog/2026/05/private-llm-deployment-guide/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/private-llm-deployment-guide/</guid><pubDate>Sat, 23 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>DORA</category><category>Kubernetes</category><category>Sovereignty</category><category>AI and ML</category><category>GPU</category><category>Multi-tenancy</category><description>The six layers of a real private LLM deployment — hardware, platform, serving, model, application, operations — with the pitfalls at each one.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/private-llm-deployment-guide.jpg" alt=""></p><p>The decision to deploy a private LLM is increasingly easy to make and surprisingly hard to execute well. The decision is easy because the trigger is usually clear: regulated data, sectoral rules, or the economics of inference at scale make a hyperscaler model API the wrong answer. Execution is hard because the supporting infrastructure — GPU scheduling, model serving, observability, multi-tenancy, audit-readiness — is more work than the LLM itself.</p>
<h2 id="when-a-private-llm-is-the-right-answer">When a private LLM is the right answer</h2>
<p>Three trigger profiles dominate the conversations that lead to private LLM deployment.</p>
<h3 id="trigger-1--regulated-data-class">Trigger 1 — regulated data class</h3>
<p>The data your AI must process is bound to a jurisdiction by regulator, sector, or contract. Examples: customer financial records under DORA / sectoral banking rules; patient-data in healthcare; classified or sensitive public-sector data; legal-privilege content; internal IP whose exposure to a third-party model would constitute a contractual or competitive breach.</p>
<p>For these workloads, sending data to a hyperscaler model API — even one with a &ldquo;private&rdquo; deployment branding — fails the substantive requirement. The data must stay in the customer&rsquo;s perimeter, and so must the model.</p>
<h3 id="trigger-2--inference-economics-at-scale">Trigger 2 — inference economics at scale</h3>
<p>For 24/7 inference workloads at sustained throughput, hyperscaler GPU economics break down. A workload that consumes a few hundred GPU-hours per day at peak burst rates can cost more on a hyperscaler than a dedicated cluster does in a year. The crossover point varies by GPU class and utilization, but moves a lot of inference workloads onto on-premise economics around 30-50% sustained utilization.</p>
<h3 id="trigger-3--auditability-and-reproducibility">Trigger 3 — auditability and reproducibility</h3>
<p>Regulator dialog increasingly requires &ldquo;exactly which model produced this output, with which weights, with which input data, on which hardware, at what time.&rdquo; Hyperscaler model APIs answer the first question (model name + version) but not the others. Private deployment makes all of them auditable.</p>
<p>If your situation matches one of these triggers strongly, or two weakly, private LLM deployment moves from &ldquo;interesting option&rdquo; to &ldquo;obvious next step.&rdquo;</p>
<h2 id="what-goes-into-a-real-private-llm-stack">What goes into a real private LLM stack</h2>
<p>A production private LLM stack has six layers, all of which need an answer:</p>
<ol>
<li><strong>Hardware</strong> — GPUs, CPUs, networking, storage</li>
<li><strong>Platform</strong> — Kubernetes, virtualization, GPU scheduling</li>
<li><strong>Serving stack</strong> — model server, batching, autoscaling</li>
<li><strong>Model layer</strong> — open-weight models, fine-tuned variants, embeddings</li>
<li><strong>Application layer</strong> — RAG, agent frameworks, LLM gateway</li>
<li><strong>Operations</strong> — observability, audit, cost management, on-call</li>
</ol>
<p>Skipping any layer — particularly platform and operations — produces a PoC that doesn&rsquo;t survive contact with production.</p>
<h2 id="layer-1-hardware">Layer 1: hardware</h2>
<p>The right GPU depends on workload, model size, and budget. The current production-ready landscape:</p>
<ul>
<li><strong>NVIDIA H100 / H200</strong> — the workhorse for fine-tuning and high-throughput inference of mid-to-large models. Expensive, available with lead time, broadly supported.</li>
<li><strong>NVIDIA Blackwell (B100/B200)</strong> — newer, higher memory bandwidth, well-suited for the largest models. Lead times longer.</li>
<li><strong>NVIDIA L40S</strong> — 48 GB GPU memory; good fit for inference of smaller models or as part of a multi-tenant inference fleet.</li>
<li><strong>NVIDIA A100</strong> — older but still cost-effective for many inference workloads; second-hand market reasonable.</li>
<li><strong>AMD MI300 / MI325</strong> — credible alternative for some workloads; ROCm tooling still maturing relative to CUDA.</li>
<li><strong>Specialized accelerators (Groq, Cerebras, etc.)</strong> — strong for specific use cases; ecosystem narrower.</li>
</ul>
<p>For most regulated-industry deployments, a fleet of H100/H200 (or A100 for cost-sensitive) plus L40S for tenant-fleet inference is the typical answer. NVIDIA data-centre GPUs are supported through the NVIDIA GPU Operator (passthrough to VMs, sharing via HAMi).</p>
<p>The CPU and storage sizing are also important — but largely follow standard practice once GPU sizing is set. Networking matters significantly for multi-GPU training (NVLink, InfiniBand, RoCE) and somewhat less for inference (where 25-100 Gbps Ethernet is usually sufficient).</p>
<h2 id="layer-2-platform">Layer 2: platform</h2>
<p>The platform sits between hardware and applications. For private LLM, the right platform answers:</p>
<ul>
<li>How GPUs are scheduled across workloads — full passthrough, MIG (NVIDIA Multi-Instance GPU), time-slicing, or virtual GPU (vGPU)</li>
<li>How VMs and containers coexist (most data-science workloads run as containers; some legacy or notebook-heavy workloads as VMs)</li>
<li>How multi-tenancy is enforced — tenant isolation, quota, scoped audit</li>
<li>How observability, identity, and storage integrate</li>
</ul>
<p>Kubernetes-native virtualization platforms (Cozystack, OpenShift Virtualization, vendor-led variants) are increasingly the default because they answer all of these in one stack.</p>
<p><a href="https://aenix.io/products/cozystack/">Cozystack</a> supports:</p>
<ul>
<li>Container-based AI workloads with Kubernetes GPU scheduling: whole-GPU allocation through the NVIDIA GPU Operator, fractional sharing (GPU memory and compute cores) through HAMi</li>
<li>VM-based AI workloads through KubeVirt with passthrough of whole GPUs or NVIDIA vGPU (requires your NVIDIA vGPU licence)</li>
<li>Multi-tenant isolation through Tenant CRD with per-tenant GPU quotas</li>
<li>MIG and time-slicing: on the roadmap, not shipped today</li>
</ul>
<p>OpenStack-based or VMware-based legacy platforms can be retrofitted to host AI workloads, but typically with more operational friction and less native Kubernetes integration.</p>
<h2 id="layer-3-serving-stack">Layer 3: serving stack</h2>
<p>The serving stack runs the model. Options:</p>
<ul>
<li><strong>vLLM</strong> — high-throughput serving for transformer models with PagedAttention. Default choice for most inference workloads.</li>
<li><strong>NVIDIA Triton Inference Server</strong> — broader model-format support; strong for mixed workloads (LLM + vision + embedding + classical ML).</li>
<li><strong>TGI (Text Generation Inference)</strong> — Hugging Face&rsquo;s serving stack; some specific feature niches.</li>
<li><strong>llama.cpp / Ollama</strong> — for smaller models, single-machine deployments, or development/PoC.</li>
<li><strong>Custom</strong> — for specialized workloads that need direct framework access.</li>
</ul>
<p>For multi-tenant production: vLLM (or Triton) with autoscaling, batched inference, and a request router (LiteLLM, Portkey, or custom). The choice between vLLM and Triton is mostly a function of the model formats and existing ML pipelines.</p>
<h2 id="layer-4-model-layer">Layer 4: model layer</h2>
<p>Open-weight models in production-ready 2026 landscape:</p>
<ul>
<li><strong>Llama (Meta)</strong> — large family covering 1B to 405B+, licence permits commercial use with caveats</li>
<li><strong>Mistral</strong> — Mixtral (MoE) and Mistral Large; commercial licence for some, Apache 2.0 for older</li>
<li><strong>Qwen (Alibaba)</strong> — strong multilingual, including non-English; Apache 2.0</li>
<li><strong>DeepSeek</strong> — strong reasoning models; licence varies</li>
<li><strong>Phi (Microsoft)</strong> — small models with surprising capability; MIT</li>
<li><strong>Gemma (Google)</strong> — small/medium-size; licence permits commercial use</li>
</ul>
<p>Selection depends on:</p>
<ul>
<li>Language requirement (multilingual vs English-only)</li>
<li>Workload type (chat / RAG / code / vision / embedding)</li>
<li>Cost-per-token target at expected throughput</li>
<li>Licence terms (commercial use, attribution, redistribution)</li>
</ul>
<p>For most regulated-industry deployments: a primary model in the 7B-70B range for general use, plus smaller models (Phi, Gemma) for cost-sensitive paths, plus an embedding model for RAG. Fine-tuning happens for domain-specific accuracy needs.</p>
<h2 id="layer-5-application-layer">Layer 5: application layer</h2>
<p>Most production deployments add:</p>
<ul>
<li><strong>LLM gateway</strong> — request routing, rate limiting, audit logging, cost tracking. Examples: LiteLLM, Portkey, custom.</li>
<li><strong>RAG infrastructure</strong> — vector database (Weaviate, Qdrant, Milvus, pgvector), document indexing pipeline, retrieval orchestration.</li>
<li><strong>Agent framework</strong> (where applicable) — LangChain, LlamaIndex, custom.</li>
<li><strong>Observability for LLM</strong> — token usage, latency p50/p95/p99, error rates, content filters.</li>
</ul>
<p>The application layer is where model-provider switching costs concentrate. A well-engineered LLM gateway makes the underlying model swappable; a poorly-engineered one locks the application to a specific model API.</p>
<h2 id="layer-6-operations">Layer 6: operations</h2>
<p>The post-deployment steady state is where most private LLM projects underperform expectations. The operational surface includes:</p>
<ul>
<li><strong>GPU utilization monitoring</strong> — under-utilized GPUs are wasted budget; over-utilized GPUs cause queueing</li>
<li><strong>Model lifecycle management</strong> — when do you upgrade Llama 3.x to 4.x? When do you re-fine-tune? Who owns the regression-test suite?</li>
<li><strong>Cost tracking by tenant / team / workload</strong> — without per-tenant tracking, cost allocation becomes unmaintainable</li>
<li><strong>Audit trail</strong> — for regulator dialog, every inference request must trace to: model + version, weights, input, output, requesting user, timestamp</li>
<li><strong>On-call</strong> — GPU failures, OOM events, queue backups</li>
<li><strong>Capacity planning</strong> — when do you buy more GPUs? When do you retire older ones?</li>
</ul>
<p>A team that under-staffs the operations layer ends up with a stack that &ldquo;works&rdquo; but isn&rsquo;t production-grade for regulator audit or business-continuity standards.</p>
<h2 id="common-deployment-patterns">Common deployment patterns</h2>
<h3 id="pattern-1-single-tenant-inference-cluster">Pattern 1: single-tenant inference cluster</h3>
<p>Smallest reasonable deployment. Single inference workload, single team, single model family. Typical hardware: 4-16 H100/H200 or 8-32 L40S. Platform: bare Kubernetes or Cozystack. Serving: vLLM. Best for: regulated workloads with one application owner; PoCs.</p>
<h3 id="pattern-2-multi-tenant-inference-fleet">Pattern 2: multi-tenant inference fleet</h3>
<p>Multi-team or multi-application deployment. Multiple model families, per-tenant quotas, shared observability. Typical hardware: 16-128 GPUs across multiple model classes. Platform: Cozystack with Tenant CRD or OpenShift with namespaces + quotas. Serving: vLLM or Triton with autoscaling and a request router. Best for: enterprise platform teams supporting multiple data-science teams.</p>
<h3 id="pattern-3-inference--fine-tuning--rag">Pattern 3: inference + fine-tuning + RAG</h3>
<p>Full AI platform. Inference fleet, dedicated fine-tuning capacity, RAG infrastructure with vector database, embedding model, LLM gateway. Typical hardware: 32-512 GPUs across roles. Platform: Cozystack or full Kubernetes with operators. Best for: financial services, healthcare, public sector with sustained AI program.</p>
<h3 id="pattern-4-air-gapped-sovereign-deployment">Pattern 4: air-gapped sovereign deployment</h3>
<p>Air-gapped or restricted-egress deployment for the most sensitive workloads. No internet egress; updates through controlled channels (Harbor mirror, internal package registry, manually distributed model weights). Typical hardware: customer-supplied. Platform: Cozystack with air-gap install. Best for: classified, defence-adjacent, healthcare-with-strict-residency.</p>
<h2 id="common-pitfalls">Common pitfalls</h2>
<h3 id="pitfall-1-skipping-the-platform-layer">Pitfall 1: skipping the platform layer</h3>
<p>Teams deploy vLLM on a couple of bare-metal boxes, call it a private LLM deployment. Works for a PoC. Falls over the first time GPU demand exceeds available capacity, or the first time multi-tenant isolation matters, or the first time the regulator asks for an audit trail.</p>
<h3 id="pitfall-2-model-api-as-private-llm">Pitfall 2: model-API-as-private-LLM</h3>
<p>Some vendors market a SaaS endpoint with privacy controls as &ldquo;private LLM.&rdquo; The data still leaves the customer&rsquo;s perimeter, even if the privacy clause is strong. For trigger-1 (regulated data) workloads, this fails the substantive requirement.</p>
<h3 id="pitfall-3-under-sizing-memory">Pitfall 3: under-sizing memory</h3>
<p>The most common mistake in initial sizing: not accounting for KV cache memory that scales with context length and batch size. A model that &ldquo;fits&rdquo; by parameter count may not fit at the operational batch size. Right-sizing requires actual benchmark with realistic context lengths.</p>
<h3 id="pitfall-4-ignoring-the-cost-allocation-problem">Pitfall 4: ignoring the cost-allocation problem</h3>
<p>Without per-tenant cost tracking, the platform&rsquo;s economics become opaque to finance. Teams over-request GPU; finance pushes back; the platform team has to retrofit cost tracking after the fact. Build it from day one.</p>
<h2 id="when-not-to-deploy-private-llm">When NOT to deploy private LLM</h2>
<p>Private LLM is the wrong answer when:</p>
<ul>
<li>Your AI workload is small (under ~10K requests/day) and not regulated. The economics of running dedicated GPU fleet don&rsquo;t beat hyperscaler on this scale.</li>
<li>Your team has no Kubernetes / platform-engineering capacity. A private LLM platform is an operational commitment.</li>
<li>Your model needs are at the absolute frontier (current GPT-class capability with all bells and whistles). The largest open-weight models close the gap rapidly, but the very latest frontier capability still trails.</li>
<li>Your data is not actually regulated, your spend is not actually growing, and the trigger is more &ldquo;we want our own thing&rdquo; than a substantive driver. The work is real and substantial; the trigger has to be real too.</li>
</ul>
<p>A good engagement is honest about these cases. The Ænix engagement specifically does not push private LLM when the alternative is the right answer.</p>
<h2 id="want-to-dig-deeper">Want to dig deeper?</h2>
<ul>
<li><strong><a href="https://aenix.io/solutions/sovereign-ai/">Sovereign AI services page</a></strong> — engagement details</li>
<li><strong><a href="https://aenix.io/solutions/data-sovereignty/">Data sovereignty</a></strong> — regulator-side trigger</li>
<li><strong><a href="https://aenix.io/solutions/dora-compliance/">DORA compliance</a></strong> — financial-services trigger</li>
<li><strong><a href="https://aenix.io/products/cozystack/">Cozystack</a></strong> — the platform we run AI workloads on</li>
</ul>
]]></content:encoded></item><item><title>Private cloud providers and platforms — a 2026 comparison</title><link>https://aenix.io/blog/2026/05/private-cloud-providers-comparison/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/private-cloud-providers-comparison/</guid><pubDate>Fri, 22 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>VMware</category><category>OpenStack</category><category>Proxmox</category><category>OpenShift</category><category>Kubernetes</category><category>Cozystack</category><description>Open-source platforms, commercial stacks, sovereign hyperscaler regions and regional providers compared — plus the migration paths between them.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/private-cloud-providers-comparison.jpg" alt=""></p><p>The private cloud landscape has shifted significantly in the last 3 years. Broadcom-induced VMware migrations, sovereignty mandates, AI workload economics, and FinOps pressure have all reshaped what &ldquo;private cloud&rdquo; means and what providers serve it.</p>
<h2 id="two-distinct-things-called-private-cloud">Two distinct things called &ldquo;private cloud&rdquo;</h2>
<p>The terminology is overloaded. &ldquo;Private cloud&rdquo; means either:</p>
<ul>
<li><strong>Private cloud platform</strong> — software you deploy on infrastructure you control. Examples: VMware VCF, Cozystack, OpenStack, OpenShift Virtualization, Proxmox VE.</li>
<li><strong>Private cloud provider</strong> — a vendor that provides dedicated infrastructure (single-tenant) which you consume. Examples: IBM Cloud Private, Oracle dedicated regions, hyperscaler &ldquo;sovereign&rdquo; regions, regional cloud providers (OVHcloud, Hetzner, or a regional provider running Ænix Public Cloud Platform).</li>
</ul>
<p>Both are valid; they answer different questions. This article focuses primarily on platforms (the software layer); providers come up where relevant.</p>
<h2 id="open-source-platforms">Open-source platforms</h2>
<h3 id="cozystack">Cozystack</h3>
<p><strong>Licence:</strong> Apache 2.0, CNCF Project.
<strong>Architecture:</strong> Kubernetes-native virtualization (KubeVirt) + Cilium and Kube-OVN networking + LINSTOR (DRBD) block storage + SeaweedFS object storage + Tenant CRD multi-tenancy + Cozystack Dashboard self-service.
<strong>Maintainer:</strong> Ænix (open-source, community-governed).
<strong>Best for:</strong> Service providers, sovereign-cloud builders, regulated multi-tenant, AI/GPU operators with sustained workloads.
<strong>Strengths:</strong> Single platform for VMs + containers + databases + S3 + GPU. Multi-tenancy structural. Light operational footprint relative to OpenStack. Open-source, no vendor lock-in.
<strong>Limits:</strong> Newer than OpenStack; smaller community.</p>
<h3 id="openstack">OpenStack</h3>
<p><strong>Licence:</strong> Apache 2.0, OpenInfra Foundation.
<strong>Architecture:</strong> Nova compute + Neutron network + Cinder block + Swift object + Keystone identity + Horizon UI + many other components.
<strong>Maintainer:</strong> OpenInfra Foundation; commercial distros from Red Hat, Canonical, Mirantis.
<strong>Best for:</strong> Large telecom operators, government clouds, OpenStack-fluent teams.
<strong>Strengths:</strong> Mature, broad community, many vendor options.
<strong>Limits:</strong> Operationally complex; harder to find OpenStack engineers in 2026; less Kubernetes-native.</p>
<h3 id="openshift-virtualization-red-hat">OpenShift Virtualization (Red Hat)</h3>
<p><strong>Licence:</strong> Red Hat commercial subscription.
<strong>Architecture:</strong> OpenShift Kubernetes + KubeVirt + Red Hat ecosystem.
<strong>Maintainer:</strong> Red Hat / IBM.
<strong>Best for:</strong> Existing Red Hat customers, enterprises with Red Hat procurement.
<strong>Strengths:</strong> Strong commercial support, mature.
<strong>Limits:</strong> Subscription pricing; tied to Red Hat / IBM relationship.</p>
<h3 id="proxmox-ve">Proxmox VE</h3>
<p><strong>Licence:</strong> AGPLv3 + commercial subscription.
<strong>Architecture:</strong> KVM + LXC + ZFS + Ceph (community).
<strong>Maintainer:</strong> Proxmox Server Solutions GmbH.
<strong>Best for:</strong> SMB virtualization, single-tenant, labs.
<strong>Strengths:</strong> Mature, easy to install, strong community.
<strong>Limits:</strong> Limited multi-tenancy; service catalog beyond VMs requires manual integration.</p>
<h3 id="apache-cloudstack">Apache CloudStack</h3>
<p><strong>Licence:</strong> Apache 2.0.
<strong>Architecture:</strong> Hypervisor-agnostic (XenServer / KVM / VMware), service-provider-oriented.
<strong>Best for:</strong> Service providers in markets where CloudStack remains established (some EU, MENA, APAC).
<strong>Strengths:</strong> Service-provider features mature; multi-tenancy native.
<strong>Limits:</strong> Smaller community than alternatives; less Kubernetes-native.</p>
<h2 id="commercial--closed-source-platforms">Commercial / closed-source platforms</h2>
<h3 id="vmware-vmware-cloud-foundation">VMware (VMware Cloud Foundation)</h3>
<p><strong>Licence:</strong> Subscription-only post-Broadcom.
<strong>Architecture:</strong> vSphere + vSAN + NSX + vCD + vRealize.
<strong>Best for:</strong> Existing VMware estates that haven&rsquo;t yet been triggered out by economics.
<strong>Strengths:</strong> Mature, well-known, extensive ecosystem.
<strong>Limits:</strong> Subscription pricing increases (2-5× observed); vendor lock-in; sovereignty concerns.</p>
<h3 id="nutanix">Nutanix</h3>
<p><strong>Licence:</strong> Subscription, multiple tiers.
<strong>Architecture:</strong> AHV (proprietary KVM-based) + Files + Volumes + Era (databases).
<strong>Best for:</strong> Existing Nutanix HCI customers, enterprises preferring appliance model.
<strong>Strengths:</strong> Operationally simple, integrated stack.
<strong>Limits:</strong> Closed source; appliance lock-in; less flexible than open alternatives.</p>
<h3 id="scale-computing-hc3">Scale Computing HC3</h3>
<p><strong>Licence:</strong> Subscription.
<strong>Architecture:</strong> KVM-based hyperconverged appliance.
<strong>Best for:</strong> ROBO / edge / SMB.
<strong>Strengths:</strong> Operationally simple.
<strong>Limits:</strong> Smaller scale ceiling; appliance lock-in.</p>
<h3 id="microsoft-azure-stack-hci">Microsoft Azure Stack HCI</h3>
<p><strong>Licence:</strong> Microsoft subscription + per-core fee.
<strong>Architecture:</strong> Hyper-V + Storage Spaces Direct + Azure Arc.
<strong>Best for:</strong> Microsoft-aligned shops with Azure relationship.
<strong>Strengths:</strong> Strong Microsoft ecosystem integration.
<strong>Limits:</strong> Locks into Microsoft licensing economics.</p>
<h3 id="oracle-cloud-native-environment--oracle-linux-virtualization-manager">Oracle Cloud Native Environment / Oracle Linux Virtualization Manager</h3>
<p><strong>Licence:</strong> Subscription / commercial.
<strong>Best for:</strong> Oracle-aligned organizations.</p>
<h2 id="sovereign-hyperscaler-regions">Sovereign hyperscaler regions</h2>
<h3 id="aws-sovereign-cloud-eu--us-gov">AWS Sovereign Cloud (EU / US Gov)</h3>
<p>Dedicated regions with sovereignty controls. Some satisfy member-state mandates; others don&rsquo;t, depending on jurisdiction.</p>
<h3 id="azure-sovereign-cloud-azure-government-azure-germany-historically">Azure Sovereign Cloud (Azure Government, Azure Germany historically)</h3>
<p>Similar pattern.</p>
<h3 id="gcp-sovereign-cloud">GCP Sovereign Cloud</h3>
<p>GCP&rsquo;s sovereign offerings, Workspace partnerships in some EU markets.</p>
<p><strong>Trade-off:</strong> these provide cloud-managed convenience but leave the service plane under hyperscaler control. For substantive sovereignty (encryption keys customer-controlled, supplier transparency, exit-readiness), customer-owned infrastructure typically wins.</p>
<h2 id="regional-cloud-providers-private-cloud-as-a-service">Regional cloud providers (private cloud as-a-service)</h2>
<p>A growing market in 2026:</p>
<ul>
<li><strong>Hetzner</strong> (Germany) — bare metal + cloud, popular in DACH</li>
<li><strong>OVHcloud</strong> (France) — strong EU sovereign positioning</li>
<li><strong>Regional providers running Ænix Public Cloud Platform</strong> — hosting providers that sell a sovereign cloud product built on Cozystack</li>
<li>Various regional providers per jurisdiction</li>
</ul>
<p>These offer private-cloud-style isolation without you operating the platform. Trade-off: provider relationship vs. direct hardware control.</p>
<h2 id="how-to-choose">How to choose</h2>
<p>Decision tree:</p>
<ol>
<li><strong>Need multi-tenant + open-source + Kubernetes-native + sovereignty?</strong> → Cozystack.</li>
<li><strong>Existing VMware estate, financial-services renewal pressure?</strong> → Plan VMware exit. Destination: typically Cozystack or OpenShift.</li>
<li><strong>OpenStack expertise + large telco / government scale?</strong> → OpenStack remains valid.</li>
<li><strong>Existing Red Hat / OpenShift commitments?</strong> → OpenShift Virtualization.</li>
<li><strong>SMB / single-tenant?</strong> → Proxmox VE.</li>
<li><strong>Don&rsquo;t want to operate the platform yourself?</strong> → Regional sovereign cloud provider (Hetzner, OVHcloud, or a regional provider running Ænix Public Cloud Platform).</li>
<li><strong>AI/GPU at scale, sustained utilization?</strong> → Cozystack or OpenShift on dedicated GPU infrastructure.</li>
<li><strong>Sovereignty + EU + low operational footprint?</strong> → Cozystack with Ænix support, or OVHcloud.</li>
</ol>
<h2 id="migration-paths">Migration paths</h2>
<p>Most modern private-cloud deployments are migrations from existing infrastructure:</p>
<ul>
<li><strong>VMware → Cozystack/OpenStack/OpenShift</strong> — KubeVirt-based migration, image conversion, multi-tenancy redesign</li>
<li><strong>Public cloud → private cloud</strong> (repatriation) — workload classification, cost honesty, destination architecture; see <strong><a href="https://aenix.io/solutions/cloud-repatriation/">Cloud repatriation</a></strong></li>
<li><strong>OpenStack → Cozystack</strong> — for teams seeking Kubernetes-native foundation; image migration is straightforward</li>
<li><strong>Hyperscaler region → sovereign region</strong> — for sovereignty-driven migrations within hyperscaler model</li>
</ul>
<h2 id="want-to-dig-deeper">Want to dig deeper?</h2>
<ul>
<li><strong><a href="https://aenix.io/products/private-cloud-platform/">Private cloud platform — Cozystack</a></strong></li>
<li><strong><a href="https://aenix.io/services/private-cloud-consulting/">Private cloud consulting</a></strong> — engineering services</li>
<li><strong><a href="https://aenix.io/alternatives/vmware-alternative/">VMware alternative</a></strong> — VMware exit</li>
<li><strong><a href="https://aenix.io/solutions/cloud-repatriation/">Cloud repatriation</a></strong> — public cloud exit</li>
<li><strong><a href="https://cozystack.io">cozystack.io</a></strong> — open-source project</li>
</ul>
]]></content:encoded></item><item><title>Private cloud architecture in 2026 — design, components, and implementation patterns</title><link>https://aenix.io/blog/2026/05/private-cloud-architecture-2026/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/private-cloud-architecture-2026/</guid><pubDate>Fri, 22 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>OpenStack</category><category>Kubernetes</category><category>KubeVirt</category><category>Sovereignty</category><category>Multi-tenancy</category><category>Financial Services</category><description>What private cloud means in 2026: the architectural layers, three patterns that work, capacity sizing, and the mistakes that recur in design reviews.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/private-cloud-architecture-2026.jpg" alt=""></p><p>Private cloud has moved from &ldquo;yesterday&rsquo;s architecture&rdquo; to &ldquo;tomorrow&rsquo;s default for regulated and cost-sensitive workloads&rdquo; within ~3 years. The Broadcom Private Cloud Outlook 2025 found 53% of organizations now prioritize private cloud for new workloads. The LSEG Global Cloud Survey reports 84% of financial-services firms have adjusted cloud strategy due to regulatory pressure. The shift is real.</p>
<p>But the architecture decisions in 2026 are different from 2018-era OpenStack-centric private clouds. The default stack has moved to Kubernetes-native virtualization (KubeVirt) plus open-source storage and networking. The operational model has matured. The trade-offs are clearer.</p>
<h2 id="what-private-cloud-means-in-2026">What &ldquo;private cloud&rdquo; means in 2026</h2>
<p>A private cloud is dedicated cloud infrastructure run for a single organization or tenant, with self-service consumption patterns matching public cloud (provision-on-demand, multi-tenancy, observability, automation) — but on infrastructure the organization controls.</p>
<p>In 2026, private cloud is one of:</p>
<ul>
<li><strong>Customer-owned, customer-operated</strong> — organization owns hardware, runs the platform.</li>
<li><strong>Customer-owned, vendor-operated</strong> — organization owns hardware, vendor runs the platform under contract.</li>
<li><strong>Vendor-owned dedicated</strong> — vendor provides dedicated infrastructure (single-tenant), organization consumes.</li>
<li><strong>Sovereign hyperscaler region</strong> — hyperscaler-operated infrastructure with sovereignty controls (some count, some don&rsquo;t).</li>
</ul>
<p>This article focuses on customer-owned options 1 and 2 — the architecture is similar; the operations differ.</p>
<h2 id="architectural-layers">Architectural layers</h2>
<p>A modern private cloud has six functional layers:</p>
<h3 id="layer-1-hardware">Layer 1: hardware</h3>
<ul>
<li><strong>Compute</strong> — x86 or ARM servers; modern generations support all relevant workloads. AI/GPU adds a separate hardware tier (NVIDIA H100/H200/L40S/Blackwell, AMD MI-series).</li>
<li><strong>Storage</strong> — replicated block storage (LINSTOR, Ceph, or vendor SAN). Object storage for backup and applications.</li>
<li><strong>Network</strong> — datacenter fabric (BGP-routed leaf-spine increasingly default), 25-100 Gbps Ethernet sufficient for most non-HPC workloads.</li>
<li><strong>Datacenter</strong> — own facility, colocation, or ROBO/edge. Power, cooling, physical security.</li>
</ul>
<h3 id="layer-2-os-and-platform-foundation">Layer 2: OS and platform foundation</h3>
<ul>
<li><strong>OS</strong> — Linux, increasingly minimal (Talos, Bottlerocket, Flatcar) for Kubernetes hosts. RHEL / Ubuntu LTS for VM hypervisors when not in Kubernetes.</li>
<li><strong>Kubernetes</strong> — vanilla, OpenShift, Cozystack, vendor-led. Distribution choice is structural.</li>
<li><strong>Virtualization</strong> — KubeVirt for VM workloads inside Kubernetes (most modern); KVM/libvirt directly (OpenStack); VMware (legacy).</li>
</ul>
<h3 id="layer-3-storage-and-networking">Layer 3: storage and networking</h3>
<ul>
<li><strong>Block storage</strong> — LINSTOR (DRBD-based, default in Cozystack), Ceph (Rook-managed), vendor SAN.</li>
<li><strong>Object storage</strong> — Ceph RGW, MinIO, SeaweedFS.</li>
<li><strong>Network virtualization</strong> — Cilium (eBPF, default in 2026), Calico, OVN.</li>
<li><strong>Load balancing</strong> — MetalLB for layer 2/3, Cilium L7, ingress controllers (NGINX / Traefik / Contour).</li>
</ul>
<h3 id="layer-4-control-plane">Layer 4: control plane</h3>
<ul>
<li><strong>Multi-tenancy</strong> — Tenant CRD (Cozystack), namespace-based (vanilla Kubernetes), vCloud-Director-style (legacy).</li>
<li><strong>Identity</strong> — Keycloak, Okta integration, AD integration. SPIFFE/SPIRE for service identity.</li>
<li><strong>GitOps</strong> — Argo CD or Flux. Cozystack uses Flux.</li>
<li><strong>Observability</strong> — VictoriaMetrics + VictoriaLogs (Cozystack default), Prometheus + Loki, vendor SaaS.</li>
<li><strong>Backup/DR</strong> — Velero + S3, per-app patterns (PostgreSQL PITR).</li>
</ul>
<h3 id="layer-5-application-and-platform-services">Layer 5: application and platform services</h3>
<ul>
<li><strong>Managed databases</strong> — PostgreSQL (CloudNativePG), MariaDB, MongoDB, Redis, Valkey, Kafka, ClickHouse, RabbitMQ, NATS, OpenSearch.</li>
<li><strong>Object storage as a service</strong> — S3-compatible.</li>
<li><strong>AI/ML platform</strong> — KubeVirt for VM-based GPU, Kubernetes-native for container-based GPU, vLLM/Triton for inference.</li>
<li><strong>Self-service portal</strong> — Backstage, Cozystack Dashboard, custom.</li>
</ul>
<h3 id="layer-6-operations">Layer 6: operations</h3>
<ul>
<li><strong>Runbooks</strong> — documented operational procedures.</li>
<li><strong>On-call rotation</strong> — incident response model.</li>
<li><strong>Capacity planning</strong> — quarterly hardware refresh, growth forecasting.</li>
<li><strong>Compliance posture</strong> — audit logging, certifications (where applicable), regulator dialog.</li>
</ul>
<h2 id="three-architectural-patterns-that-work">Three architectural patterns that work</h2>
<h3 id="pattern-1-cozystack-based-kubernetes-native-cloud">Pattern 1: Cozystack-based Kubernetes-native cloud</h3>
<p><strong>What:</strong> Single Kubernetes cluster (or fleet) with Cozystack as the platform layer. KubeVirt for VMs, Kubernetes for containers, LINSTOR for storage, Cilium for networking, Tenant CRD for multi-tenancy, Cozystack Dashboard for self-service.</p>
<p><strong>Best for:</strong> Service providers, sovereign-cloud builders, regulated multi-tenant. Greenfield deployments.</p>
<p><strong>Pros:</strong> Open-source, single-stack, multi-tenancy native, virtualization + container in one platform.</p>
<p><strong>Cons:</strong> Newer than OpenStack; smaller community than Kubernetes-only deployments.</p>
<h3 id="pattern-2-openstack-based-traditional-private-cloud">Pattern 2: OpenStack-based traditional private cloud</h3>
<p><strong>What:</strong> OpenStack as compute orchestrator, Ceph for storage, OVN for networking, Heat/Terraform for provisioning. Optional Kubernetes-on-OpenStack for container workloads.</p>
<p><strong>Best for:</strong> Organizations with deep OpenStack experience; large telco/sovereign deployments where OpenStack is the procurement standard.</p>
<p><strong>Pros:</strong> Mature, broad community, many vendor options.</p>
<p><strong>Cons:</strong> Operationally complex; harder to find OpenStack engineers in 2026; less Kubernetes-native.</p>
<h3 id="pattern-3-vmware-cloud-foundation-vcf--legacy">Pattern 3: VMware Cloud Foundation (VCF) — legacy</h3>
<p><strong>What:</strong> vSphere + vSAN + NSX + vCD + Aria (formerly vRealize). Closed source, subscription-licensed.</p>
<p><strong>Best for:</strong> Existing VMware estates that haven&rsquo;t yet been triggered out by Broadcom economics.</p>
<p><strong>Pros:</strong> Mature, well-known, extensive integration.</p>
<p><strong>Cons:</strong> Subscription-only post-Broadcom, 2-5× price increases observed, lock-in to single vendor&rsquo;s roadmap.</p>
<p>(For VMware exit guidance see <strong><a href="https://aenix.io/alternatives/vmware-alternative/">VMware alternative</a></strong>.)</p>
<h2 id="architectural-decisions-that-matter-most">Architectural decisions that matter most</h2>
<p>Five decisions with the highest long-term impact:</p>
<h3 id="decision-1-virtualization-layer">Decision 1: virtualization layer</h3>
<p>KubeVirt or traditional hypervisor? KubeVirt is the 2026 default for greenfield; traditional hypervisor (KVM/libvirt directly) is appropriate for very-large-scale OpenStack deployments where KubeVirt&rsquo;s overhead matters.</p>
<h3 id="decision-2-storage-architecture">Decision 2: storage architecture</h3>
<p>LINSTOR/DRBD for replicated block? LINSTOR is operationally simpler and Cozystack-default. Ceph is more flexible but operationally heavier. Vendor SAN is an option but contracts limit your scaling pattern.</p>
<h3 id="decision-3-networking">Decision 3: networking</h3>
<p>Cilium has become the default CNI for new deployments. NSX-equivalent (Cilium L7, service mesh) replaces VMware NSX functionality. The decision is more about your team&rsquo;s eBPF familiarity than about technical fit.</p>
<h3 id="decision-4-multi-tenancy-model">Decision 4: multi-tenancy model</h3>
<p>For service-provider model: Tenant CRD (Cozystack) is the default. For internal multi-BU: namespace-based + RBAC is sufficient. For absolute isolation: cluster-per-tenant (operationally expensive).</p>
<h3 id="decision-5-operational-model">Decision 5: operational model</h3>
<p>Customer-operated (you run it), vendor-operated (Ænix or similar runs it for you), or hybrid (you operate; vendor provides 2nd-line). Decision driven by internal team capacity and risk appetite.</p>
<h2 id="capacity-sizing">Capacity sizing</h2>
<p>A practical sizing rubric for a 100-VM-equivalent workload:</p>
<ul>
<li><strong>Compute:</strong> 6-10 servers, dual-socket, 256-512 GB RAM each. Reserve 30% headroom for failures and growth.</li>
<li><strong>Storage:</strong> ~100 TB replicated (3-replica) for steady-state plus snapshot/backup capacity. ~300 TB raw for that target.</li>
<li><strong>Network:</strong> 25 Gbps NIC per server, leaf-spine fabric, 100 Gbps backbone.</li>
<li><strong>Datacenter:</strong> ~6-10 kW per rack at modern density.</li>
<li><strong>DR site:</strong> comparable secondary footprint.</li>
</ul>
<p>Scaling beyond 1000 VMs: hardware grows roughly linearly; platform team grows sublinearly with good automation.</p>
<h2 id="operational-practices-that-matter">Operational practices that matter</h2>
<ul>
<li><strong>Quarterly capacity reviews</strong> — hardware ahead of growth.</li>
<li><strong>Twice-yearly Kubernetes upgrades</strong> — stay current on CVEs and feature parity.</li>
<li><strong>Documented incident response</strong> — incident commander, scribe, blameless post-mortems.</li>
<li><strong>Compliance posture</strong> — audit logs, certifications, regulator dialog where applicable.</li>
<li><strong>Platform-team headcount sized realistically</strong> — 1-3 engineers for a single-cluster small private cloud; 5-15 for a multi-cluster service-provider.</li>
</ul>
<h2 id="common-architectural-mistakes">Common architectural mistakes</h2>
<h3 id="mistake-1-copying-public-cloud-architecture">Mistake 1: copying public-cloud architecture</h3>
<p>Designing the private cloud to look like AWS internally. Different scale economics; public-cloud architecture patterns (massive distributed eventual-consistency systems) are over-engineering for most private clouds.</p>
<h3 id="mistake-2-vendor-led-private-cloud">Mistake 2: vendor-led private cloud</h3>
<p>Buying &ldquo;private cloud appliance&rdquo; rebuilds lock-in with new vendor. The vendor&rsquo;s roadmap becomes yours.</p>
<h3 id="mistake-3-under-engineering-operations">Mistake 3: under-engineering operations</h3>
<p>Compute and storage solved; observability/identity/backup-DR/runbook layer underinvested. Operational debt builds.</p>
<h3 id="mistake-4-skipping-multi-tenancy-at-design">Mistake 4: skipping multi-tenancy at design</h3>
<p>Single-tenant cluster scaled to multi-tenant later. Bolted-on multi-tenancy fails at scale or under regulator audit.</p>
<h3 id="mistake-5-hardware-refresh-skipped-in-budget">Mistake 5: hardware refresh skipped in budget</h3>
<p>Year 4 hardware refresh ignored in initial economics. The refresh cliff arrives unexpectedly.</p>
<h2 id="want-to-dig-deeper">Want to dig deeper?</h2>
<ul>
<li><strong><a href="https://aenix.io/services/private-cloud-consulting/">Private cloud consulting services</a></strong> — engagement details</li>
<li><strong><a href="https://aenix.io/solutions/data-sovereignty/">Data sovereignty</a></strong> — sovereignty trigger</li>
<li><strong><a href="https://aenix.io/solutions/cloud-repatriation/">Cloud repatriation</a></strong> — when coming from public cloud</li>
<li><strong><a href="https://aenix.io/alternatives/vmware-alternative/">VMware alternative</a></strong> — when coming from VMware</li>
<li><strong><a href="https://aenix.io/products/cozystack/">Cozystack</a></strong> — open-source platform foundation</li>
</ul>
]]></content:encoded></item><item><title>Platform engineering vs DevOps vs SRE — a 2026 terminology guide</title><link>https://aenix.io/blog/2026/05/platform-engineering-vs-devops-vs-sre/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/platform-engineering-vs-devops-vs-sre/</guid><pubDate>Thu, 21 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>DevOps</category><category>Platform Engineering</category><category>Observability</category><description>Where platform engineering, DevOps and SRE overlap and where they do not, what each actually builds, and the metrics that separate them.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/platform-engineering-vs-devops-vs-sre.jpg" alt=""></p><p>The terms platform engineering, DevOps, and SRE have been used interchangeably, in opposition, and as overlapping practices for nearly a decade. By 2026 the industry has roughly converged — but only roughly. Different companies still use the same words for different jobs, and the resulting org-design conversations stall because nobody quite agrees on what they&rsquo;re discussing.</p>
<p>This is the working version we use at Ænix when we engage with a customer&rsquo;s engineering organization.</p>
<h2 id="the-three-terms-in-one-paragraph-each">The three terms in one paragraph each</h2>
<p><strong>DevOps</strong> is a cultural and operational practice within product teams. Its premise: the same team that builds the software is responsible for operating it in production. Tooling supports this — CI/CD pipelines, IaC, observability — but tooling alone is not DevOps. The function lives inside product teams, not as a separate department.</p>
<p><strong>SRE (Site Reliability Engineering)</strong> is a discipline that applies software engineering to operations. The original Google formulation: SREs are software engineers tasked with keeping production reliable, with explicit error budgets, SLOs, and a rule that operational toil cannot exceed 50% of their time. SRE can sit inside product teams (embedded SRE) or as a separate function (centralized SRE).</p>
<p><strong>Platform engineering</strong> is the practice of building and operating internal platforms that product teams consume. Its premise: don&rsquo;t expect every product team to figure out infrastructure, observability, security, identity, and release engineering on their own. Build a self-service platform that handles those concerns, and let product teams ship. Platform engineering lives as a separate function — a team whose customers are other engineering teams.</p>
<h2 id="where-they-overlap">Where they overlap</h2>
<p>All three care about:</p>
<ul>
<li><strong>Reliability</strong> — SLOs, error budgets, incident response.</li>
<li><strong>Tooling</strong> — CI/CD, IaC, observability, secret management, identity.</li>
<li><strong>Velocity</strong> — making it faster for product teams to ship.</li>
<li><strong>Operational excellence</strong> — automation, runbooks, post-mortems.</li>
</ul>
<p>Where they diverge: <strong>who owns what</strong>, and <strong>whose problem is it when the abstraction breaks</strong>.</p>
<h2 id="where-they-dont-overlap">Where they don&rsquo;t overlap</h2>
<table>
  <thead>
      <tr>
          <th></th>
          <th>DevOps</th>
          <th>SRE</th>
          <th>Platform Engineering</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><strong>Who</strong></td>
          <td>Product team</td>
          <td>Reliability function (centralized or embedded)</td>
          <td>Platform team</td>
      </tr>
      <tr>
          <td><strong>Customer</strong></td>
          <td>The team itself</td>
          <td>Product teams that operate the service</td>
          <td>Other engineering teams</td>
      </tr>
      <tr>
          <td><strong>Output</strong></td>
          <td>Software in production with operational ownership</td>
          <td>SLO compliance, reliable production, error budgets</td>
          <td>Self-service paths product teams consume</td>
      </tr>
      <tr>
          <td><strong>Measured by</strong></td>
          <td>Software shipped + operational quality</td>
          <td>SLO/error budget metrics, incident rate, MTTR</td>
          <td>Time-to-environment, golden-path adoption, internal-NPS</td>
      </tr>
      <tr>
          <td><strong>Centralized?</strong></td>
          <td>No — distributed in product teams</td>
          <td>Sometimes (Google model) sometimes (embedded)</td>
          <td>Yes — separate platform team</td>
      </tr>
      <tr>
          <td><strong>Scale fit</strong></td>
          <td>All sizes</td>
          <td>Mid-to-large with dedicated reliability function</td>
          <td>Mid-to-large with multiple product teams</td>
      </tr>
      <tr>
          <td><strong>Tooling shared with</strong></td>
          <td>SRE, platform engineering</td>
          <td>DevOps, platform engineering</td>
          <td>DevOps, SRE — and provided to them</td>
      </tr>
  </tbody>
</table>
<p>The cleanest way to remember: DevOps is a <em>practice</em>, SRE is a <em>discipline</em>, platform engineering is a <em>function</em>. They can coexist; they answer different questions; they&rsquo;re not in opposition.</p>
<h2 id="when-each-is-appropriate">When each is appropriate</h2>
<h3 id="devops-without-separate-sre-or-platform-engineering">DevOps without separate SRE or platform engineering</h3>
<p>Small to mid-size organizations with 1-3 product teams. Each team operates its own services, with shared tooling but no dedicated reliability or platform function. The model works because each team can fully understand its own operational surface.</p>
<h3 id="devops--sre-without-platform-engineering">DevOps + SRE without platform engineering</h3>
<p>Mid-size organizations with reliability concerns that justify a dedicated function. Common in mid-stage startups where production reliability is becoming critical but the engineering organization is still small enough that infrastructure decisions don&rsquo;t fragment.</p>
<h3 id="devops--platform-engineering-without-separate-sre">DevOps + platform engineering without separate SRE</h3>
<p>Mid-size to large organizations where the leverage is in making infrastructure self-service rather than in keeping individual services reliable. SRE-style practices (SLOs, error budgets) get baked into the platform; reliability concerns are addressed at platform level rather than per-service.</p>
<h3 id="all-three">All three</h3>
<p>Large organizations with both reliability requirements and self-service platform needs. SREs work on reliability of critical services; platform engineering provides the substrate; product teams practice DevOps within the constraints both establish.</p>
<p>The mistake to avoid: hiring a &ldquo;DevOps engineer&rdquo; who actually does platform engineering work, or hiring an &ldquo;SRE&rdquo; who actually does general infrastructure. The titles increasingly mean specific things.</p>
<h2 id="what-platform-engineering-actually-builds">What platform engineering actually builds</h2>
<p>For organizations that need a platform engineering function, the question becomes: what does the platform team actually produce? Concretely:</p>
<h3 id="1-internal-developer-platform-idp">1. Internal developer platform (IDP)</h3>
<p>A self-service surface that product teams use without filing tickets. Common scope:</p>
<ul>
<li>Environment provisioning (dev / staging / prod)</li>
<li>Application deployment (Helm / Argo CD / Flux)</li>
<li>Observability onboarding (metrics / logs / traces)</li>
<li>Secrets management</li>
<li>Identity (workforce identity → service identity)</li>
<li>Network connectivity (mesh / ingress / load balancing)</li>
</ul>
<p>The IDP is consumed via a UI (often Backstage or similar), via APIs, or via IaC. The choice depends on the team&rsquo;s preferences — but the underlying capabilities must be self-service regardless of UI.</p>
<h3 id="2-golden-paths">2. Golden paths</h3>
<p>Opinionated, well-supported paths for common product-team needs. &ldquo;Golden&rdquo; because they&rsquo;re the recommended way; teams can deviate but pay for it in support effort. Examples:</p>
<ul>
<li>Standard service template (HTTP API, batch job, scheduled job)</li>
<li>Standard deployment pattern (canary, blue-green)</li>
<li>Standard observability stack (auto-instrumented)</li>
<li>Standard data-access patterns (managed databases via the platform)</li>
</ul>
<h3 id="3-operational-model">3. Operational model</h3>
<p>The platform itself is operated by the platform team. SLOs, on-call, runbooks, incident response — all platform-team responsibilities. Product teams should be able to assume the platform works, not contribute to fixing it.</p>
<h3 id="4-internal-product-management">4. Internal product management</h3>
<p>Platform engineering as a function has product-team customers. That requires product-management practices: roadmap, prioritization, customer interviews (your engineering teams), feedback loops, deprecation policies.</p>
<p>This is where many platform teams underperform. Engineering excellence without product orientation produces a platform nobody adopts.</p>
<h2 id="tools-that-matter">Tools that matter</h2>
<p>Platform engineering tools fall in five buckets:</p>
<h3 id="compute-and-orchestration">Compute and orchestration</h3>
<p>Kubernetes is the de facto orchestration layer for most modern platforms. Distributions: vanilla, OpenShift, Rancher, Cozystack (open-source Kubernetes-native virtualization), vendor-led variants. The choice is largely determined by the operational model.</p>
<h3 id="infrastructure-as-code">Infrastructure-as-Code</h3>
<p>Terraform, OpenTofu, Pulumi — for cloud and infrastructure provisioning. Crossplane — for Kubernetes-native infrastructure abstraction. Choice depends on team familiarity and target environments.</p>
<h3 id="gitops">GitOps</h3>
<p>Argo CD and Flux are the two production-grade GitOps engines. Both work; Flux is closer to the upstream Kubernetes way; Argo CD has stronger UI ergonomics. Cozystack uses Flux as the default.</p>
<h3 id="internal-developer-portal">Internal developer portal</h3>
<p>Backstage (CNCF Incubating) is the dominant choice. Alternatives: Port, Cortex, Compass, custom. Important note: Backstage is a portal (catalog + UI), not a platform. The platform sits underneath; the portal exposes it.</p>
<h3 id="observability">Observability</h3>
<p>VictoriaMetrics + VictoriaLogs (open-source, low-overhead — Cozystack default), Prometheus + Loki, or commercial (Datadog, New Relic). Self-hosted matters for sovereignty-sensitive workloads.</p>
<h3 id="secrets-and-identity">Secrets and identity</h3>
<p>External Secrets Operator, HashiCorp Vault, customer-controlled HSMs. Workload identity via SPIFFE/SPIRE for service-to-service auth.</p>
<p>The toolset is converging. The differentiation is in how these tools are wired together as a coherent platform — that&rsquo;s where platform engineering earns its keep.</p>
<h2 id="metrics-that-matter">Metrics that matter</h2>
<p>For evaluating a platform engineering function:</p>
<ul>
<li><strong>Time-to-environment</strong> — from product-team request to a working environment. Target: hours, not weeks.</li>
<li><strong>Golden-path adoption</strong> — fraction of new services using the recommended template / deployment / observability pattern.</li>
<li><strong>Platform-team-to-product-engineer ratio</strong> — typically 1:10 to 1:20 in mature organizations.</li>
<li><strong>Internal NPS / customer satisfaction</strong> — from product teams about the platform.</li>
<li><strong>Toil ratio</strong> — fraction of platform-team time spent on tickets vs. golden-path work. Target: keep below 50%.</li>
<li><strong>Platform availability SLO</strong> — yes, the platform itself has an SLO. Production-grade is 99.9%+ for IDP services.</li>
</ul>
<p>DORA metrics (deployment frequency, lead time, change-failure rate, time-to-restore) are product-team metrics — but the platform&rsquo;s job is to make those metrics achievable.</p>
<h2 id="common-pitfalls-for-the-platform-team">Common pitfalls (for the platform team)</h2>
<h3 id="pitfall-1-building-for-engineers-not-for-product-teams">Pitfall 1: building for engineers, not for product teams</h3>
<p>The platform team&rsquo;s customers are product teams, not other platform engineers. The architecture should optimize for the product team&rsquo;s experience, not for engineering elegance.</p>
<h3 id="pitfall-2-backstage-as-the-platform">Pitfall 2: Backstage as the platform</h3>
<p>Backstage is an excellent IDP frontend. It&rsquo;s not a platform. Buying Backstage without an underlying opinionated platform produces a catalog over the same operational mess.</p>
<h3 id="pitfall-3-too-many-opinions-too-rigid">Pitfall 3: too many opinions, too rigid</h3>
<p>Golden paths should be golden, not gold-plated. Product teams need to be able to deviate when their case is special. The platform&rsquo;s authority comes from being genuinely useful, not from policy enforcement.</p>
<h3 id="pitfall-4-under-staffed">Pitfall 4: under-staffed</h3>
<p>Platform teams that absorb both platform-build and on-call duties for shared services tend to spend their time on tickets. Capacity for golden-path work disappears. The function stalls.</p>
<h3 id="pitfall-5-vendor-led-platform-in-a-box">Pitfall 5: vendor-led platform-in-a-box</h3>
<p>Several vendors sell &ldquo;complete platform engineering solutions.&rdquo; These work for narrow customer profiles but rebuild the lock-in problem. The vendor&rsquo;s roadmap becomes your roadmap.</p>
<h2 id="a-maturity-progression">A maturity progression</h2>
<p>A typical organization moves through these stages:</p>
<ol>
<li><strong>Pre-platform</strong> — each team owns its own infrastructure. Tooling is fragmented. Shared services emerge ad-hoc.</li>
<li><strong>Shared infrastructure</strong> — central team owns shared infrastructure (Kubernetes, CI/CD, observability). No self-service yet; teams file tickets.</li>
<li><strong>Self-service primitives</strong> — central team exposes APIs / IaC modules. Teams self-serve common operations. Limited golden paths.</li>
<li><strong>Internal developer platform</strong> — opinionated platform with golden paths. Most operations are self-service. Platform team has product-team customers.</li>
<li><strong>Mature platform engineering</strong> — platform is a coherent product. SLOs, internal product management, customer-driven roadmap, controlled deprecation.</li>
</ol>
<p>Most organizations sit between stages 2 and 3. The leverage of moving to stage 4 is large but requires intentional investment.</p>
<h2 id="when-to-engage-external-help">When to engage external help</h2>
<p>External platform engineering help is the right call when:</p>
<ul>
<li>A specific deadline (regulator, board mandate, scaling moment) makes internal-only build too slow.</li>
<li>Your team has the long-term capacity but lacks specific senior platform-engineering experience to start cleanly.</li>
<li>You want a structured external assessment to validate or challenge internal direction.</li>
<li>A repatriation or migration program needs platform engineering velocity that exceeds internal hiring rate.</li>
</ul>
<p>It&rsquo;s the wrong call when:</p>
<ul>
<li>The engineering organization isn&rsquo;t ready (no platform-team headcount, no internal customer, no problem clearly named).</li>
<li>The decision is already made and the engagement is meant to validate it.</li>
</ul>
<h2 id="want-to-dig-deeper">Want to dig deeper?</h2>
<ul>
<li><strong><a href="https://aenix.io/services/platform-engineering/">Platform engineering services page</a></strong> — engagement details</li>
<li><strong><a href="https://aenix.io/services/internal-developer-platform/">Internal developer platform</a></strong> — IDP-specific engagement</li>
<li><strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong> — assessment methodology</li>
<li><strong><a href="https://aenix.io/products/cozystack/">Cozystack</a></strong> — the platform we typically build on</li>
</ul>
]]></content:encoded></item><item><title>OpenStack vs Cozystack — modernization options for OpenStack operators in 2026</title><link>https://aenix.io/blog/2026/05/openstack-vs-cozystack-modernization/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/openstack-vs-cozystack-modernization/</guid><pubDate>Thu, 21 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>OpenStack</category><category>Kubernetes</category><category>Cozystack</category><category>Sovereignty</category><category>Migration</category><description>Where OpenStack still wins, where the operational pressure comes from, and the modernization paths available to OpenStack-trained teams.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/openstack-vs-cozystack-modernization.jpg" alt=""></p><p>OpenStack remains widely deployed in telecom and government infrastructure. It also faces structural pressure: shrinking pool of OpenStack engineers, operational complexity that grows with deployment age, and competition from Kubernetes-native alternatives that didn&rsquo;t exist when OpenStack was designed.</p>
<h2 id="where-openstack-still-wins">Where OpenStack still wins</h2>
<ul>
<li><strong>Telecom-scale deployments</strong> — running thousands of nodes with telecom-specific features (NFV, DPDK integration, SR-IOV, high-throughput networking). Established expertise; vendor distros mature.</li>
<li><strong>Government / sovereign clouds</strong> — large public-sector OpenStack deployments where OpenStack is the procurement-mandated platform.</li>
<li><strong>Existing investments</strong> — organizations 5+ years into OpenStack with deep expertise. Migration cost may exceed continuing operations cost.</li>
<li><strong>Specific features</strong> — some OpenStack capabilities (deep network programmability, specific telco features) don&rsquo;t have direct Kubernetes equivalents yet.</li>
</ul>
<h2 id="where-the-pressure-is">Where the pressure is</h2>
<ul>
<li><strong>Engineer scarcity</strong> — OpenStack expertise pool is shrinking. New engineers are trained on Kubernetes, not OpenStack.</li>
<li><strong>Component sprawl</strong> — 30+ services in a typical OpenStack deployment, each with its own lifecycle, upgrade cadence, integration tests.</li>
<li><strong>Upgrade pain</strong> — major-version OpenStack upgrades remain operationally heavy.</li>
<li><strong>Vendor distro fragmentation</strong> — Red Hat OSP, Mirantis, Canonical, others — each with different opinions and support models.</li>
</ul>
<h2 id="modernization-paths">Modernization paths</h2>
<p>For OpenStack operators considering modernization:</p>
<h3 id="path-1-stay-and-tune">Path 1: stay and tune</h3>
<p>Continue OpenStack; invest in operational practices (Helm-based deployments, GitOps, automation). Right when OpenStack expertise is deep and migration cost exceeds value.</p>
<h3 id="path-2-kubernetes-on-openstack">Path 2: Kubernetes-on-OpenStack</h3>
<p>Run Kubernetes platforms (Cozystack or other) on top of OpenStack as a tenant. Adds another platform layer; some teams find this practical for gradual transition.</p>
<h3 id="path-3-parallel-deployment">Path 3: parallel deployment</h3>
<p>Build Cozystack alongside OpenStack on new hardware. Migrate workloads cohort by cohort. Decommission OpenStack as cohorts complete. Most common path for full modernization.</p>
<h3 id="path-4-full-lift-and-shift-to-kubernetes">Path 4: full lift-and-shift to Kubernetes</h3>
<p>Aggressive — replace OpenStack control plane with Kubernetes-based equivalent (Cozystack). Higher risk; faster outcome.</p>
<h2 id="cozystack-architecture-for-openstack-trained-teams">Cozystack architecture for OpenStack-trained teams</h2>
<p>Some helpful translations:</p>
<table>
  <thead>
      <tr>
          <th>OpenStack</th>
          <th>Cozystack</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Nova</td>
          <td>KubeVirt</td>
      </tr>
      <tr>
          <td>Neutron</td>
          <td>Cilium</td>
      </tr>
      <tr>
          <td>Cinder</td>
          <td>LINSTOR (DRBD-replicated block)</td>
      </tr>
      <tr>
          <td>Swift</td>
          <td>SeaweedFS (S3-compatible)</td>
      </tr>
      <tr>
          <td>Keystone</td>
          <td>Kubernetes RBAC + workforce-IdP integration</td>
      </tr>
      <tr>
          <td>Glance</td>
          <td>KubeVirt CDI image registry</td>
      </tr>
      <tr>
          <td>Magnum</td>
          <td>Native — Kubernetes is the platform</td>
      </tr>
      <tr>
          <td>Heat</td>
          <td>Kubernetes operators + GitOps</td>
      </tr>
      <tr>
          <td>Horizon</td>
          <td>Cozystack Dashboard</td>
      </tr>
      <tr>
          <td>Ceilometer</td>
          <td>VictoriaMetrics + VictoriaLogs</td>
      </tr>
      <tr>
          <td>Trove</td>
          <td>Cozystack managed databases</td>
      </tr>
      <tr>
          <td>Designate</td>
          <td>External-DNS operator</td>
      </tr>
      <tr>
          <td>Octavia</td>
          <td>MetalLB / ingress + Cilium L7</td>
      </tr>
  </tbody>
</table>
<p>Most OpenStack engineers find Cozystack&rsquo;s operational model simpler — fewer moving parts, more declarative, integrated observability.</p>
<h2 id="practical-migration-approach">Practical migration approach</h2>
<p>For mid-size (50-500 hosts) OpenStack to Cozystack:</p>
<ol>
<li><strong>Assessment (14 or 28 days)</strong> — current OpenStack deployment, workload classification, target Cozystack architecture.</li>
<li><strong>Cozystack foundation</strong> — parallel deployment on new or repurposed hardware, overlapping the first cohorts.</li>
<li><strong>Migration cohorts</strong> — workloads move cohort by cohort. Image migration via KVM→KubeVirt.</li>
<li><strong>OpenStack decommission</strong> — staged as cohorts complete.</li>
</ol>
<p>Total elapsed: 4-12 months for a mid-size deployment; 12-18 months with complex provider networks or tenant-facing OpenStack APIs.</p>
]]></content:encoded></item><item><title>OpenStack migration — a cohort-based playbook for moving to Cozystack in 2026</title><link>https://aenix.io/blog/2026/05/openstack-migration-cozystack-cohort-playbook/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/openstack-migration-cozystack-cohort-playbook/</guid><pubDate>Wed, 20 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>OpenStack</category><category>Cozystack</category><category>Migration</category><category>Multi-tenancy</category><category>Kubernetes</category><description>Cohort-based playbook for migrating production OpenStack to Cozystack: component mapping, image conversion, networking redesign, handover, and timeline.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/openstack-migration-cozystack-cohort-playbook.jpg" alt=""></p><p>OpenStack remains widely deployed in telecom and large-enterprise
infrastructure. Modernization is not a one-size-fits-all conversation:
mature telco-scale OpenStack with deep vendor distro support is a
different migration than a mid-size enterprise running upstream
OpenStack with a small ops team. Both can move to Cozystack; the
phasing and risk profile differ substantially.</p>
<h2 id="where-openstack-still-works-and-well-say-so">Where OpenStack still works (and we&rsquo;ll say so)</h2>
<p>Before discussing migration, the honest reverse case. OpenStack remains
the right answer for:</p>
<ul>
<li><strong>Tier-1 telco NFV environments</strong> where vendor distro (Red Hat OSP,
Mirantis, Canonical, Wind River) still has support runway and VNF
certification is OpenStack-specific</li>
<li><strong>Very large estates (&gt;1,000 nodes)</strong> with deep OpenStack expertise
where operational complexity is already amortised</li>
<li><strong>Government clouds</strong> where OpenStack is the procurement-mandated
platform (some EU member-state and APAC public-sector tenders)</li>
<li><strong>Existing investments at year 2-3 of a 5-year deployment</strong> where
migration cost would exceed continuing operations cost</li>
</ul>
<p>For everyone else, modernization typically warrants serious evaluation.</p>
<h2 id="what-pushes-openstack-running-organisations-to-migrate">What pushes OpenStack-running organisations to migrate</h2>
<p>Three pressures dominate the 2026 conversation:</p>
<h3 id="1-vendor-distro-lifecycle">1. Vendor distro lifecycle</h3>
<p>Red Hat is steering OSP (OpenStack Platform) customers toward
OpenShift-based offerings. Other distributions (Mirantis, Canonical)
remain on the market, but the vendor landscape for OpenStack
distributions is consolidating. Check the support dates in your own
vendor contract: for many mid-2020s deployments they fall within the
next few years.</p>
<h3 id="2-engineer-scarcity">2. Engineer scarcity</h3>
<p>OpenStack expertise is a shrinking pool. New engineers are trained on
Kubernetes, not on OpenStack&rsquo;s Nova/Neutron/Cinder/Keystone component
model. Operators with deep OpenStack expertise are aging out or moving
to other roles. Hiring is harder; retention is harder.</p>
<h3 id="3-service-catalog-ceiling">3. Service-catalog ceiling</h3>
<p>OpenStack&rsquo;s primary scope is IaaS (compute, network, storage). Managed
databases (Trove), container orchestration (Magnum), and modern
service families (managed Kafka, S3-equivalent at scale, GPU-as-a-
service, AI inference) either bolt on awkwardly or live outside the
platform entirely. For operators whose customers or internal teams
increasingly expect platform-grade services, OpenStack falls behind
without significant additional engineering.</p>
<h2 id="component-mapping">Component mapping</h2>
<p>For OpenStack operators evaluating Cozystack, the canonical translation:</p>
<table>
  <thead>
      <tr>
          <th>OpenStack</th>
          <th>Cozystack equivalent</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><strong>Nova</strong> (compute)</td>
          <td>KubeVirt on Talos</td>
      </tr>
      <tr>
          <td><strong>Neutron</strong> (networking)</td>
          <td>Cilium (eBPF)</td>
      </tr>
      <tr>
          <td><strong>Cinder</strong> (block storage)</td>
          <td>LINSTOR (DRBD-replicated block, via the Piraeus operator)</td>
      </tr>
      <tr>
          <td><strong>Swift</strong> (object storage)</td>
          <td>SeaweedFS (S3-compatible, managed Bucket service)</td>
      </tr>
      <tr>
          <td><strong>Keystone</strong> (identity)</td>
          <td>Kubernetes RBAC + workforce IdP federation (Keycloak / Okta / AD)</td>
      </tr>
      <tr>
          <td><strong>Glance</strong> (image registry)</td>
          <td>KubeVirt CDI + container image registry</td>
      </tr>
      <tr>
          <td><strong>Magnum</strong> (managed Kubernetes)</td>
          <td>Native — Kubernetes IS the platform</td>
      </tr>
      <tr>
          <td><strong>Heat</strong> (orchestration)</td>
          <td>Kubernetes operators + GitOps (Flux / Argo CD)</td>
      </tr>
      <tr>
          <td><strong>Horizon</strong> (UI)</td>
          <td>Cozystack Dashboard</td>
      </tr>
      <tr>
          <td><strong>Ceilometer / Telemetry</strong></td>
          <td>VictoriaMetrics + VictoriaLogs</td>
      </tr>
      <tr>
          <td><strong>Trove</strong> (DBaaS)</td>
          <td>Cozystack managed databases (PostgreSQL, MariaDB, MongoDB, Redis, Valkey, Kafka, NATS, RabbitMQ, ClickHouse, OpenSearch, Qdrant, FoundationDB)</td>
      </tr>
      <tr>
          <td><strong>Designate</strong> (DNS)</td>
          <td>external-dns operator + customer DNS provider</td>
      </tr>
      <tr>
          <td><strong>Octavia</strong> (load balancing)</td>
          <td>MetalLB + Cilium L7 + ingress controllers</td>
      </tr>
      <tr>
          <td><strong>Manila</strong> (file share)</td>
          <td>RWX volumes on DRBD-backed LINSTOR storage classes (Cozystack v1.0+); optional <code>nfs-driver</code> package for external NFS exports</td>
      </tr>
      <tr>
          <td><strong>Project / Domain / Role</strong> (tenancy)</td>
          <td>Tenant CRD + nested tenants</td>
      </tr>
  </tbody>
</table>
<p>Most OpenStack engineers find Cozystack&rsquo;s operational model simpler
once they cross the Kubernetes learning curve — fewer moving parts,
more declarative, integrated observability. The learning curve itself
is real: 4-8 weeks per engineer for focused training, longer for
domain depth.</p>
<h2 id="cohort-based-migration-phases">Cohort-based migration phases</h2>
<h3 id="phase-0--assessment-14-or-28-days">Phase 0 — Assessment (14 or 28 days)</h3>
<p>Inventory the OpenStack deployment:</p>
<ul>
<li><strong>Service-by-service usage</strong> — which OpenStack services do you
actually use (Nova / Neutron / Cinder almost always; everything
else is variable)</li>
<li><strong>Workload inventory</strong> — instance count, OS mix, vCPU/RAM/disk
profiles, criticality tier</li>
<li><strong>Network inventory</strong> — subnets, security groups, floating IPs,
external network configurations, Octavia load balancers</li>
<li><strong>Storage inventory</strong> — Cinder volume types, snapshot retention,
Swift bucket usage</li>
<li><strong>Multi-tenancy</strong> — project / domain hierarchy, custom roles, RBAC
policies</li>
<li><strong>Integrations</strong> — vendor VNF certifications, observability hooks,
CI/CD pipelines, ITSM tooling</li>
</ul>
<p>Output: migration plan with workload buckets (migrate-now /
migrate-later / stay / re-architect), risk flags, phasing options.</p>
<h3 id="phase-1--cozystack-foundation">Phase 1 — Cozystack foundation</h3>
<p>Hardware procurement (or repurpose of OpenStack-freed capacity in
later phases). Cozystack platform deployed on new hardware in parallel
to existing OpenStack. Cilium networking validated against OpenStack
network configurations expected to translate. LINSTOR storage
operationalised at scale. Federated identity (Keycloak + customer
IdP).</p>
<p>Tenant CRD model designed to map cleanly from OpenStack project /
domain hierarchy. Each top-level OpenStack project typically becomes
a Cozystack Tenant; nested projects become nested Tenants.</p>
<p>End state: Cozystack platform up, internally validated, ready for
workload onboarding.</p>
<h3 id="phase-2--operational-tooling">Phase 2 — Operational tooling</h3>
<p>Observability stack (VictoriaMetrics + VictoriaLogs) integrated with
customer SIEM. Backup/DR with Velero + per-app patterns. Runbook
library. On-call rotation. Incident response process. GitOps
deployment workflow (Flux default; Argo CD if customer prefers).</p>
<p>End state: production-operations-grade tooling in place; team
training in progress.</p>
<h3 id="phase-3--workload-migration-cohorts">Phase 3 — Workload migration cohorts</h3>
<p>Cohorts of 50-200 instances migrating at a time. Per cohort:</p>
<ol>
<li><strong>Image conversion</strong> — OpenStack Glance images converted to
KubeVirt-compatible format. Most KVM-based OpenStack images
convert with <code>qemu-img convert</code> + minor metadata adjustment.
Windows instances get virtio driver injection.</li>
<li><strong>Network mapping</strong> — OpenStack subnets to Cilium ClusterPool +
NetworkPolicies, security groups to NetworkPolicies, Octavia load
balancers to MetalLB + ingress controllers.</li>
<li><strong>Storage migration</strong> — Cinder volume data migrated to LINSTOR.
Snapshot history preserved or pruned per retention policy.
For Swift-stored data, S3 API compatibility allows direct
migration to SeaweedFS.</li>
<li><strong>Validation window</strong> — workload runs in parallel on Cozystack
for typically 7-14 days. Application owner sign-off required
before final cutover.</li>
<li><strong>Final cutover</strong> — DNS / load balancer flip to Cozystack
endpoint. OpenStack instance kept available for 7-30 days
rollback window.</li>
</ol>
<h3 id="phase-4--operational-handover-parallel-to-phase-3">Phase 4 — Operational handover (parallel to Phase 3)</h3>
<p>Ænix engineers reduce direct involvement. Customer operations team
absorbs first- and second-line incidents. Ænix support (Plus or
Enterprise tier for 24×7) continues for escalation. Documentation handover. Knowledge transfer
sessions.</p>
<h3 id="phase-5--openstack-decommission">Phase 5 — OpenStack decommission</h3>
<p>As migration cohorts complete, OpenStack capacity is repurposed into
the Cozystack cluster. Hardware is the same commodity x86; OpenStack
software stack is retired. Vendor distro contracts wound down per
their lifecycle.</p>
<h2 id="where-openstack-to-cozystack-migrations-stumble">Where OpenStack-to-Cozystack migrations stumble</h2>
<h3 id="1-network-model-redesign">1. Network model redesign</h3>
<p>OpenStack Neutron&rsquo;s network model has historically been highly
configurable — provider networks, tenant networks, security groups,
floating IPs, FWaaS, VPNaaS, complex routing topologies. Cilium&rsquo;s
eBPF model is fundamentally different: L4/L7 NetworkPolicies, eBPF-
based routing, Hubble for observability.</p>
<p>Translating elaborate Neutron configurations to Cilium requires
careful architecture work in Phase 0-1. Skipping this produces
production incidents in Phase 3 where customer workloads expected
specific network behaviour that doesn&rsquo;t translate directly.</p>
<h3 id="2-multi-tenant-policy-translation">2. Multi-tenant policy translation</h3>
<p>OpenStack&rsquo;s role-based access (Keystone roles, project policies)
doesn&rsquo;t map 1:1 to Kubernetes RBAC + Tenant CRD. Custom roles
defined for specific OpenStack APIs may not have direct Kubernetes
equivalents. Plan time for policy redesign rather than mechanical
translation.</p>
<h3 id="3-vnf-certification">3. VNF certification</h3>
<p>Tier-1 telco deployments running certified VNFs face a different
challenge: the VNF vendor certifies on specific OpenStack distros,
not on Cozystack. Three approaches:</p>
<ul>
<li><strong>Run VNFs on KubeVirt under Cozystack</strong> — the VNF runs as a VM;
vendor certification may or may not extend to this configuration.
Vendor dialogue required.</li>
<li><strong>Keep certified VNFs on OpenStack</strong> — parallel platforms for the
lifecycle of the certified VNF; modernization track for new VNFs
on Cozystack.</li>
<li><strong>Negotiate cloud-native equivalent</strong> — many VNF vendors are
migrating toward Cloud-Native Network Functions (CNFs) on
Kubernetes; the migration may align with vendor&rsquo;s own CNF
modernization.</li>
</ul>
<p>In practice all three patterns appear, often in the same operator,
depending on which VNF vendor and which generation.</p>
<h3 id="4-operational-culture-shift">4. Operational culture shift</h3>
<p>OpenStack operators are used to imperative APIs (CLI commands, REST
calls, console actions). Cozystack expects GitOps for production
changes. This is a culture shift, not just a tool shift. Engineers
who built OpenStack expertise around imperative workflows need 4-8
weeks of focused training plus 3-6 months of practice to internalise
the GitOps discipline.</p>
<p>Ænix engagement includes training as a workstream; customer
investment in the cultural transition is also required.</p>
<h2 id="timeline-realities">Timeline realities</h2>
<p>Mid-size enterprise (200-500 nodes, simple multi-tenancy, mostly
default networking):</p>
<ul>
<li>Phase 0: 14 or 28 days</li>
<li>Phases 1-2: foundation and tooling, overlapping the first cohorts</li>
<li>Phase 3: the bulk of the elapsed time</li>
<li>Phases 4-5: alongside the last cohorts</li>
</ul>
<p><strong>Total: 4-12 months for a mid-size deployment; 12-18 months with
complex provider networks or tenant-facing OpenStack APIs</strong></p>
<p>Tier-1 telco (1,000-5,000 nodes, complex multi-tenancy, certified VNF
environments, NFV-specific networking):</p>
<ul>
<li>Phase 0: 2-3 months</li>
<li>Phase 1: 4-6 months</li>
<li>Phase 2: 2-3 months</li>
<li>Phase 3: 12-24 months (multiple parallel cohorts)</li>
<li>Phase 4-5: 6-12 months</li>
<li>VNF modernization track: 18-36 months in parallel</li>
</ul>
<p><strong>Total: 24-48 months for full modernization; first production
workloads on Cozystack within 12-18 months</strong></p>
<h2 id="when-openstack-to-cozystack-migration-is-the-right-answer">When OpenStack-to-Cozystack migration is the right answer</h2>
<p>Strong fit:</p>
<ul>
<li>Vendor distro lifecycle is forcing a decision within 24 months</li>
<li>Engineer scarcity is starting to hurt operational quality</li>
<li>Customer demand for managed databases / S3 / containers / GPU
services exceeds what OpenStack natively handles</li>
<li>Modernisation budget is available across 2-4 years</li>
</ul>
<p>Marginal fit:</p>
<ul>
<li>Stable, mature OpenStack deployment with deep team expertise and
no service-catalog pressure — modernisation can wait for vendor
lifecycle</li>
<li>Very large estates (&gt;5,000 nodes) where modernisation cost is in
the multi-tens-of-millions range — phased multi-year programme
required</li>
</ul>
<p>Poor fit:</p>
<ul>
<li>Recently-deployed OpenStack at year 1 of a 5-year programme —
finish what you started; modernise at lifecycle</li>
<li>OpenStack-on-government-procurement-mandate without flexibility to
change platforms</li>
</ul>
<h2 id="engagement-structure">Engagement structure</h2>
<ul>
<li><strong>Discovery call</strong> (30 min, free)</li>
<li><strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong>
(fixed price, 14 days focused or 28 days full) — workload buckets,
phasing options, risk flags</li>
<li><strong>Pilot deployment</strong> (2-3 months) — Cozystack stood up, 50-100
workloads migrated, billing / operational workflows validated</li>
<li><strong>Cohort migration</strong> — workload migration in cohorts; 4-12 months
in total for a mid-size deployment, 12-18 months with complex
provider networks or tenant-facing OpenStack APIs, longer at
telco scale</li>
<li><strong>OpenStack decommission</strong> (parallel to cohort migration) — staged
as cohorts complete</li>
<li><strong>Support subscription</strong> (ongoing) — Plus or Enterprise support tier
for 24×7 escalation (see <a href="https://aenix.io/pricing/">/pricing/</a>)</li>
</ul>
<h2 id="where-to-dig-deeper">Where to dig deeper</h2>
<ul>
<li><strong><a href="https://aenix.io/migration/openstack/">OpenStack migration hub</a></strong> — commercial
landing</li>
<li><strong><a href="https://aenix.io/blog/2026/05/openstack-vs-cozystack-modernization/">OpenStack vs Cozystack modernization</a></strong> —
the modernization path analysis</li>
<li><strong><a href="https://aenix.io/alternatives/openstack-alternative/">OpenStack alternative</a></strong> —
alternative-focused commercial landing</li>
<li><strong><a href="https://aenix.io/products/public-cloud-platform/">Public Cloud Platform product page</a></strong> —
common target product for hosting-provider OpenStack migrations</li>
</ul>
]]></content:encoded></item><item><title>OpenShift vs Cozystack — comparison for KubeVirt-based platform decisions</title><link>https://aenix.io/blog/2026/05/openshift-vs-cozystack-comparison/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/openshift-vs-cozystack-comparison/</guid><pubDate>Tue, 19 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>OpenShift</category><category>Kubernetes</category><category>Cozystack</category><category>KubeVirt</category><category>Cilium</category><category>LINSTOR</category><description>Two KubeVirt-based platforms compared: shared foundations, where they genuinely differ, when OpenShift wins, and what migration between them involves.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/openshift-vs-cozystack-comparison.jpg" alt=""></p><p>OpenShift Virtualization (Red Hat) and Cozystack (Ænix / CNCF Project) are the two most-mature KubeVirt-based platforms in 2026. They share architectural foundations but differ in commercial model, operational footprint, and vendor relationship.</p>
<h2 id="shared-foundations">Shared foundations</h2>
<p>Both platforms run KubeVirt for VM workloads alongside Kubernetes containers. Both inherit Kubernetes operational patterns (declarative config, GitOps, RBAC, observability). Both support production VM workloads with live migration, snapshots, and multi-tenant isolation.</p>
<h2 id="where-they-differ">Where they differ</h2>
<h3 id="commercial-model">Commercial model</h3>
<p><strong>OpenShift Virtualization:</strong> Red Hat commercial subscription. Per-core or per-socket pricing. Includes Red Hat support, certification, and ecosystem access.</p>
<p><strong>Cozystack:</strong> Apache 2.0 open source. Ænix offers commercial support tiers; you can self-deploy with no commercial relationship.</p>
<p>For organizations standardized on Red Hat procurement, OpenShift is administratively simpler. For organizations seeking open-source-first or service-provider economics, Cozystack fits better.</p>
<h3 id="operational-footprint">Operational footprint</h3>
<p><strong>OpenShift:</strong> broad surface area — OpenShift Container Platform plus Virtualization plus Service Mesh plus Pipelines plus other addons. Operationally rich; team needs OpenShift-specific expertise.</p>
<p><strong>Cozystack:</strong> focused stack — KubeVirt + Cilium + Kube-OVN + LINSTOR + Cozystack Dashboard + observability. Lighter operational footprint; standard Kubernetes operational expertise transfers directly.</p>
<h3 id="multi-tenancy">Multi-tenancy</h3>
<p><strong>OpenShift:</strong> namespace-based with Project CRD; RBAC and quotas at namespace level. Works for enterprise multi-BU.</p>
<p><strong>Cozystack:</strong> Tenant CRD with nested tenants, scoped audit, billing-friendly model. Works for service-provider multi-customer plus enterprise multi-BU.</p>
<h3 id="vendor-relationship">Vendor relationship</h3>
<p><strong>OpenShift:</strong> Red Hat / IBM relationship. Roadmap shaped by Red Hat&rsquo;s commercial decisions.</p>
<p><strong>Cozystack:</strong> open-source community-governed (CNCF Project). Ænix is the largest contributor but not the owner. Roadmap shaped by community + commercial users.</p>
<h3 id="ecosystem-integration">Ecosystem integration</h3>
<p><strong>OpenShift:</strong> integrates with broader Red Hat ecosystem (Ansible, Satellite, Identity Management, etc.).</p>
<p><strong>Cozystack:</strong> integrates with broader CNCF ecosystem (Prometheus stack alternatives, Argo, Crossplane, etc.).</p>
<h2 id="when-openshift-wins">When OpenShift wins</h2>
<ul>
<li>Existing Red Hat / OpenShift commitments</li>
<li>Enterprise procurement standardized on Red Hat</li>
<li>Need integrated Red Hat ecosystem (Ansible Automation Platform, etc.)</li>
<li>Procurement that requires a Red Hat contract specifically</li>
</ul>
<h2 id="when-cozystack-wins">When Cozystack wins</h2>
<ul>
<li>SLA-backed support from the maintainers — published response times,
24×7 on the Plus and Enterprise tiers (<a href="https://aenix.io/pricing/">pricing</a>); AENIX
s.r.o. holds <a href="https://aenix.io/compliance/iso-27001/">ISO/IEC 27001</a></li>
<li>Open-source-first procurement</li>
<li>Service-provider model (multi-customer cloud)</li>
<li>Sovereignty / regulator requirements where open-source matters</li>
<li>Cost-sensitive at scale (no per-core subscription)</li>
<li>Greenfield with no existing Red Hat relationship</li>
<li>Need lighter operational footprint than full OpenShift</li>
</ul>
<h2 id="migration-between-them">Migration between them</h2>
<p>Both KubeVirt-based, so VM-level migration is straightforward (image-level compatibility). The architectural delta is in:</p>
<ul>
<li>Multi-tenancy model (Project CRD vs Tenant CRD)</li>
<li>Networking (OVN-Kubernetes, which replaced the deprecated OpenShift SDN, vs Cilium)</li>
<li>Storage (OpenShift Data Foundation / Ceph vs LINSTOR / DRBD)</li>
<li>Operational tooling (OpenShift CLI/Console vs Cozystack Dashboard/standard kubectl)</li>
</ul>
<p>Realistic migration timeline: 3-9 months for mid-size deployment.</p>
<h2 id="how-to-choose">How to choose</h2>
<p>If your situation matches OpenShift&rsquo;s strengths (Red Hat ecosystem, enterprise procurement, broad capability surface) — OpenShift Virtualization. If your situation matches Cozystack&rsquo;s strengths (open-source, service-provider, lighter footprint) — Cozystack.</p>
<p>For a specific evaluation, the assessment phase of either engagement helps clarify. See <strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong>.</p>
]]></content:encoded></item><item><title>Nutanix vs Cozystack vs VMware — choosing your virtualization platform in 2026</title><link>https://aenix.io/blog/2026/05/nutanix-vs-cozystack-vs-vmware/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/nutanix-vs-cozystack-vs-vmware/</guid><pubDate>Tue, 19 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>VMware</category><category>Nutanix</category><category>Kubernetes</category><category>Cozystack</category><category>KubeVirt</category><category>Cilium</category><description>Nutanix HCI with AHV, VMware after Broadcom, and Cozystack compared on architecture, where each wins, and the migration economics between them.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/nutanix-vs-cozystack-vs-vmware.jpg" alt=""></p><p>In 2026 the realistic shortlist for production virtualization platforms includes (among others) Nutanix AHV, VMware Cloud Foundation, and Cozystack. Each represents a different architectural philosophy.</p>
<h2 id="architectural-philosophies">Architectural philosophies</h2>
<p><strong>Nutanix:</strong> vendor-led integrated HCI appliance. Operational simplicity and integrated support are the core value proposition.</p>
<p><strong>VMware:</strong> mature legacy stack with deep ecosystem integration. Vendor-managed roadmap; subscription-led economics.</p>
<p><strong>Cozystack:</strong> open-source Kubernetes-native platform. Customer-controlled architecture; community-governed roadmap.</p>
<h2 id="detailed-comparison">Detailed comparison</h2>
<table>
  <thead>
      <tr>
          <th></th>
          <th>Nutanix AHV</th>
          <th>VMware (VCF)</th>
          <th>Cozystack</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><strong>Licence</strong></td>
          <td>Subscription</td>
          <td>Subscription only</td>
          <td>Apache 2.0</td>
      </tr>
      <tr>
          <td><strong>Open source</strong></td>
          <td>No</td>
          <td>No</td>
          <td>Full</td>
      </tr>
      <tr>
          <td><strong>Foundation</strong></td>
          <td>Proprietary KVM (AHV)</td>
          <td>vSphere/ESXi</td>
          <td>KubeVirt on Kubernetes</td>
      </tr>
      <tr>
          <td><strong>Multi-tenancy</strong></td>
          <td>Limited</td>
          <td>vCloud Director</td>
          <td>Tenant CRD</td>
      </tr>
      <tr>
          <td><strong>Storage</strong></td>
          <td>Distributed (proprietary)</td>
          <td>vSAN</td>
          <td>LINSTOR (DRBD)</td>
      </tr>
      <tr>
          <td><strong>Network</strong></td>
          <td>AHV networking</td>
          <td>NSX</td>
          <td>Cilium</td>
      </tr>
      <tr>
          <td><strong>Containers</strong></td>
          <td>NKP (separate)</td>
          <td>Tanzu (separate)</td>
          <td>Native</td>
      </tr>
      <tr>
          <td><strong>Hardware</strong></td>
          <td>Nutanix NX or certified OEM (Dell, HPE, Lenovo, others)</td>
          <td>x86</td>
          <td>Commodity x86</td>
      </tr>
      <tr>
          <td><strong>Best for</strong></td>
          <td>HCI-focused enterprises</td>
          <td>Existing VMware estates</td>
          <td>Service providers + sovereign cloud</td>
      </tr>
  </tbody>
</table>
<h2 id="when-each-wins">When each wins</h2>
<h3 id="nutanix-wins">Nutanix wins</h3>
<ul>
<li>Existing Nutanix HCI investment with operational expertise</li>
<li>Strong preference for integrated appliance + commercial support</li>
<li>VM-only workload portfolio</li>
<li>Mid-size enterprise with integrated procurement</li>
</ul>
<h3 id="vmware-wins">VMware wins</h3>
<ul>
<li>Existing VMware estate where renewal economics are still tolerable</li>
<li>Deep vSphere expertise that&rsquo;s hard to migrate</li>
<li>Specific VMware-only features (some niche advanced networking, storage)</li>
<li>(Increasingly rare in 2026 due to Broadcom pricing)</li>
</ul>
<h3 id="cozystack-wins">Cozystack wins</h3>
<ul>
<li>Service-provider or multi-tenant cloud-builder model</li>
<li>Sovereignty / regulator requirements</li>
<li>Open-source procurement preference</li>
<li>Mixed VM + container workloads on one platform</li>
<li>AI/GPU at scale with Kubernetes-native tooling</li>
</ul>
<h2 id="migration-economics">Migration economics</h2>
<p>Migrating between these platforms is not free. Realistic cost estimates:</p>
<ul>
<li><strong>VMware → Cozystack:</strong> about 8-12 months for ~100 VMs and 18-24 months for ~1,000 VMs, including planning and migration waves; assessment + destination build + cohort migration. Net positive after Year 2 typically.</li>
<li><strong>VMware → Nutanix:</strong> Similar timeline; uses Nutanix Move tooling.</li>
<li><strong>Nutanix → Cozystack:</strong> full-estate migration typically 9-18 months, by scope; KVM image compatibility helps.</li>
<li><strong>Cozystack → VMware/Nutanix:</strong> Rare in 2026 (reverse migration).</li>
</ul>
<h2 id="how-to-decide">How to decide</h2>
<p>The decision tree:</p>
<ol>
<li><strong>Existing platform with deep expertise + economics still work?</strong> → Stay.</li>
<li><strong>Multi-tenant or service-provider model?</strong> → Cozystack.</li>
<li><strong>Sovereignty / open-source-first procurement?</strong> → Cozystack.</li>
<li><strong>HCI appliance preference + Nutanix relationship?</strong> → Nutanix.</li>
<li><strong>VMware estate, no triggers to leave?</strong> → VMware (with eye on next renewal).</li>
<li><strong>Greenfield + Kubernetes-fluent team?</strong> → Cozystack.</li>
</ol>
]]></content:encoded></item><item><title>NIS2 requirements for cloud infrastructure — a checklist for in-scope entities in 2026</title><link>https://aenix.io/blog/2026/05/nis2-requirements-cloud-infrastructure-checklist/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/nis2-requirements-cloud-infrastructure-checklist/</guid><pubDate>Mon, 18 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>NIS2</category><category>Financial Services</category><category>Compliance</category><description>NIS2 Articles 21, 23, 28 and 12 mapped to concrete cloud architecture controls, with a working checklist and the architectural failures that recur.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/nis2-requirements-cloud-infrastructure-checklist.jpg" alt=""></p><p>The Network and Information Security Directive 2 (Directive (EU) 2022/2555 — NIS2) replaced the original NIS Directive in 2023. Transposition into national law was due by 17 October 2024. Some EU member states completed transposition on time; others ran late. Either way, by mid-2025 NIS2 is operational across the EU, with competent authorities in each member state and the European Cybersecurity Agency (ENISA) playing a coordination role.</p>
<p>For in-scope entities — and the ICT third parties serving them — the architectural work is concrete, even if the regulatory text is broad.</p>
<h2 id="who-is-in-scope">Who is in scope</h2>
<p>NIS2 applies to two categories of entities, with sectoral scoping:</p>
<h3 id="annex-i--sectors-of-high-criticality">Annex I — sectors of high criticality</h3>
<ul>
<li>Energy (electricity, gas, oil, district heating/cooling, hydrogen)</li>
<li>Transport (air, rail, water, road)</li>
<li>Banking</li>
<li>Financial market infrastructures</li>
<li>Healthcare</li>
<li>Drinking water</li>
<li>Wastewater</li>
<li><strong>Digital infrastructure</strong> — IXPs, DNS service providers, TLD name registries, cloud providers, datacenter providers, CDN providers, public electronic communications networks/services</li>
<li>ICT service management (B2B) — MSPs, MSSPs</li>
<li>Public administration (central + at member state&rsquo;s discretion regional)</li>
<li>Space</li>
</ul>
<h3 id="annex-ii--other-critical-sectors">Annex II — other critical sectors</h3>
<ul>
<li>Postal and courier services</li>
<li>Waste management</li>
<li>Chemical (manufacture, production, distribution)</li>
<li>Food production, processing, distribution</li>
<li>Manufacturing of critical products (medical devices, computers, electronic equipment, machinery, motor vehicles)</li>
<li><strong>Digital service providers</strong> — online marketplaces, search engines, social networking platforms</li>
<li>Research</li>
</ul>
<p>Whether an entity is essential or important is decided by Article 3, not by the annex alone. Large entities in Annex I sectors are essential; medium-sized entities in Annex I sectors and entities in Annex II sectors are important, unless a member state designates them otherwise. Some entities are in scope regardless of size (e.g., DNS service providers and TLD registries, which are essential).</p>
<h2 id="article-21--risk-management-measures">Article 21 — risk management measures</h2>
<p>The substantive heart of NIS2 cybersecurity requirements. Member states must ensure essential and important entities take appropriate technical, operational, and organisational measures across:</p>
<ol>
<li>Policies on risk analysis and information system security</li>
<li>Incident handling</li>
<li>Business continuity (backup management, disaster recovery, crisis management)</li>
<li>Supply chain security (including security-related aspects in relationships with direct suppliers and service providers)</li>
<li>Security in network and information systems acquisition, development, and maintenance (incl. vulnerability handling and disclosure)</li>
<li>Policies and procedures to assess effectiveness of cybersecurity risk-management measures</li>
<li>Basic cyber hygiene practices and cybersecurity training</li>
<li>Policies and procedures regarding the use of cryptography and, where appropriate, encryption</li>
<li>Human resources security, access control policies, asset management</li>
<li>Use of multi-factor authentication / continuous authentication, secured voice/video/text comms, secured emergency comms</li>
</ol>
<p>For cloud architecture, that maps to specific controls.</p>
<h2 id="article-21--architecture-mapping">Article 21 → architecture mapping</h2>
<h3 id="risk-analysis-and-information-system-security">Risk analysis and information system security</h3>
<ul>
<li>Risk register documented per critical workload</li>
<li>Threat model annual review</li>
<li>Network segmentation (zero-trust by default)</li>
<li>Pod Security Standards / Kubernetes Network Policies enforced</li>
</ul>
<h3 id="incident-handling">Incident handling</h3>
<ul>
<li>24×7 detection capability for critical workloads</li>
<li>Documented incident response with named roles</li>
<li>Blameless post-mortems with action items</li>
</ul>
<h3 id="business-continuity">Business continuity</h3>
<ul>
<li>Backup with tested recovery (Velero + per-app PITR for databases)</li>
<li>DR site in separate region or jurisdiction</li>
<li>BCP exercises at least annually with documented results</li>
<li>RTO/RPO documented and tested per critical workload</li>
</ul>
<h3 id="supply-chain-security">Supply chain security</h3>
<ul>
<li>ICT third-party inventory complete</li>
<li>Critical-function dependencies mapped to second hop</li>
<li>Contractual security clauses in critical-function agreements</li>
<li>Continuous monitoring of supplier security posture</li>
<li>Concentration-risk position documented</li>
</ul>
<h3 id="acquisition-development-maintenance-security">Acquisition, development, maintenance security</h3>
<ul>
<li>SAST / DAST in CI</li>
<li>Container image scanning + SBOM</li>
<li>Vulnerability handling SLA defined and met</li>
<li>Disclosure policy published</li>
</ul>
<h3 id="effectiveness-assessment">Effectiveness assessment</h3>
<ul>
<li>Annual external assessment recommended for essential entities</li>
<li>Internal review quarterly</li>
<li>Metrics tracked over time</li>
</ul>
<h3 id="cyber-hygiene-and-training">Cyber hygiene and training</h3>
<ul>
<li>Mandatory cybersecurity training for all staff</li>
<li>Phishing simulation</li>
<li>Privileged-user training</li>
<li>Management body cybersecurity training (Article 20)</li>
</ul>
<h3 id="cryptography">Cryptography</h3>
<ul>
<li>TLS 1.2+ for transit</li>
<li>Encryption at rest for sensitive data classes</li>
<li>Customer-controlled keys for the most sensitive workloads</li>
<li>Key rotation and emergency access documented</li>
</ul>
<h3 id="access-control--human-resources">Access control / human resources</h3>
<ul>
<li>Workload identity (SPIFFE/SPIRE or equivalent)</li>
<li>Privileged access management</li>
<li>Joiner-mover-leaver process automated where possible</li>
<li>Asset register complete and current</li>
</ul>
<h3 id="mfa--secured-comms">MFA / secured comms</h3>
<ul>
<li>MFA for all privileged accounts</li>
<li>MFA for end-user access where reasonable</li>
<li>Secured comms for incident response</li>
</ul>
<h2 id="article-23--incident-reporting">Article 23 — incident reporting</h2>
<p>NIS2 mandates a three-stage incident reporting process for significant incidents:</p>
<ul>
<li><strong>Early warning</strong> within 24 hours of becoming aware</li>
<li><strong>Incident notification</strong> within 72 hours, including initial assessment of severity, impact, and (where applicable) cross-border implications</li>
<li><strong>Final report</strong> within one month, with detailed description, type of threat or root cause, applied/ongoing mitigation, and (where applicable) cross-border impact</li>
</ul>
<p>For cloud architecture this requires:</p>
<ul>
<li>Detection telemetry capable of recognizing significant incidents at a 24-hour timeline (often the bottleneck — alert fatigue masks signal)</li>
<li>Incident classification process (what counts as significant?)</li>
<li>Incident response procedures with clear escalation triggers</li>
<li>Communication channels with national CSIRT / competent authority pre-established</li>
<li>Documentation discipline so the 72-hour and one-month reports are evidence-based</li>
</ul>
<h2 id="article-3-and-article-27--registration">Article 3 and Article 27 — registration</h2>
<p>Many essential and important entities must register with their national competent authority. The architecture decision: ensure your registration data (including DNS, IP ranges, contact details) reflects what&rsquo;s actually deployed.</p>
<h2 id="article-12--coordinated-vulnerability-disclosure">Article 12 — coordinated vulnerability disclosure</h2>
<p>Member states must designate a CSIRT to coordinate vulnerability disclosures. Entities should publish a vulnerability disclosure policy. Architecture: a structured CSAF feed or Hugo-equivalent for security advisories, with named contact points.</p>
<h2 id="a-working-nis2-cloud-architecture-checklist">A working NIS2 cloud architecture checklist</h2>
<p>Use this during architecture review or production-readiness assessment.</p>
<h3 id="risk-management">Risk management</h3>
<ul>
<li><input disabled="" type="checkbox"> Risk register documented per critical workload</li>
<li><input disabled="" type="checkbox"> Annual risk review completed</li>
<li><input disabled="" type="checkbox"> ICT asset register complete and current</li>
<li><input disabled="" type="checkbox"> Threat model documented</li>
<li><input disabled="" type="checkbox"> Cybersecurity policy approved by management body</li>
</ul>
<h3 id="incident-handling-1">Incident handling</h3>
<ul>
<li><input disabled="" type="checkbox"> 24×7 detection capability for critical workloads</li>
<li><input disabled="" type="checkbox"> Documented incident response with named roles</li>
<li><input disabled="" type="checkbox"> Incident classification process documented</li>
<li><input disabled="" type="checkbox"> Pre-established communication channels with CSIRT/authority</li>
<li><input disabled="" type="checkbox"> Reporting templates aligned to 24/72-hour/one-month structure</li>
<li><input disabled="" type="checkbox"> Annual incident-response exercise</li>
</ul>
<h3 id="business-continuity-1">Business continuity</h3>
<ul>
<li><input disabled="" type="checkbox"> BCP plan documented per critical workload</li>
<li><input disabled="" type="checkbox"> RTO/RPO documented per critical workload</li>
<li><input disabled="" type="checkbox"> Backup tested with recovery in past 12 months</li>
<li><input disabled="" type="checkbox"> DR site operational and tested</li>
<li><input disabled="" type="checkbox"> BCP exercise within past 12 months</li>
</ul>
<h3 id="supply-chain">Supply chain</h3>
<ul>
<li><input disabled="" type="checkbox"> ICT third-party inventory complete</li>
<li><input disabled="" type="checkbox"> Critical-function dependencies mapped to second hop</li>
<li><input disabled="" type="checkbox"> Contractual security clauses in critical-function agreements</li>
<li><input disabled="" type="checkbox"> Continuous supplier monitoring documented</li>
<li><input disabled="" type="checkbox"> Concentration-risk position documented</li>
</ul>
<h3 id="cybersecurity-hygiene">Cybersecurity hygiene</h3>
<ul>
<li><input disabled="" type="checkbox"> MFA for privileged accounts (universal)</li>
<li><input disabled="" type="checkbox"> MFA for end-user access where reasonable</li>
<li><input disabled="" type="checkbox"> Cybersecurity training mandatory for all staff</li>
<li><input disabled="" type="checkbox"> Management body trained</li>
<li><input disabled="" type="checkbox"> Privileged-access management deployed</li>
<li><input disabled="" type="checkbox"> Joiner-mover-leaver automated</li>
</ul>
<h3 id="vulnerability-management">Vulnerability management</h3>
<ul>
<li><input disabled="" type="checkbox"> CVE response SLA documented</li>
<li><input disabled="" type="checkbox"> SAST/DAST in CI pipelines</li>
<li><input disabled="" type="checkbox"> Container scanning + SBOM</li>
<li><input disabled="" type="checkbox"> Disclosure policy published</li>
<li><input disabled="" type="checkbox"> Coordinated vulnerability response process</li>
</ul>
<h3 id="cryptography-1">Cryptography</h3>
<ul>
<li><input disabled="" type="checkbox"> TLS 1.2+ for transit</li>
<li><input disabled="" type="checkbox"> Encryption at rest for sensitive data</li>
<li><input disabled="" type="checkbox"> Customer-controlled keys where applicable</li>
<li><input disabled="" type="checkbox"> Key rotation documented</li>
</ul>
<h3 id="audit-and-effectiveness">Audit and effectiveness</h3>
<ul>
<li><input disabled="" type="checkbox"> Annual external assessment (essential entities)</li>
<li><input disabled="" type="checkbox"> Quarterly internal review</li>
<li><input disabled="" type="checkbox"> Audit log retention meets longest applicable requirement</li>
<li><input disabled="" type="checkbox"> Metrics tracked over time</li>
<li><input disabled="" type="checkbox"> Management body kept informed</li>
</ul>
<h2 id="common-architectural-failures">Common architectural failures</h2>
<h3 id="failure-1-detection-telemetry-tuned-for-performance-not-security">Failure 1: detection telemetry tuned for performance, not security</h3>
<p>Observability built for SLO compliance; not tuned to alert on security-significant events. NIS2 requires the latter.</p>
<h3 id="failure-2-bcp-plan-never-tested">Failure 2: BCP plan never tested</h3>
<p>NIS2 requires BCP. Most plans have never been exercised under realistic failure.</p>
<h3 id="failure-3-supply-chain-visible-only-to-first-hop">Failure 3: supply chain visible only to first hop</h3>
<p>Article 21(2)(d) requires supply-chain security including direct suppliers and service providers. Most entities cannot enumerate beyond first hop.</p>
<h3 id="failure-4-incident-reporting-process-undocumented">Failure 4: incident-reporting process undocumented</h3>
<p>24/72-hour/one-month timeline is tight. Without pre-established process and templates, the first real incident becomes a scramble.</p>
<h3 id="failure-5-management-body-untrained">Failure 5: management body untrained</h3>
<p>Article 20 requires management body to follow cybersecurity training. Many compliance programs treat NIS2 as IT-only.</p>
<h2 id="how-to-assess-where-you-stand">How to assess where you stand</h2>
<p>A structured NIS2 readiness assessment is the cheapest insurance before regulator dialog. Ænix runs this as part of <strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong> with sovereignty/regulator workstream emphasized.</p>
<p>For details see the <strong><a href="https://aenix.io/solutions/nis2-compliance/">NIS2 compliance services page</a></strong>.</p>
<hr>
<h2 id="want-to-dig-deeper">Want to dig deeper?</h2>
<ul>
<li><strong><a href="https://aenix.io/solutions/nis2-compliance/">NIS2 compliance services</a></strong> — engagement details</li>
<li><strong><a href="https://aenix.io/solutions/dora-compliance/">DORA compliance</a></strong> — financial-services regulator</li>
<li><strong><a href="https://aenix.io/solutions/data-sovereignty/">Data sovereignty</a></strong> — adjacent trigger</li>
<li><strong><a href="https://aenix.io/products/cozystack/">Cozystack</a></strong> — sovereign-by-architecture platform</li>
</ul>
]]></content:encoded></item><item><title>MSP cloud platform modernization — branded cloud as managed-service offering</title><link>https://aenix.io/blog/2026/05/msp-cloud-platform-modernization/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/msp-cloud-platform-modernization/</guid><pubDate>Mon, 18 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Cozystack</category><category>Multi-tenancy</category><category>Hosting</category><category>Observability</category><description>Architecture pattern, reseller economics, and engagement sequencing for MSPs adding a multi-tenant cloud platform to a managed-services business.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/msp-cloud-platform-modernization.jpg" alt=""></p><h2 id="why-msps-need-cloud-now">Why MSPs need cloud now</h2>
<p>Enterprise customers increasingly expect cloud capabilities from their MSPs. The traditional MSP offering (managed Microsoft 365, infrastructure management, security operations) is necessary but no longer sufficient. Cloud as part of the offering — branded with the MSP, run on shared or dedicated infrastructure — is what compounds value.</p>
<h2 id="architecture-pattern">Architecture pattern</h2>
<ul>
<li><strong>Multi-tier Tenant CRD</strong> — Ænix → MSP → MSP customers</li>
<li><strong>Per-tier isolation</strong> — RBAC, quotas, observability scope, billing</li>
<li><strong>Branded customer-facing portal</strong> — Cozystack Dashboard customized per MSP</li>
<li><strong>WHMCS-integrated billing</strong> — flows through MSP&rsquo;s customer-management</li>
<li><strong>Service catalog</strong> — MSP can curate (e.g., expose only PostgreSQL, hide Kafka if not supported)</li>
</ul>
<h2 id="reseller-economics">Reseller economics</h2>
<p>For mid-size MSP (50-500 customers):</p>
<ul>
<li><strong>Platform cost</strong> — Ænix engagement + hardware + ops</li>
<li><strong>Per-customer cost</strong> — incremental hardware/storage/bandwidth</li>
<li><strong>MSP customer pricing</strong> — typically 30-50% above raw platform cost</li>
<li><strong>Margin</strong> — covers MSP support, sales, operations</li>
</ul>
<p>Break even at 30-50 paying customers if the platform and its tooling are what you are covering. Budget 50-100 if you are also funding a dedicated on-call rota from day one — that is the same business with a different cost base, not a different answer.</p>
<h2 id="engagement-sequencing">Engagement sequencing</h2>
<ol>
<li><strong>Discovery</strong> — MSP customer base, service catalog priorities</li>
<li><strong>Cozystack pilot</strong> — internal validation</li>
<li><strong>Initial customer cohort</strong> — 5-10 customers with full white-label experience</li>
<li><strong>Operations workflow</strong> — support escalation, SLA management</li>
<li><strong>Scale</strong> — open to broader MSP customer base</li>
</ol>
<p>Total elapsed: 6-12 months.</p>
]]></content:encoded></item><item><title>Launch a customer-facing cloud product — playbook for hosting providers, telcos, and regional operators</title><link>https://aenix.io/blog/2026/05/launch-customer-facing-cloud-product/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/launch-customer-facing-cloud-product/</guid><pubDate>Sun, 17 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>VMware</category><category>Kubernetes</category><category>Sovereignty</category><category>AI and ML</category><category>Multi-tenancy</category><category>Hosting</category><description>The six layers of a customer-facing cloud product, the architectural decisions specific to public cloud, and where launches stumble commercially.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/launch-customer-facing-cloud-product.jpg" alt=""></p><p>Regional and specialty cloud is having a moment in 2026. Hyperscaler economics, sovereignty pressure, and post-Broadcom market dynamics have all opened space for non-hyperscaler cloud products that didn&rsquo;t make sense to launch 5 years ago. Sovereign cloud products from regional providers in the EU, Central Asia and MENA are visible examples. Many more are in stealth or early stages.</p>
<h2 id="why-now">Why now</h2>
<p>Three independent dynamics support new cloud product launches:</p>
<p><strong>Hyperscaler economics breakdown for some workloads.</strong> Sustained inference, high-egress workloads, and regulated workloads are increasingly expensive on hyperscalers relative to dedicated infrastructure. A regional cloud serving these workloads has unit-economic advantages.</p>
<p><strong>Sovereignty as competitive advantage.</strong> EU-member-state, Kazakhstan, several APAC jurisdictions have explicit sovereign-cloud mandates. Regional providers serving these mandates have regulatory moat that hyperscalers cannot access.</p>
<p><strong>Post-Broadcom virtualization market disruption.</strong> VMware-exit-driven workloads have to land somewhere; regional providers can absorb a portion if they ship the right product.</p>
<h2 id="six-layers-of-a-cloud-product">Six layers of a cloud product</h2>
<p>A working customer-facing cloud product has six layers, all of which need engineering:</p>
<h3 id="1-hardware">1. Hardware</h3>
<p>Compute servers, storage, network fabric, datacenter / colocation. Sized for initial customer cohort plus growth headroom.</p>
<h3 id="2-platform">2. Platform</h3>
<p>Multi-tenant Kubernetes-native platform with KubeVirt for VMs, Cilium and Kube-OVN for networking, LINSTOR (DRBD-replicated block) for storage. Cozystack is the open-source default for this pattern.</p>
<h3 id="3-service-catalog">3. Service catalog</h3>
<p>What customers can self-provision: VMs, K8s clusters, managed databases (PostgreSQL, MariaDB, MongoDB, Redis, Valkey, Kafka, ClickHouse, OpenSearch, etc.), S3 buckets, GPU instances, networking primitives.</p>
<h3 id="4-customer-facing-portal">4. Customer-facing portal</h3>
<p>Self-service UI (Cozystack Dashboard or custom). Catalog browsing, provisioning, monitoring, billing visibility.</p>
<h3 id="5-billing">5. Billing</h3>
<p>WHMCS production-ready integration in two modes — shipped by Ænix as the <a href="https://aenix.io/products/whmcs-integration/">WHMCS integration</a> rather than as part of open-source Cozystack. Custom billing for specific markets.</p>
<h3 id="6-operations">6. Operations</h3>
<p>24×7 NOC, customer support, SLA management, observability per tenant, incident response.</p>
<h2 id="architectural-decisions-specific-to-public-cloud-products">Architectural decisions specific to public cloud products</h2>
<p>A public cloud product differs architecturally from an internal platform in several ways:</p>
<p><strong>Multi-tenant isolation must be hard.</strong> Customers don&rsquo;t trust each other; regulatory audits are real. Tenant CRD pattern with strong isolation is necessary, not optional.</p>
<p><strong>Self-service must be polished.</strong> Internal platforms can have &ldquo;ask the platform team&rdquo; as escape hatch. Customer-facing products can&rsquo;t.</p>
<p><strong>Billing must be accurate from day one.</strong> Customers will dispute the first bill; the system has to support that conversation.</p>
<p><strong>Observability for customers, not just for ops.</strong> Customers want their own metrics, not just SLA data.</p>
<p><strong>Compliance posture is the product.</strong> Sovereignty, data residency, audit-readiness are differentiation features, not afterthoughts.</p>
<h2 id="go-to-market-sequencing">Go-to-market sequencing</h2>
<p>Launch in cohorts:</p>
<ol>
<li><strong>Beta with 3-5 friendly customers</strong> — fix the rough edges before paying customers see them.</li>
<li><strong>Limited GA with 10-50 customers</strong> — learn billing and support patterns at small scale.</li>
<li><strong>General availability</strong> — open to general market.</li>
<li><strong>Specialty expansion</strong> — add specific services (more GPU classes, AI services, etc.) based on observed demand.</li>
</ol>
<p>Timing: at provider scale the platform goes live in weeks once hardware is ready, using the productized installer; the beta and limited-GA cohorts then run at your commercial pace. Multi-region national or operator programmes take longer: a 3-6 month pilot, then 9-18 months to full multi-region operation.</p>
<h2 id="where-launches-stumble">Where launches stumble</h2>
<h3 id="stumble-1-under-engineered-multi-tenancy">Stumble 1: under-engineered multi-tenancy</h3>
<p>&ldquo;We&rsquo;ll add proper isolation in v2.&rdquo; Customer #1 finds the gap; reputation damage is hard to recover from.</p>
<h3 id="stumble-2-billing-as-afterthought">Stumble 2: billing as afterthought</h3>
<p>Billing integration in the last month before launch. Doesn&rsquo;t work; customers don&rsquo;t pay; revenue collection is a multi-quarter project.</p>
<h3 id="stumble-3-operational-under-investment">Stumble 3: operational under-investment</h3>
<p>Platform built; ops team sized for 50 customers; they sign 200 in first quarter. Service quality collapses.</p>
<h3 id="stumble-4-copying-hyperscaler-architecture-wholesale">Stumble 4: copying hyperscaler architecture wholesale</h3>
<p>Designed at hyperscaler scale; over-engineered for actual customer count. Operational complexity exceeds revenue.</p>
<h3 id="stumble-5-undifferentiated-commodity-offering">Stumble 5: undifferentiated commodity offering</h3>
<p>Generic cloud product with no differentiation from hyperscaler. Customers default to AWS / Azure / GCP. Specialty / sovereignty / regional matters as differentiator.</p>
<h2 id="ænix-engagement">Ænix engagement</h2>
<p>Ænix has built customer-facing cloud products end-to-end on Cozystack for hosting providers and regional cloud operators. The engagement structure:</p>
<ul>
<li><strong>Free 30-minute discovery call</strong>, then a fixed-price <a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a> (14 days focused or 28 days full)</li>
<li><strong>Platform build</strong> — platform live in weeks once hardware is ready; portal, billing, operations workflows and first-cohort onboarding follow. Multi-region programmes: 3-6 month pilot, then 9-18 months to full multi-region</li>
<li><strong>Ongoing (optional)</strong> — support subscription from the <a href="https://aenix.io/pricing/">published tiers</a>, plus managed services during ramp</li>
</ul>
<p>For details see the <strong><a href="https://aenix.io/services/public-cloud-builder/">public cloud builder service</a></strong> and <strong><a href="https://aenix.io/products/public-cloud-platform/">Ænix Public Cloud Platform</a></strong>.</p>
]]></content:encoded></item><item><title>Industry 4.0 platform — cloud + edge architecture for manufacturing in 2026</title><link>https://aenix.io/blog/2026/05/manufacturing-cloud-industry-40-edge/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/manufacturing-cloud-industry-40-edge/</guid><pubDate>Sun, 17 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>NIS2</category><category>Cozystack</category><category>Sovereignty</category><category>AI and ML</category><category>Compliance</category><description>Industry 4.0 architecture in 2026: edge-to-core patterns, sovereignty for industrial IP, and the NIS2 controls manufacturers are now in scope for.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/manufacturing-cloud-industry-40-edge.jpg" alt=""></p><h2 id="what-industry-40-actually-means-in-2026">What Industry 4.0 actually means in 2026</h2>
<p>The term has accumulated marketing weight. Practically, Industry 4.0 manufacturing means:</p>
<ul>
<li>IoT instrumentation across production floors</li>
<li>Real-time data collection from machinery</li>
<li>AI-driven quality control and predictive maintenance</li>
<li>Digital twins of production lines</li>
<li>Supply-chain integration across systems</li>
<li>OT/IT convergence</li>
</ul>
<p>All of this requires infrastructure — typically a mix of edge compute (close to machinery), regional/HQ cloud (for analytics, ML, and integration), and integration with existing enterprise systems.</p>
<h2 id="architectural-pattern">Architectural pattern</h2>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">HQ Cloud (Cozystack)
</span></span><span class="line"><span class="cl">   ├── Analytics platform
</span></span><span class="line"><span class="cl">   ├── ML training
</span></span><span class="line"><span class="cl">   ├── Enterprise integration
</span></span><span class="line"><span class="cl">   └── Data warehouse
</span></span><span class="line"><span class="cl">        ↓ (data flow)
</span></span><span class="line"><span class="cl">Regional sites (Cozystack)
</span></span><span class="line"><span class="cl">   ├── Regional aggregation
</span></span><span class="line"><span class="cl">   ├── Production planning
</span></span><span class="line"><span class="cl">   └── Quality systems
</span></span><span class="line"><span class="cl">        ↓
</span></span><span class="line"><span class="cl">Production-floor edge (Cozystack)
</span></span><span class="line"><span class="cl">   ├── Real-time control
</span></span><span class="line"><span class="cl">   ├── IoT data ingestion
</span></span><span class="line"><span class="cl">   ├── Local AI inference
</span></span><span class="line"><span class="cl">   └── OT/IT interface
</span></span></code></pre></div><p>Cozystack runs at all three layers with consistent operational model.</p>
<h2 id="sovereignty-for-industrial-ip">Sovereignty for industrial IP</h2>
<p>Industrial IP — design data, formulations, process specifications — has higher sovereignty requirements than typical enterprise data. The architectural answer is an air-gap-capable platform with opt-in volume encryption at rest (LINSTOR and LUKS) and a passphrase the manufacturer holds; the key-management process is designed with you.</p>
<h2 id="nis2-compliance">NIS2 compliance</h2>
<p>Manufacturing of critical products (medical devices, computers, electronic equipment, machinery, motor vehicles) is in NIS2 scope. Architectural implications same as broader NIS2 — see <strong><a href="https://aenix.io/solutions/nis2-compliance/">NIS2 compliance</a></strong>.</p>
]]></content:encoded></item><item><title>Production Kubernetes cluster setup — architecture decisions, sizing, and operations in 2026</title><link>https://aenix.io/blog/2026/05/kubernetes-cluster-setup-production-architecture/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/kubernetes-cluster-setup-production-architecture/</guid><pubDate>Sat, 16 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>OpenShift</category><category>Kubernetes</category><category>Cozystack</category><category>KubeVirt</category><category>Talos</category><category>Sovereignty</category><description>Ten architecture decisions behind a production Kubernetes cluster — distribution, tenancy, CNI, storage, GitOps, DR — and the readiness failures that recur.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/kubernetes-cluster-setup-production-architecture.jpg" alt=""></p><p>Most Kubernetes-setup tutorials get you a cluster. They don&rsquo;t get you a production-ready cluster. The difference is the dozen architectural decisions that don&rsquo;t show up in a &ldquo;kubectl apply&rdquo; workflow but that determine whether your cluster operates well at scale or becomes a permanent maintenance burden.</p>
<h2 id="decision-1-distribution-choice">Decision 1: distribution choice</h2>
<p>The question &ldquo;which Kubernetes distribution&rdquo; looks like a tooling decision. It&rsquo;s actually an architectural decision because it determines:</p>
<ul>
<li>Operational model (managed vs self-managed control plane)</li>
<li>Multi-tenancy approach</li>
<li>Storage and networking integration</li>
<li>Vendor relationship (open-source vs commercial)</li>
<li>Long-term cost trajectory</li>
</ul>
<p>Common 2026 options:</p>
<ul>
<li><strong>Vanilla Kubernetes (kubeadm / Cluster API)</strong> — maximum control, maximum operational responsibility. Right for organizations with strong platform-engineering capability.</li>
<li><strong>Cozystack</strong> — open-source, CNCF Project, multi-tenant + virtualization (KubeVirt). Right for service providers, sovereign-cloud builders, regulated multi-tenant.</li>
<li><strong>OpenShift / OKD</strong> — Red Hat commercial / community. Right for enterprises already on Red Hat with established procurement relationship.</li>
<li><strong>Rancher / RKE2</strong> — SUSE-backed, multi-cluster management. Right for distributed cluster fleets.</li>
<li><strong>Talos</strong> — minimal Linux + Kubernetes only. Right as the OS layer underneath Cozystack or vanilla Kubernetes.</li>
<li><strong>Hyperscaler-managed (EKS / AKS / GKE)</strong> — control plane managed; you operate workloads. Right when sovereignty isn&rsquo;t a constraint and operational simplicity beats cost.</li>
</ul>
<p>Distribution choice is structural; getting it wrong is expensive to undo.</p>
<h2 id="decision-2-multi-tenancy">Decision 2: multi-tenancy</h2>
<p>Kubernetes ships with namespaces, not multi-tenancy. The question is how to bridge the gap.</p>
<p>Patterns:</p>
<ul>
<li><strong>Soft multi-tenancy (namespaces + RBAC)</strong> — single cluster, multiple namespaces, RBAC to separate. Simplest, least isolation. Acceptable for trusting tenants (BUs of one organization).</li>
<li><strong>Hard multi-tenancy (Tenant CRD)</strong> — Kubernetes-native tenant abstraction, with namespace, quotas, RBAC, observability scope, audit trail. Right for service-provider model or regulated multi-tenancy. <strong>Cozystack</strong> ships this out of the box.</li>
<li><strong>Cluster per tenant</strong> — physical isolation. Highest assurance, highest operational cost. Right for customers requiring full isolation (high-stakes regulated, classified).</li>
</ul>
<p>The choice depends on the trust model and isolation requirements. Most multi-tenant deployments in 2026 use Tenant CRD pattern — strong isolation without per-tenant cluster overhead.</p>
<h2 id="decision-3-networking-cni">Decision 3: networking (CNI)</h2>
<p>CNI choice in 2026 has converged on:</p>
<ul>
<li><strong>Cilium</strong> — eBPF-based, network policies + observability + service mesh in one. Default for new deployments. CNCF Graduated.</li>
<li><strong>Calico</strong> — long-standing, BGP-friendly, strong on network policies.</li>
<li><strong>Flannel</strong> — simpler, less feature-rich.</li>
<li><strong>Cloud-provider CNIs</strong> — AWS VPC CNI, Azure CNI, GCP — when running in cloud.</li>
</ul>
<p>For production multi-tenant deployments, Cilium is increasingly the default — its network policies integrate with the multi-tenancy model and observability is native.</p>
<h2 id="decision-4-storage">Decision 4: storage</h2>
<p>Storage in Kubernetes is a separate operational discipline. Options:</p>
<ul>
<li><strong>LINSTOR (DRBD)</strong> — replicated block storage. Production-grade for stateful workloads. Cozystack default.</li>
<li><strong>Rook + Ceph</strong> — object + block + file storage. Heavier operational footprint; powerful when justified.</li>
<li><strong>Longhorn</strong> — Rancher-led replicated block storage. Easier than Ceph, less mature than LINSTOR.</li>
<li><strong>Vendor SAN / hyperconverged</strong> — VMware vSAN, NetApp, Pure Storage. Operational handoff to vendor; cost ceiling.</li>
<li><strong>Cloud-provider storage</strong> — EBS, Azure Disks, GCP PD — when cloud-managed.</li>
</ul>
<p>For production stateful workloads on bare metal, LINSTOR/DRBD are the realistic choices. Cozystack ships LINSTOR; it&rsquo;s been validated in production.</p>
<h2 id="decision-5-identity-and-secrets">Decision 5: identity and secrets</h2>
<p>Workload identity in Kubernetes:</p>
<ul>
<li><strong>SPIFFE/SPIRE</strong> — open-source, workload identity standard. Increasingly default for service-to-service auth.</li>
<li><strong>Cloud-provider workload identity</strong> — AWS IAM Roles for Service Accounts, Azure AD Workload Identity, GCP Workload Identity. Locks identity to the cloud provider.</li>
<li><strong>OIDC + service accounts</strong> — Kubernetes-native, federated with workforce identity (Keycloak, Okta).</li>
</ul>
<p>Secrets management:</p>
<ul>
<li><strong>External Secrets Operator</strong> + cloud-provider secret stores (AWS Secrets Manager, Azure Key Vault).</li>
<li><strong>HashiCorp Vault</strong> — for organizations already on Vault.</li>
<li><strong>Sealed Secrets</strong> — for GitOps-friendly encrypted-at-rest secrets in Git.</li>
</ul>
<p>The right combination depends on your existing identity infrastructure.</p>
<h2 id="decision-6-gitops-engine">Decision 6: GitOps engine</h2>
<p>Two production-grade options:</p>
<ul>
<li><strong>Argo CD</strong> — UI-rich, plugin ecosystem, multi-tenancy via projects.</li>
<li><strong>Flux</strong> — closer to upstream Kubernetes, Helm-controller-native, lighter operational footprint.</li>
</ul>
<p>Both are CNCF Graduated. Cozystack uses Flux. The choice depends on team familiarity and operational style preferences.</p>
<h2 id="decision-7-observability-stack">Decision 7: observability stack</h2>
<p>Three serious options for self-hosted:</p>
<ul>
<li><strong>VictoriaMetrics + VictoriaLogs</strong> — low-overhead, single-binary, performant at scale. Cozystack default.</li>
<li><strong>Prometheus + Loki</strong> — Grafana ecosystem standard. More memory-intensive; complex at scale.</li>
<li><strong>Commercial (Datadog, New Relic, Splunk)</strong> — SaaS, sovereignty implications.</li>
</ul>
<p>For sovereignty-sensitive or cost-sensitive deployments, VictoriaMetrics + VictoriaLogs has been our default recommendation since ~2024.</p>
<h2 id="decision-8-backup-and-dr">Decision 8: backup and DR</h2>
<ul>
<li><strong>Velero</strong> — Kubernetes-native backup, restore to any cluster. CNCF Project.</li>
<li><strong>Stash / Kanister</strong> — backup operators with stronger app-aware features.</li>
<li><strong>Cloud-provider backup</strong> — managed backup services on hyperscaler clusters.</li>
</ul>
<p>For DR specifically: per-app patterns (PostgreSQL PITR, etc.) on top of Velero infrastructure is the production pattern.</p>
<h2 id="decision-9-ingress-and-load-balancing">Decision 9: ingress and load balancing</h2>
<ul>
<li><strong>Cloud LB + ingress controller (NGINX / Traefik / Contour)</strong> — when running in cloud.</li>
<li><strong>MetalLB + ingress controller</strong> — bare metal layer-2 load balancing.</li>
<li><strong>Cilium L7 load balancing</strong> — when Cilium is the CNI.</li>
<li><strong>Service mesh (Istio / Linkerd)</strong> — for advanced traffic management.</li>
</ul>
<p>Bare-metal deployments default to MetalLB + Cilium today.</p>
<h2 id="decision-10-lifecycle-management">Decision 10: lifecycle management</h2>
<p>How clusters get upgraded, scaled, and recovered:</p>
<ul>
<li><strong>Cluster API</strong> — Kubernetes-native cluster lifecycle. Used by Cozystack.</li>
<li><strong>Kubeadm + manual upgrades</strong> — works for small fleet.</li>
<li><strong>Vendor-managed lifecycle</strong> — OpenShift, Rancher, vendor-led.</li>
<li><strong>Hyperscaler-managed</strong> — control plane upgrades automatic.</li>
</ul>
<p>For production fleets &gt;5 clusters, Cluster API or vendor-managed lifecycle pay back the operational cost.</p>
<h2 id="operational-practices-that-matter">Operational practices that matter</h2>
<p>Beyond architecture, the practices that determine whether a cluster operates well:</p>
<h3 id="slos-and-error-budgets">SLOs and error budgets</h3>
<p>Define cluster-level SLOs (control plane availability, ingress availability, etc.). Define workload-level SLOs per service. Use error budgets to drive prioritization.</p>
<h3 id="documented-runbooks">Documented runbooks</h3>
<p>Common operational scenarios documented: pod failure, node failure, control-plane recovery, certificate rotation, storage failure, network failure. Practice them.</p>
<h3 id="capacity-planning">Capacity planning</h3>
<p>Compute, memory, and storage capacity tracked. Forecasted against workload growth. Decisions made before resource exhaustion.</p>
<h3 id="upgrade-discipline">Upgrade discipline</h3>
<p>Kubernetes minor versions every quarter; cluster upgrades quarterly or twice-yearly. Skip-version upgrades are not supported by upstream — keep current.</p>
<h3 id="incident-response">Incident response</h3>
<p>On-call rotation, incident commander role, blameless post-mortems with documented action items.</p>
<h3 id="security-posture">Security posture</h3>
<p>Pod Security Standards enforced; network policies as default-deny; secrets in cloud secret stores or Vault; image scanning in CI; runtime threat detection where applicable.</p>
<h2 id="common-production-readiness-failures">Common production-readiness failures</h2>
<h3 id="failure-1-dev-cluster-scaled-to-prod">Failure 1: dev cluster scaled to prod</h3>
<p>Cluster started for dev; workloads moved to prod without re-architecting. Falls over the first time prod load arrives.</p>
<h3 id="failure-2-no-observability-budget">Failure 2: no observability budget</h3>
<p>Prometheus runs out of memory at production scale; metrics retention dropped to 7 days; observability becomes useless. Should have been planned upfront.</p>
<h3 id="failure-3-security-as-afterthought">Failure 3: security as afterthought</h3>
<p>Pod Security Standards not enforced; network policies absent; secrets in plain ConfigMaps. Audit findings drive emergency remediation.</p>
<h3 id="failure-4-upgrade-debt">Failure 4: upgrade debt</h3>
<p>Kubernetes 3-4 minor versions behind; upgrades are big-bang projects; technical debt compounds.</p>
<h3 id="failure-5-no-platform-team">Failure 5: no platform team</h3>
<p>Cluster is &ldquo;owned&rdquo; by everyone, operated by no one. Drift accumulates; nobody can confidently change anything.</p>
<h2 id="how-to-assess-what-you-have">How to assess what you have</h2>
<p>Before building or scaling, an architecture review is the cheapest insurance. The output is a written assessment of where you stand, where the gaps are, and what production-readiness looks like for your scale.</p>
<p>Ænix runs Kubernetes architecture reviews as part of the fixed-price <strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong>: 14 days for a focused review, 28 days for the full assessment.</p>
<h2 id="want-to-dig-deeper">Want to dig deeper?</h2>
<ul>
<li><strong><a href="https://aenix.io/services/kubernetes-consulting/">Kubernetes consulting services</a></strong> — engagement details</li>
<li><strong><a href="https://aenix.io/services/platform-engineering/">Platform engineering services</a></strong> — broader scope</li>
<li><strong><a href="https://aenix.io/services/internal-developer-platform/">Internal developer platform</a></strong> — IDP layer</li>
<li><strong><a href="https://aenix.io/products/cozystack/">Cozystack</a></strong> — open-source platform foundation</li>
</ul>
]]></content:encoded></item><item><title>When Public Cloud Platform pays back for hosting providers</title><link>https://aenix.io/blog/2026/05/isp-edition-economics-hosting-providers/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/isp-edition-economics-hosting-providers/</guid><pubDate>Fri, 15 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Hosting</category><category>Cozystack</category><category>Multi-tenancy</category><category>Platform Engineering</category><category>Cloud</category><description>Unit economics of Ænix Public Cloud Platform for hosting providers: ARPU, infrastructure cost per tenant, platform-team capacity, payback, and where it breaks.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/isp-edition-economics-hosting-providers.jpg" alt=""></p><p>Most &ldquo;should we build our own cloud product?&rdquo; conversations at hosting
providers stop at the technology question. The harder question is the
unit economics: what does it cost per tenant, what&rsquo;s a realistic ARPU,
how many tenants until break-even, and where does the model fail.</p>
<p>This article is the working version of that conversation. It assumes
the technology decision is settled (the Cozystack-based Ænix Public
Cloud Platform) and focuses on whether the economics fit <em>your</em> hosting
business — not the abstract one.</p>
<h2 id="what-public-cloud-platform-actually-delivers">What Public Cloud Platform actually delivers</h2>
<p>Before economics, scope. Public Cloud Platform is a complete public-cloud
product Ænix sells to hosting providers, MSPs, regional clouds, and small-to-
mid data centres. It includes:</p>
<ul>
<li><strong>Multi-tenant Cozystack platform</strong> running on customer-controlled
bare metal (KubeVirt + Cilium + Kube-OVN + LINSTOR + Tenant CRD).</li>
<li><strong>Cozystack Dashboard</strong> — customer-facing self-service portal, brandable to
your hosting brand.</li>
<li><strong>WHMCS integration</strong> — billing flows through the customer-management
system most hosting providers already operate.</li>
<li><strong>Service catalog</strong> — VMs, tenant Kubernetes clusters, managed
databases (PostgreSQL, MariaDB, MongoDB, Redis, Valkey, Kafka, ClickHouse, etc.), S3-compatible
object storage, GPU services. Curatable per provider.</li>
<li><strong>Tenant lock / suspension</strong> — operational hooks for non-payment
and policy enforcement.</li>
<li><strong>Migration tooling</strong> — productized patterns for VMware, OpenStack,
Virtuozzo, Proxmox sources.</li>
</ul>
<p>What it is <em>not</em>: a hyperscaler. It is a sovereign, multi-tenant cloud
product for hosting providers who want to compete on regional presence,
sovereignty, and pricing flexibility — not on hyperscaler-scale catalog
depth.</p>
<h2 id="pricing-model">Pricing model</h2>
<p>Public Cloud Platform subscriptions use the published support tiers,
priced per 10 physical nodes per month on annual billing: <strong>Basic
$1,250</strong>, <strong>Standard $3,000</strong>, <strong>Plus $5,500</strong>; Enterprise is quoted
individually. Higher tiers add faster response times, unlimited
incidents, 24×7 support (Plus and Enterprise) and a wider support
scope — the full matrix is on <a href="https://aenix.io/pricing/">/pricing/</a>. Ænix does not charge per VM,
per CPU, or per GB — the Cozystack platform itself is free under
Apache 2.0; what you pay for is engagement, support, and operational
assurance. Every support tier includes the proprietary Ænix commercial
modules (billing system and WHMCS integration).</p>
<p>For a typical mid-size hosting provider running 30-100 customer-facing
nodes, that is $3,750-12,500/month on Basic or $9,000-30,000/month on
Standard at list price. Compare it with the recurring licence and
subscription cost you pay VMware or an OpenStack distribution vendor
today.</p>
<h2 id="unit-economics--per-tenant-view">Unit economics — per-tenant view</h2>
<p>The economics question every hosting CFO asks: <em>what does it cost us to
serve one tenant, and what can we realistically charge?</em></p>
<h3 id="cost-per-tenant">Cost per tenant</h3>
<p>Infrastructure cost per tenant is dominated by the underlying compute,
storage, and bandwidth, not by Cozystack itself. Cozystack overhead is
~5-10% of node capacity (typical Kubernetes-platform overhead, well
within acceptable for production). For a tenant consuming roughly:</p>
<ul>
<li>2 vCPU</li>
<li>4 GB RAM</li>
<li>50 GB block storage (replicated 3×)</li>
<li>100 GB egress / month</li>
</ul>
<p>Direct infrastructure cost (amortised hardware + colocation + bandwidth)
in 2026 European pricing is typically €15-30/month. Cozystack platform
overhead allocation (the platform team&rsquo;s salary spread across all
tenants) on a mid-size provider running 500 tenants adds another €5-10.</p>
<p>So <strong>all-in cost per typical tenant: €20-40/month</strong> at the lower end of
the resource consumption profile.</p>
<h3 id="what-you-can-charge">What you can charge</h3>
<p>Hosting-provider ARPU for a comparable resource profile in 2026 EU
markets typically lands in €40-80/month. Higher in DACH and Western
Europe, lower in Central / Eastern Europe and Central Asia. Add managed
services (managed PostgreSQL, managed S3, GPU access) and ARPU lifts
to €80-200+/month per tenant.</p>
<p>That puts margin at roughly <strong>2-3× cost</strong> at the low end, <strong>4-6× cost</strong>
on managed-service-heavy tenants. Not hyperscaler-margins; not VMware
reseller margins either. Closer to traditional hosting margins in the
post-Broadcom 2026 reality.</p>
<h2 id="break-even-math">Break-even math</h2>
<p>The other CFO question: <em>how many customers until we make money?</em></p>
<p>The fixed cost stack for a mid-size hosting provider on Public Cloud Platform:</p>
<table>
  <thead>
      <tr>
          <th>Item</th>
          <th>Monthly</th>
          <th>Annual</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Ænix support (Standard tier, 50 nodes = 5 × $3,000)</td>
          <td>$15k</td>
          <td>$180k</td>
      </tr>
      <tr>
          <td>Platform operations (ISP calculator default model: about 1.3 FTE at 10 nodes, about 2.6 at 40; more at 50 nodes and for 24×7 on-call)</td>
          <td>€20-35k</td>
          <td>€240-420k</td>
      </tr>
      <tr>
          <td>Hardware amortisation (50 nodes)</td>
          <td>€5-8k</td>
          <td>€60-100k</td>
      </tr>
      <tr>
          <td>Colocation / power / bandwidth</td>
          <td>€4-7k</td>
          <td>€50-85k</td>
      </tr>
      <tr>
          <td>Customer support team, separate from platform operations (2-4 FTE for cloud customers)</td>
          <td>€10-20k</td>
          <td>€120-240k</td>
      </tr>
      <tr>
          <td>Marketing / sales</td>
          <td>€5-15k</td>
          <td>€60-180k</td>
      </tr>
  </tbody>
</table>
<p><strong>Total monthly fixed: €44-85k, plus $15k for Ænix support.</strong></p>
<p>With €25-50/month margin per tenant (€40-80 ARPU after €15-30 direct
infrastructure cost), break-even sits at <strong>roughly 1,200-4,000 paying tenants</strong> depending on ARPU
mix and where you are in the salary band. This is the full-programme
case with a dedicated team; a 10-node start on the Basic tier with
existing staff breaks even far earlier — model your own numbers in the
<a href="https://aenix.io/isp-calculator/">ISP calculator</a>.</p>
<p>For providers currently running ~500 customers on legacy infrastructure
who are evaluating the move, this matters: you need a credible path to
at least double tenant count within 18-24 months for the economics to actually
work. Without growth, Public Cloud Platform is a cost reduction (modest) but not
a transformation.</p>
<p>For providers below ~300 customers, a full programme with a dedicated
platform team is often <em>premature</em> — that fixed cost
overwhelms the margin contribution. A smaller start (10 nodes on the
Basic tier, a narrow catalogue, existing staff) is usually the better
first step. We&rsquo;ll say so in a discovery call rather than push a larger
engagement.</p>
<h2 id="where-the-model-breaks">Where the model breaks</h2>
<p>Three failure patterns recur:</p>
<h3 id="1-customer-facing-portal-under-investment">1. Customer-facing portal under-investment</h3>
<p>Hosting providers historically compete on price and reliability.
The Cozystack Dashboard out of the box is functional but generic; differentiation
comes from polish (UX flows that match how <em>your</em> customers think about
ordering, configuring, paying). Providers who treat the portal as
&ldquo;good enough&rdquo; lose conversion to providers who invest in it.</p>
<p>Ænix engagement includes Cozystack Dashboard brand customization; deeper UX
work is typically a separate Phase 2.</p>
<h3 id="2-service-catalog-mismatch">2. Service-catalog mismatch</h3>
<p>Cozystack offers 20+ managed services; not all of them fit every
provider&rsquo;s customer base. Exposing all of them without operational
backing means customers ordering Kafka or ClickHouse and discovering
the provider can&rsquo;t really support them. Curate the catalog to what
you can actually operate at the SLA you promise. Service rollout in
cohorts is the standard playbook.</p>
<h3 id="3-operations-team-under-staffed-for-growth">3. Operations team under-staffed for growth</h3>
<p>The biggest single failure mode in our pipeline: Public Cloud Platform deployed,
launches successfully, signs 200 customers in the first quarter — and
then the 4-person operations team that worked at 50 customers can&rsquo;t
scale. Customer support response time degrades, SLA breaches multiply,
churn picks up.</p>
<p>Plan operations team size for 18-month-out customer count, not current.
Hire ahead.</p>
<h2 id="how-public-cloud-platform-compares-to-alternatives-for-hosting-providers">How Public Cloud Platform compares to alternatives for hosting providers</h2>
<p><strong>Versus VMware Cloud Director (vCD):</strong></p>
<p>vCD is the historical incumbent for hosting providers. Post-Broadcom,
subscription pricing has reshaped the math — 2-5× increases on
renewal, mandatory VCF bundling, end of perpetual licensing. For most
providers running vCD today, the renewal cycle is the trigger.
Ænix Public Cloud Platform goes live in weeks once hardware is ready;
moving an existing vCD estate onto it is a separate migration project,
sized by estate in the assessment.</p>
<p><strong>Versus OpenStack:</strong></p>
<p>OpenStack remains valid for providers with deep OpenStack expertise
and large-scale (&gt;500 nodes) deployments where operational complexity
is amortised. For mid-size providers, OpenStack&rsquo;s operational footprint
(50+ services, distinct upgrade lifecycles per component) overshoots
what the team can sustain. Public Cloud Platform substantially smaller surface.</p>
<p><strong>Versus building it yourself on vanilla Kubernetes + KubeVirt + Helm:</strong></p>
<p>This is the credible alternative for providers with strong platform
engineering capacity. Trade-off: 12-24 months of build time + ongoing
maintenance versus a turnkey deployment. We&rsquo;ve seen both work; the
build-it-yourself path is the right choice when you have a 10+ engineer
platform team and the components match your specific operational
preferences. For the typical mid-size provider with a 3-5 engineer
platform team, Public Cloud Platform wins on time-to-market and operational
predictability.</p>
<p><strong>Versus a hyperscaler-managed cloud product (white-label):</strong></p>
<p>Hyperscalers (AWS, Azure, GCP) sometimes offer hosting partners
white-label or co-branded cloud arrangements. Trade-off: lower
operational complexity for the provider, but per-customer margin is
typically lower, and sovereignty positioning is weaker (the provider
still depends on the hyperscaler, which European customers
increasingly view as a structural risk).</p>
<h2 id="when-public-cloud-platform-is-the-right-answer">When Public Cloud Platform is the right answer</h2>
<p>It fits when at least three of the following hold:</p>
<ol>
<li><strong>You operate today on bare metal or commercial hypervisor with
recurring licence pressure</strong> — VMware, OpenStack, Virtuozzo, or
commercial KVM distribution.</li>
<li><strong>You have direct customer relationships you can monetise</strong> — not
pure reseller of someone else&rsquo;s cloud.</li>
<li><strong>You have or can hire a 3-5 engineer platform team</strong> — Cozystack
needs operational ownership.</li>
<li><strong>You target regional, regulated, or sovereignty-sensitive
customers</strong> — where European-EU positioning matters.</li>
<li><strong>Your customer count is 300+ today or you have credible growth
path to 1,000+</strong> — for fixed-cost amortisation.</li>
<li><strong>You&rsquo;re willing to invest in customer-facing portal polish</strong> —
not just treat Cozystack Dashboard as &ldquo;good enough.&rdquo;</li>
</ol>
<p>Fewer than three: usually a different answer is better — staying on
existing infrastructure with cost optimisation, partnering with a
larger sovereign provider as their channel, or going hyperscaler-
managed-cloud-product as a smaller-margin route.</p>
<h2 id="engagement-structure">Engagement structure</h2>
<p>For providers where Public Cloud Platform fits:</p>
<ul>
<li><strong>Discovery call</strong> (30 min, free)</li>
<li><strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong>
(fixed price, 14 days focused or 28 days full) — current estate
inventory, target architecture, migration plan</li>
<li><strong>Platform live and pilot cohort</strong> — the platform goes live in weeks
on your hardware with the productized installer; 5-10 friendly
customers migrated, billing validated</li>
<li><strong>Limited GA</strong> (2-4 months) — 50-100 customers, operational
workflows stabilised</li>
<li><strong>General availability</strong> — open market launch</li>
<li><strong>Support subscription</strong> (ongoing) — one of the <a href="https://aenix.io/pricing/">published tiers</a>;
Plus or Enterprise for 24×7 coverage</li>
</ul>
<p>How long the commercial launch takes after the platform is live
depends on migration scope and team readiness. Multi-region national
or operator programmes follow a 3-6 month pilot, then 9-18 months to
full multi-region operation.</p>
<h2 id="where-to-dig-deeper">Where to dig deeper</h2>
<ul>
<li><strong><a href="https://aenix.io/products/public-cloud-platform/">Public Cloud Platform landing page</a></strong> —
feature list, pricing block, FAQ</li>
<li><strong><a href="https://aenix.io/industries/hosting-providers/">Hosting providers industry page</a></strong> —
hosting-provider-specific positioning</li>
<li><strong><a href="https://aenix.io/services/white-label-cloud/">White-label cloud services</a></strong> —
for MSP / channel-partner extensions of the model</li>
</ul>
]]></content:encoded></item><item><title>K-12 school district cloud infrastructure — when sovereignty matters more than convenience</title><link>https://aenix.io/blog/2026/05/k12-school-district-cloud-infrastructure/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/k12-school-district-cloud-infrastructure/</guid><pubDate>Fri, 15 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Kubernetes</category><category>Cozystack</category><category>Sovereignty</category><category>AI and ML</category><category>Multi-tenancy</category><category>Compliance</category><description>Why K-12 infrastructure differs from universities, when a district actually needs sovereign infrastructure, and the architecture pattern that fits.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/k12-school-district-cloud-infrastructure.jpg" alt=""></p><h2 id="why-k-12-is-different-from-universities">Why K-12 is different from universities</h2>
<p>Universities have research computing, AI/ML labs, and curriculum-driven cloud-native education needs. K-12 districts:</p>
<ul>
<li>Don&rsquo;t run research computing</li>
<li>Don&rsquo;t teach Kubernetes (mostly)</li>
<li>Have student-data privacy as primary infrastructure concern</li>
<li>Operate at consumer-facing scale (10K-100K+ students)</li>
<li>Have long budget cycles (3-5 year procurement)</li>
</ul>
<p>These different drivers mean different architectural answers.</p>
<h2 id="when-k-12-needs-sovereign-infrastructure">When K-12 needs sovereign infrastructure</h2>
<p>Three scenarios:</p>
<ol>
<li><strong>Regulatory pressure</strong> — some EU member states require student-data residency under national privacy regulations</li>
<li><strong>Liability / public concern</strong> — high-profile districts that publicly committed to local data residency</li>
<li><strong>EdTech platform development</strong> — districts building their own learning platforms</li>
</ol>
<p>Most K-12 districts fall in none of these — hyperscaler-managed services + standard EdTech tools is right.</p>
<h2 id="architecture-pattern-for-fitting-k-12">Architecture pattern for fitting K-12</h2>
<ul>
<li><strong>District-tier cluster</strong> — Cozystack at central district IT</li>
<li><strong>Per-school isolation</strong> — Tenant CRD per school</li>
<li><strong>Sovereign by architecture</strong> — student data on customer hardware, opt-in volume encryption with a passphrase the district holds</li>
<li><strong>Standard EdTech integrations</strong> — Google Classroom / Microsoft 365 federation</li>
</ul>
<p>For consortia (multi-district shared platform):</p>
<ul>
<li>Federated multi-district platform</li>
<li>Shared core, per-district isolation</li>
<li>Joint procurement, distributed operations</li>
</ul>
<h2 id="common-pitfalls-in-k-12-platform-projects">Common pitfalls in K-12 platform projects</h2>
<ul>
<li>Underestimating EdTech vendor integration complexity</li>
<li>Skipping FERPA / GDPR audit-readiness</li>
<li>Vendor-led &ldquo;education cloud&rdquo; with lock-in</li>
<li>Mid-cycle re-architecture due to budget cycle mismatch</li>
</ul>
]]></content:encoded></item><item><title>Internal developer portal vs internal developer platform — and Backstage's place in 2026</title><link>https://aenix.io/blog/2026/05/internal-developer-portal-vs-platform/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/internal-developer-portal-vs-platform/</guid><pubDate>Thu, 14 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Backstage</category><category>Kubernetes</category><category>Platform Engineering</category><category>Compliance</category><category>Observability</category><description>Portal and platform are not the same thing. Where Backstage actually fits, what the alternatives are, and how to decide whether you need a portal at all.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/internal-developer-portal-vs-platform.jpg" alt=""></p><p>The &ldquo;IDP&rdquo; acronym is overloaded. It means both:</p>
<ul>
<li><strong>Internal developer platform</strong> — the underlying capability stack (Kubernetes, IaC, observability, golden paths).</li>
<li><strong>Internal developer portal</strong> — the UI/catalog layer (Backstage, Port, custom).</li>
</ul>
<p>Most discussions confuse these. The actual question — &ldquo;do we need Backstage?&rdquo; — has different answers depending on which IDP you&rsquo;re talking about.</p>
<h2 id="platform-vs-portal">Platform vs portal</h2>
<p>A platform without a portal still works. A portal without a platform is wallpaper.</p>
<p><strong>Platform</strong> answers: &ldquo;What capabilities can product teams self-serve?&rdquo;</p>
<ul>
<li>Environment provisioning</li>
<li>Application deployment</li>
<li>Database / queue / cache provisioning</li>
<li>Observability onboarding</li>
<li>Secrets, identity, network access</li>
</ul>
<p><strong>Portal</strong> answers: &ldquo;How do product teams discover and access those capabilities?&rdquo;</p>
<ul>
<li>Service catalog</li>
<li>Documentation entry point</li>
<li>Self-service forms / actions</li>
<li>Cost / SLO dashboards per service</li>
</ul>
<p>The platform is required for self-service to work. The portal is required for discoverability when service count grows past what teams can hold in their heads.</p>
<h2 id="when-you-need-a-portal">When you need a portal</h2>
<p>A portal becomes valuable when:</p>
<ul>
<li><strong>Service count is large</strong> — 50+ services where engineers can&rsquo;t navigate all of them.</li>
<li><strong>Team size is large</strong> — many engineers, many of whom aren&rsquo;t yet platform-fluent.</li>
<li><strong>Cross-team service discovery is real</strong> — engineers from team A consume services from team B.</li>
<li><strong>Compliance / audit requires service catalog</strong> — regulator-ready service inventory.</li>
</ul>
<p>When none of these hold (small org, small service count, fluent platform team), a portal adds maintenance cost without much benefit.</p>
<h2 id="portal-options-compared">Portal options compared</h2>
<h3 id="backstage-cncf-incubating">Backstage (CNCF Incubating)</h3>
<p><strong>What:</strong> Open-source service catalog + plugin ecosystem. Originally built by Spotify, now CNCF.</p>
<p><strong>Strengths:</strong> Mature, broad plugin ecosystem, customizable, strong community.</p>
<p><strong>Weaknesses:</strong> Operational cost is real — running Backstage means maintaining plugins, integrating with internal tools, evolving with Backstage&rsquo;s release cadence. Underestimated by many adopters.</p>
<p><strong>Best for:</strong> Mid-to-large engineering orgs (200+ engineers) with platform-team capacity to maintain Backstage as a product.</p>
<h3 id="port-portio">Port (port.io)</h3>
<p><strong>What:</strong> SaaS internal developer portal.</p>
<p><strong>Strengths:</strong> No self-hosting; quicker to deploy than Backstage; opinionated.</p>
<p><strong>Weaknesses:</strong> SaaS dependency (sovereignty implications); less customizable than Backstage.</p>
<p><strong>Best for:</strong> Mid-size orgs willing to use SaaS, wanting to skip Backstage&rsquo;s operational cost.</p>
<h3 id="cortex--compass--custom-variants">Cortex / Compass / custom variants</h3>
<p>Adjacent options with different trade-offs. Cortex is engineering-effectiveness-focused; Compass is Atlassian-native; custom is &ldquo;build it yourself.&rdquo;</p>
<h3 id="no-portal">No portal</h3>
<p><strong>What:</strong> IaC repository + Markdown documentation + GitOps interface.</p>
<p><strong>Strengths:</strong> Zero operational cost beyond Git. No portal-specific maintenance.</p>
<p><strong>Weaknesses:</strong> Discoverability scales poorly past 30-50 services.</p>
<p><strong>Best for:</strong> Smaller orgs (under ~100 engineers) where the portal cost would exceed its value.</p>
<h2 id="backstages-actual-best-fit">Backstage&rsquo;s actual best fit</h2>
<p>Backstage works well when:</p>
<ul>
<li>200+ engineers, 100+ services</li>
<li>Platform team has 5+ engineers who can dedicate time to maintenance</li>
<li>Plugin ecosystem matches your tooling</li>
<li>Customization is a priority (you&rsquo;ll write or extend plugins)</li>
</ul>
<p>When those don&rsquo;t hold — the operational cost overshoots the value, and a lighter-weight option fits better.</p>
<h2 id="cnoe--the-open-source-platform-engineering-reference">CNOE — the open-source platform-engineering reference</h2>
<p>The CNOE (Cloud Native Operational Excellence) industry initiative is worth mentioning. It&rsquo;s an opinionated reference architecture combining Backstage, Argo CD, Crossplane, External Secrets, and other CNCF tools into a coherent platform pattern. For organizations that want a &ldquo;platform-in-a-box&rdquo; using CNCF projects, CNOE is the structured starting point.</p>
<p>CNOE is complementary to Cozystack: CNOE focuses on the developer portal + tooling layer; Cozystack focuses on the underlying multi-tenant Kubernetes-native platform with virtualization. Both can coexist.</p>
<h2 id="how-to-decide">How to decide</h2>
<p>A practical decision tree:</p>
<ol>
<li><strong>Do you have a working platform underneath?</strong> If no, build that first. Portal adoption assumes platform exists.</li>
<li><strong>Service count over 50? Engineering org over 100?</strong> If yes, a portal helps. If no, a portal is over-engineering.</li>
<li><strong>Backstage operational cost realistic?</strong> If yes (large team, plugin-friendly culture), Backstage. If no, Port or Cortex.</li>
<li><strong>SaaS acceptable?</strong> If yes, Port or Cortex. If no, Backstage or custom.</li>
<li><strong>Will you actually customize?</strong> If yes, Backstage. If no, Port.</li>
</ol>
<p>The decision is smaller than vendors make it sound. The bigger architectural decision is the platform underneath.</p>
]]></content:encoded></item><item><title>Internal developer platform examples — 6 architectural patterns without Backstage lock-in</title><link>https://aenix.io/blog/2026/05/internal-developer-platform-examples-without-backstage/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/internal-developer-platform-examples-without-backstage/</guid><pubDate>Thu, 14 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Backstage</category><category>Kubernetes</category><category>KubeVirt</category><category>Sovereignty</category><category>Multi-tenancy</category><category>Platform Engineering</category><description>Six internal developer platform patterns from production, the tools that show up across them, and how to pick one without defaulting to Backstage.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/internal-developer-platform-examples-without-backstage.jpg" alt=""></p><p>Most articles about internal developer platforms in 2026 are still framed as &ldquo;how to use Backstage.&rdquo; That framing is wrong. Backstage is a useful tool when the catalog discipline is mature; it is not the platform. The platform sits underneath, and the architecture decisions that matter are made there.</p>
<p>This is what working internal developer platforms actually look like.</p>
<h2 id="the-idp-vs-idp-problem">The IDP vs IDP problem</h2>
<p>The terminology is confusingly overloaded. &ldquo;IDP&rdquo; gets used for both:</p>
<ul>
<li><strong>Internal developer platform</strong> — the capability stack: compute, storage, networking, identity, observability, deployment automation, golden paths.</li>
<li><strong>Internal developer portal</strong> — the UI/catalog layer: Backstage, Port, Cortex, custom dashboards. Sits on top of the platform.</li>
</ul>
<p>A platform without a portal still works. A portal without a platform is wallpaper.</p>
<p>In practice, organizations get confused, buy a portal, and discover that adoption stalls because the underlying capabilities aren&rsquo;t actually self-service — they&rsquo;re just better-documented from the catalog.</p>
<p>This article focuses on platform; portal layer is a smaller, secondary decision.</p>
<h2 id="six-idp-patterns-from-production">Six IDP patterns from production</h2>
<h3 id="pattern-1-multi-tenant-kubernetes-native-cloud-platform">Pattern 1: Multi-tenant Kubernetes-native cloud platform</h3>
<p><strong>What:</strong> Single Kubernetes cluster (or fleet) with hard-multi-tenancy. Tenant CRD, per-tenant quotas, RBAC, observability scope. KubeVirt for VM workloads, regular containers for everything else, managed databases as Kubernetes-native services.</p>
<p><strong>Used by:</strong> Service providers (multi-customer), enterprises with strong BU separation, sovereign-cloud builders.</p>
<p><strong>Example stack:</strong> <a href="https://aenix.io/products/cozystack/">Cozystack</a> (open-source CNCF Project).</p>
<p><strong>Why it works:</strong> Single platform serves both VM-heavy and container-heavy workloads. Multi-tenancy is structural, not bolted-on.</p>
<h3 id="pattern-2-gitops-first-per-team-kubernetes">Pattern 2: GitOps-first per-team Kubernetes</h3>
<p><strong>What:</strong> Each product team gets its own Kubernetes namespace (or cluster), provisioned via IaC. GitOps engine (Argo CD or Flux) deploys applications. Platform team maintains the foundation; product teams own their namespace.</p>
<p><strong>Used by:</strong> Mid-size organizations (50-500 engineers) with disciplined GitOps culture.</p>
<p><strong>Example stack:</strong> Vanilla Kubernetes + Flux + Crossplane + Argo Workflows.</p>
<p><strong>Why it works:</strong> Lightweight platform team, high product-team autonomy, clear isolation boundaries.</p>
<h3 id="pattern-3-service-template--golden-path-platform">Pattern 3: Service-template + golden-path platform</h3>
<p><strong>What:</strong> Pre-built service templates (HTTP API, batch worker, scheduled job) with embedded observability, deployment, and alerting. New service = run a template; everything below is automatic.</p>
<p><strong>Used by:</strong> Engineering organizations with high service-creation rate.</p>
<p><strong>Example stack:</strong> Cookiecutter or similar templates + Helm charts + Argo CD + auto-instrumented observability.</p>
<p><strong>Why it works:</strong> Reduces time-to-first-production-deploy to hours; standardizes operational surface.</p>
<h3 id="pattern-4-paas-lite-layered-on-kubernetes">Pattern 4: PaaS-lite layered on Kubernetes</h3>
<p><strong>What:</strong> Heroku-like developer experience built on Kubernetes. <code>git push</code> → CI builds → image pushes → automatic deploy. No Kubernetes manifests in product-team-facing surface.</p>
<p><strong>Used by:</strong> Organizations where Kubernetes literacy isn&rsquo;t universal across product teams.</p>
<p><strong>Example stack:</strong> Knative + custom controllers, or Cloud Foundry-style abstractions on Kubernetes.</p>
<p><strong>Why it works:</strong> Lowest cognitive overhead for product teams. Trade-off: less flexibility for non-standard cases.</p>
<h3 id="pattern-5-backstage-first-with-capability-operators">Pattern 5: Backstage-first with capability operators</h3>
<p><strong>What:</strong> Backstage as the entry point; product teams pick from the catalog. Behind each catalog entry, a Kubernetes operator provisions the requested resource (database, queue, environment). Platform team maintains the operators.</p>
<p><strong>Used by:</strong> Organizations with strong product-management discipline at platform-team level.</p>
<p><strong>Example stack:</strong> Backstage + Crossplane + custom Kubernetes operators per service type.</p>
<p><strong>Why it works:</strong> Catalog-first UX is intuitive for product teams; operators handle the actual provisioning.</p>
<p><strong>Trade-off:</strong> Platform team must invest in product-management practices (catalog hygiene, deprecations, roadmap).</p>
<h3 id="pattern-6-external-services-as-platform">Pattern 6: External-services-as-platform</h3>
<p><strong>What:</strong> Platform team provides catalog of external SaaS services (managed databases, observability, identity) via standardized integration. The &ldquo;platform&rdquo; is mostly integration glue + provisioning automation.</p>
<p><strong>Used by:</strong> Cloud-native organizations comfortable with hyperscaler-managed services and SaaS dependencies.</p>
<p><strong>Example stack:</strong> Crossplane (with cloud providers + SaaS providers) + Backstage catalog.</p>
<p><strong>Why it works:</strong> Smallest platform-team footprint; rides hyperscaler reliability.</p>
<p><strong>Trade-off:</strong> Vendor lock-in, sovereignty concerns, cost ceiling.</p>
<h2 id="tools-that-show-up-across-patterns">Tools that show up across patterns</h2>
<p>The tooling landscape has converged. Across the 6 patterns above:</p>
<h3 id="compute--orchestration">Compute / orchestration</h3>
<p><strong>Kubernetes</strong> is the de facto orchestration layer. Distribution choice depends on operational model:</p>
<ul>
<li><strong>Cozystack</strong> for multi-tenant + virtualization (open-source, CNCF Project)</li>
<li><strong>OpenShift</strong> for enterprise commercial</li>
<li><strong>Vanilla Kubernetes</strong> for simplicity</li>
<li><strong>Talos</strong> as the OS underneath</li>
</ul>
<h3 id="iac">IaC</h3>
<p><strong>Terraform / OpenTofu</strong> for cloud and infrastructure. <strong>Crossplane</strong> for Kubernetes-native infrastructure abstraction. <strong>Pulumi</strong> as a TypeScript-friendly alternative.</p>
<h3 id="gitops">GitOps</h3>
<p><strong>Argo CD</strong> and <strong>Flux</strong> are the two production-grade GitOps engines. Both work; Flux is closer to upstream Kubernetes way; Argo CD has stronger UI ergonomics.</p>
<h3 id="internal-developer-portal">Internal developer portal</h3>
<p><strong>Backstage</strong> dominates. Alternatives: Port, Cortex, Compass, custom. Smaller-scale: a Markdown documentation site backed by a service-catalog YAML in Git.</p>
<h3 id="observability">Observability</h3>
<p><strong>VictoriaMetrics + VictoriaLogs</strong> (open-source, low-overhead). <strong>Prometheus + Loki</strong> (still common). Commercial: Datadog, New Relic, Splunk Cloud — but check sovereignty implications.</p>
<h3 id="secrets-and-identity">Secrets and identity</h3>
<p><strong>External Secrets Operator</strong> + cloud-provider secret stores or HashiCorp Vault. <strong>SPIFFE/SPIRE</strong> for service identity at scale.</p>
<h3 id="cicd">CI/CD</h3>
<p><strong>GitHub Actions</strong> for source-of-truth; <strong>Argo Workflows</strong> or <strong>Tekton</strong> for Kubernetes-native pipelines.</p>
<h2 id="how-to-pick-a-pattern">How to pick a pattern</h2>
<p>The decision tree:</p>
<ol>
<li>
<p><strong>Multi-tenancy needed?</strong> (Service-provider model, hard BU separation, regulated multi-tenant.)</p>
<ul>
<li>Yes → Pattern 1 (Multi-tenant Kubernetes-native, Cozystack-style).</li>
<li>No → continue.</li>
</ul>
</li>
<li>
<p><strong>Strong product-team autonomy desired?</strong> (Mature engineering culture, GitOps discipline.)</p>
<ul>
<li>Yes → Pattern 2 (GitOps-first per-team).</li>
<li>No → continue.</li>
</ul>
</li>
<li>
<p><strong>High service-creation rate?</strong> (&gt;10 new services per month, microservices-heavy.)</p>
<ul>
<li>Yes → Pattern 3 (Service-template + golden path).</li>
<li>No → continue.</li>
</ul>
</li>
<li>
<p><strong>Kubernetes literacy uneven across product teams?</strong></p>
<ul>
<li>Yes → Pattern 4 (PaaS-lite on Kubernetes).</li>
<li>No → continue.</li>
</ul>
</li>
<li>
<p><strong>Strong product-management discipline at platform-team level + appetite for catalog UX?</strong></p>
<ul>
<li>Yes → Pattern 5 (Backstage-first with capability operators).</li>
<li>No → Pattern 6 (External-services-as-platform).</li>
</ul>
</li>
</ol>
<p>This is a heuristic, not a formula. Real engagements often hybrid two patterns.</p>
<h2 id="what-about-the-portal">What about the portal?</h2>
<p>Once the platform pattern is chosen, the portal decision is smaller:</p>
<ul>
<li><strong>Backstage</strong> when service catalog discipline is real, the platform has 50+ catalog entries (services + capabilities), and a platform team can sustain Backstage&rsquo;s plugin maintenance.</li>
<li><strong>Port / Cortex / Compass</strong> when Backstage seems too much engineering effort and a SaaS portal is acceptable.</li>
<li><strong>Custom Markdown site + YAML catalog</strong> for organizations under 100 engineers; sufficient and cheaper.</li>
<li><strong>No portal</strong> for very small platforms; the IaC repo is the documentation.</li>
</ul>
<p>A portal is a service-discovery and self-service entry point. It&rsquo;s not the platform; not the differentiator.</p>
<h2 id="common-architectural-mistakes">Common architectural mistakes</h2>
<h3 id="mistake-1-starting-with-the-portal">Mistake 1: starting with the portal</h3>
<p>Buying Backstage before the platform exists puts the cart before the horse. Adoption stalls because the catalog points to weeks-long ticket processes.</p>
<h3 id="mistake-2-too-many-opinions-too-rigid">Mistake 2: too many opinions, too rigid</h3>
<p>Golden paths should be golden, not gold-plated. Product teams need escape hatches for non-standard cases.</p>
<h3 id="mistake-3-under-invested-platform-team-headcount">Mistake 3: under-invested platform-team headcount</h3>
<p>Platform team that absorbs both build and on-call duties for shared services becomes ticket support. Build velocity collapses.</p>
<h3 id="mistake-4-vendor-led-platform-in-a-box">Mistake 4: vendor-led platform-in-a-box</h3>
<p>Buying a &ldquo;complete IDP solution&rdquo; rebuilds the lock-in problem with a different vendor.</p>
<h3 id="mistake-5-building-for-engineering-elegance-not-adoption">Mistake 5: building for engineering elegance, not adoption</h3>
<p>Platform team&rsquo;s customers are product teams. Architecture optimized for engineering elegance often produces a platform nobody adopts.</p>
<h2 id="want-to-dig-deeper">Want to dig deeper?</h2>
<ul>
<li><strong><a href="https://aenix.io/services/internal-developer-platform/">Internal developer platform services</a></strong> — engagement details</li>
<li><strong><a href="https://aenix.io/services/platform-engineering/">Platform engineering services</a></strong> — broader scope</li>
<li><strong><a href="https://aenix.io/blog/2026/05/platform-engineering-vs-devops-vs-sre/">Platform engineering vs DevOps vs SRE</a></strong> — terminology</li>
<li><strong><a href="https://aenix.io/products/cozystack/">Cozystack</a></strong> — multi-tenant Kubernetes-native foundation</li>
</ul>
]]></content:encoded></item><item><title>Hybrid cloud architecture patterns 2026 — what works, what fails, and how to choose</title><link>https://aenix.io/blog/2026/05/hybrid-cloud-architecture-patterns-2026/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/hybrid-cloud-architecture-patterns-2026/</guid><pubDate>Wed, 13 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>AI and ML</category><category>Observability</category><description>Five hybrid cloud patterns that work in production, what makes them work, the failure modes to avoid, and when hybrid is the wrong answer.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/hybrid-cloud-architecture-patterns-2026.jpg" alt=""></p><p>Hybrid cloud as a term has accumulated marketing weight. Vendors describe almost any non-pure-public-cloud architecture as &ldquo;hybrid.&rdquo; The architecturally-relevant question is more specific: what coherent integration pattern connects your private and public substrates, and does it actually work for your workload portfolio?</p>
<p>This is the working version.</p>
<h2 id="what-hybrid-cloud-actually-means">What hybrid cloud actually means</h2>
<p>Three definitions, in increasing order of architectural usefulness:</p>
<ol>
<li><strong>Hybrid as workload distribution.</strong> Some workloads in public cloud, some in private cloud, no integration. Most &ldquo;hybrid&rdquo; architectures stop here. <em>(Practical but operationally fragmented.)</em></li>
<li><strong>Hybrid as data flow.</strong> Some workloads in each substrate, with explicit data flow architectures connecting them. <em>(Architecturally honest about cross-substrate dependencies.)</em></li>
<li><strong>Hybrid as unified platform.</strong> Single platform abstraction running on multiple substrates, with consistent operations, deployment, and observability. <em>(Where the leverage is.)</em></li>
</ol>
<p>We engineer toward (3) because (1) becomes operational debt, and (2) is half the work.</p>
<h2 id="five-working-hybrid-patterns">Five working hybrid patterns</h2>
<h3 id="pattern-1-steady-state-on-prem--elastic-in-public-cloud">Pattern 1: Steady-state on-prem + elastic in public cloud</h3>
<p><strong>What:</strong> Predictable steady-state workloads (databases, batch processing, internal apps) on private cloud. Elastic spike workloads (customer-facing web, ML training) in public cloud. Cross-substrate data flow is designed (not accidental).</p>
<p><strong>Best for:</strong> SaaS companies with predictable steady-state customer count + spike-y customer-facing patterns. Most common hybrid pattern in 2026.</p>
<p><strong>Operational pattern:</strong> Two substrates with shared platform abstraction (Cozystack on private + Cozystack control plane managing public-cloud Kubernetes). Workloads are portable in principle; in practice, steady-state ones don&rsquo;t move.</p>
<h3 id="pattern-2-critical-on-prem--non-critical-in-public-cloud">Pattern 2: Critical on-prem + non-critical in public cloud</h3>
<p><strong>What:</strong> Regulated workloads (banking, healthcare, public-sector) on private cloud. Auxiliary workloads (analytics, internal tooling, dev/test) in public cloud. Sovereignty drives the split.</p>
<p><strong>Best for:</strong> Financial services, healthcare, public-sector. The default when DORA / sectoral / data-residency rules drive architecture.</p>
<p><strong>Operational pattern:</strong> Stricter isolation between substrates than Pattern 1. Cross-substrate flows audited and minimized. Audit-readiness drives architecture.</p>
<h3 id="pattern-3-geographic-split">Pattern 3: Geographic split</h3>
<p><strong>What:</strong> EU workloads in EU on-prem; non-EU workloads in regional public cloud. Driven by sovereignty rather than cost or workload type.</p>
<p><strong>Best for:</strong> Multinational enterprises with strong sovereignty pressure in some markets but commercial freedom in others.</p>
<p><strong>Operational pattern:</strong> Per-jurisdiction tenant isolation. Cross-jurisdiction data flows under explicit legal basis (SCCs, BCRs, etc.).</p>
<h3 id="pattern-4-edge--core-hybrid">Pattern 4: Edge + core hybrid</h3>
<p><strong>What:</strong> Centralized core (private cloud or hyperscaler region) plus edge sites at customer / branch / factory locations. Cross-edge synchronization handled at platform level.</p>
<p><strong>Best for:</strong> Telco edge compute, retail, manufacturing, smart-grid operators.</p>
<p><strong>Operational pattern:</strong> Edge sites are smaller, simpler instances of the platform; core has full feature surface.</p>
<h3 id="pattern-5-ai-specific-split">Pattern 5: AI-specific split</h3>
<p><strong>What:</strong> AI training and inference on dedicated GPU infrastructure (private cloud); rest of business on public cloud or hybrid mix. AI workload economics drive the split.</p>
<p><strong>Best for:</strong> AI-heavy organizations where sustained GPU utilization makes hyperscaler economics untenable. Increasingly common in 2026.</p>
<p><strong>Operational pattern:</strong> AI cluster operates as a sub-platform; data and model flows between AI cluster and other workloads are designed.</p>
<h2 id="what-makes-hybrid-actually-work">What makes hybrid actually work</h2>
<p>Beyond the pattern, three architectural principles separate working hybrid from fragmented multi-cloud:</p>
<h3 id="principle-1-one-platform-abstraction">Principle 1: one platform abstraction</h3>
<p>The Kubernetes API (or equivalent) is the lingua franca across substrates. Workload deployment, observability, identity, and operational practices are consistent. The platform team operates one platform, not three.</p>
<p>Cozystack supports this directly: same control plane managing customer-hardware clusters, public-cloud-region deployments, and edge sites.</p>
<h3 id="principle-2-workload-portability-where-it-matters">Principle 2: workload portability where it matters</h3>
<p>Critical workloads use platform abstractions that work on multiple substrates. Non-critical workloads can be hyperscaler-native if that&rsquo;s the right trade-off.</p>
<h3 id="principle-3-explicit-data-flow-control">Principle 3: explicit data flow control</h3>
<p>Cross-cloud and cross-region data flows are designed, costed, and monitored. Egress is a budget line; sovereignty implications are documented; latency profiles known.</p>
<h2 id="common-failure-modes">Common failure modes</h2>
<h3 id="failure-1-hybrid-as-fragmented-teams">Failure 1: hybrid as fragmented teams</h3>
<p>Public-cloud team and on-prem team operating separately. Different tooling, different deployment patterns, no shared platform. &ldquo;Hybrid&rdquo; in name; multi-team chaos in practice.</p>
<h3 id="failure-2-cloud-bursting-nobody-uses">Failure 2: cloud bursting nobody uses</h3>
<p>Architecture supports bursting from on-prem to public cloud for capacity overflow. In production, bursting capability is theoretical — cross-substrate data movement is too slow. Architecture is over-engineered for unused capability.</p>
<h3 id="failure-3-vendor-led-hybrid-platform">Failure 3: vendor-led hybrid platform</h3>
<p>A vendor sells a &ldquo;complete hybrid solution&rdquo; running their software in your datacenter and theirs. Lock-in is structural; vendor&rsquo;s roadmap becomes yours.</p>
<h3 id="failure-4-operational-drift">Failure 4: operational drift</h3>
<p>Same workload runs differently on public cloud and on-prem because tooling diverges over time. Portability degrades; cost of moving workloads grows.</p>
<h3 id="failure-5-implicit-data-flows">Failure 5: implicit data flows</h3>
<p>Cross-substrate traffic emerges from accidental architecture decisions. Egress costs surprise; sovereignty audit finds compliance gaps.</p>
<h2 id="when-hybrid-is-wrong">When hybrid is wrong</h2>
<p>A few honest cases where pure architecture beats hybrid:</p>
<ul>
<li><strong>All workloads are elastic and not regulated</strong> — pure public cloud usually wins.</li>
<li><strong>All workloads are steady-state, regulated, and modest in scale</strong> — pure private cloud usually wins.</li>
<li><strong>Engineering organization is small</strong> — operating one substrate is hard enough; two is harder.</li>
<li><strong>Workload portability genuinely doesn&rsquo;t matter</strong> — pure hyperscaler-native architecture has its merits.</li>
</ul>
<p>The honest engagement says &ldquo;stay pure&rdquo; when that&rsquo;s the answer.</p>
<h2 id="implementation-sequence">Implementation sequence</h2>
<p>A practical sequence for moving from fragmented multi-cloud to coherent hybrid:</p>
<ol>
<li><strong>Workload classification</strong> — every workload labeled &ldquo;this substrate, this reason.&rdquo;</li>
<li><strong>Platform foundation</strong> — Cozystack (or chosen platform) deployed on private substrate; control plane reach extended to public-cloud regions.</li>
<li><strong>Migration cohort 1</strong> — handful of workloads moved to the unified platform pattern; pattern validated.</li>
<li><strong>Operational unification</strong> — observability, identity, deployment unified across substrates.</li>
<li><strong>Migration cohorts 2-N</strong> — remaining workloads aligned with the pattern.</li>
<li><strong>Steady state</strong> — single platform team, single operations model, multiple substrates.</li>
</ol>
<p>Total elapsed: typically about 8-12 months for ~100 VMs and 18-24 months for ~1,000 VMs, including planning and migration waves.</p>
<h2 id="want-to-dig-deeper">Want to dig deeper?</h2>
<ul>
<li><strong><a href="https://aenix.io/solutions/hybrid-cloud-platform/">Hybrid cloud platform services</a></strong> — engagement details</li>
<li><strong><a href="https://aenix.io/solutions/cloud-repatriation/">Cloud repatriation</a></strong> — when leaving public cloud</li>
<li><strong><a href="https://aenix.io/services/private-cloud-consulting/">Private cloud consulting</a></strong> — private side</li>
<li><strong><a href="https://aenix.io/products/cozystack/">Cozystack</a></strong> — the platform</li>
</ul>
]]></content:encoded></item><item><title>Developer self-service — the cost of developer drag, and what an internal developer platform actually pays back</title><link>https://aenix.io/blog/2026/05/idp-edition-developer-velocity-economics/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/idp-edition-developer-velocity-economics/</guid><pubDate>Wed, 13 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Platform Engineering</category><category>Cozystack</category><category>DevOps</category><category>Multi-tenancy</category><description>Time-to-environment cost, golden-path coverage, platform-team sizing, and the economic case for an IDP that pays back inside 12 months at 200+ engineers.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/idp-edition-developer-velocity-economics.jpg" alt=""></p><p>The &ldquo;should we invest in a platform team?&rdquo; conversation tends to stall
at one of two places: either the CFO can&rsquo;t see the economic case (&ldquo;we
already pay DevOps engineers, why add more headcount?&rdquo;), or the
engineering organisation tried Backstage-as-a-platform once, found it
shallow, and has lost trust in the category.</p>
<p>This article walks through both. What does the cost of developer drag
actually look like in numbers? What does an IDP that <em>works</em> deliver?
And what does the developer self-service layer of Ænix Private Cloud Platform do that you&rsquo;d otherwise have
to build?</p>
<h2 id="the-cost-of-developer-drag">The cost of developer drag</h2>
<p>The most consistent finding across our platform-readiness assessments
of 200-engineer-plus organisations is that <strong>time-to-environment is
2-6 weeks for what should be a 30-minute self-service action.</strong></p>
<p>Where does the time go? Typically:</p>
<ul>
<li>IAM: a ticket to the security team, 2-5 days</li>
<li>Network connectivity: a ticket to network engineering, 3-7 days</li>
<li>Observability onboarding: ad-hoc, often discovered as missing late</li>
<li>Compliance review: 1-3 days, sometimes longer for regulated data</li>
<li>Approval chain: 2-4 days, occasionally longer in matrixed orgs</li>
</ul>
<p>Each step requires a handoff, which means context loss, which means
re-confirmation, which means more time. The standalone steps are
small; the cumulative friction is large.</p>
<p>At 5% of engineering productivity (rough estimate from our engagements
matching the 200-engineer scale and 2-6 week time-to-environment), the
cost of drag on a 200-engineer organisation is <strong>~10 engineers&rsquo; worth
of lost throughput per year</strong> — roughly €1.5-2M at fully loaded rates.
Compressing time-to-environment to hours captures most of that.</p>
<p>The internal-developer-platform investment that delivers this
compression typically pays back inside 12 months for organisations of
this scale.</p>
<h2 id="what-platform-that-works-actually-delivers">What &ldquo;platform that works&rdquo; actually delivers</h2>
<p>For a credible IDP, five characteristics matter more than tool choice:</p>
<h3 id="1-faster-than-the-ticket-alternative">1. Faster than the ticket alternative</h3>
<p>If self-service takes 2 days and a ticket takes 3 days, teams will
take the ticket — waiting is easier than learning. The platform has
to be substantially faster for adoption to flip.</p>
<h3 id="2-reliable-enough-to-trust">2. Reliable enough to trust</h3>
<p>The self-service path works the first time, every time, for the
documented use case. If it breaks 1 in 10 times, teams stop trusting
it. Adoption stalls.</p>
<h3 id="3-documented-in-one-page-or-less">3. Documented in one page or less</h3>
<p>If the docs are longer than a page, the architecture is too complex.
Real golden paths are simple by design.</p>
<h3 id="4-owned-by-a-real-team">4. Owned by a real team</h3>
<p>A team maintains the path, fields edge cases, ships improvements.
Without ownership, paths decay. The platform-engineering function
must be a function, not a hobby.</p>
<h3 id="5-has-escape-hatches">5. Has escape hatches</h3>
<p>Product teams can deviate when their case is special. The escape is
to a real conversation with the platform team, not &ldquo;use the path or
fail.&rdquo;</p>
<h2 id="what-developer-self-service-on-private-cloud-platform-ships">What developer self-service on Private Cloud Platform ships</h2>
<p>The developer self-service layer of Ænix Private Cloud Platform is the productisation of these characteristics
on top of the Cozystack foundation. Specifically:</p>
<h3 id="a-multi-tenant-cozystack-platform-with-tenant-crd">A multi-tenant Cozystack platform with Tenant CRD</h3>
<p>Every product team gets a Tenant — a Kubernetes-native object with
its own namespace, quota, RBAC, observability scope. Per-tenant
isolation without per-team cluster overhead. Nested tenants support
business-unit hierarchies.</p>
<p>This solves the &ldquo;soft multi-tenancy is too leaky, cluster-per-team is
operationally expensive&rdquo; trilemma without compromise.</p>
<h3 id="golden-path-first-cozystack-dashboard">Golden-path-first Cozystack Dashboard</h3>
<p>The Cozystack Dashboard in Private Cloud Platform exposes opinionated paths for the 5-10
most common product-team needs: environment provisioning, application
deployment, managed-database provisioning, observability onboarding,
secrets management. Each path completes in minutes.</p>
<p>Curated, not exhaustive — what your platform team can actually back
with support, not every Cozystack capability.</p>
<h3 id="gitops-automation">GitOps automation</h3>
<p>Argo CD and Argo Workflows pre-integrated with the platform. GitLab
integration for the most common enterprise SCM. Product teams commit
IaC; the platform reacts. No clicky-clicky for production changes.</p>
<h3 id="pre-built-service-templates">Pre-built service templates</h3>
<p>Standard service templates (HTTP API, batch worker, scheduled job)
with embedded observability, deployment, alerting. New service =
template instantiation. Time-to-first-production-deploy is hours, not
weeks.</p>
<h3 id="internal-product-management-discipline-built-in">Internal-product-management discipline built in</h3>
<p>The platform team is treated as a function with product-team
customers. The engagement includes platform-team RACI,
internal-NPS metrics, deprecation policy templates, roadmap-management
patterns. The discipline is often the missing piece.</p>
<h2 id="what-stays-the-customers-responsibility">What stays the customer&rsquo;s responsibility</h2>
<p>Developer self-service on Private Cloud Platform is the platform substrate plus the operational discipline.
Several things remain yours:</p>
<ul>
<li><strong>Defining your specific golden paths</strong> — the 5-10 you build first
depend on what your product teams actually request most.</li>
<li><strong>Platform-engineering team sizing</strong> — typical mature ratio is
1 platform engineer per 10-20 product engineers; we recommend hiring
ahead of demand.</li>
<li><strong>Adoption</strong> — a great platform that nobody uses is a sunk cost.
Platform-team product-management practices (interviewing product
teams, measuring adoption per path, sunsetting paths that don&rsquo;t get
used) need to happen.</li>
</ul>
<p>Ænix&rsquo;s engagement model supports all three but doesn&rsquo;t replace
customer ownership. Platforms that customers don&rsquo;t own organisationally
don&rsquo;t outlast the engagement.</p>
<h2 id="the-economic-case-versus-alternatives">The economic case versus alternatives</h2>
<h3 id="versus-devops-only">Versus DevOps-only</h3>
<p>DevOps without separate platform engineering scales linearly with team
count — every product team figures out infrastructure, observability,
identity, release engineering independently. Above ~50 engineers the
duplication overwhelms the savings. Above ~200 engineers it&rsquo;s an
operational drag.</p>
<p>Developer self-service on Private Cloud Platform shifts the model: platform engineering compounds
sublinearly with team count. Adding the 21st product team doesn&rsquo;t add
21st-team&rsquo;s-worth of platform overhead — the existing platform absorbs
them.</p>
<h3 id="versus-building-it-yourself-on-vanilla-kubernetes--backstage">Versus building it yourself on vanilla Kubernetes + Backstage</h3>
<p>This is a credible alternative for organisations with strong platform
engineering capacity and clear architectural opinions. Trade-off:
12-24 months of build time before the platform reaches &ldquo;production
ready&rdquo; for product teams; ongoing platform-component maintenance
overhead.</p>
<p>Ænix Private Cloud Platform, which includes developer self-service,
delivers the platform substrate within a 3-12 month build, with
ongoing Ænix support. For organisations not staffed for a 12-24
month build, this is the difference between platform engineering
happening this year or in 2028.</p>
<h3 id="versus-backstage-as-platform-anti-pattern">Versus Backstage-as-platform anti-pattern</h3>
<p>Backstage is a portal, not a platform. Buying Backstage before the
underlying capabilities are self-service produces a beautiful catalog
over the same operational chaos. Adoption stalls.</p>
<p>Private Cloud Platform&rsquo;s Cozystack Dashboard can be replaced or augmented by Backstage if
the customer prefers — but the underlying capabilities (environment
provisioning, observability, secrets, identity) are self-service
because the platform is, not because the portal pretends they are.</p>
<h3 id="versus-hyperscaler-managed-kubernetes-eks--aks--gke">Versus hyperscaler-managed Kubernetes (EKS / AKS / GKE)</h3>
<p>Hyperscaler managed Kubernetes is operationally simpler than self-
managed clusters. Trade-off: vendor lock-in (control plane is the
vendor&rsquo;s; you can&rsquo;t replicate exactly elsewhere); cost ceiling
(hyperscaler economics on sustained workloads); sovereignty
implications.</p>
<p>Private Cloud Platform fits when the cost or sovereignty trade-offs make
self-managed worth the operational investment. For early-stage
companies without sovereignty pressure, hyperscaler-managed is still
the right call.</p>
<h2 id="when-developer-self-service-is-the-right-answer">When developer self-service is the right answer</h2>
<p>Strong fit:</p>
<ul>
<li>200+ engineers across 5+ product teams</li>
<li>Time-to-environment is currently 2-6 weeks; target is hours</li>
<li>Existing platform team is overwhelmed by ticket volume rather than
building paths</li>
<li>Multi-tenant isolation matters (regulated data, BU separation, or
service-provider model)</li>
<li>Investment in platform engineering is supported by engineering
leadership</li>
</ul>
<p>Marginal fit:</p>
<ul>
<li>100-200 engineers with growing platform pain but limited budget;
start with Cozystack Enterprise Support and scale into the Ænix
Private Cloud Platform as the team grows</li>
<li>Strong existing in-house platform with specific gaps — partial
engagement may fit better than a full Private Cloud Platform build</li>
</ul>
<p>Poor fit:</p>
<ul>
<li>Under 50 engineers, single product team; DevOps-only is right</li>
<li>Hyperscaler-managed Kubernetes meets all requirements; sovereignty
not a concern</li>
</ul>
<h2 id="engagement-structure">Engagement structure</h2>
<ul>
<li><strong>Discovery call</strong> (30 min, free) — fit assessment</li>
<li><strong>Platform Readiness Assessment</strong> (14- or 28-day, IDP workstream
emphasised) — current-state time-to-environment, maturity score,
recommended golden paths</li>
<li><strong>Pilot deployment</strong> (3-6 months) — Cozystack platform + 3-5 golden
paths + 2-3 pilot product teams onboarded</li>
<li><strong>Full build</strong> (within the 3-12 month Private Cloud Platform build,
depending on scope) — platform expanded to the full engineering
organisation, all targeted golden paths shipped</li>
<li><strong>Support subscription</strong> (ongoing) — Plus or Enterprise support tier
for 24×7 escalation (see <a href="https://aenix.io/pricing/">/pricing/</a>)</li>
</ul>
<p>Engagement size: project plus support subscription, quoted per RFP.</p>
<h2 id="where-to-dig-deeper">Where to dig deeper</h2>
<ul>
<li><strong><a href="https://aenix.io/products/private-cloud-platform/">Private Cloud Platform landing</a></strong> —
feature list, product-specific FAQ</li>
<li><strong><a href="https://aenix.io/services/internal-developer-platform/">Internal Developer Platform services</a></strong> —
engagement details</li>
<li><strong><a href="https://aenix.io/services/platform-engineering/">Platform Engineering services</a></strong> —
broader scope</li>
<li><strong><a href="https://aenix.io/solutions/developer-self-service/">Developer self-service solutions</a></strong> —
buyer-trigger landing</li>
<li><strong><a href="https://aenix.io/blog/2026/05/internal-developer-platform-examples-without-backstage/">Internal developer platform examples — 6 patterns</a></strong> —
the six production IDP patterns</li>
<li><strong><a href="https://aenix.io/blog/2026/05/internal-developer-portal-vs-platform/">Internal developer portal vs platform</a></strong> —
Backstage&rsquo;s place in 2026</li>
</ul>
]]></content:encoded></item><item><title>Ænix Billing — per-minute usage-based billing for Managed PostgreSQL, Redis, Kafka and ClickHouse on Cozystack</title><link>https://aenix.io/blog/2026/05/aenix-billing-per-minute-managed-services-cozystack/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/aenix-billing-per-minute-managed-services-cozystack/</guid><pubDate>Wed, 13 May 2026 00:00:00 +0000</pubDate><dc:creator>Timur Tukaev</dc:creator><category>Cozystack</category><category>Kubernetes</category><category>Multi-tenancy</category><category>Platform Engineering</category><category>Billing</category><description>Ænix Billing brings AWS-style per-minute, usage-based billing to managed Postgres, Redis, Kafka and ClickHouse on Cozystack via a Kubernetes-native API.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/aenix-billing-per-minute-managed-services-cozystack.jpg" alt="Ænix Billing — per-minute usage-based billing for Managed PostgreSQL, Redis, Kafka and ClickHouse on Cozystack" width="1200" height="630" loading="lazy" decoding="async"></p>
<p><strong>You ship a managed Postgres or ClickHouse cluster to a tenant in two minutes. Now you need to charge them — by the minute, broken down by CPU, memory and disk, reconcilable to the cent. AWS RDS does this. GCP does this. On your own metal? Until now: a custom Prometheus script, a monthly spreadsheet export, and a sales call to explain &ldquo;why is the number this big&rdquo;. Ænix Billing closes that gap.</strong></p>
<h2 id="what-ænix-billing-is">What Ænix Billing is</h2>
<p>A native Kubernetes extension API server that exposes the <code>billing.aenix.io/v1alpha1</code> API. You query it the same way you query the regular Kubernetes API — with <code>kubectl</code>, kubeconfig RBAC, your existing tooling — and it returns a structured <code>UsageReport</code> for any tenant, any workload, any time window.</p>
<p>Under the hood:</p>
<ul>
<li><strong>Cozystack Workload CRD</strong> — every managed service (Postgres primary, Redis sentinel, ClickHouse shard, KubeVirt VM, Kubernetes worker, S3 bucket) is one <code>Workload</code> object.</li>
<li><strong>Billing controller</strong> — turns <code>Workload</code> state into Prometheus metrics: operational lifetime, owner tenant, kind/type, resource reservations.</li>
<li><strong>VictoriaMetrics</strong> — stores metrics with stream aggregation. (Ænix and Cozystack standardise on VictoriaMetrics + VictoriaLogs, not Prometheus/Loki.)</li>
<li><strong>Billing API server</strong> — serves the API, computes the definite integral of reservations over the requested time window.</li>
</ul>
<p>The shape: a thin extension-API server, a small controller, your existing metrics store. No new dependency to operate. Your platform team already runs all of these components.</p>
<h2 id="what-gets-billed">What gets billed</h2>
<p>Every consumer (a single Postgres pod, a Redis instance, a ClickHouse replica) is reported with:</p>
<ul>
<li><code>vCPUHours</code> or <code>CPUHours</code> — switchable, depending on whether you bill physical or virtualised CPU</li>
<li><code>MemoryGiBHours</code> (or decimal <code>MemoryGBHours</code>)</li>
<li><code>EphemeralStorageGiBHours</code></li>
<li><code>PersistentVolumeGBHours</code> — partitioned by storage class (NVMe vs HDD vs replicated remote)</li>
<li><code>IPAddressHours</code> — derived from MetalLB IP pool labels</li>
<li><code>LifetimeHours</code> — pure on-time, independent of reservations</li>
<li><code>S3StorageGBHours</code> and <code>S3PhysicalStorageGBHours</code> — for object storage tenants</li>
</ul>
<p><code>Workload</code> metadata travels with the report — <code>kind: postgres</code>, <code>type: master</code> vs <code>replica</code>, the owning tenant, custom Cozystack labels — so pricing rules can be as fine-grained as you want. Charge replicas at half the rate of primaries. Charge NVMe-backed volumes 3× HDD-backed. Apply a 20% discount to a strategic tenant. The data is there; the policy is yours.</p>
<h2 id="how-you-query-it">How you query it</h2>
<p>A <code>UsageReport</code> is a Kubernetes object. You build a request like any other CRD:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">apiVersion</span><span class="p">:</span><span class="w"> </span><span class="l">billing.aenix.io/v1alpha1</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">kind</span><span class="p">:</span><span class="w"> </span><span class="l">UsageReport</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">query</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">tenant</span><span class="p">:</span><span class="w"> </span><span class="l">tenant-acme</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">includeSubTenants</span><span class="p">:</span><span class="w"> </span><span class="kc">true</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">startTimestamp</span><span class="p">:</span><span class="w"> </span><span class="ld">2026-04-01T00:00:00Z</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">endTimestamp</span><span class="p">:</span><span class="w">   </span><span class="ld">2026-05-01T00:00:00Z</span><span class="w">
</span></span></span></code></pre></div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">kubectl create -f report.yaml -o yaml
</span></span></code></pre></div><p>You get back the same object with <code>report.consumers[]</code> populated — one entry per Postgres replica, Redis pod, ClickHouse shard, K8s worker — each with its own consumption breakdown.</p>
<ul>
<li><code>includeSubTenants: true</code> recursively pulls every nested tenant</li>
<li><code>excludeTenants:</code> filters internal namespaces out</li>
<li><code>workload: postgres-myapp-1</code> scopes to a single cluster</li>
<li>Time-window the report to a billing month, a single day, or a tenant&rsquo;s lifetime — the API accepts any RFC3339 interval</li>
</ul>
<p>The same <code>kubectl get</code> mechanism that pulls a list of pods now pulls a tenant&rsquo;s invoice line items. RBAC works the same way — the billing team gets a kubeconfig scoped to the billing API; nothing more.</p>
<h2 id="why-this-matters-for-hosting-providers">Why this matters for hosting providers</h2>
<p>You already host managed Postgres, Redis, Kafka, ClickHouse, S3 and Kubernetes via Cozystack. With Ænix Billing you get the missing piece — accurate per-tenant, per-workload, per-resource accounting that plugs into your invoicing. No spreadsheets. No quarterly reconciliation. No &ldquo;why is my bill big&rdquo; arguments with customers.</p>
<p>Three things change for the operations team:</p>
<ol>
<li><strong>Billing is no longer a separate system.</strong> Same API as the rest of the platform. Same RBAC. Same kubectl. Nothing new to learn or run.</li>
<li><strong>Disputes drop.</strong> The line items in an invoice map one-to-one to the resources Cozystack actually scheduled. The customer can run the same query, see the same numbers, reconcile to the cent.</li>
<li><strong>Pricing experiments are cheap.</strong> Want to A/B test charging by <code>MemoryGiBHours</code> instead of <code>vCPUHours</code> for a specific service tier? Don&rsquo;t change billing pipeline code — change the policy that consumes the same report.</li>
</ol>
<h2 id="how-it-differs-from-aws-rds--gcp-billing">How it differs from AWS RDS / GCP billing</h2>
<p>AWS RDS and Cloud SQL bill per-minute against a managed-service price list — the provider owns the price list and the customer accepts the line items as authoritative. <strong>You are that provider</strong> on Cozystack. Ænix Billing gives you the same per-minute granularity, with the pricing model and the resource taxonomy under your control. The data is exposed, the policy is yours.</p>
<p>For hosting providers running on bare metal — Hetzner, OVH, regional data centres, sovereign cloud builders — this is the difference between competing with hyperscalers on capability vs. competing on price. The capability has been the gap.</p>
<h2 id="distribution">Distribution</h2>
<p>Ænix Billing is a proprietary Ænix module, delivered alongside Cozystack to customers with an Ænix subscription (for example <a href="https://aenix.io/products/public-cloud-platform/">Ænix Public Cloud Platform</a>). The Cozystack platform stays Apache-2.0 and CNCF-governed; the billing layer is not part of it.</p>
<p>If you run a hosting business or a private cloud on Cozystack and want billing working out of the box, <a href="https://aenix.io/contact/">book a discovery call</a> — we&rsquo;ll talk scope, pricing, and rollout.</p>
<h2 id="join-the-community">Join the community</h2>
<ul>
<li><strong>Cozystack</strong> — <a href="https://cozystack.io">cozystack.io</a></li>
<li><strong>GitHub</strong> — <a href="https://github.com/cozystack">github.com/cozystack</a></li>
<li><strong>Telegram</strong> — <a href="https://t.me/cozystack">t.me/cozystack</a></li>
<li><strong>Kubernetes Slack</strong> — <a href="https://kubernetes.slack.com/archives/C06L3CPRVN1">#cozystack</a> (need an invite? <a href="https://slack.kubernetes.io">slack.kubernetes.io</a>)</li>
<li><strong>Community meetings</strong> — <a href="https://cozystack.io/community/">calendar</a></li>
</ul>
]]></content:encoded></item><item><title>Hosting provider platform modernization — from VPS to cloud product</title><link>https://aenix.io/blog/2026/05/hosting-provider-platform-modernization/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/hosting-provider-platform-modernization/</guid><pubDate>Tue, 12 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Kubernetes</category><category>Cozystack</category><category>Sovereignty</category><category>AI and ML</category><category>GPU</category><category>Multi-tenancy</category><description>Architectural starting point, migration sequencing, and unit economics for hosting providers modernizing onto a Kubernetes-native multi-tenant platform.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/hosting-provider-platform-modernization.jpg" alt=""></p><h2 id="the-hosting-provider-opportunity">The hosting provider opportunity</h2>
<p>In 2026, hosting providers have a structural advantage hyperscalers can&rsquo;t easily replicate: customer relationships, regional presence, pricing flexibility, sovereignty positioning. They lack the cloud product to monetize this at scale.</p>
<p>Cozystack-based modernization closes the gap.</p>
<h2 id="architectural-starting-point">Architectural starting point</h2>
<p>Most hosting providers in 2026 have:</p>
<ul>
<li>Bare-metal or VPS (commercial hypervisor or vanilla KVM)</li>
<li>Per-customer manual provisioning workflows</li>
<li>Limited service catalog (VMs, maybe managed databases)</li>
<li>Custom billing or WHMCS</li>
</ul>
<p>The modernization target:</p>
<ul>
<li>Kubernetes-native multi-tenant platform (Cozystack)</li>
<li>Self-service customer-facing portal (Cozystack Dashboard or custom)</li>
<li>Expanded service catalog (VMs, K8s, managed DBs, S3, GPU)</li>
<li>WHMCS-integrated billing</li>
<li>Per-customer observability and audit</li>
</ul>
<h2 id="migration-sequencing">Migration sequencing</h2>
<ol>
<li><strong>Assessment</strong> — current platform, customer profile, service catalog gap analysis (the <a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a> is fixed-price: 14 days focused or 28 days full)</li>
<li><strong>Platform live</strong> — parallel deployment with the productized installer, internal validation</li>
<li><strong>Beta customer cohort</strong> — 3-5 friendly customers; rough edges fixed</li>
<li><strong>Limited GA</strong> — 10-50 customers, billing patterns validated</li>
<li><strong>General availability</strong> — open to market</li>
<li><strong>Specialty expansion</strong> — GPU, AI services, regional sovereignty positioning</li>
</ol>
<p>The platform itself goes live in weeks once hardware is ready. How fast steps 3-6 follow depends on your sales pace, not on the platform build. Multi-region national or operator programmes are different: plan a 3-6 month pilot, then 9-18 months to full multi-region operation.</p>
<h2 id="economics">Economics</h2>
<p>For mid-size hosting provider (1000-10000 customers):</p>
<ul>
<li><strong>Platform investment</strong> — assessment + Cozystack build + WHMCS integration</li>
<li><strong>Hardware</strong> — repurpose existing or new compute; storage; network</li>
<li><strong>Platform operations</strong> — the <a href="https://aenix.io/isp-calculator/">ISP calculator</a>&rsquo;s default model comes to about 1.3 full-time engineers at 10 nodes and about 2.6 at 40; round-the-clock on-call needs more people, or the 24×7 coverage of the Plus tier. Customer support is a separate headcount</li>
<li><strong>Customer pricing</strong> — typically 30-50% above platform raw cost</li>
</ul>
<p>Break-even depends on node count, staffing and ARPU — model it in the <a href="https://aenix.io/isp-calculator/">ISP calculator</a>. A small start on existing staff breaks even far earlier than a full programme with a dedicated team; positive economics grow as catalog adoption grows.</p>
<h2 id="common-pitfalls">Common pitfalls</h2>
<ul>
<li>Underinvesting in customer-facing portal polish</li>
<li>Inadequate billing accuracy from day 1</li>
<li>Operations team sized for 50 customers; signs 200 in Q1</li>
<li>Generic catalog instead of differentiation</li>
</ul>
]]></content:encoded></item><item><title>Financial-services cloud platforms — what TLPT readiness actually looks like in 2026</title><link>https://aenix.io/blog/2026/05/financial-services-cloud-tlpt-readiness/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/financial-services-cloud-tlpt-readiness/</guid><pubDate>Mon, 11 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Financial Services</category><category>DORA</category><category>Compliance</category><category>Sovereignty</category><category>Cozystack</category><description>What TLPT readiness under DORA actually looks like in 2026 for platform engineers at banks, insurers, and payment institutions facing a real supervisor cycle.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/financial-services-cloud-tlpt-readiness.jpg" alt=""></p><p>The DORA conversation at most financial-services organisations split
into two halves in 2024-2025. The first half — governance, policy,
ICT risk management documentation — went to legal, compliance, and
CISO teams. By the time DORA went into force on 17 January 2025,
most regulated entities had that documentation in reasonable shape.</p>
<p>The second half — making the cloud architecture <em>demonstrably</em>
DORA-compliant when a real TLPT cycle hits — has been quieter. That
silence is starting to break in 2026. Supervisor TLPT exercises are
reaching architectures that were assumed compliant when DORA went
live but had never been tested under realistic regulator scrutiny.</p>
<h2 id="what-tlpt-means-and-why-it-matters-now">What TLPT means and why it matters now</h2>
<p>Threat-Led Penetration Testing (TLPT) is required for significant
financial entities under DORA. Every three years a structured red-
team exercise is run against the live production environment by an
external test provider, with the financial entity&rsquo;s CSIRT / SOC
treated as a real defender.</p>
<p>The point is not to find vulnerabilities (any pen test does that).
The point is to demonstrate operational resilience under realistic
attack scenarios — that the entity&rsquo;s people, processes, and
infrastructure can detect, respond, and recover within DORA&rsquo;s
expected windows.</p>
<p>For platform engineers, three TLPT-readiness questions matter:</p>
<ol>
<li><strong>Detection on the reporting clock.</strong> DORA Articles 17-19 put an
initial notification of a major ICT-related incident on a short
fuse, and NIS2 Article 23 fixes a 24-hour early warning for
entities in scope for both. If your detection telemetry is tuned
for performance and not security, either window is fictional.</li>
<li><strong>Audit-trail completeness.</strong> DORA Article 6 requires a documented
ICT risk management framework you can demonstrate to a supervisor
<em>with evidence</em> — what controls were in place at the time of the
incident. Documentation isn&rsquo;t enough; running-system evidence is
the bar.</li>
<li><strong>Containment and recovery.</strong> TLPT exercises injection of real-
world attack patterns. The platform&rsquo;s network policies, identity
model, and isolation boundaries must survive realistic lateral
movement and persistence attempts.</li>
</ol>
<h2 id="what-dora-ready-architecture-actually-looks-like">What &ldquo;DORA-ready architecture&rdquo; actually looks like</h2>
<p>A defensible cloud architecture for financial services has six
properties supervisors care about:</p>
<h3 id="1-detection-telemetry-tuned-for-security">1. Detection telemetry tuned for security</h3>
<p>Most banks we engage with have rich performance telemetry and
alert-fatigued security telemetry. The alert ratio is wrong: too many
performance noise alerts, not enough signal-to-noise on the security
events that map to a reportable-incident trigger under DORA
Articles 18-19 (and NIS2 Article 23 where both apply).</p>
<p>Fix patterns:</p>
<ul>
<li>Curated alert rules tuned to MITRE ATT&amp;CK techniques relevant to
cloud infrastructure</li>
<li>Logs flowing from VictoriaLogs / VictoriaMetrics (self-hosted, in
jurisdiction) to a SIEM the SOC actually monitors</li>
<li>24×7 detection coverage with documented escalation</li>
<li>Alert hygiene as a recurring task</li>
</ul>
<p>Ænix Private Cloud Platform ships VictoriaMetrics + VictoriaLogs
configured for security-grade telemetry by default, plus security-
focused alert rules. The customer SIEM integration is engagement work.</p>
<h3 id="2-workload-identity-thats-not-a-wide-open-service-account">2. Workload identity that&rsquo;s not a wide-open service account</h3>
<p>Default Kubernetes service accounts are wide-open. For DORA-aligned
architecture, every workload has an identity, every identity is
narrowly scoped, every service-to-service call is authenticated and
authorised.</p>
<p>Pattern: SPIFFE/SPIRE for workload identity, External Secrets Operator
backed by customer HSM, Pod Security Standards enforced as policy.
Ænix engagement covers the integration; the customer&rsquo;s IdP /
workforce identity remains customer-controlled.</p>
<h3 id="3-tenant-crd-isolation-as-the-structural-answer-to-concentration-risk">3. Tenant CRD isolation as the structural answer to concentration risk</h3>
<p>Article 29 assesses concentration risk on the substantive condition
of resilience, not on contractual diversity. The Tenant CRD model provides namespace-level
isolation per business function, per data class, per criticality
tier — with per-tenant quotas, RBAC scope, observability scope,
audit-trail scope.</p>
<p>This is not &ldquo;soft multi-tenancy with hope.&rdquo; Tenant CRD is a Kubernetes-
native object reconciled by cozystack-controller; the running state
matches the spec or the operator surfaces the drift.</p>
<h3 id="4-audit-isolated-environments">4. Audit-isolated environments</h3>
<p>Critical workloads run in tenants that are explicitly audit-isolated:
separate clusters from non-production workloads, tamper-evident
logging that the customer&rsquo;s audit team can replay independently, no
shared infrastructure between audit-relevant and audit-irrelevant
workloads.</p>
<p>This costs more in infrastructure. The cost is the price of
defensible audit evidence.</p>
<h3 id="5-exit-readiness-with-tested-exit-drills">5. Exit-readiness with tested exit drills</h3>
<p>Article 28(8) requires a tested exit plan for critical-function
arrangements. Supervisors increasingly expect a partial exit drill
within the past 24 months.</p>
<p>The Cozystack-based architecture makes exit drills mechanically
simpler: workloads are standard KubeVirt VMs and Kubernetes resources.
The exit destination can be &ldquo;the same Kubernetes API on different
hardware or provider.&rdquo; Ænix&rsquo;s engagement model includes a documented
exit-drill playbook that customers run annually.</p>
<h3 id="6-supplier-chain-transparency-to-second-hop">6. Supplier-chain transparency to second hop</h3>
<p>Article 30(2)(a) demands visibility into the ICT supply chain at
least to the second hop. For the platform vendor relationship, Ænix is on the
hook — we provide an attested supplier-disclosure document that maps
upstream open-source components, security disclosure channels, and
operational dependencies.</p>
<p>Beyond the platform vendor, the customer is responsible for mapping
their own supplier chain. Ænix engagements include tooling to
inventory it but the inventory itself is customer-side work.</p>
<h2 id="where-most-financial-services-cloud-architectures-still-fall-short">Where most financial-services cloud architectures still fall short</h2>
<p>Four recurring patterns in our 2025-2026 assessments:</p>
<h3 id="gap-1-observability-quietly-leaving-the-regulator-perimeter">Gap 1: observability quietly leaving the regulator perimeter</h3>
<p>The production database is in an EU region matching the regulatory
mandate. The application running on top sends logs to a SaaS
observability vendor whose data-processing region defaults to US.
Every minute the application runs, application logs containing
transaction details, customer identifiers, protected data move to
non-compliant jurisdiction.</p>
<p>Most banks catch this only after a supervisor flag. By then, the
remediation involves either replacing the SaaS vendor (multi-quarter
project) or negotiating a regional data-processing arrangement
(possible with some vendors, slow with others).</p>
<p>The Cozystack-based architectural answer: self-hosted VictoriaMetrics +
VictoriaLogs running on the same infrastructure as the workloads.
Eliminates the residency leak.</p>
<h3 id="gap-2-exit-plan-exists-on-paper-never-tested">Gap 2: exit plan exists on paper, never tested</h3>
<p>The exit plan was written for the supervisor; nobody has rehearsed
it. Time-to-exit estimates are tabletop, not calibrated against a
drill. When the supervisor asks &ldquo;when did you last test the exit
plan?&rdquo; the answer is silence.</p>
<p>Fix: annual exit-drill rehearsal with documented outcome. Ænix
engagement provides the playbook; the customer runs the rehearsal.</p>
<h3 id="gap-3-concentration-risk-treated-as-procurement-question">Gap 3: concentration risk treated as procurement question</h3>
<p>&ldquo;We use AWS in two regions; we have a contract clause requiring
geographic diversification of our backup.&rdquo; Both true. Neither
addresses what Article 29 actually assesses — substantive
architectural resilience against single-provider failure.</p>
<p>Substantive answer: workloads use platform abstractions (Kubernetes,
KubeVirt, S3-compatible storage, standard relational databases) that
exist on multiple substrates. The exit destination is named at the
architecture level, not the legal level.</p>
<h3 id="gap-4-sub-contractor-risk-invisible-past-first-hop">Gap 4: sub-contractor risk invisible past first hop</h3>
<p>The contracted hyperscaler is documented. Its data-centre operators,
network connectivity providers, shared platform services beneath are
not. Article 30(2)(a) requires visibility to second hop.</p>
<p>For Cozystack-based architecture, the platform vendor (Ænix) discloses
upstream component sourcing. The hardware vendor is the next hop;
beyond that, customer responsibility.</p>
<h2 id="the-ænix-engagement-model-for-financial-services">The Ænix engagement model for financial services</h2>
<p>We approach the financial-services engagement differently from other
verticals, because the regulator-driver and audit-readiness drivers
shape the work.</p>
<h3 id="phase-0--discovery-and-dora-scoping">Phase 0 — Discovery and DORA scoping</h3>
<p>Confirm regulatory scope (DORA + national overlays + sectoral rules).
Confirm criticality classification of workloads. Sponsor and
supervisor-engagement contacts on customer side. Engagement model
(typical: Ænix provides advisory and escalation support; customer runs
production operations).</p>
<h3 id="phase-1--platform-readiness-assessment-with-dora-workstream">Phase 1 — Platform Readiness Assessment with DORA workstream</h3>
<p>14- or 28-day fixed-price assessment. Control-by-control architecture
review against DORA Article 6 and Articles 28-30 expectations. Output: 30-50
page report with gap analysis, prioritised remediation, timing.</p>
<h3 id="phase-2--pilot-slice-of-private-cloud-platform">Phase 2 — Pilot slice of Private Cloud Platform</h3>
<p>The first part of the build. A defined slice of critical-function
workloads migrated to Ænix Private Cloud Platform. Supervisor evidence catalogue
partially built. TLPT-readiness validated against the pilot scope.</p>
<h3 id="phase-3--full-private-cloud-platform-build">Phase 3 — Full Private Cloud Platform build</h3>
<p>Pilot and full build together take 3-12 months, depending on workload
scope and multi-DC structure; TLPT readiness then follows the bank&rsquo;s
TLPT cycle. Production-grade deployment with full compliance documentation
deliverables. Ænix participates in TLPT preparation; the test itself
is run by accredited red-team providers.</p>
<h3 id="phase-4--support-subscription">Phase 4 — Support subscription</h3>
<p>Ænix advisory and escalation under a Plus or Enterprise support tier
(see <a href="https://aenix.io/pricing/">/pricing/</a>). Access is the bank&rsquo;s choice: GitOps PR
review needs no access to production, and remote access to clusters
happens only with the bank&rsquo;s approval. Critical for bank governance.</p>
<h2 id="when-this-engagement-model-fits">When this engagement model fits</h2>
<p>Strong fit:</p>
<ul>
<li>Tier-1 or tier-2 European bank with active DORA programme</li>
<li>Insurance carrier with multi-jurisdictional regulatory exposure</li>
<li>Payment institution under DORA + PSD2/PSD3 overlap</li>
<li>Market infrastructure (CSDs, CCPs) under DORA + CSDR</li>
<li>Existing TLPT cycle that needs architecture readiness work</li>
</ul>
<p>Marginal fit:</p>
<ul>
<li>Smaller banks where the budget envelope for a multi-year programme
isn&rsquo;t yet sized for it; Public Cloud Platform with sovereignty-focused
architecture may bridge</li>
</ul>
<p>Poor fit:</p>
<ul>
<li>Banks that have already committed to a multi-year hyperscaler
programme and aren&rsquo;t reopening that decision — Ænix can advise on
specific DORA architecture gaps within the hyperscaler context, but
the full Private Cloud Platform isn&rsquo;t a fit</li>
</ul>
<h2 id="where-to-dig-deeper">Where to dig deeper</h2>
<ul>
<li><strong><a href="https://aenix.io/industries/financial-services/">Financial-services industry page</a></strong> —
the trigger-led commercial landing</li>
<li><strong><a href="https://aenix.io/solutions/dora-compliance/">DORA compliance services</a></strong> —
buyer-trigger DORA landing</li>
<li><strong><a href="https://aenix.io/products/private-cloud-platform/">Private Cloud Platform product page</a></strong> —
the product for regulated enterprises</li>
<li><strong><a href="https://aenix.io/blog/2026/05/dora-compliance-checklist-cloud-architecture/">A DORA compliance checklist for cloud infrastructure</a></strong> —
architecture-level DORA walkthrough</li>
<li><strong><a href="https://aenix.io/blog/2026/05/enterprise-edition-dora-cloud-architecture/">Private Cloud Platform — DORA and NIS2 obligations mapped to architecture</a></strong> —
product-level architectural detail</li>
<li><strong><a href="https://aenix.io/resources/dora-compliance-checklist/">DORA compliance checklist resource</a></strong> —
downloadable controls checklist</li>
</ul>
]]></content:encoded></item><item><title>Enterprise platform engineering — org design, headcount, and the failure modes at 1,000+ engineers</title><link>https://aenix.io/blog/2026/05/enterprise-platform-engineering-org-design/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/enterprise-platform-engineering-org-design/</guid><pubDate>Mon, 11 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Platform Engineering</category><category>Cozystack</category><category>Multi-tenancy</category><category>DevOps</category><description>Org design, headcount math, governance, and recurring failure modes for building a platform-engineering function at 1,000+-engineer organisations.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/enterprise-platform-engineering-org-design.jpg" alt=""></p><p>Platform engineering at 200-500 engineers is mostly a question of
&ldquo;do it well&rdquo; — define golden paths, build the IDP capability stack,
hire a platform team, ship. Platform engineering at 1,000+ engineers
is a different problem: governance across business units, consistency
without rigidity, multi-region operational coordination, regulator-
graded change management, and the political dynamics of cross-BU
infrastructure decisions.</p>
<h2 id="what-changes-at-1000-engineers">What changes at 1,000+ engineers</h2>
<p>Three structural shifts:</p>
<h3 id="1-multiple-platform-engineering-teams">1. Multiple platform-engineering teams</h3>
<p>Single platform team scales to roughly 50-100 product teams or
500-1,000 product engineers. Above that, the platform function
fragments — by domain (data platform, ML platform, infra platform),
by BU (consumer cloud, enterprise cloud, internal IT), or by
geography (region-specific platform teams under regulatory
constraints).</p>
<p>Coordination across multiple platform teams becomes its own
discipline — platform-of-platforms governance, shared standards,
escalation paths for cross-platform decisions.</p>
<h3 id="2-governance-overhead-becomes-substantial">2. Governance overhead becomes substantial</h3>
<p>Architecture decisions affect thousands of engineers and millions
of euros of recurring cost. Decision-making cannot be ad-hoc; it
needs governance — Architecture Review Boards, technology radar
processes, deprecation policies that respect 2-3 year transition
windows, cross-BU vetting.</p>
<p>At 200 engineers, &ldquo;the platform team decides&rdquo; works. At 1,000+
engineers, &ldquo;the platform team unilaterally decides&rdquo; creates
political backlash that slows adoption more than the decision
saves.</p>
<h3 id="3-regulator-graded-change-management">3. Regulator-graded change management</h3>
<p>For regulated enterprises (banks, insurers, public sector, telco,
energy, healthcare), enterprise-scale platforms operate under
audit-graded change management. Production changes go through
documented approval, are reproducible from artefacts, generate
evidence supervisors can consume.</p>
<p>The platform itself becomes a regulator-relevant object —
DORA Article 6 controls live in platform code.</p>
<h2 id="org-design-patterns">Org-design patterns</h2>
<p>Three patterns we see at enterprise scale:</p>
<h3 id="pattern-a-domain-aligned-platform-fragmentation">Pattern A: Domain-aligned platform fragmentation</h3>
<p>Separate platform teams per major engineering domain:</p>
<ul>
<li>Infrastructure platform — compute, networking, storage, identity</li>
<li>Data platform — data warehousing, ETL, real-time streams</li>
<li>ML / AI platform — GPU scheduling, model serving, feature stores</li>
<li>Application platform — runtime services, deployment automation</li>
</ul>
<p>Each platform team has its own customers (product engineering teams
that consume the relevant domain). Shared standards across via the
governance function.</p>
<p>Fits: organisations with clear engineering-domain boundaries (most
fintech, most consumer-tech at this scale).</p>
<h3 id="pattern-b-bu-aligned-platform-federation">Pattern B: BU-aligned platform federation</h3>
<p>Separate platform teams per business unit:</p>
<ul>
<li>Consumer-bank platform</li>
<li>Enterprise-bank platform</li>
<li>Wealth-management platform</li>
<li>Group-shared services platform</li>
</ul>
<p>Each BU operates its own platform with shared substrate. Federation
through governance — Architecture Review Board, technology radar,
deprecation policies.</p>
<p>Fits: organisations with strong BU autonomy (most established
financial services, most large industrial conglomerates).</p>
<h3 id="pattern-c-hybrid--shared-substrate--domain-extensions">Pattern C: Hybrid — shared substrate + domain extensions</h3>
<p>Single foundational platform (compute, networking, storage,
identity) shared across BUs. Domain-specific platform extensions
(ML platform, data platform) layered on top, operated by domain-
specialised teams.</p>
<p>Fits: organisations balancing consistency (foundational substrate)
with domain specialisation (ML, data).</p>
<p>In all three patterns, the platform-of-platforms governance is the
load-bearing piece. Without it, fragmentation produces 50 different
platforms with 50 different operational models — substantially
worse than one well-governed platform.</p>
<h2 id="headcount-math">Headcount math</h2>
<p>For enterprise-scale platform engineering, useful rule of thumb:</p>
<ul>
<li><strong>Foundational substrate platform team</strong> — 1 engineer per 50-100
product engineers. At 2,000 product engineers, that&rsquo;s 20-40
platform engineers split across infra / data / app / SRE
sub-teams.</li>
<li><strong>Domain platform teams</strong> — 5-15 engineers each, scaling with
domain-specific complexity and customer count.</li>
<li><strong>Governance + architecture function</strong> — 3-8 engineers (architects,
technology radar maintainers, deprecation managers).</li>
<li><strong>Operations and on-call</strong> — separate from build engineering;
scaled to incident volume and SLA tier.</li>
</ul>
<p>Total platform-engineering function: roughly 5-10% of total
engineering headcount in mature platform organisations. Below that,
the platform is structurally underfunded.</p>
<h2 id="governance-models">Governance models</h2>
<p>The Architecture Review Board (ARB) pattern works at enterprise
scale when designed deliberately:</p>
<h3 id="arb-membership">ARB membership</h3>
<p>Senior platform engineering leads (1 per platform team) + senior
product engineering leads (rotational, ~5 at a time) + security
lead + compliance lead + chief architect (chair).</p>
<p>The board reviews architectural decisions that affect multiple
teams, set deprecation timelines for retired capabilities, approve
introductions of new vendor / open-source dependencies, set
technology radar (Adopt / Trial / Assess / Hold).</p>
<h3 id="decision-cadence">Decision cadence</h3>
<p>Monthly ARB meeting for routine decisions. Quarterly for strategic
review. Asynchronous decision-making between meetings for time-
sensitive choices.</p>
<h3 id="what-arb-does-not-do">What ARB does NOT do</h3>
<p>ARB does not micromanage individual product team architecture
decisions within their own scope. Those remain product-team
decisions. ARB owns cross-team decisions only.</p>
<p>This boundary matters: ARBs that overreach create political
friction that undermines the governance function itself.</p>
<h2 id="where-enterprise-platform-engineering-fails">Where enterprise platform engineering fails</h2>
<p>Five failure modes recur:</p>
<h3 id="1-platform-function-as-cost-centre-not-value-driver">1. Platform function as cost centre, not value driver</h3>
<p>Platform team budgeted as overhead. Headcount restricted on cost
grounds. Investment in golden paths deferred to &ldquo;after we deliver
feature X.&rdquo; Six months later, platform velocity has degraded, but
the cost-savings narrative continues to drive the budget.</p>
<p>Fix: measure platform team value in product-engineering-velocity
terms (time-to-environment, time-to-production, error rate of
deployments). Show the impact in same metrics the CFO uses for
product team productivity.</p>
<h3 id="2-backstage-as-platform-anti-pattern">2. Backstage-as-platform anti-pattern</h3>
<p>Buy Backstage; declare platform problem solved. Backstage as a
portal works only when the underlying capabilities are actually
self-service. Enterprise organisations that buy Backstage before
the platform substrate produces beautiful catalogs over operational
chaos. Adoption stalls.</p>
<p>Fix: Backstage as the user-facing layer after the platform
substrate is real. Ænix Private Cloud Platform&rsquo;s developer self-service can be paired with Backstage
where the customer prefers; Cozystack Dashboard also works.</p>
<h3 id="3-fragmentation-without-governance">3. Fragmentation without governance</h3>
<p>Multiple platform teams emerge organically (each BU builds its own).
No shared standards. Cross-BU workload movement is hard or
impossible. Engineering hires can&rsquo;t transfer between BUs without
substantial retraining.</p>
<p>Fix: invest in governance early. ARB doesn&rsquo;t need to be heavy; it
just needs to exist and have authority.</p>
<h3 id="4-vendor-led-platform-in-a-box">4. Vendor-led &ldquo;platform-in-a-box&rdquo;</h3>
<p>Big vendor sells the customer a complete platform-engineering
solution. Customer accepts. 18-24 months later, the platform works
for the vendor&rsquo;s reference customers but not for this specific
organisation. Vendor lock-in is structural; replacement cost is
huge.</p>
<p>Fix: open-source substrate (Cozystack, vanilla Kubernetes, etc.)
with optional commercial support. Customer retains architectural
ownership.</p>
<h3 id="5-optimizing-for-engineering-elegance-rather-than-product-team">5. Optimizing for engineering elegance rather than product-team</h3>
<p>adoption</p>
<p>Architecturally beautiful platform that product teams don&rsquo;t want
to use. Adoption stalls. Platform team blames product teams; product
teams blame platform team.</p>
<p>Fix: product-team interviews as platform team&rsquo;s recurring discipline.
Measure adoption per golden path. Sunset paths that don&rsquo;t adopt.
Build paths that match what product teams actually request.</p>
<h2 id="what-ænix-enterprise-platform-engineering-delivers">What Ænix enterprise platform engineering delivers</h2>
<p>A typical engagement covers:</p>
<h3 id="workstream-1--current-state-assessment">Workstream 1 — Current-state assessment</h3>
<p>Inventory existing platform investment: teams, technology stack,
governance function, adoption metrics, time-to-environment baselines.
Often the first useful artefact is the inventory itself — most
enterprise organisations don&rsquo;t have a single document mapping the
full platform footprint.</p>
<h3 id="workstream-2--target-state-design">Workstream 2 — Target-state design</h3>
<p>Org design recommendation per the patterns above. Headcount
projections per platform team. Governance function design (ARB
charter, decision cadence, escalation paths). Technology radar
initial setup.</p>
<h3 id="workstream-3--cozystack-based-platform-substrate-where-applicable">Workstream 3 — Cozystack-based platform substrate (where applicable)</h3>
<p>Foundational substrate built on Ænix Private Cloud Platform
(for regulated organisations) or developer self-service on Private Cloud Platform (for product-focused
organisations). Multi-region, multi-DC, audit-isolated environments,
DORA / NIS2 alignment where applicable.</p>
<p>This workstream isn&rsquo;t always part of the engagement — some
customers retain existing substrate and engage Ænix for governance
and discipline work only.</p>
<h3 id="workstream-4--golden-path-roadmap">Workstream 4 — Golden path roadmap</h3>
<p>Identify the 5-15 highest-leverage golden paths for the customer&rsquo;s
specific product-team needs. Sequence them. Stage rollout. Adoption
metrics.</p>
<h3 id="workstream-5--capability-transfer-and-operational-handover">Workstream 5 — Capability transfer and operational handover</h3>
<p>Ænix engineers reduce direct involvement over time. Customer
platform engineering function absorbs ownership. Ænix support
continues for advisory and escalation (Plus or Enterprise tier for
24×7, see <a href="https://aenix.io/pricing/">/pricing/</a>).</p>
<h2 id="when-this-engagement-fits">When this engagement fits</h2>
<p>Strong fit:</p>
<ul>
<li>1,000+ engineers across multiple BUs or domains</li>
<li>Existing platform-engineering function but governance, adoption,
or cross-team consistency problems</li>
<li>Board / executive sponsorship for multi-year platform investment</li>
<li>Regulator-driven obligations (financial services, public sector,
telco, energy, healthcare)</li>
<li>Multi-region or multi-jurisdiction operational footprint</li>
</ul>
<p>Marginal fit:</p>
<ul>
<li>500-1,000 engineers — may fit a lighter developer self-service scope rather
than full enterprise platform engineering engagement</li>
</ul>
<p>Poor fit:</p>
<ul>
<li>Smaller organisations — a lighter developer self-service scope or Platform Engineering
services are the right scope</li>
<li>Single-BU organisations regardless of engineering count — the
governance overhead doesn&rsquo;t pay back</li>
</ul>
<h2 id="where-to-dig-deeper">Where to dig deeper</h2>
<ul>
<li><strong><a href="https://aenix.io/services/enterprise-platform-engineering/">Enterprise platform engineering services</a></strong> —
the commercial landing</li>
<li><strong><a href="https://aenix.io/services/platform-engineering/">Platform engineering services</a></strong> —
smaller-scope scope</li>
<li><strong><a href="https://aenix.io/services/internal-developer-platform/">Internal Developer Platform services</a></strong> —
the IDP-layer engagement</li>
<li><strong><a href="https://aenix.io/solutions/developer-self-service/">Developer self-service solution page</a></strong> —
for product-engineering-focused organisations</li>
<li><strong><a href="https://aenix.io/products/private-cloud-platform/">Private Cloud Platform product page</a></strong> —
for regulated organisations</li>
<li><strong><a href="https://aenix.io/blog/2026/05/internal-developer-platform-examples-without-backstage/">Internal developer platform — 6 patterns without Backstage lock-in</a></strong> —
six production patterns</li>
<li><strong><a href="https://aenix.io/blog/2026/05/platform-engineering-vs-devops-vs-sre/">Platform engineering maturity model</a></strong> —
five-stage, eight-dimension maturity model</li>
<li><strong><a href="https://aenix.io/blog/2026/05/idp-edition-developer-velocity-economics/">Developer self-service — developer velocity economics</a></strong> —
the IDP economic case</li>
<li><strong><a href="https://aenix.io/blog/2026/05/build-private-cloud-90-day-playbook/">Build private cloud — 90-day playbook</a></strong> —
for the substrate-build workstream</li>
</ul>
]]></content:encoded></item><item><title>Private Cloud Platform for regulated cloud — DORA and NIS2 obligations mapped to running architecture</title><link>https://aenix.io/blog/2026/05/enterprise-edition-dora-cloud-architecture/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/enterprise-edition-dora-cloud-architecture/</guid><pubDate>Sun, 10 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>DORA</category><category>Financial Services</category><category>Compliance</category><category>Sovereignty</category><category>Multi-tenancy</category><category>Cozystack</category><description>Mapping DORA ICT-risk and third-party obligations and the NIS2 Article 21(2) measures onto a defensible cloud architecture for regulated enterprises.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/enterprise-edition-dora-cloud-architecture.jpg" alt=""></p><p>Regulated-enterprise cloud architecture in 2026 is a different
conversation than it was in 2022. DORA went into force on 17 January
2025. NIS2 transposition deadline passed in October 2024. Supervisory
expectations are sharpening with each TLPT cycle. The pre-DORA pattern
— hyperscaler region with a few contractual clauses — does not survive
realistic audit scrutiny.</p>
<h2 id="the-ten-measure-areas--nis2-article-212-and-where-dora-meets-them">The ten measure areas — NIS2 Article 21(2), and where DORA meets them</h2>
<p>NIS2 Article 21(2) requires &ldquo;appropriate and proportionate technical,
operational and organisational measures&rdquo; across ten enumerated areas.
DORA reaches most of the same ground by a different route: its ICT risk
management framework (Articles 5-16, and Article 6 in particular),
incident management and reporting (Articles 17-19), and ICT third-party
risk (Articles 28-30). A regulated enterprise in scope for both is not
running two architectures — it is running one, evidenced twice. The ten
areas below are the NIS2 enumeration, with the DORA article that governs
the same control named where it differs.</p>
<p>For platform engineers, the ten areas map to concrete architecture
decisions:</p>
<h3 id="1-policies-on-risk-analysis-and-information-system-security">1. Policies on risk analysis and information system security</h3>
<p>Architecture implication: every critical-function workload has a
documented risk register entry. Threat-model artifacts are stored
alongside the workload (not in a separate &ldquo;compliance system&rdquo; that
drifts). Pod Security Standards and Kubernetes Network Policies
enforce the technical controls the risk register names.</p>
<h3 id="2-incident-handling">2. Incident handling</h3>
<p>Detection must operate within the NIS2 Article 23 reporting windows:
24-hour early warning, 72-hour incident notification, one-month final
report. DORA sets its own regime in Articles 17-19 — classification
under Article 18, reporting of major ICT-related incidents to the
competent authority under Article 19 — on deadlines fixed by the
implementing standards rather than by the article text. Build detection
to the tighter of the two; the architecture is the same either way.
The bottleneck in practice is detection telemetry — alert fatigue
masks signal. Private Cloud Platform ships with VictoriaMetrics +
VictoriaLogs and curated alert rules tuned for security, not just
performance.</p>
<h3 id="3-business-continuity">3. Business continuity</h3>
<p>RTO and RPO are documented per critical workload, <em>and tested annually</em>
with telemetry that proves the test outcome. Backup-only is
insufficient. Private Cloud Platform includes Velero + per-application
patterns (PostgreSQL PITR, Kafka snapshots, etc.) plus chaos-engineering
hooks for controlled failure injection.</p>
<h3 id="4-supply-chain-security">4. Supply chain security</h3>
<p>The most demanding requirement in DORA&rsquo;s third-party chapter
(Articles 28-30): ICT third-party arrangements inventoried, classified
by criticality, with the sub-contracting chain mapped to the <em>second
hop</em> under Article 30(2)(a). Private Cloud Platform gives you the
provider relationship Ænix can attest to (we are the platform vendor);
the open-source substrate (Cozystack) gives you transparency to the
upstream-component level. Beyond that, the customer&rsquo;s responsibility.</p>
<h3 id="5-security-in-acquisition-development-maintenance">5. Security in acquisition, development, maintenance</h3>
<p>SAST/DAST in CI, container scanning + SBOM, vulnerability handling
with documented SLA, disclosure policy published. Private Cloud Platform
provides the operator-side discipline; customer pipelines integrate.</p>
<h3 id="6-effectiveness-assessment">6. Effectiveness assessment</h3>
<p>Annual external assessment is the supervisor expectation for essential
entities. Private Cloud Platform documentation aligns with the working
formats supervisors consume — control-level mapping, evidence
catalogue, audit-trail completeness.</p>
<h3 id="7-cyber-hygiene-and-training">7. Cyber hygiene and training</h3>
<p>Out of scope for the platform itself; in scope for the customer&rsquo;s
operations team. Ænix can deliver Kubernetes Deep Dive Course as
part of the engagement (separate product).</p>
<h3 id="8-cryptography">8. Cryptography</h3>
<p>Supervisors expect customer-controlled encryption for production
storage, backups, observability data and audit logs, with key rotation
and emergency-access procedures documented. On Private Cloud Platform,
volume encryption is opt-in per storage class, and the key-management
process — including integration with a key store or HSM you already
operate — is designed with you during the engagement.</p>
<h3 id="9-human-resources-security-access-control-asset-management">9. Human resources security, access control, asset management</h3>
<p>Workload identity via SPIFFE/SPIRE or equivalent. Privileged access
management. Joiner-mover-leaver process automated. Asset register
complete and current. Private Cloud Platform ships with cozystack-controller
that maintains the asset register as a Kubernetes-native object — no
drift between policy and running state.</p>
<h3 id="10-mfa--continuous-authentication--secured-comms">10. MFA / continuous authentication / secured comms</h3>
<p>MFA on all privileged accounts at minimum, including any vendor access
you grant — Ænix support engineers included. Secured emergency comms
via the support channel itself.</p>
<h2 id="dora-articles-28-30--the-supplier-risk-dimension-that-breaks-most-setups">DORA Articles 28-30 — the supplier-risk dimension that breaks most setups</h2>
<p>DORA&rsquo;s third-party chapter is where most banks and insurers we engage with have the
biggest gaps. Three patterns recur:</p>
<h3 id="pattern-1--observability-data-quietly-leaving-the-regulator-perimeter">Pattern 1 — observability data quietly leaving the regulator perimeter</h3>
<p>The production database is hosted in an EU region matching the
regulatory mandate. The application running on top sends logs and
metrics to a SaaS observability vendor — Datadog, New Relic, Splunk
Cloud — whose data-processing region defaults to US. Application logs
containing transaction details, customer identifiers, and protected
data move to non-compliant jurisdiction every minute the application
runs.</p>
<p>DORA&rsquo;s third-party requirements apply to the <em>entire ICT
third-party arrangement</em> — observability tools included. Supervisor
audits increasingly catch this; legacy &ldquo;data classification policies&rdquo;
do not. Private Cloud Platform replaces SaaS observability with self-hosted
VictoriaMetrics + VictoriaLogs running on the same customer
infrastructure. Eliminates the most common residency leak.</p>
<h3 id="pattern-2--exit-plans-on-paper-never-tested">Pattern 2 — exit plans on paper, never tested</h3>
<p>Article 28(8) requires a documented exit plan for critical-function
arrangements with <em>tested feasibility</em>. Most entities have a plan;
fewer have rehearsed it. Supervisors are now asking for the rehearsal
within the last 24 months.</p>
<p>Private Cloud Platform&rsquo;s open-source substrate makes the exit-test
mechanically simpler: the workloads are standard KubeVirt VMs and
Kubernetes resources. The exit destination can be &ldquo;the same
Kubernetes API, different underlying hardware or provider&rdquo; rather
than a wholesale migration. Ænix&rsquo;s engagement model includes a
documented exit-drill playbook customers run annually.</p>
<h3 id="pattern-3--concentration-risk-treated-as-a-procurement-question">Pattern 3 — concentration risk treated as a procurement question</h3>
<p>Concentration risk often gets flagged then &ldquo;mitigated&rdquo; by contractual
diversity clauses. The substantive condition — workloads architecturally
diversified across multiple providers — typically isn&rsquo;t met. Article 29
assesses concentration risk on the substantive condition, not on the
procurement formality.</p>
<p>Private Cloud Platform&rsquo;s architecture inherently breaks the concentration:
the cloud provider relationship becomes hardware and bandwidth, not
platform services. Sub-contractor mapping shortens dramatically. The
sovereignty story becomes architectural rather than contractual.</p>
<h2 id="what-a-regulated-private-cloud-platform-build-adds">What a regulated Private Cloud Platform build adds</h2>
<p>Several layers that matter specifically to regulated builds:</p>
<ul>
<li><strong>Air-gap install</strong> documented and supported as a first-class
deployment mode. Updates flow through controlled channels (Harbor
mirror, customer-side artifact registry, manual approval). Suited to
classified-data and the most sensitive banking workloads.</li>
<li><strong>Multi-DC designs</strong> — regulated builds typically span two or more
datacentres, with cross-DC replication tuned for RTO/RPO targets.
VM failover between sites is a rehearsed runbook, not an automatic
switch.</li>
<li><strong>Access on your terms</strong> — advisory, runbooks and GitOps PR review
need no access to your production cluster. Where your support tier
includes it, remote access to your clusters happens only with your
approval. Critical for banks where vendor-side access is a
structural risk.</li>
<li><strong>Audit-isolated environments</strong> — separate clusters for production,
audit, and forensic copy. Audit-log retention is configurable, and
logs can be shipped to your own immutable store.</li>
<li><strong>Compliance documentation deliverables</strong> — at engagement close,
the customer receives a control-by-control evidence catalogue
aligned to DORA / NIS2 supervisor expectations.</li>
</ul>
<h2 id="what-stays-the-customers-responsibility">What stays the customer&rsquo;s responsibility</h2>
<p>Private Cloud Platform is <em>architecture-aligned with DORA / NIS2</em>. It is
not a certification. Several obligations remain on the customer:</p>
<ul>
<li><strong>Internal governance</strong> — board reporting, RM committee, ICT risk
function. Out of platform scope.</li>
<li><strong>Sectoral overlays</strong> — banking secrecy, insurance regulation,
healthcare data laws. Customer&rsquo;s responsibility to interpret
alongside DORA.</li>
<li><strong>The audit itself</strong> — Private Cloud Platform gives you defensible
architecture and evidence; running the audit cycle is your audit
team&rsquo;s work.</li>
<li><strong>Workload-specific risk decisions</strong> — Private Cloud Platform supplies
the substrate; which workloads are critical, how they&rsquo;re classified,
and what residual risk you accept is your call.</li>
</ul>
<h2 id="when-private-cloud-platform-is-the-right-answer">When Private Cloud Platform is the right answer</h2>
<p>Strong fit:</p>
<ul>
<li>You operate in a regulated sector (financial services, public sector,
healthcare, energy, telco) with DORA, NIS2, or sectoral overlay
obligations.</li>
<li>You have a board-level decision to bring critical-function workloads
off hyperscaler.</li>
<li>You can budget for a platform programme, sized at scoping.</li>
<li>You have or can hire a 5-10 engineer platform team to operate the
infrastructure.</li>
<li>You have sectoral pressure on TLPT, supplier-chain audit, or
exit-readiness in the next 12-18 months.</li>
</ul>
<p>Marginal fit:</p>
<ul>
<li>Mid-size organisations where the regulatory pressure is real but
the budget for a full programme isn&rsquo;t yet there. A narrower Private
Cloud Platform scope, or self-run Cozystack with Ænix enterprise
support, may bridge.</li>
</ul>
<p>Poor fit:</p>
<ul>
<li>Organisations without regulatory pressure. Use self-run Cozystack
with <a href="https://aenix.io/products/cozystack-enterprise-support/">Ænix enterprise support</a> — Private Cloud Platform&rsquo;s
compliance overhead doesn&rsquo;t pay back without the regulator driver.</li>
</ul>
<h2 id="engagement-structure">Engagement structure</h2>
<ul>
<li><strong>Discovery call</strong> (30 min, free)</li>
<li><strong>Platform Readiness Assessment</strong> (14- or 28-day, DORA / NIS2
workstream emphasised) — control-level gap analysis against current
architecture</li>
<li><strong>Private Cloud Platform build</strong> (3-12 months, depending on scope) —
starts with a defined pilot slice, then production-grade multi-DC
deployment with compliance documentation</li>
<li><strong>Support subscription</strong> (ongoing) — advisory, runbooks, GitOps PR
review and incident response; Plus or Enterprise tier for 24×7 (see
<a href="https://aenix.io/pricing/">/pricing/</a>)</li>
</ul>
<p>TLPT readiness and the supervisor evidence work at a tier-1 bank
continue after the platform is in production; plan them as a separate
track rather than as part of the build.</p>
<h2 id="where-to-dig-deeper">Where to dig deeper</h2>
<ul>
<li><strong><a href="https://aenix.io/products/private-cloud-platform/">Private Cloud Platform landing</a></strong> —
feature list, product-specific FAQ, customer evidence</li>
<li><strong><a href="https://aenix.io/solutions/dora-compliance/">DORA compliance services</a></strong> —
DORA-aligned engagement details</li>
<li><strong><a href="https://aenix.io/solutions/nis2-compliance/">NIS2 compliance services</a></strong> —
NIS2-aligned engagement details</li>
<li><strong><a href="https://aenix.io/resources/dora-compliance-checklist/">DORA compliance checklist resource</a></strong> —
free downloadable controls checklist</li>
<li><strong><a href="https://aenix.io/blog/2026/05/dora-compliance-checklist-cloud-architecture/">A DORA compliance checklist for cloud infrastructure</a></strong> —
longer architecture-level DORA walkthrough</li>
</ul>
]]></content:encoded></item><item><title>A DORA compliance checklist for cloud infrastructure — framework, controls, and what to demonstrate in 2026</title><link>https://aenix.io/blog/2026/05/dora-compliance-checklist-cloud-architecture/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/dora-compliance-checklist-cloud-architecture/</guid><pubDate>Sun, 10 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>DORA</category><category>Financial Services</category><category>Compliance</category><description>A working DORA checklist for cloud architecture: what Articles 21 and 28 require, where current setups fall short, and how to assess where you stand.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/dora-compliance-checklist-cloud-architecture.jpg" alt=""></p><p>The Digital Operational Resilience Act (DORA) has been in force since 17 January 2025. Across the EU&rsquo;s financial sector — banks, insurers, investment firms, payment institutions, crypto-asset service providers, and the third-party ICT providers serving them — DORA replaced a fragmented patchwork of supervisory expectations with a single, directly applicable regulation.</p>
<p>Most coverage of DORA so far has concentrated on policy and procedural requirements: incident reporting timelines, governance, risk management process. That work matters, and your CISO and legal team have been doing it for two years.</p>
<p>This article is for the people on the other side of that conversation — the platform engineers, cloud architects, and infrastructure leads who now have to translate DORA&rsquo;s expectations into running systems. Specifically: what does DORA require of your <em>cloud infrastructure</em>, what does a DORA-compliant cloud architecture actually look like, and where do current public-cloud setups typically fall short?</p>
<h2 id="what-dora-is-very-briefly">What DORA is, very briefly</h2>
<p>DORA (Regulation (EU) 2022/2554) sets binding rules across five areas:</p>
<ol>
<li><strong>ICT risk management</strong> — governance, identification, protection, detection, response, recovery.</li>
<li><strong>ICT-related incident management, classification, and reporting</strong> — including major incident notifications to competent authorities.</li>
<li><strong>Digital operational resilience testing</strong> — including, for significant entities, threat-led penetration testing (TLPT) every three years.</li>
<li><strong>Management of ICT third-party risk</strong> — Articles 28-44, the part most relevant to cloud.</li>
<li><strong>Information and intelligence sharing</strong> — voluntary frameworks for sharing cyber-threat intelligence.</li>
</ol>
<p>It applies to a broad set of in-scope financial entities and, indirectly, to ICT third-party service providers serving them. From October 2024, the European Supervisory Authorities (ESAs) finalized a long set of Regulatory Technical Standards (RTS) and Implementing Technical Standards (ITS) that put numerical and structural detail on the framework. Some of these continue to be refined.</p>
<p>For practical purposes, DORA enforcement is <em>now</em>. Compliance is not a future planning exercise.</p>
<h2 id="why-cloud-architecture-sits-at-the-centre-of-dora">Why cloud architecture sits at the centre of DORA</h2>
<p>Of the five DORA pillars, three apply directly to the cloud infrastructure layer:</p>
<ul>
<li><strong>Pillar 1 (ICT risk management)</strong> — your cloud setup is &ldquo;ICT&rdquo; by definition. Detection, recovery, segmentation, and resilience controls live there.</li>
<li><strong>Pillar 3 (operational resilience testing)</strong> — TLPT and resilience scenarios are run against the live infrastructure. The architecture has to make resilience demonstrable, not just claimed.</li>
<li><strong>Pillar 4 (third-party risk)</strong> — every hyperscaler, SaaS, and managed service in your stack is an &ldquo;ICT third-party service provider&rdquo; that brings DORA Article 28-44 obligations.</li>
</ul>
<p>Pillar 4 is where most of the new technical work concentrates. In practice, DORA forces a financial entity to look at every ICT supplier — including the public-cloud platforms its workloads run on — through a contractual, architectural, and exit-readiness lens.</p>
<h2 id="dora-article-28-and-the-supplier-risk-dimension">DORA Article 28 and the supplier-risk dimension</h2>
<p>Article 28 sets the general principles for ICT third-party risk management. Several of its requirements have direct architectural implications:</p>
<ul>
<li><strong>Risk assessment before contracting</strong> — the financial entity must assess concentration risk, sub-contractor risk, and the criticality of services received from each ICT third party.</li>
<li><strong>Monitoring throughout the contract</strong> — ongoing assessment of performance, security posture, and resilience of the third-party arrangement.</li>
<li><strong>Exit strategies</strong> — for arrangements supporting <em>critical or important functions</em>, a documented exit plan with feasibility tests.</li>
<li><strong>Contractual content</strong> — Article 30 specifies clauses that must be present, including SLAs, audit rights, sub-contracting rules, security and resilience requirements, and termination grounds.</li>
</ul>
<p>Then Article 29, and the related RTS, target <strong>concentration risk</strong> specifically. A financial entity may not concentrate critical-function ICT supply on a single third party, or on multiple third parties that share underlying infrastructure dependencies, where this concentration would impair the entity&rsquo;s resilience.</p>
<p>For most banks and insurers operating in 2026, that wording maps directly onto a familiar question: <em>we run our critical workloads in one hyperscaler region — does that satisfy DORA?</em> The honest answer is &ldquo;depends&rdquo; — but the work to demonstrate it didn&rsquo;t exist before DORA, and is non-trivial.</p>
<h2 id="what-dora-compliant-cloud-infrastructure-actually-means">What &ldquo;DORA-compliant cloud infrastructure&rdquo; actually means</h2>
<p>There is no DORA certification stamp. DORA defines obligations that must be satisfied; how a financial entity demonstrates that satisfaction is open, subject to ESAs&rsquo; supervisory expectations and the entity&rsquo;s own risk profile.</p>
<p>That said, supervisors and most large financial entities have converged on a similar set of architectural attributes that a DORA-aligned cloud setup should have.</p>
<h3 id="1-workload-portability-across-providers">1. Workload portability across providers</h3>
<p>A <em>critical or important function</em> must have a documented, tested exit plan. In practice this means the workload either runs in a way that can be moved between providers within a defined time window, or the financial entity accepts higher residual concentration risk and documents why.</p>
<p>Architectural implications:</p>
<ul>
<li>Workloads use platform abstractions that exist on at least two cloud providers (Kubernetes, KubeVirt, S3-compatible storage, standard relational databases) rather than provider-proprietary services.</li>
<li>Data formats and APIs are standardized.</li>
<li>Identity, secrets, and observability are not locked to a single provider&rsquo;s platform service.</li>
</ul>
<h3 id="2-concentration-risk-transparency">2. Concentration-risk transparency</h3>
<p>The financial entity must be able to enumerate, for each critical function, the underlying ICT supply chain — including sub-contractors and shared infrastructure dependencies (e.g., two SaaS providers running on the same hyperscaler region).</p>
<p>Architectural implications:</p>
<ul>
<li>Service catalogue maps each critical function to its underlying infrastructure layer.</li>
<li>Multi-region or multi-provider architecture for the most critical workloads, where the regulator&rsquo;s risk assessment requires it.</li>
<li>Documentation of where each data class actually resides — including backups, observability data, and CI/CD artifacts.</li>
</ul>
<h3 id="3-demonstrable-resilience">3. Demonstrable resilience</h3>
<p>Operational resilience must be tested, not just declared. For significant financial entities, TLPT exercises every three years inject real attack scenarios into live production environments. For all in-scope entities, scenario-based resilience testing is expected at least annually.</p>
<p>Architectural implications:</p>
<ul>
<li>Architecture supports realistic failure injection (chaos engineering) without unacceptable customer impact.</li>
<li>Recovery time and recovery point objectives are testable, with telemetry to prove them.</li>
<li>Backup and disaster recovery work across regions and, ideally, across providers — not just within a single hyperscaler region.</li>
</ul>
<h3 id="4-sovereignty-and-supervisory-access">4. Sovereignty and supervisory access</h3>
<p>Supervisors must have access to the information required to oversee the financial entity&rsquo;s ICT arrangements — including arrangements that involve cross-border data transfers. For some critical functions and certain regulators, that means data must remain in EU member states.</p>
<p>Architectural implications:</p>
<ul>
<li>Data residency is enforceable at the storage, backup, observability, and CI/CD-artifact layers — not just at the production storage layer.</li>
<li>Encryption posture is documented and key custody is the financial entity&rsquo;s responsibility, not the cloud provider&rsquo;s.</li>
<li>Logging and audit trails are exportable to formats supervisors can consume.</li>
</ul>
<h3 id="5-audit-and-exit-readiness">5. Audit and exit-readiness</h3>
<p>Article 30 requires audit rights for the financial entity and the supervisor. Article 28(8) requires exit plans with documented feasibility for critical-function arrangements.</p>
<p>Architectural implications:</p>
<ul>
<li>Architecture supports independent audit — security, performance, resilience — by parties other than the cloud provider&rsquo;s own first-line.</li>
<li>Exit drills (full or partial) have been performed at least once and documented.</li>
<li>The exit plan names the destination architecture, the migration sequence, and the data-portability mechanism.</li>
</ul>
<h2 id="where-current-cloud-setups-typically-fall-short">Where current cloud setups typically fall short</h2>
<p>In our work running platform readiness assessments for financial-services organizations across the EU, four gaps recur in nearly every engagement.</p>
<h3 id="gap-1--observability-data-quietly-leaves-the-regulators-perimeter">Gap 1 — observability data quietly leaves the regulator&rsquo;s perimeter</h3>
<p>The production database may be compliant. The SaaS observability stack collecting application logs from that database probably isn&rsquo;t, and the financial entity often cannot say with certainty where those logs are processed, replicated, or retained. DORA Article 28&rsquo;s data-residency expectations apply to the entire ICT third-party arrangement — observability tools included.</p>
<h3 id="gap-2--the-exit-plan-exists-on-paper-but-has-never-been-tested">Gap 2 — the exit plan exists on paper but has never been tested</h3>
<p>Article 28(8) requires an exit plan; many financial entities have one. But if the plan has never been rehearsed, the time-to-exit is fictional. Supervisors increasingly expect at least a partial exit drill within the past 24 months for critical-function arrangements.</p>
<h3 id="gap-3--concentration-risk-is-treated-as-a-procurement-question-not-an-architecture-question">Gap 3 — concentration risk is treated as a procurement question, not an architecture question</h3>
<p>Concentration risk often gets flagged, then mitigated by contractual diversity language — without any architectural change to how workloads actually depend on a single provider&rsquo;s compute, storage, identity, and network. Article 28 requires the substantive condition of resilience, not just the procurement formality.</p>
<h3 id="gap-4--sub-contractor-risk-is-invisible">Gap 4 — sub-contractor risk is invisible</h3>
<p>Hyperscalers run on data centres and connectivity providers; SaaS providers run on hyperscalers; managed services depend on shared underlying infrastructure. Article 30(2)(a) requires the financial entity to know the chain. Most do not, beyond the first hop.</p>
<h2 id="a-dora-compliance-checklist-for-cloud-architecture">A DORA compliance checklist for cloud architecture</h2>
<p>For platform engineering and infrastructure leads, this is the working checklist we use during a DORA-aligned readiness assessment. It maps to the requirements above, in plain operational language.</p>
<h3 id="workload-portability-and-exit-readiness">Workload portability and exit-readiness</h3>
<ul>
<li><input disabled="" type="checkbox"> Each critical-function workload has a named exit destination (not &ldquo;another provider&rdquo; but &ldquo;this specific architecture on this specific provider/on-prem&rdquo;).</li>
<li><input disabled="" type="checkbox"> Each critical-function workload uses platform abstractions that exist on ≥2 providers.</li>
<li><input disabled="" type="checkbox"> At least one full or substantial exit drill has been performed for each critical-function arrangement in the last 24 months.</li>
<li><input disabled="" type="checkbox"> Exit-time estimates are calibrated against the drill, not against a tabletop.</li>
<li><input disabled="" type="checkbox"> Data-portability path is documented at the schema, format, and tooling level.</li>
</ul>
<h3 id="concentration-risk">Concentration risk</h3>
<ul>
<li><input disabled="" type="checkbox"> Service catalogue maps each critical function to the cloud provider, region, and shared dependencies.</li>
<li><input disabled="" type="checkbox"> Concentration-risk position is documented per critical function and reviewed annually.</li>
<li><input disabled="" type="checkbox"> Where concentration is accepted, the residual-risk justification names the mitigations (multi-region, multi-AZ, multi-provider, on-prem fallback).</li>
<li><input disabled="" type="checkbox"> Sub-contractor chain is mapped to the second hop for all critical-function ICT third parties.</li>
</ul>
<h3 id="operational-resilience">Operational resilience</h3>
<ul>
<li><input disabled="" type="checkbox"> Recovery time objective (RTO) and recovery point objective (RPO) are documented per critical function.</li>
<li><input disabled="" type="checkbox"> RTO and RPO are tested annually with results reported.</li>
<li><input disabled="" type="checkbox"> Production architecture supports controlled failure injection without unacceptable customer impact.</li>
<li><input disabled="" type="checkbox"> Telemetry exists to prove RTO and RPO during real or simulated failure.</li>
<li><input disabled="" type="checkbox"> Backup and DR work across at least two regions; for the most critical functions, across two providers or to on-prem.</li>
</ul>
<h3 id="sovereignty-and-supervisory-access">Sovereignty and supervisory access</h3>
<ul>
<li><input disabled="" type="checkbox"> Data-residency stance is documented per data class — production, backup, observability, CI/CD artifacts.</li>
<li><input disabled="" type="checkbox"> Encryption is at-rest, in-transit, and (for the most sensitive data) in-use.</li>
<li><input disabled="" type="checkbox"> Encryption keys are under the financial entity&rsquo;s control, with documented rotation and emergency access.</li>
<li><input disabled="" type="checkbox"> Audit logs are exportable in standard formats, retained per regulator&rsquo;s requirement, and tamper-evident.</li>
<li><input disabled="" type="checkbox"> Supervisor access process is documented and tested.</li>
</ul>
<h3 id="third-party-risk-and-contracting">Third-party risk and contracting</h3>
<ul>
<li><input disabled="" type="checkbox"> All ICT third-party arrangements are inventoried, categorized as critical / non-critical, and reviewed.</li>
<li><input disabled="" type="checkbox"> Article 30 contractual content is in place for all critical-function arrangements.</li>
<li><input disabled="" type="checkbox"> Continuous monitoring of third-party performance and security posture is in place.</li>
<li><input disabled="" type="checkbox"> Concentration-risk thresholds are agreed and breach-detection is automated where possible.</li>
</ul>
<p>A real-world DORA readiness assessment goes deeper than this checklist on each line, but the checklist captures the architectural surface area an infrastructure lead will be asked about.</p>
<h2 id="how-cozystack-and-similar-kubernetes-native-architectures-help">How Cozystack and similar Kubernetes-native architectures help</h2>
<p>A DORA-aligned cloud architecture does not require a specific product. It requires the architectural attributes above — portability, concentration-risk transparency, testable resilience, sovereignty, audit-readiness.</p>
<p>Several architectural patterns make those attributes structurally easier rather than reliant on heroic operational discipline. Kubernetes-native virtualization platforms are one of them; Cozystack, the open-source CNCF Project Ænix builds, is one example.</p>
<p>What that pattern does for DORA:</p>
<ul>
<li><strong>Portability is structural.</strong> Workloads run as KubeVirt VMs and standard Kubernetes resources. The exit destination is &ldquo;the same Kubernetes API, on different underlying hardware or cloud&rdquo; — a substantially smaller migration than from a hyperscaler-proprietary service.</li>
<li><strong>Sovereignty is in the architecture, not in policy.</strong> The platform runs on the financial entity&rsquo;s chosen hardware, in the chosen jurisdiction, with the entity holding encryption keys and root cluster access. There is no provider-controlled control plane to reason about.</li>
<li><strong>Concentration risk is reduced</strong> — the cloud provider relationship becomes hardware and bandwidth, not platform services. Sub-contractor mapping shortens.</li>
<li><strong>Operational resilience is testable</strong> — the same Kubernetes primitives that allow controlled failure injection (chaos engineering) in development apply in production.</li>
<li><strong>Audit and supervisory access</strong> are simpler because the financial entity owns the platform stack: telemetry, audit trails, configuration history, and change-management records all live in the entity&rsquo;s tooling.</li>
</ul>
<p>Note what this does <em>not</em> claim: it does not say a Cozystack-based architecture is automatically DORA-compliant, or that a hyperscaler-based architecture cannot be. Both can be made to satisfy DORA&rsquo;s substantive requirements. The difference is in how much of the work is architectural-and-then-routine versus contractual-and-then-continuously-monitored.</p>
<h2 id="how-to-assess-where-you-are">How to assess where you are</h2>
<p>DORA compliance is not &ldquo;have we written a policy&rdquo; but &ldquo;can we demonstrate each control to a supervisor with evidence from the running system.&rdquo; Most financial entities want to know, before committing to a multi-year remediation, where they actually stand.</p>
<p>A focused DORA-aligned platform readiness assessment, run against your existing cloud architecture, covers four workstreams:</p>
<ol>
<li><strong>Inventory and platform maturity</strong> — workload-level mapping, including which workloads support critical functions.</li>
<li><strong>DORA gap analysis</strong> — control-by-control review of cloud-related DORA requirements against current architecture.</li>
<li><strong>Concentration and exit-feasibility</strong> — supplier-chain mapping and time-to-exit calibration.</li>
<li><strong>Resilience-testing readiness</strong> — whether your architecture supports the testing supervisors expect.</li>
</ol>
<h2 id="where-this-sits-in-the-broader-compliance-picture">Where this sits in the broader compliance picture</h2>
<p>DORA does not stand alone. In practice, financial entities operating in the EU are addressing DORA in parallel with:</p>
<ul>
<li><strong>NIS2</strong> — the EU directive on cybersecurity for essential and important entities, transposed across member states with sector-specific scoping. Many of NIS2&rsquo;s controls overlap with DORA Pillar 1 and Pillar 4; a single architectural posture can satisfy both, with mapping work.</li>
<li><strong>GDPR</strong> — data residency and processing rules that pre-date DORA but interact with it on cross-border arrangements.</li>
<li><strong>Sectoral rules</strong> — banking secrecy, insurance regulation, payments services. These vary by member state and overlay DORA without replacing it.</li>
<li><strong>Sovereign-cloud initiatives</strong> — France&rsquo;s SecNumCloud, Germany&rsquo;s BSI C5, the EU&rsquo;s emerging EUCS scheme. Where applicable, these add specific certification expectations to DORA&rsquo;s outcome-based requirements.</li>
</ul>
<p>A useful working principle: design for the most demanding overlay, then map the same controls back to DORA. The architectural surface is largely shared.</p>
<h2 id="timing--why-this-matters-in-2026-specifically">Timing — why this matters in 2026 specifically</h2>
<p>DORA&rsquo;s regulatory technical standards continue to be refined; ESA guidance is published in waves. Supervisor expectations are sharpening. The exam cycle for major financial entities — including TLPT exercises — is now reaching architectures that were assumed compliant when DORA went live in January 2025 but had never been tested under realistic regulator scrutiny.</p>
<p>For an infrastructure leader in 2026, the practical answer is to treat DORA-aligned architecture as a non-optional layer of the platform, planned and tested, rather than as a once-a-year compliance project.</p>
<hr>
<h2 id="want-to-dig-deeper">Want to dig deeper?</h2>
<ul>
<li><strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment — the engagement that includes a DORA workstream</a></strong></li>
<li><strong><a href="https://aenix.io/solutions/data-sovereignty/">Data sovereignty in 2026 — what European and APAC enterprises actually need</a></strong></li>
<li><strong><a href="https://aenix.io/products/cozystack/">Cozystack — the open-source platform we typically recommend for sovereign architectures</a></strong></li>
</ul>
]]></content:encoded></item><item><title>DevOps best practices for 2026 — beyond the slide-deck era</title><link>https://aenix.io/blog/2026/05/devops-best-practices-2026/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/devops-best-practices-2026/</guid><pubDate>Sat, 09 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Kubernetes</category><category>GitOps</category><category>AI and ML</category><category>DevOps</category><category>Observability</category><description>The eight DevOps practices that compound in 2026, what is still contested, the failure modes that recur, and how to place your team on the maturity curve.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/devops-best-practices-2026.jpg" alt=""></p><p>DevOps as a term has been around long enough that it has accumulated muddy meanings. By 2026, the actual practices that mature engineering organizations run have converged. The slide-deck era — where DevOps consulting was about transformation roadmaps drawn from Gartner reports — has moved to a smaller corner of the industry. What&rsquo;s left is engineering-grade practice, with measurable outcomes.</p>
<h2 id="what-devops-actually-is-in-2026">What DevOps actually is in 2026</h2>
<p>Three working definitions, in increasing order of usefulness:</p>
<ol>
<li>
<p><strong>DevOps as a culture.</strong> Product teams own their software end-to-end, including production operations. The cultural premise is the unit; tooling supports it. <em>(Useful for setting expectations; not useful for building anything.)</em></p>
</li>
<li>
<p><strong>DevOps as a tooling stack.</strong> CI/CD, IaC, observability, secrets management. <em>(Useful for procurement; misses the cultural and architectural pieces that make tooling effective.)</em></p>
</li>
<li>
<p><strong>DevOps as integrated engineering practice.</strong> Cultural premise + tooling + platform support + reliability discipline + organizational structure. <em>(This is what we engineer in customer engagements.)</em></p>
</li>
</ol>
<p>The mature practitioner thinks in version 3. We do too.</p>
<h2 id="the-eight-practices-that-compound">The eight practices that compound</h2>
<p>After running DevOps engagements across service providers, regulated enterprises, AI operators, and telecom operators, eight practices reliably show up as the levers that compound returns. Here&rsquo;s each, in approximate order of foundational importance.</p>
<h3 id="practice-1-everything-as-code">Practice 1: Everything-as-code</h3>
<p>All infrastructure, configuration, and operational logic in version control. No clicky-clicky in cloud consoles for production changes. No SSH-ing into machines to fix things permanently.</p>
<p>The discipline:</p>
<ul>
<li>IaC for cloud infrastructure (Terraform / OpenTofu / Crossplane)</li>
<li>GitOps for Kubernetes (Argo CD or Flux)</li>
<li>Configuration-as-code for application config</li>
<li>Policy-as-code for security and compliance (OPA / Kyverno)</li>
<li>Drift detection and self-healing where possible</li>
</ul>
<p>Returns: change auditability, rollback discipline, reduced operational variance.</p>
<h3 id="practice-2-trunk-based-development-with-continuous-deployment">Practice 2: Trunk-based development with continuous deployment</h3>
<p>Short-lived branches; mainline always deployable; production deploys multiple times per day per service. Long-lived feature branches and quarterly releases are operational debt disguised as a process.</p>
<p>The discipline:</p>
<ul>
<li>Small commits, frequently merged</li>
<li>Feature flags for incomplete features</li>
<li>Automated tests at every stage</li>
<li>Continuous deployment to non-prod automatically</li>
<li>Automated promotion to prod with human gate where needed</li>
</ul>
<p>Returns: faster feedback, smaller blast radius per change, dramatically reduced incident rate.</p>
<h3 id="practice-3-observability-over-monitoring">Practice 3: Observability over monitoring</h3>
<p>Monitoring tells you what you decided to look at; observability lets you ask new questions of your system after the fact. The 2026 standard is logs + metrics + traces, with high cardinality, queryable in production.</p>
<p>The discipline:</p>
<ul>
<li>VictoriaMetrics + VictoriaLogs (open-source, low-overhead) or Prometheus + Loki, or commercial (with sovereignty-conscious data-residency)</li>
<li>Auto-instrumentation where supported (OpenTelemetry)</li>
<li>SLOs defined per service, with error budgets</li>
<li>Alert hygiene as a recurring task — alert fatigue is operational debt</li>
</ul>
<p>Returns: faster incident diagnosis, capacity planning grounded in data, SLO discipline.</p>
<h3 id="practice-4-sre-practices--slos-error-budgets-incident-response">Practice 4: SRE practices — SLOs, error budgets, incident response</h3>
<p>The Google SRE book (2016) and SRE Workbook (2018) remain the operational reference. The 2026 reality: SRE practices are widely adopted but inconsistently. Most organizations have SLOs in some places and incident response in others; the mature ones have both as a coherent system.</p>
<p>The discipline:</p>
<ul>
<li>SLOs defined collaboratively with product teams</li>
<li>Error budgets that affect prioritization decisions</li>
<li>Blameless post-mortems with documented action items</li>
<li>Incident response with clear roles (incident commander, scribe, communicator)</li>
<li>Capacity planning informed by SLO trajectory</li>
</ul>
<p>Returns: production reliability becomes a system rather than a habit; risk is managed, not avoided.</p>
<h3 id="practice-5-security-as-a-parallel-discipline">Practice 5: Security as a parallel discipline</h3>
<p>Security shifted left for a decade; the current state is &ldquo;security shifted everywhere.&rdquo; Not a separate gate at the end; not just at design time; integrated through the lifecycle.</p>
<p>The discipline:</p>
<ul>
<li>SAST / DAST in CI</li>
<li>Container scanning + SBOM</li>
<li>Secrets scanning in code</li>
<li>Policy-as-code for runtime (OPA / Kyverno / Gatekeeper)</li>
<li>Workload identity (SPIFFE/SPIRE) for service-to-service auth</li>
<li>Supply-chain security (Sigstore, in-toto attestations)</li>
</ul>
<p>Returns: security debt that doesn&rsquo;t compound; regulator dialog easier; production posture visible.</p>
<h3 id="practice-6-platform-engineering-as-the-substrate">Practice 6: Platform engineering as the substrate</h3>
<p>DevOps within product teams scales linearly with team count. Platform engineering compounds it sublinearly: shared platform, shared golden paths, product teams ship faster on common substrate.</p>
<p>The discipline:</p>
<ul>
<li>Internal developer platform with golden paths</li>
<li>Self-service primitives for environment provisioning, deployment, observability</li>
<li>Platform team with product-engineering customers</li>
<li>Time-to-environment as a measurable platform metric</li>
</ul>
<p>Returns: linear-to-sublinear scaling of platform investment; product team velocity compounds.</p>
<p>(For depth see <strong><a href="https://aenix.io/blog/2026/05/platform-engineering-vs-devops-vs-sre/">Platform engineering vs DevOps vs SRE</a></strong>.)</p>
<h3 id="practice-7-finops-integrated-not-bolted-on">Practice 7: FinOps integrated, not bolted on</h3>
<p>Cost as a non-functional requirement, present at architecture review and continuous in production. Not a quarterly fire drill.</p>
<p>The discipline:</p>
<ul>
<li>Cost attribution per team / namespace / workload</li>
<li>Budget gates in CI/CD pipelines</li>
<li>Cost shown in deployment dashboards</li>
<li>Quarterly cost review at platform-team level</li>
<li>Repatriation evaluation as a permanent function, not a project</li>
</ul>
<p>Returns: cost trajectory predictable; FinOps doesn&rsquo;t become an emergency.</p>
<h3 id="practice-8-continuous-improvement-as-a-function">Practice 8: Continuous improvement as a function</h3>
<p>DORA metrics (deployment frequency, lead time, change-failure rate, time-to-restore). Tracked, reviewed, used to drive prioritization.</p>
<p>The discipline:</p>
<ul>
<li>Metrics dashboard visible to engineering leadership</li>
<li>Quarterly retrospective at organization level</li>
<li>Improvements tied to measurable outcomes</li>
<li>Maturity assessment annually (internal or external)</li>
</ul>
<p>Returns: improvement that compounds rather than oscillates with leadership attention.</p>
<h2 id="whats-still-contested">What&rsquo;s still contested</h2>
<p>A few practices that mature organizations disagree on:</p>
<ul>
<li><strong>Mono-repo vs poly-repo.</strong> Both work at scale. The decision is downstream of organizational structure and tooling choices.</li>
<li><strong>Centralized SRE vs embedded SRE.</strong> Both work. Centralized scales reliability discipline; embedded keeps SREs close to product context.</li>
<li><strong>Internal portal (Backstage, Port) vs IaC-first DX.</strong> Both work. Backstage is great when catalog discipline is mature; IaC-first is simpler for smaller orgs.</li>
<li><strong>Build vs buy on observability.</strong> Open-source self-hosted vs commercial SaaS. Sovereignty concerns are pushing more organizations toward self-hosted.</li>
<li><strong>Kubernetes everywhere vs Kubernetes for some workloads.</strong> Both work. The question is whether the workload mix justifies the platform investment.</li>
</ul>
<p>These are real architecture decisions, not best-practice settled answers.</p>
<h2 id="common-failure-modes">Common failure modes</h2>
<h3 id="tool-driven-devops">Tool-driven &ldquo;DevOps&rdquo;</h3>
<p>Buying Jenkins, Datadog, and Argo and calling that DevOps. Tools without practices become operational debt.</p>
<h3 id="big-bang-transformation">Big-Bang transformation</h3>
<p>&ldquo;Year-long DevOps transformation&rdquo; that delivers nothing measurable for 6+ months. Effective transformation runs in 4-week increments with named outcomes.</p>
<h3 id="devops-without-organizational-change">DevOps without organizational change</h3>
<p>Trying to instill DevOps practice in product teams without changing reporting lines, on-call ownership, or incentive structure. Practices erode under the previous incentives.</p>
<h3 id="senior-partner--junior-delivery">Senior partner / junior delivery</h3>
<p>Big-4 pattern: senior partner sells, junior consultants deliver. The customer pays senior rates for junior outputs.</p>
<h3 id="no-knowledge-transfer">No knowledge transfer</h3>
<p>Engagement ends, consultants leave, customer team can&rsquo;t operate what was built. Dependency on follow-on engagement.</p>
<h2 id="maturity-progression">Maturity progression</h2>
<p>A practical 5-stage maturity model:</p>
<ol>
<li><strong>Pre-DevOps</strong> — Dev throws over the wall, Ops catches. Manual everything. (Disappearing in 2026 but not extinct.)</li>
<li><strong>Tool-driven DevOps</strong> — CI/CD installed, IaC partially. Practices uneven.</li>
<li><strong>Practiced DevOps</strong> — Trunk-based development, automated deploys, basic observability, SRE practices in some teams.</li>
<li><strong>Platform-supported DevOps</strong> — Internal developer platform with golden paths; SRE practices systemic; FinOps integrated.</li>
<li><strong>Mature platform engineering</strong> — Platform is a coherent product; DevOps and SRE practices baked in; continuous improvement as a function.</li>
</ol>
<p>Most organizations sit at stage 2 or 3. Moving to stage 4 is the work that compounds returns.</p>
<h2 id="how-to-assess-where-you-are">How to assess where you are</h2>
<p>Annual external assessment is worth doing if your organization has more than ~50 engineers. The honest reasons:</p>
<ul>
<li>Internal teams underestimate their own technical debt (familiarity bias)</li>
<li>Industry standard is moving; what was best practice 3 years ago may not be now</li>
<li>Board / leadership benefits from independent perspective</li>
<li>Platform investment decisions need a defensible data point</li>
</ul>
<p>Ænix runs this as a <strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong> with DevOps maturity workstream emphasized. The output is a written report that names, per practice, where you stand and where the leverage is.</p>
<h2 id="want-to-dig-deeper">Want to dig deeper?</h2>
<ul>
<li><strong><a href="https://aenix.io/services/devops-consulting/">DevOps consulting services</a></strong> — engagement details</li>
<li><strong><a href="https://aenix.io/services/platform-engineering/">Platform engineering services</a></strong> — when DevOps reaches platform stage</li>
<li><strong><a href="https://aenix.io/blog/2026/05/platform-engineering-vs-devops-vs-sre/">Platform engineering vs DevOps vs SRE</a></strong> — terminology</li>
<li><strong><a href="https://aenix.io/products/cozystack/">Cozystack</a></strong> — open-source platform foundation</li>
</ul>
]]></content:encoded></item><item><title>Developer experience platforms — building self-service paths that actually get used</title><link>https://aenix.io/blog/2026/05/developer-experience-platform-self-service-paths/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/developer-experience-platform-self-service-paths/</guid><pubDate>Sat, 09 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>Backstage</category><category>Kubernetes</category><description>The ten golden paths most worth building, the five characteristics that make them work, and the architectural decisions that shape self-service.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/developer-experience-platform-self-service-paths.jpg" alt=""></p><p>Most &ldquo;developer experience&rdquo; articles in 2026 stop at &ldquo;use Backstage.&rdquo; That&rsquo;s not the answer; that&rsquo;s a tooling decision that comes after the architectural decisions. The architectural decisions determine whether self-service paths actually work.</p>
<h2 id="why-self-service-is-the-leverage">Why self-service is the leverage</h2>
<p>For a 200-engineer organization, time-to-environment of 2-5 weeks costs ~5% of engineering productivity (rough estimate from our engagements). Compressing that to hours is worth the platform investment.</p>
<p>But the savings only realize if product teams actually use the self-service paths. A platform that&rsquo;s technically self-service but operationally clunky doesn&rsquo;t get adopted; teams keep filing tickets because tickets feel reliable.</p>
<p>Adoption is the metric. Architecture decisions cascade from there.</p>
<h2 id="five-characteristics-of-golden-paths-that-work">Five characteristics of golden paths that work</h2>
<h3 id="1-faster-than-the-ticket-alternative">1. Faster than the ticket alternative</h3>
<p>The self-service path produces a working result faster than the ticket alternative. If a ticket takes 3 days and self-service takes 2 days, teams will choose the ticket because waiting is easier than learning.</p>
<h3 id="2-reliable-enough-to-trust">2. Reliable enough to trust</h3>
<p>The self-service path works the first time, every time, for the documented use case. If it breaks 1 in 10 times, teams stop trusting it.</p>
<h3 id="3-documented-in-1-page-or-less">3. Documented in 1 page or less</h3>
<p>Documentation longer than a page suggests architecture is wrong. Real golden paths are simple by design.</p>
<h3 id="4-owned-by-a-real-team">4. Owned by a real team</h3>
<p>A team that maintains the path, fields edge cases, ships improvements. Without ownership, paths decay.</p>
<h3 id="5-has-escape-hatches">5. Has escape hatches</h3>
<p>Product teams can deviate when their case is special. The escape is to a real conversation with platform team, not &ldquo;you must use the path or fail.&rdquo;</p>
<h2 id="the-10-paths-most-worth-building">The 10 paths most worth building</h2>
<p>Roughly in order of leverage:</p>
<ol>
<li><strong>Environment provisioning</strong> — dev / staging / preview / prod. The biggest single self-service win.</li>
<li><strong>Application deployment</strong> — image push → automated deploy.</li>
<li><strong>Database provisioning</strong> — managed PostgreSQL most common; MySQL, Redis next.</li>
<li><strong>Observability onboarding</strong> — auto-instrumented metrics, logs, traces.</li>
<li><strong>Object storage bucket</strong> — S3-compatible, with lifecycle policies.</li>
<li><strong>Secrets management</strong> — secret creation and access binding.</li>
<li><strong>Network access</strong> — to legacy services, shared databases, partners.</li>
<li><strong>CI/CD pipeline setup</strong> — service template + automated pipeline.</li>
<li><strong>Identity / SSO integration</strong> — workforce identity to service identity.</li>
<li><strong>Backup/DR setup</strong> — for stateful workloads.</li>
</ol>
<p>The first three (environment + deployment + database) cover ~60% of typical product-team request volume. Build them first.</p>
<h2 id="architectural-decisions-that-shape-self-service">Architectural decisions that shape self-service</h2>
<h3 id="provisioning-model">Provisioning model</h3>
<ul>
<li><strong>GitOps + IaC</strong> — product teams commit IaC manifests; platform reacts. Most common in 2026.</li>
<li><strong>API-driven</strong> — REST/RPC API exposed by platform; product teams call directly or via portal.</li>
<li><strong>Operator-driven</strong> — Kubernetes operators per resource type; product teams create CRD instances.</li>
<li><strong>Click-ops via portal</strong> — Backstage / Port / custom UI exposing operator-backed actions.</li>
</ul>
<p>Most organizations use a mix: GitOps + IaC as the source of truth, portal as the discoverability layer for non-Kubernetes-fluent teams.</p>
<h3 id="tenant-model">Tenant model</h3>
<ul>
<li><strong>Namespace per team</strong> — soft isolation, simplest, default for trust-each-other organizations.</li>
<li><strong>Cluster per team</strong> — hard isolation, operationally expensive, necessary for high-isolation cases.</li>
<li><strong>Tenant CRD per team</strong> — Kubernetes-native isolation with shared cluster, operationally efficient. Cozystack default.</li>
</ul>
<h3 id="identity-model">Identity model</h3>
<p>Workforce identity (Keycloak / Okta / Azure AD) federates to platform identity. Service identity (SPIFFE/SPIRE or service accounts) handles service-to-service. Joining the two is platform-team work.</p>
<h2 id="what-goes-wrong">What goes wrong</h2>
<h3 id="path-1-backstage-as-the-platform">Path 1: Backstage as the platform</h3>
<p>Buying Backstage before the underlying capabilities are self-service produces a beautiful catalog over the same operational mess. Adoption stalls.</p>
<h3 id="path-2-too-rigid">Path 2: too rigid</h3>
<p>Golden paths must serve 80% of cases. The remaining 20% need escape hatches; without them, teams that have specialized needs route around the platform entirely.</p>
<h3 id="path-3-unstaffed-paths">Path 3: unstaffed paths</h3>
<p>Platform team builds paths and moves on. Paths decay; bug reports pile up. Teams lose trust.</p>
<h3 id="path-4-optimizing-for-engineering-elegance">Path 4: optimizing for engineering elegance</h3>
<p>Architecturally beautiful self-service that&rsquo;s operationally clunky. Product teams care about results, not architecture.</p>
<h3 id="path-5-skipping-the-discovery">Path 5: skipping the discovery</h3>
<p>Building paths that platform team thinks are needed; turns out the actual top-10 requests are different. The fix: ask product teams what they request most often, build for those.</p>
<h2 id="implementation-sequence">Implementation sequence</h2>
<ol>
<li><strong>Inventory current ticket flow</strong> — what 10 things do product teams request most? Per-team interviews surface this.</li>
<li><strong>Pick top 3-5</strong> — covering ~60% of volume.</li>
<li><strong>Build golden path for each</strong> — typically 1-2 months elapsed per path with adequate platform-team capacity.</li>
<li><strong>Iterate on adoption</strong> — collect feedback; fix friction; add documentation.</li>
<li><strong>Add path 4-10 over subsequent quarters</strong> — based on observed demand.</li>
</ol>
<h2 id="how-to-assess-and-start">How to assess and start</h2>
<p>If self-service is on the table, the structured assessment names the top requests, where current paths fail, and what the priority build sequence looks like. Ænix runs this as part of <strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong>.</p>
<p>For details see <strong><a href="https://aenix.io/solutions/developer-self-service/">developer self-service services</a></strong> and <strong><a href="https://aenix.io/services/internal-developer-platform/">internal developer platform services</a></strong>.</p>
]]></content:encoded></item><item><title>Data residency requirements in 2026 — a practical guide for cloud architecture</title><link>https://aenix.io/blog/2026/05/data-residency-requirements-2026/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/data-residency-requirements-2026/</guid><pubDate>Fri, 08 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>DORA</category><category>NIS2</category><category>Sovereignty</category><category>Financial Services</category><category>Compliance</category><category>Backup and DR</category><description>What data residency actually requires at control level, why most cloud setups fail on inspection, and the architectural patterns that hold up.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/data-residency-requirements-2026.jpg" alt=""></p><p>Most coverage of data residency stops at &ldquo;production storage is in the right region.&rdquo; That&rsquo;s the easy part. The hard part is everywhere else — backups, observability data, CI/CD artifacts, managed-service telemetry, cross-border replication, sub-contractor processing — and it&rsquo;s where regulator audits increasingly land.</p>
<h2 id="what-data-residency-actually-means">What &ldquo;data residency&rdquo; actually means</h2>
<p>Data residency is the requirement that specified data — usually personal data, financial data, or sectoral-regulated data — be stored and processed within a defined jurisdiction. It is one component of the broader concept of data sovereignty, which also covers control of encryption keys, supplier transparency, and audit-readiness.</p>
<p>Residency requirements come from several places:</p>
<ul>
<li><strong>GDPR</strong> — for personal data of EU data subjects, with cross-border transfer rules under Articles 44-50.</li>
<li><strong>DORA</strong> — for financial-services ICT arrangements, with implications for backup, replication, and observability data.</li>
<li><strong>NIS2</strong> — for essential and important entities, with sectoral data-handling overlays.</li>
<li><strong>Sectoral rules</strong> — banking secrecy, insurance regulation, healthcare data laws (e.g., German Sozialgesetzbuch, French CNIL guidance, Italian Codice della privacy supplements).</li>
<li><strong>National data-localization laws</strong> — India&rsquo;s DPDP Act, China&rsquo;s PIPL, Russia&rsquo;s 152-FZ, Brazil&rsquo;s LGPD, several EU member-state-specific overlays.</li>
<li><strong>Procurement-mandated sovereignty</strong> — France&rsquo;s SecNumCloud-required services, Germany&rsquo;s BSI C5, Kazakhstan&rsquo;s procurement-portal sovereignty clauses, several APAC public-sector mandates.</li>
</ul>
<p>These don&rsquo;t always agree. An organization operating in 5 jurisdictions can face residency requirements that conflict on the same data class. The architectural answer is per-jurisdiction, with explicit handling for cross-border flows.</p>
<h2 id="why-most-cloud-setups-fail-residency-on-inspection">Why most cloud setups fail residency on inspection</h2>
<p>In our work running platform readiness assessments for organizations under residency pressure, four failure modes recur:</p>
<h3 id="failure-1--production-is-right-observability-is-wrong">Failure 1 — production is right, observability is wrong</h3>
<p>The production database is hosted in the EU region the regulator specified. The application running on top of it sends logs and metrics to a SaaS observability vendor — Datadog, New Relic, Splunk Cloud, etc. — whose data-processing region is, by default, US-based. Application logs containing personal data, financial transaction details, or regulated content move to a non-compliant jurisdiction every minute the application runs.</p>
<p>This is the single most common residency failure. Compliance teams catch the production storage but miss the telemetry plane. Regulator audits, increasingly, do not.</p>
<h3 id="failure-2--backups-dont-match-production-residency">Failure 2 — backups don&rsquo;t match production residency</h3>
<p>Backup storage tier is configured for the production region. The backup retention policy includes secondary copies for DR. The secondary copies, by default, replicate cross-region — often to a US region for cost reasons. Backups containing the same regulated data classes as production now live in two jurisdictions.</p>
<h3 id="failure-3--cicd-pipeline-processes-regulated-data">Failure 3 — CI/CD pipeline processes regulated data</h3>
<p>Test environments often contain anonymized or sampled production data. The CI/CD pipeline that creates them runs in whatever region the build infrastructure was deployed into — frequently a different region than production. Anonymization that&rsquo;s good enough for internal use may not satisfy residency rules at the regulator level.</p>
<h3 id="failure-4--sub-contractor-chain-is-invisible">Failure 4 — sub-contractor chain is invisible</h3>
<p>The hyperscaler is in the contract and named on the data-residency attestation. The hyperscaler&rsquo;s data-centre operator, network connectivity providers, and shared platform services are not. Article 30(2)(a) of DORA, equivalent NIS2 provisions, and several sectoral rules require this transparency to the second hop. Most organizations cannot answer it.</p>
<h2 id="a-control-level-checklist">A control-level checklist</h2>
<p>For each in-scope data class, residency must be explicit at every layer. Use this checklist as the working architecture audit.</p>
<h3 id="production-storage-and-processing">Production storage and processing</h3>
<ul>
<li><input disabled="" type="checkbox"> Production storage region is defined per data class and matches regulator requirements.</li>
<li><input disabled="" type="checkbox"> Compute that processes the data runs in the same region (or in a region permitted by cross-border transfer rules).</li>
<li><input disabled="" type="checkbox"> Managed services (databases, queues, caches, search) are configured to the same region.</li>
<li><input disabled="" type="checkbox"> Region configuration is enforced via IaC, not just runtime defaults — drift detection in place.</li>
</ul>
<h3 id="backup-and-disaster-recovery">Backup and disaster recovery</h3>
<ul>
<li><input disabled="" type="checkbox"> Backup storage tier is in a permitted region.</li>
<li><input disabled="" type="checkbox"> Cross-region backup replication is documented and either disabled or restricted to permitted regions.</li>
<li><input disabled="" type="checkbox"> DR site is in a permitted region.</li>
<li><input disabled="" type="checkbox"> Backup retention policies don&rsquo;t generate copies to non-permitted regions.</li>
<li><input disabled="" type="checkbox"> Backup access is audited.</li>
</ul>
<h3 id="observability-logging-monitoring">Observability, logging, monitoring</h3>
<ul>
<li><input disabled="" type="checkbox"> Log destination region is in scope.</li>
<li><input disabled="" type="checkbox"> Metrics and traces destination region is in scope.</li>
<li><input disabled="" type="checkbox"> SaaS observability vendor data-processing regions are documented and confirmed.</li>
<li><input disabled="" type="checkbox"> If using a SaaS vendor, data-processing-agreement covers residency in the contract, not just the marketing page.</li>
<li><input disabled="" type="checkbox"> Sampled / anonymized production data in observability tools is documented.</li>
</ul>
<h3 id="cicd-and-development-environments">CI/CD and development environments</h3>
<ul>
<li><input disabled="" type="checkbox"> Build and test infrastructure runs in a permitted region.</li>
<li><input disabled="" type="checkbox"> Test data with regulated content is anonymized to a documented standard, with the anonymization process auditable.</li>
<li><input disabled="" type="checkbox"> Dev / staging environments do not contain unmasked production data.</li>
<li><input disabled="" type="checkbox"> Source-code-management region is documented (especially for sectoral rules that include source code).</li>
</ul>
<h3 id="cross-border-flows">Cross-border flows</h3>
<ul>
<li><input disabled="" type="checkbox"> All cross-border data flows are inventoried, with the legal basis (e.g., GDPR Standard Contractual Clauses, BCRs) per flow.</li>
<li><input disabled="" type="checkbox"> Cross-border flows are minimized — happens only when justified by a documented business need.</li>
<li><input disabled="" type="checkbox"> Cross-border flow audit is auditable through telemetry.</li>
</ul>
<h3 id="encryption-and-key-custody">Encryption and key custody</h3>
<ul>
<li><input disabled="" type="checkbox"> Encryption at-rest is in place for all data classes, with documented key-management strategy.</li>
<li><input disabled="" type="checkbox"> Encryption keys are held by the customer, not by the cloud provider, where regulators require this.</li>
<li><input disabled="" type="checkbox"> Key rotation, emergency access, and audit are documented.</li>
<li><input disabled="" type="checkbox"> Encryption-in-transit covers all data flows, including internal cluster traffic.</li>
</ul>
<h3 id="supplier-chain-transparency">Supplier-chain transparency</h3>
<ul>
<li><input disabled="" type="checkbox"> All ICT third-party providers are inventoried and classified by criticality.</li>
<li><input disabled="" type="checkbox"> Sub-contractor chain is mapped to the second hop for critical-function arrangements.</li>
<li><input disabled="" type="checkbox"> Concentration-risk position is documented and reviewed.</li>
<li><input disabled="" type="checkbox"> Supplier-chain residency is checked when adding new vendors.</li>
</ul>
<h3 id="audit-and-supervisory-access">Audit and supervisory access</h3>
<ul>
<li><input disabled="" type="checkbox"> Audit logs are tamper-evident and exportable.</li>
<li><input disabled="" type="checkbox"> Retention period meets the longest applicable regulatory requirement.</li>
<li><input disabled="" type="checkbox"> Supervisor access process is documented and tested.</li>
<li><input disabled="" type="checkbox"> Cross-border supervisor cooperation processes are documented where applicable.</li>
</ul>
<p>A structured assessment goes deeper on each line. The list above captures the architectural surface area that an infrastructure lead will be asked about during regulator audits.</p>
<h2 id="common-architectural-patterns-that-work">Common architectural patterns that work</h2>
<p>A handful of architectural patterns satisfy residency requirements consistently across regulators.</p>
<h3 id="region-aligned-virtualization-with-controlled-replication">Region-aligned virtualization with controlled replication</h3>
<p>Production VMs run on infrastructure pinned to the regulator-required region. Replication is configured at the storage layer (LINSTOR, Ceph, vendor SAN) within the region or to a designated DR site in a permitted region. KubeVirt provides a Kubernetes-native abstraction that&rsquo;s easier to audit than legacy hypervisor sprawl.</p>
<h3 id="self-hosted-observability-stack">Self-hosted observability stack</h3>
<p>Replace SaaS observability vendors with a self-hosted stack — VictoriaMetrics for metrics, VictoriaLogs for logs, on the same infrastructure as production workloads. Eliminates the most common residency leak. We default to this in regulated-customer engagements.</p>
<h3 id="customer-controlled-key-management">Customer-controlled key management</h3>
<p>Encryption keys held in customer-operated HSM (or sovereign-cloud HSM with documented controls), with rotation and audit-trail enforced by the platform — not by the cloud provider. Key escrow for emergency access is documented and tested.</p>
<h3 id="multi-region-tenancy-with-explicit-cross-border-controls">Multi-region tenancy with explicit cross-border controls</h3>
<p>For multinational enterprises, per-tenant region pinning. Cozystack&rsquo;s Tenant CRD model supports nested tenants with per-region scoping, allowing single-platform multi-region operation under unified policy enforcement.</p>
<h3 id="air-gapped-or-restricted-egress-architecture">Air-gapped or restricted-egress architecture</h3>
<p>For the most sensitive workloads — public-sector classified, healthcare — the platform itself runs without internet egress, with software updates delivered through controlled channels. KubeVirt + Cozystack supports air-gapped deployments out of the box.</p>
<h2 id="what-residency-does-not-solve">What residency does not solve</h2>
<p>A few honest limits worth naming:</p>
<ul>
<li><strong>Residency does not solve sovereignty fully.</strong> Even with all data in the right region, encryption keys held by a non-sovereign provider, supplier-chain dependencies in a non-sovereign jurisdiction, or audit-process gaps can still fail the broader sovereignty test.</li>
<li><strong>Residency does not solve performance for distributed customers.</strong> A single-region architecture serving customers across continents will have latency issues. Multi-region architecture with explicit cross-border controls is the answer; &ldquo;everything in one region&rdquo; is not.</li>
<li><strong>Residency does not eliminate cross-border legal risk.</strong> US-headquartered cloud providers operating EU regions remain subject to US-government data requests under CLOUD Act and similar regimes. Some regulators consider this a residual risk that requires sovereign-cloud arrangements rather than hyperscaler regions.</li>
</ul>
<h2 id="country-specific-notes">Country-specific notes</h2>
<p>The residency landscape differs by jurisdiction. A few headline patterns:</p>
<h3 id="eu-and-eea">EU and EEA</h3>
<p>GDPR + DORA + NIS2 + sectoral rules. Member-state-specific overlays (German BSI C5, French SecNumCloud, Italian sectoral). The EU Cloud Certification Scheme (EUCS) is in finalization. Practical default: data and processing within EU/EEA, customer-controlled encryption, supplier-chain to second hop, with explicit cross-border-transfer documentation under SCCs or BCRs.</p>
<h3 id="uk">UK</h3>
<p>UK GDPR + sectoral rules (FCA, PRA for financial services). Adequacy decision with EU as of writing. Some divergence from EU under future UK Data Protection regimes.</p>
<h3 id="united-states">United States</h3>
<p>Sectoral, not general. HIPAA for health, GLBA for financial, FedRAMP for federal. State-level laws (California CCPA, Virginia VCDPA, etc.) increasingly impose data-handling rules. No single national data-residency mandate.</p>
<h3 id="kazakhstan-and-central-asia">Kazakhstan and Central Asia</h3>
<p>Procurement-mandated sovereignty for public-sector and quasi-public organizations. Practical procurement portal channels: goszakup.gov.kz, mitwork.kz, zakup.sk.kz.</p>
<h3 id="india">India</h3>
<p>DPDP Act 2023 introduces explicit data-localization for sensitive data classes, with implementing rules being finalized.</p>
<h3 id="china-russia">China, Russia</h3>
<p>Strong national data-localization regimes with explicit residency mandates and limited cross-border exceptions. Operating in these markets requires architecture-level segmentation.</p>
<h3 id="brazil">Brazil</h3>
<p>LGPD with cross-border transfer rules; localization is selective rather than universal.</p>
<p>For a multinational organization, the residency landscape is not a single rule — it&rsquo;s a matrix of jurisdictions with overlapping and sometimes contradictory requirements. The architecture answer is per-jurisdiction tenant boundaries with explicit cross-border controls, not a global policy.</p>
<h2 id="how-to-assess-where-you-actually-stand">How to assess where you actually stand</h2>
<p>Most organizations have a residency policy. Few can demonstrate it under audit-grade scrutiny. The structured next step is a focused assessment of the running architecture — not the policy.</p>
<p>A residency-emphasized engagement covers:</p>
<ol>
<li><strong>Data-class inventory and residency mapping</strong> — every regulated data class to its actual residency at every layer.</li>
<li><strong>Layer-by-layer audit</strong> — production, backup, observability, CI/CD, encryption, supplier-chain.</li>
<li><strong>Cross-border flow inventory</strong> — what crosses borders, why, under what legal basis.</li>
<li><strong>Gap analysis</strong> — where current architecture fails residency, prioritized.</li>
<li><strong>Remediation plan</strong> — what to fix, in what order, with effort estimates.</li>
</ol>
<p>Ænix runs this as part of the <strong><a href="https://aenix.io/services/platform-readiness-assessment/">Platform Readiness Assessment</a></strong>, with the sovereignty-and-regulator-gap workstream emphasized for residency scope. The output is a written report that names, per data class and per layer, what&rsquo;s compliant, what&rsquo;s not, and what an architecture-level remediation looks like.</p>
<hr>
<h2 id="want-to-dig-deeper">Want to dig deeper?</h2>
<ul>
<li><strong><a href="https://aenix.io/solutions/data-sovereignty/">Data sovereignty services page</a></strong> — engagement details</li>
<li><strong><a href="https://aenix.io/solutions/dora-compliance/">DORA compliance for cloud infrastructure</a></strong> — regulatory-adjacent trigger</li>
<li><strong><a href="https://aenix.io/solutions/cloud-repatriation/">Cloud repatriation</a></strong> — when sovereignty + cost align</li>
<li><strong><a href="https://aenix.io/products/cozystack/">Cozystack</a></strong> — sovereign-by-architecture platform we typically recommend</li>
</ul>
]]></content:encoded></item><item><title>Cozystack vs VMware — deep-dive comparison for platform engineers</title><link>https://aenix.io/blog/2026/05/cozystack-vs-vmware-deep-dive/</link><guid isPermaLink="true">https://aenix.io/blog/2026/05/cozystack-vs-vmware-deep-dive/</guid><pubDate>Thu, 07 May 2026 00:00:00 +0000</pubDate><dc:creator>Aenix Team</dc:creator><category>VMware</category><category>Kubernetes</category><category>Cozystack</category><category>KubeVirt</category><category>Cilium</category><category>LINSTOR</category><description>Cozystack against VMware layer by layer — compute, storage, network, multi-tenancy — with the operational implications and migration patterns for each.</description><content:encoded><![CDATA[<p><img src="https://aenix.io/img/blog/covers/cozystack-vs-vmware-deep-dive.jpg" alt=""></p><p>This article assumes familiarity with both platforms. For broader VMware exit guidance see <strong><a href="https://aenix.io/alternatives/vmware-alternative/">VMware alternative</a></strong> or <strong><a href="https://aenix.io/migration/vmware/">VMware migration</a></strong>.</p>
<h2 id="compute-layer">Compute layer</h2>
<p><strong>VMware vSphere/ESXi:</strong> mature type-1 hypervisor. Strong VM lifecycle, live migration with shared storage, vMotion. Tight VMware Tools integration with guests.</p>
<p><strong>Cozystack KubeVirt:</strong> qemu/KVM wrapped in Pods. KVM itself is a kernel-mode hypervisor, so the guest still runs on hardware virtualization extensions — the Pod is a scheduling and lifecycle wrapper, not an extra emulation layer. Live migration (CPU; GPU live migration is industry-wide limitation). Standard QEMU/KVM under the hood; broad guest OS support.</p>
<p>In practice both deliver production-grade VM workloads. The KubeVirt model adds Kubernetes operational integration (declarative VM config, GitOps lifecycle, native ingress, observability).</p>
<h2 id="storage-layer">Storage layer</h2>
<p><strong>VMware vSAN:</strong> software-defined storage built into vSphere. Operationally smooth; tight integration. Tied to VMware licensing.</p>
<p><strong>Cozystack LINSTOR:</strong> open-source replicated block storage, deployed through the Piraeus operator. LINSTOR uses DRBD for synchronous replication; object storage is a separate layer (SeaweedFS). More operational responsibility; more architectural flexibility.</p>
<p>For most workloads, LINSTOR matches vSAN in operational characteristics. Where S3-style object storage is also needed, Cozystack ships SeaweedFS as the managed Bucket service.</p>
<h2 id="network-layer">Network layer</h2>
<p><strong>VMware NSX:</strong> software-defined networking. L2 distributed virtual switches, L3 routing, micro-segmentation, edge gateway. Mature; complex.</p>
<p><strong>Cozystack Cilium:</strong> eBPF-based CNI with L4/L7 policies, observability, service mesh integration, MetalLB / BGP. Newer architecture; often simpler.</p>
<p>Migration from NSX-heavy environment to Cilium requires policy redesign — not a 1:1 mapping. The architectural model is different.</p>
<h2 id="multi-tenancy-layer">Multi-tenancy layer</h2>
<p><strong>VMware vCloud Director:</strong> mature multi-tenant overlay on vSphere. Service-provider features (organization, vDC, catalogs).</p>
<p><strong>Cozystack Tenant CRD:</strong> Kubernetes-native multi-tenant abstraction. Nested tenants, per-tenant quotas, scoped audit, billing-friendly.</p>
<p>Tenant model is conceptually different — vCD organizations vs Tenant CRD instances. Migration requires re-mapping tenant structure to Kubernetes-native equivalent.</p>
<h2 id="operational-implications">Operational implications</h2>
<h3 id="daily-operations">Daily operations</h3>
<p><strong>VMware:</strong> vCenter UI for ad-hoc operations; PowerCLI / Ansible for automation. SSH not the default model.</p>
<p><strong>Cozystack:</strong> kubectl + GitOps as the default model. Cozystack Dashboard UI for tenant operations. GitOps PR review for change-management.</p>
<p>The shift from vCenter-centric to kubectl-centric is a real operational learning curve for VMware-trained teams. Most engineers ramp in 4-8 weeks with focused training.</p>
<h3 id="upgrades">Upgrades</h3>
<p><strong>VMware:</strong> vCenter upgrade, then ESXi upgrade per host (rolling). Mature process.</p>
<p><strong>Cozystack:</strong> Talos OS upgrade + Kubernetes upgrade + Cozystack operator upgrade. GitOps-driven. Rolling per host.</p>
<p>Both rolling-upgrade. Operationally similar in spirit; tooling different.</p>
<h3 id="backup--dr">Backup / DR</h3>
<p><strong>VMware Site Recovery Manager:</strong> mature DR orchestration. Tested at scale.</p>
<p><strong>Cozystack Velero + per-app PITR:</strong> Velero handles cluster-level backup; per-app patterns (PostgreSQL PITR, etc.) on top. More moving parts; more flexibility.</p>
<p>For mission-critical DR, both work. The pattern is different — SRM is plug-and-play vendor-managed orchestration; the Velero stack is backup and restore with rehearsed runbooks, more transparent and tunable, and there is no SRM-style orchestrated cross-site failover.</p>
<h2 id="migration-patterns">Migration patterns</h2>
<p>VMware → Cozystack migration in production:</p>
<ol>
<li><strong>Discovery</strong> — vSphere/VCF inventory; workload classification.</li>
<li><strong>Cozystack foundation</strong> — parallel deployment; not a tenant of VMware.</li>
<li><strong>Image migration</strong> — KubeVirt CDI imports VMDK or qcow2 images. Windows VMs get VMware Tools cleanup before first KubeVirt boot.</li>
<li><strong>Network cutover</strong> — VLAN mapping into Cilium; policy parity validated against NSX rules.</li>
<li><strong>Storage cutover</strong> — vSAN → LINSTOR (DRBD); data migration during cohort cutover.</li>
<li><strong>DR cutover</strong> — Velero backup and restore with rehearsed runbooks takes over from SRM-based plans (no orchestrated cross-site failover); tested per cohort.</li>
<li><strong>VMware decommission</strong> — staged as cohorts complete.</li>
</ol>
<p>Typical elapsed time: about 8-12 months for ~100 VMs and 18-24 months for ~1,000 VMs, including planning and migration waves; mid-size estates fall in between, depending on dependencies. The driver is rarely raw copy speed — it is regression testing and the parallel-run windows application owners will agree to.</p>
<h2 id="when-the-comparison-matters">When the comparison matters</h2>
<p>This level of detail is useful when:</p>
<ul>
<li>Architecture review is in progress</li>
<li>Phase 2 implementation planning</li>
<li>Specific operational concerns (storage performance, network latency, etc.)</li>
<li>Team training planning</li>
</ul>
<p>For higher-level evaluation, <strong><a href="https://aenix.io/alternatives/vmware-alternative/">VMware alternative</a></strong> is more appropriate starting point.</p>
]]></content:encoded></item></channel></rss>