Own the baseline, rent only the peaks. Cloud bursting lets you run steady GPU workloads on hardware you control and spill inference or training spikes into public or sovereign clouds on demand — then tear the extra capacity down. Ænix builds this as GPU-as-a-service on a single Kubernetes platform, so your teams get elastic GPU without hyperscaler lock-in, opaque billing, or a full migration.
Pairs with: Ænix AI Platform — multi-tenant GPU scheduling, fractional sharing and ready blueprints for inference and fine-tuning. For the elastic self-service cloud underneath it, combine with Public Cloud Platform. Model the numbers with the ROI & TCO calculators.
What you get
GPU cloud bursting on the Ænix platform is one elastic GPU pool spread across the infrastructure you already have and the clouds you want to reach.
- Burst to public and sovereign clouds. Baseline workloads run on owned bare metal. When demand spikes, capacity is added in a public hyperscaler, a sovereign cloud, or both — and released afterwards. A sovereign cloud can be a first-class burst target when a regulator binds GPU processing to a jurisdiction or when its GPUs are simply cheaper.
- Fractional GPU sharing. With HAMi on top of the NVIDIA GPU-operator, several jobs share one physical card. A notebook, a small inference endpoint and a batch job can co-exist on a single GPU instead of each pinning a whole device.
- One Cluster API. Bare metal, hyperscaler and sovereign cloud sit behind a single Cluster API. Teams request GPU the same way everywhere; the platform decides where it lands.
- Autoscaling that respects GPUs. The Cluster Autoscaler adds GPU nodes when pods are unschedulable and removes them when the peak passes — so you pay for peak capacity only while the peak lasts.
- An encrypted mesh across sites. A WireGuard mesh stitches every site and cloud into one pod and service network, with new nodes auto-registering as they come up.
- Per-tenant isolation. Every tenant gets its own hosted control plane, so arbitrary user code and multi-tenant GPU sharing don’t compromise the platform.
Who is this for?
AI/ML teams with spiky training and inference demand, research institutions and universities running shared GPU for classes and experiments, and platform operators who want to offer GPU-as-a-service without reselling a hyperscaler. If your GPU demand is flat and predictable, you may not need bursting — buy for the baseline and stop. If it spikes, bursting is where the economics live.
How it works
The pattern is standard Kubernetes primitives, assembled and operated end-to-end.
- Cluster Autoscaler watches for GPU pods that cannot be scheduled and provisions nodes on the right target — bare metal, hyperscaler or sovereign cloud — through the Cluster API, Kubernetes’ declarative standard for lifecycle-managing clusters and machines. When the queue drains, the nodes are removed.
- Cilium plus a WireGuard mesh (Kilo) provide the CNI and an encrypted overlay that spans clouds. Freshly autoscaled nodes advertise themselves into the mesh and reach shared storage with no manual steps — the Kubernetes networking model treats them as if they were local.
- NVIDIA GPU-operator handles driver installation, device discovery and passthrough on each node, and HAMi adds fractional sharing so one card serves several pods.
- Talos Linux and Kamaji form the base: an immutable, API-managed OS for the nodes and hosted control planes for tenant clusters, so each tenant is isolated by design.
This is the same class of open, CNCF-aligned building blocks the cloud-native ecosystem standardizes on — no proprietary orchestration layer, no per-GPU control-plane tax.
The economics
GPU is the scarce, expensive resource, and its price is under pressure: GPU prices are volatile and have spiked sharply in short windows. Owning the baseline and bursting the peaks — rather than renting GPU 24/7 in a hyperscaler — is precisely where that pressure is absorbed.
In the academic multi-cloud case study, a European academic-computing SaaS moved its backend and user workloads off a public hyperscaler onto owned bare metal on Cozystack, kept a single Cluster API across bare metal, a hyperscaler and a sovereign Swiss OpenStack cloud, and burst GPU on demand. GPU on the sovereign cloud came out roughly 5x cheaper than the prior hyperscaler setup — with fractional sharing and per-tenant isolation intact, and no downtime for thousands of active users.
Your mix of baseline, peak and burst target decides the saving. Model it with the ROI & TCO calculators before you commit to hardware or a burst-target contract.
Ænix is the team behind Cozystack — a CNCF project (Sandbox today; Incubating expected late summer 2026), Apache 2.0. Ænix commercializes it as Ænix Platform, as three platforms on one engine — Public Cloud, Private Cloud and AI — that combine rather than exclude each other. We build multi-cloud GPU platforms for AI/ML, research and platform-operator organizations across the EU and DACH.