Turn your NVIDIA servers into a GPU cloud you can sell. Tenants order GPU virtual machines, Kubernetes clusters with GPU nodes, managed databases and S3 storage from your own branded portal. Each tenant stays isolated, and its usage goes to WHMCS or your own billing system. It runs on hardware you own, with an open-source CNCF project underneath.
Pairs with: Ænix Public Cloud Platform, the platform for organisations that sell cloud, and Ænix AI Platform for model serving and AI services on top. Try the customer portal in the live demo.
Who is this for?
- Data centres with GPU racks that want to sell GPU cloud, not only colocation or dedicated servers.
- New GPU clouds (neoclouds) with a few hundred to a few thousand NVIDIA GPUs and a customer base that expects self-service.
- Hosting providers and telcos adding GPU to an existing cloud product.
- National and regional cloud programmes that need GPU capacity to stay in the country.
If you only rent whole servers to a handful of customers on long contracts, a bare-metal provisioning tool may be enough. This platform is for when customers order, scale and pay for GPU themselves.
What does a GPU cloud need beyond the GPUs?
Tenants and isolation Every customer is a tenant with its own quotas, access rights, network isolation and monitoring. Resellers get nested tenants for their own customers.
Self-service ordering A branded customer portal with registration, team management and support tickets. Customers create GPU VMs, Kubernetes clusters and databases without writing YAML or opening a ticket.
Usage you can bill Usage is recorded per tenant and handed to your billing. Overdue tenants can be suspended automatically, without an engineering ticket.
Services next to the GPU Customers training or serving models also need storage, databases and queues. Selling them on the same platform raises revenue per GPU customer.
How are GPUs allocated to tenants?
Each mode has a different level of isolation, so choose it per product rather than per cluster.
| Mode | How it works | Isolation | Status |
|---|---|---|---|
| Whole GPU to a VM | One or more GPUs passed through (VFIO) to a tenant’s KubeVirt virtual machine | The tenant has the card to itself | Shipping |
| NVIDIA vGPU to a VM | A card split into vGPU profiles for several VMs | A separate vGPU per VM; requires your NVIDIA vGPU licence | Shipping |
| GPU nodes in tenant Kubernetes | Kubernetes node groups with GPUs, driver managed by the NVIDIA GPU Operator | Per tenant cluster | Shipping |
| Fractional sharing in Kubernetes | HAMi lets several containers share one GPU with memory and compute limits | Shared card; compute limits need container images with glibc older than 2.34 | Shipping (opt-in) |
| MIG partitions | Hardware partitions of one card | Hardware-level | Roadmap |
| Time-slicing | Containers take turns on one card | None between workloads | Roadmap |
GPU support covers NVIDIA data-centre GPUs through the NVIDIA GPU Operator. Ænix submitted the GPU Operator stack for NVIDIA partner validation in October 2026; the review is pending. For other accelerators, PCI passthrough to VMs is the supported path.
What do your customers get?
- GPU virtual machines with Linux or Windows, including custom images and templates.
- Managed Kubernetes with GPU node groups. Each tenant cluster has its own control plane.
- Managed databases and queues: PostgreSQL, MariaDB, Valkey, Kafka, ClickHouse, RabbitMQ, NATS, MongoDB, OpenSearch, and Qdrant as a vector database.
- S3-compatible object storage for datasets and model checkpoints.
- AI services delivered with Ænix AI Platform: model serving and the AI stack around it. A telecom operator and integrator, for example, runs NVIDIA Dynamo inference and RAG on Qdrant on its platform, packaged as Cozystack services (case study).
Kubernetes for AI, with third-party proof
Cozystack is a CNCF Certified Kubernetes distribution. In September 2026 it was accepted into the CNCF Kubernetes AI Conformance program for Kubernetes v1.35, which checks that a platform runs AI workloads the way the Kubernetes community specifies. Conformance details and how to reproduce the runs are on the Kubernetes conformance page.
How do billing and the portal work?
- WHMCS. The Ænix WHMCS integration, a proprietary Ænix module, sells GPU VMs, Kubernetes, databases and storage as WHMCS products. It works in two modes: WHMCS as the customer storefront, or the Ænix portal as the storefront with WHMCS as the billing back-end. More on WHMCS →
- Your own billing. Providers with their own system take per-tenant usage from the platform. A Swiss provider on Cozystack bills dedicated resources hourly from its in-house system (case study).
- Ænix billing. Ænix Public Cloud Platform also includes a billing back-end and front-end with payment processing, for providers that have neither.
You set the price per GPU, per hour or per bundle. The platform supplies usage per tenant; it does not dictate your price list.
Sovereignty and control
- Your hardware, your jurisdiction. The platform runs on your servers. Customer data and model weights stay in your data centre.
- Air-gapped installation is supported, and telemetry is off unless you turn it on.
- Open source underneath. Cozystack is Apache 2.0. If you stop the Ænix subscription, the platform keeps running.
- Ænix as a supplier: AENIX s.r.o. holds ISO/IEC 27001:2022 certification.
How does a launch run?
Discovery call
30 minutes, free. Your GPU fleet, your customers, what you sell today.
Platform Readiness Assessment
Fixed price, 14 or 28 days. Target architecture, network and storage design, GPU product modes, risk register.
Install
Productized installer: live in weeks once hardware is ready. Multi-region programmes start with a 3–6 month pilot.
First tenants
Portal branding, service catalogue, billing connection, then onboarding.
Run and grow
Support under SLA. Multi-region programmes reach full scale in 9–18 months after the pilot.
Network fabric design for multi-node training (InfiniBand, RoCE) is scoped in the assessment for your hardware, not sold as a packaged feature.
Where is it running?
These deployments are written up in full, with the customers anonymised as their contracts require:
- A sovereign public cloud on bare metal: a Swiss provider sells VMs, Kubernetes and GPUs from three data centres, with billing from its own system.
- 8×H100 inference on your own bare metal: all eight GPUs passed through to one isolated tenant VM, about two months to production.
- Cozystack as a universal installer: a telecom operator and integrator runs GPU passthrough into VMs and clusters, NVIDIA Dynamo and geo-distributed GPU.
- From public cloud to bare metal, bursting on demand: fractional GPU sharing and GPU cost about five times lower than the previous hyperscaler setup.
The largest GPU deployment written up here is a single 8×H100 node. For a larger fleet, we scope a proof of concept on your own hardware.
Pricing
Ænix sells a subscription (support, commercial modules and services), not a licence. Public Cloud Platform uses the published support tiers, priced per 10 physical nodes per month; support for GPU sharing is included from the Standard tier up, and Plus and Enterprise add 24×7 support and a full proof-of-concept package. AI services such as model serving, and multi-region operator programmes, are quoted per RFP. Out-of-scope work is billed at $150 per hour. Full tier details are on the pricing page.
Start with a call
Bring your GPU count and models, your current stack and what you sell today. An Ænix engineer will tell you which GPU modes fit your product and what a launch would take.
Ænix created Cozystack and maintains it together with maintainers from other companies. Cozystack is a CNCF Sandbox project, Apache 2.0; its CNCF Incubation application is in due diligence. Ænix delivers it as three platforms on one engine (Public Cloud, Private Cloud and AI) that combine rather than exclude each other.