Ænix AI Platform — sovereign AI and GPU infrastructure

Open-source Cozystack (a CNCF project we create and maintain) Ænix Platform, the supported commercial distribution Aenix builds, operates and migrates it.

Aenix AI Platform is turnkey, self-hosted AI infrastructure for AI-heavy and regulated organizations that need to run inference, fine-tuning, and RAG workloads on their own GPUs instead of hyperscaler AI APIs. Built on Cozystack (Apache 2.0, CNCF project), it bundles multi-tenant GPU scheduling with GPU-class awareness, pre-integrated model serving (vLLM-compatible), vector databases, object storage, ready-to-use open-weight models, service APIs, and sovereignty controls such as customer-controlled encryption keys and air-gapped deployment. Aenix, the open-core company behind Cozystack, productizes and delivers it as a project plus optional managed retainer, letting AI teams reach production faster while keeping model weights, training data, and operations fully under customer control.
Aenix AI Platform console

Quick facts

  • What it is Turnkey, self-hosted multi-tenant AI infrastructure for inference, fine-tuning, and RAG on customer-controlled GPUs
  • License Apache 2.0 (no per-CPU / per-core licensing)
  • Status Cozystack is a CNCF project (Sandbox since 2025-02-28; Incubating expected late summer 2026)
  • For AI-native organizations at scale, regulated AI deployments, GPU-heavy product companies, telcos and enterprises running internal AI platforms
  • GPU support NVIDIA H100, H200, A100, L40S, B100/B200 (Blackwell); CPU-only and alternative accelerators (AMD MI series, Intel Gaudi) supported
  • Foundation Cozystack with KubeVirt (VMs + containers on one Kubernetes API), Cilium (eBPF) networking, LINSTOR/DRBD storage, Tenant CRD multi-tenancy
  • Engagement 3-6 months for a typical inference fleet; 6-12 months for full inference + fine-tuning + RAG; optional managed retainer

One engine, three platforms. AI Platform runs on the same substrate as Public Cloud Platform and Private Cloud Platform, and combines with either: a provider sells GPU-as-a-Service through the billing surface it already has, a regulated enterprise runs its own inference under the key custody its auditor already accepted.

A pre-integrated stack for inference, fine-tuning and RAG on your own GPUs: multi-tenant GPU scheduling, model-serving and fine-tuning APIs, vector databases, object storage, open-weight models, and sovereignty controls that keep weights and training data inside your perimeter. For AI-native organizations and regulated AI deployments at scale.


What’s included

Ready-to-use blueprints

Built patterns for common AI workload types:

  • Single-tenant inference cluster — for one customer, one workload class
  • Multi-tenant inference fleet — shared GPU pool with logical tenant isolation
  • Inference + fine-tuning + RAG — full-stack pattern with heterogeneous GPU pools
  • Air-gapped sovereign deployment — for defence, isolated industrial, sovereign-cloud customers

(See Sovereign AI Decision Guide for blueprint detail.)

Multi-tenant GPU scheduling

Per-tenant GPU pools and GPU-class-aware scheduling (L40S for inference, H100 for fine-tuning) on the NVIDIA GPU Operator, with HAMi for fractional sharing of a card across tenants. Quotas, RBAC and observability per tenant. MIG-partitioned multi-tenancy is on the roadmap, not shipping today; size accordingly.

Models, databases, apps included

Pre-deployed open-weight models (Llama, Mistral, Qwen, DeepSeek, Phi, Gemma families). Vector DB (pgvector via PostgreSQL operator, or Qdrant). Managed databases (PostgreSQL, MariaDB, Valkey, Kafka, ClickHouse, RabbitMQ). Object storage (S3-compatible) for training data + model checkpoints.

Service APIs

Inference (vLLM-compatible by default; Triton supported), fine-tuning jobs, embedding generation, RAG retrieval, vector indexes and evaluation harnesses — platform primitives, multi-tenant-aware, rather than a bespoke MLOps build per workload.

Sovereignty controls

Customer-controlled encryption keys for model weights at rest, training data, vector indexes. Supplier transparency to second hop. Audit-isolated environment. Provider personnel access logged + time-limited. Air-gap deployment supported.

GPU sizing reference

Practical sizing tables for common workload profiles (Llama 7B / 13B / 70B / 405B, Mistral, Qwen, DeepSeek, Phi, Gemma — single-card / multi-card / multi-node configurations). Ænix engagement includes capacity planning for sustained workloads.

Hosting panel + admin interface

Branded admin dashboard for the AI platform operator. Service-creation wizards for end users (ML engineers, data scientists, app teams).

Observability for AI workloads

Inference latency / throughput metrics. GPU utilisation per tenant. Model-serving SLOs. Cost-per-token tracking. Anomaly detection for inference quality drift.

Migration tooling and expertise

Productized patterns for migration from hyperscaler AI (AWS Bedrock, Azure OpenAI Service, GCP Vertex AI) to sovereign AI infrastructure. Particularly for organisations with sustained inference workloads where economics no longer fit hyperscaler API pricing.


Who buys AI Platform

BuyerTypical engagement
AI-native startup at scaleSovereign inference fleet, replacing hyperscaler API spend
Regulated AI deployment (bank / public sector / healthcare)Sovereignty-required AI infrastructure with customer-controlled keys
GPU-heavy product companyMulti-tenant GPU platform with strict cost discipline
Telco / large enterprise running AIInternal AI platform shared across BUs

Why AI Platform over alternatives

Vs.Why AI Platform
Hyperscaler AI APIs (Bedrock, Azure OpenAI, Vertex)Sovereign — customer controls weights, data, operations. Sustained-utilization economics typically beat hyperscaler API pricing. Fine-tuning ownership. Auditability.
Building it yourself on Kubernetes and GPU driversMulti-tenant GPU scheduling, observability, sovereignty controls, blueprints and service APIs arrive together instead of as a 12-24 month platform-engineering project.
Closed-source MLOps platformsOpen-source foundation (Cozystack, Apache 2.0) — no per-engineer or per-model licensing, and the substrate stays yours if the contract ends.
Run:ai (NVIDIA)Run:ai is a GPU scheduler and quota layer that assumes a Kubernetes platform already exists underneath — cluster lifecycle, storage, networking, tenancy and the VM estate are still yours to build and run. AI Platform brings the platform itself: fractional GPU sharing, KubeVirt for the workloads that never containerized, LINSTOR/DRBD storage, Tenant CRD isolation. It is also Apache 2.0 with no per-GPU subscription and no NVIDIA-only hardware assumption. If you already run a mature Kubernetes platform and only need scheduling, Run:ai is a narrower and reasonable purchase.
KubeflowKubeflow is an ML toolchain — pipelines, notebooks, training operators, serving — not an infrastructure platform, and running it is itself a platform-engineering project. AI Platform supplies what Kubeflow assumes: multi-tenant GPU scheduling, managed databases and vector stores, object storage, observability, isolation per team. The two are complementary: teams run Kubeflow, or Dynamo, or plain vLLM, as tenant workloads on top.

Pricing

Project plus managed retainer, quoted per RFP. Discovery call to scope.

Discuss AI Platform →


Engagement structure

  • Discovery call (30 min, free)
  • Sovereign AI architecture review (1-2 weeks, fixed-price) — using the Sovereign AI Decision Guide framework + Ænix expertise
  • Pilot engagement (3-6 months) — defined slice (one workload class, one tenant, one model family)
  • Full AI Platform build (6-12 months) — production AI infrastructure with all targeted workload types
  • Managed retainer (optional, ongoing) — Ænix runs the AI platform under SLA

Ready to scope your build? Book a call →

Customer evidence

AI Platform customers are NDA-protected. AI-native organizations and regulated AI deployments are in production; reference calls can be arranged under NDA for an active opportunity.


Combine it with the other platforms

The three Ænix platforms are one engine with different surfaces switched on. AI Platform is not a separate installation — it is GPU tenancy, model serving and the data services around them, running on the same substrate as everything else you operate.

  • Private Cloud Platform — the usual pairing for regulated buyers. DORA / NIS2 architecture, customer-managed keys and audit-ready logging extend over the AI estate: model weights at rest fall under the same key custody as the primary datastore, and GPU workloads sit inside the Tenant CRD boundary the auditor already reviewed. Taking AI Platform with Private Cloud features is a configuration decision, not a second contract.
  • Public Cloud Platform — for providers selling GPU-as-a-Service. Billing, metering and the customer portal come from that side; GPU-class-aware scheduling and fractional sharing come from this one. Providers commonly start with VMs and databases and switch GPU on when demand appears, on hardware they already run.

How to start

Book a discovery call. Bring your AI workload profile (steady inference / training / fine-tuning / RAG / mix), regulatory scope, and target deployment model. We’ll discuss AI Platform fit and engagement scope.


Ænix AI Platform is built on Cozystack — a CNCF project we created and maintain (currently CNCF Sandbox; CNCF Incubating expected late summer 2026). Apache 2.0. Ænix is the open-core company.

Frequently asked questions

How is AI Platform different from running open-source Cozystack with our own AI stack?

Cozystack provides the multi-tenant Kubernetes and GPU foundation. AI Platform adds pre-integrated inference (vLLM), fine-tuning and RAG patterns, GPU-class-aware multi-tenant scheduling, vector DB and object storage, ready-to-use models and blueprints, AI service APIs, bundled sovereignty controls, GPU sizing expertise, and Aenix delivery experience, saving teams the MLOps build effort.

Which open-weight models are supported?

Open-weight families including Llama 3.x, Mistral / Mixtral, Qwen, DeepSeek (incl. V3), Phi, and Gemma, with new models added as the landscape evolves. Proprietary closed-weight models can be integrated via an API gateway pattern but are not run on customer infrastructure.

Which GPU classes do you support?

NVIDIA H100 and H200 for flagship inference and fine-tuning, A100 for general-purpose work, L40S for cost-effective inference, and B100/B200 (Blackwell) for large training and inference. CPU-only is viable for small models and RAG, and AMD MI series and Intel Gaudi are supported for sovereignty and supply-continuity scenarios.

Can we run this air-gapped?

Yes. Air-gapped deployment is one of the four standard reference architectures: open-weight models, a self-contained registry, customer-controlled HSM-backed keys, and customer-side audit logging. Operational overhead is higher, but sovereignty is maximal.

Is sovereign inference cheaper than hyperscaler AI APIs?

For sustained inference (steady production load or millions of tokens per day), running on owned or leased GPU infrastructure typically delivers a significantly lower cost per token than per-token API pricing. The breakeven depends on your workload pattern, which a discovery call scopes.

Can we fine-tune on customer data and keep ownership?

Yes. Fine-tuning is a first-class workload supporting LoRA, QLoRA, and full or partial multi-GPU runs. Training data and the resulting models stay customer-controlled, with an audit-isolated environment available for regulated training data.

How does this fit with our existing observability stack?

AI Platform ships VictoriaMetrics and VictoriaLogs for the AI-specific signals — inference latency and throughput, GPU utilisation per tenant, cost per token, model-serving SLOs — and exports them over standard exporters, so Datadog, Splunk or an existing Prometheus estate consumes them without a parallel stack.

Ready to talk?

Book a 30-minute discovery call — no commitment. We confirm fit, the right platform, and the next steps.