One engine, three platforms. AI Platform runs on the same substrate as Public Cloud Platform and Private Cloud Platform, and combines with either: a provider sells GPU-as-a-Service through the billing surface it already has, a regulated enterprise runs its own inference under the key custody its auditor already accepted.
A pre-integrated stack for inference, fine-tuning and RAG on your own GPUs: multi-tenant GPU scheduling, model-serving and fine-tuning APIs, vector databases, object storage, open-weight models, and sovereignty controls that keep weights and training data inside your perimeter. For AI-native organizations and regulated AI deployments at scale.
What’s included
Ready-to-use blueprints
Built patterns for common AI workload types:
- Single-tenant inference cluster — for one customer, one workload class
- Multi-tenant inference fleet — shared GPU pool with logical tenant isolation
- Inference + fine-tuning + RAG — full-stack pattern with heterogeneous GPU pools
- Air-gapped sovereign deployment — for defence, isolated industrial, sovereign-cloud customers
(See Sovereign AI Decision Guide for blueprint detail.)
Multi-tenant GPU scheduling
Per-tenant GPU pools and GPU-class-aware scheduling (L40S for inference, H100 for fine-tuning) on the NVIDIA GPU Operator, with HAMi for fractional sharing of a card across tenants. Quotas, RBAC and observability per tenant. MIG-partitioned multi-tenancy is on the roadmap, not shipping today; size accordingly.
Models, databases, apps included
Pre-deployed open-weight models (Llama, Mistral, Qwen, DeepSeek, Phi, Gemma families). Vector DB (pgvector via PostgreSQL operator, or Qdrant). Managed databases (PostgreSQL, MariaDB, Valkey, Kafka, ClickHouse, RabbitMQ). Object storage (S3-compatible) for training data + model checkpoints.
Service APIs
Inference (vLLM-compatible by default; Triton supported), fine-tuning jobs, embedding generation, RAG retrieval, vector indexes and evaluation harnesses — platform primitives, multi-tenant-aware, rather than a bespoke MLOps build per workload.
Sovereignty controls
Customer-controlled encryption keys for model weights at rest, training data, vector indexes. Supplier transparency to second hop. Audit-isolated environment. Provider personnel access logged + time-limited. Air-gap deployment supported.
GPU sizing reference
Practical sizing tables for common workload profiles (Llama 7B / 13B / 70B / 405B, Mistral, Qwen, DeepSeek, Phi, Gemma — single-card / multi-card / multi-node configurations). Ænix engagement includes capacity planning for sustained workloads.
Hosting panel + admin interface
Branded admin dashboard for the AI platform operator. Service-creation wizards for end users (ML engineers, data scientists, app teams).
Observability for AI workloads
Inference latency / throughput metrics. GPU utilisation per tenant. Model-serving SLOs. Cost-per-token tracking. Anomaly detection for inference quality drift.
Migration tooling and expertise
Productized patterns for migration from hyperscaler AI (AWS Bedrock, Azure OpenAI Service, GCP Vertex AI) to sovereign AI infrastructure. Particularly for organisations with sustained inference workloads where economics no longer fit hyperscaler API pricing.
Who buys AI Platform
| Buyer | Typical engagement |
|---|---|
| AI-native startup at scale | Sovereign inference fleet, replacing hyperscaler API spend |
| Regulated AI deployment (bank / public sector / healthcare) | Sovereignty-required AI infrastructure with customer-controlled keys |
| GPU-heavy product company | Multi-tenant GPU platform with strict cost discipline |
| Telco / large enterprise running AI | Internal AI platform shared across BUs |
Why AI Platform over alternatives
| Vs. | Why AI Platform |
|---|---|
| Hyperscaler AI APIs (Bedrock, Azure OpenAI, Vertex) | Sovereign — customer controls weights, data, operations. Sustained-utilization economics typically beat hyperscaler API pricing. Fine-tuning ownership. Auditability. |
| Building it yourself on Kubernetes and GPU drivers | Multi-tenant GPU scheduling, observability, sovereignty controls, blueprints and service APIs arrive together instead of as a 12-24 month platform-engineering project. |
| Closed-source MLOps platforms | Open-source foundation (Cozystack, Apache 2.0) — no per-engineer or per-model licensing, and the substrate stays yours if the contract ends. |
| Run:ai (NVIDIA) | Run:ai is a GPU scheduler and quota layer that assumes a Kubernetes platform already exists underneath — cluster lifecycle, storage, networking, tenancy and the VM estate are still yours to build and run. AI Platform brings the platform itself: fractional GPU sharing, KubeVirt for the workloads that never containerized, LINSTOR/DRBD storage, Tenant CRD isolation. It is also Apache 2.0 with no per-GPU subscription and no NVIDIA-only hardware assumption. If you already run a mature Kubernetes platform and only need scheduling, Run:ai is a narrower and reasonable purchase. |
| Kubeflow | Kubeflow is an ML toolchain — pipelines, notebooks, training operators, serving — not an infrastructure platform, and running it is itself a platform-engineering project. AI Platform supplies what Kubeflow assumes: multi-tenant GPU scheduling, managed databases and vector stores, object storage, observability, isolation per team. The two are complementary: teams run Kubeflow, or Dynamo, or plain vLLM, as tenant workloads on top. |
Pricing
Project plus managed retainer, quoted per RFP. Discovery call to scope.
Engagement structure
- Discovery call (30 min, free)
- Sovereign AI architecture review (1-2 weeks, fixed-price) — using the Sovereign AI Decision Guide framework + Ænix expertise
- Pilot engagement (3-6 months) — defined slice (one workload class, one tenant, one model family)
- Full AI Platform build (6-12 months) — production AI infrastructure with all targeted workload types
- Managed retainer (optional, ongoing) — Ænix runs the AI platform under SLA
Customer evidence
AI Platform customers are NDA-protected. AI-native organizations and regulated AI deployments are in production; reference calls can be arranged under NDA for an active opportunity.
Combine it with the other platforms
The three Ænix platforms are one engine with different surfaces switched on. AI Platform is not a separate installation — it is GPU tenancy, model serving and the data services around them, running on the same substrate as everything else you operate.
- Private Cloud Platform — the usual pairing for regulated buyers. DORA / NIS2 architecture, customer-managed keys and audit-ready logging extend over the AI estate: model weights at rest fall under the same key custody as the primary datastore, and GPU workloads sit inside the Tenant CRD boundary the auditor already reviewed. Taking AI Platform with Private Cloud features is a configuration decision, not a second contract.
- Public Cloud Platform — for providers selling GPU-as-a-Service. Billing, metering and the customer portal come from that side; GPU-class-aware scheduling and fractional sharing come from this one. Providers commonly start with VMs and databases and switch GPU on when demand appears, on hardware they already run.
How to start
Book a discovery call. Bring your AI workload profile (steady inference / training / fine-tuning / RAG / mix), regulatory scope, and target deployment model. We’ll discuss AI Platform fit and engagement scope.
Ænix AI Platform is built on Cozystack — a CNCF project we created and maintain (currently CNCF Sandbox; CNCF Incubating expected late summer 2026). Apache 2.0. Ænix is the open-core company.
