One internal platform for two things a large organisation usually builds twice: data — analytics, lakes and marts, model training — and AI/ML services, from development through training to serving. Underneath sits AI-ready infrastructure: GPU resource pools with time-slicing and per-tenant quotas, a single scheduler placing both pods and virtual machines, and usage metrics detailed enough to charge teams and to see where capacity actually goes. The platform is in rollout: GPU support is live, the AI-services MVP is most of the way through its phase.
About the project
The client is building one internal platform to cover data management and AI/ML work for the whole organisation, instead of letting each function accumulate its own stack. On the data side: analytics, data lakes and data marts, and model training. On the AI/ML side: running, developing and training models as a service other teams consume.
The two are usually treated as separate programmes and then spend years copying datasets to each other. Here they share infrastructure, tenancy and quotas from the start.
Goals and objectives
- One internal platform for data management and for AI/ML services, rather than two operating models over one estate.
- GPU capacity that can be shared between teams without either starving anyone or leaving cards idle.
- Usage measured precisely enough to charge internal consumers and to make capacity decisions on evidence.
- Multi-tenancy strong enough that teams work in isolation on shared hardware.
- Automation of the GPU lifecycle end to end, from provisioning to decommissioning.
Proposed solution
AI-ready infrastructure: GPUs for both Kubernetes and VMs.
- The GPU infrastructure layer — GPU resource pools, time-slicing, quotas per tenant and per project.
- One scheduler — a single planner for pods and virtual machines, with utilisation metrics that feed both chargeback and deep analytics.
- Data and pipelines — S3-compatible storage, databases and model artefacts, with pipelines automated GitOps-style.
Full GPU lifecycle management. Automated GPU provisioning, passthrough into both VMs and Kubernetes, and driver management, joined up rather than scripted per case:
- Full autopilot for NVIDIA cards; GPU passthrough for other vendors.
- Automatic driver installation and GPU resource management inside Kubernetes.
- Lifecycle and resource management, autoscaling and on-demand provisioning.
- Security and multi-tenancy, decommissioning and rolling upgrades.
What the platform already does
- Monitoring — utilisation, load, and the idle capacity nobody was accounting for.
- Provisioning — fast allocation of GPU or vGPU to an ML project.
- Multi-tenancy — isolated working spaces per team.
- Billing and quotas — limits, tariffs and consumption accounting.
- Dynamic allocation — GPUs redistributed as they are released instead of staying reserved.
- Inventory — every card, its location and its state, in one register.
Roadmap
- Phase 1 — complete. GPU support. Autopilot for NVIDIA, passthrough for other vendors, monitoring and resource accounting. This is the infrastructure layer.
- Phase 2 — two months, about 70% done. First AI-services MVP. AI services surfaced in the platform dashboard, support for other vendors’ GPUs, popular self-hosted models.
- Phase 3 — three months. Expanded services. Better orchestration of AI services, an enterprise-grade GPU toolset, MIG support.
- Phase 4 — expansion. Fine-tuning of the AI platform and automation for non-NVIDIA GPUs.
Why this case matters
Data and AI on one platform
Same storage, same tenancy, same quotas. No second operating model, and no copy of every dataset moving between two stacks.
Pods and VMs, one scheduler
Which is what makes chargeback possible at all — two schedulers each believing they own the cards cannot produce a number anyone will sign.
Idle GPUs are a measurable line
Utilisation, load and unused capacity are reported per tenant, and freed cards go back into the pool instead of staying booked.
A roadmap with a finished phase in it
GPU support is done and running; the AI-services layer is being built on top of it. Published in progress, not in retrospect.
This case study describes an engagement in rollout and is published in anonymized form (Tier-3 evidence): the customer is described by profile, not by name. A customer reference is available under NDA on request — talk to Ænix sales.
Ænix is the team behind Cozystack — a CNCF project (Sandbox today; Incubating expected late summer 2026), Apache 2.0. Ænix commercializes it as Ænix Platform, as three platforms on one engine — Public Cloud, Private Cloud and AI — that combine rather than exclude each other.