Sovereign AI infrastructure — GenAI and inference on data that can't leave the perimeter

Open-source Cozystack (a CNCF project we create and maintain) Ænix Platform, the supported commercial distribution Aenix builds, operates and migrates it.

Sovereign AI infrastructure runs GenAI, inference, fine-tuning, and RAG on hardware the customer owns or controls, in the customer’s chosen jurisdiction, under the customer’s encryption keys — with model weights, prompts, completions, and embeddings never leaving the perimeter. It is built for regulated organizations (financial services, healthcare, public sector) and AI/GPU operators whose data class, regulator, or inference economics make hyperscaler AI services unviable. Aenix designs, builds, and operates these platforms on Cozystack, an Apache 2.0 CNCF project that combines KubeVirt VMs and Kubernetes inference workloads on one API, with GPU scheduling through the NVIDIA GPU Operator: HAMi fractional sharing and PCI passthrough for containers, NVIDIA vGPU for VMs. Aenix has no model-provider bias and recommends the open-weight model — Llama, Mistral, Qwen, DeepSeek, Phi — that fits the data class and economics.

Quick facts

  • What it is AI inference, fine-tuning, and RAG running on customer-controlled hardware, in the customer’s jurisdiction, under the customer’s keys, with data never leaving the perimeter
  • Who it's for Regulated financial services, healthcare, public sector, and AI/GPU operators where data class, regulator, or inference economics rule out hyperscaler AI services
  • License Apache 2.0 (no per-CPU / per-core licensing)
  • Status Cozystack is a CNCF project (Sandbox since 2025-02-28; Incubating expected late summer 2026)
  • Platform Cozystack — KubeVirt for VMs and Kubernetes for inference on one API; NVIDIA GPU Operator with HAMi fractional sharing and PCI passthrough for containers, NVIDIA vGPU for VMs
  • Validated GPUs NVIDIA A100, H100, H200, L40S, and Blackwell; specific model-to-hardware fit established during the assessment
  • Engagement 14- or 28-day fixed-price Platform Readiness Assessment, then Aenix-delivered Phase 2 implementation (typically 3-9 months); air-gapped deployment supported

For regulated workloads, AI is no longer a hyperscaler-only conversation. Sensitive data classes, sectoral rules, and the economics of inference at scale are pushing financial services, healthcare, public sector, and AI-platform operators toward sovereign AI infrastructure — GenAI, inference, and analytics on the customer’s own hardware, in the customer’s chosen jurisdiction, under the customer’s encryption keys.

Ænix builds and operates these platforms end-to-end: an architecture, a deployment, and an operations model your team can actually run.

Pairs with: Ænix AI Platform — AI platform automation out of the box (multi-tenant GPU scheduling, inference/fine-tuning/RAG blueprints, vector DB + object storage, sovereignty controls); add Private Cloud Platform for a broader sovereign cloud, or the free Sovereign AI Decision Guide →.

NVIDIA-validated GPU stack · Apache 2.0 platform · EU engineers · Air-gapped deployment supported

Who needs sovereign AI

Sovereign AI is not for every workload. It is the right answer when at least three of the following hold:

  • Data class is sensitive — regulated personal data, financial records, healthcare records, classified information, internal IP that cannot be exposed to model providers.
  • Regulator binds AI processing to jurisdiction — DORA, NIS2, sectoral rules, sovereign-cloud mandates (EU member states, Kazakhstan, several APAC).
  • Inference at scale is economically painful in hyperscaler — GPU pricing, egress costs, and unpredictable spend make 24/7 inference workloads better suited to dedicated infrastructure.
  • Model behavior must be reproducible and auditable — regulator dialog requires “exactly which model produced this output, with which weights, with which input data.”
  • Air-gap or restricted-egress is required — public-sector classified, defence-adjacent, or critical-infrastructure workloads.

If you have none of these, sovereign AI is over-engineering. If you have three or more, the question is not whether — it’s how, by when, and at what cost.


What sovereign AI actually means

1. The model runs on your hardware Inference (and training, where applicable) on GPUs you own or operate, not on a hyperscaler’s GPU instances or model API. NVIDIA H100 / H200 / L40S / Blackwell for the mainstream path; AMD MI-series and Intel Gaudi where supply continuity or sovereignty argues for them.

2. The data never leaves the perimeter Training data, prompts, completions, embeddings, and any derivative artifacts stay within the customer-controlled environment. No traffic to model-provider endpoints; no observability data to SaaS vendors that process outside the perimeter.

3. The model weights are in your control Open-weight models (Llama, Mistral, Qwen, DeepSeek, Phi, etc.) running locally; or fine-tuned variants whose weights you own. Not a model API with prompt-routing into a third-party model.

4. The platform is operated by you, under your governance Kubernetes-native AI platform with clear ownership of GPU scheduling, autoscaling, model management, and audit trails. Not a black-box appliance with vendor-controlled operations.

This is not “private AI” as a label for a SaaS endpoint with a privacy clause. It’s an architecturally sovereign stack with named components and demonstrable controls.


Where common AI-platform approaches fail the sovereignty test

“Private deployment” of a SaaS model API Model provider runs the inference; data flows to the provider’s endpoint. Privacy clause notwithstanding, the data has left the perimeter. Sovereignty failed.

Hyperscaler-managed GPU with proprietary services GPU is in the right region, but model orchestration, observability, and storage hooks lock the workload into proprietary services. Exit cost grows; concentration risk grows.

Single-tenant SaaS in a “sovereign” hyperscaler region The region is sovereign, but the service plane is operated by the hyperscaler. Encryption keys, control-plane access, and software-update channels remain with a non-sovereign vendor.

Self-hosted LLM with no platform underneath A team runs vLLM or llama.cpp on a couple of bare-metal boxes, calls it private AI. Works for a PoC. Fails on multi-tenancy, GPU autoscaling, audit-readiness, or operational availability for production.

The honest answer is usually a Kubernetes-native AI platform on customer-controlled hardware, with a defined operations model.


How Ænix helps

AI workloads
InferenceFine-tuningRAG
scheduled on
Cozystack / Ænix
KubeVirt VMs + KubernetesGPU Operator: HAMi, passthrough, vGPUCustomer keys and hardware
delivers
Sovereign AI
Data never leaves the perimeterNo model-provider endpoints

The sovereign-AI engagement runs as part of our Platform Readiness Assessment with sovereignty + AI-platform workstreams emphasized. Where the engagement leads to implementation, Ænix delivers the platform end-to-end.

The assessment phase produces:

  • Architecture options — concrete platform designs for inference / training / fine-tuning at your scale, with hardware sizing.
  • Sovereignty controls — data-residency, key-custody, and audit-trail design specific to AI workloads.
  • GPU strategy — NVIDIA / AMD / alternatives sizing, model-to-hardware fit, scaling assumptions.
  • Operations model — who runs the platform, what self-service surface product / data-science teams get, what the on-call model looks like.
  • Phase 2 implementation roadmap — Ænix-delivered build, with timeline, effort estimates, and success criteria.

The implementation phase delivers:

  • Cozystack-based AI platform with KubeVirt for VMs, Kubernetes for inference workloads, NVIDIA vGPU for VM-based GPU workloads, and the NVIDIA GPU Operator with HAMi fractional sharing or PCI passthrough for container-based GPU workloads.
  • Validated model serving — vLLM, Triton, or alternatives matched to model architecture.
  • Self-service for data-science teams — provisioning paths, observability, audit trails.
  • Air-gapped deployment where the regulator requires it.

Why Ænix specifically

  • AI infrastructure is what we run. Cozystack is in production with AI / GPU operators across the EU and Central Asia. We have shipped GPU platforms supporting inference and fine-tuning workloads end-to-end.
  • No model-provider bias. We do not have a commercial relationship with a specific LLM provider. The architecture recommends the open-weight model that fits your data class, regulator, and economics — Llama, Mistral, Qwen, DeepSeek, Phi, or fine-tuned variants — and the serving stack to match.
  • Open-source platform foundation. Cozystack is a CNCF Project running on the customer’s chosen hardware in the chosen jurisdiction. Cluster-level access stays with the customer; we operate under your governance, not in spite of it.

Ready to scope your build? Book a call →

What the engagement looks like

Day 0 is a free 30-minute discovery call that fixes the scope. Days 1-13 (or 1-27) run four parallel workstreams with sovereignty and AI-platform emphasized. Day 14 (or 28) is a 60-90 minute executive readout against the written report — architecture options, sovereignty controls, GPU strategy, operations model and Phase 2 roadmap. Phase 2 is the Ænix-delivered build, typically 3-9 months to a production platform and handover. Full day-by-day methodology: Platform Readiness Assessment.


Sovereign AI platforms we’ve built

We have built and operated AI platforms for AI / GPU operators, financial-services organizations, and public-sector initiatives across the EU and Central Asia. Workload patterns include inference-at-scale (24/7), fine-tuning, RAG pipelines, and multi-tenant model serving.


Pricing and engagement scope

The sovereign-AI engagement runs in two phases.

Assessment (14- or 28-day)

Architecture options, GPU strategy, sovereignty controls, operations model, Phase 2 roadmap. Fixed-price. On request

Phase 2 implementation

Ænix-delivered build of the sovereign AI platform. Fixed-scope or time-and-materials, depending on workload count and complexity. Typical 3-9 months elapsed. On request

If Phase 2 follows assessment, the assessment cost is credited against the implementation engagement subject to scope.

We accept RFI / RFP through standard procurement channels in EU member states and Kazakhstan.



Start with a 30-minute discovery call

We confirm fit, narrow the scope to your data class and regulator, and name the 14-day or 28-day variant.

Or read more:


Ænix is the company behind Cozystack — a CNCF Project, Kubernetes Certified Distribution, OpenSSF Best Practices. We build sovereign AI platforms for AI / GPU operators, financial services, and public-sector organizations across the EU, DACH, and Central Asia.

Frequently asked questions

Is sovereign AI the same as private AI?

No. Private AI is used both for SaaS endpoints with a privacy clause and for true on-prem deployments. Sovereign AI specifically requires the model running on customer hardware, data staying inside the customer perimeter, and the platform operated under customer governance.

Which open-weight LLMs does Aenix support?

The current production-ready landscape includes Llama, Mistral, Qwen, DeepSeek, Phi, and Gemma, plus specialized code, vision, and embedding models. Specific selection happens during the assessment based on data class, language requirements, and inference economics.

Does sovereign AI cover training, or only inference?

Both. Inference is the more common entry point; most regulated organizations start there and add fine-tuning of open-weight models later. Full pre-training of frontier models is rare in this segment.

Which GPUs are validated for the platform?

NVIDIA A100, H100, H200, L40S, and Blackwell. Container GPU workloads use the NVIDIA GPU Operator with HAMi fractional sharing or PCI passthrough; VM-based GPU workloads use NVIDIA vGPU. Specific model-to-hardware fit is established during the assessment.

Can the platform run air-gapped?

Yes. Air-gapped, restricted-egress deployment is supported for public-sector classified, defence-adjacent, and critical-infrastructure workloads where the regulator requires it.

Does Aenix have a model-provider bias?

No. Aenix has no commercial relationship with any LLM provider. The architecture recommends the open-weight model and serving stack — vLLM, Triton, or alternatives — that fit the customer’s data class, regulator, and inference economics.

Ready to talk?

Book a 30-minute discovery call — no commitment. We confirm fit, the right platform, and the next steps.