Business continuity is not a line in a vendor contract — it is an outcome you have to be able to prove. Disaster recovery as a service (DRaaS) on a sovereign, self-operated platform gives you cross-data-centre synchronous replication, immutable backups, and failover that is tested rather than assumed. Ænix builds and operates these platforms on Cozystack, so your recovery-time and recovery-point objectives are architecture you own and evidence you can hand to a regulator.
Pairs with: Ænix Private Cloud Platform for the regulated cloud foundation that DR sits on; DORA compliance for the operational-resilience obligations DR helps you meet. Start with a Platform Readiness Assessment →.
What does DRaaS actually have to guarantee?
Every disaster-recovery conversation reduces to two numbers, and most vendor pitches quietly avoid them.
- Recovery-time objective (RTO) — how long you are allowed to be down. This is a function of how fast you can bring the second site into service, not of how big your backup is.
- Recovery-point objective (RPO) — how much data you can afford to lose, expressed as time. Nightly backups imply an RPO of up to 24 hours; synchronous replication targets an RPO close to zero for the protected tier.
A credible DR capability commits to both numbers per workload tier and then demonstrates them in a drill. On a sovereign platform, the replication topology, the backup immutability, and the drill records are all things you hold and can inspect — you are not trusting a hyperscaler’s opaque SLA to describe a failure mode you will never see documented.
How synchronous cross-data-centre replication works
The protected tier of a sovereign DR platform is built on synchronous block replication, so a committed write exists in more than one data centre before the application is told it succeeded.
On the reference architecture, Cozystack runs a compute cluster geo-distributed across three data centres. Volumes are replicated synchronously with LINSTOR/DRBD at replication factor three — one replica per site — and etcd, the Kubernetes cluster state store, is geo-distributed across the same three sites. The result is an architecture designed to survive the loss of an entire data centre with no data loss, because both the persistent data and the control-plane state already live elsewhere.
This is standard, open, CNCF-aligned Kubernetes infrastructure rather than proprietary DR appliances. The Kubernetes storage model treats the replicated volumes as ordinary persistent volumes, so applications do not need bespoke DR integration to benefit from cross-site durability.
Why immutable backups matter more than ever
Synchronous replication protects against hardware and site failure, but it faithfully replicates a ransomware encryption event too. That is why DR and backup are separate layers.
Backups on the platform are written to object storage with S3 Object Lock and versioning, producing immutable copies that an attacker who has compromised the primary environment cannot alter or delete within the retention window. Platform- and tenant-level Velero backups capture Kubernetes objects and volume snapshots, and a deletion-protection webhook guards critical objects — volumes, namespaces, load balancers — against accidental or malicious removal. At-rest LUKS encryption and encrypted inter-DC replication keep the recovery copies confidential as well as durable.
The distinction matters for regulators: operational-resilience frameworks increasingly expect a recovery path that is provably isolated from the blast radius of the primary incident.
Tested failover, not paper failover
A DR plan that has never been exercised is a hypothesis. The platforms Ænix operates are drilled for real.
On the anchor engagement, the client regularly powers nodes off to test resilience deliberately, which surfaces the non-obvious cascades a tabletop exercise never finds. Upgrades are rehearsed on staging on the record, then repeated on production; non-declarative commands are dropped in favour of GitOps; and each scenario has a ready runbook — DRBD recovery, cluster upgrade, storage failover. This is what converts an RTO from a marketing figure into a number you can defend.
Evidence: a 20-hour incident, zero data loss
The clearest proof of a DR posture is how it behaves on the worst day. In our anonymized sovereign public cloud case study, a multi-tenant provider hit a cascading storage failure during a major upgrade — a DRBD race, lost patches at an intermediate step, and a breaking change in the network layer. The team worked the incident for roughly 20 hours and recovered the cloud with zero data loss, then pushed the underlying bugs upstream into LINSTOR and its CSI driver. The same three-DC replication, geo-distributed etcd, and immutable-backup pattern a bank or insurer would deploy carried a real production cloud through a real disaster.
For DORA-scoped entities specifically, this is the shape of evidence DORA (Regulation (EU) 2022/2554) expects: tested resilience, documented recovery, and objectives you can show rather than assert.
Not every workload needs the same recovery tier
Treating every system as mission-critical is how DR budgets explode and drills become unmanageable. A working DR posture tiers the estate first.
- Tier 0 — synchronous. Systems where an RPO above near-zero is unacceptable — core banking ledgers, order books, patient records. These sit on synchronous cross-DC replication and are the reason the three-DC topology exists.
- Tier 1 — asynchronous plus frequent backups. Important but tolerant of minutes of data loss. Frequent immutable backups and asynchronous replication keep the cost proportionate to the risk.
- Tier 2 — backup and rebuild. Stateless or easily reconstructed services recovered from immutable backups and infrastructure-as-code, with an RTO measured in hours rather than seconds.
Tiering is the first output of the assessment, because it decides where the expensive synchronous capacity goes and where a cheaper recovery path is honestly sufficient.
How Ænix engages on disaster recovery
The engagement runs as a Platform Readiness Assessment with DR-weighted workstreams: current RTO/RPO posture per workload tier, replication and geo-topology design, backup immutability and ransomware isolation, and drill-process maturity. Output is a written report plus a Phase 2 implementation roadmap. Where the DR platform doubles as the production platform — the usual case — it pairs naturally with data sovereignty and DORA-alignment work, so continuity, residency, and compliance are engineered together rather than bolted on.
Ænix is the team behind Cozystack — a CNCF project (Sandbox today; Incubating expected late summer 2026), Apache 2.0. Ænix commercializes it as Ænix Platform, as three platforms on one engine — Public Cloud, Private Cloud and AI — that combine rather than exclude each other. We build sovereign disaster-recovery and business-continuity platforms for regulated organizations across the EU and DACH.