A Swiss cloud provider runs one Cozystack cluster stretched across three datacenters, with topology-aware placement and synchronous storage. Together we ran real DR drills — powering off a whole datacenter on purpose — and turned what broke into platform defaults.
Free live webinar · Online · Wednesday 30 September 2026 · 16:00 CEST (14:00 UTC)
The cluster that survives a datacenter outage
One hour with Andrei Kvapil, the creator of Cozystack: build real geo-resilience on your own hardware — metro-stretch, two datacenters plus a witness, and remote-DC DR — from iron and network to storage, GPU and databases that fail over on their own. Shown on a real production cluster across three datacenters.
- Free with registration
- 60 minutes, live Q&A
- Recording to every registrant
- Bring your stack — questions answered live

One datacenter is one risk. What happens when it goes dark?
Customers and regulators increasingly demand geo-resilience, and "we have backups" is not an answer when a whole site goes offline. But a distributed cluster is not a checkbox — it is three different architectures for three levels of latency, and the difference between them is the difference between RPO=0 and a split-brain outage. This session is the honest engineering: where the seams are, and how to hold them.
Where this usually starts
Everything runs in a single datacenter, and a power, network or cooling failure takes the whole business with it — while the DR plan is a runbook nobody has rehearsed. Or there is a second site, but resilience means VMware vSAN stretched and SRM licensing, or an OpenStack build that needs a full team just to stay alive.
None of these calls for a rebuild. They call for geo-resilience on hardware you own — without per-socket licensing or a platform you can't operate.
What holds it together
Cozystack is an open-source cloud platform and CNCF Sandbox project that runs one Kubernetes API across sites — virtual machines, managed databases, storage and networking as declarative resources, from three nodes to three datacenters. Nodes join through the API, storage and topology are resources, and DR is described as code.
On top of it sits the difference that matters: defaults and runbooks forged in real disaster-recovery drills with a production provider — synchronous storage, quorum tuning, topology-aware placement — so you start from what already survived a real outage.
- Metro-stretch · two-DC + witness · remote-DC
- Synchronous storage · DRBD or Ceph
- Live migration · VMs, databases, workloads
- GPU across sites & cloud-burst
- Quorum & witness defaults
- Topology-aware placement
- Runbooks from real DR drills
Distributed, in production — not on slides
Real platforms serving real customers, built on the same foundation the session walks through.
A European academic-computing platform keeps a single Cluster API spanning owned bare metal, a public hyperscaler and a sovereign Swiss cloud — bursting workloads, including GPU, across sites on demand and pulling them back when the spike passes.
Both are anonymised at the customer's request. Andrei walks through what each of them actually did — and what he would do differently on your stack.
What we'll cover
Live demos, not slides — then your questions.
- 01
Three topologies, and when each applies. Metro-stretch (RPO=0), two datacenters plus a witness, and remote-DC DR — including when stretching is an anti-pattern.
- 02
Quorum math you can trust. How many nodes survive losing one datacenter, why two sites need a witness, and how etcd behaves when the network cuts.
- 03
Storage across sites. Synchronous DRBD vs Ceph, dedicated storage networks, and the tuning that keeps latency from triggering false failovers.
- 04
Moving live workloads. VM live-migration, database switchover, pod rescheduling and queue rebalancing — and the honest limits for GPU workloads (a VM with a passthrough GPU can't live-migrate).
- 05
GPU and cloud-burst. Pooling GPUs within a site, and bursting to other sites or a public cloud — ordering GPU nodes straight into your cluster.
- 06
DR drills with a real provider. How we deliberately powered off a datacenter, what broke, what we tuned, and how it became a product default.
The session ends with a live Q&A. Questions submitted at registration get priority — and this part only happens live.
What you'll leave with
A map of the three topologies — and when each one fits, including when stretching is forbidden.
The quorum math for surviving one or two lost datacenters, including the two-DC + witness pattern.
A clear view of storage and live migration under real network latency — and the knobs that tune them.
A readiness checklist to score your own datacenters against before you build.
Who should attend
Enterprise infrastructure teams that need geo-resilience they can prove, and clouds, hosters and data centre operators that want to sell geo-redundant services. If you're comparing VMware vSAN stretched and SRM, an OpenStack build, or multi-AZ on a hyperscaler, the session is built around your situation.
- Enterprise infrastructure teams
- Cloud providers
- Data centre operators
- Hosting providers
- MSPs
Especially the people who sign off on resilience and DR:
- Architects
- SREs
- CTOs
- Infrastructure leaders

Your speaker
Andrei created Cozystack, the open-source cloud platform and CNCF Sandbox project, after more than fifteen years of building clouds and high-load infrastructure. He contributes to Kubernetes, KubeVirt, Cilium and LINSTOR, and speaks at KubeCon and other industry events. At Aenix, he helps providers across Europe build geo-resilient infrastructure on hardware they own.
Registration
Wednesday 30 September 2026 · 16:00 CEST (14:00 UTC) · online. Attendance is free — with registration: you get the calendar invite and the recording.
About the webinar
Quick facts
- Format A live online webinar, about 60 minutes: a practical walkthrough with live demos, followed by a live Q&A with the speaker
- Date Wednesday 30 September 2026, 16:00 CEST (14:00 UTC) — online. Register to get the calendar invite and the recording.
- Price Free with registration; every registrant receives the recording
- Language English
- Who it's for Architects, SREs, CTOs and infrastructure leaders who sign off on DR — and providers selling geo-redundant services
- Host Andrei Kvapil — creator and maintainer of Cozystack (CNCF Sandbox project), founder of Aenix
- After the webinar The recording, plus a map of the three topologies and a readiness checklist you can score your own datacenters against
Frequently asked questions
Is this a real production system or a lab demo?
Do I need three datacenters to get value?
How is this different from VMware vSAN stretched + SRM?
Can you really lose a datacenter with zero data loss?
What about GPU and databases across sites?
Which storage — DRBD or Ceph?
Will there be a recording?
Free live webinar · Online · Wednesday 30 September 2026 · 16:00 CEST (14:00 UTC)
Bring your datacenters to the Q&A
Wednesday 30 September 2026 · 16:00 CEST (14:00 UTC) · online. Attendance is free — with registration; every registrant gets the calendar invite and the recording.