Free live webinar · Online · Wednesday 30 September 2026 · 16:00 CEST (14:00 UTC)

The cluster that survives a datacenter outage

One hour with Andrei Kvapil, the creator of Cozystack: build real geo-resilience on your own hardware — metro-stretch, two datacenters plus a witness, and remote-DC DR — from iron and network to storage, GPU and databases that fail over on their own. Shown on a real production cluster across three datacenters.

  • Free with registration
  • 60 minutes, live Q&A
  • Recording to every registrant
  • Bring your stack — questions answered live
Timur Tukaev — workshop host

One datacenter is one risk. What happens when it goes dark?

Customers and regulators increasingly demand geo-resilience, and "we have backups" is not an answer when a whole site goes offline. But a distributed cluster is not a checkbox — it is three different architectures for three levels of latency, and the difference between them is the difference between RPO=0 and a split-brain outage. This session is the honest engineering: where the seams are, and how to hold them.

Where this usually starts

Everything runs in a single datacenter, and a power, network or cooling failure takes the whole business with it — while the DR plan is a runbook nobody has rehearsed. Or there is a second site, but resilience means VMware vSAN stretched and SRM licensing, or an OpenStack build that needs a full team just to stay alive.

None of these calls for a rebuild. They call for geo-resilience on hardware you own — without per-socket licensing or a platform you can't operate.

What holds it together

Cozystack is an open-source cloud platform and CNCF Sandbox project that runs one Kubernetes API across sites — virtual machines, managed databases, storage and networking as declarative resources, from three nodes to three datacenters. Nodes join through the API, storage and topology are resources, and DR is described as code.

On top of it sits the difference that matters: defaults and runbooks forged in real disaster-recovery drills with a production provider — synchronous storage, quorum tuning, topology-aware placement — so you start from what already survived a real outage.

One API across sites
  • Metro-stretch · two-DC + witness · remote-DC
  • Synchronous storage · DRBD or Ceph
  • Live migration · VMs, databases, workloads
  • GPU across sites & cloud-burst
+ Production hardening
  • Quorum & witness defaults
  • Topology-aware placement
  • Runbooks from real DR drills
RPO 0
lose a whole datacenter and lose no data, on synchronous metro-stretch
3 datacenters
one Kubernetes API, quorum that survives losing any one site
€0
per-socket licensing — Apache 2.0, CNCF Sandbox project

Distributed, in production — not on slides

Real platforms serving real customers, built on the same foundation the session walks through.

01
3 datacenters
Switzerland · synchronous replication · 60+ tenants in production

A Swiss cloud provider runs one Cozystack cluster stretched across three datacenters, with topology-aware placement and synchronous storage. Together we ran real DR drills — powering off a whole datacenter on purpose — and turned what broke into platform defaults.

02
One API, three clouds
~11,000 active users · bare metal, hyperscaler and sovereign cloud

A European academic-computing platform keeps a single Cluster API spanning owned bare metal, a public hyperscaler and a sovereign Swiss cloud — bursting workloads, including GPU, across sites on demand and pulling them back when the spike passes.

Both are anonymised at the customer's request. Andrei walks through what each of them actually did — and what he would do differently on your stack.

What we'll cover

Live demos, not slides — then your questions.

  1. 01

    Three topologies, and when each applies. Metro-stretch (RPO=0), two datacenters plus a witness, and remote-DC DR — including when stretching is an anti-pattern.

  2. 02

    Quorum math you can trust. How many nodes survive losing one datacenter, why two sites need a witness, and how etcd behaves when the network cuts.

  3. 03

    Storage across sites. Synchronous DRBD vs Ceph, dedicated storage networks, and the tuning that keeps latency from triggering false failovers.

  4. 04

    Moving live workloads. VM live-migration, database switchover, pod rescheduling and queue rebalancing — and the honest limits for GPU workloads (a VM with a passthrough GPU can't live-migrate).

  5. 05

    GPU and cloud-burst. Pooling GPUs within a site, and bursting to other sites or a public cloud — ordering GPU nodes straight into your cluster.

  6. 06

    DR drills with a real provider. How we deliberately powered off a datacenter, what broke, what we tuned, and how it became a product default.

The session ends with a live Q&A. Questions submitted at registration get priority — and this part only happens live.

What you'll leave with

01

A map of the three topologies — and when each one fits, including when stretching is forbidden.

02

The quorum math for surviving one or two lost datacenters, including the two-DC + witness pattern.

03

A clear view of storage and live migration under real network latency — and the knobs that tune them.

04

A readiness checklist to score your own datacenters against before you build.

Who should attend

Enterprise infrastructure teams that need geo-resilience they can prove, and clouds, hosters and data centre operators that want to sell geo-redundant services. If you're comparing VMware vSAN stretched and SRM, an OpenStack build, or multi-AZ on a hyperscaler, the session is built around your situation.

  • Enterprise infrastructure teams
  • Cloud providers
  • Data centre operators
  • Hosting providers
  • MSPs

Especially the people who sign off on resilience and DR:

  • Architects
  • SREs
  • CTOs
  • Infrastructure leaders
Andrei Kvapil

Your speaker

Andrei Kvapil
Creator of Cozystack · Founder of Aenix

Andrei created Cozystack, the open-source cloud platform and CNCF Sandbox project, after more than fifteen years of building clouds and high-load infrastructure. He contributes to Kubernetes, KubeVirt, Cilium and LINSTOR, and speaks at KubeCon and other industry events. At Aenix, he helps providers across Europe build geo-resilient infrastructure on hardware they own.

Registration

Wednesday 30 September 2026 · 16:00 CEST (14:00 UTC) · online. Attendance is free — with registration: you get the calendar invite and the recording.

About the webinar

This is a free live webinar for enterprise infrastructure teams and for clouds, hosters and data centre operators that need geo-resilience they can prove. Andrei Kvapil — the creator of Cozystack, an open-source cloud platform and CNCF Sandbox project — builds a distributed Kubernetes cluster live: the three topologies for three levels of latency (metro-stretch with RPO=0, two datacenters plus a witness, and remote-DC DR), the quorum math for surviving a lost site, synchronous storage across datacenters, live migration of VMs and databases, GPU across sites, and the real DR drills where a whole datacenter was powered off on purpose. Shown on a real production cluster stretched across three datacenters. Attendance is free with registration, and every registrant receives the recording.

Quick facts

  • Format A live online webinar, about 60 minutes: a practical walkthrough with live demos, followed by a live Q&A with the speaker
  • Date Wednesday 30 September 2026, 16:00 CEST (14:00 UTC) — online. Register to get the calendar invite and the recording.
  • Price Free with registration; every registrant receives the recording
  • Language English
  • Who it's for Architects, SREs, CTOs and infrastructure leaders who sign off on DR — and providers selling geo-redundant services
  • Host Andrei Kvapil — creator and maintainer of Cozystack (CNCF Sandbox project), founder of Aenix
  • After the webinar The recording, plus a map of the three topologies and a readiness checklist you can score your own datacenters against

Frequently asked questions

Is this a real production system or a lab demo?

Real production. The core case is a Cozystack cluster stretched across three datacenters, plus the actual DR drills we ran with the provider — including what broke and what we fixed.

Do I need three datacenters to get value?

No. We cover single-site teams planning their first second site, two-DC plus witness setups, and full three-site stretch — so you can place yourself on the map wherever you start.

How is this different from VMware vSAN stretched + SRM?

The same metro-stretch resilience, without per-socket licensing or vendor lock-in, on an open-source core you can run on your own hardware. We compare the approaches honestly.

Can you really lose a datacenter with zero data loss?

In a synchronous metro-stretch topology — datacenters within metro distance, roughly a couple of milliseconds round-trip — yes, RPO=0. We show it live and are explicit about where synchronous replication stops working over longer distances, and what you use instead.

What about GPU and databases across sites?

We cover both: GPU sharing within a site and cloud-burst to other sites or a public cloud, and managed databases that place replicas per zone and switch over automatically when a site is lost.

Which storage — DRBD or Ceph?

Both. We compare synchronous DRBD and Ceph across datacenters — the replication model, the quorum, dedicated storage networks and the tuning that keeps latency from triggering false failovers.

Will there be a recording?

Yes, to everyone who registers. The Q&A is the exception — that part only happens live.

Free live webinar · Online · Wednesday 30 September 2026 · 16:00 CEST (14:00 UTC)

Bring your datacenters to the Q&A

Wednesday 30 September 2026 · 16:00 CEST (14:00 UTC) · online. Attendance is free — with registration; every registrant gets the calendar invite and the recording.