Exam materials · Lesson 6
Observability and backups
Who collects metrics and logs, where to look at them, and how a backup differs from fault tolerance.
Two topics in one lesson, because both are about the same thing: what to do when something has gone wrong — and what you managed to prepare in advance.
Metrics and logs
The stack here is not the one people expect out of habit, and this is the exam’s favorite trap.
| What is collected | With what |
|---|---|
| Metrics | VictoriaMetrics |
| Logs | VictoriaLogs |
| Dashboards | Grafana |
| Alerts | Alerta |
There is no Prometheus, Loki or Elasticsearch in the platform. VictoriaMetrics does understand the PromQL query language, though — so familiar queries work, while the storage is different: more compact and cheaper at large volumes.
Metrics are collected by the vmagent agent, which lives next to the workload and sends
what it collects to the storage. The storage itself is deployed as a VMCluster — a
clustered VictoriaMetrics installation managed by an operator. Logs are collected by
fluent-bit — the same role, only for text.
The alerting chain is short, and the exam asks about its order: VMAlert evaluates rules against the metrics in VictoriaMetrics → Alerta gathers the firings in one place, removes duplicates and routes them → email, SMS, messengers. Alerta is there so that a single incident does not arrive as eight emails from eight sources.
Grafana ships with ready-made dashboards, and the exam asks about them by group name: Cluster Overview — the health of the cluster as a whole, Node Metrics — CPU, memory, disk and network per node, ETCD — the state of the etcd store, Storage — the health of LINSTOR and SeaweedFS, Tenant Applications — workloads and managed applications per tenant. There are also dashboards for entry-point traffic and for managed databases.
Remember inheritance from the second lesson: a tenant without its own monitoring sends metrics to its parent. But turning collection on after the fact is useless — there is nowhere to get records for the past from, so this should be decided when the cluster is created, not when a graph is suddenly needed.
Backups
There are two different mechanisms here, and mixing them up is a sure way to get a question wrong.
Backups of managed services (backup). Four objects are at work here, and mixing them up is a sure way to get it wrong:
| Object | What it does |
|---|---|
BackupClass | where and how to store backups. Created by the platform administrator, applies to the whole cluster |
Plan | the schedule: take backups on such-and-such a timetable |
BackupJob | a one-off run, here and now |
Backup | the result — the backup itself |
The key pair is Plan and BackupJob. You need a backup every night — that is a Plan.
You need a backup right now, before an upgrade — that is a BackupJob. Restoring is a
separate object, RestoreJob.
Starting with version 1.5, the platform has a ready-made cozy-default — a backup class that
works right away, without configuring storage.
Velero. One level up: it backs up the cluster’s own objects and volumes. This is about restoring the platform, not an individual database. Since version 1.5 it is not an option but a standard component — installed by default.
It is configured with two objects: BackupStorageLocation — where to store backups, and
VolumeSnapshotLocation — where to keep volume snapshots. For a Kubernetes cluster inside a
tenant, Velero is enabled with a single field: spec.addons.velero.enabled.
Where the boundary lies
The most valuable thing in this topic is understanding what backups do not do.
A backup of a managed application takes only the data. It does not include the application’s HelmRelease, the CR object you created in the managed applications catalog, or the secrets created by the database operator.
Hence the restore rule: the target application must exist before you restore. A restore pours data into an existing database; it does not rebuild the application from scratch.
There are no incremental backups of virtual machines: the platform does not do changed block tracking, and every backup is a full one. If you are used to incremental chains, you will have to recalculate backup windows and storage capacity.
Backups are not fault tolerance. Replicas save you from a node failure, backups from deleted data. These are different troubles, and one does not replace the other.
A backup that has never been restored is not a backup but a hope. Restores have to be tried before they are needed.
And most importantly: a backup is kept in object storage. If that storage lives in the same cluster as the data, it will not save you from losing the whole cluster. For real protection, the storage must be outside.
We have covered what the platform is made of and what it can do. One last question remains — why it is built exactly this way.
What the exam will ask
- That metrics are stored by VictoriaMetrics and logs by VictoriaLogs. Not Prometheus and not Loki.
- That VictoriaMetrics is PromQL-compatible.
- That metrics are collected by vmagent and stored by VMCluster.
- The alerting order: VMAlert → Alerta → email, SMS, messengers.
- The standard dashboard groups: Cluster Overview, Node Metrics, ETCD, Storage, Tenant Applications.
- The four objects:
BackupClass,Plan(schedule),BackupJob(one-off run),Backup(result). - That a scheduled backup is a
Plan, not aBackupJob. - That
cozy-defaultworks out of the box since version 1.5. - That Velero works at the platform level, not at the level of an individual service.
- That Velero became a standard component in version 1.5, and in a tenant cluster it is enabled
with the
spec.addons.velero.enabledfield. - That
BackupStorageLocationsets where backups are stored, andVolumeSnapshotLocationwhere volume snapshots are kept. - What a “data only” backup does not take: the HelmRelease, the CR object, the operator's secrets.
- That the target application must exist before the restore.
- That there are no incremental backups of virtual machines — every backup is a full one.
- That replicas and backups solve different problems.
More: monitoring · cluster services