Exam materials · Lesson 7

How it all actually works

Desired state, operators and GitOps — the four ideas everything else rests on.

The last lesson is about ideas, not components. They underlie everything covered so far, and the questions about them are phrased simply: “why does it work this way”.

You describe the outcome, not the actions

The familiar approach to administration is a sequence of steps: install, configure, start, check. Here it is different: you describe how things should be — the desired state — and the system decides on its own what to do to get there.

The difference shows when something breaks. A script that ran yesterday will not help today: it has already been executed. A description of the desired state, on the other hand, is in effect all the time — if reality drifts away from it, reality is brought back.

1. Read how it should be2. Look at how it is3. Close the gap4. Repeat forever
This loop runs continuously — both in Kubernetes controllers and in the platform's operators.

Hence self-healing: a deleted replica of an application comes back not because someone noticed it was missing, but because reality diverged from the description.

What a manifest is made of

The description of the desired state lives in a YAML file called a manifest. Whatever object you describe — a tenant, a database, a virtual machine — there are always four top-level fields, and mixing them up in the exam is a bad idea.

FieldWhat it holds
apiVersionwhich API version the type belongs to, for example apps.cozystack.io/v1alpha1
kindthe type itself: Tenant, Bucket, VMInstance
metadataidentity data: name, namespace, labels, annotations
specthe desired state — everything the object should become

The fifth field, status, you never write: the controller fills it in, and it says how things actually are. Hence a simple rule for reading any object: spec is your requirement, status is the platform’s answer to it. The reconciliation loop above is exactly the continuous bringing of one in line with the other.

apiVersion: apps.cozystack.io/v1alpha1
kind: Bucket
metadata:
  name: images          # what the object is called
  namespace: tenant-lab # where it lives
spec:                   # what it should become
  replicas: 2

New object types and operators

kubectl get tenants works even though Kubernetes itself has no tenants. It works because Cozystack extends the Kubernetes API, and the API server knows the Tenant type.

The API is extended in two different ways, and the exam distinguishes between them.

Defining your own type — a CRD (custom resource definition). You register a new type, and from then on its objects are stored and served by the regular API server along with the built-in ones. This is how, for example, backups (backups.cozystack.io) and gateways (gateway.cozystack.io) are implemented in the platform.

An aggregated API server. A separate program takes over an entire API group, and the main API server forwards requests to it. In Cozystack this is cozystack-api in the cozy-system namespace, and it is the one responsible for the most visible groups — apps.cozystack.io (tenants, buckets, databases, clusters), core.cozystack.io, sdn.cozystack.io. You will not find a CRD named tenants.apps.cozystack.io in the cluster: this type is not registered, it is served by the aggregated server.

Which component is responsible for which group is visible with a single command — the SERVICE column shows either the program’s address or the word Local, meaning “the regular API server, type from a CRD”:

kubectl get apiservices | grep cozystack

But a type on its own does nothing: it is a record the cluster agrees to store. The work is done by an operator — a program that watches objects of its type and brings the world into line.

Hence the practical conclusion: if an object has been created and nothing happened, the question is not for the object but for the operator. Either it is not running, or it failed.

GitOps

Since state is described as text, the text can be stored in version control. Git then becomes the single source of truth, and a dedicated program makes sure the cluster matches it.

The platform uses FluxCD and applies this approach to itself: its components are described and deployed the same way you would deploy your own application.

One caveat, so as not to overstate things. The desired state of the platform itself is defined not by your repository but by the Platform Package — the YAML configuration of the installation; FluxCD pulls the charts from an OCI registry. An agent inside the cluster continuously brings the cluster to the described state, so a change made bypassing the description does not live long.

FluxCD has two objects you need to know by name. HelmRepository is the source the charts come from. HelmRelease (short name hr) is the statement “this chart, of this version, with these values, must be installed and stay installed”.

How to temporarily disable reconciliation

Continuous reconciliation gets in the way in exactly one case: when you need to fix something by hand. A change made bypassing the description will be rolled back by the controller within a few seconds — that is what it exists for.

For this case HelmRelease has a switch, spec.suspend. For a release switched to suspend, the controller stops reconciling: it does not reinstall the chart, does not roll back changes, and does not touch the object at all. Meanwhile the workload keeps running — pods are not deleted, the application responds — and your manual changes are kept until reconciliation is switched back on.

kubectl patch hr <name> -n <namespace> --type merge -p '{"spec":{"suspend":true}}'

The reverse operation is "suspend": false. At the moment it is turned back on, the controller reconciles the object anew, and everything you fixed by hand bypassing the description is rolled back. That is why suspend is a diagnostic tool for the duration of an investigation, not a way to live with manual changes.

Helm

One object is one object. An application is usually a dozen: a deployment, a service, configuration, secrets, access rules.

Helm packages such a set into a chart — templates plus values. By changing the values, you get different installations from one package.

In the platform, Helm is not a recommendation but a load-bearing structure: every item in the managed applications catalog is a chart, and ordering a service deploys exactly that chart. The platform itself is also installed with a Helm chart — with a single helm upgrade --install command using the cozy-installer chart into the cozy-system namespace.

The result is the following chain, and it is worth memorizing in full:

object → HelmRelease → chart → operator → running pods

Everything the platform does goes through it. You ordered a database — it went through it. You created a tenant — it went through it. The platform installed itself — that too.

Kubernetes vocabulary that will be asked here too

This exam domain is described as “the seventh topic plus basic terminology”, so it is worth going over seven terms out loud. The English names are exactly the ones that will appear in the questions.

ObjectWhat it is and why
Namespacea named space that groups resources; each tenant has its own
Podthe smallest unit of workload — one container or several together
Deploymentdescribes what is desired: which image, how many replicas, how to update
ReplicaSetits executor: keeps the required number of pods; it is rarely touched by hand
Servicea stable name and address in front of a changing set of pods
PVCa claim for persistent storage that survives a pod restart
Secretstores sensitive data: passwords, tokens, keys

Service has three types: ClusterIP — the address is visible only inside the cluster, this is the default type; NodePort — opens a port on every node, good for testing; LoadBalancer — asks the platform for an external address; on your own hardware it is provided by MetalLB from lesson five.

And five commands the exam asks about by their exact spelling:

  • kubectl get pods -A — all pods in all namespaces at once
  • kubectl describe pod <name> — pod details and, most importantly, its events
  • kubectl api-resources — which resource types the cluster knows at all, including those added by the platform
  • kubectl apply -f manifest.yaml --dry-run=server — submit the manifest to the server for validation without saving anything
  • helm upgrade --install — idempotent: installs if the release does not exist, upgrades if it does

What the exam will ask

  • That it is the desired state that is described, not a sequence of actions.
  • That a controller continuously compares the desired state with the actual one and closes the gap.
  • That an object's desired state lives in spec, the actual state in status, while apiVersion, kind and metadata are responsible for the API version, the type and the name.
  • That a CRD adds a type, while the work is done by an operator.
  • That kubectl get tenants works because Cozystack extends the Kubernetes API: the Tenant type is served by the aggregated server cozystack-api, not by a CRD.
  • That suspend on a HelmRelease stops reconciliation but does not delete the workload: pods keep running, manual changes are kept until reconciliation is turned back on.
  • That the platform applies GitOps to itself via FluxCD.
  • That the platform's desired state is defined by the Platform Package, and FluxCD takes the charts from an OCI registry.
  • The two FluxCD objects: HelmRepository — the source of charts, HelmRelease — the statement about an installed chart.
  • That a chart is the unit of packaging, and every item in the managed applications catalog is a chart.
  • That Cozystack itself is installed with the cozy-installer Helm chart.
  • The chain: object → HelmRelease → chart → operator → pods.
  • What Namespace, Pod, Deployment, ReplicaSet, Service, PVC and Secret are — one sentence for each.
  • The three Service types: ClusterIP, NodePort, LoadBalancer.
  • Commands: kubectl get pods -A, kubectl describe pod for events, kubectl api-resources, --dry-run=server, helm upgrade --install.