← blog · July 28, 2026

Backing up a Kubernetes cluster with Velero and restoring it into another one

The real problem with Kubernetes backups is not etcd but cluster objects and persistent volume data being captured separately. How Velero is built, when to choose snapshots over file system backup, what actually breaks on a cross-cluster restore, and how to automate the drill and measure RPO and RTO.

Half the backup lives on a disk you forgot about

Most Kubernetes recovery plans are quietly split in two: an etcd snapshot on one side, the storage layer's own disk backups on the other. As long as those two halves are taken independently, a restore never gives you a consistent system. Object definitions come from one moment in time, disk contents from another; a PVC cannot find the PV it expects, and a database tries to open with a half-written log. The problem Velero actually solves is not "taking a backup" but capturing cluster objects and persistent volume data as one referentially intact set, in a single operation.

What Velero stores, and where

Velero runs as a controller inside the cluster and drives everything through CRDs. When a Backup object appears, the controller walks the API server, collects the resources you selected as JSON, writes them into a tarball, and uploads it to the object storage described by a BackupStorageLocation. A Restore object runs the same path in reverse: it pulls the archive down and applies resources in a fixed priority order. The default order starts with CRDs, then namespaces, StorageClasses, PVs and PVCs, RBAC objects, with pods near the end.

Two practical consequences follow. First, backup policy is a cluster object like any other, so velero backup create and a checked-in Schedule manifest behave identically and your retention rules can live in Git. Second, because the archive sits outside the cluster, losing the cluster entirely is survivable: install Velero into a fresh cluster, point it at the same storage location, and it syncs the existing backups back in as its own objects.

Volume data is not inside that tarball. There are three separate paths for it, and the one you pick decides whether a cross-cluster restore is possible at all.

Snapshots or file system backup

Provider snapshots. The CSI driver creates a VolumeSnapshot and the data stays inside the storage system behind the disk. It is fast and barely touches the running workload. The catch is that the snapshot is bound to that storage system, and usually to that region. If you lose the cluster but the snapshots survive, fine. If you want to restore into a cluster on a different provider, you are holding nothing portable.

File system backup. Velero's node-agent DaemonSet reads the volume as mounted on the node and writes it to object storage through Kopia. It is portable, needs no CSI snapshot support, and does deduplicated incremental uploads. The cost is honest: it reads a live file system, so there is no snapshot-level consistency. hostPath volumes are not supported, scanning cost becomes visible on large files, and a PVC not mounted by any pod is simply skipped.

CSI snapshot plus data movement. With --snapshot-move-data, Velero takes the CSI snapshot first, then moves its contents to object storage through DataUpload objects and releases the temporary snapshot. You get both a consistent read point and a portable archive; on the way back, DataDownload does the same job in reverse.

The decision rule is short. If your goal is fast rollback inside the same cluster, provider snapshots are enough. If your goal is restoring into another cluster or another provider, the data has to be in object storage, which means --snapshot-move-data or file system backup. If your CSI driver has no snapshot support, file system backup is the only option left and consistency becomes your job, via hooks.

apiVersion: velero.io/v1
kind: Schedule
metadata:
  name: nightly-app
  namespace: velero
spec:
  schedule: "0 2 * * *"
  template:
    includedNamespaces:
      - app
      - app-data
    snapshotMoveData: true
    ttl: 720h0m0s

Narrowing scope without cutting into bone

Selection uses --include-namespaces, --exclude-namespaces, --include-resources and label filters (--selector, or --or-selector when several selectors should match inclusively). Include criteria are additive among themselves; exclusions always win. To keep a specific object out regardless of any filter, the velero.io/exclude-from-backup=true label does it unconditionally.

The classic mistake here is writing a namespace list and calling it done. Namespace filters do not cover cluster-scoped objects: CRDs, ClusterRoles and their bindings, StorageClasses, PVs. Open that backup in a new cluster and your application comes back without its own CRD, while the operator that manages it does nothing and says nothing. Use --include-cluster-resources=true, or at minimum list the CRDs your workload depends on explicitly.

For anything that needs write consistency, use hooks. Annotating a pod with pre.hook.backup.velero.io/command and post.hook.backup.velero.io/command lets you quiesce writes right before the volume is read and release them afterwards. The default hook timeout is 30 seconds; on a large database, forgetting to raise pre.hook.backup.velero.io/timeout gets the hook cut off, and the default on-error behaviour fails the whole backup.

What actually breaks when you restore elsewhere

StorageClass names do not match. The source cluster had fast-ssd, the target has premium-rwo, and the PVC waits in Pending forever asking for a class that does not exist. Velero's mapping ConfigMap fixes this, and note that it is found by its labels rather than its name:

apiVersion: v1
kind: ConfigMap
metadata:
  name: change-storage-class-config
  namespace: velero
  labels:
    velero.io/plugin-config: ""
    velero.io/change-storage-class: RestoreItemAction
data:
  fast-ssd: premium-rwo

Immutable fields. A Service's clusterIP, a StatefulSet's volumeClaimTemplates block, the nodeAffinity rules baked into local PVs: all meaningless or rejected in the target cluster. Pass JSON patch rules through --resource-modifier-configmap to rewrite these during the restore instead of editing manifests by hand afterwards.

Node selectors and taints. Pods pinned to a zone label or a custom node label will not schedule where no node carries that label. The backup looks healthy, the recovery does not work.

Network dependencies. Load balancer addresses, ingress hostnames, external DNS records, certificates. A restore does not carry these, and would be wrong if it did. If your recovery runbook has no DNS step, your RTO number is incomplete.

Admission webhooks. A webhook that issues certificates or injects sidecars will reject the very resources it governs if those are restored before the webhook itself is running. When that dependency exists, restore the platform components in a separate, earlier restore.

Existing resources. By default Velero skips objects that already exist rather than overwriting them. Retry a half-finished restore without knowing that and you get a "completed" status while living with the old objects. --existing-resource-policy=update exists for deliberate overwrites, but it is documented as best-effort and falls back to skipping.

Proving the backup restores

A backup in Completed state proves nothing. The proof is an application that came out of that backup and answered a request. Teams that plan to do this by hand once a month stop doing it, so turn the drill into a job:

velero restore create drill-$(date +%s) \
  --from-backup $(velero backup get -o name | head -1 | cut -d/ -f2) \
  --namespace-mappings app:drill \
  --wait

--namespace-mappings lets you revive the workload in an isolated namespace on the same cluster. Then run a verification Job that does real work: connect to the database, count rows, read the timestamp of the newest record, hit the application's health endpoint. If that Job fails, the drill failed and it should page someone. When it passes, tear the temporary namespace down.

Get your numbers out of the same automation. Compute RPO from the gap between the backup's status.completionTimestamp and the newest record timestamp your verification found; that is real data loss, not the interval printed in your schedule. Compute RTO from the restore's start and completion timestamps plus DNS propagation and verification time. Record both as time series, because RTO grows quietly as the dataset grows and you will only notice it on a graph.

When not to go down this road

If the thing you need to recover is one large database, Velero is the wrong tool. The database's own continuous archiving and point-in-time recovery give you a tighter recovery window and stronger consistency guarantees; use Velero for the objects around it instead. A fully stateless cluster reconciled from a Git repository does not need backups either, since rebuilding is faster. On terabyte-scale volumes the scan cost of file system backup stops being acceptable, and storage-layer replication is the better answer. And if your RPO target is measured in minutes, no scheduled backup tuning will reach it; that is a synchronous or near-synchronous replication problem.

Velero earns its place in the middle case: many namespaces, dozens of interlinked objects, moderate persistent volumes, and a whole that has to remain movable from one cluster to another.