Skip to content

Backup & Restore — k8s Data

Protecting the Kubernetes workloads that store data on persistent volumes.

Applies once CSI exists

There is no CSI driver or StorageClass yet. Ceph was removed, and the replacement (most likely a Dell EqualLogic over iSCSI) is still to be designed. See Storage on Talos. Until then the only cluster state to protect is etcd, via talosctl etcd snapshot (Storage on Talos → Backups). The Velero approach below is the plan for when volumes arrive; the driver-specific values are placeholders.

Two things need protecting together:

  1. The volume data — PVC contents on the CSI-backed storage.
  2. The Kubernetes objects that describe each app (Deployments, Services, PVCs, ConfigMaps…), so a restore recreates the whole app, not just a disk.

The automated tool for this is Velero — it snapshots volumes via CSI, moves the data to off-cluster object storage, captures the k8s resources alongside, runs on a schedule, and restores to the same or a different cluster.

graph LR
    subgraph Cluster
      APP[App + PVC] --> VS[CSI VolumeSnapshot]
      VEL[Velero + node-agent]
    end
    VS --> VEL
    VEL -->|Kopia data mover| S3[(Off-cluster S3 bucket)]
    VEL -->|k8s resources| S3

Don't back storage up onto the same storage

A CSI snapshot alone lives on the same array — same failure domain, so it is a fast rollback, not a backup. The whole point of the Velero data mover is to copy snapshot data to an independent, off-cluster S3 target (a separate MinIO or offsite object storage). Point the backup target somewhere that survives losing both the HCI cluster and the array.

Prerequisites

  1. External-snapshotter installed. Talos/upstream k8s does not bundle it. Install the snapshot CRDs (VolumeSnapshot, VolumeSnapshotClass, VolumeSnapshotContent) and the snapshot-controller if your distro doesn't already run one.
  2. A VolumeSnapshotClass for the CSI driver, labelled so Velero picks it up:

    apiVersion: snapshot.storage.k8s.io/v1
    kind: VolumeSnapshotClass
    metadata:
      name: <driver>-snapclass
      labels:
        velero.io/csi-volumesnapshot-class: "true"   # Velero selects it via this label
    driver: <csi driver name>
    deletionPolicy: Retain
    

    The driver must support CSI snapshots. If the one chosen doesn't, Velero can still back volumes up with a file-system backup (--default-volumes-to-fs-backup), at the cost of reading every file each time. 3. An off-cluster S3 bucket for the Velero BackupStorageLocation, with credentials.

Install Velero (the auto tool)

Velero 1.14+ has CSI support built in; enable the node-agent (Kopia) for data movement to object storage.

velero install \
  --provider aws \
  --plugins velero/velero-plugin-for-aws:v1.10.0 \
  --bucket velero-k8s \
  --secret-file ./s3-credentials \
  --use-node-agent \
  --backup-location-config \
      region=minio,s3ForcePathStyle=true,s3Url=https://s3.backup.pubinvest.co.uk \
  --features=EnableCSI          # no-op on 1.14+, still needed on older Velero

./s3-credentials is an AWS-style credentials file for the bucket — a secret, kept out of git and out of this portal. On Talos the node-agent DaemonSet runs fine; it needs the privileged/host mounts the Velero manifests already request, so put Velero in a namespace labelled pod-security.kubernetes.io/enforce=privileged.

Automated scheduled backups

A schedule is the "auto" part — Velero fires it on cron, snapshots the PVCs, and moves the data to S3:

velero schedule create daily-apps \
  --schedule="0 1 * * *" \
  --include-namespaces app-tier1,app-tier2 \
  --snapshot-move-data \
  --ttl 336h0m0s                # 14-day retention
  • --snapshot-move-data takes a CSI snapshot, then the node-agent copies its data to the bucket via Kopia — so the backup is independent of the source storage.
  • --ttl gives retention; run several schedules (e.g. daily + weekly with a longer TTL) for a grandfather-father-son policy.
  • Scope with --include-namespaces / label selectors; exclude scratch namespaces.

Application-consistent backups

CSI snapshots are crash-consistent. For databases (POS backend, payments — Tier 1), that isn't enough. Use Velero backup hooks to quiesce before the snapshot:

annotations:
  pre.hook.backup.velero.io/command: '["/bin/sh","-c","fsfreeze -f /var/lib/data || true"]'
  post.hook.backup.velero.io/command: '["/bin/sh","-c","fsfreeze -u /var/lib/data || true"]'

For anything transactional, prefer a proper application dump (e.g. a scheduled pg_dump / native DB backup) written to a PVC that Velero then backs up — belt and braces for the data that matters most.

Restore

Granular or full — same command, from a named backup:

velero backup get                                  # list backups
velero restore create --from-backup daily-apps-20260719010000
velero restore describe <restore-name>             # watch progress / partial failures
  • Single namespace / app: --include-namespaces app-tier1.
  • Single PVC into a running app: restore with a namespace + label selector, or restore the PVC and re-point the workload.

Full DR to a rebuilt / second cluster

  1. Stand up the cluster and install the CSI driver, the StorageClass, the VolumeSnapshotClass, the snapshot-controller, and Velero pointed at the same S3 bucket (as read).
  2. velero restore create --from-backup <name>. The node-agent pulls the volume data from S3 into fresh volumes on the new storage, and Velero recreates the k8s objects.
  3. This works even if the storage backend differs on the new cluster — the data comes from object storage, not from an array-to-array link.

Verify & monitor (make it trustworthy)

  • Test restores on a schedule — a backup you've never restored is a hope, not a backup. Restore into a scratch namespace periodically and check the data.
  • Alert on failure: Velero exposes Prometheus metrics (velero_backup_failure_total, velero_backup_last_successful_timestamp). Alert when a scheduled backup fails or hasn't succeeded within its window — the same disaster/warning split as the rest of the estate.

Storage-native alternatives

Array-level snapshots and replication (whatever the chosen backend offers) are useful for whole-array DR but not k8s-app-aware: they don't capture Deployments/Services, so they don't give a one-command app restore. Revisit this section once the CSI backend is chosen; use Velero for the day-to-day, app-granular, scheduled, restorable backups that most recoveries actually need.

Secrets

S3 bucket credentials and any CSI driver credentials are secrets — inject via sealed-secrets / SOPS / an external secrets operator. They are not in this portal.