Backup & Restore — k8s Data¶
Protecting the Kubernetes workloads that store data on persistent volumes.
Applies once CSI exists
There is no CSI driver or StorageClass yet. Ceph was removed, and the
replacement (most likely a Dell EqualLogic over iSCSI) is still to be designed. See
Storage on Talos. Until then the
only cluster state to protect is etcd, via talosctl etcd snapshot
(Storage on Talos → Backups). The Velero approach below is
the plan for when volumes arrive; the driver-specific values are placeholders.
Two things need protecting together:
- The volume data — PVC contents on the CSI-backed storage.
- The Kubernetes objects that describe each app (Deployments, Services, PVCs, ConfigMaps…), so a restore recreates the whole app, not just a disk.
The automated tool for this is Velero — it snapshots volumes via CSI, moves the data to off-cluster object storage, captures the k8s resources alongside, runs on a schedule, and restores to the same or a different cluster.
graph LR
subgraph Cluster
APP[App + PVC] --> VS[CSI VolumeSnapshot]
VEL[Velero + node-agent]
end
VS --> VEL
VEL -->|Kopia data mover| S3[(Off-cluster S3 bucket)]
VEL -->|k8s resources| S3
Don't back storage up onto the same storage
A CSI snapshot alone lives on the same array — same failure domain, so it is a fast rollback, not a backup. The whole point of the Velero data mover is to copy snapshot data to an independent, off-cluster S3 target (a separate MinIO or offsite object storage). Point the backup target somewhere that survives losing both the HCI cluster and the array.
Prerequisites¶
- External-snapshotter installed. Talos/upstream k8s does not bundle it. Install
the snapshot CRDs (
VolumeSnapshot,VolumeSnapshotClass,VolumeSnapshotContent) and thesnapshot-controllerif your distro doesn't already run one. -
A VolumeSnapshotClass for the CSI driver, labelled so Velero picks it up:
apiVersion: snapshot.storage.k8s.io/v1 kind: VolumeSnapshotClass metadata: name: <driver>-snapclass labels: velero.io/csi-volumesnapshot-class: "true" # Velero selects it via this label driver: <csi driver name> deletionPolicy: RetainThe driver must support CSI snapshots. If the one chosen doesn't, Velero can still back volumes up with a file-system backup (
--default-volumes-to-fs-backup), at the cost of reading every file each time. 3. An off-cluster S3 bucket for the Velero BackupStorageLocation, with credentials.
Install Velero (the auto tool)¶
Velero 1.14+ has CSI support built in; enable the node-agent (Kopia) for data movement to object storage.
velero install \
--provider aws \
--plugins velero/velero-plugin-for-aws:v1.10.0 \
--bucket velero-k8s \
--secret-file ./s3-credentials \
--use-node-agent \
--backup-location-config \
region=minio,s3ForcePathStyle=true,s3Url=https://s3.backup.pubinvest.co.uk \
--features=EnableCSI # no-op on 1.14+, still needed on older Velero
./s3-credentials is an AWS-style credentials file for the bucket — a secret,
kept out of git and out of this portal. On Talos the node-agent DaemonSet runs fine;
it needs the privileged/host mounts the Velero manifests already request, so put
Velero in a namespace labelled pod-security.kubernetes.io/enforce=privileged.
Automated scheduled backups¶
A schedule is the "auto" part — Velero fires it on cron, snapshots the PVCs, and moves the data to S3:
velero schedule create daily-apps \
--schedule="0 1 * * *" \
--include-namespaces app-tier1,app-tier2 \
--snapshot-move-data \
--ttl 336h0m0s # 14-day retention
--snapshot-move-datatakes a CSI snapshot, then the node-agent copies its data to the bucket via Kopia — so the backup is independent of the source storage.--ttlgives retention; run several schedules (e.g. daily + weekly with a longer TTL) for a grandfather-father-son policy.- Scope with
--include-namespaces/ label selectors; exclude scratch namespaces.
Application-consistent backups¶
CSI snapshots are crash-consistent. For databases (POS backend, payments — Tier 1), that isn't enough. Use Velero backup hooks to quiesce before the snapshot:
annotations:
pre.hook.backup.velero.io/command: '["/bin/sh","-c","fsfreeze -f /var/lib/data || true"]'
post.hook.backup.velero.io/command: '["/bin/sh","-c","fsfreeze -u /var/lib/data || true"]'
For anything transactional, prefer a proper application dump (e.g. a scheduled
pg_dump / native DB backup) written to a PVC that Velero then backs up — belt and
braces for the data that matters most.
Restore¶
Granular or full — same command, from a named backup:
velero backup get # list backups
velero restore create --from-backup daily-apps-20260719010000
velero restore describe <restore-name> # watch progress / partial failures
- Single namespace / app:
--include-namespaces app-tier1. - Single PVC into a running app: restore with a namespace + label selector, or restore the PVC and re-point the workload.
Full DR to a rebuilt / second cluster¶
- Stand up the cluster and install the CSI driver, the StorageClass, the VolumeSnapshotClass, the snapshot-controller, and Velero pointed at the same S3 bucket (as read).
velero restore create --from-backup <name>. The node-agent pulls the volume data from S3 into fresh volumes on the new storage, and Velero recreates the k8s objects.- This works even if the storage backend differs on the new cluster — the data comes from object storage, not from an array-to-array link.
Verify & monitor (make it trustworthy)¶
- Test restores on a schedule — a backup you've never restored is a hope, not a backup. Restore into a scratch namespace periodically and check the data.
- Alert on failure: Velero exposes Prometheus metrics
(
velero_backup_failure_total,velero_backup_last_successful_timestamp). Alert when a scheduled backup fails or hasn't succeeded within its window — the same disaster/warning split as the rest of the estate.
Storage-native alternatives¶
Array-level snapshots and replication (whatever the chosen backend offers) are useful for whole-array DR but not k8s-app-aware: they don't capture Deployments/Services, so they don't give a one-command app restore. Revisit this section once the CSI backend is chosen; use Velero for the day-to-day, app-granular, scheduled, restorable backups that most recoveries actually need.
Secrets
S3 bucket credentials and any CSI driver credentials are secrets — inject via sealed-secrets / SOPS / an external secrets operator. They are not in this portal.