Skip to content

Storage on Talos

Where the Talos Kubernetes VMs keep their disks on the Proxmox HCI cluster, and why. Persistent volumes for workloads (CSI) are deferred: see below.

The short version

  • Control-plane VMs → local-lvm on the BOSS SSD mirror, because etcd needs SSD latency.
  • Worker VMs → vmdata, the node's local ZFS pool on 10k HDDs.
  • Neither is replicated or under Proxmox HA. Kubernetes provides the redundancy, with one control plane and the workers spread across all three nodes.

Layout

VM Per Proxmox node Disk storage Disk size Replication / HA
Control plane 1 (3 total → etcd quorum) local-lvm (BOSS SSD) 40 GB None
Worker 1 or more vmdata (local ZFS, HDD) 100 GB+ None

Why the control plane goes on the BOSS

etcd writes every change with fsync and expects it to finish in under ~10 ms. vmdata is spinning disks with no SSD write log, so synchronous writes run at HDD latency. Under load that shows up as etcd slow fdatasync warnings, leader elections and a flaky API server. The BOSS card is two mirrored M.2 SSDs; local-lvm on it has ~150 GB free, which is plenty for one 40 GB control-plane disk per node.

The BOSS SSDs are boot-grade drives, not write-heavy datacentre SSDs. etcd's write volume is modest, but watch their wear through iDRAC (see Monitoring). Nothing else busy belongs on local-lvm.

Why workers go on ZFS, unreplicated

Worker disks hold the Talos system, container images, logs and emptyDir. All of it can be rebuilt, so there's nothing to replicate. A worker lost with its node is replaced by the other nodes' workers. Proxmox replication would only double the space and add HDD load.

Why no Proxmox HA

  • The disks are local and unreplicated, so HA couldn't restart the VMs elsewhere anyway.
  • Kubernetes already handles a lost node: etcd keeps quorum on 2 of 3 control planes, and pods reschedule onto the surviving workers.
  • Size the workers for N+1: two nodes' workers must be able to carry the whole workload while the third node is down or in maintenance.

VM settings

For both roles:

Setting Value Why
SCSI controller VirtIO SCSI single One I/O thread per disk
Disk iothread=1, discard=on, ssd=1 TRIM passes through to the thin pool / thin LV
Cache Default (no cache) Safe with both ZFS and LVM
CPU type host Full instruction set; the VMs never live-migrate
Memory Fixed, ballooning off The kubelet plans around a fixed memory size
QEMU guest agent Enabled Needs the siderolabs/qemu-guest-agent extension (below)
Network vmbr0, VLAN tag of the k8s zone No storage NIC needed: there's no Ceph to reach
Start at boot On So a rebooted host brings its k8s nodes back

Talos installs to the VM's single SCSI disk (machine.install.disk: /dev/sda).

Talos image

Build the installer and ISO from the Talos Image Factory with at least:

  • siderolabs/qemu-guest-agent, so Proxmox can shut the VM down cleanly and see its IPs.

If the planned iSCSI storage goes ahead, it will also need siderolabs/iscsi-tools and siderolabs/util-linux-tools. Adding extensions later is a talosctl upgrade to a new installer image, not a reinstall, so there's no need to add them now.

Maintenance

VMs on local disks don't live-migrate cheaply, so a host reboot takes its k8s nodes down with it. Do it one Proxmox node at a time:

kubectl drain <worker> --ignore-daemonsets --delete-emptydir-data
kubectl drain <control-plane> --ignore-daemonsets --delete-emptydir-data
# shut the Talos VMs down from Proxmox (guest agent → clean shutdown), then reboot the host
# once the host is back and the VMs have started:
talosctl -n <control-plane-ip> etcd status        # all 3 members healthy before the next node
kubectl uncordon <control-plane> <worker>

This fits into the host's planned-reboot routine. Never take down two control planes at once: etcd loses quorum.

Backups

The Talos VMs themselves don't need vzdump backups. A node is rebuilt from its machine config in minutes. What does need backing up:

  • etcd: schedule talosctl etcd snapshot and ship it to the backup target. It's the whole cluster state.
  • Machine configs and talosconfig: in git (secrets via SOPS), never in this portal.
  • Workload data: none on the cluster yet; see Backup & restore for when CSI arrives.

Persistent volumes (CSI) — deferred

There's no StorageClass yet. Workloads that need persistent volumes wait for it. Ceph is gone, and local ZFS can't provide volumes that survive a node loss.

The likely direction is a Dell EqualLogic array over iSCSI. Things to settle when that's designed:

  • Driver: Dell's CSI drivers (PowerStore, PowerMax, PowerScale, PowerFlex, Unity) don't cover the EqualLogic PS series. The realistic options are to present EqualLogic LUNs to Proxmox as shared LVM storage and use the Proxmox CSI plugin, or to use static iSCSI PersistentVolumes.
  • Network: VLAN 922 iSCSI (10.128.16.0/24, L2 only, jumbo) is already reserved for this on the storage switch pair. EqualLogic wants MPIO over separate, unbonded host ports, not LACP, so it needs its own ports per host.
  • Talos: the iSCSI extensions above, if nodes connect to the array directly.