ZFS Storage & Replication¶
Each node keeps its VM disks on a local ZFS pool built from its own 8 data disks. There is no shared storage: VMs that need to survive a node failure are replicated to a partner node every 15 minutes and put under HA. Everything else runs local-only and relies on backups.
Why not Ceph
The 8 data disks per node are 10k SAS spinning disks (Seagate
ST2400MM0129), not SSDs. Ceph on 24 HDDs with the journal on the same disks
manages roughly 600 random-write IOPS across the whole cluster, at tens of
milliseconds latency. A ZFS pool of 4 mirrored pairs gives about the same
per node, with no network round trip, and ZFS batches ordinary writes into
sequential flushes that suit spinning disks. The original Ceph build is removed
with the one-off teardown below.
Disk layout (per node)¶
| Device | What | Use |
|---|---|---|
sda–sdh |
8× Seagate ST2400MM0129, 2.4 TB 10k SAS, 4K physical / 512 logical |
ZFS pool vmdata |
sdi |
Dell BOSS virtual disk (2× M.2 SSD, RAID1), ~224 GB | Proxmox OS + local-lvm |
Each node has 754 GiB RAM (see ARC size).
BOSS layout as built. The BOSS is 223.5 GB on every node, but nodes 1 and 3 were
installed with the default root size rather than maxroot=40, so their local-lvm
is smaller. Not worth reinstalling: shrinking an ext4 root in place isn't practical,
and 130 GB still holds a 40 GB Talos control-plane disk with room to spare.
| Node | pve-root |
pve-swap |
pve-data (local-lvm) |
|---|---|---|---|
pve-colo-1 |
65.6 GB | 8 GB | 130.2 GB |
pve-colo-2 |
40 GB | 8 GB | 154.8 GB |
pve-colo-3 |
65.6 GB | 8 GB | 130.2 GB |
The disks sit behind a Dell HBA330 (LSI SAS3008), a pure pass-through HBA, so ZFS
sees the raw disks. Never point a disk command at the BOSS, and don't rely on sdX
names — they can reorder across reboots. Every command below selects the data disks by
model (ST2400MM0129), and the pool is built from /dev/disk/by-id/wwn-* paths.
Create the pool¶
On each node, collect the 8 data disks' stable IDs by model, then create 4 mirrored pairs (the ZFS equivalent of RAID10) with the same pool name on every node:
D=(); for d in $(lsblk -dn -o NAME,MODEL | awk '$2=="ST2400MM0129"{print $1}'); do
D+=("/dev/disk/by-id/$(udevadm info -q symlink /dev/$d | tr ' ' '\n' | grep -m1 '^disk/by-id/wwn-' | cut -d/ -f3)")
done
printf '%s\n' "${D[@]}" # must print exactly 8 /dev/disk/by-id/wwn-… lines — stop if not
zpool create -o ashift=12 \
-O compression=lz4 -O atime=off -O xattr=sa -O acltype=posixacl -O dnodesize=auto \
vmdata \
mirror "${D[0]}" "${D[1]}" mirror "${D[2]}" "${D[3]}" \
mirror "${D[4]}" "${D[5]}" mirror "${D[6]}" "${D[7]}"
zpool status vmdata # 4 mirror vdevs, all ONLINE
- Mirrors, not RAIDZ. RAIDZ performs like a single disk per vdev for random VM I/O. 4 mirrors give 4 disks' worth of random writes and ~8 of reads, and a failed disk resilvers from its partner only.
ashift=12matches the disks' 4K physical sectors (they report 512 logical, which would otherwise tempt ZFS intoashift=9).- Same pool name on all three nodes is what lets one storage definition cover the cluster, and what Proxmox replication requires.
Add it to Proxmox once (the storage config is cluster-wide):
--sparse 1 makes VM disks thin-provisioned, so watch real pool usage (see
capacity), not allocated disk sizes.
ARC size¶
The root filesystem is ext4 on the BOSS, so the installer did not cap the ZFS read cache (ARC). Uncapped, it can take up to half the node's RAM from the VMs. Set it explicitly. On spinning disks the ARC is the main thing that makes reads fast — every cache hit saves a ~5–8 ms seek — so give it far more than the minimum (~2 GiB + 1 GiB per TiB of pool ≈ 11 GiB). The nodes have 754 GiB RAM; start at 128 GiB max / 32 GiB min (~17 %, leaving ~620 GiB per node for VMs). The minimum stops VM memory pressure squeezing the cache to nothing.
echo "options zfs zfs_arc_max=$((128 * 1024**3)) zfs_arc_min=$((32 * 1024**3))" > /etc/modprobe.d/zfs.conf
update-initramfs -u -k all
echo $((128 * 1024**3)) > /sys/module/zfs/parameters/zfs_arc_max # apply now, no reboot
echo $((32 * 1024**3)) > /sys/module/zfs/parameters/zfs_arc_min
Tune on real load: arc_summary | grep -A3 "ARC total accesses". A hit rate well
above 95 % means RAM can go back to VMs; lower with busy disks means raise it (the
echo lines change it live).
VM memory budget: HA VMs from a failed node restart on the survivors, so plan total
VM RAM for two nodes: roughly 2 × (754 − 128 ARC − ~16 host) ≈ 1.2 TB cluster-wide.
Sync writes¶
There is no SSD write log (SLOG), so synchronous writes (databases, fsync-heavy
guests) land on the spinning disks and run at HDD latency. Leave sync=standard.
Never set sync=disabled on the pool to speed things up: it trades the speed for
losing acknowledged writes on a crash. Latency-critical small disks go on the BOSS
instead (see Talos control-plane disks).
Debian's zfsutils-linux already scrubs every pool monthly (second Sunday), and ZED
mails root on any pool fault, so make sure root's mail is forwarded.
What goes where¶
| Workload | Storage | Replicated | HA |
|---|---|---|---|
| VMs that must survive a node failure | vmdata (ZFS) |
Yes, every 15 min | Yes |
| Other VMs | vmdata (ZFS) |
No | No (restore from backup) |
| Talos control-plane VMs (etcd) | local-lvm (BOSS SSD) |
No | No: Kubernetes provides HA |
| Talos worker VMs | vmdata (ZFS) |
No | No: Kubernetes provides HA |
Talos VM disk placement and settings are covered in Storage on Talos.
VM disk settings on vmdata: SCSI controller VirtIO SCSI single, iothread=1,
discard=on (so freed guest blocks return to the thin pool), cache Default (no
cache).
Replication + HA¶
Proxmox storage replication (pvesr) snapshots a VM's ZFS disks on a schedule and
sends only the changes to a partner node. It uses the migration network
(10.128.11.0/24, set in datacenter.cfg). It's asynchronous: if the source node
dies, HA restarts the VM on the partner from the last completed sync, so up to
15 minutes of writes are lost. Only replicate VMs where that's acceptable and a
cold restart elsewhere is useful.
For each VM that needs it (example: VM 110 on vm1, partner vm2):
pvesr create-local-job 110-0 vm2 --schedule '*/15'
pvesr status # job listed, State OK after the first full sync
Then put it under HA and pin it to the nodes that hold its replica. HA can only
restart the VM where its disks exist; if it tries a node without a replica, the start
fails. In PVE 9: Datacenter → HA → Resources → Add (vm:110), then Datacenter →
HA → Affinity Rules → Add → Node Affinity with nodes vm1, vm2 and Strict
ticked.
- A VM can have jobs to both other nodes (
110-0→vm2,110-1→vm3) if it should be able to restart anywhere, at the cost of a third copy. - The first sync copies the whole disk; after that only changes move, so live migration to the replica node is quick. Proxmox reverses the job's direction automatically after a migration or HA recovery.
- Disks that shouldn't be copied (scratch, caches) can be excluded per disk with
replicate=0. - The initial sync of large disks is a full read of HDDs. Stagger new jobs or use
--rate <MB/s>so it doesn't starve the VMs.
Capacity¶
Per node: 4 mirrors × 2.4 TB = 9.6 TB usable (8.7 TiB). Keep ZFS below ~80 % full: performance on a fuller pool drops sharply, especially on HDDs. That gives ≈7.7 TB of data per node, ≈23 TB across the cluster.
Every replicated VM uses space on two nodes. With U TB of local-only VMs and R
TB of replicated VMs, the budget is roughly U + 2R ≤ 23 TB, and each node's own share
has to fit that node. The original ~20 TB data target is only reachable if at most ~3
TB of it is replicated.
Replication also keeps one snapshot per job on each side. With 15-minute deltas that's small, but it counts against the pool.
Backup network¶
The backup target will be Proxmox Backup Server (PBS) — planned, not yet
deployed. Until it exists nothing is backed up, so hold production VMs until it is.
It lives on VLAN 910; each node reaches it via bond-backup
(10.128.10.21/.22/.23). Point the PBS storage at its 10.128.10.x address so
vzdump/PBS traffic rides the dedicated jumbo bond, never the guest network.
Backups are what protect the non-replicated VMs, and replication is not a backup: a deleted file or a ransomware-encrypted disk replicates within 15 minutes. Back up everything, replicated or not.
Who talks on what¶
| Traffic | Path |
|---|---|
| VM disk I/O | Local to each node (ZFS on the HBA). No storage network |
Replication (pvesr) + live migration |
911 dedicated SFP+ bond (bond-migrate), storage pair (sw-storage-1/2) |
| Guest traffic | vmbr0 trunk 932–936, routed/firewalled by OPNsense |
| Backups | 910 dedicated bond |
| Cluster heartbeat | 950 (VM-A) + 951 (storage pair), nothing else |
| Host mgmt (GUI/API/SSH) | 940 on the guest bond, gw OPNsense, mgmt-VPN scoped |
| (spare) | ex-Ceph public links (storage pair) and ex-migration links (VM pair): cabled, unused |
Removing the original Ceph¶
One-off: the cluster was first built with Ceph (HEALTH_OK) before the disks turned
out to be HDDs. ZFS needs the same 8 disks, so Ceph has to go first.
This destroys everything on Ceph
Check nothing you need is on it: pvesm list vmpool (and any other RBD/CephFS
storage) must be empty, or every guest on it backed up with vzdump to the backup
target on VLAN 910 so it can be restored onto vmdata afterwards.
-
Remove the Proxmox storage entries and pools (any node):
ceph osd pool ls # every pool except .mgr (that goes in step 4) pveceph pool destroy vmpool --remove_storages 1 # repeat for each pool, e.g. k8s-rbd # only if CephFS exists: stop every MDS first (on each node that has one) pveceph stop --service mds.$(hostname) pveceph fs destroy cephfs --remove-storages 1 --remove-pools 1 -
Destroy the OSDs, on each node in turn:
-
Remove the daemons. On each node:
pveceph mds destroy $(hostname)(if any MDS) andpveceph mgr destroy $(hostname). On two of the three nodes:pveceph mon destroy $(hostname). The last MON can't be removed this way; it goes with the purge. -
Delete the
.mgrpool.pveceph purgerefuses while any pool exists. With the managers gone it won't be recreated. On the node that still has a MON: -
Purge Ceph on each node, the one with the last MON last:
systemctl stop ceph.target pveceph purge --crash --logs apt purge -y ceph-mon ceph-osd ceph-mgr ceph-mds # keep ceph-common/ceph-fuse: proxmox-ve depends on them systemctl disable --now ceph-crash # crash collector (ceph-base) — nothing left to watch systemctl list-units 'ceph*' # empty ls /etc/pve/ceph.conf 2>/dev/null; grep -E '^(rbd|cephfs):' /etc/pve/storage.cfg # both empty -
Wipe the data disks — by model, never by
sdXname, so the BOSS can't be hit:lsblk -o NAME,MODEL,SIZE,TYPE,MOUNTPOINTS # 8× ST2400MM0129 bare; the DELLBOSS VD holds pve-* vgs | grep ceph # left-over ceph-* VG? vgremove -y <vg> dmsetup ls | grep ceph # should be empty; else dmsetup remove <name> for d in $(lsblk -dn -o NAME,MODEL | awk '$2=="ST2400MM0129"{print $1}'); do echo "== /dev/$d"; wipefs -a /dev/$d done # verify: no signatures, no partitions/LVs under any data disk for d in $(lsblk -dn -o NAME,MODEL | awk '$2=="ST2400MM0129"{print $1}'); do echo "== /dev/$d"; wipefs /dev/$d; lsblk -n /dev/$d | tail -n +2 donewipefs -aremoves every filesystem, LVM and partition-table signature, including the backup GPT.Device or resource busyon a data disk means a leftover Ceph LV still holds it — redo thevgremove/dmsetupstep. (The BOSS is always busy, sowipefscouldn't erase it even by mistake.) -
Remove the ex-Ceph public links — see Ex-Ceph public links below.
Then create the pool.
Ex-Ceph public links (reserved for iSCSI)¶
Only after Ceph is purged — until then they carry the MONs/OSDs on 10.128.12.x.
The ex-cluster links (bond-cephclu, switch p1–3) are different: they become
bond-migrate in the 911 cutover, so leave them.
On each host:
ifdown vmbr920 bond-cephpub
# delete the bond-cephpub + vmbr920 stanzas from /etc/network/interfaces
ifreload -a
ip -br link show nic6 nic7 # no master, no IP
On each storage switch (sw-storage-1/2, under /system safe-mode; the live names
are still the Ceph-era ones):
/interface bridge vlan remove [find vlan-ids=920]
/interface bridge port remove [find interface~"^bond-vm[123]-pub\$"]
/interface bonding remove [find name~"^bond-vm[123]-pub\$"]
/interface ethernet set [find name~"^sfp-sfpplus[4-6]\$"] disabled=yes comment="SPARE::ex-ceph-public, reserved iSCSI 922"
Check: /interface bonding print shows no -pub bonds; /interface bridge vlan print
shows no 920.
Why iSCSI, and why unbonded. These are each host's two links to the storage pair, one per switch — exactly the two independent paths a Dell EqualLogic wants for MPIO (multipath, one IP per path), which LACP bonding would defeat. When the array arrives:
- Switch: p4–6 become single-homed access ports, untagged VLAN 922, no
mlag-id; 922 tagged onbond-peerso a path on one switch reaches array ports on the other; edge ports, RX flow control on (Dell's EqualLogic guidance). Split the array's controller ports across both switches (p10–24 are free). - Hosts: one
10.128.16.xIP per NIC, MTU 9000, no gateway — on small bridges (vmbr922aonnic6,vmbr922bonnic7) so either Proxmox or Talos VMs can use them. Two NICs in one subnet neednet.ipv4.conf.all.arp_ignore=1andarp_announce=2, or replies leave via the wrong NIC. - Trade-off: backup (910) can then only move to the storage pair by recabling the freed VM-pair ports (p15–17) across.