Skip to content

ZFS Storage & Replication

Each node keeps its VM disks on a local ZFS pool built from its own 8 data disks. There is no shared storage: VMs that need to survive a node failure are replicated to a partner node every 15 minutes and put under HA. Everything else runs local-only and relies on backups.

Why not Ceph

The 8 data disks per node are 10k SAS spinning disks (Seagate ST2400MM0129), not SSDs. Ceph on 24 HDDs with the journal on the same disks manages roughly 600 random-write IOPS across the whole cluster, at tens of milliseconds latency. A ZFS pool of 4 mirrored pairs gives about the same per node, with no network round trip, and ZFS batches ordinary writes into sequential flushes that suit spinning disks. The original Ceph build is removed with the one-off teardown below.

Disk layout (per node)

Device What Use
sda–sdh 8× Seagate ST2400MM0129, 2.4 TB 10k SAS, 4K physical / 512 logical ZFS pool vmdata
sdi Dell BOSS virtual disk (2× M.2 SSD, RAID1), ~224 GB Proxmox OS + local-lvm

Each node has 754 GiB RAM (see ARC size).

BOSS layout as built. The BOSS is 223.5 GB on every node, but nodes 1 and 3 were installed with the default root size rather than maxroot=40, so their local-lvm is smaller. Not worth reinstalling: shrinking an ext4 root in place isn't practical, and 130 GB still holds a 40 GB Talos control-plane disk with room to spare.

Node pve-root pve-swap pve-data (local-lvm)
pve-colo-1 65.6 GB 8 GB 130.2 GB
pve-colo-2 40 GB 8 GB 154.8 GB
pve-colo-3 65.6 GB 8 GB 130.2 GB

The disks sit behind a Dell HBA330 (LSI SAS3008), a pure pass-through HBA, so ZFS sees the raw disks. Never point a disk command at the BOSS, and don't rely on sdX names — they can reorder across reboots. Every command below selects the data disks by model (ST2400MM0129), and the pool is built from /dev/disk/by-id/wwn-* paths.

Create the pool

On each node, collect the 8 data disks' stable IDs by model, then create 4 mirrored pairs (the ZFS equivalent of RAID10) with the same pool name on every node:

D=(); for d in $(lsblk -dn -o NAME,MODEL | awk '$2=="ST2400MM0129"{print $1}'); do
  D+=("/dev/disk/by-id/$(udevadm info -q symlink /dev/$d | tr ' ' '\n' | grep -m1 '^disk/by-id/wwn-' | cut -d/ -f3)")
done
printf '%s\n' "${D[@]}"          # must print exactly 8 /dev/disk/by-id/wwn-… lines — stop if not

zpool create -o ashift=12 \
  -O compression=lz4 -O atime=off -O xattr=sa -O acltype=posixacl -O dnodesize=auto \
  vmdata \
  mirror "${D[0]}" "${D[1]}" mirror "${D[2]}" "${D[3]}" \
  mirror "${D[4]}" "${D[5]}" mirror "${D[6]}" "${D[7]}"
zpool status vmdata              # 4 mirror vdevs, all ONLINE
  • Mirrors, not RAIDZ. RAIDZ performs like a single disk per vdev for random VM I/O. 4 mirrors give 4 disks' worth of random writes and ~8 of reads, and a failed disk resilvers from its partner only.
  • ashift=12 matches the disks' 4K physical sectors (they report 512 logical, which would otherwise tempt ZFS into ashift=9).
  • Same pool name on all three nodes is what lets one storage definition cover the cluster, and what Proxmox replication requires.

Add it to Proxmox once (the storage config is cluster-wide):

pvesm add zfspool vmdata --pool vmdata --content images,rootdir --sparse 1 --blocksize 16k

--sparse 1 makes VM disks thin-provisioned, so watch real pool usage (see capacity), not allocated disk sizes.

ARC size

The root filesystem is ext4 on the BOSS, so the installer did not cap the ZFS read cache (ARC). Uncapped, it can take up to half the node's RAM from the VMs. Set it explicitly. On spinning disks the ARC is the main thing that makes reads fast — every cache hit saves a ~5–8 ms seek — so give it far more than the minimum (~2 GiB + 1 GiB per TiB of pool ≈ 11 GiB). The nodes have 754 GiB RAM; start at 128 GiB max / 32 GiB min (~17 %, leaving ~620 GiB per node for VMs). The minimum stops VM memory pressure squeezing the cache to nothing.

echo "options zfs zfs_arc_max=$((128 * 1024**3)) zfs_arc_min=$((32 * 1024**3))" > /etc/modprobe.d/zfs.conf
update-initramfs -u -k all
echo $((128 * 1024**3)) > /sys/module/zfs/parameters/zfs_arc_max   # apply now, no reboot
echo $((32 * 1024**3))  > /sys/module/zfs/parameters/zfs_arc_min

Tune on real load: arc_summary | grep -A3 "ARC total accesses". A hit rate well above 95 % means RAM can go back to VMs; lower with busy disks means raise it (the echo lines change it live).

VM memory budget: HA VMs from a failed node restart on the survivors, so plan total VM RAM for two nodes: roughly 2 × (754 − 128 ARC − ~16 host) ≈ 1.2 TB cluster-wide.

Sync writes

There is no SSD write log (SLOG), so synchronous writes (databases, fsync-heavy guests) land on the spinning disks and run at HDD latency. Leave sync=standard. Never set sync=disabled on the pool to speed things up: it trades the speed for losing acknowledged writes on a crash. Latency-critical small disks go on the BOSS instead (see Talos control-plane disks).

Debian's zfsutils-linux already scrubs every pool monthly (second Sunday), and ZED mails root on any pool fault, so make sure root's mail is forwarded.

What goes where

Workload Storage Replicated HA
VMs that must survive a node failure vmdata (ZFS) Yes, every 15 min Yes
Other VMs vmdata (ZFS) No No (restore from backup)
Talos control-plane VMs (etcd) local-lvm (BOSS SSD) No No: Kubernetes provides HA
Talos worker VMs vmdata (ZFS) No No: Kubernetes provides HA

Talos VM disk placement and settings are covered in Storage on Talos.

VM disk settings on vmdata: SCSI controller VirtIO SCSI single, iothread=1, discard=on (so freed guest blocks return to the thin pool), cache Default (no cache).

Replication + HA

Proxmox storage replication (pvesr) snapshots a VM's ZFS disks on a schedule and sends only the changes to a partner node. It uses the migration network (10.128.11.0/24, set in datacenter.cfg). It's asynchronous: if the source node dies, HA restarts the VM on the partner from the last completed sync, so up to 15 minutes of writes are lost. Only replicate VMs where that's acceptable and a cold restart elsewhere is useful.

For each VM that needs it (example: VM 110 on vm1, partner vm2):

pvesr create-local-job 110-0 vm2 --schedule '*/15'
pvesr status                      # job listed, State OK after the first full sync

Then put it under HA and pin it to the nodes that hold its replica. HA can only restart the VM where its disks exist; if it tries a node without a replica, the start fails. In PVE 9: Datacenter → HA → Resources → Add (vm:110), then Datacenter → HA → Affinity Rules → Add → Node Affinity with nodes vm1, vm2 and Strict ticked.

  • A VM can have jobs to both other nodes (110-0 → vm2, 110-1 → vm3) if it should be able to restart anywhere, at the cost of a third copy.
  • The first sync copies the whole disk; after that only changes move, so live migration to the replica node is quick. Proxmox reverses the job's direction automatically after a migration or HA recovery.
  • Disks that shouldn't be copied (scratch, caches) can be excluded per disk with replicate=0.
  • The initial sync of large disks is a full read of HDDs. Stagger new jobs or use --rate <MB/s> so it doesn't starve the VMs.

Capacity

Per node: 4 mirrors × 2.4 TB = 9.6 TB usable (8.7 TiB). Keep ZFS below ~80 % full: performance on a fuller pool drops sharply, especially on HDDs. That gives ≈7.7 TB of data per node, ≈23 TB across the cluster.

Every replicated VM uses space on two nodes. With U TB of local-only VMs and R TB of replicated VMs, the budget is roughly U + 2R ≤ 23 TB, and each node's own share has to fit that node. The original ~20 TB data target is only reachable if at most ~3 TB of it is replicated.

Replication also keeps one snapshot per job on each side. With 15-minute deltas that's small, but it counts against the pool.

Backup network

The backup target will be Proxmox Backup Server (PBS) — planned, not yet deployed. Until it exists nothing is backed up, so hold production VMs until it is. It lives on VLAN 910; each node reaches it via bond-backup (10.128.10.21/.22/.23). Point the PBS storage at its 10.128.10.x address so vzdump/PBS traffic rides the dedicated jumbo bond, never the guest network.

Backups are what protect the non-replicated VMs, and replication is not a backup: a deleted file or a ransomware-encrypted disk replicates within 15 minutes. Back up everything, replicated or not.

Who talks on what

Traffic Path
VM disk I/O Local to each node (ZFS on the HBA). No storage network
Replication (pvesr) + live migration 911 dedicated SFP+ bond (bond-migrate), storage pair (sw-storage-1/2)
Guest traffic vmbr0 trunk 932–936, routed/firewalled by OPNsense
Backups 910 dedicated bond
Cluster heartbeat 950 (VM-A) + 951 (storage pair), nothing else
Host mgmt (GUI/API/SSH) 940 on the guest bond, gw OPNsense, mgmt-VPN scoped
(spare) ex-Ceph public links (storage pair) and ex-migration links (VM pair): cabled, unused

Removing the original Ceph

One-off: the cluster was first built with Ceph (HEALTH_OK) before the disks turned out to be HDDs. ZFS needs the same 8 disks, so Ceph has to go first.

This destroys everything on Ceph

Check nothing you need is on it: pvesm list vmpool (and any other RBD/CephFS storage) must be empty, or every guest on it backed up with vzdump to the backup target on VLAN 910 so it can be restored onto vmdata afterwards.

  1. Remove the Proxmox storage entries and pools (any node):

    ceph osd pool ls                                   # every pool except .mgr (that goes in step 4)
    pveceph pool destroy vmpool --remove_storages 1    # repeat for each pool, e.g. k8s-rbd
    # only if CephFS exists: stop every MDS first (on each node that has one)
    pveceph stop --service mds.$(hostname)
    pveceph fs destroy cephfs --remove-storages 1 --remove-pools 1
    
  2. Destroy the OSDs, on each node in turn:

    for id in $(ceph osd ls-tree $(hostname)); do
      ceph osd out $id
      systemctl stop ceph-osd@$id
      pveceph osd destroy $id --cleanup     # also wipes the disk's Ceph LVM
    done
    ceph osd tree                           # this host's OSDs gone
    
  3. Remove the daemons. On each node: pveceph mds destroy $(hostname) (if any MDS) and pveceph mgr destroy $(hostname). On two of the three nodes: pveceph mon destroy $(hostname). The last MON can't be removed this way; it goes with the purge.

  4. Delete the .mgr pool. pveceph purge refuses while any pool exists. With the managers gone it won't be recreated. On the node that still has a MON:

    ceph osd pool rm .mgr .mgr --yes-i-really-really-mean-it
    ceph osd pool ls                        # empty
    
  5. Purge Ceph on each node, the one with the last MON last:

    systemctl stop ceph.target
    pveceph purge --crash --logs
    apt purge -y ceph-mon ceph-osd ceph-mgr ceph-mds   # keep ceph-common/ceph-fuse: proxmox-ve depends on them
    systemctl disable --now ceph-crash                  # crash collector (ceph-base) — nothing left to watch
    systemctl list-units 'ceph*'                        # empty
    ls /etc/pve/ceph.conf 2>/dev/null; grep -E '^(rbd|cephfs):' /etc/pve/storage.cfg   # both empty
    
  6. Wipe the data disks — by model, never by sdX name, so the BOSS can't be hit:

    lsblk -o NAME,MODEL,SIZE,TYPE,MOUNTPOINTS   # 8× ST2400MM0129 bare; the DELLBOSS VD holds pve-*
    vgs | grep ceph                             # left-over ceph-* VG? vgremove -y <vg>
    dmsetup ls | grep ceph                      # should be empty; else dmsetup remove <name>
    
    for d in $(lsblk -dn -o NAME,MODEL | awk '$2=="ST2400MM0129"{print $1}'); do
      echo "== /dev/$d"; wipefs -a /dev/$d
    done
    
    # verify: no signatures, no partitions/LVs under any data disk
    for d in $(lsblk -dn -o NAME,MODEL | awk '$2=="ST2400MM0129"{print $1}'); do
      echo "== /dev/$d"; wipefs /dev/$d; lsblk -n /dev/$d | tail -n +2
    done
    

    wipefs -a removes every filesystem, LVM and partition-table signature, including the backup GPT. Device or resource busy on a data disk means a leftover Ceph LV still holds it — redo the vgremove/dmsetup step. (The BOSS is always busy, so wipefs couldn't erase it even by mistake.)

  7. Remove the ex-Ceph public links — see Ex-Ceph public links below.

Then create the pool.

Only after Ceph is purged — until then they carry the MONs/OSDs on 10.128.12.x. The ex-cluster links (bond-cephclu, switch p1–3) are different: they become bond-migrate in the 911 cutover, so leave them.

On each host:

ifdown vmbr920 bond-cephpub
# delete the bond-cephpub + vmbr920 stanzas from /etc/network/interfaces
ifreload -a
ip -br link show nic6 nic7          # no master, no IP

On each storage switch (sw-storage-1/2, under /system safe-mode; the live names are still the Ceph-era ones):

/interface bridge vlan remove [find vlan-ids=920]
/interface bridge port remove [find interface~"^bond-vm[123]-pub\$"]
/interface bonding remove [find name~"^bond-vm[123]-pub\$"]
/interface ethernet set [find name~"^sfp-sfpplus[4-6]\$"] disabled=yes comment="SPARE::ex-ceph-public, reserved iSCSI 922"

Check: /interface bonding print shows no -pub bonds; /interface bridge vlan print shows no 920.

Why iSCSI, and why unbonded. These are each host's two links to the storage pair, one per switch — exactly the two independent paths a Dell EqualLogic wants for MPIO (multipath, one IP per path), which LACP bonding would defeat. When the array arrives:

  • Switch: p4–6 become single-homed access ports, untagged VLAN 922, no mlag-id; 922 tagged on bond-peer so a path on one switch reaches array ports on the other; edge ports, RX flow control on (Dell's EqualLogic guidance). Split the array's controller ports across both switches (p10–24 are free).
  • Hosts: one 10.128.16.x IP per NIC, MTU 9000, no gateway — on small bridges (vmbr922a on nic6, vmbr922b on nic7) so either Proxmox or Talos VMs can use them. Two NICs in one subnet need net.ipv4.conf.all.arp_ignore=1 and arp_announce=2, or replies leave via the wrong NIC.
  • Trade-off: backup (910) can then only move to the storage pair by recabling the freed VM-pair ports (p15–17) across.