Proxmox 3-Node HCI Build — Networks, Local ZFS, Replication¶
Related pages
Design-phase build guide. Operational pages: Setup guide, ZFS storage & replication, Operations, Monitoring.
Version: 2.0 — 2026-09-29 (Ceph → local ZFS + replication)
Build guide for the colo hyperconverged cluster: 3× Proxmox nodes (vm1/2/3), 8× 2.4 TB 10k SAS HDD per node in a local ZFS pool, 10× SFP+ + 2× 1 G per host per DESIGN §7.4. Covers the separated networks (migration/replication, guest, backup, corosync ×2), the ZFS pool, and replication + HA for the VMs that need it.
Revised from Ceph
Version 1.0 of this guide built Ceph on the assumption the disks were SAS SSDs. They are 10k spinning disks (Seagate ST2400MM0129), on which 3-node Ceph manages only ~600 random-write IOPS cluster-wide, so storage moved to local ZFS + replication. The ex-Ceph links to the storage switch pair (renamed sw-storage-1/2) now carry migration + replication (911); the rest are spare.
Capacity up front: 4 mirrors × 2.4 TB = 9.6 TB usable per node; under ~80 % full that's ≈7.7 TB per node, ≈23 TB cluster-wide — minus a second copy of everything replicated (local-only + 2 × replicated ≤ ~23 TB). The ~20 TB data-volume target fits only if at most ~3 TB of it is replicated.
1. Prerequisites (per node)¶
- HBA in pass-through mode, not RAID. ZFS must see 8 raw
/dev/sdXdevices. The nodes have a Dell HBA330 (LSI SAS3008), which is pass-through; no cache, no virtual disks. - Boot/OS on a separate device (BOSS M.2 RAID1, shows as
sdi) — never on a data disk. - BIOS: performance power profile, all NICs enumerated; consistent interface naming across the 3 nodes makes the configs copy-paste.
- Same Proxmox VE version on all 3;
chronyNTP againsttime.cloudflare.com(corosync, HA and replication schedules assume agreed time).
2. Network plan (per node)¶
Ten SFP+ ports form five LACP 802.3ad bonds (layer2+3 hash); the two 1 G
RJ45 ports carry the two unbonded corosync rings. All 12 ports are cabled — no
spares. Each member is a host port from the per-node NIC map; vm3 is cabled
differently on its low ports — always read the NIC map, not just vm1.
| Bond / link | Members → switches | VLAN | Subnet (node IPs .11/.12/.13) | MTU |
|---|---|---|---|---|
bond-guest |
2 → VM-A + VM-B | trunk 932–936 (+ 940 mgmt) | mgmt 10.128.40.0/24 gw .1 (OPNsense) |
1500 (bridge 9000-capable) |
bond-backup |
2 → VM-A + VM-B | 910 | 10.128.10.0/24 — nodes .21/.22/.23 |
9000 |
bond-migrate |
2 → Storage-A + Storage-B | 911 (migration + replication) | 10.128.11.0/24 |
9000 |
| coro0 (1 G, single) | 1 → server stack (VM-A) | 950 | 10.128.50.0/24 |
1500 |
| coro1 (1 G, single) | 1 → storage stack | 951 | 10.128.51.0/24 |
1500 |
Per-node NIC map (host port → role). nic0–nic9/nic10/nic11 here span both
media — the five bonds use the ten SFP+ ports, corosync uses the two 1 G RJ45 ports
(nic0/nic1 on vm1/vm2; nic3/nic4 on vm3). Pin the names with systemd
.link files (match by MAC/PCI path) so they're stable across reboots and the vm3
differences stay explicit:
| NIC | vm1 |
vm2 |
vm3 |
|---|---|---|---|
nic0 |
Corosync ring1 (951, 1 G) | Corosync ring1 (951, 1 G) | Migration (storage pair) |
nic1 |
Corosync ring0 (950, 1 G) | Corosync ring0 (950, 1 G) | Spare (ex-migration) |
nic2 |
Migration (storage pair) | Migration (storage pair) | Backup |
nic3 |
Spare (ex-migration) | Spare (ex-migration) | Corosync ring1 (951, 1 G) |
nic4 |
Backup | Backup | Corosync ring0 (950, 1 G) |
nic5 |
Guest | Guest | Guest |
nic6 |
Reserved iSCSI 922 (ex-Ceph public) | Reserved iSCSI 922 (ex-Ceph public) | Reserved iSCSI 922 (ex-Ceph public) |
nic7 |
Reserved iSCSI 922 (ex-Ceph public) | Reserved iSCSI 922 (ex-Ceph public) | Reserved iSCSI 922 (ex-Ceph public) |
nic8 |
Guest | Guest | Guest |
nic9 |
Spare (ex-migration) | Spare (ex-migration) | Spare (ex-migration) |
nic10 |
Migration (storage pair) | Migration (storage pair) | Migration (storage pair) |
nic11 |
Backup | Backup | Backup |
Two facts about this map (as cabled):
- Migration is an MLAG bond on the storage pair (the ex-Ceph cluster links). Both members land on
the same MLAG domain, so
bond-migrateruns active-active 802.3ad at 20 G. - Corosync is one ring per stack for quorum diversity: ring0 (950) → server stack, ring1 (951) → storage stack. A single switch (or a whole stack) loss then drops only one ring and quorum holds. Corosync stays 1500 MTU — latency-sensitive and bandwidth-trivial, which is why it's on the 1 G ports.
(VLANs 911/950/951 are registered in NetBox and DESIGN §7.1; their subnets 10.128.11/50/51 are not yet NetBox prefixes. Storage & corosync VLANs have no gateway anywhere; only mgmt 940 and guest 93x route, via OPNsense.)
VLAN tagging lives on the switch, not the host. Every bond/link port is an
untagged access port in its VLAN (PVID set switch-side), so the host addresses the
bond/bridge directly — no host-side VLAN subinterface. The only tagged trunk
is bond-guest (932-936 + 940), which keeps its vlan-aware vmbr0 and the vmbr0.940
mgmt subinterface.
/etc/network/interfaces skeleton (node 1 shown; .12/.13 for the others):
# --- Migration + replication (VLAN 911, untagged access) — storage pair (ex-Ceph cluster links) ---
auto bond-migrate
iface bond-migrate inet static
address 10.128.11.11/24
bond-slaves nic10 nic2 # vm3: nic10 nic0
bond-mode 802.3ad
bond-xmit-hash-policy layer2+3
mtu 9000
# unconfigured: nic6 nic7 (storage pair, ex-Ceph public) — reserved for iSCSI 922 MPIO
# nic3 nic9 (VM pair, ex-migration; vm3: nic1 nic9)
# --- Guest (TAGGED trunk 932-936 + 940 mgmt) — the only trunk port ---
auto bond-guest
iface bond-guest inet manual
bond-slaves nic5 nic8 # all nodes
bond-mode 802.3ad
bond-xmit-hash-policy layer2+3
auto vmbr0
iface vmbr0 inet manual
bridge-ports bond-guest
bridge-vlan-aware yes
bridge-vids 932-936,940
auto vmbr0.940
iface vmbr0.940 inet static
address 10.128.40.11/24
gateway 10.128.40.1
# --- Backup (VLAN 910, untagged) — IP straight on the bond ---
auto bond-backup
iface bond-backup inet static
address 10.128.10.21/24 # backup is .21/.22/.23, not .11–.13
bond-slaves nic4 nic11 # vm3: nic2 nic11
bond-mode 802.3ad
bond-xmit-hash-policy layer2+3
mtu 9000
# corosync rings — single 1 G access-port links, address on the raw NIC, one per stack
auto nic1 # ring0 (link0), server stack (VM-A); vm3: nic4
iface nic1 inet static
address 10.128.50.11/24
auto nic0 # ring1 (link1), storage stack; vm3: nic3
iface nic0 inet static
address 10.128.51.11/24
Switch side must be up first: MLAG on both CRS326 pairs (2×40G QSFP+ ISLs (80G)), then per-bond MLAG LACP port pairs, VLANs tagged, jumbo on the storage/backup/migration ports. Bonds won't form active-active without MLAG.
Verify before anything else (from each node):
cat /proc/net/bonding/bond-migrate # both slaves up, LACP aggregated (repeat per bond)
ping -M do -s 8972 10.128.11.12 # jumbo end-to-end on every 9000 net
ping -M do -s 8972 10.128.10.22 # backup nodes are .21/.22/.23
corosync ping: ping 10.128.50.12 && ping 10.128.51.12
3. Cluster + corosync (two rings)¶
On vm1:
vm2/vm3:
pvecm add 10.128.40.11 --link0 10.128.50.12 --link1 10.128.51.12
pvecm add 10.128.40.11 --link0 10.128.50.13 --link1 10.128.51.13
corosync-cfgtool -s shows both links connected.
- 3 nodes = proper quorum; no qdevice needed.
4. Migration network¶
/etc/pve/datacenter.cfg:
bond-migrate, untagged access-911) on the storage pair → a node evacuation runs at wire speed without touching guest/backup. Storage replication (pvesr) uses this same migration network, so every 15-minute sync rides 911 too. (secure = SSH-tunnelled; on this isolated VLAN insecure is defensible and faster for huge-RAM VMs — either is fine, pick one and note it.)
5. Storage — local ZFS¶
Per node, 4 mirrored pairs from the 8 data disks, by stable ID (never sdX, never sdi = BOSS):
ls -l /dev/disk/by-id/ | grep -E 'wwn-.* -> \.\./\.\./sd[a-h]$'
zpool create -o ashift=12 \
-O compression=lz4 -O atime=off -O xattr=sa -O acltype=posixacl -O dnodesize=auto \
vmdata \
mirror /dev/disk/by-id/wwn-<sda> /dev/disk/by-id/wwn-<sdb> \
mirror /dev/disk/by-id/wwn-<sdc> /dev/disk/by-id/wwn-<sdd> \
mirror /dev/disk/by-id/wwn-<sde> /dev/disk/by-id/wwn-<sdf> \
mirror /dev/disk/by-id/wwn-<sdg> /dev/disk/by-id/wwn-<sdh>
pvesm add zfspool vmdata --pool vmdata --content images,rootdir --sparse 1 --blocksize 16k.
- Mirrors, not RAIDZ: random VM I/O on HDDs scales with vdev count.
ashift=12: 4K physical sectors behind 512 logical. Same pool name on every node: one storage entry, and replication requires it. - ARC capped explicitly (root is ext4, so nothing else caps it):
zfs_arc_max128 GiB /zfs_arc_min32 GiB in/etc/modprobe.d/zfs.conf(nodes have 754 GiB RAM; tune on hit rate). - No SLOG: sync writes run at HDD latency. Never
sync=disabled; put latency-critical small disks (Talos etcd) on the BOSSlocal-lvminstead. - Replication + HA only for VMs that need it:
pvesr create-local-job <vmid>-0 <partner> --schedule '*/15', then HA with a strict node-affinity rule to the two nodes holding the disks. Failover loses up to 15 minutes of writes.
Full reasoning, capacity maths and the Ceph removal procedure: ZFS storage & replication.
6. Backup network¶
Backup target: Proxmox Backup Server (PBS), planned — not yet deployed (until it is, nothing is backed up). It lives on VLAN 910; the host reaches it via bond-backup (untagged access-910, IP on the bond). Point storage at the target's 10.128.10.x address so vzdump/PBS traffic rides the dedicated bond, jumbo, never the guest net. Back up every VM — replication isn't a backup. Schedule within Proxmox as usual.
7. Kubernetes (Talos) VMs¶
Talos VMs don't use Proxmox replication or HA — Kubernetes provides their redundancy, one control-plane and worker(s) per node:
- Control plane: disk on
local-lvm(BOSS SSD mirror) — etcd needs lowfsynclatency that the HDD pool can't give. - Workers: disk on
vmdata(local ZFS), not replicated. - Persistent volumes (CSI): deferred; most likely a Dell EqualLogic over iSCSI later.
Details: Storage on Talos.
8. Order of operations¶
- Switch pairs: MLAG + ISLs + VLANs + jumbo → then node bonds (§2) → DF-ping matrix.
- Cluster + both corosync rings (§3);
pvecm statusclean. datacenter.cfgmigration network (§4).- Remove the original Ceph if present, then ZFS pool per node +
vmdatastorage (§5);zpool statusallONLINE. - Backup storage on 910 (§6).
- First VMs; replication + HA for the ones that need it (§5); Talos VMs per §7.
- Failure drills before production — §10.
9. Install — Proxmox ISO onto the BOSS card¶
Use the Proxmox VE ISO (PVE 9), not Debian-13-then-PVE. PVE 9 is Debian 13 (trixie) underneath, so "Debian 13 with Proxmox added" buys nothing except manual work — the ISO ships the PVE kernel, sane boot/partition layout, and no config drift across the three nodes. The Debian route is only worth it for exotic partitioning or PXE-image constraints you don't have.
Per node:
1. BOSS card: configure the 2× M.2 as RAID1 in the BOSS utility (it presents one ~224 GB virtual disk to the OS — this is why plain ext4 is fine, the mirroring is below the OS). UEFI boot.
2. Mount the PVE 9 ISO via iDRAC virtual media (BMC is on OOB — no crash cart needed).
3. Installer: target = the BOSS virtual disk; filesystem ext4 + LVM (default). Don't ZFS-mirror here — there's only one device (BOSS already mirrors). In advanced options set maxroot=40 and leave data at its default (~150 GB) — local-lvm holds the Talos control-plane disks (§7).
4. Temporary network for the installer: it can't do bonds/VLANs. Give it a temporary IP on one standalone NIC (or a switch port set untagged into VLAN 940), install, then apply the full §2 interfaces file and move to the real mgmt address 10.128.40.1x. Keep iDRAC KVM open for that transition.
5. Post-install, each node:
# repos: disable enterprise, enable no-subscription
sed -i 's/^deb/#deb/' /etc/apt/sources.list.d/pve-enterprise.list 2>/dev/null
echo "deb http://download.proxmox.com/debian/pve trixie pve-no-subscription" > /etc/apt/sources.list.d/pve-no-sub.list
apt update && apt full-upgrade -y
apt install -y ifupdown2 chrony iftop
timedatectl set-timezone Europe/London
vm1/2/3, .11/.12/.13), then proceed to §2→§8.
10. Failure drills¶
Run the full set before production — bond member, whole switch (MLAG), corosync ring, disk failure + resilver, replication, planned reboot (HA maintenance mode) and hard node kill + HA failover — with the exact commands and expected results in the Operations runbook.
11. Zabbix monitoring¶
Three layers on the existing Zabbix server: the Proxmox VE by HTTP template (read-only API token) for cluster truth, agent 2 on each node for the OS plus bond.degraded / corosync.rings.ok, and ZFS + replication UserParameters (zpool health and capacity, failed pvesr jobs). Commands and triggers: Monitoring.
12. Quick reference — who talks on what¶
| Traffic | Path |
|---|---|
| VM disk I/O | Local ZFS on each node — no storage network |
Replication (pvesr) + live migration |
911 dedicated SFP+ bond (bond-migrate), storage pair |
| Guest traffic | vmbr0 trunk 932–936, routed/firewalled by OPNsense |
| Backups | 910 dedicated bond |
| Cluster heartbeat | 950 (VM-A) + 951 (storage pair), nothing else on them |
| Host mgmt (GUI/API/SSH) | 940 on the guest bond, gw OPNsense, mgmt-VPN scoped |
| (spare) | ex-Ceph public links (storage pair) and ex-migration links (VM pair), cabled but unused |