Proxmox Setup Guide (0 → Hero)¶
A complete, sequential build of the 3-node HCI cluster — from bare metal to a
validated cluster running its first VM. Follow it top to bottom on all three nodes
(vm1 / vm2 / vm3). The deep-dive reference pages —
ZFS storage & replication, Operations,
Monitoring — are linked at the step where you need them.
Source of the site-specifics (IPs, VLANs, disk layout): Proxmox build (design).
What you're building
3× Proxmox VE 9 nodes, each with a local ZFS pool (8× 2.4 TB 10k SAS HDD as 4
mirrored pairs, ~7.7 TB usable per node). Cluster name pubinvest-hci1. VM
disks on local ZFS; VMs that need HA are replicated to a partner node every 15
minutes. 10× SFP+ + 2× 1 G per host across two
CRS326 MLAG pairs.
Step 0 — Before you start¶
Have these ready:
- [ ] iDRAC / BMC reachable on the OOB network (no crash cart needed).
- [ ] Proxmox VE 9 ISO.
- [ ] The server-room switch fabric built: both CRS326 pairs in MLAG, ISLs up, VLANs tagged, jumbo enabled on the storage/backup/migration ports. Bonds won't form active-active without MLAG.
- [ ] IPAM allocations confirmed in NetBox: node mgmt
.11/.12/.13, storage/corosync VLANs (see Step 4). - [ ] Consistent hardware: same NIC enumeration across all three nodes makes every config copy-paste.
graph LR
HW[0. Rack & firmware] --> INS[1-2. Install PVE 9 on BOSS]
INS --> POST[3. Post-install]
POST --> NET[4. Networking + jumbo verify]
NET --> CL[5. Cluster + corosync]
CL --> MIG[6. Migration net]
MIG --> ZFS[7. ZFS pool + storage]
ZFS --> VM[8. First VM + replication]
VM --> VAL[9. Validate / drills]
VAL --> D2[10. Day-2]
Step 1 — Rack & firmware (per node)¶
- BIOS/UEFI: UEFI boot mode, performance power profile, all NICs enumerated. Enable VT-x (virtualisation). Enable VT-d / IOMMU only if you intend PCIe passthrough to VMs.
- HBA in pass-through mode, not RAID — ZFS must see the 8 raw
/dev/sdXdevices. The nodes have a Dell HBA330 (LSI SAS3008), which is pass-through already; on a RAID controller you'd flash IT firmware or set every disk non-RAID. No controller cache, no virtual disks. - BOSS card: configure the 2× M.2 as RAID1 in the BOSS utility. It presents
one ~224 GB virtual disk to the OS (
sdi) — the mirroring is below the OS, so plain ext4 is fine. The OS never lives on a data disk.
Step 2 — Install Proxmox VE 9 onto the BOSS card¶
Use the PVE 9 ISO, not Debian-13-then-PVE. PVE 9 is Debian 13 ("trixie") underneath, so the ISO buys the PVE kernel, a sane partition layout and zero config drift across nodes for free. The Debian route only helps for exotic partitioning you don't have.
- Mount the ISO via iDRAC virtual media.
- Installer target = the BOSS virtual disk; filesystem ext4 + LVM (default). Don't ZFS-mirror (there's only one device — BOSS already mirrors).
- Advanced options:
maxroot=40, and leavedataat its default size (~150 GB).local-lvmon the BOSS SSDs holds the Talos control-plane VM disks (etcd needs SSD latency — see what goes where). - Temporary installer network: the installer can't do bonds/VLANs. Give it a temporary IP on one standalone NIC (or a switch port set untagged into VLAN 940). You'll apply the real networking in Step 4. Keep the iDRAC KVM open for that cutover.
- Set identical hostname/IP scheme:
vm1/vm2/vm3, mgmt.11/.12/.13.
Step 3 — Post-install (per node)¶
Disable the enterprise repo, enable no-subscription, patch, install helpers:
sed -i 's/^deb/#deb/' /etc/apt/sources.list.d/pve-enterprise.list 2>/dev/null
echo "deb http://download.proxmox.com/debian/pve trixie pve-no-subscription" \
> /etc/apt/sources.list.d/pve-no-sub.list
apt update && apt full-upgrade -y
apt install -y ifupdown2 chrony iftop
timedatectl set-timezone Europe/London
Point chrony at time.cloudflare.com — corosync, HA fencing and replication
schedules all assume the nodes agree on the time.
Step 4 — Networking¶
Each host cables 10× SFP+ + 2× 1 G RJ45. The ten SFP+ ports form five LACP
802.3ad bonds (layer2+3 hash); the two 1 G ports carry the two unbonded
corosync rings. All 12 ports are cabled — no spares.
VLAN tagging lives on the switch, not the host. Every bond/link port is an
untagged access port in its VLAN (PVID set switch-side), so the host addresses the
bond/bridge directly — no host-side VLAN subinterface. The only tagged trunk
is bond-guest (932-936 + 940), which keeps its vlan-aware vmbr0.
Ex-Ceph links re-used
The two Ceph bonds went with Ceph. Their cables to the storage switch pair
(sw-storage-1/2, formerly sw-ceph-1/2) now carry migration + replication
(911) on the ex-cluster ports, and the ex-public ports are spare. Migration's
old ports on the VM pair are spare too. No recabling.
| Bond / link | Members → switches | VLAN | Subnet (node .11/.12/.13) | MTU |
|---|---|---|---|---|
bond-guest |
VM-A + VM-B | trunk 932–936 + 940 mgmt | mgmt 10.128.40.0/24 gw .1 |
1500 |
bond-backup |
VM-A + VM-B | 910 | 10.128.10.0/24 |
9000 |
bond-migrate |
Storage-A + Storage-B | 911 (migration + replication) | 10.128.11.0/24 |
9000 |
| coro0 (1 G, single) | server stack (VM-A) | 950 | 10.128.50.0/24 |
1500 |
| coro1 (1 G, single) | storage stack | 951 | 10.128.51.0/24 |
1500 |
Storage & corosync VLANs have no gateway anywhere; only mgmt 940 and guest 93x route, via OPNsense.
Per-node NIC map. The nicN names span both media — the five bonds use the ten
SFP+ ports, corosync uses the two 1 G RJ45 ports (nic0/nic1 on vm1/vm2;
nic3/nic4 on vm3). Pin the names with systemd .link files (match by MAC/PCI
path) so they're stable across reboots. vm3 is cabled differently on its low
ports, so don't just copy vm1:
| NIC | vm1 |
vm2 |
vm3 |
|---|---|---|---|
nic0 |
Corosync ring1 (951, 1 G) | Corosync ring1 (951, 1 G) | Migration (storage pair) |
nic1 |
Corosync ring0 (950, 1 G) | Corosync ring0 (950, 1 G) | Spare (ex-migration) |
nic2 |
Migration (storage pair) | Migration (storage pair) | Backup |
nic3 |
Spare (ex-migration) | Spare (ex-migration) | Corosync ring1 (951, 1 G) |
nic4 |
Backup | Backup | Corosync ring0 (950, 1 G) |
nic5 |
Guest | Guest | Guest |
nic6 |
Reserved iSCSI 922 (ex-Ceph public) | Reserved iSCSI 922 (ex-Ceph public) | Reserved iSCSI 922 (ex-Ceph public) |
nic7 |
Reserved iSCSI 922 (ex-Ceph public) | Reserved iSCSI 922 (ex-Ceph public) | Reserved iSCSI 922 (ex-Ceph public) |
nic8 |
Guest | Guest | Guest |
nic9 |
Spare (ex-migration) | Spare (ex-migration) | Spare (ex-migration) |
nic10 |
Migration (storage pair) | Migration (storage pair) | Migration (storage pair) |
nic11 |
Backup | Backup | Backup |
How the two special links are cabled
- Migration is an MLAG bond on the storage pair — both members on the same
MLAG domain, so
bond-migrateruns active-active 802.3ad at 20 G, and replication and migration no longer share a switch pair with guest or backup traffic. - Corosync is one ring per stack for quorum diversity: ring0 (950) → server stack, ring1 (951) → storage stack. A single switch (or whole stack) loss then drops only one ring and quorum holds. Corosync stays 1500 MTU — latency- sensitive, bandwidth-trivial, which is why it's on the 1 G ports.
/etc/network/interfaces (node 1 shown — use .12/.13 for the others, and the
vm3 slave deltas noted inline):
# --- Migration + replication (VLAN 911, untagged access) — storage pair (ex-Ceph cluster links) ---
auto bond-migrate
iface bond-migrate inet static
address 10.128.11.11/24
bond-slaves nic10 nic2 # vm3: nic10 nic0
bond-mode 802.3ad
bond-xmit-hash-policy layer2+3
mtu 9000
# unconfigured: nic6 nic7 (storage pair, ex-Ceph public) — reserved for iSCSI 922 MPIO
# nic3 nic9 (VM pair, ex-migration; vm3: nic1 nic9)
# --- Guest (TAGGED trunk 932-936 + 940 mgmt) — the only trunk port ---
auto bond-guest
iface bond-guest inet manual
bond-slaves nic5 nic8 # all nodes
bond-mode 802.3ad
bond-xmit-hash-policy layer2+3
auto vmbr0
iface vmbr0 inet manual
bridge-ports bond-guest
bridge-vlan-aware yes
bridge-vids 932-936,940
auto vmbr0.940
iface vmbr0.940 inet static
address 10.128.40.11/24
gateway 10.128.40.1
# --- Backup (VLAN 910, untagged) — IP straight on the bond ---
auto bond-backup
iface bond-backup inet static
address 10.128.10.11/24
bond-slaves nic4 nic11 # vm3: nic2 nic11
bond-mode 802.3ad
bond-xmit-hash-policy layer2+3
mtu 9000
# corosync rings — single 1 G access-port links, one per stack, address on the raw NIC
auto nic1 # ring0 (link0), server stack (VM-A); vm3: nic4
iface nic1 inet static
address 10.128.50.11/24
auto nic0 # ring1 (link1), storage stack; vm3: nic3
iface nic0 inet static
address 10.128.51.11/24
Apply with ifreload -a (from ifupdown2), then move off the temporary installer
address to the real mgmt 10.128.40.1x — keep the iDRAC KVM open in case.
Verify jumbo before anything else
An MTU mismatch here becomes "replication/migration/backups mysteriously stall under load" later. From each node, prove the bonds aggregated and jumbo works end-to-end (DF ping, 8972 payload = 9000 MTU):
Do not proceed until every DF ping succeeds on every jumbo net.Step 5 — Cluster + corosync (two rings)¶
On vm1:
vm2 / vm3:
pvecm add 10.128.40.11 --link0 10.128.50.12 --link1 10.128.51.12
pvecm add 10.128.40.11 --link0 10.128.50.13 --link1 10.128.51.13
- link0 (VM-A switch) primary, link1 (storage-pair switch) the knet fallback — two rings on two different physical switches, so no single switch loss costs quorum.
- Corosync VLANs carry nothing else.
- Verify:
pvecm status(3 nodes, quorate) andcorosync-cfgtool -s(both links connected). 3 nodes = proper quorum, no qdevice needed.
Step 6 — Migration network¶
/etc/pve/datacenter.cfg:
bond-migrate, untagged access-911) on the storage
pair, so a node evacuation runs at wire speed without touching guest/backup traffic.
Storage replication uses the same network, so this setting also moves every
15-minute replication sync onto 911. (secure = SSH-tunnelled; on this isolated VLAN
insecure is defensible and faster for huge-RAM VMs — pick one and note it.)
Step 7 — ZFS pool¶
If the node still has the original Ceph build on it, remove it first: Removing the original Ceph. Full detail (layout reasoning, ARC, sync writes, capacity) is on ZFS storage & replication. The essential build, on each node:
# the 8 data disks' stable IDs, selected by model — never the BOSS
D=(); for d in $(lsblk -dn -o NAME,MODEL | awk '$2=="ST2400MM0129"{print $1}'); do
D+=("/dev/disk/by-id/$(udevadm info -q symlink /dev/$d | tr ' ' '\n' | grep -m1 '^disk/by-id/wwn-' | cut -d/ -f3)")
done
printf '%s\n' "${D[@]}" # must print exactly 8 wwn- paths — stop if not
zpool create -o ashift=12 \
-O compression=lz4 -O atime=off -O xattr=sa -O acltype=posixacl -O dnodesize=auto \
vmdata \
mirror "${D[0]}" "${D[1]}" mirror "${D[2]}" "${D[3]}" \
mirror "${D[4]}" "${D[5]}" mirror "${D[6]}" "${D[7]}"
# cap the ARC (root is ext4, so nothing else does) — see ZFS storage → ARC size
echo "options zfs zfs_arc_max=$((128 * 1024**3)) zfs_arc_min=$((32 * 1024**3))" > /etc/modprobe.d/zfs.conf
update-initramfs -u -k all
Then once, from any node:
- 4 mirrored pairs, not RAIDZ — random I/O on HDDs needs as many vdevs as possible.
ashift=12— the disks are 4K physical behind 512-byte logical sectors.- Same pool name on all three nodes — one storage entry covers the cluster, and replication requires it.
zpool status vmdatashows 4 mirrors, allONLINE, on every node before continuing.
Step 8 — First VM¶
- Upload an ISO to
local(or an NFS ISO store), or use a cloud-init image. - Create → disk on
vmdata(VirtIO SCSI single,iothread,discard=on); NIC onvmbr0with the right VLAN tag (930 Tier-1 / 931 Tier-2 / 932 DMZ). - Boot, install the guest OS, confirm it gets an address from OPNsense on its zone VLAN.
- If the VM needs to survive a node failure, replicate it to a partner node and put it under HA, pinned to those two nodes — see Replication + HA:
- Kubernetes (Talos) nodes follow a different placement — control-plane disks on
local-lvm, workers onvmdata, neither replicated — see Storage on Talos.
Step 9 — Validate before production¶
Take a baseline, then run the failure-drill runbook in full — bond, switch (MLAG), corosync ring, disk failure/resilver, replication, planned-reboot and hard-kill/HA drills, each with expected results.
fio --name=randwrite --filename=/vmdata/fio.test --size=8G --rw=randwrite --bs=16k \
--iodepth=32 --ioengine=libaio --direct=1 --runtime=60 --time_based --group_reporting
rm /vmdata/fio.test # note IOPS + latency per node
zpool status -x ; pvecm status ; corosync-cfgtool -s # pools healthy, 3 nodes, both links
The acceptance test that ties it together: pull the power on a node running a replicated HA test VM (drill 7) — the VM must come back on its partner node within a few minutes, with data up to its last replication sync.
Step 10 — Day-2¶
- Monitoring: wire up Zabbix — PVE API, agent2 (incl. the
bond.degraded/corosync.rings.okcustom checks), and the ZFS pool and replication checks. - Backups: point vzdump/PBS at the target on VLAN 910 (ZFS storage → backup); schedule in Proxmox. Back up every VM — replication isn't a backup, and non-replicated VMs have no other copy.
- HA: only for replicated VMs, and always with a strict node-affinity rule to the nodes that hold the replica (details).
- Planned reboots: see the routine — replicated VMs live-migrate quickly; local-only VMs either migrate with a full disk copy or shut down for the reboot.
- Updates: patch one node at a time using the planned-reboot procedure; never all three at once.