Skip to content

Proxmox Setup Guide (0 → Hero)

A complete, sequential build of the 3-node HCI cluster — from bare metal to a validated cluster running its first VM. Follow it top to bottom on all three nodes (vm1 / vm2 / vm3). The deep-dive reference pages — ZFS storage & replication, Operations, Monitoring — are linked at the step where you need them.

Source of the site-specifics (IPs, VLANs, disk layout): Proxmox build (design).

What you're building

3× Proxmox VE 9 nodes, each with a local ZFS pool (8× 2.4 TB 10k SAS HDD as 4 mirrored pairs, ~7.7 TB usable per node). Cluster name pubinvest-hci1. VM disks on local ZFS; VMs that need HA are replicated to a partner node every 15 minutes. 10× SFP+ + 2× 1 G per host across two CRS326 MLAG pairs.

Step 0 — Before you start

Have these ready:

  • [ ] iDRAC / BMC reachable on the OOB network (no crash cart needed).
  • [ ] Proxmox VE 9 ISO.
  • [ ] The server-room switch fabric built: both CRS326 pairs in MLAG, ISLs up, VLANs tagged, jumbo enabled on the storage/backup/migration ports. Bonds won't form active-active without MLAG.
  • [ ] IPAM allocations confirmed in NetBox: node mgmt .11/.12/.13, storage/corosync VLANs (see Step 4).
  • [ ] Consistent hardware: same NIC enumeration across all three nodes makes every config copy-paste.
graph LR
    HW[0. Rack & firmware] --> INS[1-2. Install PVE 9 on BOSS]
    INS --> POST[3. Post-install]
    POST --> NET[4. Networking + jumbo verify]
    NET --> CL[5. Cluster + corosync]
    CL --> MIG[6. Migration net]
    MIG --> ZFS[7. ZFS pool + storage]
    ZFS --> VM[8. First VM + replication]
    VM --> VAL[9. Validate / drills]
    VAL --> D2[10. Day-2]

Step 1 — Rack & firmware (per node)

  1. BIOS/UEFI: UEFI boot mode, performance power profile, all NICs enumerated. Enable VT-x (virtualisation). Enable VT-d / IOMMU only if you intend PCIe passthrough to VMs.
  2. HBA in pass-through mode, not RAID — ZFS must see the 8 raw /dev/sdX devices. The nodes have a Dell HBA330 (LSI SAS3008), which is pass-through already; on a RAID controller you'd flash IT firmware or set every disk non-RAID. No controller cache, no virtual disks.
  3. BOSS card: configure the 2× M.2 as RAID1 in the BOSS utility. It presents one ~224 GB virtual disk to the OS (sdi) — the mirroring is below the OS, so plain ext4 is fine. The OS never lives on a data disk.

Step 2 — Install Proxmox VE 9 onto the BOSS card

Use the PVE 9 ISO, not Debian-13-then-PVE. PVE 9 is Debian 13 ("trixie") underneath, so the ISO buys the PVE kernel, a sane partition layout and zero config drift across nodes for free. The Debian route only helps for exotic partitioning you don't have.

  1. Mount the ISO via iDRAC virtual media.
  2. Installer target = the BOSS virtual disk; filesystem ext4 + LVM (default). Don't ZFS-mirror (there's only one device — BOSS already mirrors).
  3. Advanced options: maxroot=40, and leave data at its default size (~150 GB). local-lvm on the BOSS SSDs holds the Talos control-plane VM disks (etcd needs SSD latency — see what goes where).
  4. Temporary installer network: the installer can't do bonds/VLANs. Give it a temporary IP on one standalone NIC (or a switch port set untagged into VLAN 940). You'll apply the real networking in Step 4. Keep the iDRAC KVM open for that cutover.
  5. Set identical hostname/IP scheme: vm1/vm2/vm3, mgmt .11/.12/.13.

Step 3 — Post-install (per node)

Disable the enterprise repo, enable no-subscription, patch, install helpers:

sed -i 's/^deb/#deb/' /etc/apt/sources.list.d/pve-enterprise.list 2>/dev/null
echo "deb http://download.proxmox.com/debian/pve trixie pve-no-subscription" \
  > /etc/apt/sources.list.d/pve-no-sub.list
apt update && apt full-upgrade -y
apt install -y ifupdown2 chrony iftop
timedatectl set-timezone Europe/London

Point chrony at time.cloudflare.com — corosync, HA fencing and replication schedules all assume the nodes agree on the time.

Step 4 — Networking

Each host cables 10× SFP+ + 2× 1 G RJ45. The ten SFP+ ports form five LACP 802.3ad bonds (layer2+3 hash); the two 1 G ports carry the two unbonded corosync rings. All 12 ports are cabled — no spares.

VLAN tagging lives on the switch, not the host. Every bond/link port is an untagged access port in its VLAN (PVID set switch-side), so the host addresses the bond/bridge directly — no host-side VLAN subinterface. The only tagged trunk is bond-guest (932-936 + 940), which keeps its vlan-aware vmbr0.

Ex-Ceph links re-used

The two Ceph bonds went with Ceph. Their cables to the storage switch pair (sw-storage-1/2, formerly sw-ceph-1/2) now carry migration + replication (911) on the ex-cluster ports, and the ex-public ports are spare. Migration's old ports on the VM pair are spare too. No recabling.

Bond / link Members → switches VLAN Subnet (node .11/.12/.13) MTU
bond-guest VM-A + VM-B trunk 932–936 + 940 mgmt mgmt 10.128.40.0/24 gw .1 1500
bond-backup VM-A + VM-B 910 10.128.10.0/24 9000
bond-migrate Storage-A + Storage-B 911 (migration + replication) 10.128.11.0/24 9000
coro0 (1 G, single) server stack (VM-A) 950 10.128.50.0/24 1500
coro1 (1 G, single) storage stack 951 10.128.51.0/24 1500

Storage & corosync VLANs have no gateway anywhere; only mgmt 940 and guest 93x route, via OPNsense.

Per-node NIC map. The nicN names span both media — the five bonds use the ten SFP+ ports, corosync uses the two 1 G RJ45 ports (nic0/nic1 on vm1/vm2; nic3/nic4 on vm3). Pin the names with systemd .link files (match by MAC/PCI path) so they're stable across reboots. vm3 is cabled differently on its low ports, so don't just copy vm1:

NIC vm1 vm2 vm3
nic0 Corosync ring1 (951, 1 G) Corosync ring1 (951, 1 G) Migration (storage pair)
nic1 Corosync ring0 (950, 1 G) Corosync ring0 (950, 1 G) Spare (ex-migration)
nic2 Migration (storage pair) Migration (storage pair) Backup
nic3 Spare (ex-migration) Spare (ex-migration) Corosync ring1 (951, 1 G)
nic4 Backup Backup Corosync ring0 (950, 1 G)
nic5 Guest Guest Guest
nic6 Reserved iSCSI 922 (ex-Ceph public) Reserved iSCSI 922 (ex-Ceph public) Reserved iSCSI 922 (ex-Ceph public)
nic7 Reserved iSCSI 922 (ex-Ceph public) Reserved iSCSI 922 (ex-Ceph public) Reserved iSCSI 922 (ex-Ceph public)
nic8 Guest Guest Guest
nic9 Spare (ex-migration) Spare (ex-migration) Spare (ex-migration)
nic10 Migration (storage pair) Migration (storage pair) Migration (storage pair)
nic11 Backup Backup Backup

How the two special links are cabled

  • Migration is an MLAG bond on the storage pair — both members on the same MLAG domain, so bond-migrate runs active-active 802.3ad at 20 G, and replication and migration no longer share a switch pair with guest or backup traffic.
  • Corosync is one ring per stack for quorum diversity: ring0 (950) → server stack, ring1 (951) → storage stack. A single switch (or whole stack) loss then drops only one ring and quorum holds. Corosync stays 1500 MTU — latency- sensitive, bandwidth-trivial, which is why it's on the 1 G ports.

/etc/network/interfaces (node 1 shown — use .12/.13 for the others, and the vm3 slave deltas noted inline):

# --- Migration + replication (VLAN 911, untagged access) — storage pair (ex-Ceph cluster links) ---
auto bond-migrate
iface bond-migrate inet static
    address 10.128.11.11/24
    bond-slaves nic10 nic2           # vm3: nic10 nic0
    bond-mode 802.3ad
    bond-xmit-hash-policy layer2+3
    mtu 9000

# unconfigured: nic6 nic7 (storage pair, ex-Ceph public) — reserved for iSCSI 922 MPIO
#                           nic3 nic9 (VM pair, ex-migration; vm3: nic1 nic9)

# --- Guest (TAGGED trunk 932-936 + 940 mgmt) — the only trunk port ---
auto bond-guest
iface bond-guest inet manual
    bond-slaves nic5 nic8            # all nodes
    bond-mode 802.3ad
    bond-xmit-hash-policy layer2+3
auto vmbr0
iface vmbr0 inet manual
    bridge-ports bond-guest
    bridge-vlan-aware yes
    bridge-vids 932-936,940
auto vmbr0.940
iface vmbr0.940 inet static
    address 10.128.40.11/24
    gateway 10.128.40.1

# --- Backup (VLAN 910, untagged) — IP straight on the bond ---
auto bond-backup
iface bond-backup inet static
    address 10.128.10.11/24
    bond-slaves nic4 nic11           # vm3: nic2 nic11
    bond-mode 802.3ad
    bond-xmit-hash-policy layer2+3
    mtu 9000

# corosync rings — single 1 G access-port links, one per stack, address on the raw NIC
auto nic1                            # ring0 (link0), server stack (VM-A); vm3: nic4
iface nic1 inet static
    address 10.128.50.11/24
auto nic0                            # ring1 (link1), storage stack; vm3: nic3
iface nic0 inet static
    address 10.128.51.11/24

Apply with ifreload -a (from ifupdown2), then move off the temporary installer address to the real mgmt 10.128.40.1x — keep the iDRAC KVM open in case.

Verify jumbo before anything else

An MTU mismatch here becomes "replication/migration/backups mysteriously stall under load" later. From each node, prove the bonds aggregated and jumbo works end-to-end (DF ping, 8972 payload = 9000 MTU):

cat /proc/net/bonding/bond-migrate    # both slaves up, LACP aggregated (repeat per bond)
ping -M do -s 8972 10.128.11.12       # migration + replication
ping -M do -s 8972 10.128.10.12       # backup
ping 10.128.50.12 && ping 10.128.51.12   # corosync reachability
Do not proceed until every DF ping succeeds on every jumbo net.

Step 5 — Cluster + corosync (two rings)

On vm1:

pvecm create pubinvest-hci1 --link0 10.128.50.11 --link1 10.128.51.11
On vm2 / vm3:
pvecm add 10.128.40.11 --link0 10.128.50.12 --link1 10.128.51.12
pvecm add 10.128.40.11 --link0 10.128.50.13 --link1 10.128.51.13

  • link0 (VM-A switch) primary, link1 (storage-pair switch) the knet fallback — two rings on two different physical switches, so no single switch loss costs quorum.
  • Corosync VLANs carry nothing else.
  • Verify: pvecm status (3 nodes, quorate) and corosync-cfgtool -s (both links connected). 3 nodes = proper quorum, no qdevice needed.

Step 6 — Migration network

/etc/pve/datacenter.cfg:

migration: secure,network=10.128.11.0/24
A dedicated 20G SFP+ jumbo bond (bond-migrate, untagged access-911) on the storage pair, so a node evacuation runs at wire speed without touching guest/backup traffic. Storage replication uses the same network, so this setting also moves every 15-minute replication sync onto 911. (secure = SSH-tunnelled; on this isolated VLAN insecure is defensible and faster for huge-RAM VMs — pick one and note it.)

Step 7 — ZFS pool

If the node still has the original Ceph build on it, remove it first: Removing the original Ceph. Full detail (layout reasoning, ARC, sync writes, capacity) is on ZFS storage & replication. The essential build, on each node:

# the 8 data disks' stable IDs, selected by model — never the BOSS
D=(); for d in $(lsblk -dn -o NAME,MODEL | awk '$2=="ST2400MM0129"{print $1}'); do
  D+=("/dev/disk/by-id/$(udevadm info -q symlink /dev/$d | tr ' ' '\n' | grep -m1 '^disk/by-id/wwn-' | cut -d/ -f3)")
done
printf '%s\n' "${D[@]}"          # must print exactly 8 wwn- paths — stop if not

zpool create -o ashift=12 \
  -O compression=lz4 -O atime=off -O xattr=sa -O acltype=posixacl -O dnodesize=auto \
  vmdata \
  mirror "${D[0]}" "${D[1]}" mirror "${D[2]}" "${D[3]}" \
  mirror "${D[4]}" "${D[5]}" mirror "${D[6]}" "${D[7]}"

# cap the ARC (root is ext4, so nothing else does) — see ZFS storage → ARC size
echo "options zfs zfs_arc_max=$((128 * 1024**3)) zfs_arc_min=$((32 * 1024**3))" > /etc/modprobe.d/zfs.conf
update-initramfs -u -k all

Then once, from any node:

pvesm add zfspool vmdata --pool vmdata --content images,rootdir --sparse 1 --blocksize 16k
  • 4 mirrored pairs, not RAIDZ — random I/O on HDDs needs as many vdevs as possible.
  • ashift=12 — the disks are 4K physical behind 512-byte logical sectors.
  • Same pool name on all three nodes — one storage entry covers the cluster, and replication requires it.
  • zpool status vmdata shows 4 mirrors, all ONLINE, on every node before continuing.

Step 8 — First VM

  1. Upload an ISO to local (or an NFS ISO store), or use a cloud-init image.
  2. Create → disk on vmdata (VirtIO SCSI single, iothread, discard=on); NIC on vmbr0 with the right VLAN tag (930 Tier-1 / 931 Tier-2 / 932 DMZ).
  3. Boot, install the guest OS, confirm it gets an address from OPNsense on its zone VLAN.
  4. If the VM needs to survive a node failure, replicate it to a partner node and put it under HA, pinned to those two nodes — see Replication + HA:
    pvesr create-local-job <vmid>-0 <partner-node> --schedule '*/15'
    
  5. Kubernetes (Talos) nodes follow a different placement — control-plane disks on local-lvm, workers on vmdata, neither replicated — see Storage on Talos.

Step 9 — Validate before production

Take a baseline, then run the failure-drill runbook in full — bond, switch (MLAG), corosync ring, disk failure/resilver, replication, planned-reboot and hard-kill/HA drills, each with expected results.

fio --name=randwrite --filename=/vmdata/fio.test --size=8G --rw=randwrite --bs=16k \
    --iodepth=32 --ioengine=libaio --direct=1 --runtime=60 --time_based --group_reporting
rm /vmdata/fio.test                            # note IOPS + latency per node
zpool status -x ; pvecm status ; corosync-cfgtool -s   # pools healthy, 3 nodes, both links

The acceptance test that ties it together: pull the power on a node running a replicated HA test VM (drill 7) — the VM must come back on its partner node within a few minutes, with data up to its last replication sync.

Step 10 — Day-2

  • Monitoring: wire up Zabbix — PVE API, agent2 (incl. the bond.degraded / corosync.rings.ok custom checks), and the ZFS pool and replication checks.
  • Backups: point vzdump/PBS at the target on VLAN 910 (ZFS storage → backup); schedule in Proxmox. Back up every VM — replication isn't a backup, and non-replicated VMs have no other copy.
  • HA: only for replicated VMs, and always with a strict node-affinity rule to the nodes that hold the replica (details).
  • Planned reboots: see the routine — replicated VMs live-migrate quickly; local-only VMs either migrate with a full disk copy or shut down for the reboot.
  • Updates: patch one node at a time using the planned-reboot procedure; never all three at once.