Skip to content

Compute / Storage

Compute and storage are delivered by a 3-node hyperconverged (HCI) cluster at the colo: Proxmox VE for virtualisation, with each node's VM disks on its own local ZFS pool. There is no shared storage — VMs that need HA are replicated to a partner node every 15 minutes; everything else relies on backups. (The cluster was first built with Ceph; the disks turned out to be 10k HDDs, so it was replaced with ZFS.)

The cluster rides the server-room network fabric (two CRS326 MLAG pairs) — see that page for the switch side and the full VLAN/zone plan. This section covers the compute and storage build itself.

Cluster at a glance

Nodes 3× Proxmox VE 9 — vm1 / vm2 / vm3 (mgmt .11 / .12 / .13)
Hypervisor Proxmox VE 9 (= Debian 13 "trixie" underneath), on BOSS M.2 RAID1
Storage Local ZFS per node: 8× 2.4 TB 10k SAS HDD as 4 mirrored pairs (vmdata); BOSS SSD local-lvm for Talos control-plane disks
Resilience Selected VMs replicated every 15 min (pvesr) + HA; Kubernetes handles its own
Cluster name pubinvest-hci1 (see naming)
Networking 10× SFP+ + 2× 1 G per host — five bonds + two corosync links across the two CRS326 pairs
Quorum 3 nodes, two corosync rings on two different switches (no qdevice)

Cluster naming

Cluster names follow pubinvest-<class><n> so the estate can hold more than one Proxmox cluster without collision. <class> is the storage architecture, <n> the instance:

Cluster Class Storage Status
pubinvest-hci1 hci Storage on the compute nodes — local ZFS + replication (this cluster) building
pubinvest-iscsi1 iscsi External shared iSCSI LUNs (next build) planned

Keep the class in the name — a bare pubinvest-1 wouldn't tell you at the CLI which storage model a node belongs to, which matters when the two clusters differ on the replication, maintenance and backup procedures. (hci1 keeps its name after the move from Ceph to ZFS: storage still lives on the compute nodes, and a Proxmox cluster can't be renamed in place.)

Capacity

4 mirrors × 2.4 TB = 9.6 TB usable per node. Keeping ZFS under ~80 % full gives ≈7.7 TB per node, ≈23 TB across the cluster — but every replicated VM uses space on two nodes, so the real budget is local-only + 2 × replicated ≤ ~23 TB. The original ~20 TB data target only fits if at most ~3 TB of it is replicated. Details: ZFS storage → capacity.

In this section

Page What's in it
Setup guide (0 → hero) Full sequential build: firmware, BOSS install, networking, cluster, ZFS, first VM
ZFS storage & replication Pool layout, ARC, replication + HA, capacity, backup, Ceph removal
Operations The failure-drill runbook — exact commands, expected results
Monitoring Zabbix: PVE API, agent2, ZFS + replication checks
Proxmox build (design) Design-phase HCI build guide: networks, storage, install, drills
Intel X710 NIC runbook Triage, firmware crossflash and SFP+ unlock for X710/XL710

Design principles

  • Storage never routes — migration/replication and corosync VLANs have no gateway anywhere. Backup (910) gateways on OPNsense, scoped to backup ↔ agents; mgmt (940) and guest (93x) also route there.
  • Separated networks — migration/replication, backup and two corosync rings are each on their own bond/VLAN so no traffic class starves another.
  • Jumbo where it counts — backup/migration at MTU 9000 end-to-end; an MTU mismatch shows up later as "replication mysteriously stalls under load".
  • Replicate only what needs it — replication halves usable capacity for that VM and can lose up to 15 minutes; Kubernetes and apps with their own replication don't need it.

Secrets excluded

The Zabbix PVE API token is not in this portal. Pages show the commands that generate them and where they're configured — the secret values live only on the cluster / in the monitoring config.