Compute / Storage¶
Compute and storage are delivered by a 3-node hyperconverged (HCI) cluster at the colo: Proxmox VE for virtualisation, with each node's VM disks on its own local ZFS pool. There is no shared storage — VMs that need HA are replicated to a partner node every 15 minutes; everything else relies on backups. (The cluster was first built with Ceph; the disks turned out to be 10k HDDs, so it was replaced with ZFS.)
The cluster rides the server-room network fabric (two CRS326 MLAG pairs) — see that page for the switch side and the full VLAN/zone plan. This section covers the compute and storage build itself.
Cluster at a glance¶
| Nodes | 3× Proxmox VE 9 — vm1 / vm2 / vm3 (mgmt .11 / .12 / .13) |
| Hypervisor | Proxmox VE 9 (= Debian 13 "trixie" underneath), on BOSS M.2 RAID1 |
| Storage | Local ZFS per node: 8× 2.4 TB 10k SAS HDD as 4 mirrored pairs (vmdata); BOSS SSD local-lvm for Talos control-plane disks |
| Resilience | Selected VMs replicated every 15 min (pvesr) + HA; Kubernetes handles its own |
| Cluster name | pubinvest-hci1 (see naming) |
| Networking | 10× SFP+ + 2× 1 G per host — five bonds + two corosync links across the two CRS326 pairs |
| Quorum | 3 nodes, two corosync rings on two different switches (no qdevice) |
Cluster naming¶
Cluster names follow pubinvest-<class><n> so the estate can hold more than one
Proxmox cluster without collision. <class> is the storage architecture, <n> the
instance:
| Cluster | Class | Storage | Status |
|---|---|---|---|
pubinvest-hci1 |
hci |
Storage on the compute nodes — local ZFS + replication (this cluster) | building |
pubinvest-iscsi1 |
iscsi |
External shared iSCSI LUNs (next build) | planned |
Keep the class in the name — a bare pubinvest-1 wouldn't tell you at the CLI which
storage model a node belongs to, which matters when the two clusters differ on the
replication, maintenance and backup procedures. (hci1 keeps its name after the move
from Ceph to ZFS: storage still lives on the compute nodes, and a Proxmox cluster
can't be renamed in place.)
Capacity¶
4 mirrors × 2.4 TB = 9.6 TB usable per node. Keeping ZFS under ~80 % full gives
≈7.7 TB per node, ≈23 TB across the cluster — but every replicated VM uses space
on two nodes, so the real budget is local-only + 2 × replicated ≤ ~23 TB. The
original ~20 TB data target only fits if at most ~3 TB of it is replicated. Details:
ZFS storage → capacity.
In this section¶
| Page | What's in it |
|---|---|
| Setup guide (0 → hero) | Full sequential build: firmware, BOSS install, networking, cluster, ZFS, first VM |
| ZFS storage & replication | Pool layout, ARC, replication + HA, capacity, backup, Ceph removal |
| Operations | The failure-drill runbook — exact commands, expected results |
| Monitoring | Zabbix: PVE API, agent2, ZFS + replication checks |
| Proxmox build (design) | Design-phase HCI build guide: networks, storage, install, drills |
| Intel X710 NIC runbook | Triage, firmware crossflash and SFP+ unlock for X710/XL710 |
Design principles¶
- Storage never routes — migration/replication and corosync VLANs have no gateway anywhere. Backup (910) gateways on OPNsense, scoped to backup ↔ agents; mgmt (940) and guest (93x) also route there.
- Separated networks — migration/replication, backup and two corosync rings are each on their own bond/VLAN so no traffic class starves another.
- Jumbo where it counts — backup/migration at MTU 9000 end-to-end; an MTU mismatch shows up later as "replication mysteriously stalls under load".
- Replicate only what needs it — replication halves usable capacity for that VM and can lose up to 15 minutes; Kubernetes and apps with their own replication don't need it.
Secrets excluded
The Zabbix PVE API token is not in this portal. Pages show the commands that generate them and where they're configured — the secret values live only on the cluster / in the monitoring config.