Monitoring¶
The HCI cluster is monitored by the same Zabbix server as the rest of the estate
(Tier-2, 10.128.36.x), in three layers: the PVE API for cluster truth, the agent
for the OS, plus ZFS pool and replication checks for storage.
Secrets excluded
The commands below generate an API token. That secret value is configured in Zabbix and not reproduced in this portal.
a) Proxmox VE by HTTP — the cluster view¶
One read-only API token, one Zabbix host:
pveum user add zabbix@pve --comment "monitoring"
pveum acl modify / --users zabbix@pve --roles PVEAuditor
pveum user token add zabbix@pve monitor --privsep 0 # note the secret once shown
{$PVE.URL.HOST}=10.128.40.11,
{$PVE.URL.PORT}=8006, {$PVE.TOKEN.ID}=zabbix@pve!monitor,
{$PVE.TOKEN.SECRET}=<secret>. Auto-discovers nodes, VMs/LXC and storages; alerts on
quorum, node down, storage %.
b) Zabbix agent 2 on each node — the OS view¶
apt install zabbix-agent2 (Zabbix repo for Debian 13), template Linux by Zabbix
agent, polled over the mgmt 940 addresses. Add two UserParameters the templates
lack — a silent redundancy loss is exactly what monitoring is for:
# /etc/zabbix/zabbix_agent2.d/cluster.conf
UserParameter=bond.degraded,grep -l 'MII Status: down' /proc/net/bonding/* 2>/dev/null | wc -l
UserParameter=corosync.rings.ok,corosync-cfgtool -s 2>/dev/null | grep -c 'link enabled:1.*link connected:1'
bond.degraded > 0 = warning ("running on one leg");
corosync.rings.ok < 2 = warning, < 1 = disaster.
c) ZFS pool + replication — the storage view¶
No Ceph means no storage plugin: the pool and replication state come from three more
agent2 UserParameters. zpool runs fine as the zabbix user; pvesr status needs
root, so allow just that one command:
# /etc/zabbix/zabbix_agent2.d/storage.conf
UserParameter=zfs.pool.health,zpool list -H -o health vmdata
UserParameter=zfs.pool.capacity,zpool list -H -o capacity vmdata | tr -d %
UserParameter=pvesr.failed,sudo -n /usr/bin/pvesr status 2>/dev/null | awk 'NR>1 && $NF!="OK"' | wc -l
echo 'zabbix ALL=(root) NOPASSWD: /usr/bin/pvesr status' > /etc/sudoers.d/zabbix-pvesr
chmod 440 /etc/sudoers.d/zabbix-pvesr
Triggers:
zfs.pool.health≠ONLINE→ high.DEGRADEDmeans a mirror is running on one disk: replace it before its partner fails.zfs.pool.capacity> 80 → warning, > 90 → high. HDD pools slow down sharply past 80 %, and--sparseVM disks can fill the pool without anyone resizing anything.pvesr.failed> 0 for 30 min → warning. A failed job means that VM's replica is getting older than 15 minutes, which is the data an HA failover would lose.
ZED also mails root on any pool fault, and the monthly scrub result, so make sure
root's mail is forwarded somewhere read.
BOSS card: its M.2 SSDs sit behind the BOSS controller, so the OS can't read their SMART data. Watch the BOSS virtual disk and M.2 health through iDRAC (Redfish/SNMP) instead — they carry the OS and the Talos etcd disks.
Template trimming¶
Same philosophy as the OPNsense monitoring:
- The PVE template's VM discovery creates items per VM — filter it (macro
{$PVE.VM.NAME.MATCHES}) to production VMs if the item count balloons. - Drop duplicate CPU/mem coverage: the API template and the Linux agent both report node CPU/memory — keep the agent's (finer), mute the API copies.
- Priority triggers to make loud: ZFS pool not
ONLINE, pool > 80 %, replication job failing, quorum lost,bond.degraded, both-corosync-rings, backup job failed. Everything else → daily digest. - Route via the same split as the edge: disaster → on-call, warning → digest.