Skip to content

Guide 02 — Zabbix 7 Platform Build (Zabbix + TimescaleDB + Grafana)

High-level build guide for the central monitoring stack at the colo. Target: three VMs on Proxmox, all covered by the monitoring they host, plus an external dead man's switch.

1. VM layout

VM Sizing (start) Runs
colo-mon-db-01 4 vCPU / 8 GB / 200 GB (SSD-backed Ceph) PostgreSQL 16 + TimescaleDB
colo-mon-zbx-01 4 vCPU / 8 GB / 50 GB Zabbix server 7.0 LTS + frontend
colo-mon-gfx-01 2 vCPU / 4 GB / 20 GB Grafana

At ~700 hosts / 300–500 values per second this is generous; grow later if needed. NetBox stays on its existing VM.

2. Database build

  1. Install PostgreSQL 16 + the TimescaleDB 2.x package (use the Zabbix docs' supported-version matrix for the exact pairing).
  2. CREATE EXTENSION timescaledb; in the zabbix database.
  3. Load the Zabbix server schema, then run the Zabbix-provided timescaledb.sql script — this converts history/trends tables to hypertables.
  4. In Zabbix frontend (Administration → Housekeeping): enable override item history/trend period, history 31d, trends 731d, and enable native TimescaleDB compression for chunks older than 7 days.
  5. Tune Postgres basics: shared_buffers ~25% RAM, max_connections sized to Zabbix pollers + frontend.

Backups: nightly pg_dump (config tables matter most; history is replaceable) plus Proxmox VM backup. Test a restore once before go-live.

3. Zabbix server build

  1. Install Zabbix server 7.0 LTS + frontend (nginx) from official repos, pointed at the DB VM.
  2. Set the frontend behind HTTPS (internal CA or Let's Encrypt via the VPN).
  3. Global secret macros for shared credentials: SNMPv3 user/auth/priv, Proxmox API token, UniFi controller creds, SQL monitoring account.
  4. Enable and size pollers modestly (defaults are close at this scale); watch zabbix[process,*,avg,busy] internal items after onboarding and tune once.
  5. Configure alerting media (email/Telegram/Slack/SMS as preferred) but leave routing to operator-only during parallel run (see Guide 06).

4. Grafana

  1. Install Grafana OSS + the Alexander Zobnin Zabbix datasource plugin (alexanderzobnin-zabbix-app); connect via the Zabbix API with a read-only user.
  2. Build two dashboards first:
  3. NOC overview — venue status grid (one cell per venue driven by the venue router trigger), colo health row (Proxmox cluster, Ceph health, SQL, OPNsense), metro link states.
  4. Venue drilldown — templated on site/host group; router WAN + tunnel, switch, per-AP status/clients, per-till status.
  5. Anonymous read-only viewer access on the NOC dashboard if it goes on a wall screen.

5. Monitoring the monitoring

  • Dead man's switch: cron on the Zabbix server VM curls a Healthchecks.io (or similar external) URL every 5 minutes; the external service emails/pushes if pings stop. This is the only alert path that doesn't depend on the colo.
  • Zabbix monitors its own DB VM, Grafana VM, and NetBox VM via agent (they're just hosts in NetBox like everything else).
  • Watch Zabbix internal health items (queue, cache usage, poller busy rates) with the built-in Zabbix server health template.

6. Order of operations

  1. DB VM → schema → TimescaleDB conversion
  2. Zabbix server + frontend → secret macros → media types
  3. Dead man's switch
  4. Grafana + datasource
  5. Then proceed to Guide 03 (sync) — do not hand-create hosts in the meantime