Skip to content

Server-Room Network

The colo server room is a new build on 10.128.0.0/16. This page covers its network fabric; the compute and storage platforms that ride it are documented under Compute / Storage and Kubernetes.

Two physically separate switch fabrics, each a CRS326 MLAG pair:

  • Storage fabric — 2× CRS326-24S+2Q (sw-storage-1/2, formerly sw-ceph-1/2): migration + Proxmox replication (911) and corosync ring1 (951)
  • VM fabric — 2× CRS326-24S+2Q (sw-vm-1/2)

MLAG peer-link on each pair is a 2×40G QSFP+ bond (80G ISL). Fabric ports run jumbo L2MTU/MTU 9092 (9000 storage jumbo + VLAN-tag headroom). See the crs-fabric device reference.

VLAN / zone plan

VLAN Zone Subnet Gateway Routed?
900 OOB / IPMI / iDRAC / switch mgmt 10.128.9.0/24 OOB gw Mgmt-VPN only
910 Backup 10.128.10.0/24 OPNsense Backup ↔ agents only
911 Migration + Proxmox replication (storage fabric) 10.128.11.0/24 none L2, jumbo
920 Deprecated — ex-Ceph public, removed from the switches 10.128.12.0/23 none —
921 Deprecated — ex-Ceph cluster, removed from the switches 10.128.14.0/23 none —
922 iSCSI / data 10.128.16.0/24 none L2 only, jumbo
932 VM AD & apps (PXE, stock…) 10.128.32.0/24 OPNsense Yes
933 VM Kubernetes 10.128.33.0/24 OPNsense Yes
934 VM Monitoring 10.128.34.0/24 OPNsense Yes
935 VM Data / SQL 10.128.35.0/24 OPNsense Yes (PII)
936 VM Development 10.128.36.0/24 OPNsense Yes
940 Hypervisor mgmt 10.128.40.0/24 OPNsense (mgmt only) Mgmt-VPN only
950 / 951 Corosync ring0 / ring1 10.128.50.0/24 / .51.0/24 none L2
952 SQL cluster / AG replication (physical SQL servers, not Proxmox) 10.128.46.0/24 none L2 only, jumbo

Principles: storage never routes (no gateway, jumbo 9000 end-to-end); every routed VLAN gateways on OPNsense (not the CCRs) for zero-trust; the storage and VM fabrics are physically separate (no link between the pairs — which is why routed backup 910 stays on the VM fabric with OPNsense). The colo CCRs have no direct link to the server switches — server traffic path is: venue → CCR → OPNsense (eBGP) → VM switch → server.

Services migrating from legacy 10.1.x: AD (10.1.84.3), RADIUS (10.1.88.11), UniFi, monitoring.

Out-of-band (OOB) management

  • A dedicated CRS326-24G-2S+RM flat L2 switch (10.201.201.0/24) with its own WAN and WireGuard VPN, source-NATing all VPN clients into the OOB subnet.
  • Structurally non-transit — no routes point through OOB anywhere, so it can never become a data path.
  • Members: the colo CCR 1G copper ports, OPNsense mgmt NICs + IPMI, server iDRAC/IPMI, CRS326 mgmt, the WAN kit (bdr-1/2, bng-1), and a console server on the serial ports.
  • A Zabbix VM (10.201.201.30) gets a second vNIC on the OOB LAN (ports with horizon=2 while device ports are horizon=1) and alerts on OOB WAN down.
  • Colo kit only — metro/roof recovery relies on ring diversity, RoMON and Safe Mode. The OOB switch is deliberately fully manual, outside Ansible: never automate the rescue path with the thing it rescues.

Guest WiFi delivery

The central captive portal runs on a dedicated guest OPNsense instance (separate from the corporate edge, with its own WAN IP and sacrificial NAT egress for reputation/PCI isolation). It serves each venue a /20 from 172.16.0.0/12, delivered per-venue over BGP-signalled VPLS (vpls-guest-<venue>): the spoke PE is the venue's metro router, the hub PE is the colo CCRs. RADIUS backend is 10.1.88.11. One VLAN per venue at the colo hand-off (~30 VLANs).