Skip to content

Routing

One protocol per job. Statics are banned in the core (except each router's blackhole anchor), and nothing is ever redistributed — every prefix is originated explicitly with a network statement plus an output filter.

Function Protocol Where
Infra reachability (loopbacks + p2p links) OSPFv2, single area 0 Transport routers only (~12)
Venue /16s into the core eBGP, private AS 64512+V Venue CE ↔ metro PE
Service prefixes, default route, L2VPN signalling iBGP AS 65500, route-reflected Transport routers are clients of the two colo CCRs
Label transport LDP Transport-router links only
L2 overlays BGP-signalled VPLS (RFC 4761) Metro/roof PEs + colo hubs
Internet eBGP OPNsense ↔ ISP
Failure detection BFD Single-hop links only — 60 GHz adjacencies, venue eBGP /31s, OSPF/core links. Not iBGP loopback sessions (multihop — see below)

OSPF — infra reachability only

  • Runs on the ~12 transport routers only. Venues run no OSPF.
  • Carries only transport loopback /32s (passive) and infra /30s (point-to-point network type). No venue subnets, no server VLANs, no default route.
  • Single area 0; redistribute=connected removed everywhere.
  • Costs encode traffic engineering (ref-bandwidth 100G):
Link Cost
100G 1
40G 2
25G / 50G-bond 4 / 2
10G fibre 10
1G / CAT6 100
60 GHz 500 (policy — any all-fibre path beats any 60 GHz hop)
Last resort 5000
Drain (planned work) 4000 both ends
  • BFD on every 60 GHz adjacency (200 ms × 3 = 600 ms) — radios keep Ethernet carrier up while the RF path is dead, so BFD is what actually detects the failure. BFD here is single-hop only (the OSPF/link adjacency over the directly-connected /31); the multihop iBGP loopback sessions carry no BFD (see below).

iBGP & route reflectors

  • CR-COLO-01 and CR-COLO-02 are the two route reflectors (cluster-id = own loopback → two independent RRs), carrying two address families: ip and l2vpn.
  • Each metro/roof PE runs exactly two loopback-sourced iBGP sessions (one to each RR). Venue routers are eBGP CEs, not iBGP clients.
  • All core iBGP is BFD-free by design (2026-07). Every iBGP session — RR↔RR and RR→client — is loopback-to-loopback, i.e. multihop, and carries no BFD (transport ibgp_use_bfd: false; RR→client use_bfd: False). Multihop BFD needs a dedicated /routing bfd configuration (UDP 4784) the core does not install; turning it on made RouterOS reject the session ("BFD forbidden for destination address") and the iBGP hung unestablished — a real outage where a metro-PE's venue /16 never reached the RRs. iBGP liveness instead comes from **single-hop link BFD on the physical hops + OSPF reconvergence (which withdraws the loopback)
  • BGP hold timers. The rule: BFD belongs on single-hop links (physical interface / directly-connected /31) — 60 GHz, venue eBGP /31, core links, edge eBGP — NOT on multihop loopback-to-loopback iBGP sessions.**
  • next-hop-self on core iBGP. The PE→RR iBGP connections set nexthop-choice=force-self, so a PE re-advertises eBGP-learned routes (a venue /16, edge prefixes) into iBGP with its own loopback as next-hop (OSPF-resolvable) rather than the eBGP /31 peer address the rest of the core can't resolve — so those prefixes install active on the RRs instead of sitting inactive.
  • The colo CCRs originate 0.0.0.0/0 into iBGP, conditional on the upstream / OPNsense session being up — if the edge dies, the default is withdrawn fleet-wide. All check-gateway=ping static defaults have been deleted.
  • Server-room subnets (10.128.0.0/16) originate at the colo CCRs (or from OPNsense via eBGP). All static routes in the core are eliminated.
  • Six community/policy filter chains drive origination and tagging; the community scheme:
Community Meaning
65500:100 Venue prefix
65500:200 Server-room prefix
65500:300 Public / PPPoE
65500:911 Starlink-eligible (payments)
65500:666 Blackhole

MPLS / LDP + VPLS

  • LDP on every transport-router interface, transport-address = loopback. Venue links carry no MPLS.
  • BGP-signalled VPLS (RFC 4761) on the l2vpn AF over the existing RR sessions. Per-instance route-distinguisher, import/export route-targets, unique site-id.
  • Route-targets give hub-and-spoke for free: spokes export guest-spoke / import guest-hub; colo hubs do the reverse. Spokes never import each other — structural split-horizon, no venue can bridge to another.
  • Multihoming: dual-homed services advertise the same site-id; BGP elects one designated forwarder and blocks the other attachment — loop-free, no STP.
  • Corporate IP traffic stays plain routed IP — no VRFs.
  • Depends critically on the MTU policy: L2MTU ≥ 1600 on every transport hop or VPLS silently blackholes big frames.

See VPLS numbering in the automation for how RTs, site-ids and hand-off VLANs are derived.

Venue eBGP CE model

  • Private AS 64512+V, eBGP over a /30 (or /31) to the metro PE. BFD is per-uplink: the fibre primary runs BFD on (200 ms × 3); the 60 GHz dish backup runs BFD off with tightened BGP timers (hold 9 s / keepalive 3 s) — the dish sits on an Ethernet port that stays up through RF fades, so aggressive BFD would false-flap and the BGP hold-timer does detection instead.
  • Announces 10.V.0.0/16 only (the loopback sits inside the /16). Receives the default route only (input.filter=default-only).
  • Single-homed venues: one link, one session.
  • Dual-homed venues (e.g. fibre from one hub + 60 GHz from another): two links, same AS. Outbound prefers fibre via higher local-pref; the backup PE applies local-pref 50 on the 60 GHz session for inbound.
  • Guest VLAN 20 rides the venue uplink as a tagged VLAN; the metro PE bridges it into the guest VPLS. The venue never sees guest traffic at L3 and needs no LDP/VPLS.
  • Hub failure (e.g. the Mathew St CCR dies) drops all its venues' fibres together; each fails over independently via 60 GHz to Temple Court (capacity-constrained, protected by QoS).

Quality of service

Queues live only on 60 GHz and Starlink egress (the constrained links). Four DSCP classes: NC (CS6, priority), CRITICAL (AF41 — POS/RADIUS/AD, 30%), BUSINESS (AF21, 40%), BEST-EFFORT (staff + guest, remainder), with max-limit set to ~90 % of rainy-day radio throughput.