Skip to content

Pub Invest Group — Configuration Approach, Protocols & Traffic Engineering

Version: 1.0 — 2026-07-05 Companion to DESIGN.md. All examples are RouterOS 7.19+ syntax and follow the conventions already visible in your live configs (AS 65500, 10.255.255.x loopbacks, 10.254.Y.0/24 p2p ranges, jump-chain venue firewalls).

Complete worked configs (all the fragments below assembled into copy-pasteable files): - example-configs/target-colo-rr.rsc — a full target-state colo core / route reflector (modelled on CR-COLO-01): inter-core bond, RR BGP, eBGP to OPNsense + Starlink, community filter chains, default/server-room origination, guest VPLS hub. - example-configs/target-metro-pe.rsc — a full target-state metro PE (modelled on CR-SEELSTREET-01): OSPF, LDP, iBGP-to-RRs, per-venue eBGP hand-offs, the BGP community scheme applied, guest VPLS spoke, input firewall.


1. Configuration Philosophy

  1. One protocol per job. OSPF = infrastructure reachability. BGP = services (venue /16s, default, server room, public pools). LDP = labels. Static routes = banned in the core (the only statics anywhere are each router's blackhole anchor for its own aggregate).
  2. Never redistribute. redistribute=connected,static is how the current network grew its static-route sprawl and why the BGP table is unpredictable. Everything is originated explicitly with network statements / output filters. If a prefix is in the BGP table, a human put it there on purpose.
  3. Templates, not snowflakes. Three device classes (core/metro CCR, roof RB5009, venue router), each a single template with a small variable set (site name, loopback, uplinks, venue number). Config generation from the template beats hand-editing — even if "generation" starts as a shell script with sed.
  4. Address-lists are the policy API. Firewall rules reference lists (management, routers, ad-servers, pos-outbound-permitted…), never raw IPs. Changing policy = changing a list. Keep list contents identical fleet-wide and distribute changes as one-line scripts.
  5. Config is code. Nightly /export from every device into a git repo (Oxidized or a cron + scp script). Every change gets a diff. RoMON + OOB path means you can always reach a router you just broke.
  6. Fail toward the backup, never toward open. Firewall default-drop everywhere; routing failover is additive (backup path pre-established, just higher cost) so failover never requires config to change.

2. Standard Protocol Configuration

2.1 OSPF (transport routers only — colo, metro, roof; venues run none)

/routing ospf instance
add name=core router-id=<LOOPBACK> version=2
/routing ospf area
add name=backbone area-id=0.0.0.0 instance=core
# Loopback — passive
/routing ospf interface-template
add area=backbone interfaces=loopback passive networks=<LOOPBACK>/32
# Each infra link — p2p, cost per link class (see §4.1)
add area=backbone interfaces=sfp-sfpplus1 type=ptp cost=10  networks=10.254.Y.N/30   ;# fibre
add area=backbone interfaces=vlan-60ghz   type=ptp cost=500 networks=10.254.Y.M/30   ;# 60GHz

Notes: - No redistribute on the instance. Remove redistribute=connected from the current configs. - Bind templates to interfaces, not bare networks — prevents accidental adjacencies on venue-facing ports. - Auth: enable OSPF MD5/SHA auth on all adjacencies (shared per-link keys, rotated annually). Cheap insurance on roof links.

/routing bfd configuration
add interfaces=vlan-60ghz min-tx=200ms min-rx=200ms multiplier=3
Reference it from the OSPF interface-template (use-bfd=yes) and BGP connections. 600 ms detection instead of 40 s dead-interval. On fibre it's optional (loss-of-light already drops the interface) but harmless.

2.3 iBGP with route reflectors

On CR-COLO-01 / CR-COLO-02 (the RRs):

/routing bgp template
add name=rr-clients as=65500 router-id=<LOOPBACK> \
    local.address=<LOOPBACK> .role=ibgp-rr \
    output.default-originate=if-installed \
    routing-table=main
/routing bgp connection
add name=<CLIENT-NAME> template=rr-clients remote.address=<CLIENT-LOOPBACK>/32 .as=65500 listen=yes connect=yes

On every client (metro + roof transport routers only) — exactly two sessions, with both address families:

/routing bgp template
add name=core as=65500 router-id=<LOOPBACK> local.address=<LOOPBACK> .role=ibgp \
    address-families=ip,l2vpn output.network=bgp-networks routing-table=main
/routing bgp connection
add name=RR1 template=core remote.address=10.255.255.255/32 .as=65500
add name=RR2 template=core remote.address=10.255.255.254/32 .as=65500

(address-families=ip,l2vpn on both RR and client templates — the l2vpn family is what signals VPLS, §2.6.)

  • RR redundancy: both RRs with cluster-id = own loopback (default when role=ibgp-rr); clients take best of two reflected copies.
  • Default route: output.default-originate=if-installed on the RRs — they only advertise 0/0 while they themselves hold one (learned via eBGP from OPNsense). Edge dies → default disappears everywhere automatically.

2.3b Venue eBGP (CE model — replaces OSPF/iBGP at venues)

Venue side (AS = 64512+V):

/ip route add dst-address=10.V.0.0/16 blackhole comment="BGP aggregate anchor"
/routing bgp connection
add name=uplink-primary as=6XXXX router-id=10.V.255.255 local.role=ebgp \
    remote.address=10.254.Y.N .as=65500 output.network=bgp-networks \
    input.filter=default-only use-bfd=yes
# dual-homed venues: second session on the backup /30, same everything
/ip firewall address-list add list=bgp-networks address=10.V.0.0/16
/routing filter rule add chain=default-only rule="if (dst == 0.0.0.0/0) { accept } reject"

Metro PE side, one connection per venue:

/routing bgp connection
add name=venue-<NAME> as=65500 local.role=ebgp remote.address=10.254.Y.M .as=6XXXX \
    input.filter=from-venue-<V> output.default-originate=always use-bfd=yes
/routing filter rule add chain=from-venue-<V> \
    rule="if (dst == 10.V.0.0/16) { set bgp-communities 65500:100; accept } reject"

The input filter accepts only that venue's /16 — a compromised or misconfigured venue router cannot inject anything else. Dual-homed: the backup hub's PE adds set bgp-local-pref 50 in its from-venue chain so the core returns traffic via fibre while it lives; the venue prefers its fibre-learned default with local-pref on the primary session. The blackhole anchor keeps the aggregate announced and kills traffic to unused /16 space at the venue.

2.4 eBGP at the edge

OPNsense (FRR) AS 65510 ↔ both colo CCRs; Starlink RB4011 AS 65502 ↔ both colo CCRs. Filters on the CCRs (the contract):

/routing filter rule
# From OPNsense: accept default + public + server-room zones only
add chain=from-edge rule="if (dst == 0.0.0.0/0) { set bgp-local-pref 200; accept }"
add chain=from-edge rule="if (dst in 10.128.0.0/16 && dst-len <= 24) { accept }"
add chain=from-edge rule="reject"
# From Starlink: accept ONLY payment-provider prefixes, never default
add chain=from-starlink rule="if (dst in payment-prefixes) { set bgp-local-pref 50; accept }"
add chain=from-starlink rule="reject"

(payment-prefixes = the address-list of card-processor destinations. Local-pref 50 < 200 ⇒ Starlink path is dormant while the main WAN advertises. See §4.4.)

2.5 MPLS / LDP

/mpls settings set dynamic-label-range=16-1048575
/mpls ldp instance add name=ldp lsr-id=<LOOPBACK> transport-addresses=<LOOPBACK> afi=ip
/mpls ldp interface add interface=sfp-sfpplus1 accept-dynamic-neighbors=yes
On every transport-router infra link (colo/metro/roof). Venue uplinks carry no MPLS. Filter LDP label distribution to loopbacks only (/mpls ldp accept-filter for 10.255.255.0/24) — keeps LFIBs tiny.

2.6 VPLS — BGP-signalled (guest example; PPPoE identical with different names)

Signalling rides the existing iBGP/RR sessions (address-families=ip,l2vpn, §2.3) — no per-peer pseudowire config, and the RRs give auto-discovery: adding a spoke touches only that spoke's PE.

Spoke — the venue's metro PE (venue guest VLAN 20 arrives tagged on the venue-facing port):

/interface bridge add name=br-guest-<V>
/interface vlan add name=guest-<V> interface=<venue-port> vlan-id=20
/interface bridge port add bridge=br-guest-<V> interface=guest-<V>
/routing bgp vpls
add name=guest-<V> bridge=br-guest-<V> site-id=<V> \
    route-distinguisher=65500:20<V> \
    export-route-targets=65500:2001 import-route-targets=65500:2000 \
    pw-type=vpls

Hub — both colo CCRs (RT direction reversed; per-venue VLAN toward the guest firewall trunk):

/routing bgp vpls
add name=guest-hub bridge=br-guest-hub site-id=1000 \
    route-distinguisher=65500:2000 \
    export-route-targets=65500:2000 import-route-targets=65500:2001
  • Hub-and-spoke by route-target: spokes export :2001, import :2000; hubs the reverse. Spokes never import each other's routes, so venues structurally cannot bridge to each other — no reliance on horizon being set correctly on every port.
  • Multihoming/redundancy: two PEs advertising the same site-id (dual-homed venue's guest VLAN on both hub PEs; each ISP POP circuit's dual colo hand-off) triggers BGP designated-forwarder election — the losing PE blocks its attachment circuit. Loop-free, automatic failback, no scripts, no spanning tree. This replaces the cold-standby-pseudowire hack that LDP signalling would force.
  • ISP POP circuits (for the separate WAN business, wan/WAN-DESIGN.md): same VPLS pattern with its own RT pair (e.g. :3000/:3001), spoke PEs at the rooftop hubs, hub PEs on the colo CCRs bridging each to a tagged hand-off port toward the WAN border routers — not a metro BNG. The metro is only the L2 carrier.
  • Verify on your RouterOS version (7.19+) before rollout: /routing bgp vpls instances up, /interface vpls print shows dynamic pseudowires, and MTU end-to-end (§6).

2.7 Device baseline (every router)

/ip service set telnet,ftp,www,api disabled=yes
/ip service set ssh port=22; set www-ssl certificate=<cert> disabled=no
/user aaa set use-radius=yes                 ;# RADIUS login → FreeRADIUS (10.128.36.11 post-migration; 10.1.88.11 today)
/snmp set enabled=yes contact="Pub Invest Group" location="<SITE>"   ;# move to SNMPv3 communities
/system ntp client set enabled=yes
/system ntp client servers add address=time.cloudflare.com
/system clock set time-zone-name=Europe/London
/tool romon set enabled=yes secrets=<romon-secret>
/system logging action add name=remote target=remote remote=10.128.36.20   ;# central syslog (Tier-2 zone)
/system logging add topics=critical,error,warning action=remote
/ip dns set servers=10.128.32.3,10.128.32.4  ;# internal resolvers, Tier-1 zone (10.1.84.3 etc. until Phase 7)

2.7b Naming scheme — role-site[-n], lowercase, unpadded

Discovery proved the old mixed scheme causes real errors (a roof transport router named RTR-CHARLOTTEST-001 misled planning; CR-LEVEL-01 vs -001 drift; LEVEL_IRISH_MASTE truncation). Target scheme:

  • Lowercase everywhere — hostname = DNS = NetBox = Zabbix = Oxidized, identical; aligns with the WAN business (bdr-1, WAN-DESIGN §11).
  • Role token = strict truth. The complete vocabulary — a device carries a token only if it genuinely performs that role:
Token Purpose Examples Notes
cr- Transport router — runs OSPF/LDP, iBGP client or RR (core, metro, roof PE) cr-colo-1/2, cr-coloroof-1, cr-mathewst-1, cr-charlotte-1, cr-charlotteroof-1, cr-seel-1, cr-holmes-1, cr-oldbankroof-1 Roof PEs get "roof" in the site token — distinct failure domain
rtr- Venue CE router — eBGP CE, VLANs/DHCP/zone firewall, nothing else rtr-dos-1, rtr-chr-1, rtr-mcc-1 Venue trigram as site token
fw- Firewall (OPNsense) fw-colo-1/2 (edge pair), fw-guest-1
sw- Switch (L2) sw-storage-1/2, sw-vm-1/2, sw-oob-1, sw-rbl-1 (venue) Replaces legacy SWI-
hv- Hypervisor (Proxmox) hv-1/2/3 All-colo → no site token
srv- Server named by service (VM or generic physical) srv-sql-1/2, srv-unifi-1, srv-zbx-1, srv-dc-1/2, srv-pbs-1 Never sequential-anonymous (srv-core-37). Name the service; the app over the platform where single-purpose. Abbreviations from a documented list; ≤15 chars for Windows (NetBIOS)
w60- / w5- PtP radio unit (60 GHz / 5 GHz) w60-coloroof-oldbank-1 ↔ w60-oldbank-coloroof-1 <near>-<far> encodes location + pointing; no master/slave. Replaces ANT-/LEVEL_*
ap- WiFi access point ap-dos-1 Mirrors into UniFi device names
gw- Small gateway appliance gw-oob-1 (OOB WireGuard entry)
con- Console server con-colo-1
pdu- Power distribution pdu-colo-1a (rack 1, feed A)
(k8s workloads) Named by Kubernetes, not this scheme — This scheme names pets (VMs/hardware); cattle are k8s's problem
bdr- / bng- / pop- / radius- / nms- WAN business (AS 204258) — border, subscriber edge, rooftop access, AAA, monitoring bdr-1/2, bng-1, pop-oldbank-1 Own scheme + DNS zone (WAN-DESIGN §11); same conventions
- Venues: keep the established trigrams (rtr-chr-1 = Cheers, rtr-ynk-1 = Yankees, rtr-mol-1 = Moloko) — they're organisational vocabulary, staff-fluent, short, and already the de-facto key. Three conditions make them safe: (1) the canonical trigram table is a first-class registry artifact (NetBox tenant slug = trigram, display name = full venue name; the table below in this repo until NetBox) — the failure mode isn't the codes, it's the mapping living only in heads; (2) collision rule: repeated brands take the site letter as the 3rd char or a 4th char (mcc = McCooleys Concert Sq, mcs = McCooleys Mathew St — as already practiced), assigned in the table, never ad-hoc; (3) the immutable venue key is venue_id (drives /16, AS, VPLS site-id) — a rebrand is a cheap planned trigram+DNS rename, never a renumber.
- Radios: w60-<near>-<far>[-n] (near = where the unit sits): w60-coloroof-oldbank-1 ↔ w60-oldbank-coloroof-1. No master/slave — the name encodes location + pointing; trailing index for parallel links. w5- for 5 GHz units.
Canonical trigram table (seed — move to NetBox; ? = confirm):
Tri Venue id Tri Venue id Tri Venue id
lvl LEVEL 1 mol Moloko 4 pck Peacock 5
brk Brooklyn Mixer 6 lag Lago 7 cel Celtic Corner 17
dos Dirty O'Sheas 23 ynk Yankees 27 rok Rocking Horse 33
hat Hatch 34 hej Heebie Jeebies 42 kel Kells of Eden 43
chr Cheers ? soh SOHO ? fus Fusion ?
ein Einstein Bier Haus ? nfl Nelly Foleys ? rbl Ruby Blues ?
smo Smokies ? ? mcc McCooleys Concert Sq ? mcs McCooleys Mathew St ?
bpl Boston Pool Loft ? blr Black Rabbit ? olb Old Bank Bar ?
cas Castle St Apts ? lor Lord St ? irh Irish House ?
tav Temple Tavern ? ? wds Wood Street ? zan Zanzibar ?
  • Rules: site token only where geography differs; unpadded -n; no underscores; short (MikroTik identity truncates).
  • Enforced from the registry: network.yaml/NetBox holds the name; Terraform sets /system identity from it; DNS/Zabbix/Oxidized derive from it; the scrape-to-NetBox diff alerts on identity drift.
  • Interface comments: ROLE::far-end (UPLINK::cr-colo-2, VENUE::rtr-dos-1, METRO::, EDGE::, ROOF::, OOB::, SRV::, WAN::, INTERCORE::) — machine-parseable.
  • Migration: opportunistic — rename at planned touches (ROS7 upgrade, eBGP conversion, move windows), paired with the loopback renumber where due. Never big-bang.

(The WAN business keeps its own documented scheme — bdr-N, bng-N, pop-<hub>-N, WAN-DESIGN §11 — same conventions, its own DNS zone.)

3. Firewall Standards

3.1 Core/metro routers — input-only self-protection (your current pattern, canonicalised)

/ip firewall filter
add chain=input connection-state=established,related action=accept
add chain=input connection-state=invalid action=drop
add chain=input protocol=icmp action=accept
add chain=input protocol=tcp dst-port=179 src-address-list=routers action=accept
add chain=input protocol=ospf in-interface-list=neighbour-routers action=accept
add chain=input protocol=udp dst-port=646 in-interface-list=neighbour-routers action=accept  ;# LDP
add chain=input protocol=tcp dst-port=646 src-address-list=routers action=accept
add chain=input protocol=udp dst-port=3784,3785 in-interface-list=neighbour-routers action=accept ;# BFD
add chain=input protocol=tcp dst-port=22,8291,443 src-address-list=management action=accept
add chain=input protocol=udp dst-port=161 src-address-list=snmp-servers action=accept
add chain=input action=drop
No forward chain in the core. Populate routers from the loopback range (10.255.255.0/24 as one entry — simpler than the current per-/32 list and can't drift).

3.2 Venue routers — zone jump-chain model (generalised from RTR-DOS-001)

Chains: wan (internet-only, blocks 10/8), pos (wan + AD + pos-outbound-permitted), trusted (wan + AD + UniFi + file/print lists), and per-interface jumps:

In-interface Jump target
VLAN 10 Office trusted
VLAN 30 POS pos
VLAN 40 Staff wan
VLAN 99 Mgmt trusted (+ UniFi inform)
VLAN 20 Guest none — bridged into VPLS, never routed

Final rule: chain=forward action=drop. Inbound-to-venue: only trusted-inbound list (mgmt VPN, server-room mgmt ranges) may initiate into LAN interfaces. Keep the existing established/related/invalid triplet at the top of both chains. This is exactly your RB4011 policy — the work is making all ~30 venues byte-identical apart from address variables.

3.3 OPNsense zones

Zones: WAN, CORE (transit to CCRs), SRV-T1, SRV-T2, DMZ, OOB, BACKUP, GUEST, MGMT-VPN. Default deny inter-zone; explicit rules per flow (documented in a flow matrix — start one in this repo). All venue→server-room flows are enforced here again (defence in depth vs venue router compromise).

4. Traffic Engineering

4.1 The cost scheme (this is 90% of your TE)

The scheme is reference-bandwidth 100G: wired cost = 100G ÷ link speed (bonds use aggregate). Wireless values are policy, not formula.

Link class OSPF cost Result
100G 1 Floor
40G 2 (2.5 truncated, as auto-cost routers do)
25G (or 2×25G bond = 50G) 4 (bond: 2) Inter-core bond = 2
10G fibre 10 Preferred metro standard
1G fibre / CAT6 metro 100 Beats wireless, loses to any 10G path ≤ 9 hops
60 GHz 500 Policy value — any all-wired path beats one wireless hop. Never recompute from bandwidth
Last-resort (e.g. temporary LTE) 5000 Absolute fallback

ECMP requires equal costs — parallel same-class links balance (Holmes 2×10G); mixed-speed pairs (25G+10G) won't, making the slower link pure backup, which is usually correct.

Rules of thumb: any all-fibre detour (even 4 hops = cost 40) beats one 60 GHz hop (500). Two 60 GHz hops (1000) still beat a dead path. Never hand-tune a single link cost to fix a traffic problem — change the class or add a link; per-link exceptions rot.

ECMP falls out for free: Holmes' two diverse 10G fibres (cost 10 each, one per colo CCR) load-balance per-connection automatically; Old Bank's 2× 60 GHz likewise until its fibre lands (then fibre cost 10 wins outright).

4.2 Verifying and simulating paths

  • /routing route print where dst-address=10.V.0.0/16 and /tool traceroute from loopback to loopback are the ground truth.
  • Before/after any cost change, snapshot ip route print on the affected routers and diff.
  • Planned-work drain: to empty a link before maintenance, set its OSPF cost to 4000 on both ends, wait for reconvergence (seconds), then work. Restore after. This is the only sanctioned manual TE knob.

4.3 Do you need RSVP-TE? (Mostly no)

RouterOS supports RSVP-TE tunnels, but with metric-classed links, ECMP, and BFD you get: deterministic primary/backup, fast reroute-ish behaviour, and drain capability — without per-LSP state on CCR2004s. Adopt RSVP-TE only if a concrete need appears (e.g. "guest VPLS must never exceed 2G on the Mathew St fibre while POS gets priority" can't be met with QoS alone). Revisit then; skip now. QoS (§7) is the right tool for contention, not TE tunnels.

Steady state: RB4011 (AS 65502) announces nothing/low-pref; card traffic follows default via OPNsense.

  • Populate payment-prefixes address-list (processor endpoints — e.g. the 34.117.150.224-style IPs already in pos-outbound-permitted) and keep it in git.
  • RB4011 statically routes those prefixes into the Starlink WAN with NAT, and announces them via eBGP with bgp-local-pref 50 (per §2.4 filter).
  • Main WAN healthy → OPNsense's default (local-pref 200) wins for everything.
  • Main WAN dead → OPNsense withdraws default (its own eBGP upstream drops, or use FRR conditional advertisement) → RRs stop originating 0/0 → for payment prefixes the only remaining route is Starlink's specifics → POS keeps trading; bulk traffic (guest, staff, backups) correctly blackholes rather than crushing the Starlink uplink.
  • Test quarterly: pull the ISP hand-off in a maintenance window, confirm a test transaction completes over Starlink, confirm guest WiFi is down (that's success, not failure).

4.5 BGP community scheme (tag at origin, act at the edge)

Community Meaning Applied where
65500:100 Venue aggregate Venue routers on their /16
65500:200 Server-room subnet Colo CCRs / OPNsense
65500:300 Corporate public routed block (Pub Invest /28 on OPNsense) edge in-filter
65500:911 Eligible for Starlink failover Anything that must survive WAN loss
65500:666 Blackhole (DDoS/incident) Manual; edge drops matching

Communities cost nothing now and make future policy ("only 65500:911 routes may resolve via Starlink") one filter rule instead of a re-architecture.

5. Zero-Trust Hardening Checklist

  • [ ] All router mgmt bound to loopbacks; management list = mgmt VPN (10.201.201.0/24) + jump hosts only
  • [ ] RADIUS login on every RouterOS device; local admin renamed, strong password, emergency-only
  • [ ] OSPF auth on all adjacencies; BGP sessions protected by input filter (§3.1); consider TCP-MD5 on eBGP
  • [ ] SNMPv3 (or at minimum per-device communities restricted to snmp-servers list)
  • [ ] RoMON secrets set; OOB VLAN 900 reachable only via mgmt VPN
  • [ ] Venue: no inter-VLAN rules exist; forward default-drop verified with a scan from each VLAN
  • [ ] Server room: storage VLANs have no gateway; confirmed by scan
  • [ ] Guest/PPPoE: confirm no route to/from 10/8 (traceroute + firewall counters); guest sources are 172.16/12 only — alert on any 172.16/12 source appearing inside the corporate routing domain (it means a bridge/VPLS misconfig)
  • [ ] Nightly config export to git; alert on diff without a change ticket
  • [ ] Central syslog + login alerting; netflow (/ip traffic-flow) on colo CCRs to the collector

6. MTU Plan (do this before enabling MPLS)

6.1 The overhead math — where the number comes from

The pseudowire carries the customer's whole Ethernet frame; MPLS adds labels + an optional control word on top. Worst case is PPPoE (RFC 4638) with a VLAN tag:

Component Bytes
Inner customer Ethernet header 14
Customer VLAN tag (sub 3xx / guest 200) 4
PPPoE + PPP headers 8
Customer IP (RFC 4638 clean 1500) 1500
inner frame carried by the PW 1526
MPLS: 2 labels (tunnel + VC) = 8 + control word 4 12
L2 payload the infra link must carry (L2MTU) 1538

Guest-only (no PPPoE) ≈ 1530. Round up for a possible 3rd label / 2nd tag → L2MTU floor = 1600.

6.2 What to set

Segment L2MTU MPLS MTU IP MTU Why
Fibre infra links 9000 (CCR2004 max ~9200) — 9000 Go jumbo — removes MTU as a concern for VPLS/Q-in-Q/future forever
60 GHz links radio max, must be ≥ 1600 — match The real constraint; verify per radio (WIRELESS-60G §2) — a sub-1600 hop can't carry full-MTU VPLS
MPLS (mpls-mtu), every infra iface — 1600 network-wide — Set to the 60 GHz floor, not fibre's 9000 — MPLS never fragments, so a jumbo labeled packet built on fibre would be dropped hitting a 60 GHz hop. Labeled packets are ≤1538, so 1600 is ample and uniform
Venue /30s & VLAN interfaces (inherit) — 1500 MPLS works at L2MTU; the routed IP MTU doesn't change
Storage VLANs (Ceph/iSCSI) 9000 — 9000 Jumbo end-to-end, but never traverse MPLS/VPLS (L2-local)

6.3 Order & verify

Per link: raise L2MTU both ends → set mpls-mtu=1600 → then build the VPLS. A pseudowire across one forgotten low-MTU link silently blackholes oversized frames (works for ping, fails for real traffic). Test each new path first: ping <far-underlay> size=1580 do-not-fragment (1580 > the 1538 requirement) — then a real 1500-DF ping through the VPLS from a customer.

Fibre at 10G is uncongested; queue only on 60 GHz egress (and Starlink). Simple 4-class model, classified by DSCP marked at the venue router mangle:

Class DSCP Traffic 60 GHz share
NC CS6 OSPF/BGP/BFD/LDP priority 1 (small, guaranteed)
CRITICAL AF41 POS subnets, RADIUS, AD auth guaranteed 30%
BUSINESS AF21 Office, mgmt, UniFi, backups (rate-capped in hours) guaranteed 40%
BEST-EFFORT 0 Staff + guest WiFi remainder, first to starve

Implementation: /queue tree on each 60 GHz interface with max-limit set to ~90% of real measured radio throughput (measure with btest, radios degrade in rain — set against a rainy-day number). During a hub-fibre failure the venue's whole load shifts to 60 GHz; this policy is what keeps tills working while guest WiFi degrades — which is the correct order.

8. Monitoring & Operations

  • LibreNMS (or current SNMP pollers — today 10.1.84.16/18, moving to Tier-2 10.128.36.x in Phase 7): all devices, alert on OSPF neighbour count change, BGP session down, 60 GHz RSSI/MCS drop, interface errors.
  • Netflow from colo CCRs → collector, for capacity planning and the "what filled the Mathew St fibre" question.
  • Oxidized for config backup/diff (RouterOS supported natively).
  • Smokeping/graph latency loopback-to-loopback per hub — 60 GHz path degradation shows in latency before it shows in loss.
  • NetBox as source of truth: sites, links, /30 assignments, venue /16 registry, VLANs, loopback registry. The addressing tables in DESIGN.md §3 seed it.

9. Migration Plan (from current state)

Each phase is independently valuable and reversible; nothing breaks the running network until its cutover step.

The phases group into two stages:

  • Stage 1 — the routed IP network (Phases 0–3 + 8): ROS7 everywhere, clean OSPF core, RR-based BGP, eBGP venues, new fibre core carrying all traffic, ROS6 hardware killed off. Complete and stable before any overlay work.
  • Stage 2 — L2 overlays (Phases 4–7): LDP, guest VPLS, ISP POP circuits (L2 carriage for the separate WAN business), server-room re-org. Purely additive — each venue/rooftop comes on independently with no change to Stage 1 routing.

Phase 0 — Foundations (no traffic impact) NetBox + git config backups; baseline configs (§2.7) everywhere; central syslog/monitoring; document current static routes per router (the kill list); run the discovery collection (discovery/collect.sh) for the as-built map.

Phase 0b — ROS7 gate Every router that stays in the target design runs ROS 7.19+ before its links move to the new core. Legacy ROS6 hardware (e.g. the CCR1016 at LEVEL) is not upgraded in place — the new ROS7 core is built alongside, links re-home onto it (backups first), and the ROS6 devices are decommissioned once nothing terminates on them. No overlay feature (LDP/VPLS/BGP-VPLS) is configured anywhere until this gate is fully passed.

Phase 1 — MTU + BFD Raise L2 MTU on all infra links (both ends, one link at a time, off-hours); enable BFD on 60 GHz adjacencies. Verify with DF pings.

Phase 2 — Clean IGP (transport routers only) Add OSPF interface-templates on colo/metro/roof infra links (p2p type, class costs per §4.1); remove redistribute=connected from OSPF instances; remove OSPF from venue links as each venue converts to eBGP in Phase 3; confirm LSDB contains only transport loopbacks + inter-hub /30s.

Phase 3 — eBGP venues + RR core, alongside the old mesh Configure RR templates (ip+l2vpn) on colo CCRs; metro/roof routers become clients. Convert venues one at a time to the CE model (§2.3b): eBGP session(s) up, /16 announced, default received, then remove that venue's OSPF/iBGP/static remnants. Old mesh sessions on transport routers stay up for comparison (/routing route print count-only where bgp); when identical-or-better, remove mesh sessions and redistribute=connected,static, then delete the static-route kill list one router at a time, checking reachability after each.

Phase 4 — MPLS/LDP Enable LDP on transport-router infra links only (harmless alongside — labels unused until a VPLS exists). Verify LDP neighbours = OSPF neighbours on every transport router.

Phase 5 — Guest VPLS pilot One venue (pick a dual-homed fibre venue, e.g. SOHO): tag VLAN 20 up its uplink, build the BGP-signalled spoke on its metro PE + hub on the colo CCRs, hand off to the guest firewall, run two weeks. Confirm DF election behaves by failing the primary uplink. Then template it out venue by venue; decommission local hotspots as each cuts over.

Phase 6 — PPPoE service Metro-side only (the BNG/RADIUS/edge is the separate WAN business — see wan/WAN-DESIGN.md §10 for that migration). Metro delivers the L2 circuits: first rooftop netPower + PtMP access VLAN bridged into its vpls-pop<n> at the hub router, dual hand-off at the colo to the WAN borders; RFC 4638 MTU (≥1520 payload) validated end-to-end through the wireless CPE; then per-rooftop rollout.

Phase 7 — Server room re-org (10.128.0.0/16) Build the VLAN/zone plan (DESIGN §7) on the CRS MLAG pairs + OPNsense zones in the new 10.128.0.0/16 space; migrate services into tiers, renumbering out of 10.1.x (run both ranges routed during the move, update DHCP-served DNS/option values fleet-wide, then withdraw the 10.1.x service routes — 10.1 reverts to being purely a venue block); enforce default-deny east-west last (log-only first, then drop).

Phase 8 — Ring closure + edge BGP Seel↔Holmes 60 GHz; Old Bank↔Fenwick fibre; ISP eBGP on OPNsense replacing static/VIP arrangement; Starlink runbook test (§4.4).


Appendix A — Per-device quick reference

Device class OSPF BGP LDP VPLS Firewall
Colo CCR (RR) area 0, all links RR (ip+l2vpn) + eBGP edge/Starlink + default-originate yes guest hub side + ISP POP hub → hand-off to WAN borders (RT hub) input-only
Metro CCR area 0, all links client ×2 (ip+l2vpn) + eBGP per venue yes guest spoke per venue; ISP POP spoke where rooftop input-only
Roof RB5009 / metro hub area 0, all links client ×2 (ip+l2vpn) + eBGP per venue yes as metro CCR input-only
Venue router none eBGP CE (AS 64512+V), announces 10.V.0.0/16, receives default none none (guest VLAN tagged up the uplink) input + zone forward chains
OPNsense — eBGP 65510 ↔ both colo CCRs (+ ISP transit) — — zone firewall for everything

(The retail-ISP BNG/border routers are a separate business — wan/WAN-DESIGN.md, AS 204258. They are not in this table; the metro only supplies their POP circuits as L2 VPLS.)