Pub Invest Group — Configuration Approach, Protocols & Traffic Engineering¶
Version: 1.0 — 2026-07-05
Companion to DESIGN.md. All examples are RouterOS 7.19+ syntax and follow the conventions already visible in your live configs (AS 65500, 10.255.255.x loopbacks, 10.254.Y.0/24 p2p ranges, jump-chain venue firewalls).
Complete worked configs (all the fragments below assembled into copy-pasteable files): -
example-configs/target-colo-rr.rsc— a full target-state colo core / route reflector (modelled on CR-COLO-01): inter-core bond, RR BGP, eBGP to OPNsense + Starlink, community filter chains, default/server-room origination, guest VPLS hub. -example-configs/target-metro-pe.rsc— a full target-state metro PE (modelled on CR-SEELSTREET-01): OSPF, LDP, iBGP-to-RRs, per-venue eBGP hand-offs, the BGP community scheme applied, guest VPLS spoke, input firewall.
1. Configuration Philosophy¶
- One protocol per job. OSPF = infrastructure reachability. BGP = services (venue /16s, default, server room, public pools). LDP = labels. Static routes = banned in the core (the only statics anywhere are each router's blackhole anchor for its own aggregate).
- Never redistribute.
redistribute=connected,staticis how the current network grew its static-route sprawl and why the BGP table is unpredictable. Everything is originated explicitly withnetworkstatements / output filters. If a prefix is in the BGP table, a human put it there on purpose. - Templates, not snowflakes. Three device classes (core/metro CCR, roof RB5009, venue router), each a single template with a small variable set (site name, loopback, uplinks, venue number). Config generation from the template beats hand-editing — even if "generation" starts as a shell script with
sed. - Address-lists are the policy API. Firewall rules reference lists (
management,routers,ad-servers,pos-outbound-permitted…), never raw IPs. Changing policy = changing a list. Keep list contents identical fleet-wide and distribute changes as one-line scripts. - Config is code. Nightly
/exportfrom every device into a git repo (Oxidized or a cron + scp script). Every change gets a diff. RoMON + OOB path means you can always reach a router you just broke. - Fail toward the backup, never toward open. Firewall default-drop everywhere; routing failover is additive (backup path pre-established, just higher cost) so failover never requires config to change.
2. Standard Protocol Configuration¶
2.1 OSPF (transport routers only — colo, metro, roof; venues run none)¶
/routing ospf instance
add name=core router-id=<LOOPBACK> version=2
/routing ospf area
add name=backbone area-id=0.0.0.0 instance=core
# Loopback — passive
/routing ospf interface-template
add area=backbone interfaces=loopback passive networks=<LOOPBACK>/32
# Each infra link — p2p, cost per link class (see §4.1)
add area=backbone interfaces=sfp-sfpplus1 type=ptp cost=10 networks=10.254.Y.N/30 ;# fibre
add area=backbone interfaces=vlan-60ghz type=ptp cost=500 networks=10.254.Y.M/30 ;# 60GHz
Notes:
- No redistribute on the instance. Remove redistribute=connected from the current configs.
- Bind templates to interfaces, not bare networks — prevents accidental adjacencies on venue-facing ports.
- Auth: enable OSPF MD5/SHA auth on all adjacencies (shared per-link keys, rotated annually). Cheap insurance on roof links.
2.2 BFD (all 60 GHz links, plus fibre where you want it)¶
Reference it from the OSPF interface-template (use-bfd=yes) and BGP connections. 600 ms detection instead of 40 s dead-interval. On fibre it's optional (loss-of-light already drops the interface) but harmless.
2.3 iBGP with route reflectors¶
On CR-COLO-01 / CR-COLO-02 (the RRs):
/routing bgp template
add name=rr-clients as=65500 router-id=<LOOPBACK> \
local.address=<LOOPBACK> .role=ibgp-rr \
output.default-originate=if-installed \
routing-table=main
/routing bgp connection
add name=<CLIENT-NAME> template=rr-clients remote.address=<CLIENT-LOOPBACK>/32 .as=65500 listen=yes connect=yes
On every client (metro + roof transport routers only) — exactly two sessions, with both address families:
/routing bgp template
add name=core as=65500 router-id=<LOOPBACK> local.address=<LOOPBACK> .role=ibgp \
address-families=ip,l2vpn output.network=bgp-networks routing-table=main
/routing bgp connection
add name=RR1 template=core remote.address=10.255.255.255/32 .as=65500
add name=RR2 template=core remote.address=10.255.255.254/32 .as=65500
(address-families=ip,l2vpn on both RR and client templates — the l2vpn family is what signals VPLS, §2.6.)
- RR redundancy: both RRs with cluster-id = own loopback (default when role=ibgp-rr); clients take best of two reflected copies.
- Default route:
output.default-originate=if-installedon the RRs — they only advertise 0/0 while they themselves hold one (learned via eBGP from OPNsense). Edge dies → default disappears everywhere automatically.
2.3b Venue eBGP (CE model — replaces OSPF/iBGP at venues)¶
Venue side (AS = 64512+V):
/ip route add dst-address=10.V.0.0/16 blackhole comment="BGP aggregate anchor"
/routing bgp connection
add name=uplink-primary as=6XXXX router-id=10.V.255.255 local.role=ebgp \
remote.address=10.254.Y.N .as=65500 output.network=bgp-networks \
input.filter=default-only use-bfd=yes
# dual-homed venues: second session on the backup /30, same everything
/ip firewall address-list add list=bgp-networks address=10.V.0.0/16
/routing filter rule add chain=default-only rule="if (dst == 0.0.0.0/0) { accept } reject"
Metro PE side, one connection per venue:
/routing bgp connection
add name=venue-<NAME> as=65500 local.role=ebgp remote.address=10.254.Y.M .as=6XXXX \
input.filter=from-venue-<V> output.default-originate=always use-bfd=yes
/routing filter rule add chain=from-venue-<V> \
rule="if (dst == 10.V.0.0/16) { set bgp-communities 65500:100; accept } reject"
The input filter accepts only that venue's /16 — a compromised or misconfigured venue router cannot inject anything else. Dual-homed: the backup hub's PE adds set bgp-local-pref 50 in its from-venue chain so the core returns traffic via fibre while it lives; the venue prefers its fibre-learned default with local-pref on the primary session. The blackhole anchor keeps the aggregate announced and kills traffic to unused /16 space at the venue.
2.4 eBGP at the edge¶
OPNsense (FRR) AS 65510 ↔ both colo CCRs; Starlink RB4011 AS 65502 ↔ both colo CCRs. Filters on the CCRs (the contract):
/routing filter rule
# From OPNsense: accept default + public + server-room zones only
add chain=from-edge rule="if (dst == 0.0.0.0/0) { set bgp-local-pref 200; accept }"
add chain=from-edge rule="if (dst in 10.128.0.0/16 && dst-len <= 24) { accept }"
add chain=from-edge rule="reject"
# From Starlink: accept ONLY payment-provider prefixes, never default
add chain=from-starlink rule="if (dst in payment-prefixes) { set bgp-local-pref 50; accept }"
add chain=from-starlink rule="reject"
(payment-prefixes = the address-list of card-processor destinations. Local-pref 50 < 200 ⇒ Starlink path is dormant while the main WAN advertises. See §4.4.)
2.5 MPLS / LDP¶
/mpls settings set dynamic-label-range=16-1048575
/mpls ldp instance add name=ldp lsr-id=<LOOPBACK> transport-addresses=<LOOPBACK> afi=ip
/mpls ldp interface add interface=sfp-sfpplus1 accept-dynamic-neighbors=yes
/mpls ldp accept-filter for 10.255.255.0/24) — keeps LFIBs tiny.
2.6 VPLS — BGP-signalled (guest example; PPPoE identical with different names)¶
Signalling rides the existing iBGP/RR sessions (address-families=ip,l2vpn, §2.3) — no per-peer pseudowire config, and the RRs give auto-discovery: adding a spoke touches only that spoke's PE.
Spoke — the venue's metro PE (venue guest VLAN 20 arrives tagged on the venue-facing port):
/interface bridge add name=br-guest-<V>
/interface vlan add name=guest-<V> interface=<venue-port> vlan-id=20
/interface bridge port add bridge=br-guest-<V> interface=guest-<V>
/routing bgp vpls
add name=guest-<V> bridge=br-guest-<V> site-id=<V> \
route-distinguisher=65500:20<V> \
export-route-targets=65500:2001 import-route-targets=65500:2000 \
pw-type=vpls
Hub — both colo CCRs (RT direction reversed; per-venue VLAN toward the guest firewall trunk):
/routing bgp vpls
add name=guest-hub bridge=br-guest-hub site-id=1000 \
route-distinguisher=65500:2000 \
export-route-targets=65500:2000 import-route-targets=65500:2001
- Hub-and-spoke by route-target: spokes export
:2001, import:2000; hubs the reverse. Spokes never import each other's routes, so venues structurally cannot bridge to each other — no reliance onhorizonbeing set correctly on every port. - Multihoming/redundancy: two PEs advertising the same site-id (dual-homed venue's guest VLAN on both hub PEs; each ISP POP circuit's dual colo hand-off) triggers BGP designated-forwarder election — the losing PE blocks its attachment circuit. Loop-free, automatic failback, no scripts, no spanning tree. This replaces the cold-standby-pseudowire hack that LDP signalling would force.
- ISP POP circuits (for the separate WAN business, wan/WAN-DESIGN.md): same VPLS pattern with its own RT pair (e.g.
:3000/:3001), spoke PEs at the rooftop hubs, hub PEs on the colo CCRs bridging each to a tagged hand-off port toward the WAN border routers — not a metro BNG. The metro is only the L2 carrier. - Verify on your RouterOS version (7.19+) before rollout:
/routing bgp vplsinstances up,/interface vpls printshows dynamic pseudowires, and MTU end-to-end (§6).
2.7 Device baseline (every router)¶
/ip service set telnet,ftp,www,api disabled=yes
/ip service set ssh port=22; set www-ssl certificate=<cert> disabled=no
/user aaa set use-radius=yes ;# RADIUS login → FreeRADIUS (10.128.36.11 post-migration; 10.1.88.11 today)
/snmp set enabled=yes contact="Pub Invest Group" location="<SITE>" ;# move to SNMPv3 communities
/system ntp client set enabled=yes
/system ntp client servers add address=time.cloudflare.com
/system clock set time-zone-name=Europe/London
/tool romon set enabled=yes secrets=<romon-secret>
/system logging action add name=remote target=remote remote=10.128.36.20 ;# central syslog (Tier-2 zone)
/system logging add topics=critical,error,warning action=remote
/ip dns set servers=10.128.32.3,10.128.32.4 ;# internal resolvers, Tier-1 zone (10.1.84.3 etc. until Phase 7)
2.7b Naming scheme — role-site[-n], lowercase, unpadded¶
Discovery proved the old mixed scheme causes real errors (a roof transport router named RTR-CHARLOTTEST-001 misled planning; CR-LEVEL-01 vs -001 drift; LEVEL_IRISH_MASTE truncation). Target scheme:
- Lowercase everywhere — hostname = DNS = NetBox = Zabbix = Oxidized, identical; aligns with the WAN business (
bdr-1, WAN-DESIGN §11). - Role token = strict truth. The complete vocabulary — a device carries a token only if it genuinely performs that role:
| Token | Purpose | Examples | Notes |
|---|---|---|---|
cr- |
Transport router — runs OSPF/LDP, iBGP client or RR (core, metro, roof PE) | cr-colo-1/2, cr-coloroof-1, cr-mathewst-1, cr-charlotte-1, cr-charlotteroof-1, cr-seel-1, cr-holmes-1, cr-oldbankroof-1 |
Roof PEs get "roof" in the site token — distinct failure domain |
rtr- |
Venue CE router — eBGP CE, VLANs/DHCP/zone firewall, nothing else | rtr-dos-1, rtr-chr-1, rtr-mcc-1 |
Venue trigram as site token |
fw- |
Firewall (OPNsense) | fw-colo-1/2 (edge pair), fw-guest-1 |
|
sw- |
Switch (L2) | sw-storage-1/2, sw-vm-1/2, sw-oob-1, sw-rbl-1 (venue) |
Replaces legacy SWI- |
hv- |
Hypervisor (Proxmox) | hv-1/2/3 |
All-colo → no site token |
srv- |
Server named by service (VM or generic physical) | srv-sql-1/2, srv-unifi-1, srv-zbx-1, srv-dc-1/2, srv-pbs-1 |
Never sequential-anonymous (srv-core-37). Name the service; the app over the platform where single-purpose. Abbreviations from a documented list; ≤15 chars for Windows (NetBIOS) |
w60- / w5- |
PtP radio unit (60 GHz / 5 GHz) | w60-coloroof-oldbank-1 ↔ w60-oldbank-coloroof-1 |
<near>-<far> encodes location + pointing; no master/slave. Replaces ANT-/LEVEL_* |
ap- |
WiFi access point | ap-dos-1 |
Mirrors into UniFi device names |
gw- |
Small gateway appliance | gw-oob-1 (OOB WireGuard entry) |
|
con- |
Console server | con-colo-1 |
|
pdu- |
Power distribution | pdu-colo-1a (rack 1, feed A) |
|
| (k8s workloads) | Named by Kubernetes, not this scheme | — | This scheme names pets (VMs/hardware); cattle are k8s's problem |
bdr- / bng- / pop- / radius- / nms- |
WAN business (AS 204258) — border, subscriber edge, rooftop access, AAA, monitoring | bdr-1/2, bng-1, pop-oldbank-1 |
Own scheme + DNS zone (WAN-DESIGN §11); same conventions |
- Venues: keep the established trigrams (rtr-chr-1 = Cheers, rtr-ynk-1 = Yankees, rtr-mol-1 = Moloko) — they're organisational vocabulary, staff-fluent, short, and already the de-facto key. Three conditions make them safe: (1) the canonical trigram table is a first-class registry artifact (NetBox tenant slug = trigram, display name = full venue name; the table below in this repo until NetBox) — the failure mode isn't the codes, it's the mapping living only in heads; (2) collision rule: repeated brands take the site letter as the 3rd char or a 4th char (mcc = McCooleys Concert Sq, mcs = McCooleys Mathew St — as already practiced), assigned in the table, never ad-hoc; (3) the immutable venue key is venue_id (drives /16, AS, VPLS site-id) — a rebrand is a cheap planned trigram+DNS rename, never a renumber. |
|||
- Radios: w60-<near>-<far>[-n] (near = where the unit sits): w60-coloroof-oldbank-1 ↔ w60-oldbank-coloroof-1. No master/slave — the name encodes location + pointing; trailing index for parallel links. w5- for 5 GHz units. |
|||
Canonical trigram table (seed — move to NetBox; ? = confirm): |
| Tri | Venue | id | Tri | Venue | id | Tri | Venue | id |
|---|---|---|---|---|---|---|---|---|
| lvl | LEVEL | 1 | mol | Moloko | 4 | pck | Peacock | 5 |
| brk | Brooklyn Mixer | 6 | lag | Lago | 7 | cel | Celtic Corner | 17 |
| dos | Dirty O'Sheas | 23 | ynk | Yankees | 27 | rok | Rocking Horse | 33 |
| hat | Hatch | 34 | hej | Heebie Jeebies | 42 | kel | Kells of Eden | 43 |
| chr | Cheers | ? | soh | SOHO | ? | fus | Fusion | ? |
| ein | Einstein Bier Haus | ? | nfl | Nelly Foleys | ? | rbl | Ruby Blues | ? |
| smo | Smokies ? | ? | mcc | McCooleys Concert Sq | ? | mcs | McCooleys Mathew St | ? |
| bpl | Boston Pool Loft | ? | blr | Black Rabbit | ? | olb | Old Bank Bar | ? |
| cas | Castle St Apts | ? | lor | Lord St | ? | irh | Irish House | ? |
| tav | Temple Tavern ? | ? | wds | Wood Street | ? | zan | Zanzibar | ? |
- Rules: site token only where geography differs; unpadded
-n; no underscores; short (MikroTik identity truncates). - Enforced from the registry:
network.yaml/NetBox holds the name; Terraform sets/system identityfrom it; DNS/Zabbix/Oxidized derive from it; the scrape-to-NetBox diff alerts on identity drift. - Interface comments:
ROLE::far-end(UPLINK::cr-colo-2,VENUE::rtr-dos-1,METRO::,EDGE::,ROOF::,OOB::,SRV::,WAN::,INTERCORE::) — machine-parseable. - Migration: opportunistic — rename at planned touches (ROS7 upgrade, eBGP conversion, move windows), paired with the loopback renumber where due. Never big-bang.
(The WAN business keeps its own documented scheme — bdr-N, bng-N, pop-<hub>-N, WAN-DESIGN §11 — same conventions, its own DNS zone.)
3. Firewall Standards¶
3.1 Core/metro routers — input-only self-protection (your current pattern, canonicalised)¶
/ip firewall filter
add chain=input connection-state=established,related action=accept
add chain=input connection-state=invalid action=drop
add chain=input protocol=icmp action=accept
add chain=input protocol=tcp dst-port=179 src-address-list=routers action=accept
add chain=input protocol=ospf in-interface-list=neighbour-routers action=accept
add chain=input protocol=udp dst-port=646 in-interface-list=neighbour-routers action=accept ;# LDP
add chain=input protocol=tcp dst-port=646 src-address-list=routers action=accept
add chain=input protocol=udp dst-port=3784,3785 in-interface-list=neighbour-routers action=accept ;# BFD
add chain=input protocol=tcp dst-port=22,8291,443 src-address-list=management action=accept
add chain=input protocol=udp dst-port=161 src-address-list=snmp-servers action=accept
add chain=input action=drop
routers from the loopback range (10.255.255.0/24 as one entry — simpler than the current per-/32 list and can't drift).
3.2 Venue routers — zone jump-chain model (generalised from RTR-DOS-001)¶
Chains: wan (internet-only, blocks 10/8), pos (wan + AD + pos-outbound-permitted), trusted (wan + AD + UniFi + file/print lists), and per-interface jumps:
| In-interface | Jump target |
|---|---|
| VLAN 10 Office | trusted |
| VLAN 30 POS | pos |
| VLAN 40 Staff | wan |
| VLAN 99 Mgmt | trusted (+ UniFi inform) |
| VLAN 20 Guest | none — bridged into VPLS, never routed |
Final rule: chain=forward action=drop. Inbound-to-venue: only trusted-inbound list (mgmt VPN, server-room mgmt ranges) may initiate into LAN interfaces. Keep the existing established/related/invalid triplet at the top of both chains. This is exactly your RB4011 policy — the work is making all ~30 venues byte-identical apart from address variables.
3.3 OPNsense zones¶
Zones: WAN, CORE (transit to CCRs), SRV-T1, SRV-T2, DMZ, OOB, BACKUP, GUEST, MGMT-VPN. Default deny inter-zone; explicit rules per flow (documented in a flow matrix — start one in this repo). All venue→server-room flows are enforced here again (defence in depth vs venue router compromise).
4. Traffic Engineering¶
4.1 The cost scheme (this is 90% of your TE)¶
The scheme is reference-bandwidth 100G: wired cost = 100G ÷ link speed (bonds use aggregate). Wireless values are policy, not formula.
| Link class | OSPF cost | Result |
|---|---|---|
| 100G | 1 | Floor |
| 40G | 2 | (2.5 truncated, as auto-cost routers do) |
| 25G (or 2×25G bond = 50G) | 4 (bond: 2) | Inter-core bond = 2 |
| 10G fibre | 10 | Preferred metro standard |
| 1G fibre / CAT6 metro | 100 | Beats wireless, loses to any 10G path ≤ 9 hops |
| 60 GHz | 500 | Policy value — any all-wired path beats one wireless hop. Never recompute from bandwidth |
| Last-resort (e.g. temporary LTE) | 5000 | Absolute fallback |
ECMP requires equal costs — parallel same-class links balance (Holmes 2×10G); mixed-speed pairs (25G+10G) won't, making the slower link pure backup, which is usually correct.
Rules of thumb: any all-fibre detour (even 4 hops = cost 40) beats one 60 GHz hop (500). Two 60 GHz hops (1000) still beat a dead path. Never hand-tune a single link cost to fix a traffic problem — change the class or add a link; per-link exceptions rot.
ECMP falls out for free: Holmes' two diverse 10G fibres (cost 10 each, one per colo CCR) load-balance per-connection automatically; Old Bank's 2× 60 GHz likewise until its fibre lands (then fibre cost 10 wins outright).
4.2 Verifying and simulating paths¶
/routing route print where dst-address=10.V.0.0/16and/tool traceroutefrom loopback to loopback are the ground truth.- Before/after any cost change, snapshot
ip route printon the affected routers and diff. - Planned-work drain: to empty a link before maintenance, set its OSPF cost to 4000 on both ends, wait for reconvergence (seconds), then work. Restore after. This is the only sanctioned manual TE knob.
4.3 Do you need RSVP-TE? (Mostly no)¶
RouterOS supports RSVP-TE tunnels, but with metric-classed links, ECMP, and BFD you get: deterministic primary/backup, fast reroute-ish behaviour, and drain capability — without per-LSP state on CCR2004s. Adopt RSVP-TE only if a concrete need appears (e.g. "guest VPLS must never exceed 2G on the Mathew St fibre while POS gets priority" can't be met with QoS alone). Revisit then; skip now. QoS (§7) is the right tool for contention, not TE tunnels.
4.4 Starlink payments failover (runbook)¶
Steady state: RB4011 (AS 65502) announces nothing/low-pref; card traffic follows default via OPNsense.
- Populate
payment-prefixesaddress-list (processor endpoints — e.g. the34.117.150.224-style IPs already inpos-outbound-permitted) and keep it in git. - RB4011 statically routes those prefixes into the Starlink WAN with NAT, and announces them via eBGP with
bgp-local-pref 50(per §2.4 filter). - Main WAN healthy → OPNsense's default (local-pref 200) wins for everything.
- Main WAN dead → OPNsense withdraws default (its own eBGP upstream drops, or use FRR conditional advertisement) → RRs stop originating 0/0 → for payment prefixes the only remaining route is Starlink's specifics → POS keeps trading; bulk traffic (guest, staff, backups) correctly blackholes rather than crushing the Starlink uplink.
- Test quarterly: pull the ISP hand-off in a maintenance window, confirm a test transaction completes over Starlink, confirm guest WiFi is down (that's success, not failure).
4.5 BGP community scheme (tag at origin, act at the edge)¶
| Community | Meaning | Applied where |
|---|---|---|
| 65500:100 | Venue aggregate | Venue routers on their /16 |
| 65500:200 | Server-room subnet | Colo CCRs / OPNsense |
| 65500:300 | Corporate public routed block (Pub Invest /28 on OPNsense) | edge in-filter |
| 65500:911 | Eligible for Starlink failover | Anything that must survive WAN loss |
| 65500:666 | Blackhole (DDoS/incident) | Manual; edge drops matching |
Communities cost nothing now and make future policy ("only 65500:911 routes may resolve via Starlink") one filter rule instead of a re-architecture.
5. Zero-Trust Hardening Checklist¶
- [ ] All router mgmt bound to loopbacks;
managementlist = mgmt VPN (10.201.201.0/24) + jump hosts only - [ ] RADIUS login on every RouterOS device; local
adminrenamed, strong password, emergency-only - [ ] OSPF auth on all adjacencies; BGP sessions protected by input filter (§3.1); consider TCP-MD5 on eBGP
- [ ] SNMPv3 (or at minimum per-device communities restricted to
snmp-serverslist) - [ ] RoMON secrets set; OOB VLAN 900 reachable only via mgmt VPN
- [ ] Venue: no inter-VLAN rules exist; forward default-drop verified with a scan from each VLAN
- [ ] Server room: storage VLANs have no gateway; confirmed by scan
- [ ] Guest/PPPoE: confirm no route to/from 10/8 (traceroute + firewall counters); guest sources are 172.16/12 only — alert on any 172.16/12 source appearing inside the corporate routing domain (it means a bridge/VPLS misconfig)
- [ ] Nightly config export to git; alert on diff without a change ticket
- [ ] Central syslog + login alerting; netflow (
/ip traffic-flow) on colo CCRs to the collector
6. MTU Plan (do this before enabling MPLS)¶
6.1 The overhead math — where the number comes from¶
The pseudowire carries the customer's whole Ethernet frame; MPLS adds labels + an optional control word on top. Worst case is PPPoE (RFC 4638) with a VLAN tag:
| Component | Bytes |
|---|---|
| Inner customer Ethernet header | 14 |
| Customer VLAN tag (sub 3xx / guest 200) | 4 |
| PPPoE + PPP headers | 8 |
| Customer IP (RFC 4638 clean 1500) | 1500 |
| inner frame carried by the PW | 1526 |
| MPLS: 2 labels (tunnel + VC) = 8 + control word 4 | 12 |
| L2 payload the infra link must carry (L2MTU) | 1538 |
Guest-only (no PPPoE) ≈ 1530. Round up for a possible 3rd label / 2nd tag → L2MTU floor = 1600.
6.2 What to set¶
| Segment | L2MTU | MPLS MTU | IP MTU | Why |
|---|---|---|---|---|
| Fibre infra links | 9000 (CCR2004 max ~9200) | — | 9000 | Go jumbo — removes MTU as a concern for VPLS/Q-in-Q/future forever |
| 60 GHz links | radio max, must be ≥ 1600 | — | match | The real constraint; verify per radio (WIRELESS-60G §2) — a sub-1600 hop can't carry full-MTU VPLS |
MPLS (mpls-mtu), every infra iface |
— | 1600 network-wide | — | Set to the 60 GHz floor, not fibre's 9000 — MPLS never fragments, so a jumbo labeled packet built on fibre would be dropped hitting a 60 GHz hop. Labeled packets are ≤1538, so 1600 is ample and uniform |
| Venue /30s & VLAN interfaces | (inherit) | — | 1500 | MPLS works at L2MTU; the routed IP MTU doesn't change |
| Storage VLANs (Ceph/iSCSI) | 9000 | — | 9000 | Jumbo end-to-end, but never traverse MPLS/VPLS (L2-local) |
6.3 Order & verify¶
Per link: raise L2MTU both ends → set mpls-mtu=1600 → then build the VPLS. A pseudowire across one forgotten low-MTU link silently blackholes oversized frames (works for ping, fails for real traffic). Test each new path first:
ping <far-underlay> size=1580 do-not-fragment (1580 > the 1538 requirement) — then a real 1500-DF ping through the VPLS from a customer.
7. QoS — protect the thin links¶
Fibre at 10G is uncongested; queue only on 60 GHz egress (and Starlink). Simple 4-class model, classified by DSCP marked at the venue router mangle:
| Class | DSCP | Traffic | 60 GHz share |
|---|---|---|---|
| NC | CS6 | OSPF/BGP/BFD/LDP | priority 1 (small, guaranteed) |
| CRITICAL | AF41 | POS subnets, RADIUS, AD auth | guaranteed 30% |
| BUSINESS | AF21 | Office, mgmt, UniFi, backups (rate-capped in hours) | guaranteed 40% |
| BEST-EFFORT | 0 | Staff + guest WiFi | remainder, first to starve |
Implementation: /queue tree on each 60 GHz interface with max-limit set to ~90% of real measured radio throughput (measure with btest, radios degrade in rain — set against a rainy-day number). During a hub-fibre failure the venue's whole load shifts to 60 GHz; this policy is what keeps tills working while guest WiFi degrades — which is the correct order.
8. Monitoring & Operations¶
- LibreNMS (or current SNMP pollers — today 10.1.84.16/18, moving to Tier-2 10.128.36.x in Phase 7): all devices, alert on OSPF neighbour count change, BGP session down, 60 GHz RSSI/MCS drop, interface errors.
- Netflow from colo CCRs → collector, for capacity planning and the "what filled the Mathew St fibre" question.
- Oxidized for config backup/diff (RouterOS supported natively).
- Smokeping/graph latency loopback-to-loopback per hub — 60 GHz path degradation shows in latency before it shows in loss.
- NetBox as source of truth: sites, links, /30 assignments, venue /16 registry, VLANs, loopback registry. The addressing tables in DESIGN.md §3 seed it.
9. Migration Plan (from current state)¶
Each phase is independently valuable and reversible; nothing breaks the running network until its cutover step.
The phases group into two stages:
- Stage 1 — the routed IP network (Phases 0–3 + 8): ROS7 everywhere, clean OSPF core, RR-based BGP, eBGP venues, new fibre core carrying all traffic, ROS6 hardware killed off. Complete and stable before any overlay work.
- Stage 2 — L2 overlays (Phases 4–7): LDP, guest VPLS, ISP POP circuits (L2 carriage for the separate WAN business), server-room re-org. Purely additive — each venue/rooftop comes on independently with no change to Stage 1 routing.
Phase 0 — Foundations (no traffic impact)
NetBox + git config backups; baseline configs (§2.7) everywhere; central syslog/monitoring; document current static routes per router (the kill list); run the discovery collection (discovery/collect.sh) for the as-built map.
Phase 0b — ROS7 gate Every router that stays in the target design runs ROS 7.19+ before its links move to the new core. Legacy ROS6 hardware (e.g. the CCR1016 at LEVEL) is not upgraded in place — the new ROS7 core is built alongside, links re-home onto it (backups first), and the ROS6 devices are decommissioned once nothing terminates on them. No overlay feature (LDP/VPLS/BGP-VPLS) is configured anywhere until this gate is fully passed.
Phase 1 — MTU + BFD Raise L2 MTU on all infra links (both ends, one link at a time, off-hours); enable BFD on 60 GHz adjacencies. Verify with DF pings.
Phase 2 — Clean IGP (transport routers only)
Add OSPF interface-templates on colo/metro/roof infra links (p2p type, class costs per §4.1); remove redistribute=connected from OSPF instances; remove OSPF from venue links as each venue converts to eBGP in Phase 3; confirm LSDB contains only transport loopbacks + inter-hub /30s.
Phase 3 — eBGP venues + RR core, alongside the old mesh
Configure RR templates (ip+l2vpn) on colo CCRs; metro/roof routers become clients. Convert venues one at a time to the CE model (§2.3b): eBGP session(s) up, /16 announced, default received, then remove that venue's OSPF/iBGP/static remnants. Old mesh sessions on transport routers stay up for comparison (/routing route print count-only where bgp); when identical-or-better, remove mesh sessions and redistribute=connected,static, then delete the static-route kill list one router at a time, checking reachability after each.
Phase 4 — MPLS/LDP Enable LDP on transport-router infra links only (harmless alongside — labels unused until a VPLS exists). Verify LDP neighbours = OSPF neighbours on every transport router.
Phase 5 — Guest VPLS pilot One venue (pick a dual-homed fibre venue, e.g. SOHO): tag VLAN 20 up its uplink, build the BGP-signalled spoke on its metro PE + hub on the colo CCRs, hand off to the guest firewall, run two weeks. Confirm DF election behaves by failing the primary uplink. Then template it out venue by venue; decommission local hotspots as each cuts over.
Phase 6 — PPPoE service
Metro-side only (the BNG/RADIUS/edge is the separate WAN business — see wan/WAN-DESIGN.md §10 for that migration). Metro delivers the L2 circuits: first rooftop netPower + PtMP access VLAN bridged into its vpls-pop<n> at the hub router, dual hand-off at the colo to the WAN borders; RFC 4638 MTU (≥1520 payload) validated end-to-end through the wireless CPE; then per-rooftop rollout.
Phase 7 — Server room re-org (10.128.0.0/16)
Build the VLAN/zone plan (DESIGN §7) on the CRS MLAG pairs + OPNsense zones in the new 10.128.0.0/16 space; migrate services into tiers, renumbering out of 10.1.x (run both ranges routed during the move, update DHCP-served DNS/option values fleet-wide, then withdraw the 10.1.x service routes — 10.1 reverts to being purely a venue block); enforce default-deny east-west last (log-only first, then drop).
Phase 8 — Ring closure + edge BGP Seel↔Holmes 60 GHz; Old Bank↔Fenwick fibre; ISP eBGP on OPNsense replacing static/VIP arrangement; Starlink runbook test (§4.4).
Appendix A — Per-device quick reference¶
| Device class | OSPF | BGP | LDP | VPLS | Firewall |
|---|---|---|---|---|---|
| Colo CCR (RR) | area 0, all links | RR (ip+l2vpn) + eBGP edge/Starlink + default-originate | yes | guest hub side + ISP POP hub → hand-off to WAN borders (RT hub) | input-only |
| Metro CCR | area 0, all links | client ×2 (ip+l2vpn) + eBGP per venue | yes | guest spoke per venue; ISP POP spoke where rooftop | input-only |
| Roof RB5009 / metro hub | area 0, all links | client ×2 (ip+l2vpn) + eBGP per venue | yes | as metro CCR | input-only |
| Venue router | none | eBGP CE (AS 64512+V), announces 10.V.0.0/16, receives default | none | none (guest VLAN tagged up the uplink) | input + zone forward chains |
| OPNsense | — | eBGP 65510 ↔ both colo CCRs (+ ISP transit) | — | — | zone firewall for everything |
(The retail-ISP BNG/border routers are a separate business — wan/WAN-DESIGN.md, AS 204258. They are not in this table; the metro only supplies their POP circuits as L2 VPLS.)