Skip to content

Firewall — rule policy (corporate edge)

The rule policy for the OPNsense corporate edge pair (fw-colo-1 / fw-colo-2). Build steps live in firewall-opnsense-setup; architecture in firewall-opnsense.

Policy model

Two different postures, applied deliberately:

Domain Posture
Server networks (10.128.0.0/16) default deny in — reachable only on named services from named sources
Everything else (venue↔venue, venue→internet, staff/office) default allow — with destination guardrails and egress hygiene

The result is a flat, low-friction user network wrapped around a tightly-scoped server estate. Users get Wi-Fi calling, VPNs, collaboration apps and general on-net access without per-service allowlisting; the server VLANs stay closed.

Rules are evaluated on the interface traffic ENTERS

A rule controlling "who may reach VM_DATA" written on the VM_DATA interface does nothing — that interface only sees traffic leaving the VLAN. Because this policy is overwhelmingly destination-oriented, it is implemented as floating rules (which match regardless of ingress interface), with per-interface rules used only for egress control.

Role policy lives at the venue router

roles/venue_ce/tasks/firewall.yml already gives each venue VLAN a zone chain (pos / trusted / mgmt / wan). The venue CE knows which VLAN is which; the edge does not. Keep role-aware policy there and destination guardrails here — do not duplicate.


1. Aliases

Networks

Alias Contents Notes
NET_SERVERS 10.128.0.0/16 all server VLANs
NET_VENUES 10.0.0.0/8 documentation only — not used by any rule. The CORE interface group is the venue source qualifier; see §3
NET_OOB 172.16.201.0/24 switch/router management plane
RFC1918 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16
NET_BACKUP 10.128.10.0/24 VLAN 910
NET_APPS 10.128.32.0/24 VLAN 932 — AD, DNS, apps
NET_K8S 10.128.33.0/24 VLAN 933
NET_MON 10.128.34.0/24 VLAN 934
NET_DATA 10.128.35.0/24 VLAN 935 — SQL/PII
NET_DEV 10.128.36.0/24 VLAN 936
NET_HYP 10.128.40.0/24 VLAN 940 — hypervisor mgmt

Hosts

Alias Contents
DNS_FILTER 10.128.32.3, 10.128.32.4
AD_SERVERS domain controllers
NTP_SERVERS internal time source
UNIFI_CONTROLLER UniFi controller
MON_SERVERS Zabbix / LibreNMS collectors
BACKUP_SERVER backup server(s) on VLAN 910
SQL_SERVERS database hosts on VLAN 935
DATA_CLIENTS the only hosts permitted to reach SQL_SERVERS — app servers, k8s nodes, BI hosts
WEB_SERVERS internal HTTP/HTTPS services
MGMT_TRUSTED trusted management range — jump hosts / admin workstations
MAIL_RELAY the only host permitted outbound on 25
GEO_BLOCK GeoIP / Spamhaus / abuse.ch / FireHOL feeds
DOH_PROVIDERS known DNS-over-HTTPS endpoints

Ports

Alias Contents
P_DNS 53 (TCP+UDP)
P_NTP 123 (UDP)
P_WEB 80, 443
P_AD 88, 135, 139, 389, 445, 464, 636, 3268, 3269
P_AD_RPC 49152:65535 — AD dynamic RPC, see note
P_UNIFI 8080, 8443, 8880, 8843, 6789, 3478/udp, 10001/udp
P_SQL your DB ports (1433 / 3306 / 5432)
P_BACKUP backup agent ports
P_MON 161/udp, 162/udp, 10050, 10051
P_MGMT 22, 3389, 443

AD dynamic RPC

Domain join, replication and some MMC operations use the ephemeral RPC range, not just the fixed ports. Either include P_AD_RPC in the AD allow rule, or pin the DCs to a narrow static RPC range and use that instead. Omitting it produces intermittent failures that look like DNS problems.


2. Floating rules

Direction in, quick enabled, evaluated top-down — first match wins. Passes come before blocks; the blocks are the backstop.

Scope these to the INTERNAL interfaces — never WAN

Set the Interface field on every floating rule to the internal set (CORE group, SERVER_VLANS group, the three non-grouped server VLANs 910 / 934 / 935, and OOB) — not "any".

Destination NAT happens before filtering in pf, so an inbound port forward reaches the filter carrying its translated internal destination (10.128.32.x). If these rules also applied to WAN, rule 15 would silently drop every published service — floating rules are evaluated before interface rules, so the WAN rule never gets a look. Keeping WAN out of the floating set means all inbound policy lives in one place (§3, WAN) and cannot be accidentally overridden.

# Action Source Destination Port Purpose
1 Pass MON_SERVERS any P_MON, ICMP monitoring reaches every network
2 Pass MGMT_TRUSTED any P_MGMT admin access everywhere
3 Pass DATA_CLIENTS NET_DATA P_SQL the only way into DATA
4 Block log any NET_DATA any everything else denied
5 Block log any NET_HYP any hypervisor mgmt closed
6 Block log any NET_OOB any management plane closed
7 Pass NET_SERVERS BACKUP_SERVER P_BACKUP servers push backups
8 Block log any NET_BACKUP any backups otherwise unreachable
9 Pass any AD_SERVERS P_AD, P_AD_RPC, P_DNS AD
10 Pass any DNS_FILTER P_DNS DNS
11 Pass any NTP_SERVERS P_NTP time sync
12 Pass any MON_SERVERS 10051 Zabbix active agents (agent → server)
13 Pass any UNIFI_CONTROLLER P_UNIFI UniFi adoption + firmware cache
14 Pass any WEB_SERVERS P_WEB internal web apps
15 Block log any NET_SERVERS any nothing else enters the server estate
16 Block NET_DEV NET_APPS, NET_DATA any dev must not reach prod
17 Block any !DNS_FILTER 53 forces DNS to the filters
18 Block any DOH_PROVIDERS any DoH bypass
19 Block !MAIL_RELAY any 25 outbound spam / reputation guard
20 Block any GEO_BLOCK any reputation feeds

Rule 15 is the load-bearing rule. Rules 9–14 name every sanctioned service into the server estate; 15 denies the rest. Anything not destined for NET_SERVERS falls through to the per-interface rules below and is allowed.

Rules 4, 5, 6 and 8 sit above 9–14 deliberately — DATA, HYP, OOB and BACKUP are not reachable even on the "common" services. Their only doors are rules 1, 2, 3 and 7.

Server-to-server is default-deny by construction

Rule 15 matches any source — including the other server VLANs. A flow from VM_K8S to VM_APPS is dropped unless it first matched one of the named services in 9–14. No extra rules are needed for east-west segmentation; it falls out of the destination-based model.

To tighten further, narrow the source on rules 9–14 (e.g. UniFi reachable only from CORE, not from every server VLAN) rather than adding new blocks.

This does NOT segment traffic within a VLAN

Two hosts on 10.128.33.0/24 talk switch-to-switch and never reach the firewall — no rule here can see or stop them. Inter-VLAN is filtered; intra-VLAN is not. If you need host-level isolation inside a VLAN that is a switch problem (private VLANs / port isolation), not a firewall one.

Rules 11 and 12 are easy to miss and fail confusingly

If NTP_SERVERS or MON_SERVERS sit inside 10.128.0.0/16, rule 15 blocks them without rule 11/12 above it. The symptoms are indirect: estate-wide clock drift (which then breaks Kerberos/AD auth), and Zabbix agents showing as unreachable while ICMP to them succeeds.


3. Per-interface rules (egress)

CORE — venue and office traffic

Use an interface GROUP, not two interface rulesets

Venue traffic arrives over two links — CORE_A (→ CR-COLO-1) and CORE_B (→ CR-COLO-2) — and ECMP means either may carry a given flow. Create Firewall → Groups → CORE containing both and write the rules once. Duplicating them per-interface guarantees drift, and the failure mode is silent: everything works until a path flips and half the rules aren't there.

Everything arriving here is venue traffic — the interface is the source qualifier, so no venue subnet alias is needed. The floating rules have already protected the server estate before these are evaluated, so this is minimal:

# Action Destination Port
1 Pass any any
2 Block log any any

Rule processing order

OPNsense evaluates floating → interface group → interface. Because floating rules 4, 5, 6, 8 and 15 block every server destination first, a broad any/any here cannot grant server access. Express "everything except X" through ordering — specific denies above broad allows — never through alias subtraction, which OPNsense does not support and which rots when maintained by hand.

That gives venue↔venue, venue→internet and the permitted server services, with Wi-Fi calling, VPN clients and collaboration apps all working.

The venue router currently blocks this

The venue CE trusted chain permits only UniFi + AD, and its wan chain blocks 10/8. Venue→colo traffic is dropped at the venue router, before it reaches the edge. trusted needs a rule permitting the colo supernet ahead of its wan jump, or none of the above takes effect.

WAN — inbound

The floating rules deliberately do not apply here (§2), so this is the sole inbound policy. Default is deny; everything permitted is listed explicitly.

# Action Source Destination Port
1 Pass WAN_BGP_PEERS This Firewall 179
2 Pass WAN_BGP_PEERS This Firewall 3784, 3785 (BFD, if enabled)
3 Pass any published service hosts their ports only
4 Pass MGMT_TRUSTED This Firewall P_MGMT — only if remote admin is required
5 Pass any This Firewall ICMP echo — optional, rate-limited
6 Block log any any any

New alias: WAN_BGP_PEERS = 185.109.43.20/31, 185.109.43.24/31 (fw-colo-1) / 185.109.43.22/31, 185.109.43.26/31 (fw-colo-2).

Let OPNsense keep the NAT and the rule in sync

Create port forwards with "Add associated filter rule" rather than writing the NAT and the WAN rule separately. The rule then follows the NAT target automatically — hand-written pairs drift the moment a service moves host, and the symptom is a port that stops working with no rule having been touched.

Rule 3 targets the INTERNAL address

Because NAT translates first, the destination in rule 3 is the server's internal IP (10.128.32.x), not the public /28 VIP. Writing the public address there matches nothing.

A group cannot be 'overridden' by a member interface

Processing order is floating → interface group → interface, and OPNsense emits group and interface rules as quick. So:

  • a terminal Block log in the group makes every member's own interface rules dead code — they are never evaluated;
  • even a group pass pre-empts a member's interface rule, so a VLAN cannot be made more restrictive than its group.

A group is therefore only valid for interfaces whose policy is identical. Anything needing different egress must be excluded from the group and given its own interface rules. Membership is the lever — not rule ordering.

SERVER_VLANS interface group — the VLANs that share one policy

Create Firewall → Groups → SERVER_VLANS containing 932, 933, 936, 940 only, and write the shared egress baseline once:

# Action Destination Port
1 Pass DNS_FILTER P_DNS
2 Pass NTP_SERVERS P_NTP
3 Pass any P_WEB
4 Block log any any

940 HYP_MGMT fits here because "DNS, NTP, HTTP/S out and nothing else" is the baseline; its inbound restriction comes from floating rule 5, not from this group. 936 VM_DEV fits because its prod isolation comes from floating rule 16.

Interfaces with their own rules — NOT group members

These three have materially different egress and must stay out of the group:

910 BACKUP — only the backup server routes at all:

# Action Source Destination Port
1 Pass BACKUP_SERVER DNS_FILTER P_DNS
2 Pass BACKUP_SERVER NTP_SERVERS P_NTP
3 Pass BACKUP_SERVER any P_WEB
4 Block log any any any

The rest of the VLAN is L2-only and never routes. Inbound is governed by floating rules 2 and 7.

934 VM_MON — collectors must reach everything:

# Action Destination Port
1 Pass any any
2 Block log any any

935 VM_DATA — no general internet:

# Action Destination Port
1 Pass DNS_FILTER P_DNS
2 Pass NTP_SERVERS P_NTP
3 Pass REPO_SERVERS P_WEB
4 Block log any any

New alias: REPO_SERVERS = the patch/update endpoints DATA is allowed to reach.

Adding a VLAN later

Only add it to SERVER_VLANS if its egress policy is exactly rules 1–4 above. If it needs one extra allow or one restriction, give it its own interface rules instead — the group's block will otherwise silently swallow everything you write on the interface, and the rule will look correct in the GUI while never being reached.


4. Firewall self / infrastructure

Easy to miss, and they break HA rather than user traffic:

Rule Interface Notes
Allow proto 112 (CARP) every VLAN with a VIP without it both nodes go MASTER
Allow pfsync PFSYNC peer-scoped
Allow TCP 179 (BGP) WAN (bdr /31s), CORE (CCR /31s)
Allow no NAT for the firewall's own /31s outbound NAT, above the pool rule see below

Outbound NAT must exclude the firewall's own transit /31s

If the outbound NAT rule has a broad source, firewall-originated traffic is SNAT'd into the public /28 pool. The border has no route back to the pool until it is accepted and installed, so BGP never establishes and the box cannot ping its own peer — while inbound still works, because replies ride existing state. Add a Do not NAT rule for 185.109.43.20/31 and 185.109.43.24/31 above the pool rule, and scope the pool rule's source to the internal networks only.


5. Applying this to the live pair

Anti-lockout — do this before anything else

The GUI/SSH listen on OOB, and floating rule 6 blocks any → NET_OOB — including the firewall's own OOB address. Your admin session survives only via floating rule 2 (MGMT_TRUSTED → any : P_MGMT).

Before enabling rule 6: populate MGMT_TRUSTED with your actual admin source address and confirm rule 2 is above rule 6. Leave OPNsense's built-in anti-lockout rule enabled until the whole set is proven. Have IPMI/BMC console access to hand — if you get this wrong the GUI is gone and only the console can recover it.

5.1 Reconciling what is already built

If you built from the earlier iterations of this design, these changed:

Built earlier Change to
SERVER_VLANS group with all 7 VLANs 4 members — 932, 933, 936, 940. Remove 910, 934, 935 from the group and give each its own interface rules (§3)
Per-VLAN "overrides" written on member interfaces Delete them — they were never evaluated. Re-create as standalone interface rules on the three non-grouped VLANs
Floating rules with Interface = "any" Re-scope to the internal interfaces only — otherwise rule 15 drops every inbound port forward
Floating rules numbered 1–18 Now 1–20 — NTP (11) and Zabbix-active (12) inserted; the load-bearing block moved 13 → 15
Two separate CORE interface rulesets One CORE group (CORE_A + CORE_B)
Outbound NAT with a broad source Add the Do not NAT rule for the firewall's own /31s above the pool rule (§4)
No WAN inbound ruleset Add §3 WAN — including WAN_BGP_PEERS

5.2 Order of application

Sequenced so nothing is denied before its permit exists:

  1. Aliases first, fully populated. An empty alias matches nothing, so a permit against one permits nothing — and you will not get an error.
  2. Create the groups empty (CORE, SERVER_VLANS), then add members. A group with rules takes effect on its members the moment they join.
  3. Floating PASS rules only — 1, 2, 3, 7, 9–14. These only add permission, so they cannot break anything.
  4. WAN inbound rules 1–5 (the passes), leaving rule 6 off for now.
  5. Egress rules on the group and the three standalone VLANs, with every terminal Block log set to pass + log.
  6. Floating BLOCK rules 4, 5, 6, 8, 15–20, also as pass + log.
  7. Observe for a week. Review Firewall → Log Files; every hit on a log-only rule is something that will break when you flip it.
  8. Flip to deny, one at a time, in this order: NET_HYP (5) → NET_DATA (4) → NET_BACKUP (8) → NET_OOB (6) → 15 → the terminal interface blocks → WAN rule 6 → CORE last.

Rule 15 is the one to be most careful with — it is the widest-reaching, and it is the rule that breaks published services if the floating scope is wrong. Re-test an inbound service immediately after flipping it.

5.3 Rollback

Every step above is a single rule toggle. If a flip breaks something, set that one rule back to pass + log — do not disable the whole ruleset, or you lose the logs that identify the cause. Note the rule number, review the logs, fix the missing permit, flip again.

6. Verification

Firewall → Rules → Floating       -> order matches §2; "quick" set on all
Firewall → Aliases                -> no empty aliases
# from a venue client:
  - reaches AD, DNS, UniFi, internal web              (floating 9-14)
  - CANNOT reach NET_DATA / NET_HYP / NET_OOB / 910   (floating 4,5,6,8)
  - reaches the internet, Wi-Fi calling works         (CORE rule 1)
# from a DATA_CLIENTS host: reaches SQL_SERVERS on P_SQL   (floating 3)
# from any other host:      SQL denied and LOGGED          (floating 4)
# from MON_SERVERS:         SNMP/ICMP/Zabbix to every VLAN (floating 1)
# from the INTERNET:
  - every published service still reachable            (WAN 3 — regression test
    this after any floating-rule change; it is the first thing rule 15 breaks)
  - nothing else answers                               (WAN 6)

Mental model — pf is not RouterOS

There is no forward chain. Every rule is evaluated on the interface the packet enters, and NAT (both directions) is applied before filtering. Two consequences that catch people moving from RouterOS:

  • "Who may reach X" rules must be written where the traffic enters (or as a floating rule), never on X's own interface.
  • Inbound-NAT rules match the translated internal address; outbound-NAT source exclusions must exist or the firewall's own traffic gets rewritten (see §4).