In Part 1, the lending batch job crossed an ExpressRoute circuit to an unmigrated file share and the four hour job became eleven. This field note is the network decision underneath that symptom: forced tunnelling, hub inspection, and a design nobody wrote down as a trade off.
SELL! is a fictional company. Any resemblance to a real platform team’s war story is the point.
Setting the Scene
Week eighteen at SELL!. The Visio on the wall looks familiar: one hub vNet, forced tunnelling of internet traffic through an on premises appliance, ExpressRoute to Auckland, everything inspected, everything centralised. It carries a permanent tax in latency, cost, and engineering time.
Ross, who runs operations from the Petone floor, asked on cutover weekend why the cloud was slower than the server room. He is still asking.
Add another UDR. Widen the firewall SKU. Keep the inherited design.
- Scale the hub firewall because inspection is "non negotiable."
- Leave egress hairpinning through on premises for SaaS calls.
- Peer ad hoc to Australia East when a PaaS service is missing in NZ North.
- Defer documenting why any of this exists.
The Pilot: The Inherited Diagram Works Perfectly
On paper the Visio is familiar, and familiarity feels like a design. Hemi files a change to let payments reach a Sydney SaaS API without the office hairpin. The request sits for three weeks because the design is “simple” and simple things are not supposed to need exceptions.
For some NZ organisations that shape is genuinely correct. For SELL!, it is a datacentre replica wearing cloud clothes.
The Visio did not move. The SaaS call still goes via Petone.
- Hemi's UDR exception is in the queue.
- The hub SKU is widened instead.
- Ross's batch is still eleven hours.
- Nobody owns the latency budget.
The Traffic Nobody Mapped
The network team owned the drawing; nobody else was consulted. Eighteen weeks in, app teams discover the hub firewall inspects everything at a throughput cost, and the “simple” architecture cannot change without a programme.
Latency Day: Ross Asks About Petone
Ross is in the room with the eleven hour batch on the screen. He asks the same question he asked on cutover weekend: why is the cloud slower than the Petone server room? The circuit and the hairpin are the cause. The Visio has no owner for that answer.
Post-Cutover: The Cracks Widen
The hub is up. Then the first real SaaS call leaves NZ North, and the cracks widen.
Egress Issues
All internet egress hairpins on premises. A workload in Azure New Zealand North talking to a SaaS API in Sydney routes via the office firewall, adding latency, cost, and a dependency that contradicts any datacentre exit story.
Segmentation Issues
One giant shared pattern because “routing is simpler.” Hemi’s payments dev subnet can still route to lending prod. Segmentation is someone’s good intentions rather than architecture.
Region Issues
No design for the trans-Tasman reality. A significant share of NZ workloads consume services only available in Australia East. The design degenerates into ad hoc peering and a mess of UDR exceptions.
| Symptom | What to do instead |
|---|---|
| The design optimises for on premises hosting. Cloud native workloads pay a latency, cost, and complexity tax they do not need. | Design for where workloads will be, not where they are today. Separate hybrid connectivity from
intra cloud patterns. Write a network decision record, including why you chose vWAN or did not. References: Why Landing Zones Fail · CAF: Network topology and connectivity · WAF: Performance Efficiency |
The Lessons We Can Learn
Azure New Zealand North Changes the Conversation, Partially
In country residency and latency to NZ users improve. Be clear eyed: NZ North does not have full service parity with Australia East. Document which region hosts what. Do not assume every PaaS service exists locally.
If you peer or connect to Australia East, understand the latency, the egress costs, and whether that traffic routes through your hub or directly.
Virtual WAN earns its cost at NZ scale when you have more than one hub, you need managed trans-Tasman routing, or you cannot staff a custom hub and spoke. A single hub and a handful of spokes does not.Decide Egress on Purpose, Not by Inertia
Forced tunnelling to on premises is a datacentre-era default. If you are a public sector agency or a CPS 234 entity, or you sit under the equivalent NZISM profile, centralised inspection may be a deliberate, defensible requirement. Build it, document the cost, and prefer Azure Firewall in the hub over backhauling to an office. If you are a cloud native commercial business like SELL!, direct internet egress with a cloud native stack is usually faster, cheaper, and more resilient.
Hybrid Connectivity Sized for NZ Geography
ExpressRoute is justified for consistent latency, high volume hybrid workloads, and regulated environments. Site to site VPN is frequently good enough at SMB scale. Do not gold plate. If your primary site is Wellington, have an honest conversation about regional disaster scenarios: a single region Azure design with on premises in the same metro shares more risk than a diagram admits.
Unwind Forced Tunnelling Without a Big Bang
SELL! cannot flip the Visio in one change window. The unwind is phased:
- Stop assigning forced tunnelling to new spokes. New work egresses via Azure Firewall in the hub, or direct where the network decision record allows it
- Move SaaS paths off the office hairpin first. Highest latency, lowest political cost
- Clean UDRs spoke by spoke. Rightsize the hub firewall last, after inspection volume actually drops
| Symptom | What to do instead |
|---|---|
| "We chose this network because that is how we have always done it." No written trade off. No owner for the latency budget. | Write a network decision record: what traffic is inspected, by what, why, and what it costs.
Centralised and distributed designs are both valid. An undocumented one is not. References: CAF: Network segmentation · WAF: Reliability design patterns |
Operations: The Missing Network Decision Record
The Visio is not the record. A network decision record names the topology, the egress path, the region matrix, and the person who owns the latency budget. Until that exists, every change request is an argument about a drawing on the wall.
| What they did | Should have done |
|---|---|
| Copied the colo hub spoke pattern into Azure and called it a landing zone. | Mapped one real workload's data flows before locking topology. |
| Forced all egress through on premises by default. | Chose egress per requirement: cloud native stack for commercial SaaS paths, central inspection only where justified and documented. |
| Discovered Australia East dependencies through outages and ad hoc peering. | Published a region matrix for PaaS availability and DR before the first production cutover. |
| Scaled the firewall to paper over chatty app-to-database traffic across the circuit. | Co located what talks frequently and treated distance as a design input, not a surprise. |
The Moral
Network diagrams inherit easily. Decisions do not. SELL! drew the old world into the new one, then wondered why the batch job remembered Petone.
The test: draw the data flows for one real workload: user to app, app to database, app to SaaS, backup to DR. If any arrow surprises you, the design was inherited, not chosen.
A topology review, a network decision record, and a region matrix are a fractional deliverable. See Fractional Cloud Architecture and Advisory.
One Block · build from here