Before You Deploy · Field Note 1 of 7

Landing Zones Without Business Context

Why lift-and-shifting a real workload onto a vanilla ALZ reference is where the missing requirements finally come due.

SR
Steve Rackham
16 min read Guides

Spurious Egregious Lending Limited (or SELL!) is a fictional company used to illustrate how things can go wrong when you rush to deploy a landing zone without understanding the business context.

Any resemblance to a real platform team’s war story is, unfortunately, the point.


Setting the Scene

Monday morning at SELL!, a Wellington fintech with eighty staff and a lending product that needs to scale. The investment deck has said “cloud-first transformation” seventeen times. Capital is raised. The board is ready.

“Board wants us in the cloud by third quarter. Make it happen.”

The CTO passes it down. The team meets once, lands on a plan, and puts it in the next sprint: deploy the reference landing zone, iterate later. Deadline first, backlog second.

Sprint plan: Ship the Reference Architecture

The reference architecture is deployed to accelerate delivery.

Week 1
  1. Reference articles are read.
  2. Code is coded and pipelines are built.
  3. Global defaults are accepted to accelerate deployment.
  4. The landing zone deploys ahead of schedule.

The Pilot: The Accelerator Works Perfectly

And without business context, it does. At the end of the sprint at week two, Spurious Lending has:

End of Sprint: Landing Zone Accelerator

The reference architecture is deployed ahead of schedule. The landing zone accelerator worked perfectly.

Week 2
  1. A Management Group hierarchy exactly like the diagrams in the reference articles.
  2. Default Azure Policy initiatives assigned at the root management group, including the denial of public IP addresses and non-approved VM SKUs.
  3. A hub-spoke Virtual Network with a central Azure Firewall protecting the network.
  4. A Log Analytics workspace for centralised logging and monitoring.
  5. A Dev/Test Subscription for the platform team to use for testing and development.

The architecture diagrams look great. The CTO shares them with the board and congratulates the team on a job well done.

The Workload Nobody Mapped

The Application

Three tier application hosting the main lending product. The application is hosted across three servers in a single datacentre, with a Fortinet firewall protecting the network.

Pre-migration
  1. The database tier is hosted on a Microsoft SQL Server 2016 virtual machine.
  2. The application tier is hosted on a Windows Server 2019 virtual machine.
  3. The web tier is hosted across two Linux web servers, with a load balancer in front.
  4. The application has an API integration layer that integrates with two credit bureaus and a payments provider.
  5. Backups are taken nightly to a local tape backup device.
  6. The application is accessed by a web browser and by a mobile app.

The mandate of “third quarter” gives no time to map the workload’s dependencies, and so, given the criticality of the application and discussion with the product team, the platform team decides to do a lift and shift. Migration will be done over a single weekend, everything at once.

The product team does not fully understand the landing zone’s architecture or policy assignments, and so the platform team does not fully understand the workload’s dependencies. The two halves of the company are about to meet for the first time.

A landing zone is an architectural expression of your organisation’s compliance and governance requirements. Microsoft make it easy to deploy one, but your governance, security, and compliance requirements are not encoded in the accelerator.

Nobody asked the only two questions that mattered: what data sovereignty or residency requirements apply, and what compliance requirements apply. The application handles credit information. The accelerator defaulted to Australia East. The FAQ below covers both, including where DR replicas should live. Those answers come due in Part 2.

Migration Day: The Weekend Cutover

On migration day, the product team’s deployment fails. Exemptions go out by hand, one policy at a time. Each failure is treated as a new problem rather than evidence of a missing intake process.

Migration Day: Cutover Issues

What stalled the weekend migration.

Migration day

The landing zone and the workload met for the first time. Neither side had been designed for the other.

  1. The Root Management Group policies deny public IPs, expect a tagging standard nobody wrote, and restrict VM sizes to a SKU list built for the platform team's own subscription. The workload violates all three.
  2. The hub addressing plan never met the 192.168.20.0/24 range the servers have used since 2017, hardcoded in connection strings, SQL Agent jobs, and a vendor appliance with no documentation.
  3. SQL moves on the second attempt, then remembers the old world: linked servers to the old domain controller, nightly jobs with hard coded hostnames, collation and compatibility surprises the lift and shift assessment never surfaced.

The weekend window was Saturday plus a Sunday recovery day. By Tuesday night, three days of troubleshooting later, two of them unplanned, the application can finally see its data. But the cutover was declared done with pieces still straddling both worlds: one application server was left on-premise “temporarily,” and the file share the batch jobs read from was never migrated at all. Both still straddle the circuit at week eighteen.

So what happened?

The accelerator deployed without issue. The workload landed on it in pieces. The landing zone was built for no particular workload, then handed one.

Post-Migration: The Cracks Widen

The workload is up. Then the first real business day begins, and the cracks widen.

Latency Issues

The on-premise SQL Server sat ten metres from the application servers. On the landing zone, the nightly batch job, a four-hour credit reconciliation running against the bureau data, now pulls its inputs over an ExpressRoute circuit from that unmigrated file share, because nobody inventoried the data flows. The four-hour batch becomes eleven hours, overrunning the bureau’s processing window. The 9am loan approval reports are late, and Ross, who runs operations from the Petone floor, asks why “the cloud” is slower than the server room.

Latency bites the interactive paths too. The application server left on-premise during the weekend cutover now makes thousands of chatty database calls across the circuit for every loan application. Approval times triple. Nobody had profiled the chatter because the assessment was a checklist, not a design exercise.

Compliance Issues

Tui, the compliance lead, catches the remediation work during its third attempt. Her questions are short and the landing zone answered none of them.

The database holding credit information falls under the Privacy Act 2020 and Credit Reporting Privacy Code, and the landing zone’s DR guidance pointed at a paired region in Australia. That was not automatically unlawful, but it was never assessed as an offshore disclosure with comparable safeguards. “Not a residency law” was treated as “any region is fine.” It was a default nobody caught. Meanwhile the policies still allowed engineers to deploy anywhere on Earth.

And then: centralised logging now ships the workload’s logs, containing customer PII, to a workspace that half the platform team can read. Privacy by design, inverted.

Then Sales signs a pilot with a government agency, and procurement forwards a questionnaire asking for an NZISM control mapping and evidence of a completed cloud risk assessment. The platform team spends the next four weeks remediating: region restrictions written after the fact, log access redesigned, the bureau integrations and on-premise data flows documented for the first time, and an unpleasant conversation with the agency’s assessor about why the architecture diagrams predate the workload they claim to host.

The board deadline is the last day of Q3. Cutover is week eight, a fortnight before that date, and still incomplete. Four weeks of Tui’s remediations push the programme into Q4. The CEO asks why the thing that “worked perfectly” in week two cannot host one application by the date the board was given.

Monitoring Issues

The workload’s alerting never made it to the new Log Analytics workspace. The monitoring that existed on-premise died with the cutover. The empty workspace from week two is still empty.

The Lessons We Can Learn

Nothing about the eventual architecture was wrong. The hub-spoke network was fine. The management group hierarchy was fine. What was wrong was the order: the landing zone was built for no particular workload, then handed one.

What they did Should have done
Built the LZ, then met the workload. Spent week one inventorying the workload: server specs, IP scheme, SQL dependencies, data flows, integrations, alerting.
Let root policies deny first, exempt later. Set policy with a documented exception process before the first migration.
Discovered the on-premise IP conflict at cutover. Reconciled the source addressing plan with the landing zone's design, and allowed for the overlap, before day one.
Assumed the SQL Server would lift as-is. Run an actual SQL assessment: linked servers, Agent jobs, compatibility level, hard coded hostnames.
Discovered latency in production. Mapped every cross-premises data flow, file shares, batch jobs, chatty app to database traffic, and profiled it before cutover.
Let DR defaults point offshore, with no region restriction. Written an allowed regions position and treated every offshore replica as an IPP 12 / Rule 12 decision.

See: Where should DR replicas live for a NZ regulated workload?

Centralised logs without classifying them. Decided what the logs contain, who may read them, and how long they are kept before shipping PII to a shared workspace.
Left the week two workspace empty after cutover. Moved on-premise alerts into Log Analytics before the weekend, and tested one page on the new path.

The accelerator never asked what workload would live on the platform. It could not. It deployed Microsoft’s defaults flawlessly, to a problem it had never been introduced to.

The Moral

A landing zone is an architectural expression of your organisation’s obligations and its actual workloads. Spurious Lending built one from defaults, then discovered both, obligations and workloads, one outage at a time. Two weeks of deployment bought them eight weeks of retrofit, a slipped deadline, and an auditor’s first impression they never fully recovered from.

The test: can you name the workload, its data flows, and the regulations covering its data before the first terraform apply? If not, you are deploying an accelerator, not designing a platform.

If you want that work structured as an engagement, see Fractional Cloud Architecture and Advisory.

Frequently Asked

Questions

Short answers to the NZ compliance questions this post usually raises. For the full reference list, start with Why Azure Landing Zones Fail Before They Begin.

View the full series

One Block

Before your next lift and shift, inventory the workload's data flows, dependencies, and address scheme, and list the three NZ regulations that apply to its data. Share both lists with the person who will own the first cutover weekend.
See all articles