The earlier field notes covered missing context, deny first governance, isolation, subscription sprawl, and datacentre networking. This one is about institutional memory: what happens when Infrastructure as Code is a repo that no longer describes reality.
SELL! is a fictional company. Any resemblance to a real platform team’s war story is the point.
Setting the Scene
Week twenty-six at SELL!. Six months in: the first hiring cycle has come and gone, and nobody new has landed. Sam, the platform lead who told Tui the golden path was ready in Part 3, has a Sydney offer. The story that follows a failed first impression across the next three jobs is now following Sam to Australia. The bus number for the platform is one, and that one is boarding a plane.
Document the portal clicks. Promise to "get back to IaC" next quarter.
- Export a few screenshots into Confluence.
- Share the personal service principal "temporarily."
- Defer drift detection until after the agency audit.
- Hope the next hire can reverse engineer the estate.
The Pilot: The Laptop Apply Works Perfectly
On paper the landing zone is in a repo. Sam’s first terraform apply from a laptop worked in week two. Screenshots in Confluence feel like a handover. Tui asks whether the agency audit can wait until Sam’s last Friday. The plan says yes.
The Repo Nobody Trusted
Hemi stopped trusting it in week fourteen, when a portal firewall change saved payments and never came back as a pull request. By week twenty-six the repository is a historical document. The truth lives in the portal.
Resignation Day: The Bus Number of One
Sam’s last week. Before they leave, the team discovers:
- A firewall rule added during an incident, never codified
- A management group moved by hand to unblock payments
- Policy tweaks applied in the portal “just for now”
- Remote state that only one identity can unlock
The notice period ends. The agency audit does not wait.
Post-Departure: The Cracks Widen
The knowledge transfer deck is done. Sam is in Sydney. Then Tui collects for the agency audit, and the evidence does not exist.
Drift Issues
A workload app changes when its team deploys. A landing zone changes when anything changes: Microsoft updates policy definitions, security requirements shift after an incident, a new NZISM control needs implementation, cost pressure demands a SKU change. High frequency change plus manual deployment equals configuration drift , guaranteed.
Continuity Issues
In NZ, the small team problem bites twice: with one or two people holding the operational knowledge, their departure is a continuity incident. IaC discipline is institutional memory.
Audit Issues
Tui needs who changed which policy, and when. The trail is Sam’s mailbox and a Confluence page of screenshots. The questionnaire that was deferred until after the audit is now the audit.
| Symptom | What to do instead |
|---|---|
| Someone's laptop is the source of truth. Within a year the repo no longer matches the portal. | Everything through pipelines from day one, including the first deployment. Treat IaC as the product.
Test policy and RBAC changes in a non production management group before
production. References: Why Landing Zones Fail · CAF: Platform automation and DevOps · WAF: Operational Excellence |
The Lessons We Can Learn
You do not need everything on day one. You need these, roughly in this order.
1. Everything through a pipeline, including the first deployment. No exceptions for temporary portal fixes. If a change happens in the portal during an incident, it becomes a change request against the repo the same day. Label those PRs portal-changes and keep a one page incident-codification checklist next to the runbook.
2. State and identity hygiene. Remote state with locking (terraform state lock, or the equivalent). Least privilege deployment identities (workload identity federation or managed identity, not a human’s account). Platform repo access separated from workload repos.
3. Environment strategy for the platform itself. terraform plan on every pull request. Test platform changes in a non production management group before production. The spend is small. A broken deny locking out production pipelines on a Friday is not.
4. Policy as code with validation. Policies versioned in the same repo. Static analysis on the platform code. A test harness that catches “this policy change would break vending” before merge.
5. Release discipline. Semantic versioning, a changelog workload teams can read, communicated deprecations. Breaking changes without notice is how platform adoption dies quietly. Azure DevOps or GitHub Actions both work; pick one and put the apply there, not on a laptop.
6. Drift detection. Scheduled runs flag when reality no longer matches the repo. In a small team, that automation is the audit function between reviews.
7. Clone and README must be enough. Document the bus number in the repo: which pipeline, which identity, which order. If git clone and the README cannot reproduce the path, the repo is a monument.
| Symptom | What to do instead |
|---|---|
| Knowledge transfer is screenshots and a shared secret. The next hire cannot reproduce the platform. | Document the bus number in the repo: which pipeline, which identity, which order. `git clone` and the
README must be enough. References: CAF: Platform automation and DevOps · WAF: Safe deployment practices |
Operations: Reconverge the Drifted Estate
The departing engineer has left. The repo is a monument. Reconvergence is a programme, not a hope:
- Treat the portal as the truth of record until each scope is imported and the plan is clean
- Freeze portal changes for that scope while you catch up
- Import one management group at a time. Do not boil the tenant
- Unlock state onto a workload identity before you need a second pair of hands
A bootstrap or a drift catch-up is a fractional deliverable. See Fractional Cloud Architecture and Advisory.
| What they did | Should have done |
|---|---|
| Applied the first landing zone with a personal credential. | Bootstrapped through a pipeline and a workload identity from day one. |
| Fixed incidents in the portal and left the repo behind. | Codified every portal change the same day, no exceptions. |
| Promoted policy changes straight into production management groups. | Tested in a non production management group with validation in the pipeline. |
| Treated drift detection as a nice to have after the audit. | Scheduled detection runs early, while the team still remembered what "correct" meant. |
The Moral
In a market where senior engineers are routinely recruited offshore, the platform that survives a key departure is the one where the repo, not a person, holds the truth. SELL! learned that the week the truth tried to resign.
The test:
git clone, follow the README, and you can reproduce the platform’s deployment path without asking anyone. If that fails, the repo is a monument, not a source of truth.
One Block · build from here