Five Things That Break in Public-Sector Cloud Migrations — In Order

I've been close to 19 cloud migrations in public sector and finance. Six finished on time, five never finished, eight in between. The eight had the same five things break in the same order.

John Baek
John Baek
Founder, CollabOps
Five Things That Break in Public-Sector Cloud Migrations — In Order

In the last six years I've been close to 19 cloud migrations in the public sector and regulated finance — Korean and Japanese agencies, EU public bodies, retail banks, insurers, defense contractors. Six finished on schedule. Five never finished. The eight in between had the exact same five things break in the same order.

Order matters. This post walks the five in the order they break. If you don't catch number one, number two follows automatically.

1. Procurement assumes infrastructure, not service

At project start, internal procurement frames cloud as renting servers. So the RFP demands predictable monthly cost + 5-year SLA + a clear exit clause.

But cloud is a service catalog. Each service has its own pricing model. SLAs come as service-level credits, not refunds. Exit means data export feasibility, not refund obligation.

Get this first decision wrong and the RFP cycle stretches from six months to eighteen. Procurement and IT have the same meeting in different languages every week.

Fix: Before writing the RFP, IT spends 30 minutes briefing procurement on the four cloud pricing models. The next meeting runs in one language.

2. Data residency interpretations conflict

Once procurement aligns, the next thing to break is data residency. Almost every public agency has a one-line policy — "data must remain domestic" — and that one line has at least four valid interpretations.

Interpretation 1:  Service instances must be in domestic regions
Interpretation 2:  Operations staff (support, SRE) must be domestic
Interpretation 3:  Encryption keys must be in domestic KMS
Interpretation 4:  Backups and DR sites must also be domestic

You assume interpretation 1 satisfies the policy and start. Six months in, the security team produces interpretation 2 and the vendor selection itself has to restart.

Fix: In the first week of kickoff, security, legal, and the CIO sign a matrix that names which interpretations apply. Without that signed matrix, no RFP goes out.

3. SSO doesn't speak the national identity standard

Third break — identity. Global cloud providers' SAML/OIDC don't natively support Korea's KISA certificates, Japan's My Number Card, or EU's eIDAS schemes.

You discover this in the PoC the moment the first real user tries to log in with their real credentials. Up until then, everyone assumed "SSO will work, right?"

The fix is usually a bespoke adapter layer — adding one to three months. Then there's the question of who maintains that layer afterward (vendor? customer? integrator?). That's the secondary break.

Fix: In the first week of the PoC, verify one real user can log in with their real credential. If not, schedule the adapter immediately.

4. Audit trail expectations diverge from what cloud provides

Fourth break. This one breaks latest — usually in a security review at month six, when everyone thought things were going well.

Regulated industries' audit expectations:

- Who accessed which system, when, and what did they change?
- Under which *originating system's* authority did the action happen?
- Was that authority *valid* at that moment?
- Are all three preserved in *tamper-evident* form?

Cloud provider audit logs (CloudTrail, Cloud Audit Logs, etc.) give enough input for these questions, but reconstructing the answers is the customer's responsibility. Realize this late and you build a log analysis layer — adding two to four months.

Fix: Specify the three audit/monitoring tools at project start (provider audit + log aggregation + tamper verification). The blueprint with these three aligned must exist before the PoC.

5. Operational handoff — who wakes up at 3am

Last break. It happens just before launch. Live-day approaches, the integrator-run environment has to transfer to customer ops.

Seven out of eight times, the handoff manual covers cloud console usage but the incident-response responsibility matrix is blank.

Who opens a ticket with the cloud provider?         — undefined
Who escalates to the vendor on SLA breach?           — undefined
Who is the *first responder* at 3am?                 — defaults to integrator

Defaulting to the integrator means integrator costs continue for six months post-launch. At that point the total cost estimate falls apart.

Fix: Before PoC ends, sign an operational RACI matrix. For each scenario, name the first responder, second responder, and escalation owner.

What the six on-time projects had in common

There was one thing — only one — that the six on-schedule projects shared: the project manager already knew about these five before kickoff. Knowing the list shaved four to six months off the timeline.

All five have known fixes. None are difficult. The trap is the order: if number one breaks, number two follows automatically.

When delaying six months is cheaper

A public-sector, finance, or defense cloud migration should start with a fix in hand for each of the five breakages above. If you're a CIO or PM about to kick one off, fill in those five boxes first. If three or more are blank, delaying the start by six months is cheaper than starting now — the PMs who knew this list before kickoff finished four to six months faster than those who didn't.

Tags#public-sector#cloud-migration#compliance#finance#government