Skip to content
Richard Cooper
Go back

The strangler fig

Every legacy system eventually produces the same meeting. Someone says “we should just rewrite it,” and everyone feels a brief illicit thrill, because a rewrite is the fantasy where the constraints go away. It stays a fantasy because a big-bang asks for two things nobody can give you: a year without feature delivery, and a re-specification, from the outside, of a system nobody fully understands any more. The real specification isn’t the documentation — it’s a decade of accumulated behaviour customers depend on, including the parts that are technically bugs.

The alternative is the strangler fig, named for the plant that germinates in the canopy, roots down around its host, and ends up standing as a hollow column where the original tree was. Put something in front of the legacy system that can route a request either way, move one capability at a time, and when nothing routes to the old system, delete it.

That’s the pattern, and it’s the right one. It’s also sold far too cheaply — having spent a while inside one, the parts that decide it aren’t the ones in the diagram.

Don’t migrate ahead of capability

The single rule that makes the difference: something only becomes eligible to move once the new system can service it completely, end to end.

That sounds obvious and it is routinely violated, because the pressure is always to show migration progress. If you move a customer whose needs the new platform only mostly covers, you haven’t migrated them — you’ve created a support problem that neither system can solve, since the old one no longer owns them and the new one can’t finish the job.

Done properly, this inverts the usual planning direction: the migration schedule becomes a consequence of the feature roadmap rather than a parallel track competing with it. Each cohort lines up behind the capability it needs. On the migration I worked on, the first cohort was deliberately the most boring population we could describe — the simplest product shape, single item, no add-ons, standard monthly billing — precisely because those were the only customers the new platform could fully serve. Everything else was explicitly gated: the customers with add-ons waited on the add-ons design, the multi-item customers on a separate cutover, and so on. Anything with work in flight was excluded until that work resolved.

The ramp within a cohort matters as much as the ordering: ten, then fifty, a hundred, and up in steps, with a rehearsal at the same size in a lower environment before each tranche and a requirement that the previous tranche be stable first. This is the last responsible moment applied to risk rather than design — you keep the population small while your information is poor, and you only widen once each step has told you something.

A narrow cohort also does something subtler than limiting blast radius: it makes whole categories of defect irrelevant for now. Our field-level mapping audit found real gaps — data the migration tool never read, a hard-coded date, a loader fetching a plausible-looking wrong row — and none could affect the first cohort, because its selection criteria excluded every shape that used those fields. Be precise about what that buys, though. Those gaps were bounded, not fixed, and every widening of the cohort brought some of them into scope.

The seam is the whole game

You can only strangle what you can intercept. The facade you put in front of the legacy system is an anti-corruption layer doing exactly the job that post describes — translating between an old model and the one you actually want. And as its follow-up argues, the seam that protects your new domain from the legacy model is the same seam that lets you swap which implementation serves a request.

Where a legacy system has no seam, you have to build one, and that work is the least glamorous and most decisive part of the whole programme. If consumers reach into the old database directly, or call in-process, or if the reporting layer reads the tables, then “route this capability elsewhere” isn’t a routing change at all — it’s a change to every caller. Plenty of migrations die here, before the interesting work starts.

The mechanism worth stealing is one we ended up with on the legacy side: a tracking table that records what has moved. It does three jobs at once. It’s the record of migration state, it’s the thing that switches the old system off for those accounts — and, read the other way, it’s the switch that turns them back on.

The seam only covers what comes through the seam

This is the one I’d most want to hand to someone starting a migration, because it cost us real damage and it isn’t in the diagram.

A facade routes requests. Scheduled work is not a request. Nothing routes a nightly job, a timer, a queue consumer, a third-party webhook or a report — they sit behind the seam and reach straight into whichever system they were written against.

On a different slice of the same programme, a nightly job picked up subscriptions that had fallen due for cancellation and flipped each parent record to cancelled. It did that correctly. What it didn’t do was publish the downstream event cancelling the line items underneath — and neither did the new path, because each had been written assuming the other owned that step. Split-brain in the literal sense: two halves of one operation, each individually correct, with a gap no test covered because no test spanned both.

It ran for nine nights, cancelling 177 parent records and leaving 271 line items live and orphaned. Nothing alerted, because from inside either system nothing had failed. The old path did its job, the new path did its job, and the operation as a whole existed nowhere as something that could succeed or fail.

Two lessons I’d now treat as non-negotiable. Inventory everything that touches the legacy store without going through the facade before you move a single account — batch jobs are where this hides, because they’re invisible on a diagram drawn around request flow. And assign every step of a multi-step operation to exactly one side of the seam, explicitly. “Both paths handle this” is the same sentence as “neither path handles this”, and you learn which you meant several nights later.

Reversibility is layered, and it expires

“We can always roll back” is the claim that makes migrations approvable, so it deserves more scrutiny than it usually gets.

In our case reversibility was real, but it came from three different mechanisms at three different levels, and only stating all three is honest:

  1. Cohort discipline bounds who is exposed at all — an escalating ramp means the worst case is always the size of the current tranche.
  2. The tracking table is the operational flip: set the state back and those accounts are served by the old system again.
  3. Point-in-time restore on the target database gives a bounded window in which the new system’s data can be wound back.

Put together, a genuine post-migration rollback means restoring the target, flipping the tracking state, and letting the old system resume. But notice what that depends on. The restore window has a duration, and once you’re outside it the data half of the rollback is gone — you’re in forward-fix territory, whatever the plan said.

So “the migration is reversible” and “rollback means forward-fix” are both wrong as blanket statements. Which one is true depends on how long ago the tranche ran. I’d go further: if you can’t say how long your reversibility lasts, you don’t have reversibility, you have an assumption. It’s a one-way door with a slow hinge, and the hinge closes on a schedule nobody has written down.

Somebody has to own the truth

The data question is the one that gets under-planned, and it isn’t really a technical question. It’s this: for any given account, at any given moment, which system is authoritative?

The answer we settled on is the simplest one available, and I’d recommend it: the old system is the system of record before migration, the new system is the system of record after, and the handover is the migration event itself. One authority at a time, with a defined moment of transfer. What you must not do is let both be partially true, because “which balance is right” becomes a question with no answer and someone will eventually be given the wrong number.

Choosing that has consequences you have to accept out loud rather than discover:

The same care applies to the schema on both sides — expand, migrate, contract, exactly as in zero-downtime by design — and to the transitional period where two systems hold overlapping state, which buys availability with consistency.

The clock starts from the last thing you move

Here’s the detail that reframed the economics for me, and I’ve not seen it discussed anywhere.

Our retention commitment was a year in read-only after the final migrated account. Not a year per account — a year from the last one. Which means every month of migration slippage extends the legacy system’s life, licences, hosting and audit obligations by a month at the far end. The cost of going slowly isn’t only the cost of running two systems while you go; it’s a directly proportional extension of the thing you’re trying to be rid of.

That inverts the intuition that a slower strangler is a safer strangler. Slower is safer per tranche and more expensive in total, and a migration that stalls halfway is the worst of every world: two systems to run, a facade that has quietly become permanent architecture, and a decommission date that has stopped approaching. Which is, I think, the real failure mode of this pattern — not a bad cutover, but a loss of organisational attention once the visible wins have been banked. The easy capabilities go first and generate the appearance of momentum. What’s left is the undocumented behaviour that three customers depend on.

A colleague made this point more economically than I’ve managed to, in a thread where the term was being used loosely for “we’ll gradually improve it”. He posted Fowler’s original article and added:

He gave it the name for a reason.

Which is the whole argument in six words. The fig doesn’t cohabit with its host, and it isn’t a metaphor for gradual improvement. It kills the tree and stands in its place. A strangler migration that never reaches that state hasn’t been cautious — it has just stopped, and taken the name of a pattern it isn’t following.

The honest limits

The one-line version

The strangler fig works, and it is the right pattern — but it is not the incremental, risk-free story it gets sold as. It asks you to build a seam you may not have, to name a single system of record and live with what that costs, to know the expiry date on your own reversibility, and above all to finish, because the whole economic case rests on a deletion that nobody will be thanked for. The pattern is sound. The hard part was never the routing.


Share this post:

Previous Post
Your people are the problem (and also the solution)
Next Post
A Design Authority without the bureaucracy