Skip to the lesson
Little Builders system design, explained small

Complex Migrations and Evolution

Imagine a school bus full of kids that needs new wheels, but it is not allowed to stop, because everyone must get to school on time. The mechanics have to swap one wheel at a time while the bus keeps rolling, and they must be ready to put an old wheel back if a new one wobbles. Changing a big computer system works the same way. People use it every second of every day, so we cannot switch it off, replace everything, and hope for the best. Instead we change it in small, safe steps while it keeps running, and we always keep a way back.

Strangler Fig pattern

Grow the new tree around the old one, branch by branch

In some forests, a vine called a strangler fig starts high up on an old tree and sends its roots down to the ground. Year after year it wraps around the tree. One day the old tree is gone, and the fig stands in its place, the same shape, but all new.

We can replace an old computer system the same way. First we put a helper in front of it, like a hallway monitor who points each kid to the right classroom. Grown-ups call this helper a routing layer or a facade. At the start it sends every request to the old system. Then we rebuild one small feature in the new system, say the toy search, and tell the monitor to send only search requests to the new room. Everything else still goes to the old room.

We repeat this, one feature at a time, until the old system gets no requests at all. Then we switch it off. If a new feature misbehaves, the monitor simply points that one feature back to the old room. So every move is small, and every move can be undone.

The price is that for months we run two systems side by side, and both often need the same information. If a kid changes their name in the new system, the old one may need to hear about it too, so the two copies of the data must be kept in step. That syncing is often the hardest part of the whole job.

A new vine wraps the old tree one ring at a time, one ring per feature moved. Raise the pointer to move more, lower it to give some back.

Remember

Put a router in front, move one feature at a time, and retire the old system only when nothing uses it.

Backward-compatible schema changes

Change the sign-up sheet so old and new readers both understand it

A schema is the shape of our saved data: which boxes each record has, and what goes in each box. Think of the sign-up sheet for a class trip, with columns for name, age, and lunch choice. Our toy shop has a record for each toy, with a column called price that holds the price in dollars.

When we update an app, we do not replace every computer at the same instant. We swap them a few at a time, so for a while old copies and new copies of the app run side by side, reading and writing the very same data. And if we have to undo an update, old copies run again on data the new copies wrote. So every change to the data must make sense to both versions at once.

Backward compatible means the new app can still read data the old app wrote. Forward compatible means the old app can still read data the new app wrote. During a rollout we need both, because each version reads what the other one saved. A few simple rules keep us safe:

  • Add, do not rename or delete in one step. A rename is really a delete plus an add, and the old app will still look for the old name.
  • Make new fields optional and give them a sensible default, because records saved by the old app will leave them empty.
  • Never change what an existing field means. If price means dollars, do not start storing cents in it, or the old app will read 1999 cents as 1,999 dollars. Add a new field, such as price_cents, instead.
  • Readers should skip fields they do not recognize instead of crashing. Then the new app can add fields without breaking the old one.
One lock, an old key and a new key. Point at either key and it opens the lock, as both must work while the update rolls out.

Remember

During a rollout, old and new versions run together, so every change must make sense to both.

Expand and contract

Put up the new name tag before you peel off the old one

Picture relabeling the cubbies in a classroom. First you stick the new name tag next to the old one. For a while both tags are there, so every kid finds their cubby whichever name they look for. Only when everyone knows the new tag do you peel the old one off. That is expand and contract, also called parallel change.

Back to our toy shop. We want prices stored in cents, in a new field called price_cents, and the old price field to go away. We get there in six small releases. Steps 1 and 2 are the expand part, steps 3 and 4 move everything over, and steps 5 and 6 are the contract part.

After every step, both the old and the new app still work, so we can pause or go back at any point. Steps 1 to 4 are easy to undo. Step 5 can still be undone by copying fresh values back into price. Step 6 throws the old data away for good, so we only take it once we are sure nothing reads price anymore.

The trade-off is time and care. A change that sounds like one line takes several releases, the data is stored twice for a while, and the code must handle both fields. That extra care is the price of never stopping the bus.

  • 1. Add price_cents, empty and optional. Nobody uses it yet.
  • 2. Write both: every save fills in price and price_cents.
  • 3. Backfill: a background job fills price_cents for all the old rows, in small batches so the database stays calm.
  • 4. Switch reads: the app now trusts price_cents.
  • 5. Stop writing price, since nothing reads it anymore.
  • 6. Remove price for good.
Relabel a cubby in six safe steps: stick the new tag on, use both, switch to the new one, then peel the old one off. Slide to step through.

Remember

Expand first, move everything over, and contract last, one safe release at a time.

Blue-green deployment

Two identical playrooms, and a door that swings between them

Imagine two identical playrooms, one painted blue and one painted green. Right now all the kids play in blue. While they play, grown-ups set up the new toys in the empty green room and test every single one.

When green is ready, they swing the door, so every kid who arrives goes to green instead. The swing happens all at once. In computer terms, a router or load balancer, which is the traffic helper that decides where each request goes, starts sending everything to the new copy. If something looks wrong, they swing the door back to blue, which is still set up and waiting. Going back takes seconds, not hours.

There are costs. You need two full playrooms, so you pay for about double the computers, at least while you switch. The new room starts cold, with nothing remembered yet, so teams often warm it up with test traffic first. If the switch is done by changing the internet's address book (DNS), some phones remember the old address for a while, so a router switch is quicker and cleaner.

Most importantly, both rooms usually share one database, the one big toy box. So the data changes must work for blue and green at the same time, which is exactly what expand and contract gives us.

Two identical playrooms and one door. Point at a room: the door swings to close the other one, and every kid moves to the room now open.

Remember

Get the idle copy ready, switch everyone at once, and switch back if anything looks wrong.

Canary releases

Let a few kids try the new slide before everyone does

Long ago, coal miners took a little canary bird down into the mine. The bird felt bad air before people did, so if the canary got sick, the miners knew to get out fast. A canary release is a small early warning in the same spirit.

Instead of sending everyone to the new version, we send a small slice, maybe 1 in every 100 people. Then we compare: does the new version make more mistakes (errors) or answer more slowly (latency) than the old one, for the same kind of visitors at the same time? If it is just as good, we widen the slice to 5 percent, then 25, then everyone. If it is worse at any step, an automatic checker sends everyone back to the old version.

The good part is that a bug only reaches a few people, and only briefly. The costs are that a rollout takes longer, both versions run together for a while (so the data rules from earlier still apply), and you need good measurements. With a tiny slice and few visitors, there may not be enough results yet to tell whether the new version is really worse, so teams wait long enough at each step to be sure.

A hundred seats, and a few try the new version first. Raise the pointer to widen the slice from 1 to 5 to 25 to every seat.

Remember

Start small, compare against the old version, and widen only while it stays healthy.

Feature flags

A light switch for each new feature

Putting new code on the computers and turning a new feature on do not have to happen at the same moment. A feature flag is like a light switch inside the app. We deliver the new code with the switch off, so nobody notices anything. Later we flip the switch, maybe first for teachers only, then for 1 percent of kids, then for everyone.

This splits two jobs apart: deploying, which means putting code on the computers, and releasing, which means letting people use it. If the new feature misbehaves, we flip the switch off in seconds, without shipping any new code. Flags pair nicely with canary releases and with the strangler fig, because they let us move people over a little at a time.

The trade-off is clutter. Every flag is a fork in the road inside the code, and each one doubles the number of ways the app can behave, which makes testing harder. Old flags that nobody removes pile up like forgotten switches on a wall, so good teams delete a flag soon after its feature is fully on.

Remember

Deploy the code switched off, then release it with a switch you can flip back.

Quick recap

  1. Big changes happen while the system keeps running, so we change it in small steps and always keep a way back.
  2. Strangler fig: put a router in front of the old system and move one feature at a time to the new one, until the old one can be retired.
  3. During any rollout, old and new versions run at the same time, so every data change must make sense to both.
  4. Expand and contract: add the new field, write both, backfill, switch reads, stop writing the old one, then remove it.
  5. Blue-green: prepare an identical copy, switch everyone at once, and switch back if needed. It costs about double the computers while you switch.
  6. Canary: send a small slice to the new version, compare it with the old one, then widen or roll back.
  7. Feature flags separate shipping code from turning features on.

Grown-up words

and what they mean in plain words

Migration
Moving a system or its data from an old way to a new way.
Strangler fig
Replacing an old system piece by piece behind a router, until nothing of it is left.
Facade (routing layer)
A helper in front of the systems that decides where each request goes.
Schema
The shape of saved data: which fields a record has and what they mean.
Backward compatible
The new version can still use data the old version made.
Backfill
Filling in a new field for all the records that already exist.
Blue-green deployment
Two identical copies of a system, with all traffic switched from one to the other at once.
Canary release
Giving a new version to a small slice of users first and watching it closely.
Rollback
Going back to the previous version when the new one misbehaves.
Feature flag
A switch in the code that turns a feature on or off without shipping new code.