System Resiliency
Think about a big school on a stormy day. The phone line to the bus company crackles, so the office calls again, but not every second. When the bus company clearly is not answering, the office stops calling for a while and tells parents the backup plan. The lunch line and the library line are kept apart, so a jam in one does not stop the other. And if the whole building loses power, there is a second building across town ready to open its doors. Resilient systems work the same way. They expect parts to break, and they plan so that one broken part does not take everything else down with it.
Retries with exponential backoff
Try again, but wait a little longer each time
Sometimes a call fails for a short, silly reason: a busy moment, or a message that got lost on the way. Trying again often works, like knocking on a door a second time when nobody heard you the first time.
But knocking again and again with no pause just annoys someone who is already busy. Exponential backoff means each wait is double the one before: wait 100 milliseconds, then 200, then 400, then 800. There is a cap, so the wait never grows past a limit like a few seconds, and a maximum number of tries. After the last try, we give up and report the error.
Only retry things that are safe to repeat. Asking for the lunch menu twice is harmless. Charging a card twice is not. An action that gives the same result no matter how many times you do it is called idempotent. For actions like payments, send an idempotency key, a unique ticket number for that request, so the server can notice it already did this one and not do it again.
Retries can also make an outage worse. If a website calls service A, which calls B, which calls C, and every layer makes 3 attempts, one broken C can receive 3 times 3 times 3, which is 27 calls for a single click. This is a retry storm. Good habits are to retry at only one layer, and to keep a retry budget, for example retries may be at most 10 percent of all calls, so retrying stops when something is clearly broken.
Remember
Retry only safe things, wait longer each time, and know when to give up.
Jitter
A little randomness so everyone does not knock at the same moment
Imagine the whole class is told: if the slide is busy, come back in exactly one minute. Then everyone comes back at the very same second, and the slide is jammed again. Waiting longer does not help if everyone waits the same amount.
When a server hiccups, thousands of clients fail at the same moment. With plain backoff they all retry at 100 milliseconds, and then all again at 300, in giant waves. A crowd rushing in together like this is called a thundering herd, and it can knock the server over again just as it was getting back up.
Jitter means adding randomness to each wait. A popular recipe called full jitter picks a random wait between zero and the current backoff. One client waits 23 milliseconds, another 87, another 51. The same number of retries arrive, but spread out in a gentle trickle that the server can handle.
The trade-off is that one request's timing becomes less predictable. Some waits come out shorter and some longer. That is a small price for not trampling the server.
Remember
Backoff spreads retries out over time, and jitter spreads them out across clients.
Circuit breakers
Stop calling a friend who is clearly not answering
A house has a circuit breaker that switches the power off when too much electricity flows, so the wires do not overheat. Software borrows the idea. A circuit breaker sits in front of calls to another service and keeps track of how often they fail.
It has three states. Closed is normal: calls go through, and failures are counted. When failures cross a limit, like half of the last 20 calls, the breaker trips to open. While open, it does not call the service at all. It fails fast, answering right away with an error or a backup answer. That saves our own workers from waiting on slow timeouts, and it gives the sick service some quiet time to recover.
After a cool-down, like 30 seconds, the breaker becomes half-open. It lets a few trial calls through. If they succeed, the breaker closes and life goes back to normal. If one fails, it opens again and waits another cool-down.
A breaker works best with a fallback, which is a plan B answer such as the weather saved earlier today, a default list of popular toys, or a polite message asking people to try again later. The hard part is tuning. Trip too easily and you block a service that was fine. Trip too slowly and your workers pile up waiting.
Remember
Closed lets calls through, open fails fast, and half-open tests the water before trusting again.
Bulkhead isolation
Walls inside a ship so one leak does not sink it
Big ships are split into watertight rooms by walls called bulkheads. If one room springs a leak, only that room floods, and the ship stays afloat.
In software, the leak is usually a slow service. Say our app has one shared pool of 50 workers (threads or connections) for every kind of call. The weather service becomes very slow, so every worker that calls it gets stuck waiting. Soon all 50 workers are stuck, and even payments, which is perfectly healthy, cannot be served because nobody is free.
With bulkheads, each dependency gets its own small pool, say 10 workers for weather and 20 for payments. When weather is slow, only its 10 workers get stuck, and payments keeps running. The same idea works at bigger sizes too: separate servers for big customers and small customers, or a separate copy of a service for each region, so one noisy group cannot hurt the others.
The trade-off is some waste. Split pools cannot lend each other spare workers, so one pool may sit half empty while another is full, and every pool must be sized with care. Bulkheads pair well with timeouts and circuit breakers, which free stuck workers sooner.
Remember
Give each risky dependency its own pool, so one slow friend cannot hold up everyone.
Multi-region failover
A second building across town, ready to open its doors
A region is a group of data centers in one part of the world. A storm, a power cut or a bad update can take a whole region down. So important systems run in more than one region.
In active-passive, one region serves everyone and another waits as a backup. A warm standby has servers already running and ready. A cold standby must be started up first, which is cheaper but slower. In active-active, every region serves users all the time. It uses every building fully and recovers fastest, but it is much harder to build, because two regions may change the same data at the same time.
To switch, a traffic director keeps checking each region with health checks, like a teacher calling roll. When a region stops answering, the director (a global load balancer, or DNS, the internet's address book) sends people to the healthy region. DNS answers are remembered by devices for a while, so some people may keep going to the old address until that memory runs out. Data is copied between regions, usually a little behind (asynchronously), so the last few moments of changes may be lost. RPO, the recovery point objective, is how much recent data you can afford to lose. RTO, the recovery time objective, is how long you can afford to be down.
Two dangers remain. A failover you never practice probably will not work, so teams hold game days where they switch over on purpose. And watch out for split brain: if both regions think they are in charge and both accept changes, the data goes two different ways and is very hard to merge. Systems prevent this by making sure only one side can be in charge at a time, for example with a tie-breaking vote from a third place.
Remember
Keep a second region ready, know your RPO and RTO, and practice the switch before you need it.
Quick recap
- Retry brief hiccups, with waits that double up to a cap, and a limit on the number of tries.
- Add jitter so crowds of clients do not all retry in the same instant.
- Only retry safe actions, or use idempotency keys, and retry at one layer to avoid retry storms.
- A circuit breaker goes from closed to open after many failures, then half-open to test if things have recovered.
- Bulkheads give each dependency its own pool, so one slow service cannot use up every worker.
- Multi-region failover moves traffic to a healthy region. RPO is the data you may lose, RTO is the time to recover.
Grown-up words
and what they mean in plain words
- Retry
- Trying the same call again after it failed.
- Exponential backoff
- Doubling the wait before each new try.
- Jitter
- Randomness added to waits so clients spread out.
- Idempotent
- Safe to do twice: the result is the same as doing it once.
- Circuit breaker
- A switch that stops calls to a failing service for a while.
- Fallback
- A plan B answer used when the real one is not available.
- Bulkhead
- A wall that keeps trouble inside one compartment.
- Failover
- Moving work to a backup when the main one fails.
- RPO
- Recovery point objective: how much recent data you can afford to lose.
- RTO
- Recovery time objective: how long you can afford to be down.