Observability and Telemetry
A doctor cannot open you up to see why you feel sick. Instead they listen with a stethoscope, read a thermometer, and ask what happened. A big computer system is the same. We cannot pause it and peek inside, so we make it leave clues as it works. Those clues are called telemetry. Observability is how well the clues let us answer brand new questions, like why checkout is slow for people in one city today, without first adding new code. Then we set clear goals for how healthy the system should be, and keep score.
Logs, metrics and traces
Three kinds of clues: a diary, a scoreboard and a map
Logs are like a diary. Each line says what happened at one moment, for example: a payment failed for order 1187 at 3:02. Logs are great for details, but a busy system can write billions of lines a day.
Metrics are like a scoreboard. They are numbers counted over time: requests per second, errors per minute, how full the disk is. They are cheap to keep and great for graphs and alarms, but they cannot tell you which exact request went wrong.
Traces are like a map of one request's trip. They show every service the request visited and how long each stop took. Together the three answer the big questions: is something wrong (metrics), where is it (traces), and why did it happen (logs).
All clues cost money to collect, send and store. So teams choose what is worth keeping, instead of saving everything forever.
Remember
Metrics tell you something is wrong, traces show where, and logs explain why.
Distributed tracing
Following one request on its trip through many services
One tap on the buy button can travel through many services: the front door, the cart, payment, the bank and email. If the tap is slow, which one is to blame? Tracing follows that one request, like a parcel with a tracking number. Every post office it passes through scans it, so later you can see exactly where it spent its time.
When a request first arrives at the edge (the front door of the system), it is given a trace ID, a random tracking number. Each service passes that number along to the next one inside the request headers, which are like the label on the outside of an envelope. Passing it along like this is called context propagation. A common standard for it is the traceparent header, defined by the W3C, a group that writes web standards.
Each service records a span: what work it did, when it started, when it ended, and which span called it (its parent). Gather all the spans with the same trace ID and you can draw a waterfall picture. Long bars show where the time went, and indented bars show who called whom.
Keeping every trace is costly, so systems sample, which means they keep only some. Head sampling decides at the very start, for example keep 1 request in 100. Tail-based sampling waits until the request has finished and then decides, so it can keep all the errors and all the slow ones, which are the interesting ones. The cost is holding every span for a short while before choosing. And if one service forgets to pass the trace ID along, the map breaks into separate pieces.
Remember
One trace ID passed along every hop turns scattered spans into one clear map.
Structured logging
Diary lines filled in like a form, not scribbled as a sentence
A free text log line reads like a sentence: payment broke for user 42 again, took forever. A person can read it, but a computer cannot easily answer how many payment errors happened this hour, because every developer writes their sentences differently.
Structured logging writes each line like a filled-in form, usually as JSON, a simple text format made of named fields and their values: time, level, service, trace_id, user_id, message. Now a computer can search, filter and count. Find every error from payment. Show all lines for one trace. Count failures per minute.
A few habits make logs much more useful. Use the same field names in every service, so trace_id is never also called traceId or tid. Always include the trace ID, so you can jump from one log line to the whole trace. Use levels such as debug, info, warn and error, so you can turn down the noise. And never log secrets or personal details like passwords, card numbers or home addresses, because logs get copied to many places and read by many people.
The trade-off is volume and cost. Structured lines are a bit bigger, and logging every tiny step can cost more than running the app itself. Log what helps you answer questions, not every heartbeat.
Remember
Write logs as forms with the same fields everywhere, include the trace ID, and leave secrets out.
SLI, SLO and SLA
Your test score, your goal, and your deal with your parents
An SLI, a service level indicator, is the measurement. It is like your score on the weekly spelling test. A good SLI measures what users actually feel, for example the percent of requests that succeed and answer in under 300 milliseconds.
An SLO, a service level objective, is the target for that measurement over a window of time. It is like your own goal: get at least 9 out of 10 words right every week. For a service it might be: 99.9 percent of requests are good, measured over 30 days. The team sets it and keeps it inside the company.
An SLA, a service level agreement, is a promise to customers that has consequences, like a deal with your parents: if your score drops below 7, no video games this weekend. For a company, breaking it usually means paying money back or giving credits. Because breaking it is costly, the SLA is set looser than the SLO, for example 99.5 percent. That way the team gets a warning, by missing its own goal, long before the promise is broken.
Aim for the right level, not for 100 percent. Every extra nine (going from 99.9 to 99.99 percent) costs a lot more work and money, and users often cannot tell the difference, because their own home internet drops more often than that.
Remember
The SLI is what you measure, the SLO is what you aim for, and the SLA is what you promise, usually a bit looser.
Error budgets
An allowance of mistakes you are allowed to spend
If the goal is 99.9 percent good, then 0.1 percent is allowed to go wrong. That leftover is the error budget, like a weekly allowance of mistakes. In other words, the error budget is 1 minus the SLO. Thirty days have 43,200 minutes, and 0.1 percent of that is 43.2 minutes of full downtime. Counting by requests instead, 1 out of every 1,000 requests may fail.
Burn rate is how fast you are spending the budget. A burn rate of 1 spends it exactly by the end of the 30 days. A burn rate of 10 would use it all up in 3 days. Good alarms watch the burn rate. A very fast burn wakes someone up right away, while a slow burn becomes a task to fix during the working day.
The budget also turns arguments into a simple rule. If plenty of budget is left, the team can ship new features faster and take more risks. If the budget is spent, risky launches are frozen and the team works on reliability until things are healthy again. Everyone agrees on this rule ahead of time, so nobody has to argue about it in the middle of a bad day.
An error budget is only as good as its SLI. If you measure the wrong thing, you can have plenty of budget left while users are unhappy.
Remember
The error budget is 1 minus the SLO. Spend it on speed while you have it, and fix reliability when it runs out.
Quick recap
- Observability means working out what is happening inside a system from the clues it gives off.
- Logs give details, metrics show trends and drive alarms, and traces show one request's whole trip.
- A trace ID travels in request headers, and each service adds a span, so you can draw a waterfall.
- Structured logs use named fields with the same names everywhere, include the trace ID, and never contain secrets.
- The SLI is the measurement, the SLO is the internal goal, and the SLA is the looser promise to customers.
- The error budget is 1 minus the SLO, about 43 minutes a month at 99.9 percent. Watch how fast it burns.
Grown-up words
and what they mean in plain words
- Telemetry
- The clues a system sends out about itself.
- Log
- A diary line about one moment.
- Metric
- A number counted over time, like errors per minute.
- Trace
- The map of one request's trip through many services.
- Span
- One stop on that trip, with a start time and an end time.
- Trace ID
- The tracking number shared by every span of one request.
- SLI
- Service level indicator: the measurement of how well things are going.
- SLO
- Service level objective: the team's goal for that measurement.
- SLA
- Service level agreement: a promise to customers, with a cost if broken.
- Error budget
- How much failure the SLO still allows, and burn rate is how fast it is used.