Skip to the lesson
Little Builders system design, explained small

API Design and Networking

In class, kids pass notes. To get an answer, you need to know who to ask, how to write the note so they understand it, and how it will travel across the room. Computers pass notes all day long. An API is like a restaurant menu: it lists exactly what you may ask for and how to ask, so the kitchen never has to guess. In this lesson we look at the roads the notes travel on (HTTP/2, HTTP/3 and gRPC), how to keep a line open for chatting (WebSockets), the helpers who stand at the door (gateways and load balancers), and the hall monitor who stops anyone from asking too often (rate limiting).

HTTP/2

Many notes share one conveyor belt, cut into small labeled pieces

HTTP is the set of rules a web browser and a server use to pass notes. With the older version, HTTP/1.1, one connection (one open road between two computers) carries one request and its answer at a time. If a big picture is coming down the road, a small note behind it has to wait. Browsers worked around this by opening about six roads to each website, which costs extra time and work.

HTTP/2 uses one road for everything. Every answer is cut into small pieces, and each piece gets a label saying which request it belongs to. One request and its answer together are called a stream. Pieces from many streams are mixed together on the same road, and the other side sorts them back out using the labels. This is called multiplexing. A big picture no longer blocks a small note, because their pieces take turns.

HTTP/2 also squeezes the headers, the little address labels that ride along with every request. Many of them repeat on every note, like which browser you use and what kinds of answers you understand. So both sides keep a shared list and send a short number instead of the whole label. This is called header compression.

The weakness is the road underneath. HTTP/2 runs on TCP, a delivery rule that promises every piece arrives in the exact order it was sent. Picture one conveyor belt carrying everyone's toy pieces. If one piece falls off, the whole belt stops until that piece is sent again, even for kids whose pieces were fine. This is called head-of-line blocking, and it hurts most on bumpy networks like a weak phone signal.

One belt carries pieces of three streams, mixed and taking turns. Dots on each lid say whose it is. Hover to slow the belt and follow one.

Remember

HTTP/2 mixes many requests on one connection, but on TCP one lost piece still makes them all wait.

HTTP/3 (QUIC)

Each note gets its own little belt, all riding in the same truck

HTTP/3 keeps the good ideas of HTTP/2 but changes the road underneath. Instead of TCP, it uses newer rules called QUIC, which are built on UDP. UDP is a very plain way to send pieces, with no promise about order or arrival. QUIC adds its own promises on top, but it keeps track of order for each stream separately.

Picture a truck carrying several small belts, one for each kid. If a piece falls off one belt, only that kid waits for a new copy. Everyone else keeps getting their pieces. So a lost piece only slows down its own stream.

QUIC is also quicker to start. Before two computers can talk safely, they say hello and agree on a secret code for locking their messages, called encryption. TCP with modern encryption needs two back and forth trips before the first real request. QUIC does both in one trip. When you come back to a site you visited before, it can even send the first request with zero waiting trips. A sneaky listener could copy and replay that very first early request, so it is only used for safe things like reading a page.

QUIC gives each conversation its own ID number instead of tying it to the phone's network address. So when you walk out of the house and your phone switches from Wi-Fi to mobile data, the conversation keeps going instead of starting over. The trade-offs: some office and school networks block or slow down UDP, so browsers keep HTTP/2 as a backup, and QUIC does more of its work in ordinary programs, which can cost servers more computing power.

Three streams ride their own little belts. When one loses a piece, only that lane waits for the copy. Point at a lane to drop a piece there.

Remember

HTTP/3 gives every stream its own lane, starts faster, and survives a network switch.

gRPC

A strict order form that both sides agreed on before anyone ordered

Imagine a pizza shop where you must order with a printed form. The form lists every box you can fill in and what goes in each box: a number for the size, a word for the topping. You and the shop got the very same form ahead of time, so nobody misreads an order. gRPC works like this. The form is a file written with Protocol Buffers (often called protobuf), and it describes every message and every question one service can ask another.

From that one form, tools write the calling code for you, in many programming languages. Asking another service then looks like asking a friend down the hall a normal question. The generated code packs each message into compact binary, a tight row of numbers instead of readable words. That is smaller and quicker to pack and unpack than JSON, the readable, labeled text most web APIs send.

gRPC rides on HTTP/2, so many calls share one connection. It also supports streaming: the server can send a long stream of answers to one question, the client can send a stream of messages, or both can talk at the same time. Each call can carry a deadline, so a slow answer is given up on instead of waited for forever.

It shines when services inside one company talk to each other, because everyone can share the form. The trade-offs: binary messages are hard for people to read while hunting bugs, and web browsers cannot speak gRPC directly, so they need a helper called gRPC-Web plus a small translator in front of the server. Also, since one connection stays open for a long time, the load balancer must spread out single calls, not just connections, or one server ends up doing all the work.

Both sides share one form: a plate with shaped holes, and blocks that fit only their own hole. Point at a hole to drop its block in.

Remember

gRPC is a shared, strict order form plus small binary messages over HTTP/2, great between services.

WebSockets

A phone call that stays open, instead of asking are we there yet over and over

A normal web request is like calling, asking one question, and hanging up. If you want to know when a new chat message arrives, you have to keep calling back: anything new? anything new? This is called polling. It is like a kid in the back seat asking are we there yet every minute. Most answers are no, and every question costs a trip.

A WebSocket starts as a normal HTTP request that says: can we switch this into a phone call? If the server agrees, the same connection is upgraded into a long-lived, two-way line. Now either side can talk whenever it wants. The server can push a new message the moment it arrives, without being asked.

This is great for chat, multiplayer games, live sports scores and shared drawing boards, where small updates fly back and forth quickly. If only the server ever needs to talk, a simpler one-way stream called Server-Sent Events can be enough.

The costs: every open line holds a little memory on a server, and a popular app may hold millions of lines open at once. Lines drop when a phone loses signal, so apps need reconnect logic that tries again and catches up on anything missed. Load balancers must allow long connections and not hang up on quiet ones, so apps send tiny heartbeat messages (ping and pong) to show they are still there. And when friends are connected to different servers, a message must be passed between servers through a shared message hub.

Two cups joined by a string: a line that stays open. Move along it and the message bead runs either way, with no new call each time.

Remember

WebSockets keep one two-way line open so updates arrive right away, but every open line costs the server something.

API Gateway

The front desk of a big school

A big app is often split into many small services, like classrooms in a big school: one for user accounts, one for orders, one for videos. Visitors should not wander the halls looking for the right room. Instead, everyone comes in through one front desk. That front desk is the API gateway.

The front desk does the shared chores once, so each classroom does not have to. It checks badges (authentication, making sure you are who you say you are). It sends you to the right classroom based on what you asked for (routing). It makes sure no visitor comes too often (rate limiting). It can gather answers from several classrooms into one reply, so a phone does not have to make five separate trips. And it writes down who came and went (logging).

The trade-offs: every request makes one extra stop, which adds a little time. Since everyone goes through the front desk, if it breaks, nobody gets in, so gateways run as several copies behind a load balancer. It is also tempting to stuff too much logic into the front desk, which turns it into a slow, crowded bottleneck. Keep the gateway for shared chores, and leave the real work to the services.

One front desk stands before many doors. Everyone comes in here first. Point at a door and the gate swings to send you through it.

Remember

A gateway is one front door that handles badges, directions and limits for every service behind it.

Layer 4 vs Layer 7 load balancing

Reading only the envelope, or opening the letter

A load balancer is like the lunch line helper who sends each kid to the counter with the shortest line, so no single counter gets swamped. In computers, many copies of a server sit behind it, and it spreads requests among them. Common rules are taking turns (round robin) or picking the copy with the fewest people already waiting (least connections). It also stops sending anyone to a copy that has stopped answering its health checks.

The layer numbers come from a standard way of stacking network jobs into floors. A layer 4 load balancer only reads the outside of the envelope: the computer's address (the IP address) and the port, which is like an apartment number. It never opens the letter. That makes it very fast and simple, and it works for any kind of traffic, even locked (encrypted) traffic it cannot read. But it cannot tell what is being asked, so it cannot send video requests one way and shopping requests another.

A layer 7 load balancer opens the letter and reads the request itself: the path like /videos, the headers and the cookies. Now it can be clever. It can send /videos to the video team and /shop to the shop team, keep a user going to the same server using a cookie (sticky sessions), and retry a failed request somewhere else. It can also unlock the encryption itself, called ending TLS, so the servers behind it do not have to. The cost is more work for every single request, and it must hold the keys that unlock the traffic.

Many big systems use both: a layer 4 balancer at the very front to spread huge amounts of traffic cheaply, and layer 7 balancers behind it to make the smart choices.

Layer 4 reads only the envelope: address and stamp. Layer 7 opens it and reads the letter. Raise the pointer to lift the flap.

Remember

Layer 4 routes by address and port and is fast. Layer 7 reads the request and is smart.

Rate limiting algorithms

A hall monitor who stops anyone from asking too often

If one visitor asks a million questions a second, everyone else gets slow answers, and the servers might fall over. Rate limiting is a hall monitor with a rule like: each kid may ask 10 questions a minute. When someone goes over, the server answers with HTTP status 429, which means too many requests, often with a Retry-After note saying how long to wait. Good clients wait and try again, leaving longer gaps each time. Limits are counted per user, per API key or per IP address, and when many servers share one limit, they keep the counts in a fast shared store such as Redis.

Token bucket: picture a jar of tokens. A helper drops in one token at a steady pace, say one each second, until the jar is full at 5. Each request must take a token, and if the jar is empty, the request is turned away. A full jar lets a quick burst of 5 through at once, but over time nobody goes faster than the drip. Leaky bucket: requests pour into a bucket with a small hole in the bottom. They drip out to the server one by one at a steady pace. If too many pour in, the bucket overflows and the extras are turned away. Bursts become a calm, even flow, but requests may wait in line.

Fixed window counter: count requests in each clock minute, and reset the count to zero when a new minute starts. It is simple and needs only one number per user. The weakness is the edge. With a limit of 10, a kid can ask 10 times at 0:59 and 10 more at 1:00, which is 20 in about two seconds.

Sliding window log: write down the time of every request, and count only the ones from the last 60 seconds. It is exact, but remembering every time stamp uses a lot of memory for busy users. Sliding window counter: keep just two counts, last minute and this minute, and blend them. If we are 30 percent of the way into this minute, count 70 percent of last minute's requests plus all of this minute's. It assumes last minute's requests were spread evenly, so it is a close guess, not exact, but it needs very little memory.

  • Token bucket: allows short bursts, keeps a steady average.
  • Leaky bucket: turns bursts into an even flow.
  • Fixed window counter: simplest, but lets a double burst through at the edge.
  • Sliding window log: exact, but hungry for memory.
  • Sliding window counter: a good estimate with tiny memory.
A jar holds five tokens and a tap drips one back each second. Every ask takes a token. Hover on the jar to ask faster than the drip.

Remember

Rate limiting keeps one busy visitor from crowding out everyone else, and turned away requests get a 429 that says try again later.

Quick recap

  1. An API is a menu: it lists what you may ask for and exactly how to ask.
  2. HTTP/2 mixes many requests on one connection, but on TCP one lost piece stalls them all. HTTP/3 (QUIC) gives each stream its own lane, starts faster and survives network switches.
  3. gRPC uses a shared, strict order form (Protocol Buffers) and small binary messages over HTTP/2. Great between services, and browsers need gRPC-Web.
  4. WebSockets keep a two-way line open for live updates, at the cost of server memory and reconnect work.
  5. An API gateway is the front desk: badges, directions, limits and logs in one place. Run several copies so it never becomes the one thing that breaks.
  6. Layer 4 load balancers read the envelope (address and port). Layer 7 load balancers read the letter (path, headers, cookies).
  7. Token bucket, leaky bucket, fixed window, sliding window log and sliding window counter all say 429, try later, to anyone asking too often.

Grown-up words

and what they mean in plain words

API
A menu of things one program may ask another program for, and how to ask.
Multiplexing
Mixing pieces of many requests on one connection, each piece labeled with its stream.
Head-of-line blocking
When one stuck piece at the front makes everything behind it wait.
QUIC
Newer delivery rules built on UDP. It is the road under HTTP/3.
Protocol Buffers
A way to describe messages as a strict form and pack them into small binary.
WebSocket
A connection that stays open so both sides can talk at any time.
API gateway
The single front door that checks, routes and limits requests for many services.
Load balancer
A helper that spreads requests across many copies of a server.
TLS
The lock that keeps web traffic private. It is the s in https.
HTTP 429
The answer that means too many requests, please wait and try again.