Skip to the lesson
Little Builders system design, explained small

Real-Time Fraud Detection

Imagine a lunch line where every kid pays with a lunch card, and the helper at the register has to decide, in less time than a blink, whether the card is being used by its real owner. The helper cannot stop to investigate. So a team in the back room keeps a running list of clues for every card, like how many times it was used in the last few minutes and in which towns, and keeps those clues on a sticky note right next to the register. The helper glances at the note, asks a well trained judge for a quick opinion, and says yes, show me your code, or no. Weeks later, when the truth about some swipes comes out, the judge learns from its mistakes.

The real-time budget

Deciding before you finish blinking

When you tap a card, the whole approval, travelling from the shop to the bank and back, usually takes a few hundred milliseconds. A millisecond is a thousandth of a second, and a blink takes roughly 100 to 400 of them. The fraud check gets only a slice of that time, often a few tens of milliseconds.

That shapes everything. There is no time to dig through years of history or wait on slow services while the shopper stands at the till. All the heavy counting must already be done before the swipe arrives, so at decision time we only look up answers that are ready and waiting.

If the check runs out of time, the payment still needs an answer. So the system has a backup plan, such as using simple rules only, or a policy chosen in advance, like approving small amounts and declining risky ones. The trade-off is that the backup is less accurate, so it should be rare.

A stopwatch: once round is a whole card payment. The fraud check gets only the small slice at the top. Turn the hand inside it.

Remember

The fraud check gets tens of milliseconds, so the slow work must be done before the swipe arrives.

Transaction streams with Kafka

A conveyor belt of swipes, with one lane per card

Every swipe is written as a small event into Kafka, a system that keeps streams of events in order and lets many programs read them. Think of a conveyor belt that never drops anything and that several workers can watch at once.

A stream in Kafka is called a topic, and it is split into partitions, which are like separate lanes on the belt. We pick the lane by card number, so every swipe from the same card always goes into the same lane. Inside one lane, order is kept, so one card's swipes are always read in the order they were written.

Lanes let many computers work side by side, one lane each, without ever needing to share a card's story. The trade-off is that order only holds inside a lane, and a very busy card or shop can make one lane much busier than the others.

Remember

Split the stream by card number, so each card's swipes stay in order in one lane.

Sliding window aggregations with Flink

Counting swipes in the last 10 minutes, again and again

Flink is a stream processor, a program that reads events from the belt and keeps counting as they go by. A sliding window is a moving time frame. Every minute we ask, how many times was this card used in the last 10 minutes, and how much was spent? Because the frame moves forward one minute at a time, the windows overlap, and each swipe is counted in ten of them.

Other useful counts are how many different shops, or how many different countries, a card was used in during the last hour. A card used in three countries within one hour is a strong clue.

Flink keeps these running counts in memory for each card, called keyed state. Every so often it saves a snapshot, called a checkpoint. If a computer crashes, Flink loads the last snapshot and replays the events since then from Kafka, so each swipe affects the counts exactly once, never lost and never doubled.

Swipes do not always arrive in order. A shop's machine might lose its connection and send a swipe late. Flink uses event time, the time written on the swipe itself, not the time it arrived. It also uses a watermark, a marker that says, I do not expect anything older than this anymore. A window is closed only after the watermark passes its end. The trade-off: waiting longer catches more late swipes, but the counts arrive a little later.

Card swipes as beads along a rail of minutes. Slide the ten-minute frame along it: the swipes inside lift, and that is the count.

Remember

Sliding windows give fresh, overlapping counts per card, and checkpoints plus watermarks keep those counts correct.

Online and offline feature stores

A sticky note by the register, and a big diary in the back room

Features are the clues a model uses, such as swipes in the last 10 minutes or how much this card usually spends. A feature store is where we keep them, and it has two parts.

The online store is the sticky note by the register. It is a very fast key-value store, which works like a row of labeled cubbies: give it a card number and it hands back that card's latest clues, usually within a few milliseconds. The streaming jobs keep it up to date as swipes arrive.

The offline store is the big diary in the back room. It keeps the full history of every clue over time, which is what we need to train new models. It is large and cheap, but far too slow to use during a swipe.

Both stores must use the same clue recipes. If training counted swipes one way and the live system counted them another way, the model would practice one game and then be asked to play a different one. This mismatch is called training-serving skew. Training data must also be point-in-time correct: for each past swipe, we use the clues exactly as they were at that moment, never peeking at what happened later. Otherwise the model looks smarter in tests than it really is.

Fresh clues sit on sticky notes in a small drawer by the register, quick to grab. The full history fills a big slow cabinet behind.

Remember

Fresh clues live in the fast online store, history lives in the offline store, and both come from the same recipes.

Low-latency model scoring

A quick look at the clues, a quick opinion, a quick answer

When a swipe arrives, the payment service asks the scoring service, is this OK? The scoring service fetches the card's clues from the online store, then adds a few clues it can only work out right now, such as how this amount compares with the card's usual amount.

Then it runs the model, a program trained on millions of past swipes, which returns a risk score, like a number between 0 and 1. Rules sit beside the model for things we always want, such as blocking cards that were reported stolen. Together they give one of three answers: approve, a step-up challenge that asks the person to prove it is them, for example with a code sent to their phone, or decline.

Every call has a strict timer. If the online store or the model is slow, the service does not wait. It falls back to rules only, or to the policy chosen in advance. Every decision is logged together with the clues that were used, so we can later check how good the decisions were.

The trade-off is between catching thieves and annoying honest customers. Set the bar too low and fraud slips through. Set it too high and real people get declined at the shop. The challenge answer is the middle path for unclear cases.

Remember

Fetch ready clues, add a few fresh ones, score with the model and rules, and always answer before the timer runs out.

Fraud labels and retraining

Learning from the swipes we got wrong

For many swipes, we only learn the truth later. A cardholder may spot a strange charge weeks afterwards and dispute it, which can lead to a chargeback, where the money is pulled back. Confirmed fraud reports like these are the labels that tell us which swipes were bad, and they arrive days or weeks after the swipe.

Those labels are joined with the logged decisions and the point-in-time clues from the offline store, and that becomes the training data for the next model. The new model is tested against the current one before it is trusted with real payments.

The trade-off is delay. Because labels arrive late, the newest fraud tricks show up in the training data only after a while. Fraudsters also change their tricks once they get blocked, so models are retrained regularly, and quick rules fill the gap in between.

Remember

Late fraud labels flow back into training, so the judge keeps learning new tricks.

Quick recap

  1. The fraud check gets tens of milliseconds inside a card payment that takes a few hundred.
  2. Kafka keeps swipes in lanes chosen by card number, so each card's events stay in order.
  3. Flink counts swipes in overlapping sliding windows, with checkpoints for safe recovery and watermarks for late events.
  4. The online feature store answers in milliseconds. The offline store keeps history for training. Both use the same recipes.
  5. Scoring combines stored clues, fresh clues, a model and rules into approve, challenge or decline, with a fallback if time runs out.
  6. Fraud labels arrive weeks later and become point-in-time correct training data for the next model.

Grown-up words

and what they mean in plain words

Kafka
A system that stores streams of events in order and lets many programs read them.
Partition
One lane of a stream. Events with the same key, like a card number, always go to the same lane.
Flink
A stream processor that keeps running counts as events flow by.
Sliding window
A time frame of fixed length that moves forward in small steps, so the frames overlap.
Watermark
A marker that says events older than this are no longer expected, so a window can close.
Checkpoint
A saved snapshot of a stream job's counts, used to recover after a crash.
Feature
A clue the model uses, like the number of swipes in the last 10 minutes.
Feature store
Where features are kept: a fast online part for decisions and a big offline part for training.
Training-serving skew
When clues are computed differently in training and in real use, so the model gets confused.
Chargeback
When a disputed card charge is pulled back from the shop. Fraud chargebacks are late labels that a swipe was bad.