A walk through the real-time fraud detection stack we rebuilt with a US mid-size bank: Kafka ingest, Flink features, a three-model ensemble, the analyst’s alert queue, which model does which job, and why the card decline you hate at the grocery store still happens.
Every fraud vendor who walks into a bank says the same thing: AI will solve your fraud problem. I spent 16 months on a fraud-platform rebuild with a mid-size US bank (4.2 million retail accounts, about 480,000 small-business accounts) whose head of fraud operations had heard that pitch a dozen times and did not believe a word of it. The brief was simple. Build the thing, then write down what AI does and what two humans and a phone still do. This is the first half, the architecture. How a card swipe becomes a fraud score and a decision in under 80 milliseconds, which models sit where, and what holds them together. No magic. Part 2 covers where it broke.

By Shri Director of Technology at GetSetLive and Bagful Cloud Hosting, with 25 years across infrastructure, datacenters, cloud platforms, and modern AI/ML. Day-to-day work focuses on consulting clients to make their existing systems AI-aware and AI-ready, and architecting AI/ML solutions tailored to each customer environment. Every article here comes from direct field exposure, written to help the next batch of professionals design and deploy more sophisticated solutions in their own careers. Read all posts by Shri →
Editor, GetSetLive Technical Blogs17 min read
Key Takeaways
- AI does not stop fraud. It cut the bank’s alert haystack from 18,000 a day to about 400, and rules plus analysts decide what happens to those 400.
- The stack is an ensemble of three narrow models (XGBoost, an autoencoder, GraphSAGE) behind a rules engine, not one big model; median swipe-to-decision is 72 ms.
- Step-up authentication, the “confirm this purchase” push, only exists because the model outputs a probability. It is the single biggest customer-experience win.
- The feedback loop from analyst labels back into monthly retraining is the part vendors never demo, and the part that keeps the model alive.
What Is Inside
- The sentence in every vendor deck
- What the bank ran before the rebuild
- What the model watches that rules could not
- Anatomy of the real-time pipeline
- Which model does which job
- One card swipe, end to end
- A word on SARs and AML
- What comes next
The Sentence in Every Vendor Deck
“AI stops fraud.” I have seen that line on roughly 70% of the fraud vendor slides I have sat through, and it is wrong in a specific, interesting way. AI does not stop fraud. What it did at the bank was shrink the alert haystack from something a human team cannot review (18,000 alerts a day before the rebuild) to something they can (around 400 a day now). Three things together stop the fraud itself: the model’s score, a rule that decides what to do with that score, and, in a small but critical slice of cases, a phone call from an analyst to the customer.
On my first day the head of fraud operations put it to me in a way I have quoted many times since. A model is a filter. It does not decide anything. The decision is a rule: if the score is above 0.83 and the merchant category is one of 14 high-risk ones, block the transaction and send a push notification. The model knows none of that. It returns a number. Everything around the number is policy, and policy is a person figuring out the bank’s risk appetite for the quarter. I have heard a version of that speech from every fraud operations lead I have worked with, and it is always the truest thing they say all day.
What the Bank Ran Before the Rebuild
Until late 2023 the bank ran fraud detection the way most regional banks still do. A rules engine, written in a mix of old SAS code and a newer Drools layer, scored every transaction against roughly 240 rules. The rules were the classics. A transaction above 500 dollars in a country the customer has never transacted in. Three card-not-present transactions inside four minutes. A velocity breach on a specific merchant category code. Each rule produced a yes or no. If any rule fired, the transaction was flagged; if enough fired, it was blocked. The fraud team wrote those rules over a decade, almost always after an incident. Someone got defrauded, a rule went in. Someone complained, a rule got tuned.
The system caught fraud. It also produced around 18,000 alerts per business day, and about 12,000 of those needed an analyst. Fourteen analysts cannot review 12,000 alerts a day, so in practice they reviewed none of them properly. The floor had 14 analysts on two shifts, so each analyst was expected to dispose of roughly 850 alerts per shift. That is 35 an hour, under 90 seconds each. Nobody investigates fraud in 90 seconds. What happened instead, and I watched it during our discovery weeks, is that analysts built reflexes. If it looks like the 300 alerts already closed today that turned out to be nothing, close it. The false positive rate on blocked transactions sat around 68%. Real customers got declined at the grocery store while a script somewhere burned through stolen card numbers unseen.
Below is the pre-AI stack as we mapped it in the first month. The red stops are where volume, accuracy, or plain blindness was costing the bank money and customers.
Free to use, share it in your presentations, blogs, or learning materials.

The pre-AI fraud stack at the bank. Red stops mark where volume, accuracy, or blind spots cost money and customers.
The failures were structural. Better rules would not have fixed them. A rule cannot see beyond a single transaction. There was no memory of “this account has been quiet for 14 months and just lit up”, because that needs a windowed feature across time and the rules engine had no compute for it. There was no view of mule accounts, because mule detection needs a graph and rules do not do graphs. Encrypted mobile banking sessions were a black box. First-party fraud, where the customer is the fraudster, is about 11% of losses and was almost invisible, because every rule assumed the customer was the victim.
And then there were the SARs. A Suspicious Activity Report is a regulatory filing a bank must make when it suspects money laundering. Each one took an analyst two to three hours to write. A senior analyst might finish two in a day. That number came up in every planning meeting we had, and it turned out to be the easiest thing in the whole programme to fix.
What the Model Watches That Rules Could Not
The rebuild took the bank’s data engineering team and us 16 months, four months longer than the plan. Twelve of those months went into data plumbing, not models. On day two I asked the data engineering lead to name the single thing a model does that a rule cannot. No hesitation: context. A rule sees one transaction on its own. A model sees it against the last 10 minutes, the last 30 days, and the links that account has to other accounts. Same data, joined differently, and the joining is the whole game. Rules are a lookup table. Models are a function that takes a vector.
Here is what the model now watches, roughly in order of how much each one moved our numbers:
- Windowed transaction features. Count in the last 10 minutes, average amount over 30 days, distinct merchant categories in 24 hours, ratio of the current amount to the 90-day average. Flink computes all of these continuously from the Kafka stream.
- Device fingerprint and behavioural biometrics. Typing cadence on mobile logins, swipe pressure, scroll pattern, and the JA4 TLS fingerprint of the app session, through a BioCatch integration. These tell the real account holder apart from someone who merely has the password.
- Graph proximity. Is this account within two hops of a known mule? Shared device, shared phone number, shared beneficiary? The graph lives in Neo4j and a PyTorch Geometric GraphSAGE model queries it.
- Merchant embedding similarity. Every merchant the customer has used gets a 64-dimensional embedding. A transaction at a merchant cluster unlike anything this customer has touched is a weak signal on its own, and a useful one in combination.
- Cross-channel signals. A physical ATM withdrawal in Chicago at 14:32 while the mobile app is active in Atlanta at 14:30 is a red flag even when each event alone looks fine.
None of these are rules. They are features. A rule says “block if X”. A feature is an input to a model that says “the probability this is fraud is 0.73, and here are the six inputs that drove it”. The difference matters because a feature-based system combines weak signals into a strong one. Three weak anomalies on a rules engine are three rules that did not fire. Three weak anomalies on a gradient-boosted tree are a 0.91 fraud score. That, more than any single model choice, is what the 16 months bought.
Anatomy of the Real-Time Pipeline
The architecture drawing on the data engineering team’s wall covers most of the wall. The diagram below is the simplified version. It reads left to right: events in on the far left, a decision out on the far right. Median time from swipe to decision is 72 milliseconds, with a 99th percentile of 118. We cannot afford to be slower, because the card networks give an issuer roughly two seconds for the whole authorisation and Visa and Mastercard use a good chunk of that themselves.
Free to use, share it in your presentations, blogs, or learning materials.

Six lanes, thirteen components. Median card-swipe-to-decision time is 72 milliseconds.
Lane 1: Event Ingest
Everything the model will ever see enters through Apache Kafka. The bank runs a Confluent Cloud cluster with five topics: card-auth (Visa and Mastercard authorisation messages), ach-events (ACH pushes and pulls), wire-events (Fedwire and SWIFT), mobile-sessions (app login and in-app behaviour), and atm-events. Peak is about 14,000 events per second, and Black Friday 2024 pushed it to 19,000 for six hours, which is the day we learned our partition count was two too low. Every event has a schema in Confluent Schema Registry. Each downstream consumer knows exactly which fields to expect, and a producer cannot quietly change one.
Lane 2: Stream Processing and Feature Engineering
Apache Flink is the workhorse. A Flink job enriches every event on Kafka with windowed aggregates computed on the fly: events in the last 10 minutes, card-not-present amount in the last 24 hours, distinct IP addresses in the last 7 days, and about 80 other features held in per-account keyed state. The state lives in RocksDB on the Flink task managers, so lookups are local and fast. When a transaction arrives, Flink computes the features and attaches them to the event. It then writes the enriched record to a Kafka topic for model serving and to the online feature store.
This lane is where our worst early incident lived. In month seven a checkpoint failure during a deploy replayed 40 minutes of the card-auth topic, and for about 25 minutes every “count in last 10 minutes” feature was roughly double. Scores drifted up. The step-up rate tripled. The care queue noticed before our dashboards did. Our fix was idempotent state updates keyed on the authorisation ID, plus an alert on feature distribution shift, which Part 2 comes back to. Nothing about it was a modelling problem.
Lane 3: Feature Store
The bank uses Tecton, the commercial feature store built by people from the Uber Michelangelo team. Its job is the one every feature store has: the features the model learned on must be the exact same features available at serving time. Online lookups hit Redis in under a millisecond. Historical features go to Snowflake for training. Every feature definition is versioned, so a feature can change or retire without breaking the model in silence. This sounds boring. It is boring. It is also the thing that separates a working ML system from a flaky demo, and I would pick it over any model upgrade.
Lane 4: Model Inference
This is where the models live, and there are three of them in the transaction path. The primary is an XGBoost classifier that takes around 240 features and outputs a fraud probability between 0 and 1. Beside it runs an autoencoder for unsupervised anomaly detection, and a GraphSAGE graph neural network that scores how close the transaction sits to known fraud in the account graph. A small logistic meta-model combines the three scores into one number. The ensemble runs inside a FastAPI container per region, with TorchServe handling the GraphSAGE inference and XGBoost serving itself. Median inference latency across all three is 9 milliseconds. Twelve Kubernetes pods sit behind an internal load balancer and scale on request rate.
We argued for a month about whether the autoencoder earned its place. It adds latency and it is hard to explain to a regulator. It stayed because in the shadow period it flagged a card-testing pattern (hundreds of one-dollar authorisations across fresh merchant IDs) two weeks before the supervised model had any labelled examples of it. That is the whole case for an unsupervised model in the ensemble: it sees shapes before anyone has named them.
Lane 5: Decision Engine
The ensemble score, plus a dozen hard rules that still exist (sanctions list hit, card reported stolen, customer on fraud hold), flow into a Drools decision engine. Drools produces one of five outcomes: allow, allow with customer notification, step-up authentication (an OTP challenge or a biometric check in the app), block and notify, or block and lock the card. The fraud operations team sets the thresholds between those outcomes and reviews them monthly; the model has no say in them. When the retail fraud loss ratio drifts above 0.04% of swipe volume, thresholds tighten. When complaints about declines spike, they loosen. It is a permanent trade between fraud loss and customer experience. Nobody solves it; they manage it.
Lane 6: Case Management and Feedback
Alerts that need a human land in Actimize, the case management system. Each case carries the transaction, the features, the top 10 SHAP values from XGBoost, the graph neighbourhood, and a one-paragraph summary written by an LLM. Analysts close the case with one of six labels, including “confirmed fraud” and “confirmed legitimate”. Those labels flow back through Kafka into the training set. The monthly retrain on the last 90 days of labelled data then learns from last month’s analyst decisions. This feedback loop is the most important piece of the whole system. Without it the model goes stale within weeks, because the people on the other side adapt faster than any vendor’s release cycle.
Which Model Does Which Job
Seven distinct models run across the bank’s fraud and AML stack, three of them in the real-time path. Every choice was deliberate. The data engineering lead had to justify each one to the model risk committee in writing, and that exercise killed two models we liked. The diagram shows the mapping; the table underneath is the short version of those justifications.
Free to use, share it in your presentations, blogs, or learning materials.

Seven jobs, seven model families. Each matched to the shape of the problem, not picked for novelty.
| Function | Model | Why this one |
|---|---|---|
| Primary fraud score at transaction time | XGBoost | Fast tabular inference, handles missing values, SHAP-explainable for regulator review |
| Unsupervised anomaly (zero-day fraud) | Autoencoder | No labels needed; catches patterns the supervised model has not seen yet |
| Mule ring and collusion detection | GraphSAGE (GNN) | Tabular models cannot see account-to-account relationships; the GNN traverses the neighbourhood |
| Behavioural biometrics on mobile | LSTM on keystroke and swipe sequences | Biometrics are time series; the LSTM captures cadence |
| Device fingerprinting | Random Forest on JA4/TLS features | Hard to evade, interpretable, cheap at scale |
| Synthetic identity detection | Entity resolution + graph embeddings | Synthetic identities are stitched across sources; the graph exposes the stitching |
| SAR narrative drafting | Fine-tuned LLM (human-reviewed) | Only model that writes regulator-grade prose; kept out of every automatic decision |
Nothing in that table is a foundation model for fraud. There is no GPT-fraud. What there is, is a set of small specialist models, each picked because its assumptions match the shape of one problem. I have now seen the same pattern at a lender, an insurer, and a bank: narrow models do narrow jobs well, and one big model does nothing well.
One Card Swipe, End to End
The best way to feel how the pipeline holds up is to follow one transaction. The example below comes from the bank’s staging replay logs, altered to protect the customer but otherwise real. A 41-year-old debit card holder who lives near the bank’s home city is at an electronics store in Charlotte on a Saturday afternoon, buying a 1,200 dollar laptop. That is unusual for this customer on three counts: never shopped in Charlotte, never at that retailer, and the amount is three standard deviations above his 90-day average.
- T+0 ms: The store terminal swipes the card. The authorisation request goes to Mastercard, which forwards it to the bank over ISO 8583.
- T+4 ms: The bank’s payment gateway normalises the authorisation into a JSON event and writes it to the card-auth topic.
- T+9 ms: Flink picks up the event and pulls the customer’s 80 windowed features from Tecton. Transactions in the last 10 minutes: 0. Thirty-day average amount: 47 dollars. Distinct cities in 30 days: 1. Features enriched and attached.
- T+14 ms: The enriched event lands on the inference topic and a FastAPI worker reads it.
- T+23 ms: XGBoost returns 0.71 (somewhat suspicious). The autoencoder returns 0.68 (high reconstruction error). GraphSAGE returns 0.12 (no fraud-graph proximity). The meta-model combines them to 0.69.
- T+28 ms: A decision rule fires: score between 0.60 and 0.80, plus high amount, plus new merchant city, equals step-up authentication. Response: challenge with a mobile push.
- T+32 ms: The bank returns a conditional response to Mastercard: approve with 3-D Secure step-up.
- T+600 ms: The customer’s phone buzzes: “Confirm $1,200 at an electronics store in Charlotte?” He taps yes; Face ID confirms.
- T+1,840 ms: The bank sends final approval to Mastercard. The terminal beeps. He walks out with the laptop.
The inference took 9 milliseconds and the decision was out 28 milliseconds after the swipe hit the gateway. The remaining 1.8 seconds was a person reacting to a push notification. Under the old rules engine this purchase would have been either fully approved (no rule caught “first time in Charlotte”) or fully blocked (a velocity rule might have fired on the amount). The middle path, where the customer is asked rather than refused, did not exist in the rules world. The model’s probability is what unlocks it.
Step-up authentication is the unsung hero of modern fraud detection. It is invisible to the customer who is real and fatal to the attacker who is not. On the bank’s numbers it now resolves about 62% of ambiguous transactions without an analyst and without a decline, which is the single line on the board deck that made the whole programme worth it.
A Word on SARs and AML
The card pipeline is the fast, real-time side. The anti-money-laundering side is slower and more regulated. In the US a Suspicious Activity Report has to be filed with FinCEN within 30 days of detection. Before the rebuild the bank filed around 3,800 SARs a year. Each one took a senior analyst two to three hours to draft. We added a fine-tuned LLM that reads the case file, the relevant transaction history, and the analyst’s preliminary notes, and produces a first draft of the narrative. An analyst reads it, corrects it, and signs it. That draft saves about 90 minutes per SAR. Today the bank files around 4,600 a year (detection got better, not worse) and the per-SAR effort is down to roughly 45 minutes of senior analyst time.
The LLM is explicitly not in the decision path. It does not decide whether a SAR gets filed. It writes a draft; a human decides. Every SAR is signed by a named compliance officer who carries personal regulatory exposure if it is wrong. That is the pattern FinCEN and the OCC have said they expect, and it is the only pattern our compliance head would sign off on. For the first month the compliance head read every single draft against the source case before trusting a paragraph of it, and I think that was the right call.
What Comes Next
The architecture is the mechanical half. The more interesting half, honestly, is what happens when this stack fails in production, and it does. The bank had three meaningful incidents in the year after go-live, each of which taught the fraud team something about how AI detection breaks at scale. Part 2 covers the named case studies that put real numbers on the scale of the problem (DBS, JPMorgan, HSBC, Danske Bank, Standard Chartered), the brutal arithmetic of false-positive economics on a 0.04% base rate, what SR 11-7 and the EU AI Act mean for a fraud team, and what the senior analyst does on the cases the model refuses to handle. Part 2 is here.
Related Reading
- What AI Actually Does in Loan Underwriting (Part 1): The Architecture: the same lens on an Indian NBFC’s underwriting stack, where the models are similar and the plumbing is different.
- What AI Actually Does in Loan Underwriting (Part 2): Where It Breaks and Who Still Signs Off: thin files, festival drift, and proxy discrimination, the lending versions of the failures Part 2 covers for banking.
- 7 Times AI Gave the Wrong Answer (with Proof): concrete examples of AI confidently getting it wrong, which matters when the model is touching money.
References
- Kai Waehner, Fraud Detection with Apache Kafka, KSQL and Apache Flink
- NVIDIA, Supercharging Fraud Detection with Graph Neural Networks
- Fintech Global, How AI Is Reshaping AML Compliance in Modern Banking, 2026
- Conduktor, Real-Time Fraud Detection with Streaming
- Ampcome, Agentic AI Examples in Banking: Compliance & Risk, 2026
Frequently Asked Questions
Does AI actually stop bank fraud?
Not on its own. The models produce a fraud probability for each transaction. Rules set by the fraud operations team decide what to do with that score: allow, challenge with step-up authentication, or block. The model narrows the field; rules and humans make the call.
Which machine learning model is used for real-time fraud detection?
Most banks run an ensemble. XGBoost is the primary supervised classifier on tabular features, an autoencoder catches unsupervised anomalies, and a graph neural network such as GraphSAGE catches mule rings. A small meta-model combines the three scores, and the whole ensemble answers in under 15 milliseconds.
What is step-up authentication in fraud detection?
The middle-ground response between approve and decline. When the fraud score is ambiguous (roughly 0.60 to 0.80), the bank asks the customer to confirm the transaction through a mobile push, OTP, or biometric check. It catches fraud without declining legitimate purchases, and it only works because the model produces a probability rather than a yes or no.
How long does an AI fraud decision take?
Model inference runs in 5 to 15 milliseconds. End to end, from card swipe to bank response, the median is 70 to 90 milliseconds. Card networks allow about two seconds for the whole authorisation, so there is headroom, but the pipeline is tuned hard because latency compounds across hops.
What is a feature store and why does a fraud system need one?
A feature store guarantees that the features a model was trained on are computed the same way at serving time. Without one, training and production drift apart silently and the model scores wrong with full confidence. The bank in this article uses Tecton with Redis for online lookups and Snowflake for training data.
Why keep an autoencoder if XGBoost is more accurate?
Because the autoencoder needs no labels. It flags transaction shapes the supervised model has never been trained on, such as a new card-testing pattern, weeks before enough labelled examples exist to retrain XGBoost. It costs a few milliseconds and some explainability, and it earns its place on zero-day fraud.
Can AI replace fraud analysts?
No. AI cut the daily alert volume from a number no team can review to a number that fits in a shift. The analysts who remain handle the hard cases, phone customers, confirm fraud rings, and draft Suspicious Activity Reports. AI does the bulk filtering; humans do the judgement and carry the regulatory exposure.
The post What AI Actually Does in Bank Fraud Detection (Part 1): The Architecture appeared first on GetSetLive Blogs.
Top comments (0)