DEV Community

Cover image for Fraud Detection: Why Your Rules Engine Has a Ceiling
Emmanuel R for CobuildX AI

Posted on Originally published at cobuildx.ai

Fraud Detection: Why Your Rules Engine Has a Ceiling

Rules-based fraud detection is fast to deploy and easy to explain. It also has a ceiling: fraudsters adapt to rules, false positive rates are hard to reduce, and new fraud patterns take time to codify. ML adds a measurable lift above that ceiling — here is what that looks like in practice.

Most financial institutions run fraud detection on a rules engine. Block transactions over a certain dollar amount to unusual recipients. Flag velocity patterns above threshold. Decline cards used in multiple geographies within a short window. These rules work. They catch a significant fraction of fraud and they have the advantage of being fast to implement, easy to explain to regulators, and transparent to the operations team. They also have a ceiling that most institutions have already hit.

A rules engine is a static representation of known fraud patterns. Fraudsters adapt to it. Once a particular attack pattern triggers enough rule-based blocks, the fraud operation modifies its behaviour to stay below the thresholds — lower transaction amounts, slower velocity, more geographically plausible patterns. The rules that were catching fraud last year catch less of it this year, and adding new rules to close the gap increases false positives on legitimate transactions.

Key insight: Rules catch the fraud patterns you have already seen and codified. ML catches the patterns that are statistically anomalous but have not yet been written into a rule — which is where new fraud appears.

"Every time we add a new rule to close a fraud vector, our false positive rate goes up and our customer service team starts getting calls about blocked legitimate transactions."

What Rules-Based Detection Does Well

Rules-based fraud detection has genuine strengths that explain why it remains the foundation of most fraud stacks.

It is deterministic and auditable. When a transaction is blocked, the operations team can tell the customer exactly which rule triggered the decline. This transparency matters for customer relations and for regulatory examinations.

It is fast. Rule evaluation is computationally trivial — microseconds per transaction. At the transaction volumes that large financial institutions process, latency matters.

It handles known, well-defined attack patterns reliably. Card-not-present fraud with high velocity on new accounts, account takeover with credential stuffing patterns, check kiting with specific inter-account timing signatures — these patterns can be codified into rules that catch them with high precision.

Rules are also the right tool for hard business constraints that are not about statistical patterns: geographic restrictions, product eligibility requirements, regulatory holds. These belong in a rules engine regardless of how sophisticated the ML layer is.

Rules are the right tool for hard business constraints and well-defined known patterns — they are not the right tool for detecting fraud patterns that have not been codified yet

Where the Ceiling Is

The ceiling on rules-based detection shows up in three ways.

First, new fraud patterns. When a new attack vector appears — a new synthetic identity structure, a new card-not-present exploit, a new account funding scheme — it takes time to detect the pattern, analyse it, write the rule, test it, and deploy it. During that window, fraud losses accumulate. The detection gap between when new fraud appears and when a rule catches it is typically weeks to months.

Second, the precision-recall tradeoff. Adding rules to catch more fraud always increases false positives on legitimate transactions. At some point, the marginal fraud caught by the next rule is less costly than the customer friction and manual review cost of the additional false positives it creates. Most mature rules engines are already at or past this point.

Third, adversarial adaptation. Fraud rings that are large enough to observe their own block rates modify their behaviour to avoid triggered rules. This is well-documented in card fraud, account opening fraud, and money movement fraud. Rules that were effective eighteen months ago are less effective today because the fraud behaviour has changed to avoid them.

New patterns, false positive ceilings, and adversarial adaptation are the three vectors where rules-based detection consistently underperforms

What ML Adds

ML fraud detection models — gradient boosting classifiers, neural networks, graph-based anomaly detection — catch fraud through a different mechanism than rules. They learn the statistical profile of legitimate transactions and flag deviations from that profile, without requiring those deviations to match a specific predefined pattern.

This means they can catch fraud that has not been codified into a rule yet. A new synthetic identity scheme may not trigger any existing rule, but it will deviate from the statistical profile of legitimate new accounts in ways that a well-trained model detects. The fraudsters would need to perfectly mimic the full behavioural and demographic distribution of legitimate customers — not just avoid specific thresholds — to evade a model-based detector.

The measurable lift of adding ML on top of rules varies significantly by institution and fraud type. In card fraud, lift of 15–30% reduction in fraud losses at equivalent false positive rates is commonly reported. In account opening fraud and synthetic identity fraud, the lift can be larger because these fraud types are harder to codify into rules.

ML does not replace the rules engine. The two work better together than either does alone. Hard constraints belong in rules. Statistical anomaly detection belongs in the ML layer. Most production fraud stacks run both.

ML adds 15–30% fraud loss reduction at equivalent false positive rates by catching statistical anomalies that rules cannot codify — it works alongside rules, not instead of them

Model Drift in an Adversarial Environment

Fraud detection models face a challenge that most ML models do not: the distribution they are trained on is actively manipulated by the adversary once the model is deployed.

A churn prediction model trained on last year's customers is less accurate this year because customer behaviour naturally evolved. A fraud model trained on last year's fraud patterns is less accurate this year partly because fraud behaviour naturally evolved, and partly because fraud operators specifically adapted to avoid the model's detection patterns.

This means fraud models need more frequent retraining than most ML applications — quarterly at minimum for high-volume card fraud, more frequently if new attack vectors are appearing. It also means the training data pipeline needs to be maintained carefully: new labelled fraud examples need to be incorporated quickly, and the labelling process (which requires investigations and chargebacks to confirm fraud) needs to be timely enough to produce useful training signal.

Monitoring for model drift in fraud detection should track both the false positive rate (to catch model degradation on legitimate transactions) and the fraud detection rate (to catch model degradation on fraud). Both can degrade independently if the fraud distribution shifts.

Fraud models require more frequent retraining than most ML applications because the adversary actively adapts — quarterly retraining is a minimum, not a target


Originally published on the CobuildX blog.

Top comments (0)