"Let's add AI to this" is a sentence that skips past the actual engineering decision, which is: does this problem need a model that generalizes from data, or does it need a well-written set of rules? Those two paths diverge hard on architecture, ops burden, and failure modes. Here's how I actually think through it, with the tradeoffs made concrete.
Start with the decision function, not the tech stack
Before touching a stack, run the problem through this:
def needs_ml(problem):
if problem.has_clear_deterministic_rules and not problem.pattern_drifts_over_time:
return False # a rules engine will outperform ML here: cheaper, explainable, no training data needed
if problem.requires_generalizing_to_unseen_patterns:
return True # fraud detection, anomaly detection, anything adversarial
if problem.improves_meaningfully_with_more_data_over_time:
return True # personalization, recommendation, forecasting
return False # default to the simpler system until proven otherwise
This isn't pseudocode you'd actually ship, but writing the decision out this explicitly forces a real conversation instead of "AI sounds right for this."
Architecture A: Rules Engine
Client Request
-> API Gateway
-> Rules Service (deterministic logic, versioned rule set)
-> Database (direct query or cached lookup)
-> Response
Characteristics:
- Latency: low and predictable — no model inference step
- Explainability: total — every output traces to a specific rule
- Failure mode: a wrong rule is a bug; it's deterministic and reproducible
- Ops burden: version-control the rule set, standard CI/CD, no retraining pipeline
- Cost: low, scales like any standard CRUD service
This is the right architecture for: fixed business logic, compliance-driven decisions that need to be auditable, anything where a wrong output has to be traceable to an exact cause (this matters a lot in regulated domains — a "the model decided" answer doesn't satisfy an auditor the way "rule #14 triggered" does).
Architecture B: Real ML Pipeline
Client Request
-> API Gateway
-> Feature Store (real-time + batch features)
-> Model Serving Layer (versioned model, A/B-tested)
-> Post-processing / business rule overlay
-> Response
-> Logging -> Feedback Loop -> Retraining Pipeline -> Model Registry
Characteristics:
- Latency: inference adds real overhead — needs its own SLA and monitoring
- Explainability: partial at best — SHAP/LIME-style tooling helps but doesn't fully solve it
- Failure mode: silent degradation (model drift) is the dangerous one, not the loud crash
- Ops burden: feature store maintenance, retraining schedule, drift monitoring, model versioning, rollback strategy
- Cost: meaningfully higher — training infra, serving infra, and the team time to maintain both
This is the right architecture for: fraud/anomaly detection, recommendation systems that need to keep improving, forecasting problems with too many variables for hand-written rules to track.
The part most teams skip: monitoring for drift
A rules engine doesn't quietly get worse. A model does, and it's the single most common way "AI features" fail in production without anyone noticing for months.
def check_model_drift(predictions_window, baseline_distribution):
current_distribution = compute_distribution(predictions_window)
drift_score = compute_psi(baseline_distribution, current_distribution)
if drift_score > DRIFT_THRESHOLD:
alert_team("Model drift detected — accuracy likely degrading")
trigger_retraining_pipeline()
If this monitoring doesn't exist, the "AI feature" isn't production-ready regardless of how good the model was at launch. This is the piece that's almost always missing from a rushed "let's add AI" implementation, because it doesn't show up in a demo — it only shows up three months later as a slow, invisible accuracy decline.
A concrete example: fraud detection done both ways
Rules-based (works until it doesn't):
def flag_transaction(tx):
if tx.amount > THRESHOLD:
return True
if tx.country not in tx.user.usual_countries:
return True
return False
Fast to ship, fully explainable, and fraud patterns evolve past it within weeks because it can only catch what someone already thought to write a rule for.
ML-based (higher cost, but built to adapt):
def flag_transaction(tx):
features = feature_store.get_features(tx)
risk_score = fraud_model.predict(features)
if risk_score > MODEL_THRESHOLD:
return True
return False
This generalizes to new fraud patterns it wasn't explicitly told about — the entire reason to accept the added complexity. If the use case doesn't need that generalization, this is overengineering for the sake of a "we use AI" line in a pitch deck.
The actual takeaway
Neither architecture is "better" in the abstract. The decision function at the top is the whole point: pick based on whether the problem needs to generalize and adapt, not based on which one sounds more impressive in a product meeting. Most of the cost and maintenance burden of the ML path is invisible at demo time and shows up entirely in month three through six — which is exactly when teams that skipped this decision find out the hard way.
If you're scoping something like this and want a second opinion on which side of that line a specific feature falls on, happy to talk through it: 7Pillars – AI Application Development
Genuinely curious what others have seen: has anyone had a rules-based system get rebuilt into ML mid-flight because it hit a wall the rules couldn't handle? What was the tell that it needed to change?
Top comments (0)