DEV Community

Sundeep Mann
Sundeep Mann

Posted on Originally published at 7pillars.com.au

When to Actually Build ML vs. When a Rules Engine Is Fine (With Architecture)

"Let's add AI to this" is a sentence that skips past the actual engineering decision, which is: does this problem need a model that generalizes from data, or does it need a well-written set of rules? Those two paths diverge hard on architecture, ops burden, and failure modes. Here's how I actually think through it, with the tradeoffs made concrete.

Start with the decision function, not the tech stack

Before touching a stack, run the problem through this:

def needs_ml(problem):
    if problem.has_clear_deterministic_rules and not problem.pattern_drifts_over_time:
        return False  # a rules engine will outperform ML here: cheaper, explainable, no training data needed

    if problem.requires_generalizing_to_unseen_patterns:
        return True  # fraud detection, anomaly detection, anything adversarial

    if problem.improves_meaningfully_with_more_data_over_time:
        return True  # personalization, recommendation, forecasting

    return False  # default to the simpler system until proven otherwise
Enter fullscreen mode Exit fullscreen mode

This isn't pseudocode you'd actually ship, but writing the decision out this explicitly forces a real conversation instead of "AI sounds right for this."

Architecture A: Rules Engine

Client Request
    -> API Gateway
    -> Rules Service (deterministic logic, versioned rule set)
    -> Database (direct query or cached lookup)
    -> Response
Enter fullscreen mode Exit fullscreen mode

Characteristics:

  • Latency: low and predictable — no model inference step
  • Explainability: total — every output traces to a specific rule
  • Failure mode: a wrong rule is a bug; it's deterministic and reproducible
  • Ops burden: version-control the rule set, standard CI/CD, no retraining pipeline
  • Cost: low, scales like any standard CRUD service

This is the right architecture for: fixed business logic, compliance-driven decisions that need to be auditable, anything where a wrong output has to be traceable to an exact cause (this matters a lot in regulated domains — a "the model decided" answer doesn't satisfy an auditor the way "rule #14 triggered" does).

Architecture B: Real ML Pipeline

Client Request
    -> API Gateway
    -> Feature Store (real-time + batch features)
    -> Model Serving Layer (versioned model, A/B-tested)
    -> Post-processing / business rule overlay
    -> Response
    -> Logging -> Feedback Loop -> Retraining Pipeline -> Model Registry
Enter fullscreen mode Exit fullscreen mode

Characteristics:

  • Latency: inference adds real overhead — needs its own SLA and monitoring
  • Explainability: partial at best — SHAP/LIME-style tooling helps but doesn't fully solve it
  • Failure mode: silent degradation (model drift) is the dangerous one, not the loud crash
  • Ops burden: feature store maintenance, retraining schedule, drift monitoring, model versioning, rollback strategy
  • Cost: meaningfully higher — training infra, serving infra, and the team time to maintain both

This is the right architecture for: fraud/anomaly detection, recommendation systems that need to keep improving, forecasting problems with too many variables for hand-written rules to track.

The part most teams skip: monitoring for drift

A rules engine doesn't quietly get worse. A model does, and it's the single most common way "AI features" fail in production without anyone noticing for months.

def check_model_drift(predictions_window, baseline_distribution):
    current_distribution = compute_distribution(predictions_window)
    drift_score = compute_psi(baseline_distribution, current_distribution)

    if drift_score > DRIFT_THRESHOLD:
        alert_team("Model drift detected — accuracy likely degrading")
        trigger_retraining_pipeline()
Enter fullscreen mode Exit fullscreen mode

If this monitoring doesn't exist, the "AI feature" isn't production-ready regardless of how good the model was at launch. This is the piece that's almost always missing from a rushed "let's add AI" implementation, because it doesn't show up in a demo — it only shows up three months later as a slow, invisible accuracy decline.

A concrete example: fraud detection done both ways

Rules-based (works until it doesn't):

def flag_transaction(tx):
    if tx.amount > THRESHOLD:
        return True
    if tx.country not in tx.user.usual_countries:
        return True
    return False
Enter fullscreen mode Exit fullscreen mode

Fast to ship, fully explainable, and fraud patterns evolve past it within weeks because it can only catch what someone already thought to write a rule for.

ML-based (higher cost, but built to adapt):

def flag_transaction(tx):
    features = feature_store.get_features(tx)
    risk_score = fraud_model.predict(features)
    if risk_score > MODEL_THRESHOLD:
        return True
    return False
Enter fullscreen mode Exit fullscreen mode

This generalizes to new fraud patterns it wasn't explicitly told about — the entire reason to accept the added complexity. If the use case doesn't need that generalization, this is overengineering for the sake of a "we use AI" line in a pitch deck.

The actual takeaway

Neither architecture is "better" in the abstract. The decision function at the top is the whole point: pick based on whether the problem needs to generalize and adapt, not based on which one sounds more impressive in a product meeting. Most of the cost and maintenance burden of the ML path is invisible at demo time and shows up entirely in month three through six — which is exactly when teams that skipped this decision find out the hard way.


If you're scoping something like this and want a second opinion on which side of that line a specific feature falls on, happy to talk through it: 7Pillars – AI Application Development

Genuinely curious what others have seen: has anyone had a rules-based system get rebuilt into ML mid-flight because it hit a wall the rules couldn't handle? What was the tell that it needed to change?

Top comments (0)