DEV Community

RobustTrueTry
RobustTrueTry

Posted on

Why Your Hand‑Coded Decision Model Crashes on Missing Values

When you build a decision model by hand, a single missing value can turn the whole prediction into a crash or a biased guess. This happens because naïve if‑else trees assume every feature is present and try to compare it directly with a threshold.

What you'll learn:

  • Why missing data breaks naïve if‑else trees
  • How to add safe‑split handling without external libraries
  • When pruning helps and when it hurts interpretability

Understanding the Failure Mode

A hand‑coded decision model often looks like a series of nested comparisons. If any feature used in a comparison is None, the comparison raises a TypeError and the model fails. Even if you catch the exception, falling back to a default branch can systematically mis‑predict minority cases.

Building a Minimal Hand‑Coded Decision Model

Below is a tiny model that decides loan approval based on annual income and credit score. It works fine as long as both fields are filled.

def loan_approved(income, credit_score):
    if income is None or credit_score is None:
        raise ValueError("Missing input")
    if income > 50000:
        if credit_score > 700:
            return True
        else:
            return False
    else:
        return False

## Example usage

print(loan_approved(60000, 720))  # True
Enter fullscreen mode Exit fullscreen mode

The function explicitly checks for missing inputs and raises an error. In a real pipeline you might let the error propagate, causing the whole batch to abort.

Adding Safe‑Split Handling for Missing Data

Instead of aborting, we can route missing values to a surrogate split that uses the other feature. This keeps the model running while preserving as much signal as possible.

def loan_approved_safe(income, credit_score):
    # If income is missing, decide based on credit score alone
    if income is None:
        return credit_score > 650
    # If credit score is missing, decide based on income alone
    if credit_score is None:
        return income > 45000
    # Normal logic when both are present
    if income > 50000:
        return credit_score > 700
    return False

## Examples

print(loan_approved_safe(None, 680))   # True (uses credit score)
print(loan_approved_safe(40000, None)) # False (uses income)
print(loan_approved_safe(60000, 720))  # True
Enter fullscreen mode Exit fullscreen mode

The surrogate splits are simple thresholds chosen from domain knowledge. In practice you would derive them from the training data, but the key idea is to avoid crashing while still making a usable prediction.

Trade‑offs: Pruning vs. Overfitting

Hand‑coded models are prone to overfitting when you keep adding rules for edge cases. Pruning simplifies the tree but may hide real patterns. The table below compares three strategies qualitatively.

Strategy Interpretability Overfitting Risk Implementation Effort
No pruning High (every rule visible) High (fits noise) Low (just add rules)
Pre‑pruning (max depth) Medium (depth limit) Medium Medium (need depth check)
Post‑pruning (cost‑complexity) Medium‑High (pruned after build) Low‑Medium High (requires validation set)

Choosing a strategy depends on how much you value transparency versus predictive stability. For quick prototypes, a modest depth limit often works well.

When the Model Still Breaks

Even with safe splits and pruning, the model can fail if:

  • A feature has an unexpected type (e.g., a string where a number is expected)
  • The data distribution shifts and the surrogate thresholds become meaningless
  • You add so many rules that the model becomes a lookup table, defeating the purpose of a decision tree

Regularly validating on a hold‑out set and updating thresholds helps mitigate these issues.

Key Takeaways

  • Missing values will crash naïve if‑else decision models unless you handle them explicitly.
  • Safe‑split routing lets the model keep working while preserving predictive signal.
  • Pruning reduces overfitting but adds complexity; pick a strategy that matches your interpretability needs.

Source

Build your own decision model – I added concrete failure analysis, working code with safe‑split handling, and a qualitative trade‑off table that the original post did not cover.

Support this work

These write-ups are researched and published with no paywall, sponsor, or tracking. If one saved you an afternoon, a small tip keeps them coming.

USDT, USDC or USDD · TRC-20 (Tron)

TFTNsfyomKrnUutRjBTGVULp19ByW29KbY
Enter fullscreen mode Exit fullscreen mode

Top comments (0)