DEV Community

joeschatzman
joeschatzman

Posted on

Why Your Backtest Is Lying to You (and How to Close the Backtest-to-Live Gap)

I once built a strategy with a beautiful backtest. Smooth equity curve, healthy Sharpe, a win rate that made me feel clever. I put it live.

It proceeded to lose money at roughly an 8% win rate and cost me five figures before I pulled the plug.

Nothing about the code was "wrong." The backtest was just lying to me — quietly, in the specific ways backtests always do. This post is the checklist I wish I'd run before going live. None of it is exotic; all of it is the difference between a number that flatters you and a process you can trust.

Why backtests overstate reality

A backtest is a simulation, and every simulation makes assumptions. The optimistic ones stack up:

  • Overfitting. If you tune parameters until the curve looks great, you haven't found an edge — you've memorized the noise in one slice of history. The more knobs you turn, the more certain this becomes.
  • Lookahead bias. Using information that wouldn't have been available at decision time. The classic version: computing a signal on today's close and then "buying at the close." In reality you'd act on the next bar. It's easy to leak the future without noticing.
  • Slippage and fills. Backtests love to fill you at the exact price you wanted. Live markets don't. On anything less than deeply liquid instruments, the gap between assumed and actual fills eats returns.
  • Survivorship bias. Testing on today's index members ignores every ticker that got delisted. Your universe is quietly pre-filtered for winners.
  • Regime dependence. A strategy tuned on a 2023–2024 bull run has never seen a real drawdown. It looks robust because it was never stressed.

Individually, each nudges results up a little. Together, they can turn a break-even system into a "genius" backtest.

How to actually close the gap

The goal isn't a prettier backtest. It's a strategy that behaves live roughly like it did in the test. Here's the sequence that gets you there.

1. Split your data — and mean it

Tune on one stretch of history, then evaluate on data the strategy has never touched. The simplest version: build on 2023–2024, validate on 2025. Better: walk-forward — repeatedly optimize on a window, test on the next unseen window, and roll it forward. If performance falls off a cliff out-of-sample, you found an artifact, not an edge.

# Conceptual walk-forward loop
for train_window, test_window in rolling_windows(data):
    params = optimize(train_window)      # fit on the past
    result = run(test_window, params)    # judge ONLY on unseen data
    oos_results.append(result)
# The out-of-sample curve is the one that matters.
Enter fullscreen mode Exit fullscreen mode

2. Stress the trade order, not just the trades

A single equity curve is one lucky (or unlucky) ordering of your trades. Monte Carlo it: reshuffle the trade sequence thousands of times and look at the distribution of outcomes. If the 5th-percentile path is a catastrophe, your "safe" strategy isn't — you just got a friendly ordering the first time.

3. Model slippage and commissions on purpose

Add realistic per-trade cost and slippage assumptions and re-run. If a small, honest slippage estimate erases the edge, the edge was never real — it lived in the frictionless fantasy of a naive fill model.

4. Paper trade before real capital

Run the exact same logic against live market data with no money on the line. This is where lookahead bugs and data-feed quirks surface — the ones no historical test can catch because they only exist in the seam between "backtest engine" and "live engine." Ideally the same code runs both, so there's no second implementation to drift.

5. Put the guardrails in before you need them

Even a validated strategy meets a market it didn't expect. Decide your limits up front and let the machine enforce them:

  • position sizing rules (fixed fractional, volatility-scaled — pick one and stick to it),
  • a hard per-trade risk cap,
  • and a drawdown circuit breaker that halts the strategy automatically if it hits your max-loss line.

The point isn't to predict the bad day. It's to make sure the bad day can't compound while you're asleep.

The honest checklist

Before anything goes live, I want a "yes" to all of these:

  • [ ] Does it survive out-of-sample / walk-forward, not just the fitted window?
  • [ ] Is the Monte Carlo distribution acceptable at the 5th percentile, not just the median?
  • [ ] Does it still work with realistic slippage + commissions?
  • [ ] Did it run in paper trading and behave like the backtest?
  • [ ] Are position sizing, risk limits, and a drawdown halt actually wired in?

If any answer is "no," I don't have an edge yet. I have a hypothesis and a pretty chart.

A note on tooling

You can do all of this by hand — walk-forward loops, a Monte Carlo reshuffler, a slippage model, a paper-trading harness, a risk monitor. I eventually got tired of rebuilding that scaffolding for every idea and built a self-hosted framework where the same engine backtests and trades live, so there's no second implementation to drift. That's AlgoDeploy if you're curious — but the method above is the thing that matters, whatever you run it with.

The one-sentence version

A great backtest is easy; a real edge is hard — and the entire job of a serious process is to tell them apart before you fund the account, not after.

Backtested and hypothetical results have inherent limitations and do not guarantee future performance. Nothing here is investment advice — it's about software and process. Trade your own decisions.

Top comments (0)