DEV Community

Wataru Suda
Wataru Suda

Posted on

Your backtest is lying to you: walk-forward analysis and the Deflated Sharpe Ratio in plain Python

A backtest answers "what would have happened". The question you care about is "what will happen", and the gap between the two grows with every parameter you try.

I have a simple daily trend-following rule running with real money (small money: about ¥24,000). Before it went live I tried 108 parameter combinations. The best had a Sharpe ratio of 1.46 on the last three years. My fixed, picked-in-advance parameters scored 1.20.

(The free lite backtester for this system, public API only, no key: GMO Coin Trend Lab, source on GitHub.)

Should I switch to the winner? Two tools answer that, and both fit in a page of Python.

1. Walk-forward: only score what the optimiser has not seen

The idea: optimise on a window (in-sample), then trade the winner on the next window (out-of-sample), roll forward, repeat. Stitch the out-of-sample pieces together. That stitched curve is the only performance you are allowed to believe.

import pandas as pd
from itertools import product

grid = list(product([10, 20, 40, 55],      # breakout lookback
                    [5, 10, 20],           # exit lookback
                    [50, 100, 200],        # trend filter
                    [1.5, 2.0, 3.0]))      # ATR stop multiple  -> 108 sets

oos_returns = []
for year in range(2021, 2027):
    is_start, is_end = f"{year-2}-01-01", f"{year-1}-12-31"
    oos_start, oos_end = f"{year}-01-01", f"{year}-12-31"

    best, best_sharpe = None, -9
    for g in grid:
        res = run_backtest(data, params(g), is_start, is_end)   # in-sample only
        if res["sharpe"] > best_sharpe:
            best, best_sharpe = g, res["sharpe"]

    oos = run_backtest(data, params(best), oos_start, oos_end)  # never seen by the optimiser
    oos_returns.append(oos["equity"].pct_change().dropna())

oos = pd.concat(oos_returns)
print("stitched OOS Sharpe:", oos.mean() / oos.std() * 365 ** 0.5)
Enter fullscreen mode Exit fullscreen mode

My result, on daily bars from 2018 to 2026:

Out-of-sample year Optimised parameters Fixed parameters (20/10/50/2.0)
2021 +20.2% +30.0%
2022 −7.5% −6.2%
2023 +12.0% +14.2%
2024 +28.7% +61.9%
2025 +0.7% +1.2%
2026 YTD +1.9% −2.1%

Stitched out-of-sample Sharpe: 0.85, versus 1.11 for the full-period backtest.

Two things to read from this table. The honest number is 0.85, not 1.11 and certainly not 1.46. And the fixed parameters did as well as or better than the re-optimised ones in four years out of six. Re-optimising added nothing. That is a useful negative result: it means the rule is not fragile, and it means I should stop fiddling.

A common sanity metric is walk-forward efficiency: out-of-sample annual return divided by in-sample annual return. Above 0.5 is usually called acceptable, below 0.3 smells like overfitting.

2. Deflated Sharpe Ratio: pay for every trial you ran

Try enough random strategies and one will look brilliant. Bailey and López de Prado's Deflated Sharpe Ratio asks: given that I ran N trials, what Sharpe would the best one show by luck alone, and does mine clear that bar?

from math import sqrt, exp
from statistics import NormalDist

def deflated_sharpe(sr, n_obs, skew, kurt, n_trials, var_trial_sr):
    """sr and var_trial_sr are per-period (daily), not annualised.
    kurt is raw kurtosis (3 for a normal distribution)."""
    N = NormalDist()
    euler = 0.5772156649
    if n_trials > 1:
        expected_max = sqrt(var_trial_sr) * (
            (1 - euler) * N.inv_cdf(1 - 1 / n_trials)
            + euler * N.inv_cdf(1 - 1 / (n_trials * exp(1))))
    else:
        expected_max = 0.0
    denom = sqrt(1 - skew * sr + (kurt - 1) / 4 * sr ** 2)
    return N.cdf((sr - expected_max) * sqrt(n_obs - 1) / denom)
Enter fullscreen mode Exit fullscreen mode

Inputs you need: the daily Sharpe of the strategy you picked, the number of daily observations, the skew and kurtosis of its returns, the number of trials, and the variance of Sharpe ratios across those trials.

With n_trials=1 this collapses to the Probabilistic Sharpe Ratio, the probability that the true Sharpe is above zero given non-normal returns.

My numbers:

  • PSR against zero: 0.999. The strategy's edge over nothing is statistically solid.
  • DSR treating all 108 grid points as trials: 0.08. As "the best of 108", the result is not significant at all.

Both are true at once, and the difference is the whole point. If I had chosen my parameters by picking the grid winner, I would have to live with 0.08. I did not. The 20/10/50/2.0 set comes from the published literature and was fixed before I looked at this data, so the first number applies. If you cannot honestly say that, the second number is yours.

Rules I now follow

  1. Log every trial. The DSR needs the count. If you do not know how many things you tried, assume it is a lot.
  2. Fix the walk-forward windows before you start. Tuning the window length is just another parameter search wearing a disguise.
  3. Plan around the out-of-sample number. Mine is 0.85, about +9% a year. That is what goes into my expectations, not 16.9%.
  4. Halve anything published. McLean and Pontiff found anomaly returns drop by roughly half after publication. My rule is a published one.
  5. Change parameters only when the challenger wins out-of-sample by a wide margin. My weekly review requires a Sharpe gap above 0.5 and a win on the most recent held-out year. So far: no changes.

Try it on real data

The free lite backtester that produced the full-period numbers above pulls daily bars from the exchange's public API (no key, no account): https://wataflow1.gumroad.com/l/trend-lab-free

The walk-forward loop and the DSR function shown here, together with the live engine they guard, are in the paid kit on the same store.

Not investment advice. The most valuable output of this exercise was a reason to do nothing.

Top comments (0)