You test 11 moving-average crossovers on five years of daily prices. The best one has a Sharpe of 1.2, and a t-test on its returns says p = 0.01. Is it real?
We simulated that situation 300,000 times, in markets where by construction nothing can be predicted, and counted how often the usual tests said "edge". A t-test on the best rule said yes between 10.9% and 78.9% of the time, depending on the market. The usual fix for testing many rules, Bonferroni, held up in flat markets but said yes 8.0% to 10.3% of the time when the market simply drifted upward, and 58.5% to 59.7% of the time when stocks just had different average returns. A placebo test stayed between 4.6% and 5.5% everywhere, against a promised 5%.
This post is about why, and how to run the placebo on your own pipeline.
Two things that look like skill
The search. The best of 11 rules is the best of 11 draws. Even with no edge, the maximum is high. On flat simulated markets, an 11-rule SMA grid's best Sharpe averaged 0.30 to 0.32 with nothing to find. That is what Bonferroni and the deflated Sharpe ratio correct for, if you know how many rules you really tried.
What the market does on its own. A long-only rule earns the drift of a rising market without predicting anything. A cross-sectional ranking that buys past winners earns the gap between assets with different long-run average returns, again without predicting anything: in our simulated markets where assets only differed in their constant means, a 20-stock momentum grid's best Sharpe averaged 1.27 to 1.29. A t-test on the strategy's returns asks "did it make money?", not "did it predict?", so drift and dispersion pass for skill, and dividing alpha by the number of rules does nothing about that.
The placebo
Keep your data's returns, but put the days in a random order, the same order for every asset. Each asset keeps its own returns, fat tails and average; each day keeps its cross-section; but nothing about a day tells you anything about the next one. That is a placebo: a market with the same raw material and nothing to predict.
Now run your whole pipeline on the real data and on 19 placebos: the same search, the same filters, the same parameter choices. If the real result beats all 19, the chance of that happening by luck is 1 in 20 (p = 0.05). If it beats 15 of them, p = 0.25. Every choice the pipeline makes is counted, because it makes it again on every placebo, and drift and dispersion are in the placebos too, so they stop looking like skill.
When the days are interchangeable, this p-value is exact: the real ordering is just one of the orderings. Volatility clustering breaks that, so we measured it.
The study
We wrote the rules down before running anything: three ways of reordering (a full permutation, a permutation of blocks of 11 days, and a stationary bootstrap), a bar of at most 6% false positives at a nominal 5% in every simulated market, and the rule that picks the winner. The pre-registration was committed before the run, and a test in the repository checks the run used it.
- Markets with nothing to predict: five return shapes (normal, fat-tailed, negatively skewed, GARCH volatility clustering, volatility regimes), each flat or drifting, or for the cross-section, with equal or dispersed average returns. 1,260 days each.
- Pipelines: an 11-rule SMA crossover grid, a 6-lookback time-series momentum grid, and a 4-lookback cross-sectional momentum grid over 20 assets, each reporting its best Sharpe.
- 30 combinations, 10,000 tests each, with 19 placebos per test, plus planted trends to measure power.
The full permutation met the bar in all 30: 4.6% to 5.5%, including GARCH and regime markets. The block permutation also met it (5.8% at worst) with less power. The stationary bootstrap was far too cautious: never above 1.5%, so it would miss real edges.
How often does the placebo catch a real edge? With a planted trend that a perfect forecaster would trade at a Sharpe of about 2, the SMA grid beat all 19 placebos 45.1% of the time and the time-series grid 47.3%; at a Sharpe of about 1, both were near 14.8%. The cross-sectional grid caught its strongest planted trend 89.0% of the time. An honest caveat: with only 19 placebos, Bonferroni catches some trends more often in flat markets, where it is valid. 99 placebos give a finer p-value.
What it does not tell you
Beating the placebos shows that the real order of your returns holds more than random orders of the same returns. It says nothing about costs, capacity or whether the edge lasts. And it counts only the search this pipeline runs: if you tried and dropped ten other pipelines before this one, those trials are not in the number.
Try it
The placebo is the placebo_test tool in canli-validation-mcp, an open-source MCP server. plan writes your price file and 19 placebo files to a temporary folder; run your pipeline on each, then compare returns the p-value. It runs on your machine; your code never leaves it. In Claude Code:
claude mcp add canli-validation -- npx -y canli-validation-mcp
The study's code, pre-registration and results are in the canlicapital repository. If you run the placebo on a real pipeline, I would like to hear what it finds.
Top comments (0)