DEV Community

Cover image for I Tested 58 Trading Hypotheses. 48 Failed. Here’s the Framework That Killed Them.
Pavel Kazantsev 💎
Pavel Kazantsev 💎

Posted on Originally published at pkazantsev.com

I Tested 58 Trading Hypotheses. 48 Failed. Here’s the Framework That Killed Them.

58 hypotheses went in. 48 were rejected, 9 remained promising, and 1 could not be evaluated because the required data was unavailable.

None of the 9 earned the label "edge." They earned the right to be tested further.

Some ideas should die before a backtest is ever run. The rest should face the cheapest falsification test first. A hypothesis that cannot answer "who pays us?" and "why does this persist?" is not a hypothesis. It is a guess with a Python file attached.

The kill pipeline that processed all 58, in the order most ideas die:

hypothesis idea
    ↓
cost filter        (does it survive fees and turnover at all?)
    ↓
robustness check   (is the result a plateau or a lucky parameter?)
    ↓
null benchmark     (does it beat a structure-preserving random version?)
    ↓
untouched holdout  (does it hold on a year it never saw?)
    ↓
contamination audit (is the result real or did the code lie to me?)
    ↓
survived research → forward / shadow validation → capital
Enter fullscreen mode Exit fullscreen mode

Each gate catches a different type of false alpha. A strategy can clear the first four and die on the fifth. Passing all five is not a green light. It is permission to start collecting forward data.

Full kill pipeline — from hypothesis through five gates to locked capital

Gate 1: costs

The most common way a backtest looks good: skip costs. Cross-sectional reversal on 15-minute Bybit bars had a gross Sharpe of 7 to 13, positive in 80%+ of out-of-sample windows. The gross effect was clearly present before costs. At taker fees, rebalancing every bar, trading costs ran 5 to 10 times the size of the raw edge. After costs, zero out of 76 out-of-sample windows remained positive.

I tried entry/exit thresholds, hysteresis, coarser timeframes, slower rebalancing. No plateau anywhere near break-even. Adjacent parameter settings swung from +45% to −60%.

Run the cost test before you believe the signal, not after.

Before costs: Sharpe 7–13, 80%+ OOS positive. After costs: 0/76 windows positive

Gate 2: robustness

Daily momentum had one configuration that returned +97% net at taker fees. The lookback sweep: l=7 gave +33%, l=10 gave −23%, l=14 gave +97%, l=21 gave +35%. No plateau. The +97% was a lucky draw from a jagged surface.

Half the return came from the 2021 bull alone. 2022 was negative. A strategy that only makes money in bull markets and loses in bear is more consistent with market beta than neutral alpha.

One good result on a swept parameter is not evidence. The whole surface needs to hold up.

Momentum lookback sweep: l=7 +33%, l=10 -23%, l=14 +97%, l=21 +35% — no plateau

Gate 3: null benchmark

Funding carry looked strong in development: +48.6% net over 4.5 years, Sharpe 1.17. Before trusting it, I ran a null test: scramble which coin's funding the strategy reads, keep everything else identical. The scrambled version returned +28%.

More than half the return was the structural funding premium: exposure to perpetual market mechanics rather than signal-driven selection. The selection component showed evidence beyond this null in development (empirical permutation p=0.024), but it was much smaller than the headline return implied.

Before believing the result, ask whether a random version of the signal would also make money.

Funding carry +48.6% vs scrambled null +28%, selection component +20.6%, p=0.024

Gate 4: untouched holdout

The lockbox is a date range fixed in advance, never touched until a single opening event. No research, no tuning, no plotting inside that range until development is complete.

Funding carry passed development and validation. The lockbox was a 12-month period (2025-09 to 2026-09), sealed from day one. The frozen strategy made −0.5% on it. The premium that paid from 2021 to 2024 was absent in the lockbox year. One year of absence is not a permanent verdict. It is exactly the kind of regime dependence the lockbox is designed to expose early.

Once a lockbox is opened, it is burned. Any further confirmation requires a new, still-unseen date range. Development results are hypotheses, and only sealed, untouched data can confirm them.

Timeline 2021–2026: dev period +48.6%, lockbox 2025-09 to 2026-09 returned -0.5%

Gate 5: contamination audit

One BTC funding-trend hypothesis had a development t-stat of 2.6. The 1,248 events had 72-hour holding periods, so observations were heavily overlapping. The Newey-West corrected t-stat fell from 2.6 to 1.07, removing the apparent statistical evidence from that test. A real-looking result from a real-looking test, broken by one statistical assumption.

On a separate project, a football prediction model passed closing-line value (CLV) validation with +152bps, 95% CI excluded zero. A five-check forensic audit found two contaminated features. Four of five checks failed. The metric that reported the good result was the last place to look for the error.

1,248 overlapping 72h events, OLS t-stat 2.6 drops to 1.07 after Newey-West NW-9 correction

Where the 58 actually died

Each hypothesis is assigned to the first decisive failure that stopped further evaluation. Some failed more than one check.

Kill gate Count Example
Costs ~20 Price MR: gross Sharpe 7–13, net negative
Robustness ~10 Daily momentum: +97% was one lucky lookback
Null benchmark ~8 Funding carry null still returned +28%
Holdout / fresh OOS ~7 Funding carry −0.5% on sealed year; agent ensemble −35% OOS
No detectable signal / other rejection ~3 New-listing cohort: corr≈0, null p=0.94
Total rejected 48
Data unavailable 1 Required data unavailable
Remaining PROMISING 9 Passed development and validation gates
Benjamini–Hochberg (BH) significant at FDR=0.05 8
Independent mechanisms 4–5 roughly 4–5 distinct economic bets

9 survivors does not mean 9 strategies. Several share the same economic mechanism with different entry filters. 8 Benjamini–Hochberg (BH) significant hypotheses at FDR=0.05 collapse to roughly 4–5 independent bets on market structure by economic mechanism.

58 → 48 killed + 9 promising + 1 unavailable → 8 BH-significant → 4–5 independent mechanisms

The uncomfortable failures

Two cases that took longer than they should have.

An LLM-generated factor-residual ensemble showed +76% in development, Sharpe 0.6, tuned over multiple iterations. On a fresh holdout it reversed to −35%. The development result held inside the data it was tuned on. The fresh holdout showed what the agent had actually learned: the development period.

In a funding carry reproduction audit a month after the original publication, the holdout result changed from −0.5% to +1.3%. Same code, same data cache, fingerprints matched. I still do not have a clean explanation. I published the discrepancy table rather than quietly resolving it. I scrutinized the original negative result far less than I would have scrutinized a positive one. Negative results get a lower evidentiary standard by default. They should not.

What survived

A handful of hypotheses passed enough gates to justify forward validation. They are currently shadow-deployed: signals recorded in real time, parameters locked, no capital.

None of them earned the right to be trusted based on the backtest. The backtest is permission to collect more data, not permission to trade.

The specific signals are not published here. Forward validation is ongoing and the gate has not been reached.

Steal this framework

Before trusting any strategy result:

[ ] Realistic costs included from the start, not added after
[ ] Parameter robustness: result holds over adjacent settings (plateau, not spike)
[ ] Null benchmark: structure-preserving random version loses
[ ] Untouched holdout: a date range sealed before any research on that data
[ ] Contamination audit: no look-ahead, no survivorship, no feature leakage
[ ] Forward validation: signals collected in real time, parameters locked
Enter fullscreen mode Exit fullscreen mode

Costs and robustness killed roughly half of the 58 before a holdout was needed. Run the cheap gates first.

Research kill checklist: 6 gates from costs to forward validation

The real lesson

A high kill rate is not evidence that the pipeline is good. But a research process that rarely kills its own ideas is one I would distrust.

48 of 58 died. If the kill rate were much lower, I would not trust the process.


More experiments, engineering notes and quantitative research at pkazantsev.com

Top comments (2)

Collapse
 
deanlee profile image
Dean Lee •

The funding carry null test isolates the structural baseline that catches most cross-sectional pipelines. That 28 percent return from scrambled funding rates reflects the market-clearing premium for warehouse inventory risk. When retail pays funding to maintain levered long exposure, counterparties absorb directional tail skew. The position underwrites liquidity insurance rather than harvesting idiosyncratic selection. When selection only adds marginal basis over that baseline, the portfolio carries unhedged jump risk into deleveraging cascades.

The Newey-West correction in Gate 5 resolves the mechanical flaw that inflates overlapping holding periods. A 72-hour holding window injects moving-average errors into the residuals, but volatility clustering does the real damage. When funding spreads widen, return variance typically spikes at the exact same time. An unadjusted t-statistic treats high-volatility clustered events as independent draws and compresses the standard error. Forcing the correction early prevents mistaking regime persistence for statistical significance.

Collapse
 
pavel_kkkkazantsev profile image
Pavel Kazantsev 💎 •

The inventory risk framing is the right read on that 28%. Most pipelines that look like selection are levered exposure to the clearing mechanism — the funding receiver is short the retail panic premium, not harvesting alpha.

The volatility clustering is what made the 2.6 look credible. Overlapping windows were visible in the data; the conditional heteroskedasticity wasn't until I forced the correction. Both inflate the stat for different reasons — MA residuals are mechanical, clustering is regime-driven. An autocorrelation plot on the squared residuals would have caught it earlier.

Useful additions.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.