I built a Python paper-trading bot for crypto spot markets. The backtest looked fine. Then I tested it the strict way, and it fell apart. This is what the strict way looks like, and what it found.
The rule: lock the criteria before you look
Before each test I write down the hypothesis, the data, the metric and the pass/fail thresholds. I hash the file (sha256). The test script refuses to run if the hash does not match. After that I get one run. No tweaking the threshold after seeing the result.
It sounds bureaucratic. It is the only thing that stopped me from fooling myself, because I tried a lot of ideas and some of them looked good at first glance.
Result 1: the bot had no edge before costs
Update (same day): a reader pointed out that treating 3,314 trades as independent understates the uncertainty. Resampling whole entry days (393 clusters), the pre-cost mean of -0.07% per trade has a 95% interval of about -0.56% to +0.41%, and the net mean of -0.37% has about -0.86% to +0.10%. So the honest reading is: no evidence of an edge before costs, costs of about 0.30% per trade are certain, and the 81%/19% split below is not statistically supportable. This check was done after the fact, not pre-registered.
I replayed the bot's engine on history: 3,314 simulated trades.
- Profit factor before any costs: 0.979
- After slippage: 0.949
- After fees: 0.894
So the signal itself was already slightly negative (about -0.07% per trade), and costs of about 0.30% per round trip did the rest. About 81% of the total loss was cost, 19% was the signal.
The shape of the trades was also a trap. Only 25.3% of trades won, with a payoff of 2.63 to 1. Breakeven needed 27.5% winners. We were 2.2 points short, and no amount of stop tuning changes that if entries are no better than random.
Result 2: the "random entries beat the bot" finding was my own bug
At one point a control showed random entries doing much better than the bot: profit factor 1.69 against 0.89. I almost wrote it up as a discovery.
I locked a follow-up test to find out why. Answer: the random entries were drawn from a window of plus or minus 14 days around each bot entry. The bot buys after a price rise by construction, so random entries placed before it were buying the same rise cheaper. Random entries placed after the bot entry:
- Random before: average +0.277 R, profit factor 1.69
- Random after: average -0.063 R, profit factor 0.875
- The bot: average -0.068 R, profit factor 0.894
The advantage was look-ahead in my control group, not a flaw in the bot. A placebo test is only useful if the placebo is clean.
Result 3: survivorship explained every false positive
Momentum, rotation and several other ideas looked good on today's list of coins. On a point-in-time universe that includes coins that were later delisted, they vanished. In my own post-mortem, the choice of universe explained all of the false positives I found.
Result 4: the effect sits in a handful of episodes
When the "good" result survived a basic test, I removed the best 5% of days or the top few episodes. In several cases the effect went to zero or flipped sign. The number of independent events was in the tens, even though the table showed thousands of rows, because rows from the same day are not independent.
What I took from it
- Write the pass/fail criteria first. Hash them. One run.
- Test against a placebo, and check that the placebo is clean.
- Check the result without its best days.
- Count independent events, not rows.
- Include delisted instruments.
- Stress the costs. Most small effects are smaller than the fee.
I turned this method into a small command-line toolkit, called BotGauntlet: you lock your criteria, run the tests on your own trade list, and get a plain PASS/FAIL report. It will most likely tell you your idea does not work, which is the point. It is not signals and not financial advice. I am collecting a waitlist at https://botgauntlet.com and I would like to hear which of these checks you already run.
Disclosure: the code and analysis were developed with AI assistance (Claude). All numbers above come from my own pre-registered runs on historical data, a simulation, not live trading. Trading involves risk of loss.
Top comments (8)
The cleanest line in this post is the placebo table: random entries placed after the bot's entry average -0.063 R, the bot averages -0.068 R. That says the bot is indistinguishable from entering at random after the same rise, which is a stronger statement than "slightly negative signal".
One check on the 81/19 split. Treating the 3,314 trades as independent, the 25.3% win rate has a standard error of about 0.76 points, so the interval is roughly 23.8% to 26.8% and the 27.5% breakeven sits outside it. But you note the independent events number in the tens, and trades that open on the same day share one market move. If I read that right, the interval on the pre-cost -0.07% per trade is much wider than the 3,314 suggests, and the 19% "signal" share could have either sign. Resampling whole days (or whole episodes) with replacement and reporting the interval on the pre-cost mean and on the bot-minus-placebo difference would show whether the loss is cost plus nothing, or cost plus a real negative edge.
Do you resample by day in the tool, or does it treat each trade as one draw?
Thank you, this is exactly the check I should have shown. I reran it on the same 3,314 trades, resampling whole entry days (393 clusters) instead of single trades. Pre-cost mean per trade: -0.07%, 95% interval roughly -0.56% to +0.41%. Net mean: -0.37%, interval roughly -0.86% to +0.10%. So you are right: the sign of the pre-cost signal can't be determined, and even the net loss isn't separated from zero at the day level. What stays solid is the cost side, about 0.30% per trade, and the placebo comparison (random-after matches the bot). The 81/19 split I wrote is not supportable at these intervals, so I'm correcting the post. To your question: the tool already aggregates trades to one value per date and runs the t-test and permutation test on those, but it does not report a day-resampled interval yet. I'm adding that. This rerun was done after the fact, not pre-registered, and I'll label it that way.
Good, that's the honest reading. One more layer on the day-level interval: 393 clusters treats adjacent days as independent, and trade outcomes in a single bot often aren't. Volatility regimes and a market that trends or chops for a week move many days together. A block bootstrap by week (or by month, if you have enough of them) will usually widen the interval again, so I'd report it next to the day-level one and say which you trust.
It also tells you how much data a verdict needs. With a half-width around 0.5% per trade and a cost drag of about 0.30%, the interval has to shrink to roughly 0.15% before the pre-cost sign or the net sign is separable, and half-width falls with the square root of the number of independent clusters. That is about 10x the current clusters. Worth writing in the post as the stopping rule, so a later rerun can't be read as a result found by looking longer.
Agreed on all points, and thanks for the arithmetic. I ran the weekly block bootstrap too (73 weekly blocks): pre-cost mean -0.07% per trade, roughly -0.60% to +0.44%; net mean -0.37%, roughly -0.90% to +0.14%. Practically the same as the day-level result, so adjacent-day dependence is not changing the conclusion here. Monthly blocks I skipped, there are too few months for the interval to mean anything. Your point about the half-width is the key one: at about 0.5% per trade I could not separate even a 0.30% cost drag from noise, and the interval would need to shrink about 3x, so roughly 10x more independent blocks. That means this history cannot confirm or reject a small edge, only rule out a large one. I'm adding both a block-resampled interval and a "minimum detectable effect" line to the tool's report, so the report says how much data a verdict needs. And yes, writing the stopping rule down before looking is exactly the point of the lock step.
Weekly and daily agreeing is the useful part: it says the width comes from per-trade noise, not from clustering, so no cleverer resampling scheme will rescue the interval.
On the minimum detectable effect line, I'd define it from power rather than from the interval. A 95% half-width of about 0.52 is a standard error of about 0.27. For 80% power at a two-sided 5% test the detectable effect is roughly 2.8 standard errors, so about 0.75% per trade, more than twice the 0.30% cost drag. Getting that down to 0.30% takes about 6x the blocks, roughly 450 weeks, a bit under the 10x I quoted earlier because that figure asked the whole interval to clear the drag rather than just reach 80% power. Both are worth printing, with the stated assumption that future weeks behave like these 73.
Checked your arithmetic and I get the same numbers: half-width 0.52% -> SE about 0.265%, 2.8 SE about 0.74% per trade, and about 6x the blocks (roughly 450 weeks) to bring it down to the 0.30% cost drag. So this history can rule out a large edge, but it cannot confirm or reject a small one.
One honest gap on my side: the tool's minimum detectable effect line is currently computed from the date-level standard error, and the week-level interval is printed separately. Here the two agree, but on other data the week-level one can be wider, and then the line would be too optimistic. I'm changing it to use the wider of the two and to print the assumption (future blocks behave like the sample) next to it.
Stopping rule for a rerun: no rerun before 6x the current independent blocks, and no change to the criteria in between. Thanks for the careful reading, this thread is better than the post.
Locking criteria upfront is the right instinct, but the devil's in what you didn't lock. Did you freeze the feature set, hyperparameters, and the train/test split boundaries — or just the signal logic?
In crypto spot, the silent killers are usually:
One thing that bit me: I locked the strategy but not the universe selection. The backtest ran on "top 50 by volume" — but that list changes monthly. The paper run picked a different 50. Performance evaporated.
What did your paper-trading environment actually simulate? Slippage model? Queue position? Orderbook depth? (site: labagent .tech)
Thanks for the questions. What I lock before each run is the hypothesis, the data, the metric and the pass/fail thresholds, and the script refuses to run if the hash does not match. The bot is spot only, so funding does not apply. On the universe: the survivorship result in the post is exactly that point, the effect disappeared once delisted coins were included. Costs were applied in the replay (slippage and fees, about 0.30% per round trip), but I agree that paper fills are optimistic, so this says nothing about live execution.