<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Botgauntlet</title>
    <description>The latest articles on DEV Community by Botgauntlet (@botgauntlet).</description>
    <link>https://dev.to/botgauntlet</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4171912%2F580537ed-ddde-4d56-b4d6-c1ad79409f4f.png</url>
      <title>DEV Community: Botgauntlet</title>
      <link>https://dev.to/botgauntlet</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/botgauntlet"/>
    <language>en</language>
    <item>
      <title>I locked my test criteria before running them. My trading bot failed.</title>
      <dc:creator>Botgauntlet</dc:creator>
      <pubDate>Thu, 08 Oct 2026 18:22:38 +0000</pubDate>
      <link>https://dev.to/botgauntlet/i-locked-my-test-criteria-before-running-them-my-trading-bot-failed-n4n</link>
      <guid>https://dev.to/botgauntlet/i-locked-my-test-criteria-before-running-them-my-trading-bot-failed-n4n</guid>
      <description>&lt;p&gt;I built a Python paper-trading bot for crypto spot markets. The backtest looked fine. Then I tested it the strict way, and it fell apart. This is what the strict way looks like, and what it found.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule: lock the criteria before you look
&lt;/h2&gt;

&lt;p&gt;Before each test I write down the hypothesis, the data, the metric and the pass/fail thresholds. I hash the file (sha256). The test script refuses to run if the hash does not match. After that I get one run. No tweaking the threshold after seeing the result.&lt;/p&gt;

&lt;p&gt;It sounds bureaucratic. It is the only thing that stopped me from fooling myself, because I tried a lot of ideas and some of them looked good at first glance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 1: the bot had no edge before costs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Update (same day):&lt;/strong&gt; a reader pointed out that treating 3,314 trades as independent understates the uncertainty. Resampling whole entry days (393 clusters), the pre-cost mean of -0.07% per trade has a 95% interval of about -0.56% to +0.41%, and the net mean of -0.37% has about -0.86% to +0.10%. So the honest reading is: no evidence of an edge before costs, costs of about 0.30% per trade are certain, and the 81%/19% split below is not statistically supportable. This check was done after the fact, not pre-registered.&lt;/p&gt;

&lt;p&gt;I replayed the bot's engine on history: 3,314 simulated trades.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Profit factor before any costs: 0.979&lt;/li&gt;
&lt;li&gt;After slippage: 0.949&lt;/li&gt;
&lt;li&gt;After fees: 0.894&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the signal itself was already slightly negative (about -0.07% per trade), and costs of about 0.30% per round trip did the rest. About 81% of the total loss was cost, 19% was the signal.&lt;/p&gt;

&lt;p&gt;The shape of the trades was also a trap. Only 25.3% of trades won, with a payoff of 2.63 to 1. Breakeven needed 27.5% winners. We were 2.2 points short, and no amount of stop tuning changes that if entries are no better than random.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 2: the "random entries beat the bot" finding was my own bug
&lt;/h2&gt;

&lt;p&gt;At one point a control showed random entries doing much better than the bot: profit factor 1.69 against 0.89. I almost wrote it up as a discovery.&lt;/p&gt;

&lt;p&gt;I locked a follow-up test to find out why. Answer: the random entries were drawn from a window of plus or minus 14 days around each bot entry. The bot buys after a price rise by construction, so random entries placed before it were buying the same rise cheaper. Random entries placed after the bot entry:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Random before: average +0.277 R, profit factor 1.69&lt;/li&gt;
&lt;li&gt;Random after: average -0.063 R, profit factor 0.875&lt;/li&gt;
&lt;li&gt;The bot: average -0.068 R, profit factor 0.894&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The advantage was look-ahead in my control group, not a flaw in the bot. A placebo test is only useful if the placebo is clean.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 3: survivorship explained every false positive
&lt;/h2&gt;

&lt;p&gt;Momentum, rotation and several other ideas looked good on today's list of coins. On a point-in-time universe that includes coins that were later delisted, they vanished. In my own post-mortem, the choice of universe explained all of the false positives I found.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 4: the effect sits in a handful of episodes
&lt;/h2&gt;

&lt;p&gt;When the "good" result survived a basic test, I removed the best 5% of days or the top few episodes. In several cases the effect went to zero or flipped sign. The number of independent events was in the tens, even though the table showed thousands of rows, because rows from the same day are not independent.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I took from it
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Write the pass/fail criteria first. Hash them. One run.&lt;/li&gt;
&lt;li&gt;Test against a placebo, and check that the placebo is clean.&lt;/li&gt;
&lt;li&gt;Check the result without its best days.&lt;/li&gt;
&lt;li&gt;Count independent events, not rows.&lt;/li&gt;
&lt;li&gt;Include delisted instruments.&lt;/li&gt;
&lt;li&gt;Stress the costs. Most small effects are smaller than the fee.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I turned this method into a small command-line toolkit, called BotGauntlet: you lock your criteria, run the tests on your own trade list, and get a plain PASS/FAIL report. It will most likely tell you your idea does not work, which is the point. It is not signals and not financial advice. I am collecting a waitlist at &lt;a href="https://botgauntlet.com" rel="noopener noreferrer"&gt;https://botgauntlet.com&lt;/a&gt; and I would like to hear which of these checks you already run.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: the code and analysis were developed with AI assistance (Claude). All numbers above come from my own pre-registered runs on historical data, a simulation, not live trading. Trading involves risk of loss.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>testing</category>
      <category>crypto</category>
      <category>datascience</category>
    </item>
  </channel>
</rss>
