<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dry Run</title>
    <description>The latest articles on DEV Community by Dry Run (@thedryrun).</description>
    <link>https://dev.to/thedryrun</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4083479%2F6ca25be6-5b3f-4e96-aa32-198e4db01954.jpg</url>
      <title>DEV Community: Dry Run</title>
      <link>https://dev.to/thedryrun</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/thedryrun"/>
    <language>en</language>
    <item>
      <title>A decorrelated bot that loses money is just a diversified way to lose money</title>
      <dc:creator>Dry Run</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:15:12 +0000</pubDate>
      <link>https://dev.to/thedryrun/a-decorrelated-bot-that-loses-money-is-just-a-diversified-way-to-lose-money-10hm</link>
      <guid>https://dev.to/thedryrun/a-decorrelated-bot-that-loses-money-is-just-a-diversified-way-to-lose-money-10hm</guid>
      <description>&lt;p&gt;&lt;em&gt;I added trend-following bots to my fleet to decorrelate it. The correlation matrix said it worked. The P&amp;amp;L said it didn't. Here's the number that settles the argument.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I run a small fleet of trading bots in paper mode. Most of them are variations on the same idea — mean-reversion on crypto: buy the dip, scale in, exit in green. Running five bots that all do the same thing isn't a fleet, it's one bet with extra logging. So I deliberately added bots designed to make money when the dip-buyers struggle: trend-followers (a Donchian breakout, a Supertrend). A trend-follower next to a dip-buyer is supposed to be out of phase — one eats ranging chop, the other eats sustained moves. On paper, textbook diversification.&lt;/p&gt;

&lt;p&gt;Then I measured it. Two months of daily P&amp;amp;L snapshots, five bots. The correlation matrix lied to me in both directions, and untangling how is the most useful thing I've done to the fleet all quarter.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "diversified" actually looked like
&lt;/h2&gt;

&lt;p&gt;I take a snapshot of each bot's P&amp;amp;L every day (realized plus the mark-to-market on open positions) and correlate the daily changes. Two numbers stopped me:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The two bots I was &lt;strong&gt;most confident&lt;/strong&gt; were diversified — my aggressive dip-buyer, and a second instance of the same strategy running on a different exchange — correlated at &lt;strong&gt;0.93&lt;/strong&gt;. They're the same strategy on two venues, and the market doesn't care which venue you're on. I wasn't running two bots. I was paying twice for one bet and calling the duplicate "diversification." (The more conservative sibling of the same family came in at 0.51 — better, still the same reflex reacting to the same dips.)&lt;/li&gt;
&lt;li&gt;The bots I added &lt;strong&gt;specifically to decorrelate&lt;/strong&gt; — the trend-followers — actually delivered on the correlation axis. The breakout bot correlated with the dip-buyers at around &lt;strong&gt;0.05&lt;/strong&gt;. That's about as uncorrelated as two things trading the same asset class get. On this axis, the plan worked exactly as designed: when the mean-reverters zigged, it zagged.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the "diversifiers" I trusted were secretly the same bet, and the "diversifiers" I added on purpose were genuinely uncorrelated. Great news, right?&lt;/p&gt;

&lt;h2&gt;
  
  
  The correlation matrix is a liar of omission
&lt;/h2&gt;

&lt;p&gt;Here's the P&amp;amp;L over the same window:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Breakout bot: −445 units.&lt;/strong&gt; Its backtest over a roughly −41% market was −42%, with a negative Sharpe. A confirmed, structural money-loser.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Supertrend bot: +501 units&lt;/strong&gt; — the best raw P&amp;amp;L in the entire fleet — on a strategy whose backtest over the same cycle was &lt;strong&gt;−52%, Sharpe −5.8.&lt;/strong&gt; The paper gain was a regime mirage: two lucky months sitting on top of a strategy that bleeds across a full cycle.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now put the two facts together. My beautifully decorrelated breakout bot (r ≈ 0.05) has negative expectancy. Adding it to the fleet doesn't reduce portfolio risk — it &lt;em&gt;adds&lt;/em&gt; a stream of losses that merely happens to arrive on different days than everyone else's. That's the trap, and it deserves to be stated as a rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Decorrelation only helps a portfolio when it's multiplied by an expectancy that isn't negative. A decorrelated bot with negative edge is not a hedge. It's a diversified way to lose money.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Low correlation is &lt;em&gt;necessary&lt;/em&gt; for diversification, not &lt;em&gt;sufficient&lt;/em&gt;. The textbook version quietly assumes every leg has a non-negative return and only the &lt;em&gt;timing&lt;/em&gt; differs. Drop that assumption — as any real, unvetted strategy forces you to — and the correlation coefficient becomes exactly the kind of flattering metric that hides the failure mode. It told me the breakout bot was a great teammate. The expectancy told me it was dead weight in a different jersey.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I did about it
&lt;/h2&gt;

&lt;p&gt;I paused both trend-followers. Not because they were correlated — they weren't — but because "uncorrelated and losing" is worse than useless: it loses money &lt;em&gt;and&lt;/em&gt; adds operational surface area to babysit. Low correlation earned them a second look; the backtest and the net-of-open-positions P&amp;amp;L is the test they actually had to survive, and they failed it.&lt;/p&gt;

&lt;p&gt;That left the real question. If a second crypto strategy can't decorrelate my crypto fleet — they all ultimately trade "is crypto up or down today" — where does real decorrelation come from? Not from another strategy on the same asset class. From a different asset class entirely. The genuinely independent engine in my fleet is a dual-momentum ETF bot running on a stock broker: different market, different instruments, and a different &lt;em&gt;clock&lt;/em&gt; — it rebalances once a month instead of reacting intraday. Its P&amp;amp;L physically can't correlate 0.93 with a crypto dip-buyer, because it isn't looking at the same universe on the same timescale. That's structural decorrelation, not a strategy I bolted on and hoped.&lt;/p&gt;

&lt;h2&gt;
  
  
  The general lesson (as usual, not really about trading)
&lt;/h2&gt;

&lt;p&gt;This is a portfolio problem, but it's the same shape as every system where you add redundant components and tell yourself you've bought resilience:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Two services behind one load balancer, both calling the same downstream database, are one point of failure with two hostnames.&lt;/strong&gt; Their availability is 0.93-correlated whether the architecture diagram admits it or not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Three "independent" safety checks all fine-tuned from the same base model&lt;/strong&gt; tend to be wrong on the same inputs. You think you bought an ensemble; you're running one opinion in triplicate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A backup in the same region, same account, same blast radius&lt;/strong&gt; as the thing it protects is decorrelated on paper and perfectly correlated the one day it matters.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The engineering move is identical every time: don't count boxes, measure the correlation of their &lt;em&gt;failures&lt;/em&gt; — and then check that each box has a positive reason to exist on its own. Redundancy that shares a hidden root is fake redundancy. And a component that's genuinely independent but quietly broken doesn't harden the system; it just spreads the breakage thin enough to be hard to see.&lt;/p&gt;

&lt;p&gt;I added losing bots on purpose to decorrelate a fleet, and the measurement taught me two things I didn't want to hear: the bots I trusted were the same bet, and the bots that were truly independent weren't worth keeping. Real diversification wasn't a cleverer strategy — it was a different asset class with a different clock. Everything else was correlation theater.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;One honest caveat: two months of paper data across a single market regime is a small, ugly sample, and every number above will move. I'm not reporting constants — I'm reporting a method. Measure the correlation of failures, then demand that each leg justify itself standalone. That part survives the sample.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I document a fleet of autonomous trading agents as an engineering problem — evaluation, risk, reliability — not as trading advice. No signals, no tips; just what breaks and how I try to catch it first. Follow along if that's your thing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>programming</category>
      <category>softwareengineering</category>
      <category>ai</category>
      <category>trading</category>
    </item>
    <item>
      <title>A 100% win rate is a red flag, not a résumé</title>
      <dc:creator>Dry Run</dc:creator>
      <pubDate>Tue, 18 Aug 2026 14:52:18 +0000</pubDate>
      <link>https://dev.to/thedryrun/a-100-win-rate-is-a-red-flag-not-a-resume-50fd</link>
      <guid>https://dev.to/thedryrun/a-100-win-rate-is-a-red-flag-not-a-resume-50fd</guid>
      <description>&lt;p&gt;&lt;em&gt;What running a fleet of autonomous trading bots taught me about picking the right metric.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I run a small fleet of automated trading bots in paper mode. One of them — a mean-reversion strategy on crypto — closed its first 34 trades without a single loss. 34 out of 34. A 100% win rate.&lt;/p&gt;

&lt;p&gt;If I put that number on a landing page, it would sell. It's also close to meaningless. Here's why, and why it changed how I evaluate any autonomous system I put in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  The number is real. The conclusion isn't.
&lt;/h2&gt;

&lt;p&gt;The strategy belongs to the "buy the dip, scale in, exit in profit" family. Two mechanics produce that spotless record:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;It scales into losers.&lt;/strong&gt; When a position moves against it, it adds to the position at a lower price, dragging down the average entry. A trade that would have been a 4% loss becomes a 0.5% win after two add-ons.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It only ever closes in green.&lt;/strong&gt; There's no time-based or loss-based exit in the base logic. A position that's underwater simply... stays open. It isn't a loss until it's realized, and it's never realized at a loss.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Put those together and a 100% win rate isn't evidence of skill. It's the &lt;em&gt;definition&lt;/em&gt; of the strategy. The metric is measuring the exit rule, not the edge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the risk actually lives
&lt;/h2&gt;

&lt;p&gt;The losses don't disappear — they move to places the win rate doesn't look:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Locked capital / time under water.&lt;/strong&gt; In a bear backtest, this style held some positions for up to 121 days. Capital busy averaging down a bag is capital that isn't compounding. Win rate says nothing about that opportunity cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tail risk.&lt;/strong&gt; Scaling into a falling asset works beautifully until the asset doesn't come back. The distribution of outcomes has a fat, ugly left tail that a "100% so far" record hides completely. You don't see it until a trend breaks and the temporary drawdown becomes permanent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sequence risk.&lt;/strong&gt; 34 trades in a calm, ranging market tells you how the strategy behaves in a calm, ranging market. That's it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is a knock on the strategy — scaling in is a legitimate approach. The point is that &lt;strong&gt;win rate is the wrong lens&lt;/strong&gt; for it. It's a vanity metric here, the same way "99.9% of requests return 200" is a vanity metric if you never look at the 0.1% that time out and take the checkout flow down with them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule I run by now
&lt;/h2&gt;

&lt;p&gt;Because the flattering metric is worthless, I won't let a bot near real money on the strength of it. The go-live criteria I actually use:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A minimum number of trades&lt;/strong&gt; (I use ≥100), so I'm looking at a distribution, not an anecdote.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Survive at least one real correction in paper mode.&lt;/strong&gt; The entire risk of this strategy lives in a downtrend it can't average out of. If it hasn't lived through one, I haven't seen the number that matters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Drawdown, time-under-water, and a Monte Carlo reshuffle of the trade sequence&lt;/strong&gt; — because the order the trades happened to arrive in is one sample, and I want the P95/P99 of the drawdown, not the single lucky path.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Win rate isn't on the list. Neither is total P&amp;amp;L over a too-short window.&lt;/p&gt;

&lt;h2&gt;
  
  
  The general lesson (this isn't really about trading)
&lt;/h2&gt;

&lt;p&gt;The bot is just a clean example of a trap that shows up wherever you operate an autonomous system: &lt;strong&gt;the system will happily hand you the metric that looks best, and it's usually the one that hides the failure mode.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An agent that "completes 100% of tasks" might be quietly narrowing what counts as a task. A pipeline with "zero errors" might be swallowing them. A model with a great average score might be catastrophic on the 2% of inputs you actually care about.&lt;/p&gt;

&lt;p&gt;The engineering job isn't to collect the flattering number. It's to design the evaluation so the failure mode has to show itself &lt;em&gt;before&lt;/em&gt; it costs you. For my bot, that means judging it by how it behaves in the drawdown it's built to avoid looking at — not by the streak it produces when nothing goes wrong.&lt;/p&gt;

&lt;p&gt;A 100% win rate didn't make me trust the bot. It's the thing that made me go looking for what it was hiding.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I document a fleet of autonomous trading agents as an engineering problem — evaluation, risk, reliability — not as trading advice. No signals, no tips; just what breaks and how I try to catch it first. Follow along if that's your thing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>machinelearning</category>
      <category>trading</category>
    </item>
  </channel>
</rss>
