DEV Community

Cover image for A cheap trend-following ETH/USDC bot with two models in the loop: what paper trading showed
Dimon
Dimon

Posted on Originally published at dimonb19a.hashnode.dev AI-assisted

A cheap trend-following ETH/USDC bot with two models in the loop: what paper trading showed

I built a small trading bot and ran it on paper for a few days. It did not make money, and I do not think it can in its current form. I am publishing it anyway, with the code, the sessions and the arithmetic, because most "trading bot" posts share neither.

Repository: github.com/dimonb19a/trend-switcher (MIT). The session summaries are in SESSIONS.md there.

What it is

The bot holds one position, all of it either in ETH or in USDC, and switches sides when several independent checks agree. It trades ETH/USDC on Base through Uniswap v3. By default it runs on paper: live prices and live on-chain quotes, but a virtual capital and simulated fills.

It is not a signal service, not a backtest and not a stop-loss. Every limit it has stops new actions; none of them exits a held position. If it holds ETH and ETH falls, it keeps the drawdown until the machinery decides to switch.

How a decision is made

Every 60 seconds:

  1. Code computes features: EMA20 against EMA50 on 5-minute candles, recent returns, volatility, and the real cost of switching right now (a DEX quote for the whole position plus gas).
  2. A judge answers four fixed questions. The judge is a model that returns probabilities for fixed-choice questions over an HTTP API, not a chat model writing prose. The questions: the multi-hour regime (trend up, down, range), the direction over the next 15 minutes, the quality of the moment, and the risk-off probability. You bring your own provider; the repository names none.
  3. Votes. A judgment votes for a side or it does not. The default preset trend needs five agreeing judgments out of the last six, within nine minutes. trend-fast needs three of four. forecast votes only on a 15-minute move big enough to beat the switching cost.
  4. Filters and a veto. The candidate must agree with the EMA trend. Then a second model (DeepSeek) reads the same state and may veto, at most once per five minutes.
  5. Hard limits, checked right before any effect: switches per day, minimum hold, daily loss halt, total loss kill, inference budget, data freshness. A halt latches until an operator resets it.

Models only vote or veto. Code decides, and every limit holds whatever a model says.

What it costs

Inference is the only running cost, and it is a fixed number of dollars per month, not a percentage of the capital. At a 60-second tick, both models together cost me a few dollars a month. So the question is simple division:

Capital Inference as a share of capital (at $5 a month)
$100 5 %
$1,000 0.5 %
$10,000 0.05 %

On $100 the models eat the capital faster than any plausible edge. On $10,000 inference is dust, and the only question left is whether there is an edge at all. That question does not depend on the capital, and the sessions below answer it for what I tried: no.

How it loses

A trend follower earns when it catches moves bigger than its switching bill and pays the bill every time the trend was not there. The bill has two parts.

The round trip itself is small: the pool fee twice, the slippage twice, gas on Base is cents. About 0.15 % of the position.

The real cost is the move between a flip and its reversal. A regime call flips at the edges of a sideways range by construction: near the bottom the judge sees a trend down and the bot sells; the range holds; near the top the judge sees a trend up and the bot buys back. The bot has sold low and bought high by the width of the range it mistook for a trend. In my sessions each such episode cost 0.6–0.9 % against simply holding.

The forecast rule has the opposite problem: a 15-minute call must beat the round trip with room to spare, so the judge would have to be right on direction well over two thirds of the time. It was not, and it knew it. It almost never rated a 15-minute move as likely enough, so under that rule the bot never traded.

The paper sessions

Paper, a virtual capital of $1,000, starting in ETH, a 60-second tick. "Result" is the end of the window after simulated execution costs; "hold" is what keeping ETH would have done over the same window.

# Window (UTC) Preset · second model Slots Switches Result Hold Ended
1 2026-09-29, 30 min forecast · forecast 30 / 30 0 −0.295 % same (no switch) smoke run
2 2026-09-30 00:01 → 04:08 forecast · forecast 248 / 720 0 not reported (one judge bill stayed unknown) — stopped by the unknown-bill guard, as designed
3 2026-09-30 09:31 → 21:31 forecast · forecast 720 / 720 0 −0.08 % same (no switch) deadline
4a 2026-09-30 19:00 → 10-01 03:00 trend (5 of 6) · forecast 480 / 480 2 −0.64 % +0.28 % deadline
4b same window trend-fast (3 of 4) · forecast 480 / 480 2 −0.58 % +0.28 % deadline
5a 2026-10-01 15:33 → 20:33 trend · forecast 300 / 300 2 −0.09 % +0.83 % deadline
5b same window trend · no second model 300 / 300 2 −0.07 % +0.88 % deadline
5c same window trend · regime 300 / 300 2 −0.10 % +0.82 % deadline
6 2026-10-02 00:11 → 00:41 trend · forecast 30 / 30 0 −0.043 % same (no switch) smoke of the release

Three things the table says:

  • Under the forecast rule the bot never traded in sixteen hours of judgments. No judgment qualified as a vote. That is the rule working, not a measurement of the model.
  • Under the trend rule every arm switched exactly twice and every arm lost to holding. Session 4 is the textbook whipsaw: sold near the night's low, bought back near its high, 0.9 % behind hold. The faster window changed the timing by minutes, not the outcome. Session 5 repeated it in a narrower range.
  • The second model did not change the outcome. With it, without it, and with a different question, the arms made the same two switches. The lever is the regime call itself flipping at the edges of a range, not the veto.

Session 5 was launched three times; the first two halted at the first switch on a public RPC error, which is recorded in the repository. The arms of sessions 4 and 5 shared the hours but not identical inputs, so they are exploratory comparisons, not controlled trials. And six sessions over four days are a sample of a few market regimes, not a distance.

What the models actually do, from watching them

The judge never sees a chart. It sees a few hundred words of numbers (the price, the two EMAs, the recent returns, the volatility, the cost of switching) and answers four fixed questions with probabilities. After thousands of answers, two habits stand out.

On the 15-minute question it is cautious almost all the time. It rarely rates a move as likely enough to beat the cost, and on a flat day that caution is simply correct: there was no move to call. On the regime question it is decisive, and that is where the whipsaw comes from. A regime opinion asked every minute will flip at the edges of a range, because from inside a range the edge looks exactly like the start of a trend. The second model, asked whether it agrees with a switch, mostly said no to buying back and yes to selling; it delayed switches by a few minutes and changed none of the outcomes.

None of this is a score for either model. It is what a one-minute cadence does to any forecaster, human or not: it asks for a trend opinion far more often than trends change. The fix is in the question and the cadence, not in a better model.

Where the upside is

The honest statement first: the bot has no proven edge, and I am not claiming one. But the part that was expensive to build is not the strategy.

The foundation is the feed, the on-chain quotes, the paper fills at real quotes, a ledger that cannot invent a bill or a slot, the hard limits, the judge integration, the presets and the timed-session harness. None of it cares which strategy sits on top. The strategy is the smallest part of the repository: one policy that turns the judge's answers into a vote, the questions themselves, and a handful of thresholds. Replace that, and everything else keeps working and keeps accounting.

That is where the money could be, in theory. Trend following pays when moves are larger than the switching bill and last longer than the bot's reaction time. At a 60-second cadence on a sideways day there are no such moves. At a horizon of hours or days there may be, and the same machinery can be pointed at that horizon by changing the judge's question and the tick, not by rewriting the bot. With a capital where inference is dust, the strategy is the only variable left, and the paper harness lets anyone test a candidate over a real distance with the same accounting I used, before a single real dollar moves.

Three things I would try first, as hypotheses:

  • A longer horizon. A 5-minute tick, a direction question over hours, holds of hours. Same round trip, bigger moves to catch, twenty times fewer model calls.
  • Partial positions, so a wrong flip costs a fraction of the range.
  • A range filter that sits out when the regime is range instead of trading its edges.

None of these is tested. The knobs are in the repository (VOTE_WINDOW, VOTE_MIN, the thresholds, the second model's question, the limits), so anyone can run the experiment they believe in, on paper, with their own keys.

The repository

MIT, Node 24, one SQLite ledger per run, an offline test suite. Every model call is journaled before it is sent and its bill stays UNKNOWN until the provider confirms it; a timed session records planned, attempted and missed slots, so a sleeping laptop produces missed slots, not invented ones. To run on paper you need your own judge provider (an API origin, a pinned model id, a key) and a DeepSeek key. Live mode is capped by code and has never been armed; the repository does not encourage it.

If you find a configuration that beats holding over a distance, I would like to hear about it, with the ledger.

Top comments (1)

Collapse
 
arhancanli profile image
Arhan Canli •

The "five of the last six within nine minutes" rule is worth a second look, because it probably gives less protection than it seems. Six judgments a minute apart see almost the same state (the 5-minute candles and the EMAs barely change between ticks), so their errors are highly correlated. Five agreeing votes is then closer to one opinion repeated five times than to five independent checks, which would explain why trend and trend-fast made the same two switches a few minutes apart.

If you try again, the cheaper fix is probably a price condition rather than more votes: switch only after the price has moved past the recent range by more than the episode cost you measured (0.6-0.9%), or require the regime call to persist across several 5-minute candles rather than several 1-minute ticks. Both make the votes depend on new information. Your table already gives the target: a flip has to be worth more than about 0.9% on average to beat holding, and that number is a good bar to write down before the next paper run.