The backtest validation engine source code is available here
I took two paid Telegram signal channels, read a month of their posts (August 2026) through an LLM, and replayed every signal on one-minute candles with commissions and slippage. The first channel produced +10.3% for the month, the second â18.1%. Both figures are without leverage. The difference comes down to the post format, and it's visible even before running the backtest.
Why bother
Imagine a trader who is about to pay for Telegram signals and follow them. What does he need insurance against?
A stop-loss by itself doesn't protect anything. You can set a stop that gets hit 100% of the time and lose your deposit while formally following risk management. Whether the author's rules actually work can only be shown by a replay on history: real posts, one-minute candles, commissions, slippage. The output is a number you can verify.
That's what I did. I picked channels with opposite formats:
- Channel A publishes a complete trading plan in every post.
- Channel B builds its pitch on personal discipline: "I take 1% of the deposit per trade", "I accept losses and move on."
Results
| Channel A | Channel B | |
|---|---|---|
| Signal format | entry range + 5 targets + stop | "shorting BTC" + a screenshot |
| Trades per month | 21 | 24 |
| Result (no leverage, net of commissions) | +10.3% | â18.1% |
| Winning trades (win rate) | 76% (16 of 21) | 29% (7 of 24) |
| Profit factor | 1.78 | 0.57 |
| Average position duration | ~17 hours | ~44 hours |
| Worst trade | â4.6% (author's stop) | â7.9% (our stop; the author has none) |
Same market, same month, same model in the pipeline â and a 28-percentage-point difference. Below I'll break down how it was computed and where the gap comes from.
The LLM parses the posts, the code makes the decisions
The key design decision in this project: the model makes no trading decisions. It got the task LLMs are actually good at â reading unstructured data.
A post like "Open LONG in the $78600 - $79400 range, targets..., stop $77400" is text with typos, emojis, and ad inserts. Sometimes the signal isn't text at all but a terminal screenshot with the word "Short" in the image. Turning that into typed JSON is the model's job. Deciding whether to enter a position and where to cut losses is the code's job.
import { generateObject, InferenceName, type FormatModel } from "json-inference";
const TradingPositionFormat = {
type: "object",
required: ["id", "symbol", "position", "entryRange", "targets", "stoploss", "reasoning"],
properties: {
id: {
type: "number",
description: "ID of the source message of the signal. If there is no signal, return -1",
},
position: {
type: "string",
enum: ["long", "short", "wait"],
description: "Position type. If there is no position, return wait",
},
entryRange: {
type: "object",
required: ["from", "to"],
properties: { from: { type: "number" }, to: { type: "number" } },
},
targets: {
type: "array",
description: "Position targets, 5 levels. If there is no signal, return []",
items: { type: "number" },
},
stoploss: { type: "number" },
reasoning: { type: "string", description: "Rationale - for debugging" },
},
} satisfies FormatModel;
How reliable is this
I compared two generations of the cache on identical inputs. The verdicts and all numeric fields matched bit for bit. Only the free text in reasoning differed, down to the language of the response. The grammar holds the structure together; the model copies the prices straight from the post. For a parser, nothing more is needed.
The pipeline: post -> signal -> trade
The strategy ticks every minute on every symbol. Calling the LLM and Telegram on every tick would produce thousands of unnecessary requests and a flood ban. So signal retrieval is wrapped in Cache.file: for Channel A, the channel is checked once every 5 minutes per symbol.
The cache is stored in Mongo. There's a useful side effect: the measure-items collection becomes a dataset of all model responses for the month.
import { Cache } from "backtest-kit";
import { scrapeLookback } from "telegram-reader";
const getSignal = Cache.file(
async (symbol: string, when: Date) => {
const coin = `#${symbol.replace("USDT", "")}`;
// a 15-minute window back from the backtest's "now".
// A post published after `when` is invisible to the strategy
let messages = await scrapeLookback({
channel: "channel_a",
limit: 15,
dimension: "minute",
when,
});
messages = messages.filter(({ content }) => content.includes(coin));
if (!messages.length) return { entry: null }; // we don't call the LLM
const entry = await generateObject(
InferenceName.OllamaInference,
{
format: TradingPositionFormat,
messages: [
{ role: "user", content: PARSER_PROMPT }, // prompt string with post-parsing rules
...messages.map(({ id, date, content }) => ({
role: "user" as const,
content: `ID ${id}\n\n${content}\n\n[${date.toISOString()}]`,
})),
],
},
"gemma4:31b-cloud",
);
return { entry, message };
},
{ interval: "5m", name: "channel_a_entry_v3" },
);
Channel A: discipline in the post structure
Every signal post is a complete trading plan:
SIGNAL #BTC/USDT
đ Open LONG in the $78600 - $79400 range - 6% of deposit, X15 leverage
đ Targets: $80300 / $80900 / $82000 / $83200 / $84500
âī¸ STOP LOSS: $77400
The range, the five targets, and the stop are present in every post throughout the month. A signal like this can be executed by code with nothing left to guess. That's what discipline looks like in machine-readable form.
How the strategy executes it
- Entry â at market on any post no older than 15 minutes. The signal's edge lives close to the moment of publication.
- Take-profit â at the furthest target.
- Stop-loss â from the post.
- Position management â "level ratcheting". Every reached target becomes a floor. If price pulls back beyond the last reached level by more than 30% of the step between targets, the position is closed.
The step is computed from the target grid of the specific post, so the tolerance adjusts itself: PENGU has steps of roughly 0.3â0.9%, while BTC's are hundreds of dollars.
listenActivePing(async ({ data, currentPrice }) => {
const levels = data.payload.levels as number[]; // targets from the post
const currentLevel = getCurrentLevel(data.position, levels, currentPrice);
const { lastLevel } = await LEVEL_STATE.getState(); // state of the specific position
if (currentLevel > lastLevel) {
await LEVEL_STATE.setState({ lastLevel: currentLevel }); // new floor
return;
}
if (!lastLevel) return; // no target reached yet - waiting for stop or take
const drift = getLevelDrift(data.position, levels, lastLevel, currentPrice);
const step = getLevelStep(levels, lastLevel); // step between adjacent targets
if (drift > step * LEVEL_DRIFT_RATIO) {
await commitClosePending(data.symbol); // pulled back beyond a reached level - exit
}
});
Outcome
21 trades, +10.3% without leverage net of commissions, 16 of 21 in profit.
- The author's stop fired 3 times: from â3.8% to â4.6% â a reasonable size.
- Price reached the furthest take-profit once.
- The ratchet closed all the remaining trades.
Where the profit came from
I did not copy the author's own money management. His plan â scaling out across five targets with X15â25 leverage â is unprofitable on this very data. The distant targets are almost never reached, and a single stop at that leverage burns 60â100% of margin.
The profit emerged at the intersection: the author's structured signal plus our risk management. The author's direction makes money; his leverage burns the deposit.
Channel B: discipline in words only
Channel B's format is the opposite. There are no ranges, targets, or stops. There are posts like:
Opening a position
Working BTC - short, entered with a possible add-on in mind
Closing a position
So, we're shorting BTC, ETH, SOL at 1% of the deposit each,
BTC and ETH at 100x leverage, Solana at 50x.
The direction and ticker are often visible only in an image: a position screenshot, a TradingView chart, or an exchange trade card. So the model receives not just the post text but the image as well:
const signal = await generateObject(
InferenceName.OllamaInference,
{
format: PositionOpenFormat, // { id, symbol, position: long|short|wait, reasoning }
messages: [
{ role: "user", content: promptOpen },
...messages.map(({ id, date, content, photo }) => ({
role: "user" as const,
content: `ID ${id}\n\n${content}\n\n[${date.toISOString()}]`,
...(photo && { images: [photo] }), // the post's image in base64 - part of the signal
})),
],
},
"gemma4:31b-cloud",
);
I only wrote the image-recognition rules after reviewing the channel's actual screenshots. There was one trap the model fell into regularly: a green P&L number on a short means profit, not "long". The direction is determined only by the word "Long/Short" or by the position of the take and stop zones on the chart. That rule is now spelled out explicitly in the prompt.
How the strategy executes it
The author publishes no stops, so the stop is ours: 7.5% from the entry price.
For this channel, entry and exit are separate decisions made at different times. So there are two pipelines:
- while there is no position, the model searches the posts for an entry;
- while a position is open, the model searches for the post where the author closes it.
addStrategySchema({ strategyName: "main_strategy" }); // empty schema, no getSignal
listenIdlePing(async ({ symbol, when, currentPrice }) => { // ticks while there is no position
const { signal } = await getOpenSignal(symbol, when); // Cache.file, 4h
if (!signal || signal.position === "wait") return;
await commitCreateSignal(symbol, {
id: `${signal.id}-${signal.symbol.toLowerCase()}`,
symbol: signal.symbol,
...Position.moonbag({
position: signal.position,
currentPrice,
percentStopLoss: 7.5, // our own stop - the author doesn't publish any
}),
});
});
listenActivePing(async ({ symbol, when }) => { // ticks while a position is open
const { signal } = await getCloseSignal(symbol, when);
if (signal?.action === "close") await commitClosePending(symbol);
});
Outcome: 24 trades, â18.1%, 7 of 24 in profit.
Conclusions
Run the numbers yourself!
Before paying for signals â let alone following them â verify the author on history. It's cheaper than any insurance. The entire analysis above is built on two backtest.jsonl files. Each line is one event in a position's lifecycle. The basic metrics take about ten lines to compute:
import { readFileSync } from "node:fs";
const trades = readFileSync("backtest_channel_a.jsonl", "utf8")
.split("\n").filter(Boolean).map(JSON.parse)
.filter((e) => e.data.action === "closed");
const pnl = trades.reduce((sum, e) => sum + e.data.pnl, 0);
const wins = trades.filter((e) => e.data.pnl > 0).length;
const grossProfit = trades.filter((e) => e.data.pnl > 0).reduce((s, e) => s + e.data.pnl, 0);
const grossLoss = trades.filter((e) => e.data.pnl <= 0).reduce((s, e) => s + e.data.pnl, 0);
console.log({
trades: trades.length,
totalPnl: pnl.toFixed(2), // +10.26
winRate: (wins / trades.length * 100).toFixed(0) + "%", // 76%
profitFactor: (grossProfit / -grossLoss).toFixed(2), // 1.78
});
Every closed event contains:
- entry and exit prices;
- the close reason:
take_profit,stop_loss,closed, ortime_expired; - the maximum profit and maximum drawdown of the position;
- in the
notefield â the model's full response and the source post.
Any trade can be rewound to the specific Telegram post, and the model's reasoning can be read. The method ports to any other channel in an evening.
Tools
- backtest-kit â a backtest and live-trading engine with look-ahead protection.
- telegram-reader â reading Telegram channels via MTProto.
- json-inference â structured output for ollama.
- backtest-ollama-casual â the repository with both strategies and the analysis scripts.
Top comments (0)