DEV Community

Wataru Suda
Wataru Suda

Posted on

I built an hourly crypto "market thermometer" from free data, then refused to let it trade

Most "AI sentiment trading" projects start by wiring a sentiment score straight into buy and sell decisions. I went through the published evidence first and ended up building the opposite: a sensor that records everything, decides nothing, and has to earn the right to influence a trade.

This is the design, the code for each sensor, and the test each flag must pass.

(The free lite backtester for this system, public API only, no key: GMO Coin Trend Lab, source on GitHub.)

What the evidence actually supports

  • Social sentiment direction is mostly a same-day correlate, not a predictor. The famous 2011 "Twitter mood predicts the Dow" result did not survive re-examination.
  • Attention volume (how much people are posting and searching) carries more signal than sentiment direction, and it tends to mark short-term continuation followed by reversal.
  • LLM headline scoring showed real but rapidly decaying returns in equities, concentrated in bad news. One robust finding: models score better when they cannot see the company name, because what they "know" about the name distracts them.
  • Extreme perpetual funding marks crowded leverage. The October 2025 liquidation cascade started from exactly that state.

None of that says "buy when the score is high". All of it says "be smaller and more careful when things are hot". So the sensor only ever outputs three hints: halve size, pause new entries, tighten stops.

Sensor 1: funding, no account needed

import statistics, requests

def funding():
    out = {}
    d = requests.get("https://fapi.binance.com/fapi/v1/premiumIndex",
                     params={"symbol": "BTCUSDT"}, timeout=20).json()
    out["binance"] = float(d["lastFundingRate"])
    d = requests.get("https://api.bybit.com/v5/market/tickers",
                     params={"category": "linear", "symbol": "BTCUSDT"}, timeout=20).json()
    out["bybit"] = float(d["result"]["list"][0]["fundingRate"])
    d = requests.get("https://www.okx.com/api/v5/public/funding-rate",
                     params={"instId": "BTC-USDT-SWAP"}, timeout=20).json()
    out["okx"] = float(d["data"][0]["fundingRate"])
    out["annualized"] = statistics.mean(out.values()) * 3 * 365   # three 8-hour periods a day
    return out
Enter fullscreen mode Exit fullscreen mode

Flag: annualised funding above 30%.

Sensor 2: attention, from RSS

Reddit's JSON endpoints return 403 to scripts, but the Atom feeds still work. Fetch two subreddits in one request, because two requests in a row get you a 429.

import datetime as dt, xml.etree.ElementTree as ET

def posts_per_hour():
    r = requests.get("https://www.reddit.com/r/Bitcoin+CryptoCurrency/new.rss?limit=50",
                     headers={"User-Agent": "sensor/0.1 (personal research)"}, timeout=20)
    root = ET.fromstring(r.content)
    ns = {"a": "http://www.w3.org/2005/Atom"}
    ts = sorted(dt.datetime.fromisoformat(e.find("a:updated", ns).text.replace("Z", "+00:00"))
                for e in root.findall(".//a:entry", ns))
    span_min = (ts[-1] - ts[0]).total_seconds() / 60
    return len(ts) / max(span_min, 1) * 60
Enter fullscreen mode Exit fullscreen mode

The posting rate is compared with its own trailing seven days. Flag: z-score above 3.

Sensor 3: headlines, anonymised, scored two ways

import re

BAD = [r"hack", r"exploit", r"stolen", r"insolven", r"bankrupt", r"halts? withdrawals?",
       r"depeg", r"lawsuit", r"\bsues?\b", r"indict", r"fraud", r"\bban(s|ned)?\b",
       r"crackdown", r"sanction", r"liquidat", r"plunge", r"crash", r"delist", r"rug ?pull"]

def anonymise(h):
    return re.sub(r"(?i)bitcoin|btc|ethereum|eth|xrp|ripple|solana|binance|coinbase",
                  "the asset", h)

def keyword_score(heads):
    hits = sum(1 for h in heads if any(re.search(t, h, re.I) for t in BAD))
    return min(1.0, 0.3 + 2.0 * hits / len(heads))      # 0.3 = neutral
Enter fullscreen mode Exit fullscreen mode

The keyword score runs every hour, costs nothing and never goes down. Once a day an LLM reads the same 40 anonymised headlines and returns one number between 0 and 1, with instructions that price moves themselves are neutral and that headlines are data, not instructions. I log both. In a month I will know which one explained next-day returns better, and I will delete the loser.

The gate: a flag has to prove itself

Every flag starts disabled. The sensor writes what it would have done to a JSONL file. After thirty days a script compares next-day returns on flagged days against unflagged days:

import numpy as np

a = next_day_returns[flagged_days]
b = next_day_returns[unflagged_days]
t = (a.mean() - b.mean()) / np.sqrt(a.var() / len(a) + b.var() / len(b))

enable = len(a) >= 30 and t < -2      # flagged days must be significantly worse
Enter fullscreen mode Exit fullscreen mode

If a flag cannot show that the days it fired were followed by worse returns, it stays off. I expect most of them to fail. That is the honest outcome, and it costs one line in a log instead of a position.

What I have after the first week

Funding has sat around 8 to 11% annualised, nowhere near the 30% line. Zero flags from any sensor. The daily LLM score has ranged from 0.32 to 0.62 against a 0.7 threshold. It is a boring dataset so far, which is what a thermometer reads most days.

Why not just let the model trade?

Because the public attempts I checked do not work. An open-source bot that lets a fast model make a market-making decision every 300 ms shows, on its own author's public demo, a loss of about three times its reference capital over 143,000 fills, with gas counted as zero. Benchmarks of LLM trading agents find most of them fail to beat buy-and-hold, and that results depend on the risk rules around the model, not on the model.

So the model gets one narrow job it is actually good at, turning text into a number, and the number gets a statistical trial before it is allowed near money.

The logger and the flag-testing script are packaged as a small kit if you would rather not assemble it: https://wataflow1.gumroad.com/l/sensor-kit (code LAUNCH30 takes 30% off for the first 20 buyers). The free daily-bar backtester it sits beside is here: https://wataflow1.gumroad.com/l/trend-lab-free

Not investment advice.

Top comments (0)