Results from evaluating Google's TimesFM 3.0 on real market data. Everything strictly causal, everything session confined.
Data: 118,857 bars of Nasdaq-100 futures at 3-second resolution, five sessions, with bid/ask volume. Two targets: price midpoint, and an order-flow series (cumulative delta minus volume signed by candle direction).
1. Direct forecasting. 2,674 anchors, 36-bar horizon.
Price: 7.4% worse than naive, 50.0% direction.
Order flow: 7.4% worse than naive, 54.4% direction.
2. Coarser bars. 1-minute bars, 36-minute horizon. Price 5.4% worse at 50.2%. Order flow 19.2% worse at 44.9%. Aggregation does not create predictability.
3. An anomaly reframing built to suit the model. Fair objection: anomaly detection needs a model of normal, not a forecast of a random walk. The two series mirror each other, so the sum of their rolling z-scores is mean reverting, structurally much closer to the seasonal series where this model scores 98.7%.
Result: the residual was 5.6% worse than naive at 51.6% direction. No more forecastable than the raw series.
4. A detector from the quantiles. Each forecast returns nine deciles, so you can estimate P(up) for each series and combine them.
Agreement with the incumbent method: +0.000. Not weak. None.
Worse: when the detector scored highest confidence, the realized hit rate was 24.7% against a 34.3% base rate. Anti-informative when confident, which is more dangerous than useless.
One apparent exception survived my first round of checking and turned out to be a sampling artifact. That is the next post, and it is the most transferable thing in the series.
Top comments (0)