DEV Community

Michael Hairetis
Michael Hairetis

Posted on

TimesFM 3.0 on 118,857 bars of futures tape: four framings, all worse than naive

Results from evaluating Google's TimesFM 3.0 on real market data. Everything strictly causal, everything session confined.

Data: 118,857 bars of Nasdaq-100 futures at 3-second resolution, five sessions, with bid/ask volume. Two targets: price midpoint, and an order-flow series (cumulative delta minus volume signed by candle direction).

1. Direct forecasting. 2,674 anchors, 36-bar horizon.
Price: 7.4% worse than naive, 50.0% direction.
Order flow: 7.4% worse than naive, 54.4% direction.

2. Coarser bars. 1-minute bars, 36-minute horizon. Price 5.4% worse at 50.2%. Order flow 19.2% worse at 44.9%. Aggregation does not create predictability.

3. An anomaly reframing built to suit the model. Fair objection: anomaly detection needs a model of normal, not a forecast of a random walk. The two series mirror each other, so the sum of their rolling z-scores is mean reverting, structurally much closer to the seasonal series where this model scores 98.7%.

Result: the residual was 5.6% worse than naive at 51.6% direction. No more forecastable than the raw series.

4. A detector from the quantiles. Each forecast returns nine deciles, so you can estimate P(up) for each series and combine them.

Agreement with the incumbent method: +0.000. Not weak. None.

Worse: when the detector scored highest confidence, the realized hit rate was 24.7% against a 34.3% base rate. Anti-informative when confident, which is more dangerous than useless.

One apparent exception survived my first round of checking and turned out to be a sampling artifact. That is the next post, and it is the most transferable thing in the series.

https://michaelhairetis.medium.com/thousands-of-bars-of-futures-tape-several-framings-one-consistent-result-a65e89073ca8

Top comments (0)