Before pointing Google's TimesFM 3.0 at anything real, I ran a calibration baseline. Two synthetic series with known properties, same model, same settings, held out.
Seasonal with drift: 98.7% better than the naive baseline, 100% direction accuracy.
Random walk: 9.8% worse than naive, 47% direction.
The naive baseline is the crudest forecast available: next value equals last value.
Row two is the important one, and it is not a defect.
A random walk is unpredictable by construction. Each step is independent of every step before it. A model that appeared to forecast one would be reporting structure that does not exist. In a backtest that looks like skill. In production it looks like losing money.
The correct behaviour on an unpredictable series is to fail, ideally about as badly as naive. That is what happened.
Read together the two rows are a map.
Where it works: demand forecasting, energy load, web traffic, call volume, inventory, capacity planning, sensor telemetry. Domains where next week resembles last week plus a trend.
The commercial argument there is not really accuracy, it is that you do not fit a model per series. Ten thousand SKUs, ten thousand zero-shot forecasts, about 8 ms each.
Where it does not: anything closer to a random walk than to a seasonal series.
I then spent two days confirming that market tape sits firmly on the second side of that line.
Top comments (0)