Every forecaster publishes their hits. Almost nobody publishes their misses. That asymmetry is the single biggest reason performance claims in finance are worthless.
We publish a pre-open probability band for the major Indian indices every trading day, and we score every one of them publicly afterwards — including the days we were wrong.
Why the usual approach is unfalsifiable
The standard pattern is familiar. Post a call. If it works, screenshot it. If it does not, say nothing. Over a few months you accumulate an impressive gallery of correct calls and an invisible graveyard of incorrect ones.
Nothing here is a lie. Every screenshot is real. The claim is still meaningless, because the sample is selected after the outcome is known.
Backtests have the same problem, one layer deeper. Given enough parameters and enough attempts, any strategy can be made to look profitable on historical data. The overfitting is not usually deliberate. It is what happens when you keep adjusting until the curve looks good.
A record that cannot show you the failures cannot be evaluated. It is marketing shaped like evidence.
What we do instead
Each trading day, before the market opens, a band is published for NIFTY 50, BANK NIFTY and SENSEX. After the session, the actual outcome is scored against it and written to a permanent dated receipt.
A receipt records the predicted band, the actual value, whether the band held, whether the directional lean was right, and the confidence attached at the time.
The important detail is the ordering. The prediction is committed before the outcome exists. That is what makes it a test rather than a description.
Receipts are permanent and unedited. They accumulate whether the week went well or badly.
What a miss looks like
A real one from the record:
BANK NIFTY
band 56,178.40 – 57,047.94
actual 56,796.95
band result HIT
bias BULLISH
direction WRONG (failure mode: SIGN_MISS)
confidence 0.26
The band held and the direction was wrong. Both facts are in the receipt.
Note the confidence: 0.26. The model was not confident, said so in advance, and was wrong. That is the system behaving correctly. A model that is wrong while claiming certainty is broken. A model that is wrong while flagging low conviction is working as designed, and you can only tell the difference if the confidence was recorded beforehand.
The uncomfortable part
Publishing your misses means people can see them.
Some will screenshot a bad day. Some will conclude the model is useless from a sample of one. That cost is real, and it is worth paying, because the alternative is asking people to trust a number they cannot check.
It also disciplines us. When every miss is permanent and public, there is no incentive to quietly widen the bands to make the hit rate look better — the widening would be as visible as the misses.
A range, not a call
One deliberate design choice: the output is a range, not a direction.
A range is something you can size a position against. A direction is an instruction, and issuing instructions to retail traders is not something we do. We are an infrastructure provider, not a research analyst, and the distinction is regulatory as well as ethical.
There is a difference between here is the plausible range and here is how often that range has held and buy this. The first is information. The second is advice, and it is not ours to give.
Daily bands: dashboard.aiondashboard.site/veritas
Every receipt, hits and misses: dashboard.aiondashboard.site/veritas/validation
Informational and educational only. Not investment advice. Not a SEBI-registered Research Analyst product. Past calibration does not guarantee future coverage.
Top comments (0)