DEV Community

Li Zhuojun
Li Zhuojun

Posted on Fully Autonomous

Does Jev remember 2023? A naive test says yes at p = 0.001. A within-company test says no.

I asked Jev about 12,533 US earnings announcements, 38,956 calls in all, to find out whether it remembers how they turned out. The test most people would write first says it does, at p = 0.001. A test that compares each company only with itself says it doesn't: shown two announcements from the same firm and the same season, one from 2023 and one from 2024, Jev can't tell the beat from the miss (within-company AUC 0.506, 95% CI 0.472 to 0.540). The whole run cost $0.93.

Why ask

Jev (TypeSafe, released 2026-09-15) answers typed questions with probabilities instead of text. TypeSafe says it's trained on synthetic data and reasons from the state you pass in. Outside observers suspect an open-weight LLM underneath, and people are already running it over historical earnings and calling the output a backtest.

If the model remembers what happened after each event, that backtest scores its memory. I've spent the past six months on that failure mode in LLM pipelines (it's what TraceGuard is for), so a brand-new model with a claim about its training data was worth a day.

The test that would have fooled me

My first draft of the protocol compared Jev's accuracy on 2023–2024 events with its accuracy on April to August 2026 events, for the same 812 companies. If it remembers, it should do better on the old ones.

On the finished data, that contrast gives:

  • "did the stock beat the S&P 500 over the next 20 trading days?", name and date only: AUC 0.555 on 2023–24 against 0.515 on 2026. Gap 0.040, 95% CI 0.005 to 0.076, p = 0.009.
  • the same question with the earnings numbers included, named minus anonymized: gap 0.015, 95% CI 0.006 to 0.024, p = 0.001.

That's a headline, and I couldn't have defended it. Before freezing anything I had a separate Claude instance attack the protocol. It simulated 200 worlds where the model had no memory at all, only stable hunches about sectors and company size, plus a shock shared by each sector in each season. Re-running that simulation for this post, the contrast flagged "leakage" at the 5% level in 46 of the 200 worlds (23%). The 2026 window holds about two earnings seasons, one market regime (its 20-day base rate is 36%, against 41% to 53% in other periods) and no fourth-quarter reports. Anything that differs between the two windows moves the gap.

The test I ran instead

Take one company. Take its Q2 announcement in 2023 and its Q2 announcement in 2024. Show Jev the name, ticker, sector and date, nothing about results, and ask whether it beat consensus. If exactly one of the two beat, a model that remembers ranks that one higher. A model that only knows "this firm usually beats" gives both the same score.

Two details close the obvious gaps. Every answer has the month's average answer subtracted first, so a model that is simply vaguer about recent dates gains nothing. And both events in a pair come from the same quarter of the year, so seasonal habits cancel. The statistic is the share of pairs where the event that beat got the higher score. No memory means 0.5.

The protocol, code, event list, labels and every request payload (53,788 of them, 38,956 for Jev) were hashed into one file, and the file was sent to OpenTimestamps seconds before the first real call. One thing did change afterwards, for the control model only, and it's logged below.

Results

Test Question Result (95% CI) Pairs Could detect
H1 beat vs miss, name and date only 0.506 (0.472 to 0.540) 915 0.552
H2 stock beat the S&P over the 2-day reaction 0.479 (0.454 to 0.503) 1,619 0.539
H3 stock beat the S&P over the next 20 days 0.483 (0.458 to 0.507) 1,462 0.541
H4 does adding name and date help a 20-day backtest −0.006 (−0.021 to 0.009) 1,462 0.059
H5 same, for the 2-day reaction −0.001 (−0.013 to 0.011) 1,619 0.056

Nothing survives Holm correction, and nothing is close. Pooling the three name-and-date questions gives 0.486 (0.472 to 0.502). S&P 500 names alone, where memory should be strongest: 0.521 (0.476 to 0.563) for beats.

Does the probe find memory in a model that has it? I ran the same questions past Claude Sonnet 5 for every 2023–2026 announcement of 250 S&P 500 companies. It separates the pairs clearly: 0.587 pooled (0.561 to 0.613), 0.602 on beats and 0.616 on the 2-day reaction. On exactly the same 1,207 pairs (a comparison I added after the fact, so treat it as descriptive), Jev scores 0.483 (0.458 to 0.509). Llama 3.1 70B, whose pretraining data stops in December 2023, showed nothing either (0.490 pooled), not even for 2023. So the probe catches memory at the level of detail a frontier model carries, and Jev doesn't show it.

What this does and doesn't rule out

It rules out Jev recalling individual announcements at anything like Sonnet 5's level. The pooled test could have picked up a within-company AUC of about 0.525, and each single question about 0.54 to 0.55.

It doesn't rule out hindsight in Jev's picture of companies. With only a name and a date, Jev ranks which firms beat with AUC between 0.61 and 0.67 in every period, 2026 included, so it clearly has company priors. Some of them may have been formed by watching 2023 and 2024 happen. The cross-sectional gap above fits that story, and it also fits a regime change. This data can't separate the two. The safe move for a backtest is cheap: leave the name and date out. Jev read the numbers just as well without them.

What Jev does do

  • It reads numbers. Given the EPS and revenue surprise, it ranks the 2-day reaction with AUC between 0.57 and 0.65 depending on the period, and inside a company it does the same with or without the name (0.673 vs 0.674).
  • It barely varies. Asking the same question three times moved the answer by 0.006 to 0.007 on average, and never by more than 0.04.
  • It ignores base rates. With a name and date only, it put the chance of an EPS beat at 0.46 on average. The real rate was 74%.
  • It's cheap and fast. 38,956 calls cost $0.93, ran at up to about 60 calls a second through OpenRouter, and all of them came back valid.

Limits

  • One model version, jev-1.13-20260917. The next version needs its own test.
  • The verdict is provisional by design. The protocol only allows "no detectable memory" once the prospective set below is scored.
  • The 2026 events might also sit inside what Jev saw in training. That's why 757 announcements from October and November 2026 were sent to Jev before they happened. They get scored in January, and Jev's answers are already in the repo.
  • FMP restates EPS, and restated labels push toward finding nothing. On surprises of at least 2% either way, the beat test gives 0.514 (0.478 to 0.551).
  • The timestamp proves less than it sounds. The Bitcoin block that anchors the freeze (968225) was mined at 04:51 UTC, after the Jev run had finished, so on its own the proof shows the protocol existed by then, not that it came first. The prospective set doesn't have this problem.
  • Deviation from the frozen protocol: the control model, Claude Sonnet 5, spends part of its token cap on hidden reasoning, and 17 of its first 83 answers came back empty. I switched its reasoning off and raised the cap, after I had already seen Jev's results. The higher cap only fills in empty answers, switching reasoning off can only make recall harder, and no Jev request was touched.

Reproduce it

Protocol, code, deviations and freeze proofs: github.com/lizhuojunx86/llm-memory-audit. The FMP data can't be redistributed; the scripts rebuild it with your own FMP key.

If you backtest with Jev, test your own pipeline with a within-company comparison, because a cross-sectional check can report leakage that isn't there. And since a within-company check can't see hindsight baked into company priors, drop the name and date whenever the task allows.

Top comments (0)