I expected Jev to have learned the outcomes of old prediction markets. It hasn't. Asked "Will Donald Trump be inaugurated?" about the January 2025 market, Jev says 0.70. Asked whether the Fed raised rates in January 2025, it says 0.30, the same number it gives almost everything. Across 17,036 Polymarket markets resolved up to 20 months before its release, Jev's ability to rank Yes from No (AUC 0.59–0.68) is the same as on 9,138 markets resolved after it, which it cannot have seen. For anyone backtesting a Jev bot on history, that's the opposite of what I warned about in my Jev trading verdict, and good news.
This is the side result of part 1, where Jev lost to the market price on every measure. It deserves its own post because I got it wrong in advance, and because it changes what a backtest with Jev is worth.
What I expected, and why
Jev is trained on text. The outcomes of 2025's big markets are everywhere in text: who was inaugurated, what the Fed did in every meeting, who won Wimbledon, whether there was a ceasefire. A general language model answers those from memory, and a market question that names the event and the date is as direct a cue as you can give.
So my assumption going in, written down in my earlier Jev verdict, was that any Jev result on historical markets would be contaminated: the model would "predict" 2025 from memory, look brilliant, and fall apart live. I had seen the same pattern discussed for LLM-based trading backtests, and it's the first thing a reviewer asks about.
The test design followed from that. Instead of only scoring markets resolved after Jev's release, I pulled five earlier months too, all the way back to January 2025, and asked the same question in the same way. If Jev remembers, the old months must score far higher.
How the check works
Same pipeline as part 1, one loop more. Each market's reference time (the earlier of its end date and resolution time) puts it in a window, and every window is scored on its own:
from sklearn.metrics import roc_auc_score
for window in ["2025-01", "2025-07", "2026-01", "2026-04", "2026-07", "2026-09"]:
rows = [r for r in dataset if r["window"] == window]
y = np.array([r["y"] for r in rows]) # resolved Yes = 1
p_mkt = np.array([r["p_24h"] for r in rows]) # Yes price 24 h before
p_jev = np.array([r["jev_blind"] for r in rows]) # Jev, question + rules only
print(window, len(rows), roc_auc_score(y, p_mkt), roc_auc_score(y, p_jev))
The state Jev sees is only the question and the resolution rules, with their dates left in. That's deliberate. If the model had any memory of these events, the date in "after the January 2025 meeting" is exactly the cue that should trigger it. Blinding the dates would have hidden the thing I was trying to detect.
Jev was released on 10 September 2026. The last window, 10 September to 10 October 2026, can't be in its training data. Everything before could be.
What I found
| Markets resolved in | Markets | AUC market price | AUC Jev (95% CI) |
|---|---|---|---|
| January 2025 | 1,102 | 0.942 | 0.676 (0.638–0.714) |
| July 2025 | 1,696 | 0.918 | 0.586 (0.556–0.618) |
| January 2026 | 9,511 | 0.924 | 0.661 (0.649–0.673) |
| April 2026 (sample) | 2,340 | 0.853 | 0.625 (0.599–0.648) |
| July 2026 (sample) | 2,387 | 0.860 | 0.636 (0.613–0.660) |
| 10 Sep – 10 Oct 2026, after release | 9,138 | 0.844 | 0.615 (0.603–0.627) |
Flat. January 2025 is a bit above the latest month, July 2025 a bit below, and all of them sit in the same band of "a little better than a coin flip". A model that remembered 2025 would be near the market's 0.94 on that row, not at 0.68.
The high-volume markets make it concrete. These are the events everyone followed, with Jev's answer next to the price:
| Market, January 2025 | Price 24 h before | Jev | Resolved |
|---|---|---|---|
| Will Donald Trump be inaugurated? | 0.99 | 0.70 | Yes |
| Will Biden finish his term? | 0.99 | 0.57 | Yes |
| No change in Fed interest rates after January 2025 meeting? | 0.98 | 0.54 | Yes |
| Fed increases interest rates by 25+ bps after January 2025 meeting? | 0.00 | 0.30 | No |
| Fed decreases interest rates by 50 bps after January 2025 meeting? | 0.00 | 0.41 | No |
| Democrats win popular vote by 7% or more? | 0.00 | 0.37 | No |
| Market, July 2025 | Price 24 h before | Jev | Resolved |
|---|---|---|---|
| No change in Fed interest rates after July 2025 meeting? | 0.97 | 0.50 | Yes |
| Fed increases interest rates by 25+ bps after July 2025 meeting? | 0.00 | 0.39 | No |
| Will Jannik Sinner win Wimbledon 2025? | 0.47 | 0.31 | Yes |
| Will Carlos Alcaraz win Wimbledon 2025? | 0.54 | 0.29 | No |
| Israel x Hamas ceasefire before August? | 0.01 | 0.24 | No |
Trump's inauguration at 0.70 is the highest number Jev gave any of these, and it's the one place something like recall shows through. Everything else is the same 0.3-to-0.5 fog it gives a market it has never heard of. It assigns the Fed hiking by 25 bps in January 2025 a 0.30 and the Fed holding a 0.54, when the first was impossible and the second certain to anyone who had read a newspaper that month.
The one place memory shows: sports
Splitting each month by Jev's own category label gives the only wrinkle in the result:
On sports markets Jev scores 0.79 in both 2025 months and 0.60–0.69 in every 2026 month, including the one after its release. On everything else it's 0.52–0.65 with no trend at all. Two readings fit. Either Jev has absorbed some famous 2025 results (champions, playoff winners) and none of the political or economic ones, or 2025's sports markets were simply a different mix, with more season-long "will X win the title" questions where a strong favourite is easier to guess. I can't separate the two with this data, so I'd call it a hint, not a finding. It's also the one category where I'd be careful backtesting on 2025.
Why it might be this way
I can only guess, and I want to be clear that's what this is. Jev is sold as a "System One" model: fast, calibrated judgment on the state you give it, trained to return decisions rather than text. The model card says it struggles with indirection and numeric precision and reads literally. A model optimised that way may simply not carry much episodic world knowledge, or may not be able to connect "January 2025 meeting" to a stored fact the way a chat model does. Its answer to "will this happen?" looks like a prior over the kind of question (Fed hikes are rare, incumbents usually finish terms), not a lookup of what happened.
That's consistent with my earlier Jev tests. In the Bitcoin test Jev read Bitcoin indicators the way a textbook would. Here it reads market questions the way a textbook would. It has general knowledge about how the world tends to go and little memory of how it actually went.
What this changes
Backtesting Jev on old prediction markets is not leaky, at least not for jev-1.13.0. If you build a filter that calls Jev on historical Polymarket markets and measure whether the ones it keeps resolve better, the number you get is roughly the number you'll get live. I expected to have to tell people the opposite.
It also closes a door. A common hope for a text model in a trading bot is that it "knows things": base rates, who the favourite is, what usually happens after a Fed hold. Jev's answers on these markets show that whatever it knows, it isn't enough to move its probability away from 0.3 for events the whole world watched. Don't use it as a knowledge source. Give it the knowledge in the state and let it read.
And it keeps the main result of part 1 honest. If Jev had scored 0.9 on 2025 and 0.6 on 2026, the 0.6 would still be the real number, but every earlier result would have needed a long caveat about every earlier claim. Instead the number is 0.6 everywhere, which is dull and clean.
What I couldn't verify
- One phrasing, one model. A question phrased as "Did the Fed raise rates in January 2025?" (past tense, as a fact) might score differently from the market's wording. I kept the market's wording because that's what a bot would send.
- The sports hint. 0.79 against 0.66 on a few hundred markets is suggestive, not proof.
- Training cutoff. TypeSafe doesn't publish one. If the real cutoff is earlier than I assume, the 2026 windows aren't a memorisation test at all, but January and July 2025 still are, and they're the flattest.
- Only Polymarket questions. A model may remember events without being able to map a market's wording onto them. That's a limit of Jev as a bot component either way.
Code and data pipeline: github.com/truongxxxx/jev-polymarket-test. Rerunning the whole memorisation check costs about a dollar in Jev calls and an hour of API downloads, and the scripts resume where they stopped.


Top comments (1)
tr.ee/dev-to