This is a submission for the Kaggle Benchmarking Challenge
My first evaluation gave DeepSeek, Claude, and Gemini perfect reported scores across ten cases each.
That should have been reassuring. Instead, it made me question the benchmark. Were these models genuinely good at maintaining a user's information across a conversation, or had I made the cases too easy to distinguish between them?
I wanted to investigate a deceptively important failure mode: an AI assistant can make the right update and still get the user's overall state wrong.
Changing a budget is straightforward. Changing a budget without accidentally changing someone's location, career goal, available time, or long-term preferences is a different problem.
The problem: updating without losing context
Imagine a user whose learning budget is ₹5,000. Later, they confirm that their scholarship has been approved, increasing the budget to ₹10,000.
The budget should change. Their location should not. Neither should their career goal, learning preferences, or available time.
Now make the conversation less straightforward.
The user temporarily has only three hours available because they are ill. They correct an earlier career goal. They mention using a different note-taking application for a college group project. They discuss a possible internship and a move to another city, neither of which is confirmed.
A reliable assistant must distinguish temporary circumstances from lasting changes, corrections from conflicting evidence, and confirmed facts from possibilities. It must also preserve information that the latest message never asked it to change.
I call this problem LLM State Reconciliation.
My central research question is:
Can language models update a structured representation of a user's information across multiple messages without losing important facts, introducing unrelated changes, or treating uncertain events as established facts?
What makes this benchmark different?
Testing whether a model changes only the requested setting is an important problem in its own right. But user information introduces an additional challenge: the meaning of an update depends on context.
A temporary reduction in availability is not necessarily a permanent schedule change. A conditional resignation is not a completed resignation. A tool preference expressed for a team project does not automatically replace a personal preference. A possible relocation should not become a recorded fact before it happens.
These distinctions motivated my benchmark design. Rather than testing only whether unrelated fields remain unchanged, I wanted to examine whether a model understands which changes the conversation actually justifies.
That means evaluating both the update and the state left behind.
From an easy pilot to twelve harder cases
My initial experiments involved DeepSeek, Claude, and Gemini. I tested simpler single-message updates, sequential updates, and longer conversational scenarios.
The initial long-form evaluation produced perfect reported accuracy across all ten cases for each model. I did not treat this as proof that the problem was solved. I treated it as a limitation of the evaluation: the cases were not sufficiently discriminating.
That prompted a harder benchmark.
The H01–H12 suite uses a 25-field baseline and twelve cases designed to test different failure modes:
- Temporary versus permanent changes
- Explicit corrections
- Conflicting evidence
- Conditional future events
- References to earlier messages
- Multiple simultaneous updates
- Task-specific versus general preferences
- Stale information
- Personal versus project-specific preferences
- Uncertain employment outcomes
- Confirmed versus unconfirmed changes
- Compound state reconciliation
I defined the expected state for each case and separated intended updates from fields that should remain untouched.
Across the suite, the scoring design identifies nine intended field updates and 291 protected field slots. That separation matters: a model could correctly update the requested field while silently changing something else.
A metric that measures only the requested update would miss that failure.
What the early evidence tells me—and what it doesn't
One reported DeepSeek observation illustrates why preservation needs to be evaluated separately.
In the location-update case, the intended location changed to Bengaluru, but the budget remained at ₹6,000 from an earlier case rather than returning to the ₹5,000 baseline. This was recorded as a possible cross-case contamination issue.
That observation raises a methodological question as well as a model-behaviour question: was each case supposed to start independently from the same baseline, or was the benchmark testing a deliberately continuous conversation?
The answer changes how the result should be interpreted. Independent tests need independent initialisation; sequential tests need an expected state that carries forward only the changes justified by earlier messages.
The evaluation protocol must make that distinction explicit.
There is an important limitation to my current results. Although I collected additional model responses and summaries, the complete raw outputs were not preserved for every model and case. I therefore cannot independently verify the full H01–H12 preservation and collateral-mutation metrics or present a reliable comparative leaderboard.
I have kept conversation-reported observations separate from scores that can actually be calculated from raw outputs. Missing evidence is not a passing score.
The most useful finding so far is therefore methodological: a benchmark can produce perfect scores without establishing that it measures the difficult behaviour you care about.
How I would evaluate state reconciliation
I separate the evaluation into three core dimensions:
1. Target-update accuracy
Did the model make the change justified by the conversation?
2. Preservation accuracy
Did it leave unrelated fields untouched?
3. Collateral mutations
How often did it change something that should have remained unchanged?
I would also assess temporal reasoning and uncertainty handling separately. A model should not permanently overwrite a preference because of a temporary exception, nor should it treat a conditional plan as a completed event.
For a reproducible comparison, every model should receive the same prompts and baseline, each case should follow a clearly defined independent or sequential protocol, and the complete outputs should be preserved. The same deterministic scoring code should then evaluate each model against the same gold answers.
Implementation and limitations
I created the shared baseline and twelve-case gold-answer structure, implemented scoring and validation code, and added automated tests. The latest development run passed all ten tests. The existing C1–C10 evaluation also remained functional after the project changes.
These are implementation checks, not proof of model performance. The earlier C1–C10 results and the harder H01–H12 suite are distinct evaluations; I do not treat the former as a substitute for missing raw evidence from the latter.
The next iteration would recover or regenerate complete raw outputs, run cases under a documented initialisation protocol, repeat evaluations to measure consistency, and expand the range of scenarios. I would also investigate whether the same failure patterns appear in practical conversational agents that maintain user profiles over time.
My benchmark
Project: LLM State Reconciliation: Updates, Preservation, and Uncertainty
Research focus: Multi-turn reasoning, context tracking, and structured state integrity.
Models in the initial evaluation: DeepSeek, Claude, and Gemini.
Current status: The benchmark design, gold-answer structure, scoring code, and validation tests are implemented. The harder evaluation remains incomplete because the complete raw outputs required for independently verified comparative metrics are missing.
Kaggle publication: I encountered an account-verification issue while attempting to publish the benchmark publicly. A public Kaggle benchmark link is not yet available.
I would rather document these limitations than present an unverified leaderboard as a result. The benchmark's value depends on whether its measurements can be inspected and reproduced, not simply on whether the project has a compelling premise.
The question I want this benchmark to answer
As conversational assistants take on more responsibility for remembering preferences, plans, goals, and constraints, updating information correctly is only half the problem.
The other half is knowing what the update does not authorise the assistant to change.
When an AI assistant changes what it remembers, can we trust it to preserve what we never asked it to change?
That is the behaviour I want LLM State Reconciliation to measure.
Note to the organizers: A Kaggle phone-verification issue has prevented me from publishing the benchmark publicly before the deadline. I understand that a public Kaggle benchmark link is part of the challenge requirements, and I am actively trying to resolve the issue with kaggle support. I would be grateful if the organizers could consider my submission in light of this circumstance or advise whether an alternative verification process is acceptable. Thank you for your consideration.



Top comments (0)