When More Semantic Signals Made My AI Judge Worse
I have been building a small web game called Sideways.
It is a lateral-thinking puzzle game: the player sees a mysterious situation and tries to discover what really happened by asking questions.
A typical interaction should feel almost trivial:
Player:
Is it light that wakes the guest?
Game Master:
Yes.
Or:
Player:
Does the wake-up mechanism make sound?
Game Master:
No.
The interesting part is that the Game Master is powered by AI.
I chose Jev because I did not want a chatbot generating long answers. I wanted a decision layer that could take application state, make a typed judgment, and let ordinary TypeScript decide what the product should do.
Jev describes itself as a decision model rather than a chat model. Its API works with application state and typed questions such as Choice, Score, and Noul, returning structured results and probabilities rather than conversational prose.
That sounded almost perfect for this game.
But I discovered something unexpected:
The more semantic intelligence I tried to extract from the model, the worse the actual game experience became.
This article is about that failure.
And why I am now making the AI do less.
The Architecture I Wanted
From the beginning, I did not want the model to control the application directly.
The architecture was:
Player text
↓
Probabilistic semantic judgment
↓
Typed output
↓
Deterministic TypeScript policy
↓
Player-facing result
The principle was simple:
AI = semantic judgment
Code = deterministic policy
User = consequential intent
The model should understand language.
My code should decide what happens next.
This has several advantages.
The browser never needs to receive the hidden solution.
The model does not get to decide arbitrary UI behavior.
The public API can expose only simple results such as:
YES
NO
PARTLY
IRRELEVANT
UNCLEAR
while keeping the private reasoning signals on the server.
For solution attempts, the public results are similarly constrained:
SOLVED
PARTIAL
NOT_SOLVED
Conceptually, I still believe this architecture is correct.
The problem was how much semantic structure I asked the model to produce.
Version 1: Simple Evidence
The first Question Judge was relatively simple.
For a player's question, the model estimated something like:
{
wellFormed: number,
relevant: number,
evidence: {
entailed: number,
contradicted: number,
undetermined: number
}
}
TypeScript then converted these probabilities into a product result.
For example:
if (wellFormed < 0.8) return "UNCLEAR";
if (relevant <= 0.2) return "IRRELEVANT";
if (entailed >= 0.8) return "YES";
if (contradicted >= 0.8) return "NO";
return "UNCLEAR";
This had an important property:
A low probability of YES did not automatically mean NO.
An unknown fact could remain unknown.
That part worked well.
But actual gameplay revealed another problem.
Natural player language is messy.
People write things like:
light wake him?
does it make sound
maybe lamp?
They make spelling mistakes.
They omit articles.
They use pronouns.
They do not write propositions like lawyers.
So I tried to make the judge more sophisticated.
Version 2: More Semantic Intelligence
In Version 2, the Question Judge expanded into:
{
wellFormed,
relevant,
evidence: {
entailed,
contradicted,
undetermined
},
mixedClaims: {
supported,
contradicted
}
}
This allowed a new result:
PARTLY
For example:
There is no alarm. Only a lamp wakes the guest.
contains both a false claim and an important true claim.
I did not want that to be treated as simply NO.
For solution checking, I went even further.
Instead of judging a theory with a simple checklist, the model returned five semantic dimensions:
{
coreMechanism,
causalRelation,
supportingInsight,
contextualError,
competingMechanism
}
The idea was reasonable.
Consider these three theories:
There is a light.
The light wakes the guest.
There is a light, but vibration wakes the guest.
They should not receive the same result.
I wanted:
"There is a light."
→ PARTIAL
"The light wakes the guest."
→ SOLVED
"Light exists, but vibration wakes the guest."
→ NOT_SOLVED
So I separated:
- identifying the important object;
- understanding the causal mechanism;
- making a harmless contextual mistake;
- asserting a competing wrong mechanism.
Architecturally, this looked much better.
And the Offline Tests Looked Excellent
The implementation passed:
880 unit/API tests
82 desktop/mobile E2E tests
Leak scans
Lint
Type checking
Build
Full repository checks
I also created separate calibration and engineering-generalization corpora across all six puzzles.
The browser did not receive private grading data.
The hidden solution remained server-only.
The number of provider calls remained bounded.
Everything looked clean.
But there was one very important limitation.
Most of those tests proved:
semantic signal
→ parser
→ deterministic policy
→ API result
→ UI
They did not prove that the real model would infer the expected semantic signal from previously unseen player language.
That distinction turned out to matter a lot.
Then I Played the Game with the Real Model
I deployed the new judge to a protected preview environment and started typing ordinary questions.
The results were surprising.
Example 1
Player:
is it light?
Result:
Not enough information
Example 2
Player:
is it light that wakes the guest?
Result:
Not enough information
Then I submitted the exact same sentence as a theory:
Theory:
is it light that wakes the guest?
Result:
Not quite yet.
A human Game Master would almost certainly understand the intended idea.
The hidden solution is that a visual light signal wakes the sleeping guest.
Example 3
Player:
does the wake up mechanism make sound?
Result:
Please rephrase
Even this typo-heavy version:
does the wake up mechines make sound?
also produced:
Please rephrase
That was a serious UX problem.
The grammar is imperfect, but the semantic intent is obvious.
Example 4
Player:
Instead of an alarm, lights from a device wakes the guest.
Question result:
Not enough information
Theory result:
Not quite yet.
Again, the wording is imperfect.
But the important causal insight is clearly present.
Why Did More Structure Make Things Worse?
I cannot infer Jev's internal reasoning from its output alone.
So the following are hypotheses about my decision architecture, not claims about the internals of the model.
Several problems stood out.
1. The Semantic Dimensions Were Not Really Independent
I had separated concepts like:
well-formedness
relevance
truth
mixed claims
causal understanding
because it made the TypeScript architecture easier to reason about.
But humans do not necessarily understand a sentence in that order.
Take:
does the wake up mechines make sound?
A human probably performs something closer to:
"mechines" is probably "mechanism"
↓
The player means the wake-up mechanism
↓
They are asking whether it makes sound
↓
I know the answer
↓
No
This is one integrated semantic interpretation.
I had converted it into several semi-independent judgments.
That created more opportunities for one signal to disagree with the others.
2. Probabilities Became Hard Cliffs
The intention behind probabilities was to avoid brittle Boolean decisions.
But then my code contained rules such as:
if (wellFormed < 0.8) {
return "REPHRASE";
}
Imagine the model essentially understands the sentence but produces:
wellFormed = 0.77
The rest of the semantic evidence no longer matters.
The application suddenly says:
Please rephrase.
A continuous probability had become a discrete cliff.
Adding more semantic dimensions meant adding more places where such cliffs could occur.
The model did not necessarily need to be dramatically wrong.
My policy only needed one signal to fall on the wrong side of a threshold.
3. I Was Asking a Small Decision Model to Solve a Large Semantic Task
Jev's documentation emphasizes focused, well-scoped typed decisions. Choice is designed for predefined classification, while Noul represents a yes/no judgment whose probability can be interpreted by application code.
But my Question Judge was effectively asking the model to do all of this at once:
Recover imperfect language
Resolve pronouns
Determine the semantic proposition
Determine relevance
Compare with hidden truth
Detect multiple claims
Separate supported and contradicted claims
Distinguish uncertainty from contradiction
The individual fields looked small.
The total semantic task was not.
I had decomposed the output without necessarily decomposing the difficulty.
4. The Judge Knew the Scenario, but Not the "Mystery Focus"
This turned out to be especially interesting.
Consider:
is it light?
As an isolated English sentence, it is ambiguous.
What is "it"?
The alarm?
The mechanism?
The object?
But in a lateral-thinking game, a human Game Master has another piece of context:
What mystery is the player currently trying to explain?
For the quiet-alarm puzzle, that focus is roughly:
What causes the sleeping guest to wake
without sound or physical contact?
Given that context:
is it light?
has a very natural interpretation:
Is light the cause?
I had provided the scenario and hidden reference, but not an explicit representation of the question being investigated.
That may have encouraged overly literal ambiguity handling.
5. My Solution Judge May Have Confused Completeness with Correctness
The Version 2 Solution Judge required strong scores for both:
coreMechanism
causalRelation
before returning SOLVED.
That sounds reasonable.
But consider:
The light wakes the guest.
This is extremely short.
It also contains both the mechanism and causal relationship.
If the model gives:
coreMechanism = high
causalRelation = medium
the deterministic policy returns PARTIAL.
From the application's perspective, that is a false negative.
The player has solved the mystery.
The grading representation made the answer look less complete than it actually was.
Version 3: Make the Model Do Less
So I am now testing a different architecture.
Instead of asking for many independent semantic scores, each judge will return one categorical probability distribution.
For Question:
type QuestionSemanticClass =
| "SUPPORTED"
| "CONTRADICTED"
| "MIXED"
| "UNKNOWN"
| "IRRELEVANT"
| "AMBIGUOUS";
Then ordinary TypeScript maps it:
SUPPORTED
→ YES
CONTRADICTED
→ NO
MIXED
→ PARTLY
UNKNOWN
→ NOT ENOUGH INFORMATION
IRRELEVANT
→ NOT RELEVANT
AMBIGUOUS
→ PLEASE REPHRASE
There is still probabilistic information.
But there is only one semantic decision.
No chain of:
0.8
0.8
0.2
0.8
...
gates.
The Solution Judge Is Being Simplified Too
Instead of five independent dimensions, I am testing one semantic classification:
type SolutionSemanticClass =
| "CORE_CAUSAL_EXPLANATION"
| "CORE_WITH_CONTEXT_ERROR"
| "SUPPORTING_INSIGHT_ONLY"
| "COMPETING_WRONG_MECHANISM"
| "NO_MEANINGFUL_INSIGHT";
TypeScript then maps:
CORE_CAUSAL_EXPLANATION
→ SOLVED
CORE_WITH_CONTEXT_ERROR
→ PARTIAL
SUPPORTING_INSIGHT_ONLY
→ PARTIAL
COMPETING_WRONG_MECHANISM
→ NOT_SOLVED
NO_MEANINGFUL_INSIGHT
→ NOT_SOLVED
The important distinction remains.
There is a light.
should not equal:
The light wakes the guest.
And neither should equal:
The vibration wakes the guest.
But the model no longer needs to independently estimate five continuous semantic dimensions before my application can make that distinction.
I Am Also Adding a "Mystery Focus"
Each puzzle will have a small server-only field such as:
What causes the sleeping guest to wake
without sound or physical contact?
This is not the hidden answer.
It simply represents the question raised by the public scenario.
The hope is that it helps resolve natural shorthand:
is it light?
does it make sound?
was he covered?
without adding phrase-specific exceptions.
That distinction matters.
I do not want:
if text === "is it light?"
I want:
short player language
+
current mystery
→ recoverable semantic proposition
And If Version 3 Still Fails?
Then I plan to simplify again.
Possibly all the way to:
YES
OTHER
For Question judging:
YES
would mean:
This player proposition is supported by the hidden explanation.
Everything else would collapse into:
OTHER
That would intentionally stop distinguishing among:
NO
UNKNOWN
IRRELEVANT
AMBIGUOUS
This loses information.
It may still produce a better game.
That is the part of this experiment I find most interesting.
The Product Does Not Need the Maximum Amount of AI Information
It is tempting to think:
More semantic signals
→ more information
→ better product decisions
My experience so far suggests something more like:
More semantic signals
→ more uncertain boundaries
→ more policy interactions
→ potentially worse UX
The best AI interface may not be the one that extracts the most information from a model.
It may be the one that asks the model for the smallest decision the product actually needs.
Another Lesson: Offline Tests and Semantic Accuracy Are Different Things
I had hundreds of passing tests.
They were useful.
They caught:
- schema mistakes;
- API regressions;
- accidental hidden-data leaks;
- extra provider calls;
- incorrect deterministic mappings;
- broken UI behavior.
I absolutely want those tests.
But:
synthetic semantic signals
→ correct product output
does not prove:
previously unseen human language
→ correct semantic signal
Those are different systems.
For AI products, both need testing.
I now think of them as two separate validation layers:
Layer 1:
Software correctness
Layer 2:
Model behavior under real language
A green CI pipeline proves the first.
It does not automatically prove the second.
Why I Still Like the Typed-Decision Architecture
None of this has made me want to replace the system with a free-form chatbot.
Actually, the opposite.
I still want:
AI
→ bounded semantic decision
TypeScript
→ application policy
Jev's API is explicitly designed around application state and typed decisions, with probability distributions that software can consume.
The lesson for me is not:
Typed decisions are too simple.
It is:
I should respect their simplicity.
If the application needs one classification, asking for one well-designed Choice may be better than constructing a small semantic ontology and connecting it with threshold gates.
Where the Experiment Stands
At the time of writing:
v1
Simple probabilistic evidence
↓
v2
More semantic dimensions
↓
Real-model testing exposed brittle behavior
↓
v3
One categorical semantic distribution per judge
+ explicit mystery focus
↓
Currently being tested
If v3 performs well on previously unseen wording, I will freeze the Judge and move on to production hardening.
If it does not, I will test the YES / OTHER design rather than adding more semantic complexity.
That constraint is intentional.
At some point, engineering discipline means stopping optimization.
Final Takeaway
The most useful lesson from this project so far is surprisingly simple:
Do not ask the AI to make a more complicated decision than your product actually needs.
A sophisticated semantic representation can be elegant in TypeScript.
It can have clean types.
It can have excellent unit tests.
It can even look more theoretically correct.
And still produce a worse experience for the person actually playing the game.
Sometimes the better AI architecture is not:
Understand more.
It is:
Decide less.
I will update this article once the simplified Version 3 Judge has been tested against the real model.
About the project
Sideways is an experimental lateral-thinking puzzle game built with Next.js, TypeScript, and Jev.
The project explores a specific architecture:
Probabilistic Semantic Judgment
→ Typed Output
→ Deterministic TypeScript Policy
→ Product Result
The next question is how small that semantic judgment can become while still producing a good game.
Top comments (0)