I have been building a small lateral-thinking puzzle game called Sideways.
The basic interaction is simple.
The player sees a mysterious situation and asks questions until they discover what really happened.
For example:
Player:
Is it light?
Game Master:
Yes.
Or:
Player:
Does the wake-up mechanism make sound?
Game Master:
No.
The Game Master is powered by Jev.
I deliberately did not want a general-purpose chatbot generating explanations.
I wanted something much closer to this:
Natural-language input
↓
Probabilistic semantic decision
↓
Typed result
↓
Deterministic TypeScript
↓
Game behavior
After several iterations, I realized that this is almost like building probabilistic IF statements over natural language.
And I also learned something I did not expect:
Making those semantic IF statements more sophisticated actually made the game worse.
The fix was not to make the AI smarter.
The fix was to make the decision smaller.
The original idea: AI understands meaning, code controls behavior
The architecture follows a simple rule:
AI = semantic judgment
Code = deterministic policy
User = consequential intent
The model does not directly decide which UI to show.
The browser does not receive the hidden puzzle solution.
The model produces a bounded semantic judgment, and TypeScript turns that judgment into a product result.
For questions, the public result can be:
YES
NO
PARTLY
IRRELEVANT
UNCLEAR
For theory submissions:
SOLVED
PARTIAL
NOT_SOLVED
This separation still feels right to me.
The mistake was not the architecture itself.
The mistake was how much semantic structure I tried to extract from one player sentence.
Version 2: the judge became too clever
In Version 2, a Question judgment looked roughly like this:
{
wellFormed,
relevant,
evidence: {
entailed,
contradicted,
undetermined
},
mixedClaims: {
supported,
contradicted
}
}
Most of those values were probabilities between 0 and 1.
Then TypeScript applied rules such as:
if (wellFormed < 0.8) {
return "REPHRASE";
}
if (relevant <= 0.2) {
return "IRRELEVANT";
}
if (relevant < 0.8) {
return "NOT_ENOUGH_INFORMATION";
}
if (
mixedClaims.supported >= 0.8 &&
mixedClaims.contradicted >= 0.8
) {
return "PARTLY";
}
if (evidence.entailed >= 0.8) {
return "YES";
}
if (evidence.contradicted >= 0.8) {
return "NO";
}
The Solution Judge was even more detailed:
{
coreMechanism,
causalRelation,
supportingInsight,
contextualError,
competingMechanism
}
Again, deterministic TypeScript combined these scores.
The goal was reasonable.
I wanted to distinguish:
"There is a light."
→ PARTIAL
from:
"The light wakes the guest."
→ SOLVED
and from:
"There is a light, but vibration wakes the guest."
→ NOT_SOLVED
Conceptually, Version 2 was much more expressive.
The code also looked disciplined.
And the software tests looked excellent
Version 2 passed:
880 unit/API tests
82 desktop/mobile E2E tests
Leak scanning
Lint
Type checking
Production build
I had calibration data.
I had a separate engineering-generalization corpus.
Hidden grading metadata stayed server-side.
Provider calls remained bounded.
Everything looked good.
Then I used the real model.
The real game felt worse
One puzzle is about a guest waking without sound or physical contact.
The hidden mechanism is a visual light signal.
I tried this:
is it light?
Version 2:
Not enough information
Then:
is it light that wakes the guest?
Version 2:
Not enough information
I submitted the exact idea as a theory:
is it light that wakes the guest?
Version 2:
Not quite yet.
That was already worrying.
Then I tried:
does the wake up mechanism make sound?
Result:
Please rephrase
And even:
does the wake up mechines make sound?
produced:
Please rephrase
The word mechines is obviously a typo.
But the intended meaning is still easy for a human to recover.
At this point the architecture was technically clean but the actual Game Master felt strangely rigid.
What may have gone wrong
There was probably no single cause.
In fact, Version 3 changed several things at once, so I cannot claim that any one of them independently fixed the problem.
But four design problems became obvious.
1. The 0.8 thresholds may have been too conservative
This was my first suspicion.
Suppose Jev had actually returned something like:
wellFormed = 0.74
relevant = 0.92
contradicted = 0.89
for:
does the wake up mechines make sound?
That would mean the system understood quite a lot.
But my code did this first:
if (wellFormed < 0.8) {
return "REPHRASE";
}
So all the useful information below that gate became irrelevant.
The probability was continuous.
My application policy turned it into a cliff.
0.7999 → rejected
0.8000 → accepted
A lower threshold might have improved the result.
I still think this is a plausible explanation.
But Version 2 had several thresholds, so lowering one would not necessarily solve the whole problem.
2. Several uncertain decisions were connected with hard gates
The deeper problem was not just that 0.8 might have been too high.
There were several such boundaries.
The system effectively became:
IF recoverable enough
AND relevant enough
AND supported enough
AND not contradicted enough
AND ...
THEN YES
Each semantic dimension could be individually reasonable.
But the final behavior depended on all of them interacting correctly.
I had replaced one uncertain AI decision with several uncertain AI decisions and connected them using deterministic cliffs.
That made the overall product more brittle.
3. I decomposed things that humans understand together
Consider:
does the wake up mechines make sound?
A human probably does not consciously evaluate:
grammar quality
→ relevance
→ referent resolution
→ semantic truth
as separate numerical decisions.
We do something more like:
"mechines" probably means "mechanism"
↓
They mean the wake-up mechanism
↓
They are asking whether it makes sound
↓
No
It is one semantic interpretation.
Version 2 decomposed that interpretation into multiple scores.
That was elegant from a TypeScript perspective.
It may have been unnatural from a semantic-decision perspective.
4. The judge knew the scenario, but not explicitly what mystery was being solved
This became one of the most useful changes.
Take:
is it light?
As isolated English, it is ambiguous.
But this is not isolated English.
The player is solving a specific mystery.
For the quiet-alarm puzzle, the actual question being investigated is approximately:
What causes the sleeping guest to wake
when the alarm makes no sound
and nothing touches the guest?
A human Game Master naturally interprets:
is it light?
as:
Is light what causes the guest to wake?
Version 2 had the scenario and hidden answer, but it did not have an explicit representation of the mystery focus.
That changed in Version 3.
Version 3: one semantic decision per judge
Instead of trying to improve Version 2 by adding more examples, more thresholds, or more scores, I removed complexity.
The Question Judge now makes one categorical semantic decision.
Conceptually:
type QuestionSemanticClass =
| "SUPPORTED"
| "CONTRADICTED"
| "MIXED"
| "UNKNOWN"
| "IRRELEVANT"
| "AMBIGUOUS";
Jev returns one probability distribution over those choices.
Then TypeScript performs a deterministic mapping:
SUPPORTED
→ YES
CONTRADICTED
→ NO
MIXED
→ PARTLY
UNKNOWN
→ NOT ENOUGH INFORMATION
IRRELEVANT
→ NOT RELEVANT
AMBIGUOUS
→ PLEASE REPHRASE
There is no active:
wellFormed >= 0.8
gate anymore.
There is no chain of independent semantic Nouls.
There is one semantic classification.
The Solution Judge was simplified in the same way
Version 2 had five independent signals.
Version 3 has one classification:
type SolutionSemanticClass =
| "CORE_CAUSAL_EXPLANATION"
| "CORE_WITH_CONTEXT_ERROR"
| "SUPPORTING_INSIGHT_ONLY"
| "COMPETING_WRONG_MECHANISM"
| "NO_MEANINGFUL_INSIGHT";
The product mapping remains deterministic:
CORE_CAUSAL_EXPLANATION
→ SOLVED
CORE_WITH_CONTEXT_ERROR
→ PARTIAL
SUPPORTING_INSIGHT_ONLY
→ PARTIAL
COMPETING_WRONG_MECHANISM
→ NOT_SOLVED
NO_MEANINGFUL_INSIGHT
→ NOT_SOLVED
So I did not abandon semantic distinctions.
I reduced the number of separate decisions required to produce them.
I also added mysteryFocus
Every puzzle now includes a small server-only description of what the player is trying to explain.
For example:
What causes the sleeping guest to wake
when the alarm makes no sound
and nothing touches the guest?
This does not contain the answer.
It only gives the judge the same contextual frame that a human Game Master naturally has.
That helps with short language such as:
is it light?
does it buzz?
is it moving?
without creating hard-coded phrase exceptions.
I shortened the shared rubric
Another change was less visible but important.
Earlier rubrics had gradually accumulated semantic instructions and special distinctions.
Version 3 intentionally moved back toward a smaller question:
Which semantic category best describes this player's statement in this puzzle?
The rubric still tells the model to tolerate:
non-native English
minor spelling mistakes
missing articles
telegraphic wording
ordinary shorthand
natural pronoun resolution
But it does not ask for several independent measurements of those properties.
Again:
Decide less.
What happened after the simplification?
The difference was immediate.
Here are real protected-preview results.
Before: Version 2
is it light?
→ Not enough information
is it light that wakes the guest?
→ Not enough information
does the wake up mechanism make sound?
→ Please rephrase
does the wake up mechines make sound?
→ Please rephrase
is it light that wakes the guest?
[Theory]
→ Not quite yet.
Now compare that with Version 3.
After: Version 3
is it light?
→ Yes
is brightness involved?
→ Yes
something bright wake him?
→ Yes
This is especially important.
The grammar is poor:
something bright wake him?
But the intended meaning is recoverable.
The new judge treated it that way.
Then:
does the wake up mechines make sound?
→ No
The typo no longer caused the system to reject the whole question.
Similarly:
does it buzz?
→ No
and:
is there vibration?
→ No
Short questions also became usable.
Mixed statements started behaving correctly too
I tried:
something bright wake him?
does the wake up mechines make sound?
The result was:
Partly
with the UI:
Some of that is right, but another part isn't.
This is exactly what PARTLY was intended to mean.
Not uncertainty.
Not closeness.
Actual mixed semantic truth.
Most importantly, concise correct theories started solving the puzzle
Question:
does light wake him?
Result:
Yes
Then I reused the exact same text as a theory:
does light wake him?
Result:
You got it.
Another phrasing:
the room gets bright and that wakes him
Question:
Yes
Theory:
You got it.
That is much closer to how a human Game Master should behave.
Wrong causal mechanisms still failed
Simplification did not mean blindly accepting anything containing the correct keyword.
For example:
lamp is there but vibration wakes him
Question:
No
Theory:
Not quite yet.
That matters because false SOLVED results are much worse for this game than conservative partial judgments.
So far, Version 3 became more tolerant without immediately destroying the distinction between correct and incorrect causal explanations.
The behavior also generalized to another puzzle
I tested a different puzzle involving an apparently empty frame whose contents seem to change.
Some examples:
is the frame big
→ Not relevant
is it an animal
→ No
is it a TV
→ No
is it a machine
→ No
Then:
it changes its color with sun light
→ Yes
Submitted as a theory:
it changes its color with sun light
→ You got it.
Again, the grammar is not perfect.
But the causal idea is correct.
And the judge accepted it.
Version 3 is not perfect
There is still some instability around boundary cases.
For example, during one session:
is it the sunrise and sunset
→ Yes
while similar wording later produced:
is it sunrise and sunset
→ Not enough information
That tells me the system is not magically deterministic at the semantic level.
Nor should I expect it to be.
The important difference is that the major player-facing failures changed from:
"I clearly understand what you're asking,
but please rephrase."
to occasional disagreement around genuinely fuzzy boundaries.
For a proof-of-concept game, that is a much better failure mode.
So what actually fixed it?
The honest answer is:
I do not know which individual change fixed it.
Version 3 changed several things together:
Removed multiple continuous semantic gates
Removed the 0.8/0.2 decision chain
Replaced many Nouls with one Choice distribution
Added mysteryFocus
Shortened and generalized the rubric
Made language recovery part of the single semantic classification
Any one of those may have contributed.
The improvement is evidence that the overall architecture became better.
It is not a controlled experiment proving that 0.8 alone was the problem.
A useful future experiment would compare Version 2 under several threshold policies.
That could answer:
Was Jev already producing useful semantic probabilities that my application was throwing away?
I think that is entirely possible.
Jev started feeling less like "adding AI" and more like building probabilistic IF statements
This project changed how I think about decision models.
I am not really building a chatbot.
I am building something conceptually closer to:
if (semanticCondition) {
doSomething();
}
Except the condition is not:
x > 10
It is:
Does this player's sentence semantically express
a proposition supported by the hidden explanation?
Jev estimates that semantic condition.
TypeScript decides what happens next.
That is why I increasingly think of this pattern as:
probabilistic IF statements over natural language
or:
an AI-powered switch/case
Version 2 accidentally turned that simple idea into a giant conditional expression.
Version 3 moved it back toward a small semantic switch.
There is also a hidden engineering cost: calibration
Another lesson was that API pricing is not the only cost of using a decision model.
For this project I needed to repeatedly:
design semantic categories
test real player wording
test spelling mistakes
test short questions
test false mechanisms
test incomplete solutions
check false positives
check false negatives
change policies
retest on protected Preview
That is human labor.
Jev reduced one kind of complexity for me.
I did not have to parse arbitrary assistant prose into application state.
But that complexity did not disappear.
Some of it moved into:
decision design
calibration
threshold selection
test-play
behavioral evaluation
I would therefore not say:
Jev requires more engineering work than a traditional LLM.
I have not run a controlled comparison that proves that.
A more accurate statement is:
In this project, Jev moved a meaningful part of the engineering effort from output handling to decision design and calibration.
That is an important cost to account for.
849 tests still did not replace test-play
Version 3 currently passes:
849 unit/API tests
82 desktop/mobile E2E tests
Leak scanning
Lint
Type checking
Build
Those tests are valuable.
They prove things such as:
the parser rejects invalid distributions
the deterministic mapping works
private grading data does not reach the browser
provider call counts remain bounded
the UI still behaves correctly
But they still cannot fully prove:
A human writes an unseen sentence
↓
The model understands it as intended
That requires actual behavioral testing.
For AI systems I now think about validation as two separate layers:
Layer 1
Software correctness
Layer 2
Model behavior
A green CI pipeline is excellent evidence for Layer 1.
It is not sufficient evidence for Layer 2.
Why I am stopping here
There are still things I could tune.
I could add more categories.
I could add confidence gates.
I could add more examples.
I could create exceptions for phrases like:
sunrise and sunset
I am deliberately not doing that.
That path is exactly how Version 2 became complicated.
For this PoC, Version 3 is now a strong freeze candidate.
The remaining requirement is not:
perfect semantic classification
It is:
Natural short player language usually works.
Minor grammar and spelling errors usually work.
Correct causal explanations can solve the puzzle.
Wrong causal explanations do not solve it.
Hidden information stays hidden.
The game remains fun.
That is enough.
What if this still breaks later?
I already have a fallback.
Simplify even further:
YES
OTHER
OTHER would intentionally collapse:
NO
UNKNOWN
IRRELEVANT
AMBIGUOUS
possibly MIXED
That would lose information.
But if it produced a better game experience, I would seriously consider it.
Again:
The goal is not to extract the maximum amount of semantic information from the model.
The goal is to ask for the smallest useful decision.
Final takeaway
The most surprising lesson from building Sideways has been this:
More semantic structure does not automatically produce better AI behavior.
Version 2 looked more sophisticated.
It had more probabilities.
More semantic axes.
More explicit thresholds.
More detailed grading logic.
And worse real gameplay.
Version 3 asked Jev to do less:
One semantic Choice
↓
One probability distribution
↓
Deterministic TypeScript switch
And the Game Master got noticeably better.
So the principle I am taking forward is:
Do not ask AI to make a more complicated decision than your product actually needs.
Or even shorter:
Understand enough.
Decide less.
That turned out to be a much better architecture for this game.
Top comments (0)