DEV Community

Cover image for When Simpler AI Decisions Worked Better: What I Changed in My Jev-Powered Game Judge
494
494

Posted on

When Simpler AI Decisions Worked Better: What I Changed in My Jev-Powered Game Judge

I have been building a small lateral-thinking puzzle game called Sideways.

The basic interaction is simple.

The player sees a mysterious situation and asks questions until they discover what really happened.

For example:

Player:
Is it light?

Game Master:
Yes.
Enter fullscreen mode Exit fullscreen mode

Or:

Player:
Does the wake-up mechanism make sound?

Game Master:
No.
Enter fullscreen mode Exit fullscreen mode

The Game Master is powered by Jev.

I deliberately did not want a general-purpose chatbot generating explanations.

I wanted something much closer to this:

Natural-language input
        ↓
Probabilistic semantic decision
        ↓
Typed result
        ↓
Deterministic TypeScript
        ↓
Game behavior
Enter fullscreen mode Exit fullscreen mode

After several iterations, I realized that this is almost like building probabilistic IF statements over natural language.

And I also learned something I did not expect:

Making those semantic IF statements more sophisticated actually made the game worse.

The fix was not to make the AI smarter.

The fix was to make the decision smaller.


The original idea: AI understands meaning, code controls behavior

The architecture follows a simple rule:

AI = semantic judgment
Code = deterministic policy
User = consequential intent
Enter fullscreen mode Exit fullscreen mode

The model does not directly decide which UI to show.

The browser does not receive the hidden puzzle solution.

The model produces a bounded semantic judgment, and TypeScript turns that judgment into a product result.

For questions, the public result can be:

YES
NO
PARTLY
IRRELEVANT
UNCLEAR
Enter fullscreen mode Exit fullscreen mode

For theory submissions:

SOLVED
PARTIAL
NOT_SOLVED
Enter fullscreen mode Exit fullscreen mode

This separation still feels right to me.

The mistake was not the architecture itself.

The mistake was how much semantic structure I tried to extract from one player sentence.


Version 2: the judge became too clever

In Version 2, a Question judgment looked roughly like this:

{
  wellFormed,
  relevant,
  evidence: {
    entailed,
    contradicted,
    undetermined
  },
  mixedClaims: {
    supported,
    contradicted
  }
}
Enter fullscreen mode Exit fullscreen mode

Most of those values were probabilities between 0 and 1.

Then TypeScript applied rules such as:

if (wellFormed < 0.8) {
  return "REPHRASE";
}

if (relevant <= 0.2) {
  return "IRRELEVANT";
}

if (relevant < 0.8) {
  return "NOT_ENOUGH_INFORMATION";
}

if (
  mixedClaims.supported >= 0.8 &&
  mixedClaims.contradicted >= 0.8
) {
  return "PARTLY";
}

if (evidence.entailed >= 0.8) {
  return "YES";
}

if (evidence.contradicted >= 0.8) {
  return "NO";
}
Enter fullscreen mode Exit fullscreen mode

The Solution Judge was even more detailed:

{
  coreMechanism,
  causalRelation,
  supportingInsight,
  contextualError,
  competingMechanism
}
Enter fullscreen mode Exit fullscreen mode

Again, deterministic TypeScript combined these scores.

The goal was reasonable.

I wanted to distinguish:

"There is a light."
→ PARTIAL
Enter fullscreen mode Exit fullscreen mode

from:

"The light wakes the guest."
→ SOLVED
Enter fullscreen mode Exit fullscreen mode

and from:

"There is a light, but vibration wakes the guest."
→ NOT_SOLVED
Enter fullscreen mode Exit fullscreen mode

Conceptually, Version 2 was much more expressive.

The code also looked disciplined.


And the software tests looked excellent

Version 2 passed:

880 unit/API tests
82 desktop/mobile E2E tests
Leak scanning
Lint
Type checking
Production build
Enter fullscreen mode Exit fullscreen mode

I had calibration data.

I had a separate engineering-generalization corpus.

Hidden grading metadata stayed server-side.

Provider calls remained bounded.

Everything looked good.

Then I used the real model.


The real game felt worse

One puzzle is about a guest waking without sound or physical contact.

The hidden mechanism is a visual light signal.

I tried this:

is it light?
Enter fullscreen mode Exit fullscreen mode

Version 2:

Not enough information
Enter fullscreen mode Exit fullscreen mode

Then:

is it light that wakes the guest?
Enter fullscreen mode Exit fullscreen mode

Version 2:

Not enough information
Enter fullscreen mode Exit fullscreen mode

I submitted the exact idea as a theory:

is it light that wakes the guest?
Enter fullscreen mode Exit fullscreen mode

Version 2:

Not quite yet.
Enter fullscreen mode Exit fullscreen mode

That was already worrying.

Then I tried:

does the wake up mechanism make sound?
Enter fullscreen mode Exit fullscreen mode

Result:

Please rephrase
Enter fullscreen mode Exit fullscreen mode

And even:

does the wake up mechines make sound?
Enter fullscreen mode Exit fullscreen mode

produced:

Please rephrase
Enter fullscreen mode Exit fullscreen mode

The word mechines is obviously a typo.

But the intended meaning is still easy for a human to recover.

At this point the architecture was technically clean but the actual Game Master felt strangely rigid.


What may have gone wrong

There was probably no single cause.

In fact, Version 3 changed several things at once, so I cannot claim that any one of them independently fixed the problem.

But four design problems became obvious.


1. The 0.8 thresholds may have been too conservative

This was my first suspicion.

Suppose Jev had actually returned something like:

wellFormed = 0.74
relevant = 0.92
contradicted = 0.89
Enter fullscreen mode Exit fullscreen mode

for:

does the wake up mechines make sound?
Enter fullscreen mode Exit fullscreen mode

That would mean the system understood quite a lot.

But my code did this first:

if (wellFormed < 0.8) {
  return "REPHRASE";
}
Enter fullscreen mode Exit fullscreen mode

So all the useful information below that gate became irrelevant.

The probability was continuous.

My application policy turned it into a cliff.

0.7999 → rejected
0.8000 → accepted
Enter fullscreen mode Exit fullscreen mode

A lower threshold might have improved the result.

I still think this is a plausible explanation.

But Version 2 had several thresholds, so lowering one would not necessarily solve the whole problem.


2. Several uncertain decisions were connected with hard gates

The deeper problem was not just that 0.8 might have been too high.

There were several such boundaries.

The system effectively became:

IF recoverable enough
AND relevant enough
AND supported enough
AND not contradicted enough
AND ...
THEN YES
Enter fullscreen mode Exit fullscreen mode

Each semantic dimension could be individually reasonable.

But the final behavior depended on all of them interacting correctly.

I had replaced one uncertain AI decision with several uncertain AI decisions and connected them using deterministic cliffs.

That made the overall product more brittle.


3. I decomposed things that humans understand together

Consider:

does the wake up mechines make sound?
Enter fullscreen mode Exit fullscreen mode

A human probably does not consciously evaluate:

grammar quality
→ relevance
→ referent resolution
→ semantic truth
Enter fullscreen mode Exit fullscreen mode

as separate numerical decisions.

We do something more like:

"mechines" probably means "mechanism"
        ↓
They mean the wake-up mechanism
        ↓
They are asking whether it makes sound
        ↓
No
Enter fullscreen mode Exit fullscreen mode

It is one semantic interpretation.

Version 2 decomposed that interpretation into multiple scores.

That was elegant from a TypeScript perspective.

It may have been unnatural from a semantic-decision perspective.


4. The judge knew the scenario, but not explicitly what mystery was being solved

This became one of the most useful changes.

Take:

is it light?
Enter fullscreen mode Exit fullscreen mode

As isolated English, it is ambiguous.

But this is not isolated English.

The player is solving a specific mystery.

For the quiet-alarm puzzle, the actual question being investigated is approximately:

What causes the sleeping guest to wake
when the alarm makes no sound
and nothing touches the guest?
Enter fullscreen mode Exit fullscreen mode

A human Game Master naturally interprets:

is it light?
Enter fullscreen mode Exit fullscreen mode

as:

Is light what causes the guest to wake?
Enter fullscreen mode Exit fullscreen mode

Version 2 had the scenario and hidden answer, but it did not have an explicit representation of the mystery focus.

That changed in Version 3.


Version 3: one semantic decision per judge

Instead of trying to improve Version 2 by adding more examples, more thresholds, or more scores, I removed complexity.

The Question Judge now makes one categorical semantic decision.

Conceptually:

type QuestionSemanticClass =
  | "SUPPORTED"
  | "CONTRADICTED"
  | "MIXED"
  | "UNKNOWN"
  | "IRRELEVANT"
  | "AMBIGUOUS";
Enter fullscreen mode Exit fullscreen mode

Jev returns one probability distribution over those choices.

Then TypeScript performs a deterministic mapping:

SUPPORTED
→ YES

CONTRADICTED
→ NO

MIXED
→ PARTLY

UNKNOWN
→ NOT ENOUGH INFORMATION

IRRELEVANT
→ NOT RELEVANT

AMBIGUOUS
→ PLEASE REPHRASE
Enter fullscreen mode Exit fullscreen mode

There is no active:

wellFormed >= 0.8
Enter fullscreen mode Exit fullscreen mode

gate anymore.

There is no chain of independent semantic Nouls.

There is one semantic classification.


The Solution Judge was simplified in the same way

Version 2 had five independent signals.

Version 3 has one classification:

type SolutionSemanticClass =
  | "CORE_CAUSAL_EXPLANATION"
  | "CORE_WITH_CONTEXT_ERROR"
  | "SUPPORTING_INSIGHT_ONLY"
  | "COMPETING_WRONG_MECHANISM"
  | "NO_MEANINGFUL_INSIGHT";
Enter fullscreen mode Exit fullscreen mode

The product mapping remains deterministic:

CORE_CAUSAL_EXPLANATION
→ SOLVED

CORE_WITH_CONTEXT_ERROR
→ PARTIAL

SUPPORTING_INSIGHT_ONLY
→ PARTIAL

COMPETING_WRONG_MECHANISM
→ NOT_SOLVED

NO_MEANINGFUL_INSIGHT
→ NOT_SOLVED
Enter fullscreen mode Exit fullscreen mode

So I did not abandon semantic distinctions.

I reduced the number of separate decisions required to produce them.


I also added mysteryFocus

Every puzzle now includes a small server-only description of what the player is trying to explain.

For example:

What causes the sleeping guest to wake
when the alarm makes no sound
and nothing touches the guest?
Enter fullscreen mode Exit fullscreen mode

This does not contain the answer.

It only gives the judge the same contextual frame that a human Game Master naturally has.

That helps with short language such as:

is it light?
does it buzz?
is it moving?
Enter fullscreen mode Exit fullscreen mode

without creating hard-coded phrase exceptions.


I shortened the shared rubric

Another change was less visible but important.

Earlier rubrics had gradually accumulated semantic instructions and special distinctions.

Version 3 intentionally moved back toward a smaller question:

Which semantic category best describes this player's statement in this puzzle?

The rubric still tells the model to tolerate:

non-native English
minor spelling mistakes
missing articles
telegraphic wording
ordinary shorthand
natural pronoun resolution
Enter fullscreen mode Exit fullscreen mode

But it does not ask for several independent measurements of those properties.

Again:

Decide less.


What happened after the simplification?

The difference was immediate.

Here are real protected-preview results.

Before: Version 2

is it light?
→ Not enough information
Enter fullscreen mode Exit fullscreen mode
is it light that wakes the guest?
→ Not enough information
Enter fullscreen mode Exit fullscreen mode
does the wake up mechanism make sound?
→ Please rephrase
Enter fullscreen mode Exit fullscreen mode
does the wake up mechines make sound?
→ Please rephrase
Enter fullscreen mode Exit fullscreen mode
is it light that wakes the guest?
[Theory]
→ Not quite yet.
Enter fullscreen mode Exit fullscreen mode

Now compare that with Version 3.


After: Version 3

is it light?
→ Yes
Enter fullscreen mode Exit fullscreen mode
is brightness involved?
→ Yes
Enter fullscreen mode Exit fullscreen mode
something bright wake him?
→ Yes
Enter fullscreen mode Exit fullscreen mode

This is especially important.

The grammar is poor:

something bright wake him?
Enter fullscreen mode Exit fullscreen mode

But the intended meaning is recoverable.

The new judge treated it that way.


Then:

does the wake up mechines make sound?
→ No
Enter fullscreen mode Exit fullscreen mode

The typo no longer caused the system to reject the whole question.

Similarly:

does it buzz?
→ No
Enter fullscreen mode Exit fullscreen mode

and:

is there vibration?
→ No
Enter fullscreen mode Exit fullscreen mode

Short questions also became usable.


Mixed statements started behaving correctly too

I tried:

something bright wake him?
does the wake up mechines make sound?
Enter fullscreen mode Exit fullscreen mode

The result was:

Partly
Enter fullscreen mode Exit fullscreen mode

with the UI:

Some of that is right, but another part isn't.
Enter fullscreen mode Exit fullscreen mode

This is exactly what PARTLY was intended to mean.

Not uncertainty.

Not closeness.

Actual mixed semantic truth.


Most importantly, concise correct theories started solving the puzzle

Question:

does light wake him?
Enter fullscreen mode Exit fullscreen mode

Result:

Yes
Enter fullscreen mode Exit fullscreen mode

Then I reused the exact same text as a theory:

does light wake him?
Enter fullscreen mode Exit fullscreen mode

Result:

You got it.
Enter fullscreen mode Exit fullscreen mode

Another phrasing:

the room gets bright and that wakes him
Enter fullscreen mode Exit fullscreen mode

Question:

Yes
Enter fullscreen mode Exit fullscreen mode

Theory:

You got it.
Enter fullscreen mode Exit fullscreen mode

That is much closer to how a human Game Master should behave.


Wrong causal mechanisms still failed

Simplification did not mean blindly accepting anything containing the correct keyword.

For example:

lamp is there but vibration wakes him
Enter fullscreen mode Exit fullscreen mode

Question:

No
Enter fullscreen mode Exit fullscreen mode

Theory:

Not quite yet.
Enter fullscreen mode Exit fullscreen mode

That matters because false SOLVED results are much worse for this game than conservative partial judgments.

So far, Version 3 became more tolerant without immediately destroying the distinction between correct and incorrect causal explanations.


The behavior also generalized to another puzzle

I tested a different puzzle involving an apparently empty frame whose contents seem to change.

Some examples:

is the frame big
→ Not relevant
Enter fullscreen mode Exit fullscreen mode
is it an animal
→ No
Enter fullscreen mode Exit fullscreen mode
is it a TV
→ No
Enter fullscreen mode Exit fullscreen mode
is it a machine
→ No
Enter fullscreen mode Exit fullscreen mode

Then:

it changes its color with sun light
→ Yes
Enter fullscreen mode Exit fullscreen mode

Submitted as a theory:

it changes its color with sun light
→ You got it.
Enter fullscreen mode Exit fullscreen mode

Again, the grammar is not perfect.

But the causal idea is correct.

And the judge accepted it.


Version 3 is not perfect

There is still some instability around boundary cases.

For example, during one session:

is it the sunrise and sunset
→ Yes
Enter fullscreen mode Exit fullscreen mode

while similar wording later produced:

is it sunrise and sunset
→ Not enough information
Enter fullscreen mode Exit fullscreen mode

That tells me the system is not magically deterministic at the semantic level.

Nor should I expect it to be.

The important difference is that the major player-facing failures changed from:

"I clearly understand what you're asking,
but please rephrase."
Enter fullscreen mode Exit fullscreen mode

to occasional disagreement around genuinely fuzzy boundaries.

For a proof-of-concept game, that is a much better failure mode.


So what actually fixed it?

The honest answer is:

I do not know which individual change fixed it.

Version 3 changed several things together:

Removed multiple continuous semantic gates

Removed the 0.8/0.2 decision chain

Replaced many Nouls with one Choice distribution

Added mysteryFocus

Shortened and generalized the rubric

Made language recovery part of the single semantic classification
Enter fullscreen mode Exit fullscreen mode

Any one of those may have contributed.

The improvement is evidence that the overall architecture became better.

It is not a controlled experiment proving that 0.8 alone was the problem.

A useful future experiment would compare Version 2 under several threshold policies.

That could answer:

Was Jev already producing useful semantic probabilities that my application was throwing away?

I think that is entirely possible.


Jev started feeling less like "adding AI" and more like building probabilistic IF statements

This project changed how I think about decision models.

I am not really building a chatbot.

I am building something conceptually closer to:

if (semanticCondition) {
  doSomething();
}
Enter fullscreen mode Exit fullscreen mode

Except the condition is not:

x > 10
Enter fullscreen mode Exit fullscreen mode

It is:

Does this player's sentence semantically express
a proposition supported by the hidden explanation?
Enter fullscreen mode Exit fullscreen mode

Jev estimates that semantic condition.

TypeScript decides what happens next.

That is why I increasingly think of this pattern as:

probabilistic IF statements over natural language

or:

an AI-powered switch/case

Version 2 accidentally turned that simple idea into a giant conditional expression.

Version 3 moved it back toward a small semantic switch.


There is also a hidden engineering cost: calibration

Another lesson was that API pricing is not the only cost of using a decision model.

For this project I needed to repeatedly:

design semantic categories
test real player wording
test spelling mistakes
test short questions
test false mechanisms
test incomplete solutions
check false positives
check false negatives
change policies
retest on protected Preview
Enter fullscreen mode Exit fullscreen mode

That is human labor.

Jev reduced one kind of complexity for me.

I did not have to parse arbitrary assistant prose into application state.

But that complexity did not disappear.

Some of it moved into:

decision design
calibration
threshold selection
test-play
behavioral evaluation
Enter fullscreen mode Exit fullscreen mode

I would therefore not say:

Jev requires more engineering work than a traditional LLM.

I have not run a controlled comparison that proves that.

A more accurate statement is:

In this project, Jev moved a meaningful part of the engineering effort from output handling to decision design and calibration.

That is an important cost to account for.


849 tests still did not replace test-play

Version 3 currently passes:

849 unit/API tests
82 desktop/mobile E2E tests
Leak scanning
Lint
Type checking
Build
Enter fullscreen mode Exit fullscreen mode

Those tests are valuable.

They prove things such as:

the parser rejects invalid distributions
the deterministic mapping works
private grading data does not reach the browser
provider call counts remain bounded
the UI still behaves correctly
Enter fullscreen mode Exit fullscreen mode

But they still cannot fully prove:

A human writes an unseen sentence
        ↓
The model understands it as intended
Enter fullscreen mode Exit fullscreen mode

That requires actual behavioral testing.

For AI systems I now think about validation as two separate layers:

Layer 1
Software correctness

Layer 2
Model behavior
Enter fullscreen mode Exit fullscreen mode

A green CI pipeline is excellent evidence for Layer 1.

It is not sufficient evidence for Layer 2.


Why I am stopping here

There are still things I could tune.

I could add more categories.

I could add confidence gates.

I could add more examples.

I could create exceptions for phrases like:

sunrise and sunset
Enter fullscreen mode Exit fullscreen mode

I am deliberately not doing that.

That path is exactly how Version 2 became complicated.

For this PoC, Version 3 is now a strong freeze candidate.

The remaining requirement is not:

perfect semantic classification
Enter fullscreen mode Exit fullscreen mode

It is:

Natural short player language usually works.

Minor grammar and spelling errors usually work.

Correct causal explanations can solve the puzzle.

Wrong causal explanations do not solve it.

Hidden information stays hidden.

The game remains fun.
Enter fullscreen mode Exit fullscreen mode

That is enough.


What if this still breaks later?

I already have a fallback.

Simplify even further:

YES
OTHER
Enter fullscreen mode Exit fullscreen mode

OTHER would intentionally collapse:

NO
UNKNOWN
IRRELEVANT
AMBIGUOUS
possibly MIXED
Enter fullscreen mode Exit fullscreen mode

That would lose information.

But if it produced a better game experience, I would seriously consider it.

Again:

The goal is not to extract the maximum amount of semantic information from the model.

The goal is to ask for the smallest useful decision.


Final takeaway

The most surprising lesson from building Sideways has been this:

More semantic structure does not automatically produce better AI behavior.

Version 2 looked more sophisticated.

It had more probabilities.

More semantic axes.

More explicit thresholds.

More detailed grading logic.

And worse real gameplay.

Version 3 asked Jev to do less:

One semantic Choice
        ↓
One probability distribution
        ↓
Deterministic TypeScript switch
Enter fullscreen mode Exit fullscreen mode

And the Game Master got noticeably better.

So the principle I am taking forward is:

Do not ask AI to make a more complicated decision than your product actually needs.

Or even shorter:

Understand enough.

Decide less.
Enter fullscreen mode Exit fullscreen mode

That turned out to be a much better architecture for this game.

Top comments (0)