DEV Community

494
494

Posted on

When More Semantic Signals Made My AI Judge Worse: Building a Lateral Thinking Game with Jev

When More Semantic Signals Made My AI Judge Worse

I have been building a small web game called Sideways.

It is a lateral-thinking puzzle game: the player sees a mysterious situation and tries to discover what really happened by asking questions.

A typical interaction should feel almost trivial:

Player:
Is it light that wakes the guest?

Game Master:
Yes.
Enter fullscreen mode Exit fullscreen mode

Or:

Player:
Does the wake-up mechanism make sound?

Game Master:
No.
Enter fullscreen mode Exit fullscreen mode

The interesting part is that the Game Master is powered by AI.

I chose Jev because I did not want a chatbot generating long answers. I wanted a decision layer that could take application state, make a typed judgment, and let ordinary TypeScript decide what the product should do.

Jev describes itself as a decision model rather than a chat model. Its API works with application state and typed questions such as Choice, Score, and Noul, returning structured results and probabilities rather than conversational prose.

That sounded almost perfect for this game.

But I discovered something unexpected:

The more semantic intelligence I tried to extract from the model, the worse the actual game experience became.

This article is about that failure.

And why I am now making the AI do less.


The Architecture I Wanted

From the beginning, I did not want the model to control the application directly.

The architecture was:

Player text
    ↓
Probabilistic semantic judgment
    ↓
Typed output
    ↓
Deterministic TypeScript policy
    ↓
Player-facing result
Enter fullscreen mode Exit fullscreen mode

The principle was simple:

AI = semantic judgment
Code = deterministic policy
User = consequential intent
Enter fullscreen mode Exit fullscreen mode

The model should understand language.

My code should decide what happens next.

This has several advantages.

The browser never needs to receive the hidden solution.

The model does not get to decide arbitrary UI behavior.

The public API can expose only simple results such as:

YES
NO
PARTLY
IRRELEVANT
UNCLEAR
Enter fullscreen mode Exit fullscreen mode

while keeping the private reasoning signals on the server.

For solution attempts, the public results are similarly constrained:

SOLVED
PARTIAL
NOT_SOLVED
Enter fullscreen mode Exit fullscreen mode

Conceptually, I still believe this architecture is correct.

The problem was how much semantic structure I asked the model to produce.


Version 1: Simple Evidence

The first Question Judge was relatively simple.

For a player's question, the model estimated something like:

{
  wellFormed: number,
  relevant: number,
  evidence: {
    entailed: number,
    contradicted: number,
    undetermined: number
  }
}
Enter fullscreen mode Exit fullscreen mode

TypeScript then converted these probabilities into a product result.

For example:

if (wellFormed < 0.8) return "UNCLEAR";
if (relevant <= 0.2) return "IRRELEVANT";
if (entailed >= 0.8) return "YES";
if (contradicted >= 0.8) return "NO";

return "UNCLEAR";
Enter fullscreen mode Exit fullscreen mode

This had an important property:

A low probability of YES did not automatically mean NO.

An unknown fact could remain unknown.

That part worked well.

But actual gameplay revealed another problem.

Natural player language is messy.

People write things like:

light wake him?
Enter fullscreen mode Exit fullscreen mode
does it make sound
Enter fullscreen mode Exit fullscreen mode
maybe lamp?
Enter fullscreen mode Exit fullscreen mode

They make spelling mistakes.

They omit articles.

They use pronouns.

They do not write propositions like lawyers.

So I tried to make the judge more sophisticated.


Version 2: More Semantic Intelligence

In Version 2, the Question Judge expanded into:

{
  wellFormed,
  relevant,
  evidence: {
    entailed,
    contradicted,
    undetermined
  },
  mixedClaims: {
    supported,
    contradicted
  }
}
Enter fullscreen mode Exit fullscreen mode

This allowed a new result:

PARTLY
Enter fullscreen mode Exit fullscreen mode

For example:

There is no alarm. Only a lamp wakes the guest.
Enter fullscreen mode Exit fullscreen mode

contains both a false claim and an important true claim.

I did not want that to be treated as simply NO.

For solution checking, I went even further.

Instead of judging a theory with a simple checklist, the model returned five semantic dimensions:

{
  coreMechanism,
  causalRelation,
  supportingInsight,
  contextualError,
  competingMechanism
}
Enter fullscreen mode Exit fullscreen mode

The idea was reasonable.

Consider these three theories:

There is a light.
Enter fullscreen mode Exit fullscreen mode
The light wakes the guest.
Enter fullscreen mode Exit fullscreen mode
There is a light, but vibration wakes the guest.
Enter fullscreen mode Exit fullscreen mode

They should not receive the same result.

I wanted:

"There is a light."
→ PARTIAL

"The light wakes the guest."
→ SOLVED

"Light exists, but vibration wakes the guest."
→ NOT_SOLVED
Enter fullscreen mode Exit fullscreen mode

So I separated:

  • identifying the important object;
  • understanding the causal mechanism;
  • making a harmless contextual mistake;
  • asserting a competing wrong mechanism.

Architecturally, this looked much better.


And the Offline Tests Looked Excellent

The implementation passed:

880 unit/API tests
82 desktop/mobile E2E tests
Leak scans
Lint
Type checking
Build
Full repository checks
Enter fullscreen mode Exit fullscreen mode

I also created separate calibration and engineering-generalization corpora across all six puzzles.

The browser did not receive private grading data.

The hidden solution remained server-only.

The number of provider calls remained bounded.

Everything looked clean.

But there was one very important limitation.

Most of those tests proved:

semantic signal
→ parser
→ deterministic policy
→ API result
→ UI
Enter fullscreen mode Exit fullscreen mode

They did not prove that the real model would infer the expected semantic signal from previously unseen player language.

That distinction turned out to matter a lot.


Then I Played the Game with the Real Model

I deployed the new judge to a protected preview environment and started typing ordinary questions.

The results were surprising.

Example 1

Player:
is it light?

Result:
Not enough information
Enter fullscreen mode Exit fullscreen mode

Example 2

Player:
is it light that wakes the guest?

Result:
Not enough information
Enter fullscreen mode Exit fullscreen mode

Then I submitted the exact same sentence as a theory:

Theory:
is it light that wakes the guest?

Result:
Not quite yet.
Enter fullscreen mode Exit fullscreen mode

A human Game Master would almost certainly understand the intended idea.

The hidden solution is that a visual light signal wakes the sleeping guest.


Example 3

Player:
does the wake up mechanism make sound?

Result:
Please rephrase
Enter fullscreen mode Exit fullscreen mode

Even this typo-heavy version:

does the wake up mechines make sound?
Enter fullscreen mode Exit fullscreen mode

also produced:

Please rephrase
Enter fullscreen mode Exit fullscreen mode

That was a serious UX problem.

The grammar is imperfect, but the semantic intent is obvious.


Example 4

Player:
Instead of an alarm, lights from a device wakes the guest.

Question result:
Not enough information

Theory result:
Not quite yet.
Enter fullscreen mode Exit fullscreen mode

Again, the wording is imperfect.

But the important causal insight is clearly present.


Why Did More Structure Make Things Worse?

I cannot infer Jev's internal reasoning from its output alone.

So the following are hypotheses about my decision architecture, not claims about the internals of the model.

Several problems stood out.


1. The Semantic Dimensions Were Not Really Independent

I had separated concepts like:

well-formedness
relevance
truth
mixed claims
causal understanding
Enter fullscreen mode Exit fullscreen mode

because it made the TypeScript architecture easier to reason about.

But humans do not necessarily understand a sentence in that order.

Take:

does the wake up mechines make sound?
Enter fullscreen mode Exit fullscreen mode

A human probably performs something closer to:

"mechines" is probably "mechanism"
        ↓
The player means the wake-up mechanism
        ↓
They are asking whether it makes sound
        ↓
I know the answer
        ↓
No
Enter fullscreen mode Exit fullscreen mode

This is one integrated semantic interpretation.

I had converted it into several semi-independent judgments.

That created more opportunities for one signal to disagree with the others.


2. Probabilities Became Hard Cliffs

The intention behind probabilities was to avoid brittle Boolean decisions.

But then my code contained rules such as:

if (wellFormed < 0.8) {
  return "REPHRASE";
}
Enter fullscreen mode Exit fullscreen mode

Imagine the model essentially understands the sentence but produces:

wellFormed = 0.77
Enter fullscreen mode Exit fullscreen mode

The rest of the semantic evidence no longer matters.

The application suddenly says:

Please rephrase.
Enter fullscreen mode Exit fullscreen mode

A continuous probability had become a discrete cliff.

Adding more semantic dimensions meant adding more places where such cliffs could occur.

The model did not necessarily need to be dramatically wrong.

My policy only needed one signal to fall on the wrong side of a threshold.


3. I Was Asking a Small Decision Model to Solve a Large Semantic Task

Jev's documentation emphasizes focused, well-scoped typed decisions. Choice is designed for predefined classification, while Noul represents a yes/no judgment whose probability can be interpreted by application code.

But my Question Judge was effectively asking the model to do all of this at once:

Recover imperfect language
Resolve pronouns
Determine the semantic proposition
Determine relevance
Compare with hidden truth
Detect multiple claims
Separate supported and contradicted claims
Distinguish uncertainty from contradiction
Enter fullscreen mode Exit fullscreen mode

The individual fields looked small.

The total semantic task was not.

I had decomposed the output without necessarily decomposing the difficulty.


4. The Judge Knew the Scenario, but Not the "Mystery Focus"

This turned out to be especially interesting.

Consider:

is it light?
Enter fullscreen mode Exit fullscreen mode

As an isolated English sentence, it is ambiguous.

What is "it"?

The alarm?

The mechanism?

The object?

But in a lateral-thinking game, a human Game Master has another piece of context:

What mystery is the player currently trying to explain?

For the quiet-alarm puzzle, that focus is roughly:

What causes the sleeping guest to wake
without sound or physical contact?
Enter fullscreen mode Exit fullscreen mode

Given that context:

is it light?
Enter fullscreen mode Exit fullscreen mode

has a very natural interpretation:

Is light the cause?
Enter fullscreen mode Exit fullscreen mode

I had provided the scenario and hidden reference, but not an explicit representation of the question being investigated.

That may have encouraged overly literal ambiguity handling.


5. My Solution Judge May Have Confused Completeness with Correctness

The Version 2 Solution Judge required strong scores for both:

coreMechanism
causalRelation
Enter fullscreen mode Exit fullscreen mode

before returning SOLVED.

That sounds reasonable.

But consider:

The light wakes the guest.
Enter fullscreen mode Exit fullscreen mode

This is extremely short.

It also contains both the mechanism and causal relationship.

If the model gives:

coreMechanism = high
causalRelation = medium
Enter fullscreen mode Exit fullscreen mode

the deterministic policy returns PARTIAL.

From the application's perspective, that is a false negative.

The player has solved the mystery.

The grading representation made the answer look less complete than it actually was.


Version 3: Make the Model Do Less

So I am now testing a different architecture.

Instead of asking for many independent semantic scores, each judge will return one categorical probability distribution.

For Question:

type QuestionSemanticClass =
  | "SUPPORTED"
  | "CONTRADICTED"
  | "MIXED"
  | "UNKNOWN"
  | "IRRELEVANT"
  | "AMBIGUOUS";
Enter fullscreen mode Exit fullscreen mode

Then ordinary TypeScript maps it:

SUPPORTED
→ YES

CONTRADICTED
→ NO

MIXED
→ PARTLY

UNKNOWN
→ NOT ENOUGH INFORMATION

IRRELEVANT
→ NOT RELEVANT

AMBIGUOUS
→ PLEASE REPHRASE
Enter fullscreen mode Exit fullscreen mode

There is still probabilistic information.

But there is only one semantic decision.

No chain of:

0.8
0.8
0.2
0.8
...
Enter fullscreen mode Exit fullscreen mode

gates.


The Solution Judge Is Being Simplified Too

Instead of five independent dimensions, I am testing one semantic classification:

type SolutionSemanticClass =
  | "CORE_CAUSAL_EXPLANATION"
  | "CORE_WITH_CONTEXT_ERROR"
  | "SUPPORTING_INSIGHT_ONLY"
  | "COMPETING_WRONG_MECHANISM"
  | "NO_MEANINGFUL_INSIGHT";
Enter fullscreen mode Exit fullscreen mode

TypeScript then maps:

CORE_CAUSAL_EXPLANATION
→ SOLVED

CORE_WITH_CONTEXT_ERROR
→ PARTIAL

SUPPORTING_INSIGHT_ONLY
→ PARTIAL

COMPETING_WRONG_MECHANISM
→ NOT_SOLVED

NO_MEANINGFUL_INSIGHT
→ NOT_SOLVED
Enter fullscreen mode Exit fullscreen mode

The important distinction remains.

There is a light.
Enter fullscreen mode Exit fullscreen mode

should not equal:

The light wakes the guest.
Enter fullscreen mode Exit fullscreen mode

And neither should equal:

The vibration wakes the guest.
Enter fullscreen mode Exit fullscreen mode

But the model no longer needs to independently estimate five continuous semantic dimensions before my application can make that distinction.


I Am Also Adding a "Mystery Focus"

Each puzzle will have a small server-only field such as:

What causes the sleeping guest to wake
without sound or physical contact?
Enter fullscreen mode Exit fullscreen mode

This is not the hidden answer.

It simply represents the question raised by the public scenario.

The hope is that it helps resolve natural shorthand:

is it light?
Enter fullscreen mode Exit fullscreen mode
does it make sound?
Enter fullscreen mode Exit fullscreen mode
was he covered?
Enter fullscreen mode Exit fullscreen mode

without adding phrase-specific exceptions.

That distinction matters.

I do not want:

if text === "is it light?"
Enter fullscreen mode Exit fullscreen mode

I want:

short player language
+
current mystery
→ recoverable semantic proposition
Enter fullscreen mode Exit fullscreen mode

And If Version 3 Still Fails?

Then I plan to simplify again.

Possibly all the way to:

YES
OTHER
Enter fullscreen mode Exit fullscreen mode

For Question judging:

YES
Enter fullscreen mode Exit fullscreen mode

would mean:

This player proposition is supported by the hidden explanation.

Everything else would collapse into:

OTHER
Enter fullscreen mode Exit fullscreen mode

That would intentionally stop distinguishing among:

NO
UNKNOWN
IRRELEVANT
AMBIGUOUS
Enter fullscreen mode Exit fullscreen mode

This loses information.

It may still produce a better game.

That is the part of this experiment I find most interesting.


The Product Does Not Need the Maximum Amount of AI Information

It is tempting to think:

More semantic signals
→ more information
→ better product decisions
Enter fullscreen mode Exit fullscreen mode

My experience so far suggests something more like:

More semantic signals
→ more uncertain boundaries
→ more policy interactions
→ potentially worse UX
Enter fullscreen mode Exit fullscreen mode

The best AI interface may not be the one that extracts the most information from a model.

It may be the one that asks the model for the smallest decision the product actually needs.


Another Lesson: Offline Tests and Semantic Accuracy Are Different Things

I had hundreds of passing tests.

They were useful.

They caught:

  • schema mistakes;
  • API regressions;
  • accidental hidden-data leaks;
  • extra provider calls;
  • incorrect deterministic mappings;
  • broken UI behavior.

I absolutely want those tests.

But:

synthetic semantic signals
→ correct product output
Enter fullscreen mode Exit fullscreen mode

does not prove:

previously unseen human language
→ correct semantic signal
Enter fullscreen mode Exit fullscreen mode

Those are different systems.

For AI products, both need testing.

I now think of them as two separate validation layers:

Layer 1:
Software correctness

Layer 2:
Model behavior under real language
Enter fullscreen mode Exit fullscreen mode

A green CI pipeline proves the first.

It does not automatically prove the second.


Why I Still Like the Typed-Decision Architecture

None of this has made me want to replace the system with a free-form chatbot.

Actually, the opposite.

I still want:

AI
→ bounded semantic decision

TypeScript
→ application policy
Enter fullscreen mode Exit fullscreen mode

Jev's API is explicitly designed around application state and typed decisions, with probability distributions that software can consume.

The lesson for me is not:

Typed decisions are too simple.

It is:

I should respect their simplicity.

If the application needs one classification, asking for one well-designed Choice may be better than constructing a small semantic ontology and connecting it with threshold gates.


Where the Experiment Stands

At the time of writing:

v1
Simple probabilistic evidence
        ↓

v2
More semantic dimensions
        ↓

Real-model testing exposed brittle behavior
        ↓

v3
One categorical semantic distribution per judge
+ explicit mystery focus
        ↓

Currently being tested
Enter fullscreen mode Exit fullscreen mode

If v3 performs well on previously unseen wording, I will freeze the Judge and move on to production hardening.

If it does not, I will test the YES / OTHER design rather than adding more semantic complexity.

That constraint is intentional.

At some point, engineering discipline means stopping optimization.


Final Takeaway

The most useful lesson from this project so far is surprisingly simple:

Do not ask the AI to make a more complicated decision than your product actually needs.

A sophisticated semantic representation can be elegant in TypeScript.

It can have clean types.

It can have excellent unit tests.

It can even look more theoretically correct.

And still produce a worse experience for the person actually playing the game.

Sometimes the better AI architecture is not:

Understand more.
Enter fullscreen mode Exit fullscreen mode

It is:

Decide less.
Enter fullscreen mode Exit fullscreen mode

I will update this article once the simplified Version 3 Judge has been tested against the real model.


About the project

Sideways is an experimental lateral-thinking puzzle game built with Next.js, TypeScript, and Jev.

The project explores a specific architecture:

Probabilistic Semantic Judgment
→ Typed Output
→ Deterministic TypeScript Policy
→ Product Result
Enter fullscreen mode Exit fullscreen mode

The next question is how small that semantic judgment can become while still producing a good game.

Top comments (0)