<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: 494</title>
    <description>The latest articles on DEV Community by 494 (@kitadaro).</description>
    <link>https://dev.to/kitadaro</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4161578%2F3606890d-2529-4f0c-aac9-06eb9b6604ee.jpg</url>
      <title>DEV Community: 494</title>
      <link>https://dev.to/kitadaro</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kitadaro"/>
    <language>en</language>
    <item>
      <title>When Simpler AI Decisions Worked Better: What I Changed in My Jev-Powered Game Judge</title>
      <dc:creator>494</dc:creator>
      <pubDate>Sun, 04 Oct 2026 16:43:52 +0000</pubDate>
      <link>https://dev.to/kitadaro/when-simpler-ai-decisions-worked-better-what-i-changed-in-my-jev-powered-game-judge-a8n</link>
      <guid>https://dev.to/kitadaro/when-simpler-ai-decisions-worked-better-what-i-changed-in-my-jev-powered-game-judge-a8n</guid>
      <description>&lt;p&gt;I have been building a small lateral-thinking puzzle game called &lt;strong&gt;Sideways&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The basic interaction is simple.&lt;/p&gt;

&lt;p&gt;The player sees a mysterious situation and asks questions until they discover what really happened.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Player:
Is it light?

Game Master:
Yes.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Player:
Does the wake-up mechanism make sound?

Game Master:
No.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Game Master is powered by Jev.&lt;/p&gt;

&lt;p&gt;I deliberately did not want a general-purpose chatbot generating explanations.&lt;/p&gt;

&lt;p&gt;I wanted something much closer to this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Natural-language input
        ↓
Probabilistic semantic decision
        ↓
Typed result
        ↓
Deterministic TypeScript
        ↓
Game behavior
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After several iterations, I realized that this is almost like building &lt;strong&gt;probabilistic IF statements over natural language&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;And I also learned something I did not expect:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Making those semantic IF statements more sophisticated actually made the game worse.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The fix was not to make the AI smarter.&lt;/p&gt;

&lt;p&gt;The fix was to make the decision smaller.&lt;/p&gt;




&lt;h1&gt;
  
  
  The original idea: AI understands meaning, code controls behavior
&lt;/h1&gt;

&lt;p&gt;The architecture follows a simple rule:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI = semantic judgment
Code = deterministic policy
User = consequential intent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model does not directly decide which UI to show.&lt;/p&gt;

&lt;p&gt;The browser does not receive the hidden puzzle solution.&lt;/p&gt;

&lt;p&gt;The model produces a bounded semantic judgment, and TypeScript turns that judgment into a product result.&lt;/p&gt;

&lt;p&gt;For questions, the public result can be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;YES
NO
PARTLY
IRRELEVANT
UNCLEAR
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For theory submissions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SOLVED
PARTIAL
NOT_SOLVED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This separation still feels right to me.&lt;/p&gt;

&lt;p&gt;The mistake was not the architecture itself.&lt;/p&gt;

&lt;p&gt;The mistake was how much semantic structure I tried to extract from one player sentence.&lt;/p&gt;




&lt;h1&gt;
  
  
  Version 2: the judge became too clever
&lt;/h1&gt;

&lt;p&gt;In Version 2, a Question judgment looked roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;wellFormed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;relevant&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;entailed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;contradicted&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;undetermined&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="nx"&gt;mixedClaims&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;supported&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;contradicted&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Most of those values were probabilities between 0 and 1.&lt;/p&gt;

&lt;p&gt;Then TypeScript applied rules such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;wellFormed&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;REPHRASE&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;relevant&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;IRRELEVANT&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;relevant&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;NOT_ENOUGH_INFORMATION&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;mixedClaims&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;supported&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt;
  &lt;span class="nx"&gt;mixedClaims&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;contradicted&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;PARTLY&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;entailed&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;YES&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;contradicted&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;NO&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Solution Judge was even more detailed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;coreMechanism&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;causalRelation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;supportingInsight&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;contextualError&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;competingMechanism&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Again, deterministic TypeScript combined these scores.&lt;/p&gt;

&lt;p&gt;The goal was reasonable.&lt;/p&gt;

&lt;p&gt;I wanted to distinguish:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"There is a light."
→ PARTIAL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"The light wakes the guest."
→ SOLVED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"There is a light, but vibration wakes the guest."
→ NOT_SOLVED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Conceptually, Version 2 was much more expressive.&lt;/p&gt;

&lt;p&gt;The code also looked disciplined.&lt;/p&gt;




&lt;h1&gt;
  
  
  And the software tests looked excellent
&lt;/h1&gt;

&lt;p&gt;Version 2 passed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;880 unit/API tests
82 desktop/mobile E2E tests
Leak scanning
Lint
Type checking
Production build
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I had calibration data.&lt;/p&gt;

&lt;p&gt;I had a separate engineering-generalization corpus.&lt;/p&gt;

&lt;p&gt;Hidden grading metadata stayed server-side.&lt;/p&gt;

&lt;p&gt;Provider calls remained bounded.&lt;/p&gt;

&lt;p&gt;Everything looked good.&lt;/p&gt;

&lt;p&gt;Then I used the real model.&lt;/p&gt;




&lt;h1&gt;
  
  
  The real game felt worse
&lt;/h1&gt;

&lt;p&gt;One puzzle is about a guest waking without sound or physical contact.&lt;/p&gt;

&lt;p&gt;The hidden mechanism is a visual light signal.&lt;/p&gt;

&lt;p&gt;I tried this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it light?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Version 2:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Not enough information
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it light that wakes the guest?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Version 2:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Not enough information
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I submitted the exact idea as a theory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it light that wakes the guest?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Version 2:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Not quite yet.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That was already worrying.&lt;/p&gt;

&lt;p&gt;Then I tried:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does the wake up mechanism make sound?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Result:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Please rephrase
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And even:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does the wake up mechines make sound?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;produced:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Please rephrase
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The word &lt;code&gt;mechines&lt;/code&gt; is obviously a typo.&lt;/p&gt;

&lt;p&gt;But the intended meaning is still easy for a human to recover.&lt;/p&gt;

&lt;p&gt;At this point the architecture was technically clean but the actual Game Master felt strangely rigid.&lt;/p&gt;




&lt;h1&gt;
  
  
  What may have gone wrong
&lt;/h1&gt;

&lt;p&gt;There was probably no single cause.&lt;/p&gt;

&lt;p&gt;In fact, Version 3 changed several things at once, so I cannot claim that any one of them independently fixed the problem.&lt;/p&gt;

&lt;p&gt;But four design problems became obvious.&lt;/p&gt;




&lt;h1&gt;
  
  
  1. The 0.8 thresholds may have been too conservative
&lt;/h1&gt;

&lt;p&gt;This was my first suspicion.&lt;/p&gt;

&lt;p&gt;Suppose Jev had actually returned something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;wellFormed = 0.74
relevant = 0.92
contradicted = 0.89
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does the wake up mechines make sound?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That would mean the system understood quite a lot.&lt;/p&gt;

&lt;p&gt;But my code did this first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;wellFormed&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;REPHRASE&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So all the useful information below that gate became irrelevant.&lt;/p&gt;

&lt;p&gt;The probability was continuous.&lt;/p&gt;

&lt;p&gt;My application policy turned it into a cliff.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0.7999 → rejected
0.8000 → accepted
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A lower threshold might have improved the result.&lt;/p&gt;

&lt;p&gt;I still think this is a plausible explanation.&lt;/p&gt;

&lt;p&gt;But Version 2 had several thresholds, so lowering one would not necessarily solve the whole problem.&lt;/p&gt;




&lt;h1&gt;
  
  
  2. Several uncertain decisions were connected with hard gates
&lt;/h1&gt;

&lt;p&gt;The deeper problem was not just that &lt;code&gt;0.8&lt;/code&gt; might have been too high.&lt;/p&gt;

&lt;p&gt;There were several such boundaries.&lt;/p&gt;

&lt;p&gt;The system effectively became:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;IF recoverable enough
AND relevant enough
AND supported enough
AND not contradicted enough
AND ...
THEN YES
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each semantic dimension could be individually reasonable.&lt;/p&gt;

&lt;p&gt;But the final behavior depended on all of them interacting correctly.&lt;/p&gt;

&lt;p&gt;I had replaced one uncertain AI decision with several uncertain AI decisions and connected them using deterministic cliffs.&lt;/p&gt;

&lt;p&gt;That made the overall product more brittle.&lt;/p&gt;




&lt;h1&gt;
  
  
  3. I decomposed things that humans understand together
&lt;/h1&gt;

&lt;p&gt;Consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does the wake up mechines make sound?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A human probably does not consciously evaluate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;grammar quality
→ relevance
→ referent resolution
→ semantic truth
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;as separate numerical decisions.&lt;/p&gt;

&lt;p&gt;We do something more like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"mechines" probably means "mechanism"
        ↓
They mean the wake-up mechanism
        ↓
They are asking whether it makes sound
        ↓
No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is one semantic interpretation.&lt;/p&gt;

&lt;p&gt;Version 2 decomposed that interpretation into multiple scores.&lt;/p&gt;

&lt;p&gt;That was elegant from a TypeScript perspective.&lt;/p&gt;

&lt;p&gt;It may have been unnatural from a semantic-decision perspective.&lt;/p&gt;




&lt;h1&gt;
  
  
  4. The judge knew the scenario, but not explicitly what mystery was being solved
&lt;/h1&gt;

&lt;p&gt;This became one of the most useful changes.&lt;/p&gt;

&lt;p&gt;Take:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it light?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;As isolated English, &lt;code&gt;it&lt;/code&gt; is ambiguous.&lt;/p&gt;

&lt;p&gt;But this is not isolated English.&lt;/p&gt;

&lt;p&gt;The player is solving a specific mystery.&lt;/p&gt;

&lt;p&gt;For the quiet-alarm puzzle, the actual question being investigated is approximately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What causes the sleeping guest to wake
when the alarm makes no sound
and nothing touches the guest?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A human Game Master naturally interprets:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it light?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Is light what causes the guest to wake?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Version 2 had the scenario and hidden answer, but it did not have an explicit representation of the &lt;strong&gt;mystery focus&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That changed in Version 3.&lt;/p&gt;




&lt;h1&gt;
  
  
  Version 3: one semantic decision per judge
&lt;/h1&gt;

&lt;p&gt;Instead of trying to improve Version 2 by adding more examples, more thresholds, or more scores, I removed complexity.&lt;/p&gt;

&lt;p&gt;The Question Judge now makes one categorical semantic decision.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;QuestionSemanticClass&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;SUPPORTED&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;CONTRADICTED&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;MIXED&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;UNKNOWN&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;IRRELEVANT&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;AMBIGUOUS&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Jev returns one probability distribution over those choices.&lt;/p&gt;

&lt;p&gt;Then TypeScript performs a deterministic mapping:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SUPPORTED
→ YES

CONTRADICTED
→ NO

MIXED
→ PARTLY

UNKNOWN
→ NOT ENOUGH INFORMATION

IRRELEVANT
→ NOT RELEVANT

AMBIGUOUS
→ PLEASE REPHRASE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no active:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;wellFormed &amp;gt;= 0.8
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;gate anymore.&lt;/p&gt;

&lt;p&gt;There is no chain of independent semantic Nouls.&lt;/p&gt;

&lt;p&gt;There is one semantic classification.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Solution Judge was simplified in the same way
&lt;/h1&gt;

&lt;p&gt;Version 2 had five independent signals.&lt;/p&gt;

&lt;p&gt;Version 3 has one classification:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;SolutionSemanticClass&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;CORE_CAUSAL_EXPLANATION&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;CORE_WITH_CONTEXT_ERROR&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;SUPPORTING_INSIGHT_ONLY&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;COMPETING_WRONG_MECHANISM&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;NO_MEANINGFUL_INSIGHT&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The product mapping remains deterministic:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CORE_CAUSAL_EXPLANATION
→ SOLVED

CORE_WITH_CONTEXT_ERROR
→ PARTIAL

SUPPORTING_INSIGHT_ONLY
→ PARTIAL

COMPETING_WRONG_MECHANISM
→ NOT_SOLVED

NO_MEANINGFUL_INSIGHT
→ NOT_SOLVED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So I did not abandon semantic distinctions.&lt;/p&gt;

&lt;p&gt;I reduced the number of separate decisions required to produce them.&lt;/p&gt;




&lt;h1&gt;
  
  
  I also added &lt;code&gt;mysteryFocus&lt;/code&gt;
&lt;/h1&gt;

&lt;p&gt;Every puzzle now includes a small server-only description of what the player is trying to explain.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What causes the sleeping guest to wake
when the alarm makes no sound
and nothing touches the guest?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This does not contain the answer.&lt;/p&gt;

&lt;p&gt;It only gives the judge the same contextual frame that a human Game Master naturally has.&lt;/p&gt;

&lt;p&gt;That helps with short language such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it light?
does it buzz?
is it moving?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;without creating hard-coded phrase exceptions.&lt;/p&gt;




&lt;h1&gt;
  
  
  I shortened the shared rubric
&lt;/h1&gt;

&lt;p&gt;Another change was less visible but important.&lt;/p&gt;

&lt;p&gt;Earlier rubrics had gradually accumulated semantic instructions and special distinctions.&lt;/p&gt;

&lt;p&gt;Version 3 intentionally moved back toward a smaller question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Which semantic category best describes this player's statement in this puzzle?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The rubric still tells the model to tolerate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;non-native English
minor spelling mistakes
missing articles
telegraphic wording
ordinary shorthand
natural pronoun resolution
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But it does not ask for several independent measurements of those properties.&lt;/p&gt;

&lt;p&gt;Again:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Decide less.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  What happened after the simplification?
&lt;/h1&gt;

&lt;p&gt;The difference was immediate.&lt;/p&gt;

&lt;p&gt;Here are real protected-preview results.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before: Version 2
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it light?
→ Not enough information
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it light that wakes the guest?
→ Not enough information
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does the wake up mechanism make sound?
→ Please rephrase
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does the wake up mechines make sound?
→ Please rephrase
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it light that wakes the guest?
[Theory]
→ Not quite yet.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now compare that with Version 3.&lt;/p&gt;




&lt;h1&gt;
  
  
  After: Version 3
&lt;/h1&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it light?
→ Yes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is brightness involved?
→ Yes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;something bright wake him?
→ Yes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is especially important.&lt;/p&gt;

&lt;p&gt;The grammar is poor:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;something bright wake him?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But the intended meaning is recoverable.&lt;/p&gt;

&lt;p&gt;The new judge treated it that way.&lt;/p&gt;




&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does the wake up mechines make sound?
→ No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The typo no longer caused the system to reject the whole question.&lt;/p&gt;

&lt;p&gt;Similarly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does it buzz?
→ No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is there vibration?
→ No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Short questions also became usable.&lt;/p&gt;




&lt;h1&gt;
  
  
  Mixed statements started behaving correctly too
&lt;/h1&gt;

&lt;p&gt;I tried:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;something bright wake him?
does the wake up mechines make sound?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The result was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Partly
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;with the UI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Some of that is right, but another part isn't.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is exactly what &lt;code&gt;PARTLY&lt;/code&gt; was intended to mean.&lt;/p&gt;

&lt;p&gt;Not uncertainty.&lt;/p&gt;

&lt;p&gt;Not closeness.&lt;/p&gt;

&lt;p&gt;Actual mixed semantic truth.&lt;/p&gt;




&lt;h1&gt;
  
  
  Most importantly, concise correct theories started solving the puzzle
&lt;/h1&gt;

&lt;p&gt;Question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does light wake him?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Result:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Yes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then I reused the exact same text as a theory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does light wake him?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Result:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You got it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Another phrasing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;the room gets bright and that wakes him
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Yes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Theory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You got it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is much closer to how a human Game Master should behave.&lt;/p&gt;




&lt;h1&gt;
  
  
  Wrong causal mechanisms still failed
&lt;/h1&gt;

&lt;p&gt;Simplification did not mean blindly accepting anything containing the correct keyword.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;lamp is there but vibration wakes him
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Theory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Not quite yet.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That matters because false &lt;code&gt;SOLVED&lt;/code&gt; results are much worse for this game than conservative partial judgments.&lt;/p&gt;

&lt;p&gt;So far, Version 3 became more tolerant without immediately destroying the distinction between correct and incorrect causal explanations.&lt;/p&gt;




&lt;h1&gt;
  
  
  The behavior also generalized to another puzzle
&lt;/h1&gt;

&lt;p&gt;I tested a different puzzle involving an apparently empty frame whose contents seem to change.&lt;/p&gt;

&lt;p&gt;Some examples:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is the frame big
→ Not relevant
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it an animal
→ No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it a TV
→ No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it a machine
→ No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;it changes its color with sun light
→ Yes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Submitted as a theory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;it changes its color with sun light
→ You got it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Again, the grammar is not perfect.&lt;/p&gt;

&lt;p&gt;But the causal idea is correct.&lt;/p&gt;

&lt;p&gt;And the judge accepted it.&lt;/p&gt;




&lt;h1&gt;
  
  
  Version 3 is not perfect
&lt;/h1&gt;

&lt;p&gt;There is still some instability around boundary cases.&lt;/p&gt;

&lt;p&gt;For example, during one session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it the sunrise and sunset
→ Yes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;while similar wording later produced:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it sunrise and sunset
→ Not enough information
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That tells me the system is not magically deterministic at the semantic level.&lt;/p&gt;

&lt;p&gt;Nor should I expect it to be.&lt;/p&gt;

&lt;p&gt;The important difference is that the major player-facing failures changed from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"I clearly understand what you're asking,
but please rephrase."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to occasional disagreement around genuinely fuzzy boundaries.&lt;/p&gt;

&lt;p&gt;For a proof-of-concept game, that is a much better failure mode.&lt;/p&gt;




&lt;h1&gt;
  
  
  So what actually fixed it?
&lt;/h1&gt;

&lt;p&gt;The honest answer is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I do not know which individual change fixed it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Version 3 changed several things together:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Removed multiple continuous semantic gates

Removed the 0.8/0.2 decision chain

Replaced many Nouls with one Choice distribution

Added mysteryFocus

Shortened and generalized the rubric

Made language recovery part of the single semantic classification
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Any one of those may have contributed.&lt;/p&gt;

&lt;p&gt;The improvement is evidence that the overall architecture became better.&lt;/p&gt;

&lt;p&gt;It is not a controlled experiment proving that &lt;code&gt;0.8&lt;/code&gt; alone was the problem.&lt;/p&gt;

&lt;p&gt;A useful future experiment would compare Version 2 under several threshold policies.&lt;/p&gt;

&lt;p&gt;That could answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Was Jev already producing useful semantic probabilities that my application was throwing away?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I think that is entirely possible.&lt;/p&gt;




&lt;h1&gt;
  
  
  Jev started feeling less like "adding AI" and more like building probabilistic IF statements
&lt;/h1&gt;

&lt;p&gt;This project changed how I think about decision models.&lt;/p&gt;

&lt;p&gt;I am not really building a chatbot.&lt;/p&gt;

&lt;p&gt;I am building something conceptually closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;semanticCondition&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;doSomething&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Except the condition is not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Does this player's sentence semantically express
a proposition supported by the hidden explanation?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Jev estimates that semantic condition.&lt;/p&gt;

&lt;p&gt;TypeScript decides what happens next.&lt;/p&gt;

&lt;p&gt;That is why I increasingly think of this pattern as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;probabilistic IF statements over natural language&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;an AI-powered switch/case&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Version 2 accidentally turned that simple idea into a giant conditional expression.&lt;/p&gt;

&lt;p&gt;Version 3 moved it back toward a small semantic switch.&lt;/p&gt;




&lt;h1&gt;
  
  
  There is also a hidden engineering cost: calibration
&lt;/h1&gt;

&lt;p&gt;Another lesson was that API pricing is not the only cost of using a decision model.&lt;/p&gt;

&lt;p&gt;For this project I needed to repeatedly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;design semantic categories
test real player wording
test spelling mistakes
test short questions
test false mechanisms
test incomplete solutions
check false positives
check false negatives
change policies
retest on protected Preview
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is human labor.&lt;/p&gt;

&lt;p&gt;Jev reduced one kind of complexity for me.&lt;/p&gt;

&lt;p&gt;I did not have to parse arbitrary assistant prose into application state.&lt;/p&gt;

&lt;p&gt;But that complexity did not disappear.&lt;/p&gt;

&lt;p&gt;Some of it moved into:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;decision design
calibration
threshold selection
test-play
behavioral evaluation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I would therefore not say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Jev requires more engineering work than a traditional LLM.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I have not run a controlled comparison that proves that.&lt;/p&gt;

&lt;p&gt;A more accurate statement is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;In this project, Jev moved a meaningful part of the engineering effort from output handling to decision design and calibration.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is an important cost to account for.&lt;/p&gt;




&lt;h1&gt;
  
  
  849 tests still did not replace test-play
&lt;/h1&gt;

&lt;p&gt;Version 3 currently passes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;849 unit/API tests
82 desktop/mobile E2E tests
Leak scanning
Lint
Type checking
Build
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those tests are valuable.&lt;/p&gt;

&lt;p&gt;They prove things such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;the parser rejects invalid distributions
the deterministic mapping works
private grading data does not reach the browser
provider call counts remain bounded
the UI still behaves correctly
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But they still cannot fully prove:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A human writes an unseen sentence
        ↓
The model understands it as intended
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That requires actual behavioral testing.&lt;/p&gt;

&lt;p&gt;For AI systems I now think about validation as two separate layers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Layer 1
Software correctness

Layer 2
Model behavior
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A green CI pipeline is excellent evidence for Layer 1.&lt;/p&gt;

&lt;p&gt;It is not sufficient evidence for Layer 2.&lt;/p&gt;




&lt;h1&gt;
  
  
  Why I am stopping here
&lt;/h1&gt;

&lt;p&gt;There are still things I could tune.&lt;/p&gt;

&lt;p&gt;I could add more categories.&lt;/p&gt;

&lt;p&gt;I could add confidence gates.&lt;/p&gt;

&lt;p&gt;I could add more examples.&lt;/p&gt;

&lt;p&gt;I could create exceptions for phrases like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sunrise and sunset
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I am deliberately not doing that.&lt;/p&gt;

&lt;p&gt;That path is exactly how Version 2 became complicated.&lt;/p&gt;

&lt;p&gt;For this PoC, Version 3 is now a strong freeze candidate.&lt;/p&gt;

&lt;p&gt;The remaining requirement is not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;perfect semantic classification
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Natural short player language usually works.

Minor grammar and spelling errors usually work.

Correct causal explanations can solve the puzzle.

Wrong causal explanations do not solve it.

Hidden information stays hidden.

The game remains fun.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is enough.&lt;/p&gt;




&lt;h1&gt;
  
  
  What if this still breaks later?
&lt;/h1&gt;

&lt;p&gt;I already have a fallback.&lt;/p&gt;

&lt;p&gt;Simplify even further:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;YES
OTHER
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;OTHER&lt;/code&gt; would intentionally collapse:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;NO
UNKNOWN
IRRELEVANT
AMBIGUOUS
possibly MIXED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That would lose information.&lt;/p&gt;

&lt;p&gt;But if it produced a better game experience, I would seriously consider it.&lt;/p&gt;

&lt;p&gt;Again:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The goal is not to extract the maximum amount of semantic information from the model.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The goal is to ask for the smallest useful decision.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final takeaway
&lt;/h1&gt;

&lt;p&gt;The most surprising lesson from building Sideways has been this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;More semantic structure does not automatically produce better AI behavior.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Version 2 looked more sophisticated.&lt;/p&gt;

&lt;p&gt;It had more probabilities.&lt;/p&gt;

&lt;p&gt;More semantic axes.&lt;/p&gt;

&lt;p&gt;More explicit thresholds.&lt;/p&gt;

&lt;p&gt;More detailed grading logic.&lt;/p&gt;

&lt;p&gt;And worse real gameplay.&lt;/p&gt;

&lt;p&gt;Version 3 asked Jev to do less:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;One semantic Choice
        ↓
One probability distribution
        ↓
Deterministic TypeScript switch
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the Game Master got noticeably better.&lt;/p&gt;

&lt;p&gt;So the principle I am taking forward is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Do not ask AI to make a more complicated decision than your product actually needs.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Or even shorter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Understand enough.

Decide less.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That turned out to be a much better architecture for this game.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>typescript</category>
      <category>nextjs</category>
      <category>jev</category>
    </item>
    <item>
      <title>When More Semantic Signals Made My AI Judge Worse: Building a Lateral Thinking Game with Jev</title>
      <dc:creator>494</dc:creator>
      <pubDate>Sun, 04 Oct 2026 12:48:10 +0000</pubDate>
      <link>https://dev.to/kitadaro/when-more-semantic-signals-made-my-ai-judge-worse-building-a-lateral-thinking-game-with-jev-2424</link>
      <guid>https://dev.to/kitadaro/when-more-semantic-signals-made-my-ai-judge-worse-building-a-lateral-thinking-game-with-jev-2424</guid>
      <description>&lt;h1&gt;
  
  
  When More Semantic Signals Made My AI Judge Worse
&lt;/h1&gt;

&lt;p&gt;I have been building a small web game called &lt;strong&gt;Sideways&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It is a lateral-thinking puzzle game: the player sees a mysterious situation and tries to discover what really happened by asking questions.&lt;/p&gt;

&lt;p&gt;A typical interaction should feel almost trivial:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Player:
Is it light that wakes the guest?

Game Master:
Yes.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Player:
Does the wake-up mechanism make sound?

Game Master:
No.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The interesting part is that the Game Master is powered by AI.&lt;/p&gt;

&lt;p&gt;I chose &lt;strong&gt;Jev&lt;/strong&gt; because I did not want a chatbot generating long answers. I wanted a decision layer that could take application state, make a typed judgment, and let ordinary TypeScript decide what the product should do.&lt;/p&gt;

&lt;p&gt;Jev describes itself as a decision model rather than a chat model. Its API works with application state and typed questions such as &lt;code&gt;Choice&lt;/code&gt;, &lt;code&gt;Score&lt;/code&gt;, and &lt;code&gt;Noul&lt;/code&gt;, returning structured results and probabilities rather than conversational prose.&lt;/p&gt;

&lt;p&gt;That sounded almost perfect for this game.&lt;/p&gt;

&lt;p&gt;But I discovered something unexpected:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The more semantic intelligence I tried to extract from the model, the worse the actual game experience became.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This article is about that failure.&lt;/p&gt;

&lt;p&gt;And why I am now making the AI do less.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Architecture I Wanted
&lt;/h2&gt;

&lt;p&gt;From the beginning, I did not want the model to control the application directly.&lt;/p&gt;

&lt;p&gt;The architecture was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Player text
    ↓
Probabilistic semantic judgment
    ↓
Typed output
    ↓
Deterministic TypeScript policy
    ↓
Player-facing result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The principle was simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI = semantic judgment
Code = deterministic policy
User = consequential intent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model should understand language.&lt;/p&gt;

&lt;p&gt;My code should decide what happens next.&lt;/p&gt;

&lt;p&gt;This has several advantages.&lt;/p&gt;

&lt;p&gt;The browser never needs to receive the hidden solution.&lt;/p&gt;

&lt;p&gt;The model does not get to decide arbitrary UI behavior.&lt;/p&gt;

&lt;p&gt;The public API can expose only simple results such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;YES
NO
PARTLY
IRRELEVANT
UNCLEAR
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;while keeping the private reasoning signals on the server.&lt;/p&gt;

&lt;p&gt;For solution attempts, the public results are similarly constrained:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SOLVED
PARTIAL
NOT_SOLVED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Conceptually, I still believe this architecture is correct.&lt;/p&gt;

&lt;p&gt;The problem was how much semantic structure I asked the model to produce.&lt;/p&gt;




&lt;h2&gt;
  
  
  Version 1: Simple Evidence
&lt;/h2&gt;

&lt;p&gt;The first Question Judge was relatively simple.&lt;/p&gt;

&lt;p&gt;For a player's question, the model estimated something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;wellFormed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;relevant&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nl"&gt;entailed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;contradicted&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;undetermined&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;TypeScript then converted these probabilities into a product result.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;wellFormed&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;UNCLEAR&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;relevant&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;IRRELEVANT&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;entailed&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;YES&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;contradicted&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;NO&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;UNCLEAR&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This had an important property:&lt;/p&gt;

&lt;p&gt;A low probability of &lt;code&gt;YES&lt;/code&gt; did not automatically mean &lt;code&gt;NO&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;An unknown fact could remain unknown.&lt;/p&gt;

&lt;p&gt;That part worked well.&lt;/p&gt;

&lt;p&gt;But actual gameplay revealed another problem.&lt;/p&gt;

&lt;p&gt;Natural player language is messy.&lt;/p&gt;

&lt;p&gt;People write things like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;light wake him?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does it make sound
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;maybe lamp?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;They make spelling mistakes.&lt;/p&gt;

&lt;p&gt;They omit articles.&lt;/p&gt;

&lt;p&gt;They use pronouns.&lt;/p&gt;

&lt;p&gt;They do not write propositions like lawyers.&lt;/p&gt;

&lt;p&gt;So I tried to make the judge more sophisticated.&lt;/p&gt;




&lt;h2&gt;
  
  
  Version 2: More Semantic Intelligence
&lt;/h2&gt;

&lt;p&gt;In Version 2, the Question Judge expanded into:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;wellFormed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;relevant&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;entailed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;contradicted&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;undetermined&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="nx"&gt;mixedClaims&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;supported&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;contradicted&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This allowed a new result:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PARTLY
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;There is no alarm. Only a lamp wakes the guest.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;contains both a false claim and an important true claim.&lt;/p&gt;

&lt;p&gt;I did not want that to be treated as simply &lt;code&gt;NO&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;For solution checking, I went even further.&lt;/p&gt;

&lt;p&gt;Instead of judging a theory with a simple checklist, the model returned five semantic dimensions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;coreMechanism&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;causalRelation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;supportingInsight&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;contextualError&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;competingMechanism&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The idea was reasonable.&lt;/p&gt;

&lt;p&gt;Consider these three theories:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;There is a light.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The light wakes the guest.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;There is a light, but vibration wakes the guest.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;They should not receive the same result.&lt;/p&gt;

&lt;p&gt;I wanted:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"There is a light."
→ PARTIAL

"The light wakes the guest."
→ SOLVED

"Light exists, but vibration wakes the guest."
→ NOT_SOLVED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So I separated:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;identifying the important object;&lt;/li&gt;
&lt;li&gt;understanding the causal mechanism;&lt;/li&gt;
&lt;li&gt;making a harmless contextual mistake;&lt;/li&gt;
&lt;li&gt;asserting a competing wrong mechanism.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Architecturally, this looked much better.&lt;/p&gt;




&lt;h2&gt;
  
  
  And the Offline Tests Looked Excellent
&lt;/h2&gt;

&lt;p&gt;The implementation passed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;880 unit/API tests
82 desktop/mobile E2E tests
Leak scans
Lint
Type checking
Build
Full repository checks
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I also created separate calibration and engineering-generalization corpora across all six puzzles.&lt;/p&gt;

&lt;p&gt;The browser did not receive private grading data.&lt;/p&gt;

&lt;p&gt;The hidden solution remained server-only.&lt;/p&gt;

&lt;p&gt;The number of provider calls remained bounded.&lt;/p&gt;

&lt;p&gt;Everything looked clean.&lt;/p&gt;

&lt;p&gt;But there was one very important limitation.&lt;/p&gt;

&lt;p&gt;Most of those tests proved:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;semantic signal
→ parser
→ deterministic policy
→ API result
→ UI
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;They did &lt;strong&gt;not&lt;/strong&gt; prove that the real model would infer the expected semantic signal from previously unseen player language.&lt;/p&gt;

&lt;p&gt;That distinction turned out to matter a lot.&lt;/p&gt;




&lt;h2&gt;
  
  
  Then I Played the Game with the Real Model
&lt;/h2&gt;

&lt;p&gt;I deployed the new judge to a protected preview environment and started typing ordinary questions.&lt;/p&gt;

&lt;p&gt;The results were surprising.&lt;/p&gt;

&lt;h3&gt;
  
  
  Example 1
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Player:
is it light?

Result:
Not enough information
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Example 2
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Player:
is it light that wakes the guest?

Result:
Not enough information
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then I submitted the exact same sentence as a theory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Theory:
is it light that wakes the guest?

Result:
Not quite yet.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A human Game Master would almost certainly understand the intended idea.&lt;/p&gt;

&lt;p&gt;The hidden solution is that a visual light signal wakes the sleeping guest.&lt;/p&gt;




&lt;h3&gt;
  
  
  Example 3
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Player:
does the wake up mechanism make sound?

Result:
Please rephrase
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Even this typo-heavy version:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does the wake up mechines make sound?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;also produced:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Please rephrase
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That was a serious UX problem.&lt;/p&gt;

&lt;p&gt;The grammar is imperfect, but the semantic intent is obvious.&lt;/p&gt;




&lt;h3&gt;
  
  
  Example 4
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Player:
Instead of an alarm, lights from a device wakes the guest.

Question result:
Not enough information

Theory result:
Not quite yet.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Again, the wording is imperfect.&lt;/p&gt;

&lt;p&gt;But the important causal insight is clearly present.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Did More Structure Make Things Worse?
&lt;/h2&gt;

&lt;p&gt;I cannot infer Jev's internal reasoning from its output alone.&lt;/p&gt;

&lt;p&gt;So the following are hypotheses about &lt;strong&gt;my decision architecture&lt;/strong&gt;, not claims about the internals of the model.&lt;/p&gt;

&lt;p&gt;Several problems stood out.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. The Semantic Dimensions Were Not Really Independent
&lt;/h2&gt;

&lt;p&gt;I had separated concepts like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;well-formedness
relevance
truth
mixed claims
causal understanding
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;because it made the TypeScript architecture easier to reason about.&lt;/p&gt;

&lt;p&gt;But humans do not necessarily understand a sentence in that order.&lt;/p&gt;

&lt;p&gt;Take:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does the wake up mechines make sound?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A human probably performs something closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"mechines" is probably "mechanism"
        ↓
The player means the wake-up mechanism
        ↓
They are asking whether it makes sound
        ↓
I know the answer
        ↓
No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is one integrated semantic interpretation.&lt;/p&gt;

&lt;p&gt;I had converted it into several semi-independent judgments.&lt;/p&gt;

&lt;p&gt;That created more opportunities for one signal to disagree with the others.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Probabilities Became Hard Cliffs
&lt;/h2&gt;

&lt;p&gt;The intention behind probabilities was to avoid brittle Boolean decisions.&lt;/p&gt;

&lt;p&gt;But then my code contained rules such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;wellFormed&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;REPHRASE&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Imagine the model essentially understands the sentence but produces:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;wellFormed = 0.77
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rest of the semantic evidence no longer matters.&lt;/p&gt;

&lt;p&gt;The application suddenly says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Please rephrase.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A continuous probability had become a discrete cliff.&lt;/p&gt;

&lt;p&gt;Adding more semantic dimensions meant adding more places where such cliffs could occur.&lt;/p&gt;

&lt;p&gt;The model did not necessarily need to be dramatically wrong.&lt;/p&gt;

&lt;p&gt;My policy only needed one signal to fall on the wrong side of a threshold.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. I Was Asking a Small Decision Model to Solve a Large Semantic Task
&lt;/h2&gt;

&lt;p&gt;Jev's documentation emphasizes focused, well-scoped typed decisions. &lt;code&gt;Choice&lt;/code&gt; is designed for predefined classification, while &lt;code&gt;Noul&lt;/code&gt; represents a yes/no judgment whose probability can be interpreted by application code.&lt;/p&gt;

&lt;p&gt;But my Question Judge was effectively asking the model to do all of this at once:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Recover imperfect language
Resolve pronouns
Determine the semantic proposition
Determine relevance
Compare with hidden truth
Detect multiple claims
Separate supported and contradicted claims
Distinguish uncertainty from contradiction
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The individual fields looked small.&lt;/p&gt;

&lt;p&gt;The total semantic task was not.&lt;/p&gt;

&lt;p&gt;I had decomposed the output without necessarily decomposing the difficulty.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. The Judge Knew the Scenario, but Not the "Mystery Focus"
&lt;/h2&gt;

&lt;p&gt;This turned out to be especially interesting.&lt;/p&gt;

&lt;p&gt;Consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it light?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;As an isolated English sentence, &lt;code&gt;it&lt;/code&gt; is ambiguous.&lt;/p&gt;

&lt;p&gt;What is "it"?&lt;/p&gt;

&lt;p&gt;The alarm?&lt;/p&gt;

&lt;p&gt;The mechanism?&lt;/p&gt;

&lt;p&gt;The object?&lt;/p&gt;

&lt;p&gt;But in a lateral-thinking game, a human Game Master has another piece of context:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What mystery is the player currently trying to explain?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For the quiet-alarm puzzle, that focus is roughly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What causes the sleeping guest to wake
without sound or physical contact?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Given that context:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it light?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;has a very natural interpretation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Is light the cause?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I had provided the scenario and hidden reference, but not an explicit representation of the &lt;strong&gt;question being investigated&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That may have encouraged overly literal ambiguity handling.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. My Solution Judge May Have Confused Completeness with Correctness
&lt;/h2&gt;

&lt;p&gt;The Version 2 Solution Judge required strong scores for both:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;coreMechanism
causalRelation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;before returning &lt;code&gt;SOLVED&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That sounds reasonable.&lt;/p&gt;

&lt;p&gt;But consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The light wakes the guest.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is extremely short.&lt;/p&gt;

&lt;p&gt;It also contains both the mechanism and causal relationship.&lt;/p&gt;

&lt;p&gt;If the model gives:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;coreMechanism = high
causalRelation = medium
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the deterministic policy returns &lt;code&gt;PARTIAL&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;From the application's perspective, that is a false negative.&lt;/p&gt;

&lt;p&gt;The player has solved the mystery.&lt;/p&gt;

&lt;p&gt;The grading representation made the answer look less complete than it actually was.&lt;/p&gt;




&lt;h1&gt;
  
  
  Version 3: Make the Model Do Less
&lt;/h1&gt;

&lt;p&gt;So I am now testing a different architecture.&lt;/p&gt;

&lt;p&gt;Instead of asking for many independent semantic scores, each judge will return &lt;strong&gt;one categorical probability distribution&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For Question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;QuestionSemanticClass&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;SUPPORTED&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;CONTRADICTED&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;MIXED&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;UNKNOWN&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;IRRELEVANT&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;AMBIGUOUS&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then ordinary TypeScript maps it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SUPPORTED
→ YES

CONTRADICTED
→ NO

MIXED
→ PARTLY

UNKNOWN
→ NOT ENOUGH INFORMATION

IRRELEVANT
→ NOT RELEVANT

AMBIGUOUS
→ PLEASE REPHRASE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is still probabilistic information.&lt;/p&gt;

&lt;p&gt;But there is only one semantic decision.&lt;/p&gt;

&lt;p&gt;No chain of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0.8
0.8
0.2
0.8
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;gates.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Solution Judge Is Being Simplified Too
&lt;/h2&gt;

&lt;p&gt;Instead of five independent dimensions, I am testing one semantic classification:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;SolutionSemanticClass&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;CORE_CAUSAL_EXPLANATION&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;CORE_WITH_CONTEXT_ERROR&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;SUPPORTING_INSIGHT_ONLY&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;COMPETING_WRONG_MECHANISM&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;NO_MEANINGFUL_INSIGHT&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;TypeScript then maps:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CORE_CAUSAL_EXPLANATION
→ SOLVED

CORE_WITH_CONTEXT_ERROR
→ PARTIAL

SUPPORTING_INSIGHT_ONLY
→ PARTIAL

COMPETING_WRONG_MECHANISM
→ NOT_SOLVED

NO_MEANINGFUL_INSIGHT
→ NOT_SOLVED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important distinction remains.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;There is a light.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;should not equal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The light wakes the guest.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And neither should equal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The vibration wakes the guest.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But the model no longer needs to independently estimate five continuous semantic dimensions before my application can make that distinction.&lt;/p&gt;




&lt;h2&gt;
  
  
  I Am Also Adding a "Mystery Focus"
&lt;/h2&gt;

&lt;p&gt;Each puzzle will have a small server-only field such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What causes the sleeping guest to wake
without sound or physical contact?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is not the hidden answer.&lt;/p&gt;

&lt;p&gt;It simply represents the question raised by the public scenario.&lt;/p&gt;

&lt;p&gt;The hope is that it helps resolve natural shorthand:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is it light?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;does it make sound?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;was he covered?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;without adding phrase-specific exceptions.&lt;/p&gt;

&lt;p&gt;That distinction matters.&lt;/p&gt;

&lt;p&gt;I do not want:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;if text === "is it light?"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I want:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;short player language
+
current mystery
→ recoverable semantic proposition
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  And If Version 3 Still Fails?
&lt;/h1&gt;

&lt;p&gt;Then I plan to simplify again.&lt;/p&gt;

&lt;p&gt;Possibly all the way to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;YES
OTHER
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For Question judging:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;YES
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;would mean:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This player proposition is supported by the hidden explanation.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Everything else would collapse into:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OTHER
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That would intentionally stop distinguishing among:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;NO
UNKNOWN
IRRELEVANT
AMBIGUOUS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This loses information.&lt;/p&gt;

&lt;p&gt;It may still produce a better game.&lt;/p&gt;

&lt;p&gt;That is the part of this experiment I find most interesting.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Product Does Not Need the Maximum Amount of AI Information
&lt;/h1&gt;

&lt;p&gt;It is tempting to think:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;More semantic signals
→ more information
→ better product decisions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;My experience so far suggests something more like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;More semantic signals
→ more uncertain boundaries
→ more policy interactions
→ potentially worse UX
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The best AI interface may not be the one that extracts the most information from a model.&lt;/p&gt;

&lt;p&gt;It may be the one that asks the model for the &lt;strong&gt;smallest decision the product actually needs&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  Another Lesson: Offline Tests and Semantic Accuracy Are Different Things
&lt;/h1&gt;

&lt;p&gt;I had hundreds of passing tests.&lt;/p&gt;

&lt;p&gt;They were useful.&lt;/p&gt;

&lt;p&gt;They caught:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;schema mistakes;&lt;/li&gt;
&lt;li&gt;API regressions;&lt;/li&gt;
&lt;li&gt;accidental hidden-data leaks;&lt;/li&gt;
&lt;li&gt;extra provider calls;&lt;/li&gt;
&lt;li&gt;incorrect deterministic mappings;&lt;/li&gt;
&lt;li&gt;broken UI behavior.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I absolutely want those tests.&lt;/p&gt;

&lt;p&gt;But:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;synthetic semantic signals
→ correct product output
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;does not prove:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;previously unseen human language
→ correct semantic signal
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those are different systems.&lt;/p&gt;

&lt;p&gt;For AI products, both need testing.&lt;/p&gt;

&lt;p&gt;I now think of them as two separate validation layers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Layer 1:
Software correctness

Layer 2:
Model behavior under real language
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A green CI pipeline proves the first.&lt;/p&gt;

&lt;p&gt;It does not automatically prove the second.&lt;/p&gt;




&lt;h1&gt;
  
  
  Why I Still Like the Typed-Decision Architecture
&lt;/h1&gt;

&lt;p&gt;None of this has made me want to replace the system with a free-form chatbot.&lt;/p&gt;

&lt;p&gt;Actually, the opposite.&lt;/p&gt;

&lt;p&gt;I still want:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI
→ bounded semantic decision

TypeScript
→ application policy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Jev's API is explicitly designed around application state and typed decisions, with probability distributions that software can consume.&lt;/p&gt;

&lt;p&gt;The lesson for me is not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Typed decisions are too simple.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I should respect their simplicity.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If the application needs one classification, asking for one well-designed &lt;code&gt;Choice&lt;/code&gt; may be better than constructing a small semantic ontology and connecting it with threshold gates.&lt;/p&gt;




&lt;h1&gt;
  
  
  Where the Experiment Stands
&lt;/h1&gt;

&lt;p&gt;At the time of writing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;v1
Simple probabilistic evidence
        ↓

v2
More semantic dimensions
        ↓

Real-model testing exposed brittle behavior
        ↓

v3
One categorical semantic distribution per judge
+ explicit mystery focus
        ↓

Currently being tested
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If v3 performs well on previously unseen wording, I will freeze the Judge and move on to production hardening.&lt;/p&gt;

&lt;p&gt;If it does not, I will test the &lt;code&gt;YES / OTHER&lt;/code&gt; design rather than adding more semantic complexity.&lt;/p&gt;

&lt;p&gt;That constraint is intentional.&lt;/p&gt;

&lt;p&gt;At some point, engineering discipline means stopping optimization.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Takeaway
&lt;/h1&gt;

&lt;p&gt;The most useful lesson from this project so far is surprisingly simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Do not ask the AI to make a more complicated decision than your product actually needs.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A sophisticated semantic representation can be elegant in TypeScript.&lt;/p&gt;

&lt;p&gt;It can have clean types.&lt;/p&gt;

&lt;p&gt;It can have excellent unit tests.&lt;/p&gt;

&lt;p&gt;It can even look more theoretically correct.&lt;/p&gt;

&lt;p&gt;And still produce a worse experience for the person actually playing the game.&lt;/p&gt;

&lt;p&gt;Sometimes the better AI architecture is not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Understand more.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Decide less.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I will update this article once the simplified Version 3 Judge has been tested against the real model.&lt;/p&gt;




&lt;h2&gt;
  
  
  About the project
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Sideways&lt;/strong&gt; is an experimental lateral-thinking puzzle game built with Next.js, TypeScript, and Jev.&lt;/p&gt;

&lt;p&gt;The project explores a specific architecture:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Probabilistic Semantic Judgment
→ Typed Output
→ Deterministic TypeScript Policy
→ Product Result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The next question is how small that semantic judgment can become while still producing a good game.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>typescript</category>
      <category>nextjs</category>
      <category>jev</category>
    </item>
  </channel>
</rss>
