<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: horio</title>
    <description>The latest articles on DEV Community by horio (@forifor).</description>
    <link>https://dev.to/forifor</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4122314%2Ffdbc052c-9c38-4e2f-be6a-fd085a34a889.png</url>
      <title>DEV Community: horio</title>
      <link>https://dev.to/forifor</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/forifor"/>
    <language>en</language>
    <item>
      <title>Five ways a voice agent tells you it booked a table when it didn't</title>
      <dc:creator>horio</dc:creator>
      <pubDate>Tue, 22 Sep 2026 03:24:48 +0000</pubDate>
      <link>https://dev.to/forifor/five-ways-a-voice-agent-tells-you-it-booked-a-table-when-it-didnt-1lpj</link>
      <guid>https://dev.to/forifor/five-ways-a-voice-agent-tells-you-it-booked-a-table-when-it-didnt-1lpj</guid>
      <description>&lt;p&gt;I maintain &lt;a href="https://github.com/FORIFOR/oathra" rel="noopener noreferrer"&gt;Oathra&lt;/a&gt;, an open-source runtime for AI agents that make&lt;br&gt;
phone calls. Its one opinion is that the model does not get to decide whether the call succeeded. A&lt;br&gt;
separate piece of code reads the other party's words and decides.&lt;/p&gt;

&lt;p&gt;That sounds like belt and braces until you look at what clerks actually say. Here are five replies to&lt;br&gt;
the same request. Only one of them is a booking. A model asked "did this succeed?" says yes to at least&lt;br&gt;
four, because all five &lt;em&gt;sound&lt;/em&gt; like a yes.&lt;/p&gt;

&lt;p&gt;The request, every time:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI: Could I book a table for two at 7:30 pm on September 25? The name is Tanaka.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  1. The pencil
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Them: Sure, I will pencil you in for September 25 at 7:30 pm and call you back to confirm.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Verdict: &lt;strong&gt;incomplete&lt;/strong&gt;, &lt;code&gt;{ date, time, partySize }&lt;/code&gt; extracted, &lt;code&gt;confirmed&lt;/code&gt; missing.&lt;/p&gt;

&lt;p&gt;"Sure" is agreement. The date, the time and the party size are all there, and they are all correct.&lt;br&gt;
The only thing missing is the thing you called about. A completion check that counts filled fields&lt;br&gt;
passes this. A check that requires the callee to have committed does not.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. The explicit hold
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Them: We can hold September 25 at 7:30 pm for now, but it is not confirmed yet.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Verdict: &lt;strong&gt;incomplete&lt;/strong&gt;, nothing extracted at all.&lt;/p&gt;

&lt;p&gt;This one is interesting because the clerk is being maximally clear, and a naive extractor still walks&lt;br&gt;
away with &lt;code&gt;time = 19:30&lt;/code&gt;. The clause carrying the time is the one being negated. If you extract values&lt;br&gt;
per-utterance instead of per-clause, "not confirmed yet" and "7:30 pm" end up in different variables and&lt;br&gt;
the negation is lost on the way.&lt;/p&gt;
&lt;h2&gt;
  
  
  3. The counter-offer
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Them: The only thing left that evening is 9 pm. Would that do?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Verdict: &lt;strong&gt;incomplete&lt;/strong&gt;, nothing extracted.&lt;/p&gt;

&lt;p&gt;A number was said. It was not offered as your booking; it was offered as a question. The trap here is&lt;br&gt;
that the agent's &lt;em&gt;next&lt;/em&gt; turn is usually "9 pm works, thank you" — and if your extractor took &lt;code&gt;21:00&lt;/code&gt;&lt;br&gt;
from the clerk's turn and your agent then says something agreeable, you have a fully populated booking&lt;br&gt;
that nobody ever agreed to. The clerk's turn has to stay a proposal until the caller accepts it and the&lt;br&gt;
clerk acknowledges the acceptance.&lt;/p&gt;
&lt;h2&gt;
  
  
  4. The retraction
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Them: Yes, we have you down for two at 7:30 pm on September 25.
Them: Sorry, that day is fully booked after all.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Verdict: &lt;strong&gt;incomplete&lt;/strong&gt;, &lt;code&gt;{ date, time, partySize }&lt;/code&gt; extracted, &lt;code&gt;confirmed&lt;/code&gt; gone.&lt;/p&gt;

&lt;p&gt;This is the one I got wrong. Until yesterday, Oathra reported this as &lt;strong&gt;completed&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;My retraction pattern required the refusal to name the booking — "we cannot take the reservation",&lt;br&gt;
"the booking is cancelled". A clerk who simply says the slot is gone does not phrase it that way. They&lt;br&gt;
say the table is taken, the day is private-hire, they are closed that day. The booking is equally dead&lt;br&gt;
and my code called it a success.&lt;/p&gt;

&lt;p&gt;The fix is narrow on purpose: an &lt;em&gt;availability&lt;/em&gt; word from the callee (full, private hire, closed,&lt;br&gt;
"fully booked") revokes a confirmation spoken strictly earlier. Not any refusal — "we can't take cards"&lt;br&gt;
after a booking is a payment remark, not a cancellation. And only &lt;em&gt;earlier&lt;/em&gt;, so the ordinary "7 pm is&lt;br&gt;
full but 7:30 is free" that happens &lt;strong&gt;before&lt;/strong&gt; a booking is untouched.&lt;/p&gt;

&lt;p&gt;It went out as &lt;a href="https://github.com/FORIFOR/oathra/pull/37" rel="noopener noreferrer"&gt;PR #37&lt;/a&gt; with tests in both directions.&lt;/p&gt;
&lt;h2&gt;
  
  
  5. The actual yes
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Them: You are all set for September 25 at 7:30 pm, party of two.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Verdict: &lt;strong&gt;completed&lt;/strong&gt;, &lt;code&gt;{ date: 2026-09-25, time: 19:30, partySize: 2, confirmed: true }&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This one also failed until today, in a way I find more embarrassing than #4: it returned &lt;strong&gt;nothing at&lt;br&gt;
all&lt;/strong&gt;. Not "unconfirmed" — empty.&lt;/p&gt;

&lt;p&gt;The engine had two separate ideas, "the callee agreed to a value the caller proposed" and "the callee&lt;br&gt;
confirmed the booking", and "you are all set" was in the second list but not the first. So nothing the&lt;br&gt;
caller had proposed was ever verified; and because a confirmation is bound to the values that were&lt;br&gt;
settled when it was spoken, a confirmation with no settled values behind it is stale and gets dropped.&lt;br&gt;
Two lists that should have overlapped, and the result is a blank screen for a perfectly normal sentence.&lt;/p&gt;

&lt;p&gt;I only found it because I pasted an English log into my own public checker while writing a comment on&lt;br&gt;
someone else's thread. Three of five natural English confirmations worked. The Japanese side, which I&lt;br&gt;
use daily, was fine. The lesson is not about regexes; it is that the language you don't test in is the&lt;br&gt;
language that's broken.&lt;/p&gt;
&lt;h2&gt;
  
  
  How this is kept honest
&lt;/h2&gt;

&lt;p&gt;Every release runs 10,000 seeded adversarial dialogues — the five shapes above plus voicemail,&lt;br&gt;
transfers, hold-then-reply, dialect confirmations, wrong restatements — against a hard gate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;False Completion: 0 / 10000 adversarial runs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A false completion is when the runtime reports "completed" and the callee's own ground truth says they&lt;br&gt;
never committed. Zero is the only passing number. It does not prove the checker is right about&lt;br&gt;
everything; it proves the specific ways I know a call can lie are all covered, and it fails loudly the&lt;br&gt;
moment a change reopens one.&lt;/p&gt;

&lt;p&gt;The inverse error — reporting "incomplete" when the table really was booked — is allowed to happen. It&lt;br&gt;
is a phone call you make again. The other direction is a customer standing outside a restaurant.&lt;/p&gt;
&lt;h2&gt;
  
  
  Watch the verdict change
&lt;/h2&gt;

&lt;p&gt;Twenty-six seconds of the checker being used, no sound. Each time the shop's reply is swapped, the same&lt;br&gt;
code runs again and the fields re-settle. The camera moves, the cursor and the captions were added in the&lt;br&gt;
edit; the screen and the verdicts were not.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://forifor.github.io/oathra/en/#verdict-video" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyndmpvksonkstitn9bbz.png" alt="The agent has said " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Try it on your own log
&lt;/h2&gt;

&lt;p&gt;If you run a voice agent, the fastest version of this is to paste one of your own transcripts into the&lt;br&gt;
checker. No install, no key, nothing leaves the tab:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://forifor.github.io/oathra/en/check.html" rel="noopener noreferrer"&gt;https://forifor.github.io/oathra/en/check.html&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One turn per line, with a speaker:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI: Could I book a table for two at 7:30 pm on September 25? The name is Tanaka.
Them: Sure, I will pencil you in and call you back to confirm.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If it reads your clerk's yes as a no, that is a bug I want — the supported phrasings are a list, and&lt;br&gt;
lists are always short somewhere. The repo is Apache-2.0:&lt;br&gt;
&lt;a href="https://github.com/FORIFOR/oathra" rel="noopener noreferrer"&gt;FORIFOR/oathra&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>typescript</category>
      <category>programming</category>
    </item>
    <item>
      <title>Four regexes and a staleness rule: how my phone agent refuses to call 'probably fine' a booking</title>
      <dc:creator>horio</dc:creator>
      <pubDate>Sun, 13 Sep 2026 15:38:00 +0000</pubDate>
      <link>https://dev.to/forifor/four-regexes-and-a-staleness-rule-how-my-phone-agent-refuses-to-call-probably-fine-a-booking-3ech</link>
      <guid>https://dev.to/forifor/four-regexes-and-a-staleness-rule-how-my-phone-agent-refuses-to-call-probably-fine-a-booking-3ech</guid>
      <description>&lt;p&gt;In &lt;a href="https://dev.to/forifor/i-let-an-ai-make-phone-calls-then-took-the-word-booked-away-from-it-5484"&gt;the launch post&lt;/a&gt; I said Oathra decides whether a call succeeded from what the &lt;em&gt;callee&lt;/em&gt; said, not from the model. This is the part that does it. It lives in &lt;a href="https://github.com/FORIFOR/oathra/tree/main/packages/evidence" rel="noopener noreferrer"&gt;packages/evidence&lt;/a&gt; and uses no LLM.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three lines that fooled the model
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;"Probably fine, but it's not confirmed yet" — a hedge&lt;/li&gt;
&lt;li&gt;"7 pm is full. 7:30 is open" — a refusal and an offer in one breath&lt;/li&gt;
&lt;li&gt;"Got it. But the price will be 23,500 yen" — acceptance, then a change of terms&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Asked "was the reservation made?", an LLM would say yes to 1 and 3 often enough to matter. The conversation &lt;em&gt;flows&lt;/em&gt; like a yes. So completion moved out of the model into four rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rule 1: split into clauses before reading values
&lt;/h2&gt;

&lt;p&gt;"7 pm is full but 7:30 works" as one sentence yields two candidate times. The clause is split right after a contrast word (Japanese ですが/ますが/けど, English &lt;em&gt;but&lt;/em&gt;), and a clause carrying a negative (full, unavailable, can't…) never produces an offer.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sub&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;(?&amp;lt;&lt;/span&gt;&lt;span class="sr"&gt;=ですが|ますが|けど|けれど|but&lt;/span&gt;&lt;span class="se"&gt;\s)&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;7:00 drops out. 7:30 survives as a &lt;em&gt;callee offer&lt;/em&gt;, which stays pending until the agent accepts it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rule 2: one hedge word disqualifies the whole utterance
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;HEDGE_RE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="sr"&gt;/と思います|たぶん|多分|おそらく|かもしれません|確認します|調べてみ|probably|maybe|perhaps|I think|let me check|not sure|I'll check|might be/i&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Agreement is &lt;code&gt;AGREEMENT_RE &amp;amp;&amp;amp; !REFUSAL_RE &amp;amp;&amp;amp; !HEDGE_RE&lt;/code&gt;, so "probably fine" is not an agreement even though "fine" matches. The same guard sits in front of the confirmation check: "I think… you're booked" never sets &lt;code&gt;confirmed&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The calling side follows the same rule. The scripted agent used to restate its request after a hedge until it gave up. Since v0.1.1 it asks, at most twice, "can you confirm the reservation for the 12th, 7 pm, two people, or is it still tentative?"&lt;/p&gt;

&lt;h2&gt;
  
  
  Rule 3: &lt;code&gt;confirmed&lt;/code&gt; has exactly two entry points
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The callee says an explicit confirmation ("your table is booked", 「ご予約承りました」) — &lt;code&gt;CONFIRMATION_RE&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;The agent asks a yes/no confirmation question ("can you confirm the booking?") and the callee's reply &lt;em&gt;starts&lt;/em&gt; with an affirmative — &lt;code&gt;CONFIRM_REQUEST_RE&lt;/code&gt; then &lt;code&gt;AFFIRMATIVE_RE&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A reply containing a question mark is not an answer. A re-quote ("a non-smoking room would be 19,900 a night") is not an answer. And nothing the agent says counts: only nodes whose speaker is &lt;code&gt;callee&lt;/code&gt; are read.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rule 4: a confirmation is bound to the terms at the moment it was spoken
&lt;/h2&gt;

&lt;p&gt;When "you're booked" is followed by "the rate is 23,500", the earlier confirmation goes stale. The engine snapshots the settled values at the time of the confirmation plus whatever the callee restated in that utterance, and drops &lt;code&gt;confirmed&lt;/code&gt; if any of them differs from the final state.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;stale&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Object&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;entries&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;snapshot&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;(([&lt;/span&gt;&lt;span class="nx"&gt;field&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="kc"&gt;undefined&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;out&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;field&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;stale&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;confirmed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent then asks for the confirmation again.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does with the three lines
&lt;/h2&gt;

&lt;p&gt;Fed straight into the engine (&lt;a href="https://github.com/FORIFOR/oathra/blob/main/docs/launch/miscompletion-cases.md" rel="noopener noreferrer"&gt;log&lt;/a&gt;):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Callee line&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"Probably fine, but it's not confirmed yet"&lt;/td&gt;
&lt;td&gt;incomplete, no &lt;code&gt;confirmed&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"7 pm is full. 7:30 is open"&lt;/td&gt;
&lt;td&gt;incomplete, 7:00 never taken, 7:30 pending&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;…then "7:30 it is" / "booked for two at 7:30"&lt;/td&gt;
&lt;td&gt;complete, time=19:30&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Got it. But the price will be 23,500"&lt;/td&gt;
&lt;td&gt;incomplete&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Booked at 18,000" then "sorry, 23,500"&lt;/td&gt;
&lt;td&gt;incomplete, confirmation stale&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;You can try them in the browser: &lt;code&gt;npx oathra demo&lt;/code&gt;, pick Play (you answer the phone), and the lines are one-click buttons under the input.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest limit
&lt;/h2&gt;

&lt;p&gt;This depends on phrase coverage. CI requires 0 false completions over 10,000 mutated callees, but that is "zero within the mutations I wrote", not a real-call number. If you find a phrasing that slips through, &lt;a href="https://github.com/FORIFOR/oathra/issues" rel="noopener noreferrer"&gt;open an issue&lt;/a&gt;. It is usually one more alternation in a regex, and it is the most useful contribution.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>typescript</category>
      <category>opensource</category>
      <category>showdev</category>
    </item>
    <item>
      <title>I built an AI team that ships real work — and shows you the conversation</title>
      <dc:creator>horio</dc:creator>
      <pubDate>Sat, 12 Sep 2026 15:44:18 +0000</pubDate>
      <link>https://dev.to/forifor/i-built-an-ai-team-that-ships-real-work-and-shows-you-the-conversation-251a</link>
      <guid>https://dev.to/forifor/i-built-an-ai-team-that-ships-real-work-and-shows-you-the-conversation-251a</guid>
      <description>&lt;p&gt;Most "multi-agent" demos are bots narrating to each other. The log is the product; the work is secondary. I wanted the opposite: one request in, real files out, and a trail I could audit afterwards.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agent Team&lt;/strong&gt; is an open-source (MIT), local-first runtime where a Master plans, a Researcher / Builder / Reviewer actually do the work, and you get the artifacts &lt;strong&gt;plus&lt;/strong&gt; the real bot-to-bot messages, a timeline, and verification bound to each artifact revision.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/FORIFOR/Multibot" rel="noopener noreferrer"&gt;https://github.com/FORIFOR/Multibot&lt;/a&gt; · Site + 59s intro: &lt;a href="https://forifor.github.io/Multibot/" rel="noopener noreferrer"&gt;https://forifor.github.io/Multibot/&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What makes it different
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No scripted chat.&lt;/strong&gt; The chat panel is a projection of &lt;code&gt;message.sent&lt;/code&gt; events — messages that were really delivered to another bot's mailbox. A question from the Builder wakes the Researcher, which answers with &lt;code&gt;reply_to&lt;/code&gt;. Acknowledgements never wake a model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evidence, not vibes.&lt;/strong&gt; Checks and review verdicts are events bound to an artifact revision hash. The final report is compiled from the event log; a model summary cannot upgrade "started" to "done".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime-enforced limits.&lt;/strong&gt; Tool scope, write scope, budget reservation (spent + reserved for in-flight calls), approvals with hash + nonce, cancel / resume / fork — enforced in code, not prompt wording. Unknown model prices refuse to start.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-bot configuration.&lt;/strong&gt; Each bot inherits a default connection and model and can override endpoint, model, effort and system prompt (lockable). Configured vs. provider-reported model are both shown. No silent fallbacks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Replay never calls a model.&lt;/strong&gt; Fork from a checkpoint with a different model for one bot and compare. Export the whole run as JSONL.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  No API key needed (Claude Code)
&lt;/h2&gt;

&lt;p&gt;The default connection runs every agent session through the local &lt;code&gt;claude -p&lt;/code&gt;. Claude Code owns the loop for one session; the team's tools (&lt;code&gt;send_message&lt;/code&gt;, &lt;code&gt;publish_artifact&lt;/code&gt;, &lt;code&gt;run_check&lt;/code&gt;, …) are exposed to it as MCP tools through a small stdio proxy that forwards each call to the runtime's ToolGateway. Policy, budget and the event log are identical to the API path; Claude Code's own built-in tools are disabled for these sessions. Cost and usage are read from the CLI's JSON result, and the model it actually used comes from &lt;code&gt;modelUsage&lt;/code&gt;, never from the model's own claims.&lt;/p&gt;

&lt;p&gt;Claude API (official SDK), any OpenAI-compatible chat endpoint and local Ollama work the same way, per bot.&lt;/p&gt;

&lt;h2&gt;
  
  
  A real run, unedited
&lt;/h2&gt;

&lt;p&gt;On 2026-09-13 I ran this request through the local Claude Code CLI: &lt;em&gt;"From this product description, build a Japanese launch page and three social-post drafts. Record assumptions for anything missing. Stop before publishing. Have the Reviewer verify."&lt;/em&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model (reported by the provider)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;claude-opus-5&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plan&lt;/td&gt;
&lt;td&gt;Master chose 1 builder task + 1 reviewer task and skipped the researcher; 8 recorded assumptions (no prices, no invented numbers, placeholder URLs only)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deliverables&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;index.html&lt;/code&gt; (single-file page), &lt;code&gt;posts.md&lt;/code&gt;, &lt;code&gt;HANDOFF.md&lt;/code&gt;, &lt;code&gt;final-report.md&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Verification&lt;/td&gt;
&lt;td&gt;10 programmatic checks → all pass; reviewer verdict 6/6, plus 4 optional findings sent back as a real message&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Usage&lt;/td&gt;
&lt;td&gt;39 model turns · 35 tool calls · $1.66 list-price equivalent · 18 min 37 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Status&lt;/td&gt;
&lt;td&gt;completed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The generated files, the final report and the 77-event JSONL log are committed under &lt;code&gt;docs/evidence/&lt;/code&gt; without edits. Four runs total: run 1 finished &lt;em&gt;partial&lt;/em&gt; and exposed a bug (a reviewer verifying two tasks had only its last verdict applied — fixed), runs 2–3 hit limits that were then tuned. One request is not a benchmark, and the prompts are original seeds.&lt;/p&gt;

&lt;p&gt;The demo video on the site drives the real UI and runtime with a &lt;strong&gt;scripted test provider&lt;/strong&gt; (labelled on screen) so it is deterministic and free to reproduce.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;One request → a checked plan.&lt;/strong&gt; The Master turns the request into deliverables, assumptions and a task DAG (structured output). The runtime validates schema, cycles, owners, tools, write scopes and limits before anything runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Independent bots, real delivery.&lt;/strong&gt; Each bot has its own conversation state, mailbox, task scope, tools and workspace. A worker receives its task, input artifact refs and its own inbox — not the whole history.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify, revise, report from evidence.&lt;/strong&gt; The Reviewer runs checks against a specific revision. Fail → the Builder revises → re-review, bounded. The final report is compiled from the event log.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Stack: Python 3.12 / FastAPI / SQLite (WAL, append-only events, SSE), React + TypeScript. macOS seatbelt sandbox for builder commands (no network, writes only inside the task workspace); elsewhere a plain subprocess that says so in every result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/FORIFOR/Multibot &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;Multibot
&lt;span class="nb"&gt;cd &lt;/span&gt;backend &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; uv venv .venv &lt;span class="nt"&gt;--python&lt;/span&gt; 3.12 &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; uv pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--python&lt;/span&gt; .venv/bin/python &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="s1"&gt;'.[dev]'&lt;/span&gt;
&lt;span class="nb"&gt;cd&lt;/span&gt; ../frontend &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; pnpm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; pnpm build &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; ../backend
.venv/bin/agentteam probe    &lt;span class="c"&gt;# real capability check through `claude -p`&lt;/span&gt;
.venv/bin/agentteam serve    &lt;span class="c"&gt;# http://127.0.0.1:8787&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I'd love feedback on two things: whether the review → revise → re-review loop bound to revisions holds up on your tasks, and where "fewer messages, more evidence" (workers get their task, artifact refs and inbox — not the history) breaks down.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>python</category>
      <category>agents</category>
    </item>
    <item>
      <title>I let an AI make phone calls, then took the word "booked" away from it</title>
      <dc:creator>horio</dc:creator>
      <pubDate>Sat, 12 Sep 2026 15:39:30 +0000</pubDate>
      <link>https://dev.to/forifor/i-let-an-ai-make-phone-calls-then-took-the-word-booked-away-from-it-5484</link>
      <guid>https://dev.to/forifor/i-let-an-ai-make-phone-calls-then-took-the-word-booked-away-from-it-5484</guid>
      <description>&lt;p&gt;Getting an AI to place a phone call takes an evening. The trouble starts after that.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It said "your table is booked" when the restaurant had said no such thing.&lt;/li&gt;
&lt;li&gt;It talked to a voicemail greeting for two minutes, politely asking for a reservation.&lt;/li&gt;
&lt;li&gt;It accepted a price above the budget because the conversation had a nice flow.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All three happened on real calls to my own phone. So I built a runtime that &lt;strong&gt;takes the completion decision away from the model and gives it to code&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/FORIFOR/oathra" rel="noopener noreferrer"&gt;https://github.com/FORIFOR/oathra&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx oathra demo
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two AIs start a phone call in your browser, no API key. One plays a restaurant, the other wants a table. 7 pm is full, the restaurant offers 7:30, the caller takes it, gives a name, and only when the restaurant says "you're all set" does the run become MISSION COMPLETE.&lt;/p&gt;

&lt;p&gt;The result is not a summary. Every field is anchored to something the &lt;em&gt;other party&lt;/em&gt; said.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"field"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"time"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"19:30"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"source"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"callee"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"transcript"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"7 is fully booked, but we do have 7:30."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"span"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"7:30"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"explicit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"verified"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Completion is a formula, not an opinion:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;complete = connected ∧ date.verified ∧ time.verified ∧ partySize.verified ∧ confirmed.verified ∧ constraints hold
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the caller says "great, that's confirmed", it is recorded and ignored. Only the callee's "you're booked" fills &lt;code&gt;confirmed&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The evidence rules
&lt;/h2&gt;

&lt;p&gt;This part is deterministic parsers and rules. No LLM.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"7 is full, but 7:30 works" → no evidence for 7:00 (negated clause)&lt;/li&gt;
&lt;li&gt;"the 13th, sorry, the 14th" → the 14th survives, with a supersede trail&lt;/li&gt;
&lt;li&gt;"probably fine" → neither agreement nor confirmation&lt;/li&gt;
&lt;li&gt;price changes after "you're booked" → the confirmation goes stale and must be re-obtained&lt;/li&gt;
&lt;li&gt;"OK, I'll look elsewhere" → not an acceptance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is a command that generates ten thousand hostile clerks to break this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;oathra &lt;span class="nb"&gt;eval&lt;/span&gt; &lt;span class="nt"&gt;--adversarial&lt;/span&gt; 10000
&lt;span class="c"&gt;# never-confirm / wrong-restate / negate-then-offer / silent-hangup / caller-echo-trap&lt;/span&gt;
&lt;span class="c"&gt;# False Completion: 0 / 10000 adversarial runs&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fuzzer caught real bugs. The first one: "¥21,100" was split at the comma, became "¥100", and the agent happily accepted. More recently I let GPT-4o mini and Gemini &lt;em&gt;play the clerk&lt;/em&gt; (&lt;code&gt;oathra eval --callee openai&lt;/code&gt;), which found refusals that quoted a number being read as offers, and product names being parsed as serial numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real phone calls
&lt;/h2&gt;

&lt;p&gt;The same runtime drives real calls.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;oathra setup phone      &lt;span class="c"&gt;# pick a voice engine and a carrier&lt;/span&gt;
oathra phone doctor     &lt;span class="c"&gt;# which layer is broken&lt;/span&gt;
oathra call &lt;span class="nt"&gt;--to&lt;/span&gt; +81…   &lt;span class="c"&gt;# dial&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Carriers (Twilio direct, or Plivo / your own SIP through a LiveKit gateway) and voice engines (GPT-Live, OpenAI Realtime, Deepgram + LLM + TTS) are chosen independently.&lt;/p&gt;

&lt;p&gt;Calls to my own phone:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Stack&lt;/th&gt;
&lt;th&gt;What happened&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Deepgram + GPT-4o-mini + TTS&lt;/td&gt;
&lt;td&gt;11.8 s to first reply. It kept asking "can you hear me?"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Same, streaming TTS&lt;/td&gt;
&lt;td&gt;2.2 s. Talked to voicemail for 2 minutes; noise interrupted it 13 times&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;GPT-Live (full duplex)&lt;/td&gt;
&lt;td&gt;~0.4 s. A 6 min 25 s chat with no errors&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Call three was when it felt usable. GPT-Live is $0.05/min for the session; Twilio to a Japanese mobile is ¥28.78/min.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it honestly stands
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Twilio direct and GPT-Live are verified on real calls&lt;/li&gt;
&lt;li&gt;Plivo and custom SIP are implemented from provider docs; PSTN not yet verified&lt;/li&gt;
&lt;li&gt;MCP server is v0.2&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every number above was measured by me and can be reproduced with the commands in the README. Scenarios are YAML and welcome as PRs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://forifor.github.io/oathra/en/" rel="noopener noreferrer"&gt;https://forifor.github.io/oathra/en/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>typescript</category>
      <category>voice</category>
    </item>
  </channel>
</rss>
