<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Aditi Gupta</title>
    <description>The latest articles on DEV Community by Aditi Gupta (@aditi_gupta_8d81622a592aa).</description>
    <link>https://dev.to/aditi_gupta_8d81622a592aa</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4030300%2F33933e04-a6d7-4476-a1b5-adccbc9082c1.jpg</url>
      <title>DEV Community: Aditi Gupta</title>
      <link>https://dev.to/aditi_gupta_8d81622a592aa</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aditi_gupta_8d81622a592aa"/>
    <language>en</language>
    <item>
      <title>Claude Opus 5.5 Explained: Coding, Cost, and What Developers Need to Change</title>
      <dc:creator>Aditi Gupta</dc:creator>
      <pubDate>Wed, 30 Sep 2026 08:22:14 +0000</pubDate>
      <link>https://dev.to/aditi_gupta_8d81622a592aa/claude-opus-55-explained-coding-cost-and-what-developers-need-to-change-5162</link>
      <guid>https://dev.to/aditi_gupta_8d81622a592aa/claude-opus-55-explained-coding-cost-and-what-developers-need-to-change-5162</guid>
      <description>&lt;p&gt;An AI coding model can generate a convincing patch in seconds. The more useful question is how long it takes to produce a patch you can confidently merge.&lt;/p&gt;

&lt;p&gt;Did it find the relevant code? Preserve existing behavior? Test the failure case? Notice that its fix breaks under concurrent requests?&lt;/p&gt;

&lt;p&gt;That is the lens through which to evaluate &lt;strong&gt;Claude Opus 5.5&lt;/strong&gt;. Anthropic’s release promises improvements in coding capability, speed, and efficiency. Early community feedback is enthusiastic, but the practical value depends on the complete workflow.&lt;/p&gt;

&lt;p&gt;This guide explains what Opus 5.5 is, how its pricing works, what changes in the API, and how to evaluate it on a realistic coding task.&lt;/p&gt;

&lt;p&gt;Find full Review here :&lt;br&gt;
  &lt;iframe src="https://www.youtube.com/embed/NjxxIPhK4G8" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Claude Opus 5.5?
&lt;/h2&gt;

&lt;p&gt;Claude Opus 5.5 is an AI model from Anthropic, released on September 22, 2026, designed for long-running coding and knowledge work.&lt;/p&gt;

&lt;p&gt;It processes text and images and produces text. When connected to tools through an application, it can inspect files, request code changes, run commands, and use the results to decide what to do next.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Specification&lt;/th&gt;
&lt;th&gt;Opus 5.5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude API model ID&lt;/td&gt;
&lt;td&gt;&lt;code&gt;claude-opus-5-5&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context window&lt;/td&gt;
&lt;td&gt;1 million tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard maximum output&lt;/td&gt;
&lt;td&gt;128,000 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;Text and images&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;Text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thinking&lt;/td&gt;
&lt;td&gt;Adaptive, always enabled&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Default effort&lt;/td&gt;
&lt;td&gt;&lt;code&gt;medium&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Knowledge cutoff&lt;/td&gt;
&lt;td&gt;June 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are the specifications in Anthropic’s &lt;a href="https://platform.claude.com/docs/en/models/opus-5-5/overview" rel="noopener noreferrer"&gt;official model overview&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The model and the coding application play different roles. Opus generates responses and tool requests. The application provides repository access, executes permitted actions, and maintains the conversation.&lt;/p&gt;

&lt;p&gt;Calling the API alone does not give Claude access to your filesystem or turn it into an autonomous coding agent.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does “agentic coding” mean in practice?
&lt;/h2&gt;

&lt;p&gt;Consider a webhook handler that occasionally processes the same payment event twice.&lt;/p&gt;

&lt;p&gt;A single code-generation request might produce a replacement function. An agentic workflow can investigate the surrounding system:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Read the webhook handler and database schema
    ↓
Inspect queue behavior and existing tests
    ↓
Reproduce duplicate processing
    ↓
Implement a fix
    ↓
Run sequential and concurrent delivery tests
    ↓
Inspect failures and revise the patch
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The challenge is broader than writing syntax. The agent must discover whether duplicate delivery comes from retries, simultaneous requests, or a failure between the side effect and the acknowledgement.&lt;/p&gt;

&lt;p&gt;It also needs to understand what the system means by “processed.”&lt;/p&gt;

&lt;p&gt;That makes this a useful example for evaluating a coding model: a plausible local edit can still be an incorrect system-level fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  What has improved over Opus 5?
&lt;/h2&gt;

&lt;p&gt;Anthropic reports that Opus 5.5 generates output more than 30% faster than Opus 5 and costs approximately 40% less on typical workloads at default settings. The cost estimate combines lower token prices with reduced token consumption.&lt;/p&gt;

&lt;p&gt;Its published coding results include:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Evaluation&lt;/th&gt;
&lt;th&gt;Opus 5&lt;/th&gt;
&lt;th&gt;Opus 5.5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-Bench 4.0&lt;/td&gt;
&lt;td&gt;52.3%&lt;/td&gt;
&lt;td&gt;66.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FrontierCode v1.1, Main&lt;/td&gt;
&lt;td&gt;48.0%&lt;/td&gt;
&lt;td&gt;54.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CursorBench 4.0&lt;/td&gt;
&lt;td&gt;46.6%&lt;/td&gt;
&lt;td&gt;57.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are Anthropic-reported results under specified configurations. Most listed Opus 5.5 results use maximum effort; Terminal-Bench uses &lt;code&gt;xhigh&lt;/code&gt;. Anthropic also discloses fallback-model use when safeguards intervened in certain evaluations. &lt;a href="https://www.anthropic.com/claude-opus-5-5" rel="noopener noreferrer"&gt;Source: launch announcement&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The results justify testing the model. They do not predict how reliably it will fix your webhook handler.&lt;/p&gt;

&lt;p&gt;Also distinguish &lt;strong&gt;output speed&lt;/strong&gt; from &lt;strong&gt;task completion time&lt;/strong&gt;. Faster text generation helps, but a coding session includes repository searches, commands, tests, and repair attempts.&lt;/p&gt;

&lt;h2&gt;
  
  
  What developers are noticing
&lt;/h2&gt;

&lt;p&gt;Early discussions emphasize faster responses, lower subscription usage, and clearer explanations. Some developers also report stronger bug fixes and more effective delegation to subagents. &lt;a href="https://www.reddit.com/r/ClaudeCode/comments/1wnt4d3/" rel="noopener noreferrer"&gt;Community discussion&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;These observations suggest useful evaluation questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does the model need fewer correction prompts?&lt;/li&gt;
&lt;li&gt;Does it explain the actual behavioral change?&lt;/li&gt;
&lt;li&gt;Does it inspect relevant dependencies before editing?&lt;/li&gt;
&lt;li&gt;Are its final reports easier to verify?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Treat reports such as “this used 4% of my allowance instead of 10%” as individual experiences. Subscription allowances are not a direct measurement of API spending.&lt;/p&gt;

&lt;p&gt;Creative demos provide another perspective. An interactive 3D project reportedly consumed about $1,874 in tokens, while comments identified existing assets and prebuilt systems used in the scene. Such examples show what a combined workflow can produce, but do not isolate the model’s contribution or establish typical costs. &lt;a href="https://www.reddit.com/r/TopologyAI/comments/1wp6clb/this_is_what_1874_of_opus_55_tokens_looks_like_in/" rel="noopener noreferrer"&gt;Project discussion&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pricing: why “half the token price” is not half the task cost
&lt;/h2&gt;

&lt;p&gt;Standard API pricing is:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Token category, per million tokens&lt;/th&gt;
&lt;th&gt;Opus 5.5&lt;/th&gt;
&lt;th&gt;Sonnet 5.5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Uncached input&lt;/td&gt;
&lt;td&gt;$4&lt;/td&gt;
&lt;td&gt;$2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;td&gt;$10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache reads&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Five-minute cache writes&lt;/td&gt;
&lt;td&gt;$5&lt;/td&gt;
&lt;td&gt;$2.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One-hour cache writes&lt;/td&gt;
&lt;td&gt;$8&lt;/td&gt;
&lt;td&gt;$4&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Opus 5.5 fast mode has premium input/output rates of $8/$40 per million tokens. Batch processing discounts standard input and output pricing by 50%. &lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;Source: official pricing&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Notice that &lt;strong&gt;cache reads cost the same for both models&lt;/strong&gt;. A session that repeatedly reads cached context will not necessarily become half as expensive when switched to Sonnet.&lt;/p&gt;

&lt;p&gt;For an illustrative Opus request without caching:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;100,000 input tokens × $4 / 1,000,000  = $0.40
 10,000 output tokens × $20 / 1,000,000 = $0.20

Model cost = $0.60
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That excludes tool charges and other pricing modifiers.&lt;/p&gt;

&lt;p&gt;For the webhook task, count the investigation, implementation, testing, and repair attempts together:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cost per accepted task =
    total spending across all evaluation runs
    ÷ number of runs that meet acceptance criteria
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Track human review time separately. A low API bill is less impressive if a developer spends an hour correcting the patch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Opus 5.5 versus Sonnet 5.5
&lt;/h2&gt;

&lt;p&gt;Sonnet 5.5’s published Terminal-Bench 4.0 score is 70.6%, compared with Opus 5.5’s 66.4%. Opus leads on other coding evaluations, including CursorBench 4.0.&lt;/p&gt;

&lt;p&gt;Anthropic positions Sonnet for well-scoped everyday tasks and Opus for complex, open-ended work requiring sustained judgment. It also notes that higher Sonnet effort settings can produce task costs closer to Opus. &lt;a href="https://www.anthropic.com/claude-sonnet-5-5" rel="noopener noreferrer"&gt;Source: Sonnet 5.5 announcement&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;A benchmark score does not cleanly separate “reasoning” from “implementation.” It measures performance on a particular set of tasks under a particular setup.&lt;/p&gt;

&lt;p&gt;For our example, I would test Sonnet on implementing a clearly specified deduplication mechanism. I would test Opus on investigating an ambiguous duplicate-processing incident that crosses the handler, database, and queue.&lt;/p&gt;

&lt;p&gt;Those are evaluation hypotheses, not guaranteed model rankings.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical API example: reviewing the webhook handler
&lt;/h2&gt;

&lt;p&gt;Here is a simplified handler with a concurrency problem:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;handle_webhook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;payments&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;was_processed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;already_processed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;payments&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;apply_credit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mark_processed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;processed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two requests can both pass &lt;code&gt;was_processed()&lt;/code&gt; before either records completion. A process can also crash after applying credit but before marking the event as processed.&lt;/p&gt;

&lt;p&gt;Ask Opus to review these failure modes before requesting a patch.&lt;/p&gt;

&lt;p&gt;Install the SDK:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-U&lt;/span&gt; anthropic
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Set &lt;code&gt;ANTHROPIC_API_KEY&lt;/code&gt; in your environment, then save the handler as &lt;code&gt;webhook.py&lt;/code&gt; and run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;source&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;webhook.py&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;read_text&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;review_request&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
Review this webhook handler for duplicate side effects.

Analyze:
- Sequential delivery of the same event.
- Concurrent delivery of the same event.
- A crash after applying credit but before recording completion.

Identify what depends on the database and payment API.
Do not invent transaction guarantees or available methods.
Propose a minimal design and focused regression tests.
Explain why the design handles each failure case.
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-5-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;4096&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;output_config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;effort&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;medium&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;review_request&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;```
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="n"&gt;endraw&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
python&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;source&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="n"&gt;raw&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
```&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;block&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;block&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;type&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;block&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Stop reason:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stop_reason&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Usage:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;model_dump&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The request follows Anthropic’s documented API pattern. Responses must be read by block type, and &lt;code&gt;max_tokens&lt;/code&gt; includes thinking plus visible output. Inspect the stop reason for truncation. &lt;a href="https://platform.claude.com/docs/en/models/opus-5-5/migration-guide" rel="noopener noreferrer"&gt;Source: migration guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This example requests analysis. It does not execute tests or implement a fix.&lt;/p&gt;

&lt;p&gt;The design review should establish whether credit updates and deduplication can share a database transaction. If the side effect occurs in an external service, determine whether that service supports an idempotency key. A unique event record alone does not resolve every crash-recovery scenario.&lt;/p&gt;

&lt;p&gt;That distinction is a useful test of the model’s engineering judgment.&lt;/p&gt;

&lt;h2&gt;
  
  
  API migration changes to check
&lt;/h2&gt;

&lt;p&gt;Existing integrations may need more than a new model ID.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Change&lt;/th&gt;
&lt;th&gt;What can break&lt;/th&gt;
&lt;th&gt;What to do&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Thinking is always enabled&lt;/td&gt;
&lt;td&gt;Disabled thinking or manual thinking-budget requests&lt;/td&gt;
&lt;td&gt;Omit &lt;code&gt;thinking&lt;/code&gt; or use adaptive thinking; tune effort&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Forced tool selection is unsupported&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;tool_choice&lt;/code&gt; values &lt;code&gt;any&lt;/code&gt; or &lt;code&gt;tool&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Use &lt;code&gt;auto&lt;/code&gt; or &lt;code&gt;none&lt;/code&gt;; handle the possibility of no tool call&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thinking blocks are bound to conversation context&lt;/td&gt;
&lt;td&gt;Replaying blocks after editing earlier instructions or tools&lt;/td&gt;
&lt;td&gt;Preserve blocks and follow documented history-update rules&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Progress updates arrive in thinking blocks&lt;/td&gt;
&lt;td&gt;Interfaces rendering only text may appear silent&lt;/td&gt;
&lt;td&gt;Configure a supported display mode and update rendering&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Older computer-use tool unsupported on Claude API and Google Cloud&lt;/td&gt;
&lt;td&gt;Requests using &lt;code&gt;computer_20251124&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Migrate to the supported computer toolset&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Refusals can return HTTP 200&lt;/td&gt;
&lt;td&gt;Applications treating every HTTP success as task success&lt;/td&gt;
&lt;td&gt;Inspect &lt;code&gt;stop_reason&lt;/code&gt; and refusal details&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;See Anthropic’s &lt;a href="https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5" rel="noopener noreferrer"&gt;API and behavior changes&lt;/a&gt;. When migrating from older models, also review sampling parameters and assistant-prefill restrictions in the migration guide.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build completion checks into the agent
&lt;/h2&gt;

&lt;p&gt;A coding agent can end a turn while work remains.&lt;/p&gt;

&lt;p&gt;Anthropic specifically documents that Opus 5.5 may return a progress report with &lt;code&gt;stop_reason: "end_turn"&lt;/code&gt; during a longer task. It recommends tracking open work, using bounded continuations, and waiting for pending commands or subagents. &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5" rel="noopener noreferrer"&gt;Source: unattended-agent guidance&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;For the webhook fix, define completion before execution:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Task: prevent duplicate webhook side effects.

Inspect:
- Handler and persistence code.
- Payment-service interface.
- Queue retry behavior.
- Existing tests.

Completion requires:
- Sequential duplicate-delivery coverage.
- Concurrent duplicate-delivery coverage.
- A tested recovery path for interrupted processing.
- Existing public response behavior preserved.
- Relevant tests executed and results reported.

Identify blockers explicitly.
Avoid unrelated refactors.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your application should verify test outcomes and pending work before accepting completion. Keep runtime, spending, and continuation limits as well.&lt;/p&gt;

&lt;p&gt;For a small fix, one agent may be sufficient. If you delegate test implementation to a second agent, provide an explicit contract, isolate file changes, and validate the integrated result. Include coordination and repair costs when comparing that approach with a single-agent run.&lt;/p&gt;

&lt;h2&gt;
  
  
  A reproducible evaluation you can run
&lt;/h2&gt;

&lt;p&gt;The following is a proposed experiment, not a report of measured results.&lt;/p&gt;

&lt;p&gt;Use one small repository with a webhook handler, a persistent test database, and a controllable fake payment service. Prepare three task variants:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Agent receives&lt;/th&gt;
&lt;th&gt;Acceptance criteria&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sequential duplicates&lt;/td&gt;
&lt;td&gt;Reproduction showing the same event delivered twice&lt;/td&gt;
&lt;td&gt;One credit applied; retry receives the expected response&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Concurrent duplicates&lt;/td&gt;
&lt;td&gt;A test triggering overlapping deliveries&lt;/td&gt;
&lt;td&gt;One credit applied under forced overlap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interrupted processing&lt;/td&gt;
&lt;td&gt;A fault injected after the external side effect&lt;/td&gt;
&lt;td&gt;Retrying safely completes recovery without a second credit&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Keep acceptance tests outside the agent’s editable workspace. For the concurrency test, use synchronization to force overlap; merely launching two requests may fail to exercise the race.&lt;/p&gt;

&lt;p&gt;Compare Opus 5.5 at medium effort with Sonnet 5.5 at medium effort. Repeat each task three times per configuration from a clean checkout. That gives 18 runs—a small exploratory evaluation, not enough to establish a universal ranking.&lt;/p&gt;

&lt;p&gt;Hold these conditions constant:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Repository state, instructions, tools, and permissions.&lt;/li&gt;
&lt;li&gt;Dependencies and test environment.&lt;/li&gt;
&lt;li&gt;Per-run time and spending limits.&lt;/li&gt;
&lt;li&gt;Acceptance criteria and review rubric.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Record the results without filling gaps with estimates:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Configuration&lt;/th&gt;
&lt;th&gt;Accepted runs&lt;/th&gt;
&lt;th&gt;Total cost&lt;/th&gt;
&lt;th&gt;Median completion time&lt;/th&gt;
&lt;th&gt;Median review time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Opus 5.5, medium&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonnet 5.5, medium&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Inspect rejected patches as carefully as successful ones. Did the agent miss the race, invent a database capability, weaken a test, or stop before validation?&lt;/p&gt;

&lt;p&gt;If you publish the experiment, include the repository commit, prompts, execution dates, SDK version, effort settings, and failed runs. That gives other developers enough context to challenge or reproduce your findings.&lt;/p&gt;

&lt;p&gt;Opus 5.5’s reported gains make it worth evaluating. The strongest evidence for adopting it will be a change in your own workflow: more accepted patches, fewer repair rounds, or less time spent verifying the result.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>claude</category>
    </item>
    <item>
      <title>Jev by TypeSafe AI: The Complete Developer Guide</title>
      <dc:creator>Aditi Gupta</dc:creator>
      <pubDate>Thu, 24 Sep 2026 17:32:23 +0000</pubDate>
      <link>https://dev.to/aditi_gupta_8d81622a592aa/jev-by-typesafe-ai-the-complete-developer-guide-51a9</link>
      <guid>https://dev.to/aditi_gupta_8d81622a592aa/jev-by-typesafe-ai-the-complete-developer-guide-51a9</guid>
      <description>&lt;p&gt;Jev is a model from TypeSafe AI that returns typed decisions with probabilities instead of text. You send data plus questions with fixed answer options, and get back one answer per question in roughly 70-500 ms, at $0.042 per million input tokens with output free. It's great for routing, classification, scoring, guardrails, and LLM-as-judge work. It can't write, can't do math, has no vision, and isn't open weights. Treat it as a smart &lt;code&gt;if&lt;/code&gt; statement, test it on your own data, and ignore the hype posts.&lt;/p&gt;

&lt;p&gt;In the past ten days, Jev has taken over r/LocalLLaMA, r/LLMDevs, r/AI_Agents, and half of dev Twitter. Some call it the biggest thing since ChatGPT. Others call it "just a classifier" with a $40M marketing budget. Both camps are partly right, and the useful answer sits in the details.&lt;/p&gt;

&lt;p&gt;This guide covers what Jev is, how the API works, real code, pricing, where it breaks, what people are building, and straight answers to the questions that keep coming up on Reddit.&lt;/p&gt;

&lt;p&gt;Find full video here: &lt;br&gt;
  &lt;iframe src="https://www.youtube.com/embed/MwwA8ee-u-k" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  What Jev actually is
&lt;/h2&gt;

&lt;p&gt;Jev is a model that makes decisions and does not write text. You send it some data (the "state") plus a set of typed questions, and it sends back one answer per question with probabilities attached.&lt;/p&gt;

&lt;p&gt;It comes from &lt;a href="https://typesafe.ai" rel="noopener noreferrer"&gt;TypeSafe AI&lt;/a&gt;, a San Francisco lab that came out of stealth on September 15, 2026 with $40M in seed funding led by DCVC. The founder is Diogo Almeida, who worked at OpenAI on RLHF and InstructGPT. Jev is their first public model, and TypeSafe calls it a "System One model."&lt;/p&gt;

&lt;p&gt;The best mental model is a smart &lt;code&gt;if&lt;/code&gt; statement. Normal code branches on things it can compute:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;total&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nf"&gt;applyDiscount&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;order&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That falls apart when the condition needs judgment. Is this ticket angry? Is this email about billing? Which of these 40 buttons continues checkout? Before Jev you had three options: write brittle rules, train a custom classifier with labeled data, or ask an LLM for JSON and hope. Jev takes questions you define at runtime like an LLM, but returns a constrained probability distribution like a classifier.&lt;/p&gt;

&lt;p&gt;The name has two references. "Jev" comes from William Stanley Jevons, whose paradox says that cheaper resources lead to more consumption. "System One" comes from Daniel Kahneman's fast, intuitive System 1 thinking, as opposed to slow, deliberate System 2 reasoning. TypeSafe's bet is simple: make an AI decision cost a tiny fraction of a cent and you'll put decisions in places you'd never call an LLM.&lt;/p&gt;

&lt;p&gt;If you want the ELI5 version for your non-dev friends: Jev is a very fast multiple-choice test taker. You hand it some information and a few questions with fixed answer options, and it circles one answer per question and tells you how sure it is. It can't write an essay, but it can take thousands of these tests per second for pennies.&lt;/p&gt;
&lt;h2&gt;
  
  
  How it works under the hood
&lt;/h2&gt;

&lt;p&gt;TypeSafe hasn't published much about internals, so be careful with anyone claiming to know the exact architecture. What's public from the &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" rel="noopener noreferrer"&gt;launch post&lt;/a&gt;: a new model architecture, a parallel sampler, and a training method called &lt;strong&gt;Reinforcement Learning for Calibrated Decisions (RLCD)&lt;/strong&gt;. Jev is transformer-based but not an LLM. No weights, parameter count, or architecture paper have been released.&lt;/p&gt;

&lt;p&gt;The key difference from an LLM with structured outputs is how the answer gets produced. An LLM generates &lt;code&gt;{"category": "billing"}&lt;/code&gt; token by token, and a schema only constrains that process. Generation can still fail, stop early, or drift. Jev evaluates every question independently and in parallel against the same state, and outputs a probability distribution over the options you defined. A successful response can never contain a value outside your schema.&lt;/p&gt;

&lt;p&gt;RLCD is the training side. TypeSafe argues that RLHF tuned models for human preference, which produced good chat but also overconfidence and mode dropping. RLCD instead aims for probabilities that match outcomes: across many predictions, answers given 90% probability should be right about 90% of the time. That calibration claim is the most important one, and it's also the least independently tested so far.&lt;/p&gt;

&lt;p&gt;The "it's just Qwen plus logits" theory from r/LocalLLaMA deserves a fair hearing. Yes, you can approximate the behavior by taking an open LLM, doing a single forward pass, and reading the probabilities over a fixed answer set. People have done this for years, and several "Jev in 25 lines of Python" posts show it. Whether Jev does something meaningfully better depends on RLCD and the parallel sampler, and without weights nobody outside TypeSafe can verify that. Test it on your data instead of arguing about it.&lt;/p&gt;
&lt;h2&gt;
  
  
  The API: three primitives, one endpoint
&lt;/h2&gt;

&lt;p&gt;Everything goes through one endpoint:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST https://api.typesafe.ai/v1/systemone
Authorization: Bearer &amp;lt;API_KEY&amp;gt;
Content-Type: application/json
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The body carries &lt;code&gt;model&lt;/code&gt;, &lt;code&gt;state&lt;/code&gt;, and a map of &lt;code&gt;questions&lt;/code&gt;. There are three question types:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;What it asks&lt;/th&gt;
&lt;th&gt;What comes back&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;noul&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Is this true?&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;noul&lt;/code&gt;: probability from 0 to 1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;choice&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Which of these options?&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;choice&lt;/code&gt;, &lt;code&gt;probabilities&lt;/code&gt;, &lt;code&gt;confidence&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;score&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Where on this ordered scale?&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;score&lt;/code&gt;, &lt;code&gt;legend&lt;/code&gt;, &lt;code&gt;probabilities&lt;/code&gt;, &lt;code&gt;confidence&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;"Noul" is TypeSafe's name for a yes/no question. The Vercel AI SDK calls the same thing &lt;code&gt;boolean&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Here is a raw request for support ticket triage:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://api.typesafe.ai/v1/systemone &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$TYPESAFE_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "jev-latest",
    "state": {
      "ticket": "Our webhook stopped firing after the update. We are losing orders. Please fix today."
    },
    "questions": {
      "team": {
        "type": "choice",
        "instructions": "Which team should handle `ticket`?",
        "criteria": {
          "billing": "Charges, invoices, refunds",
          "engineering": "Bugs, outages, integrations",
          "sales": "Pricing, upgrades, new accounts",
          "other": "Anything else"
        }
      },
      "severity": {
        "type": "score",
        "instructions": "How severe is the problem in `ticket`?",
        "criteria": [
          "Cosmetic, nothing is blocked",
          "Feature broken but a workaround exists",
          "Blocking, customer is losing money or data"
        ]
      },
      "is_urgent": {
        "type": "noul",
        "instructions": "Does `ticket` ask for action today?"
      }
    }
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The response always has the same shape: an &lt;code&gt;answers&lt;/code&gt; object keyed by your question IDs, the versioned &lt;code&gt;model&lt;/code&gt; that answered, and &lt;code&gt;usage&lt;/code&gt; with token counts. A Choice answer looks like this:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"choice"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"choice"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"engineering"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"probabilities"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"billing"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.02&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"engineering"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.95&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"sales"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"other"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.03&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.91&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;A Score returns the probability-weighted mean of the level indexes. A &lt;code&gt;1.7&lt;/code&gt; on the severity scale above means "mostly blocking, some weight on workaround exists."&lt;/p&gt;

&lt;p&gt;A few things trip people up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The question ID is not sent to the model.&lt;/strong&gt; &lt;code&gt;team&lt;/code&gt; means nothing to Jev, so the full question must live in &lt;code&gt;instructions&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use backticks around field paths&lt;/strong&gt; like &lt;code&gt;`ticket`&lt;/code&gt; or &lt;code&gt;`order.items[0]`&lt;/code&gt; so the model knows which part of the state you mean.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Always add an &lt;code&gt;other&lt;/code&gt; option&lt;/strong&gt; to a Choice when your list might not cover every input. Jev has to pick something.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  Python SDK
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;typesafe-sdk   &lt;span class="c"&gt;# Python 3.10+&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;typesafe_sdk&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Choice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Noul&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;TypeSafeClient&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;TypeSafeClient&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# reads TYPESAFE_API_KEY
&lt;/span&gt;
&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;system_one&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ticket&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ticket_text&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;questions&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;team&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Choice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;instructions&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Which team should handle `ticket`?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;criteria&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;billing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Charges, invoices, refunds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;engineering&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bugs, outages, integrations&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;other&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;is_urgent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Noul&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;instructions&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Does `ticket` ask for action today?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;team&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;answers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;team&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;team&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;send_to_human&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ticket_text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;team&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choice&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;engineering&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;answers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;is_urgent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;noul&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;page_oncall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ticket_text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;There's also an &lt;code&gt;AsyncTypeSafeClient&lt;/code&gt; with configurable retries.&lt;/p&gt;
&lt;h3&gt;
  
  
  JavaScript / TypeScript SDK
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; @typesafe-ai/sdk   &lt;span class="c"&gt;# Node 20+&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;choice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;noul&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;TypeSafeClient&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@typesafe-ai/sdk&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;TypeSafeClient&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;answers&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;systemOne&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;state&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;ticket&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;questions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;team&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;choice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Which team should handle `ticket`?&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;billing&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Charges, invoices, refunds&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;engineering&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Bugs, outages, integrations&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;other&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}),&lt;/span&gt;
    &lt;span class="na"&gt;is_urgent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;noul&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Does `ticket` ask for action today?&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// answers.team.choice is typed as "billing" | "engineering" | "other"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The TypeScript SDK infers answer types from your question definitions, so you get autocomplete and exhaustive &lt;code&gt;switch&lt;/code&gt; checks for free. Keep calls on the server, since the SDK blocks browser use to protect your API key.&lt;/p&gt;
&lt;h3&gt;
  
  
  Other ways to call it
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Vercel AI Gateway:&lt;/strong&gt; available as &lt;code&gt;typesafe-ai/jev&lt;/code&gt; at the same price, no waitlist. AI SDK 7 has &lt;code&gt;experimental_evaluate&lt;/code&gt; for this kind of model. Note that TypeSafe's &lt;code&gt;confidence&lt;/code&gt; lives in &lt;code&gt;result.providerMetadata.typesafe.confidence&lt;/code&gt; on that path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LangChain:&lt;/strong&gt; exposed through &lt;code&gt;TypeSafeClassifier&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloudflare Workers AI:&lt;/strong&gt; users report it listed as &lt;code&gt;typesafe/jev&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  Limits and errors
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Limit&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;State + all questions&lt;/td&gt;
&lt;td&gt;~64,000 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;State + longest single question&lt;/td&gt;
&lt;td&gt;~32,000 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Choice options&lt;/td&gt;
&lt;td&gt;up to 255&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Score levels&lt;/td&gt;
&lt;td&gt;2 to 10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;text only (string, JSON object, or JSON array)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rate limits (early access)&lt;/td&gt;
&lt;td&gt;250,000 tokens/sec, 1,200 requests/min&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Errors: &lt;code&gt;401&lt;/code&gt; bad key, &lt;code&gt;422&lt;/code&gt; validation failure (the response names the field), &lt;code&gt;429&lt;/code&gt; rate limited, &lt;code&gt;529&lt;/code&gt; overloaded. Retry the last two with exponential backoff; the SDKs do this for you.&lt;/p&gt;

&lt;p&gt;The current model is &lt;code&gt;jev-1.13.0&lt;/code&gt;, and &lt;code&gt;jev-latest&lt;/code&gt; points at the stable release. Log the versioned model ID from every response, and pin it once you've tuned thresholds against it.&lt;/p&gt;
&lt;h2&gt;
  
  
  Confidence, fan-out, and the patterns that matter
&lt;/h2&gt;

&lt;p&gt;The probabilities are the real product. Every Choice and Score answer carries a &lt;code&gt;confidence&lt;/code&gt; from 0 to 1, computed from the shape of the distribution. If one option dominates, confidence is high. If probability is spread across options, it's low. An answer can win at 0.84 and still have confidence around 0.6 because a second option holds meaningful weight.&lt;/p&gt;

&lt;p&gt;That gives you a clean three-band pattern:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;team&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;answers&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;team&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.9&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;autoRoute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;team&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choice&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;          &lt;span class="c1"&gt;// act&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;team&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;routeWithReview&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;team&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choice&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;    &lt;span class="c1"&gt;// act, but flag for review&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;sendToHuman&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;                   &lt;span class="c1"&gt;// don't trust it&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Set thresholds by the cost of being wrong. Labeling a dashboard event can run at 0.5. Deleting a file should need 0.9 and probably a human anyway. Your risk tolerance lives in plain numbers in your code, which is far easier to review than a prompt.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Speculative fan-out.&lt;/strong&gt; With LLMs you ask one question, then decide the next. With Jev, questions run in parallel and each extra one only costs its own tokens, so ask everything you might need in one call and let code decide which answers to use. In one TypeSafe cookbook, 13 questions in a single call were 12.2x cheaper and 10x faster than 13 sequential calls. Most of that saving comes from sending a large state once, so it matters most with big inputs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Composite scoring.&lt;/strong&gt; When a judgment depends on several factors, ask one Score per factor and combine them in code:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// normalize each score to 0..1 by dividing by (number of levels - 1)&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;norm&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ans&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;levels&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;ans&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;score&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;levels&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;priority&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="mf"&gt;0.6&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;answers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;severity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="mf"&gt;0.3&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;answers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;frustration&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="mf"&gt;0.1&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;answers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;report_quality&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;If rankings feel off, change a coefficient and rerun. No prompt rewriting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Routing.&lt;/strong&gt; This is where most real savings show up. Put Jev in front of your stack and let it decide whether a request needs a database lookup, a cheap model, an expensive model, or a person. The order-status lookup never touches an LLM, and only the hard cases reach your frontier model. The same idea works as a model router inside an agent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Using coding agents to integrate it.&lt;/strong&gt; Install TypeSafe's agent skill first, because agents trained on LLM APIs tend to ask one question per call and invent request fields.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Claude Code&lt;/span&gt;
claude plugin marketplace add typesafe-ai/skills
claude plugin &lt;span class="nb"&gt;install &lt;/span&gt;typesafe@typesafe-ai

&lt;span class="c"&gt;# Cursor, Codex, others&lt;/span&gt;
npx skills add typesafe-ai/skills &lt;span class="nt"&gt;--skill&lt;/span&gt; typesafe-ai
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h2&gt;
  
  
  Pricing and speed, with the fine print
&lt;/h2&gt;

&lt;p&gt;Jev costs &lt;strong&gt;$0.042 per million input tokens&lt;/strong&gt; (TypeSafe markets it as $42 per billion), and output tokens are free because there's almost nothing to meter. For comparison, TypeSafe quotes typical LLM input prices at $0.20 to $10 per million. A 300-token support ticket costs about a hundredth of a cent. TypeSafe reports 70-500 ms end-to-end, measured from the US West Coast, so add your own network latency if you're elsewhere.&lt;/p&gt;

&lt;p&gt;The famous "193.6x faster, 444.6x cheaper" numbers need context. They come from TypeSafe's own &lt;a href="https://evals.typesafe.ai/" rel="noopener noreferrer"&gt;workflow evals&lt;/a&gt;, where the reference answers were generated by frontier models and the workflows were written by TypeSafe's team. TypeSafe itself says these gains sit at the high end of real use, and that it can't prove the price isn't subsidized.&lt;/p&gt;

&lt;p&gt;Independent checks are more modest but still strong:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An r/accelerate test on multiple-choice benchmarks found Jev roughly Terra-tier on System 1 tasks at about 18x lower cost.&lt;/li&gt;
&lt;li&gt;Bryo AI's CTO found Gemini slightly more accurate for email triage, but 10-20x more expensive.&lt;/li&gt;
&lt;li&gt;One r/LocalLLaMA user got 94.6% with a model trained on their own data versus 68.1% for Jev over the API.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's the realistic picture. Jev is a strong general-purpose default, and a narrow model trained on your data can still beat it.&lt;/p&gt;

&lt;p&gt;For your own bill, the math is simple:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;savings = (share of AI spend on decisions) x (how much cheaper Jev is on those calls)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;If 60% of your spend is classification, routing, scoring, and verification, you might cut the bill by more than half. If it's 10%, you'll save at most 10%. Include retries, human review, and any LLM steps before or after Jev in the calculation.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where Jev breaks
&lt;/h2&gt;

&lt;p&gt;TypeSafe is fairly honest about limits and publishes a "jaggedness" page per model version. These are the failure modes that matter in production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It can't do math.&lt;/strong&gt; Counting, arithmetic, date comparisons, and hex color distances are unreliable. Do those in code and pass in the result. To count items that match a meaning, ask one Noul per item and sum in code:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;questions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Object&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fromEntries&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;items&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;`item_&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;noul&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Is &lt;/span&gt;&lt;span class="se"&gt;\`&lt;/span&gt;&lt;span class="s2"&gt;items[&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;]&lt;/span&gt;&lt;span class="se"&gt;\`&lt;/span&gt;&lt;span class="s2"&gt; a fruit?`&lt;/span&gt;&lt;span class="p"&gt;)])&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;answers&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;systemOne&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;state&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;items&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="nx"&gt;questions&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;items&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;answers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;`item_&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;noul&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;strong&gt;It reads you literally.&lt;/strong&gt; Scoping words, negations, and implied conditions get taken at face value. If you catch yourself explaining what you "really meant" after a wrong answer, that explanation belongs in the instruction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Score is not a measurement.&lt;/strong&gt; A 1.4 doesn't mean "40% of the way between levels." Use scores for thresholds and ranking only. Describe situations in your criteria ("blocking, no workaround exists") instead of degrees ("very severe"), or the model spreads probability across levels.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context rot.&lt;/strong&gt; Accuracy drops as the state fills with irrelevant content. Filter in code first, or ask a Noul per chunk ("is this relevant?") and drop the rest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt injection.&lt;/strong&gt; State is treated as data, but adversarial text can still shift answers. Test with hostile inputs before exposing it publicly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Zero hallucinations" is a narrow claim.&lt;/strong&gt; It means Jev can't return an option you didn't define. It can still pick the wrong one. The 0% figure is a structural guarantee about the output shape, not an accuracy number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No vision, no writing.&lt;/strong&gt; Images, audio, and video need converting to text first. People have forced it to "chat" by picking one character at a time, but it's slow and bad at it. The working rule: with Jev you pick a card from the deck, you don't ask it to name one. If your instinct is "extract X," rephrase it as "here are the candidates for X, which one is it?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It's closed.&lt;/strong&gt; Hosted API only, early access behind a waitlist, no weights and no self-hosting option.&lt;/p&gt;
&lt;h2&gt;
  
  
  What people are building with it
&lt;/h2&gt;

&lt;p&gt;The community moved fast. &lt;a href="https://www.shipwithjev.com/" rel="noopener noreferrer"&gt;shipwithjev.com&lt;/a&gt;, an independent catalog, already lists hundreds of builds. Some highlights with real numbers:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Browser and phone agents.&lt;/strong&gt; Browser Use's jev-ultrafast ran a Zurich to London Google Flights search in 7.1 seconds. A planner LLM sets the goal, and Jev picks the next element to click as a Choice over the page's interactive elements. Droidrun's mobile-jev drove Uber on a real Android phone.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/browser-use" rel="noopener noreferrer"&gt;
        browser-use
      &lt;/a&gt; / &lt;a href="https://github.com/browser-use/jev-ultrafast" rel="noopener noreferrer"&gt;
        jev-ultrafast
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Fastest and cheapest web agent
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/browser-use/jev-ultrafast/docs/banner.svg"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fbrowser-use%2Fjev-ultrafast%2FHEAD%2Fdocs%2Fbanner.svg" alt="Jev Ultrafast · Browser Use × TypeSafe" width="100%"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Jev Ultrafast ⚡&lt;/h1&gt;
&lt;/div&gt;

&lt;div class="markdown-alert markdown-alert-important"&gt;
&lt;p class="markdown-alert-title"&gt;Important&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Browser Use Cloud waitlist is open.&lt;/strong&gt; Get early access to ultrafast browser agents in the cloud
&lt;strong&gt;&lt;a href="https://browser-use.com/ultrafast?utm_source=github&amp;amp;utm_medium=readme&amp;amp;utm_campaign=jev-ultrafast" rel="nofollow noopener noreferrer"&gt;Join the waitlist →&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;A browser agent with a dynamic, indexed action space.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Give it one goal. &lt;a href="https://docs.typesafe.ai/introduction" rel="nofollow noopener noreferrer"&gt;TypeSafe's Jev&lt;/a&gt; picks an operation and an element. A small LLM writes text only when the operation is &lt;code&gt;TYPE_TEXT&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zürich → London on Google Flights in 7.1 seconds.&lt;/strong&gt; One natural-language goal, actual text generation, and loading waits included.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/browser-use/jev-ultrafast/docs/demo.mp4" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fbrowser-use%2Fjev-ultrafast%2FHEAD%2Fdocs%2Fdemo.gif" alt="A real Google Flights search at 1× speed, with generated city names and dynamic operation/target decisions" width="100%"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/browser-use/jev-ultrafast/docs/demo.mp4" rel="noopener noreferrer"&gt;Watch the MP4&lt;/a&gt; · &lt;a href="https://github.com/browser-use/jev-ultrafast/docs/performance.md" rel="noopener noreferrer"&gt;Measurements&lt;/a&gt; · &lt;a href="https://github.com/browser-use/jev-ultrafast/jev_ultrafast/agent.py" rel="noopener noreferrer"&gt;Read the loop&lt;/a&gt;&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;The action space&lt;/h2&gt;
&lt;/div&gt;

&lt;p&gt;Every observation produces a new element table:&lt;/p&gt;

&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;&lt;pre class="notranslate"&gt;&lt;code&gt;[1] button    Change ticket type · Round trip
[2] combobox  Where from?        · San Francisco
[3] combobox  Where to?          · empty
[4] textbox   Departure          · empty
...
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The operations are &lt;code&gt;CLICK&lt;/code&gt;, &lt;code&gt;TYPE_TEXT&lt;/code&gt;, &lt;code&gt;SELECT&lt;/code&gt;, &lt;code&gt;SCROLL_UP&lt;/code&gt;, &lt;code&gt;SCROLL_DOWN&lt;/code&gt;, &lt;code&gt;WAIT&lt;/code&gt;, &lt;code&gt;DONE&lt;/code&gt;, and &lt;code&gt;BLOCKED&lt;/code&gt;. Only supported operations and targets are offered.&lt;/p&gt;
&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;…&lt;/div&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/browser-use/jev-ultrafast" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Coding agent guardrails.&lt;/strong&gt; jev-guard rates each tool call as deny, ask, or allow. Vercel CEO Guillermo Rauch reported Jev up to 18x faster at p95 than GPT Luna for command safety checks, and more accurate. Others use it for Pi's auto mode, instant context compaction (score each old tool call and drop the irrelevant ones), and as a "subconscious" filter that handles small checks before the main model sees anything.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/leepokai" rel="noopener noreferrer"&gt;
        leepokai
      &lt;/a&gt; / &lt;a href="https://github.com/leepokai/jev-guard" rel="noopener noreferrer"&gt;
        jev-guard
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Auto mode for every coding agent, built on Jev: risk-scores every tool call with session context (deny / ask / allow), flags prompt injection in results, checks skills and plugins. Claude Code, Codex, Copilot, Gemini, Cursor, pi, OpenCode, ACP.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div&gt;
  &lt;a rel="noopener noreferrer" href="https://github.com/leepokai/jev-guard/assets/icon.svg"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fleepokai%2Fjev-guard%2FHEAD%2Fassets%2Ficon.svg" width="112" alt="jev-guard"&gt;&lt;/a&gt;
  &lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;jev-guard&lt;/h1&gt;
&lt;/div&gt;
  &lt;p&gt;&lt;strong&gt;A security hook for coding agents, powered by &lt;a href="https://typesafe.ai/" rel="nofollow noopener noreferrer"&gt;Jev&lt;/a&gt;.&lt;/strong&gt;&lt;/p&gt;
  &lt;a rel="noopener noreferrer" href="https://github.com/leepokai/jev-guard/assets/works-with.svg"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fleepokai%2Fjev-guard%2FHEAD%2Fassets%2Fworks-with.svg" alt="Works with Claude Code, Codex, Copilot CLI, Gemini CLI, Cursor, pi, OpenCode, ACP"&gt;&lt;/a&gt;
  &lt;p&gt;
    &lt;a href="https://www.npmjs.com/package/jev-guard" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/df5ea80ad1ce5930982f9efbcb18e90df977b458f27ad203d3f8e7cc677fe223/68747470733a2f2f696d672e736869656c64732e696f2f6e706d2f762f6a65762d67756172643f636f6c6f723d323536334542266c6162656c3d6e706d" alt="npm"&gt;&lt;/a&gt;
    &lt;a rel="noopener noreferrer nofollow" href="https://camo.githubusercontent.com/830324359ddedf6f9a65f156364a18b73ccc8ceb90c89bf0786f7dd3b5422752/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6e6f64652d25453225383925413532302e332d333339393333"&gt;&lt;img src="https://camo.githubusercontent.com/830324359ddedf6f9a65f156364a18b73ccc8ceb90c89bf0786f7dd3b5422752/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6e6f64652d25453225383925413532302e332d333339393333" alt="node 20.3+"&gt;&lt;/a&gt;
    &lt;a rel="noopener noreferrer nofollow" href="https://camo.githubusercontent.com/044cf4fc28bc8a677da8fcbb71c1403e42f33445e21af936c8ab5671a34da9ec/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f646570656e64656e636965732d302d304631373241"&gt;&lt;img src="https://camo.githubusercontent.com/044cf4fc28bc8a677da8fcbb71c1403e42f33445e21af936c8ab5671a34da9ec/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f646570656e64656e636965732d302d304631373241" alt="zero dependencies"&gt;&lt;/a&gt;
    &lt;a href="https://github.com/leepokai/jev-guard/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/7b7cfaa7682c7bbf9210a1e2faae568dd291059c3dcf9b06ca1df8bc19bc386e/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d4d49542d363437343842" alt="MIT"&gt;&lt;/a&gt;
  &lt;/p&gt;
&lt;/div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Auto mode, for every coding agent&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;Claude Code's &lt;a href="https://code.claude.com/docs/en/permission-modes#eliminate-prompts-with-auto-mode" rel="nofollow noopener noreferrer"&gt;auto mode&lt;/a&gt; is described as: &lt;em&gt;"A separate classifier model reviews actions before they run, blocking anything that escalates beyond your request, targets unrecognized infrastructure, or appears driven by hostile content Claude read."&lt;/em&gt; That is exactly the job jev-guard does — as three typed questions to Jev (&lt;code&gt;risk&lt;/code&gt;, &lt;code&gt;user_requested&lt;/code&gt;, &lt;code&gt;from_untrusted&lt;/code&gt;) instead of a proprietary classifier — and it does it for Codex, Copilot, Gemini, Cursor, pi, OpenCode and ACP editors too, with the same policy and the same session memory everywhere. If you want auto mode outside Claude Code, or a second opinion inside it, this is the build.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h3 class="heading-element"&gt;Why Jev: price and speed, with sources&lt;/h3&gt;
&lt;/div&gt;
&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;

&lt;th&gt;Figure&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Price&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;$0.042 per 1M input tokens, $0 output&lt;/strong&gt; — a typical jev-guard call is ~1k tokens, so &lt;strong&gt;≈ $0.00004 per&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;…&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/leepokai/jev-guard" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;&lt;strong&gt;Production usage.&lt;/strong&gt; Metaview, a recruiting platform, says it shipped Jev into every agent on its platform over a weekend, and candidate searches went from minutes to seconds at the same accuracy and lower cost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bulk labeling.&lt;/strong&gt; A teardown of 724 live ads from 37 brands took about 40 seconds and nine cents. Resume screening, email triage, listing classification, and feature engineering for classical ML models all fit the same "label every row" shape.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Search and databases.&lt;/strong&gt; JevQL and pg-jev let you add plain-language filters to Postgres queries, like &lt;code&gt;WHERE jev(meetings, 'could have been an email')&lt;/code&gt;. jevsearch re-ranks keyword hits by intent with no vector database. For memory retrieval in agents, the pattern is similar: pull a larger candidate set with vector search, then ask Jev per item whether it's relevant. Results in r/Rag were mixed, so expect precision gains, not magic.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/kylemclaren" rel="noopener noreferrer"&gt;
        kylemclaren
      &lt;/a&gt; / &lt;a href="https://github.com/kylemclaren/jevql" rel="noopener noreferrer"&gt;
        jevql
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Semantic SQL for Postgres, powered by Jev
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;
  
    
    
    &lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fkylemclaren%2Fjevql%2FHEAD%2F.github%2Freadme%2Fheader-light.png" class="article-body-image-wrapper"&gt;&lt;img alt="jevQL: a SQL editor, a terminal and an agent, all pointed at one Postgres, with jev() judging the rows" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fkylemclaren%2Fjevql%2FHEAD%2F.github%2Freadme%2Fheader-light.png" width="100%"&gt;&lt;/a&gt;
  
&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;jevQL&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;Semantic SQL for &lt;strong&gt;vanilla PostgreSQL&lt;/strong&gt;. One extra family of functions, &lt;code&gt;jev()&lt;/code&gt;, works in any query, from the CLI, from a shared HTTP and MCP node, or from the Go, TypeScript and Python SDKs.&lt;/p&gt;

&lt;div class="highlight highlight-source-sql notranslate position-relative overflow-auto js-code-highlight"&gt;
&lt;pre&gt;&lt;span class="pl-k"&gt;SELECT&lt;/span&gt; name, city, jev_prob(people, &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;could work from home&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;) &lt;span class="pl-k"&gt;AS&lt;/span&gt; p
&lt;span class="pl-k"&gt;FROM&lt;/span&gt; people
&lt;span class="pl-k"&gt;WHERE&lt;/span&gt; jev(people, &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;could work from home&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;)
  &lt;span class="pl-k"&gt;AND&lt;/span&gt; country &lt;span class="pl-k"&gt;=&lt;/span&gt; &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;PT&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;
&lt;span class="pl-k"&gt;ORDER BY&lt;/span&gt; p &lt;span class="pl-k"&gt;DESC&lt;/span&gt;
&lt;span class="pl-k"&gt;LIMIT&lt;/span&gt; &lt;span class="pl-c1"&gt;20&lt;/span&gt;;&lt;/pre&gt;

&lt;/div&gt;
&lt;p&gt;The database only ever sees ordinary SQL. &lt;code&gt;jev_*&lt;/code&gt; calls are evaluated by the
CLI using &lt;a href="https://typesafe.ai" rel="nofollow noopener noreferrer"&gt;TypeSafe&lt;/a&gt;'s System One model (Jev). No
&lt;code&gt;CREATE EXTENSION&lt;/code&gt;, no superuser, no wire-protocol proxy.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Install&lt;/h2&gt;
&lt;/div&gt;

&lt;p&gt;Homebrew (macOS and Linux):&lt;/p&gt;

&lt;div class="highlight highlight-source-shell notranslate position-relative overflow-auto js-code-highlight"&gt;
&lt;pre&gt;brew install kylemclaren/tap/jevql&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Prebuilt binaries for macOS and Linux (arm64 and amd64) are attached to each
&lt;a href="https://github.com/kylemclaren/jevql/releases" rel="noopener noreferrer"&gt;GitHub release&lt;/a&gt; as
&lt;code&gt;jevql_&amp;lt;version&amp;gt;_&amp;lt;os&amp;gt;_&amp;lt;arch&amp;gt;.tar.gz&lt;/code&gt; with a &lt;code&gt;checksums.txt&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;From source, with Go 1.23+ and a C compiler (libpg_query is bundled via
&lt;a href="https://github.com/pganalyze/pg_query_go" rel="noopener noreferrer"&gt;&lt;code&gt;pg_query_go&lt;/code&gt;&lt;/a&gt; and needs CGO):&lt;/p&gt;
&lt;div class="highlight highlight-source-shell notranslate position-relative overflow-auto js-code-highlight"&gt;
&lt;pre&gt;CGO_ENABLED=1&lt;/pre&gt;…
&lt;/div&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/kylemclaren/jevql" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;&lt;strong&gt;LLM-as-judge replacement.&lt;/strong&gt; People use Jev to grade agent traces against rubrics, check whether a cited source supports a claim, and verify RAG answers. One r/Rag benchmark found it tied on accuracy while being 187x cheaper, but it let through 23% of unsupported claims. Good as a first pass, risky as the only check.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real-time apps and games.&lt;/strong&gt; Doom at about 10 decisions per second, Pac-Man, StarCraft, Minecraft, NPC combat decisions, live tone scoring while you type, a debate "BS meter," and Home Assistant control through HA-Jev. Voice agents are a natural fit too, since turn-level decisions like intent routing, escalation, and "did the caller confirm?" need answers in well under a second.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Security and moderation.&lt;/strong&gt; Jailbreak and PII detectors under 300 ms, anti-phishing checks, secret detection, spam and auto-reply detection in the Laravel AI SDK, and moderation against a site's own rules.&lt;/p&gt;

&lt;p&gt;Two warnings. The trading bot threads (including one that reportedly lost $31,680 overnight) prove that fast decisions aren't good decisions. Jev doesn't understand markets any better than an LLM. And most of these are first-week demos, not production case studies. Every cost and latency figure above is self-reported by the builder.&lt;/p&gt;

&lt;h2&gt;
  
  
  Every Reddit question, answered
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Isn't it just a classifier?&lt;/strong&gt;&lt;br&gt;
Mostly yes, and that's the point. Normal classifiers need labeled data and training per task, while Jev takes new labels at runtime. "A zero-shot classifier with calibrated probabilities behind a fast API" is an accurate description. Whether that deserves the hype depends on how many of your LLM calls are secretly classification. The r/AI_Agents hot take nailed it: we've been using LLMs as very expensive if/else statements.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why not just use normal automation or JSON parsing?&lt;/strong&gt;&lt;br&gt;
If a rule can decide correctly, use the rule. It's free and never wrong. Jev is for conditions a rule can't express, like "is this customer about to churn" or "does this message sound like a scam." Parsing tells you what fields exist. Jev tells you what the text means.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jev wasn't first. Same work existed months earlier.&lt;/strong&gt;&lt;br&gt;
Correct. Using LLM logits for classification is old, and several researchers posted earlier open work. TypeSafe never claimed to invent the concept. Its claim is a new training method and a production-grade API, which is more of a product achievement than a research one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;They raised $40M and an open clone matched the API in 3 days. Where's the moat?&lt;/strong&gt;&lt;br&gt;
Copying the API shape is easy, since it's just state in and probabilities out. Copying the quality of the probabilities is the hard part. If RLCD really produces better-calibrated answers than a logit-grabbing hack, that's the moat. If it doesn't, the Christensen-disruption crowd is right and this becomes a commodity fast. Nobody outside TypeSafe has proven which is true yet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How is RLCD even RL if the outputs are differentiable?&lt;/strong&gt;&lt;br&gt;
Fair question, and TypeSafe hasn't published a paper to settle it. Choice, Score, and Noul outputs could be trained with plain supervised loss if you had perfect labels. The RL framing likely comes from rewarding calibration against outcomes rather than matching fixed labels. Until details are public, treat "RL" as TypeSafe's description, not a verified method.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I run it locally or self-host it?&lt;/strong&gt;&lt;br&gt;
Not Jev itself. Open alternatives appeared within days: Laya (an open-weights System One model some users report beating Jev on speed), OpenJev, Nokia's AnyJev (a training-free layer that turns open LLMs into calibrated decision models), and Contrastive-LM's CLM-8B. If data residency or cost at huge scale matters, benchmark these on your task. As for "Laya came first, why is everyone talking about Jev?", distribution and polish usually beat being first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is it a scam?&lt;/strong&gt;&lt;br&gt;
No. The API works, pricing checks out in independent tests, and the founder's track record is real. The "scam" posts mostly point to a single bias example or to aggressive marketing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jev said Taiwan isn't a country / shows sexist or racist tendencies.&lt;/strong&gt;&lt;br&gt;
Jev always picks one of your options, so it gives an answer on sensitive questions where an LLM might hedge or refuse. That makes underlying biases more visible. It doesn't prove it's a Chinese model or uniquely biased, but it does mean you must test hiring, moderation, lending, and anything involving people for fairness before shipping.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why is Reddit flooded with Jev posts?&lt;/strong&gt;&lt;br&gt;
Part genuine excitement, part low-effort SEO spam (the "use Jev for FREE" posts). The r/LocalLLaMA mod request got 1.2K upvotes for a reason. Trust posts with code and numbers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The demos are misleading (Tesla FSD in an hour, car wash question).&lt;/strong&gt;&lt;br&gt;
Partly true. The "FSD" demo fed Jev clean simulated data instead of camera input, which is the hard part of self-driving. Game and driving demos show latency, not real-world readiness. The car wash cost comparison ($6.86 for Jev vs $57.84 for DeepSeek) shows a real gap on that task, but one benchmark isn't your workload.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I use it with OpenAI, ChatGPT, Codex, or Claude?&lt;/strong&gt;&lt;br&gt;
Alongside them, not as a replacement. Jev decides, the LLM writes. Common setups use Jev to route between models, gate tool calls, or filter work before an expensive agent runs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is it good for coding?&lt;/strong&gt;&lt;br&gt;
It can't write code. It's useful around coding agents: command safety checks, "is the task done?" checks, picking which files matter, semantic linting, and routing tasks to the right model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;People used it to generate pixel art and game levels. I thought it can't generate?&lt;/strong&gt;&lt;br&gt;
It can't on its own. Those builds use an LLM or code to propose options, and Jev picks between them fast and repeatedly. The generation comes from the loop, not the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is it good for roleplay?&lt;/strong&gt;&lt;br&gt;
As a helper, maybe. It can decide mood, which character speaks next, or whether a reply stays in character, while a writing model produces the text.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jev controlling a Mac in real time: better than Astra or Fable?&lt;/strong&gt;&lt;br&gt;
Faster, not smarter. Computer control is mostly a string of small choices (which button, which menu), which suits Jev. It needs a text view of the screen like an accessibility tree, since it has no vision. For multi-step planning, reasoning models still lead, so the common pattern is LLM plans, Jev clicks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can Jev detect AI-written posts?&lt;/strong&gt;&lt;br&gt;
It can score text against criteria you write, but it has the same problem as GPTZero and other detectors: there are no reliable signals for AI writing. Don't auto-reject posts on its score alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Will 400x cheaper inference unlock ad-supported AI apps?&lt;/strong&gt;&lt;br&gt;
For decision-heavy features, the cost per user drops low enough that ads could cover it. But most consumer AI value is still generated text, which Jev can't produce.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I get access? Is there a free tier?&lt;/strong&gt;&lt;br&gt;
Join the waitlist at &lt;a href="https://typesafe.ai" rel="noopener noreferrer"&gt;typesafe.ai&lt;/a&gt; (people report getting in within a day or two), or use the Vercel AI Gateway at the same price with no waitlist. There's no known free tier beyond that, so be suspicious of "free Jev" posts.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where to start
&lt;/h2&gt;

&lt;p&gt;Pick one boring decision your code handles with a fragile regex or a costly LLM call. Replace it with one Noul or one Choice. Run it in shadow mode next to your current logic for a week, log the confidence, compare against real outcomes, and only then let it act on the low-risk path. That's how a smart &lt;code&gt;if&lt;/code&gt; statement should be adopted: one branch at a time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Useful links:&lt;/strong&gt; &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" rel="noopener noreferrer"&gt;Launch post&lt;/a&gt; · &lt;a href="https://docs.typesafe.ai" rel="noopener noreferrer"&gt;Docs&lt;/a&gt; · &lt;a href="https://github.com/typesafe-ai/skills" rel="noopener noreferrer"&gt;Agent skills&lt;/a&gt; · &lt;a href="https://www.shipwithjev.com/" rel="noopener noreferrer"&gt;Community builds&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What are you using Jev for, or what's stopping you? Drop it in the comments.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>OmniRoute AI Explained: Features, Installation, Claude Code Setup</title>
      <dc:creator>Aditi Gupta</dc:creator>
      <pubDate>Wed, 23 Sep 2026 11:03:19 +0000</pubDate>
      <link>https://dev.to/aditi_gupta_8d81622a592aa/omniroute-ai-explained-features-installation-claude-code-setup-1lbj</link>
      <guid>https://dev.to/aditi_gupta_8d81622a592aa/omniroute-ai-explained-features-installation-claude-code-setup-1lbj</guid>
      <description>&lt;p&gt;&lt;strong&gt;OmniRoute is a free, open-source AI gateway that connects applications to multiple AI providers through a shared API endpoint.&lt;/strong&gt; It centralizes provider credentials, model selection, fallback routing, token compression, and usage monitoring.&lt;/p&gt;

&lt;p&gt;You can use it with coding tools such as Claude Code and OpenCode, connect it to your own applications, or run it in front of local models.&lt;/p&gt;

&lt;p&gt;This guide explains what OmniRoute does, whether it is free, how to install it, and how it compares with OpenCode and 9router. It also addresses questions raised in OmniRoute Reddit discussions about latency, free providers, token savings, and safety.&lt;/p&gt;

&lt;p&gt;Find a summary here: &lt;br&gt;
  &lt;iframe src="https://www.youtube.com/embed/JbrM6S5l3Io" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  What does OmniRoute do?
&lt;/h2&gt;

&lt;p&gt;OmniRoute receives AI requests from a client, selects an eligible provider or model, forwards the request, and returns the response. Depending on your configuration, it can apply compression, track usage, and fall back to another connection when the first one fails.&lt;/p&gt;

&lt;p&gt;A typical workflow looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Claude Code, OpenCode, or your application
                   ↓
               OmniRoute
                   ↓
       Selected AI provider or local model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;a href="https://github.com/diegosouzapw/OmniRoute" rel="noopener noreferrer"&gt;OmniRoute GitHub repository&lt;/a&gt; contains the source code, installation instructions, documentation, and release information.&lt;/p&gt;

&lt;p&gt;The main benefit is centralized control. Several applications can use the same gateway while you manage provider connections and routing policies in one place.&lt;/p&gt;

&lt;h2&gt;
  
  
  OmniRoute AI features
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Multiple providers and local models
&lt;/h3&gt;

&lt;p&gt;OmniRoute supports several types of connections, including API keys, supported OAuth integrations, public endpoints, and local inference servers.&lt;/p&gt;

&lt;p&gt;Its provider reference includes local options such as Ollama, LM Studio, and vLLM alongside cloud providers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Provider support does not automatically grant model access.&lt;/strong&gt; Availability still depends on your account, credentials, region, quota, and the provider’s requirements. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/reference/PROVIDER_REFERENCE.md" rel="noopener noreferrer"&gt;Provider reference&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Automatic model routing
&lt;/h3&gt;

&lt;p&gt;OmniRoute provides routing aliases that express what you want to prioritize:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Routing alias&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;auto&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Balanced routing with preference for the last successful provider&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;auto/coding&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Quality-focused selection for coding&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;auto/fast&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Favor lower latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;auto/cheap&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Favor lower token costs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;auto/smart&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Quality-focused routing with greater exploration&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These aliases represent selection policies, rather than specific models. Actual results depend on the eligible models and connections available to your installation.&lt;/p&gt;

&lt;p&gt;Despite its name, &lt;code&gt;auto/offline&lt;/code&gt; prioritizes quota availability; it does not restrict requests to local models. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/routing/AUTO-COMBO.md" rel="noopener noreferrer"&gt;Auto-routing documentation&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Combos and automatic fallback
&lt;/h3&gt;

&lt;p&gt;A &lt;strong&gt;combo&lt;/strong&gt; groups models or connections under a routing policy. You can configure an ordered fallback chain or select strategies based on cost, load, quota, context size, and other factors.&lt;/p&gt;

&lt;p&gt;For example, a coding combo could try your preferred model first and then use a compatible backup when necessary.&lt;/p&gt;

&lt;p&gt;OmniRoute also includes circuit breakers that temporarily stop routing requests to repeatedly failing providers. This improves resilience, but requests can still fail when no eligible destination is available. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/architecture/RESILIENCE_GUIDE.md" rel="noopener noreferrer"&gt;Resilience guide&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Token compression
&lt;/h3&gt;

&lt;p&gt;OmniRoute includes compression engines that reduce selected content before it reaches a model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;RTK&lt;/strong&gt; targets terminal, build, test, and other tool output.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Caveman&lt;/strong&gt; condenses natural-language content.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stacked compression&lt;/strong&gt; combines engines in a configured sequence.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Compression can help with verbose coding sessions, but savings vary with the input and settings. Aggressive compression also needs quality checks because removing information can affect the answer. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/compression/COMPRESSION_ENGINES.md" rel="noopener noreferrer"&gt;Compression documentation&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Dashboard and usage visibility
&lt;/h3&gt;

&lt;p&gt;The dashboard provides controls for providers, combos, analytics, health, and costs. It gives users a central place to inspect their setup and investigate problems. &lt;a href="https://www.omniroute.online/" rel="noopener noreferrer"&gt;OmniRoute dashboard overview&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  APIs beyond chat
&lt;/h3&gt;

&lt;p&gt;OmniRoute’s documented API surface includes chat, embeddings, image generation, audio, transcription, reranking, search, video, music, and OCR, as well as Batch and Files APIs.&lt;/p&gt;

&lt;p&gt;Support depends on the selected provider and model. A unified endpoint does not make every model capable of every task. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/reference/API_REFERENCE.md" rel="noopener noreferrer"&gt;API reference&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  MCP, A2A, memory, and skills
&lt;/h3&gt;

&lt;p&gt;OmniRoute also supports more advanced agent workflows:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;What it adds&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;MCP server&lt;/td&gt;
&lt;td&gt;Structured tools for interacting with gateway functions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A2A interface&lt;/td&gt;
&lt;td&gt;Agent discovery and task interactions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Optional memory&lt;/td&gt;
&lt;td&gt;Retrieval and reuse of stored context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Skills&lt;/td&gt;
&lt;td&gt;Configurable reusable instructions&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Memory and A2A are disabled by default in the reviewed documentation. These features are optional; basic routing does not require them. See the &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/frameworks/MCP-SERVER.md" rel="noopener noreferrer"&gt;MCP&lt;/a&gt;, &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/frameworks/A2A-SERVER.md" rel="noopener noreferrer"&gt;A2A&lt;/a&gt;, and &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/frameworks/MEMORY.md" rel="noopener noreferrer"&gt;memory documentation&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is OmniRoute free?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Yes. OmniRoute’s self-hosted software is free and MIT-licensed. AI model usage and hosting can still cost money.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There are three separate costs to understand:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OmniRoute software&lt;/td&gt;
&lt;td&gt;Free, open-source software&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Upstream AI usage&lt;/td&gt;
&lt;td&gt;Depends on provider pricing, subscriptions, or free allowances&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hosting&lt;/td&gt;
&lt;td&gt;Your computer’s resources or a server bill&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;OmniRoute can route requests to providers offering free access, but it does not create unlimited credits or remove upstream restrictions.&lt;/p&gt;

&lt;p&gt;Its free-tier catalog distinguishes recurring allowances, temporary signup credits, and shared quota pools. Multiple models may share one allowance, so their advertised limits cannot always be added together. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/reference/FREE_TIERS.md" rel="noopener noreferrer"&gt;Free-tier documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;A free setup is possible when your selected providers offer suitable free access and your usage stays within their conditions. It is not a guarantee that every model or workload will be free.&lt;/p&gt;

&lt;h2&gt;
  
  
  OmniRoute install: how to set up OmniRoute
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Install the package
&lt;/h3&gt;

&lt;p&gt;For the npm installation, use a supported Node.js version. The project documents Node 24 LTS as an option.&lt;/p&gt;

&lt;p&gt;Run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; omniroute
omniroute
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;-g&lt;/code&gt; flag installs the command globally. Using only &lt;code&gt;npm install omniroute&lt;/code&gt; installs the package in the current project instead.&lt;/p&gt;

&lt;p&gt;By default, the dashboard opens at:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;http://localhost:20128
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The OpenAI-compatible API base URL is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;http://localhost:20128/v1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Follow the &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/guides/SETUP_GUIDE.md" rel="noopener noreferrer"&gt;setup guide&lt;/a&gt; for platform-specific installation details.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Connect a provider
&lt;/h3&gt;

&lt;p&gt;Open &lt;strong&gt;Providers&lt;/strong&gt; in the dashboard and configure a supported connection using the required credentials or sign-in flow.&lt;/p&gt;

&lt;p&gt;Start with one provider and confirm that it works before adding a large fallback chain.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Create your gateway API key
&lt;/h3&gt;

&lt;p&gt;Open &lt;strong&gt;Endpoints&lt;/strong&gt; and create or copy your OmniRoute API key.&lt;/p&gt;

&lt;p&gt;There are two different credentials involved:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Provider credentials:&lt;/strong&gt; allow OmniRoute to call the upstream service.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OmniRoute API key:&lt;/strong&gt; allows your client to call the gateway.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They serve different purposes and should not be confused. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/architecture/AUTHZ_GUIDE.md" rel="noopener noreferrer"&gt;Authorization guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to use OmniRoute
&lt;/h2&gt;

&lt;p&gt;Once the gateway is running, configure a compatible client with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Base URL: http://localhost:20128/v1
API key:  Your OmniRoute API key
Model:    auto, or an available provider/model identifier
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To test the connection, set &lt;code&gt;OMNIROUTE_API_KEY&lt;/code&gt; in your environment and run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:20128/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$OMNIROUTE_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "auto",
    "messages": [
      {
        "role": "user",
        "content": "Explain what an API gateway does."
      }
    ]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can inspect the model list with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:20128/v1/models &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$OMNIROUTE_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These requests use OmniRoute’s documented API interfaces. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/reference/API_REFERENCE.md" rel="noopener noreferrer"&gt;API reference&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;For readers searching “how to use omni route,” the essential sequence is: &lt;strong&gt;install the gateway, connect a provider, configure your client, and test a request.&lt;/strong&gt; Add routing policies and compression after that basic connection works.&lt;/p&gt;

&lt;h2&gt;
  
  
  OmniRoute Claude Code integration
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;OmniRoute can act as the backend gateway for Claude Code, allowing it to use configured, compatible model connections.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;With OmniRoute running and Claude Code installed, the documented interactive configuration command is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;omniroute configure claude
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Follow the prompts to select the connection and model. OmniRoute also provides process-based launchers for supported coding tools. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/guides/CLI-INTEGRATIONS.md" rel="noopener noreferrer"&gt;CLI integration guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;It helps to distinguish three terms:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code:&lt;/strong&gt; the coding application.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude:&lt;/strong&gt; Anthropic’s model family.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OmniRoute:&lt;/strong&gt; the gateway handling the configured connection.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Using Claude Code through OmniRoute does not necessarily mean a Claude model is answering. That depends on your selected route.&lt;/p&gt;

&lt;p&gt;Likewise, an “OmniRoute Claude” integration does not automatically provide free Claude access. Check which provider serves the model, what access your account includes, and whether the connection method is supported.&lt;/p&gt;

&lt;h2&gt;
  
  
  OmniRoute Docker installation
&lt;/h2&gt;

&lt;p&gt;OmniRoute can also run in Docker. This example keeps the published port accessible only from the host computer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; omniroute &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--restart&lt;/span&gt; unless-stopped &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--stop-timeout&lt;/span&gt; 40 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-p&lt;/span&gt; 127.0.0.1:20128:20128 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; omniroute-data:/app/data &lt;span class="se"&gt;\&lt;/span&gt;
  diegosouzapw/omniroute:latest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The named volume preserves application data across container replacement.&lt;/p&gt;

&lt;p&gt;Memory requirements depend on the workload. Long coding-agent sessions can require substantially more memory than light dashboard or chat use. The documentation includes larger memory settings and Compose deployment options; its single-container quick-run path is intended for users who already run Redis separately. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/guides/DOCKER_GUIDE.md" rel="noopener noreferrer"&gt;Docker guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Also remember that &lt;code&gt;localhost&lt;/code&gt; refers to the current machine or container. A client inside another container needs an address that can reach the OmniRoute service.&lt;/p&gt;

&lt;h2&gt;
  
  
  What are the key differences between OmniRoute and OpenCode?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;OmniRoute is primarily an AI gateway. OpenCode is an AI coding agent. They can be used together.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;OmniRoute&lt;/th&gt;
&lt;th&gt;OpenCode&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Main purpose&lt;/td&gt;
&lt;td&gt;Manage and route model requests&lt;/td&gt;
&lt;td&gt;Help users understand and modify code&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Main interaction&lt;/td&gt;
&lt;td&gt;Gateway API, dashboard, and management commands&lt;/td&gt;
&lt;td&gt;Coding interface in a terminal, desktop app, or IDE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Provider handling&lt;/td&gt;
&lt;td&gt;Centralizes connections for multiple clients&lt;/td&gt;
&lt;td&gt;Connects the coding agent to its chosen provider&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical responsibility&lt;/td&gt;
&lt;td&gt;Routing, fallback, compression, and monitoring&lt;/td&gt;
&lt;td&gt;Working with project files and carrying out coding tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Can it work independently?&lt;/td&gt;
&lt;td&gt;Yes, with compatible clients&lt;/td&gt;
&lt;td&gt;Yes, with directly configured providers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Can they work together?&lt;/td&gt;
&lt;td&gt;Serves as OpenCode’s gateway&lt;/td&gt;
&lt;td&gt;Sends model requests through OmniRoute&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;OpenCode’s documentation describes workflows for explaining code, planning features, making changes, and undoing changes. &lt;a href="https://opencode.ai/docs/" rel="noopener noreferrer"&gt;OpenCode documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;A combined setup looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You → OpenCode → OmniRoute → AI provider
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the supported configuration flow, use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;omniroute configure opencode
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;OmniRoute also documents a separate plugin integration. Plugin compatibility depends on the OpenCode major version, so follow the appropriate instructions for your installation. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/guides/CLI-INTEGRATIONS.md" rel="noopener noreferrer"&gt;OpenCode integration details&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  OmniRoute vs 9router
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;OmniRoute began as a fork of 9router&lt;/strong&gt;, a relationship acknowledged in its repository. Both projects provide AI gateway functionality. &lt;a href="https://github.com/diegosouzapw/OmniRoute#-acknowledgments" rel="noopener noreferrer"&gt;OmniRoute project history&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The current 9router documentation lists fallback routing, multiple accounts, format translation, custom combos, token-saving features, quota tracking, and analytics. Those capabilities should not be presented as exclusive to OmniRoute. &lt;a href="https://github.com/decolua/9router#-key-features" rel="noopener noreferrer"&gt;9router features&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;OmniRoute’s documentation additionally describes its particular auto-routing modes, compression pipelines, MCP server, A2A interface, and memory system.&lt;/p&gt;

&lt;p&gt;For an &lt;strong&gt;OmniRoute vs 9router&lt;/strong&gt; evaluation, compare the exact workflow you need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does your provider connection work reliably?&lt;/li&gt;
&lt;li&gt;Does your coding client support the integration?&lt;/li&gt;
&lt;li&gt;Are the routing controls sufficient?&lt;/li&gt;
&lt;li&gt;How much operational complexity does each setup introduce?&lt;/li&gt;
&lt;li&gt;Which project resolves issues affecting your environment?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Provider counts alone do not establish which gateway will work better for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  OmniRoute Online and cloud OmniRoute: what is the difference?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;OmniRoute Online&lt;/strong&gt; commonly refers to the project website, &lt;a href="https://www.omniroute.online/" rel="noopener noreferrer"&gt;omniroute.online&lt;/a&gt;. Visiting the website is separate from running your own gateway.&lt;/p&gt;

&lt;p&gt;The website currently presents self-hosting alongside Cheaper Inference as a hosted gateway option. Hosted offerings have their own access and billing arrangements.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;cloud OmniRoute&lt;/strong&gt; setup can also mean deploying your own instance on a VPS or server. OmniRoute documents remote contexts and scoped tokens for managing a remote installation from a local command line. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/guides/REMOTE-MODE.md" rel="noopener noreferrer"&gt;Remote-mode guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Use the endpoint issued by your chosen service or your own deployment. Do not assume an address copied from an older tutorial remains a supported public endpoint.&lt;/p&gt;

&lt;h2&gt;
  
  
  OmniRoute Reddit questions: practical answers
&lt;/h2&gt;

&lt;p&gt;The questions below reflect themes in community discussions, including the Reddit examples reviewed for this guide. They are not a statistical ranking of the most frequently asked questions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is OmniRoute safe?
&lt;/h3&gt;

&lt;p&gt;OmniRoute is open source and can run locally, but those facts alone do not establish that every configuration is safe.&lt;/p&gt;

&lt;p&gt;Cloud-routed prompts still reach the selected provider. Data handling also depends on logging, optional memory, upstream services, and gateway access controls.&lt;/p&gt;

&lt;p&gt;The project documents authorization controls and guardrails, including optional credential masking. These are useful protections, not proof that sensitive information can be sent through any provider without review. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/security/GUARDRAILS.md" rel="noopener noreferrer"&gt;Security and guardrails&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can using OmniRoute get my provider account restricted?
&lt;/h3&gt;

&lt;p&gt;The answer depends on the provider and connection method. An integration being technically available does not establish permission to use it under every subscription or account type.&lt;/p&gt;

&lt;p&gt;Check the provider’s current rules for third-party clients, automated access, and credential use. Treat Reddit reports of bans or successful access as individual reports rather than universal outcomes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does OmniRoute actually reduce token usage?
&lt;/h3&gt;

&lt;p&gt;It can reduce tokens in eligible content through compression. The size of the reduction depends on your prompts, tool output, conversation history, and compression mode.&lt;/p&gt;

&lt;p&gt;Routing to a cheaper model can reduce spending without reducing tokens. These are different forms of savings.&lt;/p&gt;

&lt;p&gt;Measure input tokens, output tokens, cost, and task success separately. Claims such as “95% savings” should not be treated as a guaranteed reduction in your entire bill. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/compression/COMPRESSION_ENGINES.md" rel="noopener noreferrer"&gt;Compression modes&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why is OmniRoute slow?
&lt;/h3&gt;

&lt;p&gt;Possible causes include a slow upstream model, exhausted quotas, request queues, repeated retries, long context, or insufficient server resources.&lt;/p&gt;

&lt;p&gt;A useful diagnostic sequence is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Test one explicitly selected model.&lt;/li&gt;
&lt;li&gt;Inspect provider health and request failures.&lt;/li&gt;
&lt;li&gt;Check whether retries or fallback attempts explain the delay.&lt;/li&gt;
&lt;li&gt;Temporarily simplify the combo.&lt;/li&gt;
&lt;li&gt;Compare a short request with the slow workload.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;OmniRoute’s resilience documentation describes queues, cooldowns, and circuit breakers that can affect routing behavior. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/architecture/RESILIENCE_GUIDE.md" rel="noopener noreferrer"&gt;Resilience guide&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  What are the best free providers for heavy coding?
&lt;/h3&gt;

&lt;p&gt;There is no permanent best free provider. Availability, quotas, model quality, and tool-calling support change.&lt;/p&gt;

&lt;p&gt;For heavy coding, evaluate a provider on a small real task before adding it to a fallback chain. Check whether it can handle your context size, call tools correctly, and sustain the required request volume.&lt;/p&gt;

&lt;p&gt;Use OmniRoute’s current free-tier catalog to find candidates, then verify the provider’s own terms and limits. &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/reference/FREE_TIERS.md" rel="noopener noreferrer"&gt;Free-tier catalog&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why does a listed model fail?
&lt;/h3&gt;

&lt;p&gt;A catalog listing does not prove that your account can currently use that model.&lt;/p&gt;

&lt;p&gt;Check the credentials, exact model identifier, quota, region, endpoint compatibility, and whether the provider still offers the model. A model that answers plain-text prompts may also fail a coding-agent request requiring tool calling.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does OmniRoute provide unlimited AI usage?
&lt;/h3&gt;

&lt;p&gt;No. Routing across eligible connections can improve availability, but each provider’s limits still apply.&lt;/p&gt;

&lt;p&gt;If every eligible connection is exhausted or unavailable, the request can still fail.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to uninstall OmniRoute
&lt;/h2&gt;

&lt;p&gt;For a global npm installation, stop the running process and remove the package:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm uninstall &lt;span class="nt"&gt;-g&lt;/span&gt; omniroute
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Removing the package does not necessarily remove stored configuration, credentials, or usage history.&lt;/p&gt;

&lt;p&gt;For the Docker container shown earlier:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker stop omniroute
docker &lt;span class="nb"&gt;rm &lt;/span&gt;omniroute
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These commands preserve the named data volume. Deleting stored data is a separate step and permanently removes the configuration it contains. Follow the instructions for your installation method in the &lt;a href="https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/docs/guides/UNINSTALL.md" rel="noopener noreferrer"&gt;OmniRoute uninstall guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who should use OmniRoute?
&lt;/h2&gt;

&lt;p&gt;OmniRoute is useful for developers who want shared provider management across tools, configurable fallback, usage visibility, or a combination of local and cloud models.&lt;/p&gt;

&lt;p&gt;A direct provider connection may be sufficient for a simple application. A gateway becomes more valuable as the number of clients, credentials, routing requirements, and operational questions grows.&lt;/p&gt;

&lt;p&gt;To evaluate it, start with one working provider and one client. Then add a compatible backup and test whether the routing, monitoring, and compression features improve your actual workflow.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>claude</category>
      <category>github</category>
    </item>
    <item>
      <title>GPT-6 Astra for Developers: When to Use It and When to Save Your Tokens</title>
      <dc:creator>Aditi Gupta</dc:creator>
      <pubDate>Mon, 14 Sep 2026 17:28:56 +0000</pubDate>
      <link>https://dev.to/aditi_gupta_8d81622a592aa/gpt-6-astra-for-developers-when-to-use-it-and-when-to-save-your-tokens-3gn4</link>
      <guid>https://dev.to/aditi_gupta_8d81622a592aa/gpt-6-astra-for-developers-when-to-use-it-and-when-to-save-your-tokens-3gn4</guid>
      <description>&lt;p&gt;I came away from testing GPT-6 Astra with a practical recommendation: use it when the difficult part of a task justifies the extra cost.&lt;/p&gt;

&lt;p&gt;For my video, I gave it three game-building prompts, with one attempt per game and no follow-up fixes. It produced a space shooter, a racing game, and a physics stacking game. All three ran at the Light setting, and together they used only a small portion of my weekly allowance.&lt;/p&gt;

&lt;p&gt;That made me interested in Astra for substantial first drafts and prototypes. It didn't tell me to replace every model in my workflow.&lt;/p&gt;

&lt;p&gt;Here's how I'd decide where to use it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A quick decision guide
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Your task&lt;/th&gt;
&lt;th&gt;Where I'd start&lt;/th&gt;
&lt;th&gt;What would justify Astra&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Small edits, boilerplate, or routine transformations&lt;/td&gt;
&lt;td&gt;A cheaper model&lt;/td&gt;
&lt;td&gt;The cheaper option repeatedly misses an important requirement&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A substantial prototype with several interacting requirements&lt;/td&gt;
&lt;td&gt;Astra at a lower effort setting&lt;/td&gt;
&lt;td&gt;A usable first version saves significant implementation and correction time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A difficult bug or architectural decision&lt;/td&gt;
&lt;td&gt;Your usual model, then Astra if needed&lt;/td&gt;
&lt;td&gt;Better diagnosis or reasoning changes the outcome&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Work across browsers, files, and applications&lt;/td&gt;
&lt;td&gt;Astra with the necessary tools connected&lt;/td&gt;
&lt;td&gt;It can complete a meaningful workflow and leave a result you can verify&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3D scenes and interactive visual concepts&lt;/td&gt;
&lt;td&gt;A focused Astra trial&lt;/td&gt;
&lt;td&gt;You need editable spatial output and can inspect it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repetitive tasks at high volume&lt;/td&gt;
&lt;td&gt;A smaller model or deterministic code&lt;/td&gt;
&lt;td&gt;Measured quality gains outweigh the additional cost&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are my starting recommendations, not results from a controlled comparison across all those tasks.&lt;/p&gt;

&lt;h2&gt;
  
  
  What my test actually showed
&lt;/h2&gt;

&lt;p&gt;The useful finding from my three builds was how little intervention they required. I didn't need to repair the main interactions through several additional prompts.&lt;/p&gt;

&lt;p&gt;For a developer exploring an idea, fewer correction cycles can be valuable. A working first version gives you something concrete to evaluate.&lt;/p&gt;

&lt;p&gt;But my test was limited: three prompts, three first attempts, and no equivalent runs against competing models. It supports trying Astra for prototyping. It doesn't establish long-term maintainability, production readiness, or a universal coding advantage.&lt;/p&gt;

&lt;p&gt;Video walkthrough:   &lt;iframe src="https://www.youtube.com/embed/SLhZOp42tRU" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I'd give Astra a harder assignment
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Problems with several constraints
&lt;/h3&gt;

&lt;p&gt;Astra is positioned for complex reasoning, coding, research, computer use, and document creation. I'd consider it when a task combines several of those capabilities. &lt;a href="https://developers.openai.com/api/docs/models/gpt-6-astra" rel="noopener noreferrer"&gt;OpenAI's model documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;For example, I'd try it on a bug that crosses multiple components, a migration plan with compatibility constraints, or an analysis that requires reconciling conflicting evidence.&lt;/p&gt;

&lt;p&gt;My evaluation would be specific: did it identify the problem, respect the constraints, and produce reasoning I can check?&lt;/p&gt;

&lt;p&gt;I wouldn't choose it just because an assignment contains code or mathematics. The difficulty and the value of getting it right matter more.&lt;/p&gt;

&lt;h3&gt;
  
  
  Work that spans applications
&lt;/h3&gt;

&lt;p&gt;Computer use lets models operate browser and desktop interfaces through a connected environment. That makes tasks such as testing a user flow or completing a sequence in an application possible, provided the necessary tools and access are available. &lt;a href="https://developers.openai.com/api/docs/guides/tools-computer-use" rel="noopener noreferrer"&gt;Computer-use documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I'd consider Astra when the work involves understanding a goal and carrying it through several steps. I'd still want an inspectable result: a saved artifact, completed checks, or a clear account of what changed.&lt;/p&gt;

&lt;h3&gt;
  
  
  3D and spatial prototyping
&lt;/h3&gt;

&lt;p&gt;This is another promising area. OpenAI has published a walkthrough of Astra building editable Blender scenes, inspecting renders, and transferring a scene into Unreal Engine. &lt;a href="https://learn.chatgpt.com/blog/architectural-visualization-with-astra" rel="noopener noreferrer"&gt;Architectural visualization example&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;That gives me a reason to test it on a visual prototype with geometry, materials, and interactions. It doesn't establish that the result is physically accurate or suitable for engineering use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I wouldn't spend the premium
&lt;/h2&gt;

&lt;p&gt;I wouldn't make Astra the default for renaming variables, formatting JSON, summarizing a short document, or producing routine boilerplate.&lt;/p&gt;

&lt;p&gt;I'd also avoid repeatedly asking a premium model to explore an idea while I'm still deciding what I want. I'd clarify the brief first, then give Astra a task with a defined finish line.&lt;/p&gt;

&lt;p&gt;If Sol or another model already handles your regular development work well, I don't see a reason to switch without a comparison on your own repository. Review the quality of the diff, the tests, the unnecessary changes, and the time you spend correcting it.&lt;/p&gt;

&lt;p&gt;For repeated structured work, I'd also consider whether a script would solve the problem more predictably.&lt;/p&gt;

&lt;h2&gt;
  
  
  Token efficiency doesn't automatically mean lower cost
&lt;/h2&gt;

&lt;p&gt;At the published base API rates:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input per 1M tokens&lt;/th&gt;
&lt;th&gt;Cached input per 1M tokens&lt;/th&gt;
&lt;th&gt;Output per 1M tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;$10&lt;/td&gt;
&lt;td&gt;$1&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;$4&lt;/td&gt;
&lt;td&gt;$0.40&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Source: &lt;a href="https://developers.openai.com/api/docs/models/compare" rel="noopener noreferrer"&gt;OpenAI's model comparison&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Here's an illustrative calculation for 100,000 uncached input tokens and 10,000 total billed output tokens:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Astra:&lt;/strong&gt; $1.00 input + $0.50 output = &lt;strong&gt;$1.50&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sol:&lt;/strong&gt; $0.40 input + $0.20 output = &lt;strong&gt;$0.60&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are calculated examples, not measurements from my games. They exclude tool fees and use base rates without caching or speed adjustments.&lt;/p&gt;

&lt;p&gt;Astra can still be economical if it needs fewer attempts or substantially less work to complete the task. OpenAI reports lower estimated API cost per task in several evaluations because Astra used fewer output tokens while achieving stronger results. That result is specific to those evaluations. &lt;a href="https://developers.openai.com/api/docs/guides/latest-model" rel="noopener noreferrer"&gt;Model guidance&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;There is also a hidden part of the output count: reasoning tokens are billed as output tokens even though they aren't shown as the final answer. A concise response can still involve substantial reasoning. &lt;a href="https://developers.openai.com/api/docs/guides/reasoning" rel="noopener noreferrer"&gt;Reasoning-token documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The number I care about is the cost of reaching an acceptable result, including retries and review.&lt;/p&gt;

&lt;h2&gt;
  
  
  Subscription limits need a separate check
&lt;/h2&gt;

&lt;p&gt;Those API calculations are not a conversion formula for your subscription's usage meter.&lt;/p&gt;

&lt;p&gt;In Codex and Work, allowance consumption depends on the model, context, complexity, reasoning, tools, retrieval, and caching. A five-hour allowance window does not promise five hours of continuous execution. &lt;a href="https://learn.chatgpt.com/docs/pricing" rel="noopener noreferrer"&gt;Subscription usage documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This is why I wouldn't use my inexpensive game runs to predict the cost of a long session inside a large repository.&lt;/p&gt;

&lt;p&gt;I would check the allowance before and after a representative task, while accounting for any other sessions running at the same time. Then I'd compare the consumption with how often I need to repeat that work.&lt;/p&gt;

&lt;p&gt;Fast mode deserves a separate decision, too: faster execution can consume credits at a higher rate. &lt;a href="https://learn.chatgpt.com/docs/pricing" rel="noopener noreferrer"&gt;Speed and usage details&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Astra still needs improvement
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Calibrating effort to the assignment.&lt;/strong&gt; OpenAI's guidance notes that Astra can perform broader testing than a small coding change requires. I'd like stronger judgment about when another check will improve confidence and when the task is already complete.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Knowing when to proceed.&lt;/strong&gt; The documentation also describes clarification pauses and sensitivity to conflicting instructions. Those behaviors can interrupt a task even when the user expects it to continue.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Producing cleaner prose by default.&lt;/strong&gt; Astra can favor detailed formatting and recurring phrases. If the output is documentation, a technical explanation, or a report, that can create editing work. &lt;a href="https://developers.openai.com/api/docs/guides/latest-model" rel="noopener noreferrer"&gt;Documented behaviors&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Making consumption easier to understand.&lt;/strong&gt; I'd like clearer explanations of which parts of a run consumed the allowance. That would make it easier to distinguish expensive reasoning from unnecessary repetition.&lt;/p&gt;

&lt;p&gt;My game test did not measure these failure modes. They are documented behaviors and product improvements I'd watch for when evaluating longer-term use.&lt;/p&gt;

&lt;h2&gt;
  
  
  The workflow I'd recommend trying
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Define the deliverable.&lt;/strong&gt; State the result, constraints, relevant files, and acceptance criteria.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Start with lower effort.&lt;/strong&gt; My builds succeeded on Light. Increase effort when the task exposes a need for deeper reasoning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep context relevant.&lt;/strong&gt; Supply enough information to solve the problem without filling the session with unrelated history.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use Astra for a bounded role.&lt;/strong&gt; A difficult plan, implementation, or review can be a sensible assignment. Splitting work across models is worth testing, but handoffs aren't free.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ask for evidence of completion.&lt;/strong&gt; For code, that might mean relevant test results and an explanation of unresolved issues.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compare total effort.&lt;/strong&gt; Include model cost, corrections, review time, and whether you could actually use the result.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I'd use Astra more often if it consistently reduced the time needed to get work into an acceptable state. If it produced roughly the same result as my usual model at a higher cost, I'd keep the usual model.&lt;/p&gt;

&lt;p&gt;That's the comparison I'd encourage other developers to make on a few tasks they understand well.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Gemini 3.8 Flash for Coding: Settings, Prompts, and Debugging Tips</title>
      <dc:creator>Aditi Gupta</dc:creator>
      <pubDate>Wed, 09 Sep 2026 17:17:14 +0000</pubDate>
      <link>https://dev.to/aditi_gupta_8d81622a592aa/gemini-38-flash-for-coding-settings-prompts-and-debugging-tips-5141</link>
      <guid>https://dev.to/aditi_gupta_8d81622a592aa/gemini-38-flash-for-coding-settings-prompts-and-debugging-tips-5141</guid>
      <description>&lt;p&gt;Gemini 3.8 Flash can help you get from an idea to a working prototype quickly. Getting reliable results, though, depends on how you define the task, manage revisions, and test the output.&lt;/p&gt;

&lt;p&gt;This guide covers a practical workflow for using it: where to start, how to write useful prompts, and what to do when a debugging conversation stops making progress.&lt;/p&gt;

&lt;p&gt;The observations come from a session building three small games in Google AI Studio. The prompt templates below are reusable examples based on those lessons.&lt;/p&gt;

&lt;h2&gt;
  
  
  The quick version
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Start with a small feature you can test.&lt;/li&gt;
&lt;li&gt;Try medium thinking first; evaluate whether higher thinking improves your particular task.&lt;/li&gt;
&lt;li&gt;Include expected behavior and constraints in the prompt.&lt;/li&gt;
&lt;li&gt;Describe visual choices explicitly.&lt;/li&gt;
&lt;li&gt;When a fix fails repeatedly, ask for a diagnosis before another patch.&lt;/li&gt;
&lt;li&gt;Check functionality separately from appearance.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The video walkthrough shows the builds and the issues that came up:&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/Eq7FE1F_UU0" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Start with a task you can verify
&lt;/h2&gt;

&lt;p&gt;Gemini 3.8 Flash is available through Google’s developer tools, including Google AI Studio. For current capabilities and availability, see the &lt;a href="https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash" rel="noopener noreferrer"&gt;official model documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;During my session, I could build using free AI Studio access without entering a credit card. Treat that as an account of the session: access limits and API billing are separate things to check before relying on it for ongoing work.&lt;/p&gt;

&lt;p&gt;For an initial project, choose something with a clear finish line:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A form that validates a few inputs.&lt;/li&gt;
&lt;li&gt;A searchable list using sample data.&lt;/li&gt;
&lt;li&gt;A single interactive page.&lt;/li&gt;
&lt;li&gt;One feature in an existing application.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;“Build a project management app” leaves many decisions unresolved. “Build a task list with add, complete, and filter actions” gives you something you can inspect.&lt;/p&gt;

&lt;p&gt;You can add complexity after the basic behavior works.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Use medium thinking as a starting point
&lt;/h2&gt;

&lt;p&gt;Higher thinking sounds like the obvious choice for coding. My debugging session gave me a reason to question that default.&lt;/p&gt;

&lt;p&gt;A racing prototype registered a collision when there was no other car on the road. On high thinking, the model repeatedly attempted a similar unsuccessful fix. After I switched to medium, it changed approach and resolved the issue.&lt;/p&gt;

&lt;p&gt;That is one observation, not proof that medium is universally better. It also doesn’t establish a fixed percentage of token savings.&lt;/p&gt;

&lt;p&gt;My practical recommendation is to start with medium for small, well-defined tasks and inspect the result. Try higher thinking when a task needs more analysis, then compare whether it produces a better outcome.&lt;/p&gt;

&lt;p&gt;Watch for repeated failure as well as usage. A long response doesn’t tell you whether the diagnosis is improving.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Give the model acceptance criteria
&lt;/h2&gt;

&lt;p&gt;A useful coding prompt explains how you will decide whether the result works.&lt;/p&gt;

&lt;p&gt;Here’s an example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Build a task list using HTML, CSS, and JavaScript.

Features:
- Add a task.
- Mark a task complete.
- Filter by all, active, or completed.
- Save tasks in localStorage.

Acceptance criteria:
- Reject empty or whitespace-only tasks.
- Preserve tasks after a page refresh.
- Keep each task's completion state when switching filters.
- Show an empty state when a filter has no matching tasks.

Constraints:
- No external dependencies.
- Support keyboard interaction.
- Make the layout usable on mobile.

Explain how to run it and how to check each acceptance criterion.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This makes several otherwise hidden decisions explicit.&lt;/p&gt;

&lt;p&gt;It also improves follow-up requests. If filtering breaks, you can point to a specific requirement instead of asking the model to “make it work.”&lt;/p&gt;

&lt;p&gt;Keep the first request focused. A small feature with clear checks is easier to evaluate than a large application with loosely defined behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Specify the design instead of asking for “modern”
&lt;/h2&gt;

&lt;p&gt;AI-generated interfaces can look convincing while feeling surprisingly similar.&lt;/p&gt;

&lt;p&gt;My three game prototypes shared visual tendencies despite being different genres. Describing the palette, viewing angle, and mood more precisely helped create variety.&lt;/p&gt;

&lt;p&gt;For a web interface, try instructions like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Visual direction:
- Warm white background and dark gray text.
- Muted blue for primary actions.
- Compact layout with readable spacing.
- Clear borders around inputs.
- Minimal decoration.
- Visible keyboard focus states.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can also describe the intended user experience:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;This is an internal tool used repeatedly throughout the day.
Prioritize scanning, readable tables, and quick access to actions.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That gives the model a reason for the design choices.&lt;/p&gt;

&lt;p&gt;Use an actual screenshot when you have a specific visual target. When you don’t, a few concrete constraints are more useful than several broad adjectives.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. When debugging stalls, ask for evidence
&lt;/h2&gt;

&lt;p&gt;Repeatedly asking “try again” can keep a conversation circling around the same assumption.&lt;/p&gt;

&lt;p&gt;A better follow-up includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The steps that reproduce the issue.&lt;/li&gt;
&lt;li&gt;What you expected.&lt;/li&gt;
&lt;li&gt;What happened instead.&lt;/li&gt;
&lt;li&gt;What the previous fix failed to change.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The previous patch did not resolve the issue.

Steps to reproduce:
1. Add two tasks.
2. Mark the first task complete.
3. Switch to the active filter.
4. Switch back to all tasks.

Expected:
The first task remains complete.

Actual:
Both tasks appear active.

Before editing:
- Trace where completion state is stored and updated.
- Identify the likely cause and the code supporting that diagnosis.
- Explain why the previous patch did not address it.

Then propose the smallest relevant fix.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The purpose is to get a diagnosis you can assess.&lt;/p&gt;

&lt;p&gt;If the model claims a fix works, ask what it actually checked. A suggested test and an executed test are different pieces of evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Keep the current project state clear
&lt;/h2&gt;

&lt;p&gt;As you revise a project, the conversation accumulates obsolete instructions, failed patches, and decisions you have changed.&lt;/p&gt;

&lt;p&gt;When the model starts repeating earlier work or losing track of requirements, summarize the current state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Current goal:
Add filtering to the existing task list.

Already working:
- Creating tasks.
- Toggling completion.
- Saving and loading tasks.

Current issue:
Switching filters changes completion state.

Preserve:
- Existing storage format.
- Existing keyboard interactions.

Relevant files:
- app.js: state and event handlers.
- index.html: filter controls.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A short summary helps make the next request unambiguous.&lt;/p&gt;

&lt;p&gt;If you start a fresh chat, include the relevant current code too. The summary explains the task; the code lets the model investigate the implementation.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Test behavior separately from appearance
&lt;/h2&gt;

&lt;p&gt;The collision bug in my racer was a useful reminder: a polished screen can still contain incorrect logic.&lt;/p&gt;

&lt;p&gt;For a small web application, check a few categories:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Area&lt;/th&gt;
&lt;th&gt;Useful checks&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Main flow&lt;/td&gt;
&lt;td&gt;Can the user complete the intended task?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;What happens with empty, invalid, or unusually long values?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;State&lt;/td&gt;
&lt;td&gt;Does behavior remain correct after filtering, editing, or refreshing?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Errors&lt;/td&gt;
&lt;td&gt;Does the interface handle a failed request or missing data?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interaction&lt;/td&gt;
&lt;td&gt;Can you use the main controls with a keyboard and on a small screen?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Choose checks that fit the feature. You don’t need an elaborate test suite for every prototype, but you do need evidence that its important behavior works.&lt;/p&gt;

&lt;p&gt;If a bug returns after a later change, a focused regression test can help catch that specific failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Flash and Pro fit
&lt;/h2&gt;

&lt;p&gt;For my workflow, Flash is useful when the task is concrete and I can inspect the result quickly. I still prefer Pro for longer planning discussions and explanations.&lt;/p&gt;

&lt;p&gt;That led to a simple personal rule: &lt;strong&gt;plan with Pro, build with Flash.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It isn’t a universal model ranking. Choose based on the next task and the quality of the results you’re getting.&lt;/p&gt;

&lt;p&gt;For a small build, the most useful loop is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Define one behavior.&lt;/li&gt;
&lt;li&gt;Generate an implementation.&lt;/li&gt;
&lt;li&gt;Run it.&lt;/li&gt;
&lt;li&gt;Report a specific failure.&lt;/li&gt;
&lt;li&gt;Verify the fix before expanding the scope.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Gemini 3.8 Flash made the first implementation easy to reach in my session. Clear requirements and careful feedback made that implementation more useful.&lt;/p&gt;

&lt;p&gt;What has helped your AI coding sessions most: better initial prompts, smaller tasks, or more specific debugging feedback?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>gemini</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Archify : Honest Review</title>
      <dc:creator>Aditi Gupta</dc:creator>
      <pubDate>Thu, 03 Sep 2026 07:20:29 +0000</pubDate>
      <link>https://dev.to/aditi_gupta_8d81622a592aa/archify-honest-review-30p8</link>
      <guid>https://dev.to/aditi_gupta_8d81622a592aa/archify-honest-review-30p8</guid>
      <description>&lt;p&gt;Architecture diagrams become unreliable when the code changes but the diagram does not.&lt;/p&gt;

&lt;p&gt;Archify approaches this problem by generating diagrams through your coding agent. The agent analyzes a repository, creates a typed JSON description of the system, and Archify validates and renders it as an interactive HTML/SVG diagram.&lt;/p&gt;

&lt;p&gt;I tested it on an unfamiliar repository to answer four practical questions:&lt;/p&gt;

&lt;p&gt;How accurately does it represent real code?&lt;/p&gt;

&lt;p&gt;Can the diagram be regenerated when the code changes?&lt;/p&gt;

&lt;p&gt;Is its architecture comparison useful during reviews?&lt;/p&gt;

&lt;p&gt;What should you verify before trusting the output?&lt;/p&gt;

&lt;p&gt;This post explains how to install Archify, generate a focused architecture diagram, check its claims against the repository, and avoid wasting model usage on an overly broad scan.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Archify actually does
&lt;/h2&gt;

&lt;p&gt;Archify is an agent skill for Claude Code, Cursor, Codex, and OpenCode.&lt;/p&gt;

&lt;p&gt;The coding agent examines your repository or system description and writes a typed JSON representation of the architecture. Archify validates that representation and compiles it into a self-contained interactive HTML/SVG diagram.&lt;/p&gt;

&lt;p&gt;It currently supports five types of diagrams:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Architecture&lt;/li&gt;
&lt;li&gt;Workflow&lt;/li&gt;
&lt;li&gt;Sequence&lt;/li&gt;
&lt;li&gt;Data flow&lt;/li&gt;
&lt;li&gt;Lifecycle&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This distinction matters: the coding agent interprets the repository, while Archify validates and renders the resulting structure.&lt;/p&gt;

&lt;p&gt;Archify can catch malformed data, invalid references, layout problems, and several kinds of misleading visual routing. It cannot guarantee that the agent understood every architectural detail correctly.&lt;/p&gt;

&lt;p&gt;You still need to review the result.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/BPVCykmq0CA" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;Archify requires Node.js 18 or newer.&lt;/p&gt;

&lt;p&gt;Install the skill globally:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add tt-a1i/archify &lt;span class="nt"&gt;-g&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Then open your project in a supported coding agent and use a scoped prompt such as:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Analyze this repository, then use Archify to create a high-level
runtime architecture diagram.

Show 8–12 core components, the primary request path, external
dependencies, storage, and trust boundaries.

Include source evidence where it is supported. Put secondary
details in cards instead of adding more edges.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The &lt;code&gt;8–12 core components&lt;/code&gt; constraint is useful. Without it, a large repository can produce a technically detailed diagram that is difficult to read.&lt;/p&gt;

&lt;p&gt;For a system that does not exist as code yet, you can also start with plain English:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Use Archify to draw this system:

Browser -&amp;gt; API gateway -&amp;gt; authentication service -&amp;gt; application API
-&amp;gt; Redis cache -&amp;gt; PostgreSQL fallback.

Show the trust boundary around the private services and database.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h2&gt;
  
  
  1. Regenerating a diagram is easier than maintaining one
&lt;/h2&gt;

&lt;p&gt;Every architecture diagram I have created manually was accurate for a limited time.&lt;/p&gt;

&lt;p&gt;Then somebody renamed a service, moved a responsibility, or introduced another queue. The code changed, but the diagram did not.&lt;/p&gt;

&lt;p&gt;An Archify diagram is not automatically synchronized with your repository. You must rerun the agent when the architecture changes.&lt;/p&gt;

&lt;p&gt;The difference is that regeneration is much cheaper than manually redrawing the system.&lt;/p&gt;

&lt;p&gt;Instead of deciding whether an old diagram is still trustworthy, I can regenerate it from the current repository and review the differences.&lt;/p&gt;

&lt;p&gt;That makes architecture documentation feel more like a build artifact and less like a drawing somebody has to remember to maintain.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. Source evidence makes the diagram easier to challenge
&lt;/h2&gt;

&lt;p&gt;In evidence-backed architecture mode, components can reference repository files and line ranges pinned to a specific commit.&lt;/p&gt;

&lt;p&gt;This is the part I found most valuable.&lt;/p&gt;

&lt;p&gt;When I select a component, I can inspect the source evidence behind it instead of accepting a convincing-looking box. If a component is missing, I can search the repository and determine whether the agent overlooked it or whether the code genuinely is not there.&lt;/p&gt;

&lt;p&gt;There is an important limitation, though.&lt;/p&gt;

&lt;p&gt;Validation does not make hallucination impossible. The agent still authors the JSON structure, and repository evidence is optional and subject to supported repository conditions.&lt;/p&gt;

&lt;p&gt;My rule is simple: I do not trust the diagram because it looks polished. I trust it only after I have checked the important nodes and relationships against the code.&lt;/p&gt;

&lt;p&gt;A useful verification pass is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do the main nodes correspond to real folders, modules, or services?&lt;/li&gt;
&lt;li&gt;Do the important edges represent actual calls or data movement?&lt;/li&gt;
&lt;li&gt;Are external dependencies shown?&lt;/li&gt;
&lt;li&gt;Are security and deployment boundaries based on evidence?&lt;/li&gt;
&lt;li&gt;Is anything suspiciously absent?&lt;/li&gt;
&lt;li&gt;Does the primary path match the application’s real entry point?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The diagram accelerates understanding. It does not replace engineering judgment.&lt;/p&gt;
&lt;h2&gt;
  
  
  3. Guided views explain the system without changing it
&lt;/h2&gt;

&lt;p&gt;Archify diagrams can contain guided views that focus the reader on existing nodes and relationships.&lt;/p&gt;

&lt;p&gt;A view might highlight:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;How a request enters the system&lt;/li&gt;
&lt;li&gt;Where authentication happens&lt;/li&gt;
&lt;li&gt;How tools are selected&lt;/li&gt;
&lt;li&gt;Where untrusted execution is isolated&lt;/li&gt;
&lt;li&gt;How the response returns to the caller&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;These views are presentation layers over the authored topology. They do not silently create additional components or connections.&lt;/p&gt;

&lt;p&gt;This makes the output useful for more than private exploration. I could see it working well for onboarding, design discussions, and explaining an unfamiliar service during a review.&lt;/p&gt;
&lt;h2&gt;
  
  
  4. The delta view could be useful in pull requests
&lt;/h2&gt;

&lt;p&gt;Archify can compare two validated architecture snapshots and produce a Before, Delta, and After view.&lt;/p&gt;

&lt;p&gt;From a checkout containing the Archify CLI, the command looks like this:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;node archify/bin/archify.mjs compare architecture &lt;span class="se"&gt;\&lt;/span&gt;
  base.json &lt;span class="se"&gt;\&lt;/span&gt;
  head.json &lt;span class="se"&gt;\&lt;/span&gt;
  architecture-delta.html &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Adjust the CLI path to match your installation.&lt;/p&gt;

&lt;p&gt;The comparison can identify authored additions, removals, changes, movement, and rerouted relationships. It also produces a machine-readable receipt.&lt;/p&gt;

&lt;p&gt;What it does &lt;strong&gt;not&lt;/strong&gt; do is equally important.&lt;/p&gt;

&lt;p&gt;The delta view does not determine:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Runtime impact&lt;/li&gt;
&lt;li&gt;Operational risk&lt;/li&gt;
&lt;li&gt;Whether a change is safe&lt;/li&gt;
&lt;li&gt;Whether a pull request should be merged&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It compares two authored architecture descriptions. It does not replace tests, observability, or human review.&lt;/p&gt;

&lt;p&gt;Even with that limitation, seeing architectural additions and removals beside a pull request is something a static diagram rarely provides.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where I would be careful
&lt;/h2&gt;

&lt;p&gt;Archify itself is open source and free to run locally. The potentially expensive part is the coding agent analyzing your repository.&lt;/p&gt;

&lt;p&gt;On a small application, that may be negligible. On a large monorepo, an unbounded request can consume significant model context and usage.&lt;/p&gt;

&lt;p&gt;I now scope repository runs by:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Naming the application or package I want analyzed&lt;/li&gt;
&lt;li&gt;Limiting the diagram to the most important components&lt;/li&gt;
&lt;li&gt;Asking for one primary runtime path&lt;/li&gt;
&lt;li&gt;Excluding generated files and vendored dependencies&lt;/li&gt;
&lt;li&gt;Requesting additional diagrams only when the first one exposes a real question&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I would not begin with “diagram the entire monorepo.”&lt;/p&gt;

&lt;p&gt;I would start with something narrower:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Map the runtime architecture of packages/payments only.

Include its public entry points, database access, queues, external
providers, and calls to other workspace packages.

Limit the result to 12 primary components.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Smaller diagrams are cheaper to generate and easier to verify.&lt;/p&gt;
&lt;h2&gt;
  
  
  When I would use Archify
&lt;/h2&gt;

&lt;p&gt;I would use it for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Exploring an unfamiliar repository&lt;/li&gt;
&lt;li&gt;Preparing an onboarding walkthrough&lt;/li&gt;
&lt;li&gt;Documenting a request or data path&lt;/li&gt;
&lt;li&gt;Reviewing a proposed architectural change&lt;/li&gt;
&lt;li&gt;Creating a diagram that needs to be regenerated regularly&lt;/li&gt;
&lt;li&gt;Turning a system description into a shareable technical artifact&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I would not rely on it alone for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Security audits&lt;/li&gt;
&lt;li&gt;Production dependency discovery&lt;/li&gt;
&lt;li&gt;Runtime performance analysis&lt;/li&gt;
&lt;li&gt;Merge-safety decisions&lt;/li&gt;
&lt;li&gt;Proving that an architecture is complete&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those require evidence beyond a diagram.&lt;/p&gt;
&lt;h2&gt;
  
  
  Video walkthrough
&lt;/h2&gt;

&lt;p&gt;I tested Archify on both a plain-English system and a real repository. I also show two failure cases and how I checked the output.&lt;/p&gt;


&lt;div class="crayons-card c-embed text-styles text-styles--secondary"&gt;
    &lt;div class="c-embed__content"&gt;
      &lt;div class="c-embed__body flex items-center justify-between"&gt;
        &lt;a href="REPLACE_WITH_YOUTUBE_URL" rel="noopener noreferrer" class="c-link fw-bold flex items-center"&gt;
          &lt;span class="mr-2"&gt;REPLACE_WITH_YOUTUBE_URL&lt;/span&gt;
          

        &lt;/a&gt;
      &lt;/div&gt;
    &lt;/div&gt;
&lt;/div&gt;



&lt;h2&gt;
  
  
  Final verdict
&lt;/h2&gt;

&lt;p&gt;Archify did not eliminate the work of understanding an unfamiliar codebase.&lt;/p&gt;

&lt;p&gt;It changed the order of that work.&lt;/p&gt;

&lt;p&gt;Instead of reading everything before I could form a useful mental model, I started with a structured map. I then used the code to confirm, correct, and deepen that map.&lt;/p&gt;

&lt;p&gt;That was much faster than starting from a blank whiteboard.&lt;/p&gt;

&lt;p&gt;The most useful feature was not that the diagrams looked good. It was that the underlying structure could be validated, regenerated, inspected, and compared.&lt;/p&gt;

&lt;p&gt;Have your architecture diagrams ever survived longer than a month, or have you accepted that they eventually lie?&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/tt-a1i/archify" rel="noopener noreferrer"&gt;Archify repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/tt-a1i/archify/blob/main/README_EN.md" rel="noopener noreferrer"&gt;Archify English README&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/tt-a1i/archify/blob/main/archify/schemas/README.md" rel="noopener noreferrer"&gt;Archify schema documentation&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>claude</category>
      <category>github</category>
    </item>
    <item>
      <title>DeepSeek Harness Explained: What It Is, When to Use It, and When Not To</title>
      <dc:creator>Aditi Gupta</dc:creator>
      <pubDate>Fri, 21 Aug 2026 08:44:36 +0000</pubDate>
      <link>https://dev.to/aditi_gupta_8d81622a592aa/deepseek-harness-explained-what-it-is-when-to-use-it-and-when-not-to-3la3</link>
      <guid>https://dev.to/aditi_gupta_8d81622a592aa/deepseek-harness-explained-what-it-is-when-to-use-it-and-when-not-to-3la3</guid>
      <description>&lt;p&gt;AI models can generate code, explain repositories, and suggest fixes. But a model alone cannot safely inspect your project, execute commands, remember a long-running task, or coordinate multiple tools.&lt;/p&gt;

&lt;p&gt;That surrounding infrastructure is called an &lt;strong&gt;agent harness&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;DeepSeek Harness also called &lt;code&gt;dsh&lt;/code&gt; is DeepSeek’s open-source implementation of that infrastructure. It provides the runtime that connects a model to files, tools, sessions, sandboxes, approval policies, workflows, and a user interface.&lt;/p&gt;

&lt;p&gt;Here is the breakdown: &lt;br&gt;
   &lt;iframe src="https://www.youtube.com/embed/l_4jI_IIDd8?start=8"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;This article explains what DeepSeek Harness is, what makes it interesting, how to try it, and—just as importantly—when it may be the wrong tool.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;DeepSeek Harness is currently a developer preview. Expect breaking changes, unfinished edges, and rapidly evolving APIs.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness is not:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A new DeepSeek model&lt;/li&gt;
&lt;li&gt;A model-training framework&lt;/li&gt;
&lt;li&gt;A replacement for Node.js, Python, or your IDE&lt;/li&gt;
&lt;li&gt;A guarantee that an AI-generated change is correct or safe&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It is an &lt;strong&gt;agent runtime&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A useful mental model is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI agent = model + harness
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model provides reasoning and language capabilities. The harness gives that model a controlled way to interact with the outside world.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Your request
    ↓
DeepSeek Harness
    ├── builds the model context
    ├── exposes approved tools
    ├── manages the workspace
    ├── executes tool calls
    ├── records the session
    ├── applies approval and sandbox policies
    └── returns results through the UI
    ↓
Configured AI model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model decides what it wants to do. The harness decides how that action is represented, executed, recorded, and controlled.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why do agents need a harness?
&lt;/h2&gt;

&lt;p&gt;Suppose you ask a regular chat model:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Find the authentication bug in this repository, fix it, and run the tests.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The model needs more than intelligence to complete that request. It needs a way to:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Discover the repository structure&lt;/li&gt;
&lt;li&gt;Read the relevant files&lt;/li&gt;
&lt;li&gt;Search for related code&lt;/li&gt;
&lt;li&gt;Edit the implementation&lt;/li&gt;
&lt;li&gt;Execute the test suite&lt;/li&gt;
&lt;li&gt;Inspect failures&lt;/li&gt;
&lt;li&gt;Make another change&lt;/li&gt;
&lt;li&gt;Preserve a record of what happened&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A harness supplies these capabilities.&lt;/p&gt;

&lt;p&gt;Without one, you have a model that can tell you what code might work. With one, you have an agent that can potentially inspect and modify a real environment subject to the permissions you give it.&lt;/p&gt;

&lt;p&gt;That last part matters. A harness makes a model more useful, but it also makes the model more capable of causing damage. Workspace boundaries, approvals, sandboxes, and human review remain essential.&lt;/p&gt;

&lt;h2&gt;
  
  
  What makes DeepSeek Harness different?
&lt;/h2&gt;

&lt;p&gt;The main design principle is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Everything is a plugin.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Models, tools, skills, sessions, storage, sandboxes, agent loops, scheduling, and even the UI are provided through plugins.&lt;/p&gt;

&lt;p&gt;At the center is &lt;strong&gt;Cordis&lt;/strong&gt;, a plugin kernel responsible for mounting plugins, resolving their dependencies, and letting them communicate through services and events.&lt;/p&gt;

&lt;p&gt;This has an important practical consequence: capabilities can be replaced or recomposed without maintaining a permanent fork of the harness.&lt;/p&gt;

&lt;p&gt;For example, a developer could theoretically swap:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One model provider for another&lt;/li&gt;
&lt;li&gt;A local shell backend for a remote sandbox&lt;/li&gt;
&lt;li&gt;The default storage implementation for a custom store&lt;/li&gt;
&lt;li&gt;One approval policy for a stricter policy&lt;/li&gt;
&lt;li&gt;The standard agent loop for a specialized workflow&lt;/li&gt;
&lt;li&gt;The browser UI for another client&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This architecture is most valuable when you want to &lt;strong&gt;build or study agent infrastructure&lt;/strong&gt;, not merely chat with an AI model.&lt;/p&gt;

&lt;h2&gt;
  
  
  DeepSeek Harness does not require a DeepSeek model
&lt;/h2&gt;

&lt;p&gt;The name can be misleading.&lt;/p&gt;

&lt;p&gt;DeepSeek Harness includes direct support for configuring DeepSeek, but it can also work with other catalog providers and custom OpenAI-compatible endpoints. The model and the harness are separate layers.&lt;/p&gt;

&lt;p&gt;That means you can evaluate questions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How does the same model behave with different tools?&lt;/li&gt;
&lt;li&gt;How do two models perform inside the same agent environment?&lt;/li&gt;
&lt;li&gt;What happens when the sandbox or approval policy changes?&lt;/li&gt;
&lt;li&gt;Can an internal model endpoint be connected to a reusable agent runtime?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For custom providers, you supply details such as the provider ID, base URL, API protocol, credentials, and model list.&lt;/p&gt;

&lt;p&gt;Be aware that “OpenAI-compatible” does not always mean perfectly compatible. Different gateways may use different roles, token-limit fields, reasoning formats, or image capabilities. DeepSeek Harness exposes compatibility settings for these cases, but connecting an unusual endpoint may require experimentation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Every run is traceable
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness uses an append-only session log.&lt;/p&gt;

&lt;p&gt;The log records the model-visible history of a run, including prompts, messages, tool calls, tool results, context injections, and agent activity. Features such as resuming, forking, searching, replaying, and inspecting a trajectory are built from this event stream.&lt;/p&gt;

&lt;p&gt;This is useful for debugging questions like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why did the agent edit this file?&lt;/li&gt;
&lt;li&gt;Which tool result changed its direction?&lt;/li&gt;
&lt;li&gt;What context did the model receive?&lt;/li&gt;
&lt;li&gt;Where did a multi-step task begin to fail?&lt;/li&gt;
&lt;li&gt;Did the problem come from the model, a tool, or the harness configuration?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Traceability is especially helpful when developing an agent system. Looking only at the final answer often hides the real failure.&lt;/p&gt;

&lt;p&gt;It also has a privacy implication: session logs may contain code, prompts, tool output, file contents, or other sensitive context. Treat stored trajectories as potentially sensitive data.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four runtime modes
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness provides several modes for different kinds of work.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;What it provides&lt;/th&gt;
&lt;th&gt;Best suited for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Standard&lt;/td&gt;
&lt;td&gt;File editing, shell access, search, skills, planning, goals, subagents, and workflows&lt;/td&gt;
&lt;td&gt;General agent-assisted development&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code&lt;/td&gt;
&lt;td&gt;Standard capabilities exposed through a code-based orchestration SDK&lt;/td&gt;
&lt;td&gt;Multi-step tool orchestration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Minimal&lt;/td&gt;
&lt;td&gt;A persistent shell and file editor&lt;/td&gt;
&lt;td&gt;Benchmarking models with minimal harness influence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Creator&lt;/td&gt;
&lt;td&gt;Runtime inspection and plugin experimentation in addition to standard capabilities&lt;/td&gt;
&lt;td&gt;Building presets and extending the harness&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Standard mode
&lt;/h3&gt;

&lt;p&gt;Start here if you want to understand the normal user experience.&lt;/p&gt;

&lt;p&gt;It provides the familiar capabilities expected from a coding agent: reading files, editing code, searching, running commands, planning work, and delegating subtasks.&lt;/p&gt;

&lt;h3&gt;
  
  
  Code mode
&lt;/h3&gt;

&lt;p&gt;Code mode lets the model combine several tool operations in a generated TypeScript program.&lt;/p&gt;

&lt;p&gt;This can reduce the overhead of repeatedly moving between the model and individual tools. It is useful for complex orchestration, but it also increases the importance of execution controls and careful review.&lt;/p&gt;

&lt;h3&gt;
  
  
  Minimal mode
&lt;/h3&gt;

&lt;p&gt;Minimal mode intentionally removes most of the surrounding machinery.&lt;/p&gt;

&lt;p&gt;It is useful when comparing models or studying how much the harness itself influences performance. It is less convenient for everyday development because many higher-level capabilities are absent by design.&lt;/p&gt;

&lt;h3&gt;
  
  
  Creator mode
&lt;/h3&gt;

&lt;p&gt;Creator mode is for developers experimenting with the harness itself.&lt;/p&gt;

&lt;p&gt;Use it to inspect the runtime, test plugins, and compose custom presets. If your goal is simply to fix an application bug, Creator mode is probably unnecessary.&lt;/p&gt;

&lt;h2&gt;
  
  
  When should you use DeepSeek Harness?
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness is a strong candidate in the following situations.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. You are building an agent platform
&lt;/h3&gt;

&lt;p&gt;If your product needs interchangeable tools, model providers, storage systems, sandboxes, or agent loops, the plugin architecture gives you an existing composition model to study or extend.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. You need inspectable agent runs
&lt;/h3&gt;

&lt;p&gt;The session event stream and trajectory view make it easier to reconstruct what an agent saw and did. This is valuable for debugging, evaluations, and failure analysis.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. You want to compare models inside the same environment
&lt;/h3&gt;

&lt;p&gt;Model comparisons are difficult when each model uses a different set of prompts, tools, and execution rules. A configurable harness helps keep more of the environment consistent.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. You are experimenting with custom tools or policies
&lt;/h3&gt;

&lt;p&gt;Because tools and execution policies are extension points, the project is relevant when testing a custom capability, approval flow, sandbox backend, or internal integration.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. You want an open-source base you can inspect
&lt;/h3&gt;

&lt;p&gt;DeepSeek Harness is released under the MIT license. You can examine the implementation, modify it, and build on it within the terms of that license.&lt;/p&gt;

&lt;p&gt;Remember that &lt;strong&gt;MIT-licensed does not mean zero operating cost&lt;/strong&gt;. A configured model provider may charge for API usage, and remote infrastructure can introduce additional costs.&lt;/p&gt;

&lt;h2&gt;
  
  
  When should you not use it?
&lt;/h2&gt;

&lt;p&gt;A new open-source agent system can be exciting, but it is not automatically the right choice.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. You only need a simple model call
&lt;/h3&gt;

&lt;p&gt;If your application sends a prompt and receives an answer, a model SDK may be enough. Adding a full harness introduces plugins, sessions, configuration, storage, and operational complexity you may not need.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. You need a stable production API today
&lt;/h3&gt;

&lt;p&gt;The project is explicitly marked as a developer preview and warns that compatibility-breaking changes will occur.&lt;/p&gt;

&lt;p&gt;That makes it suitable for learning, prototyping, and experimentation. Production adoption requires version pinning, migration planning, testing, and a willingness to follow upstream changes.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. You cannot isolate the working environment
&lt;/h3&gt;

&lt;p&gt;An agent that can edit files and execute commands should not receive unrestricted access to a sensitive machine.&lt;/p&gt;

&lt;p&gt;If you cannot provide a narrow workspace, suitable approval policies, secret isolation, and preferably a disposable environment, do not use it for autonomous changes.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Your workflow requires deterministic results
&lt;/h3&gt;

&lt;p&gt;An agent loop combines model decisions with changing context and tool output. Even with the same request, the exact path may vary.&lt;/p&gt;

&lt;p&gt;Use conventional scripts, tests, and workflow engines when deterministic execution is the primary requirement.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Your team does not need harness customization
&lt;/h3&gt;

&lt;p&gt;If a mature coding assistant already meets your needs, adopting an extensible agent runtime may create maintenance work without delivering meaningful value.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. You plan to trust the output without review
&lt;/h3&gt;

&lt;p&gt;A traceable agent can still make incorrect changes. Logs help explain a decision; they do not make that decision correct.&lt;/p&gt;

&lt;p&gt;Treat generated code like a contribution from an unfamiliar developer: review the diff, run tests, inspect security-sensitive changes, and verify the behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to try DeepSeek Harness
&lt;/h2&gt;

&lt;p&gt;The fastest path is through its local Web UI.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Check Node.js
&lt;/h3&gt;

&lt;p&gt;The repository currently declares support for Node.js &lt;code&gt;^22.19.0&lt;/code&gt; or &lt;code&gt;&amp;gt;=24.0.0&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;node &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because the project is changing quickly, verify the current requirement in the repository before installing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Start the Web UI
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @deepseek-ai/dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;By default, this starts a local server at:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;http://127.0.0.1:3080
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the command from the project directory you want to work with. The process uses its starting directory as the default filesystem location, although you still need to select a workspace in the UI.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Configure a model
&lt;/h3&gt;

&lt;p&gt;Open:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Settings → Models
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can enter a DeepSeek API key, add another supported provider, or configure a custom provider.&lt;/p&gt;

&lt;p&gt;Credentials are stored separately from normal settings, and the UI receives a redacted credential descriptor after saving rather than the literal key.&lt;/p&gt;

&lt;p&gt;You still need to protect the machine and the harness home directory. Never commit credential files to a repository.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Choose a workspace
&lt;/h3&gt;

&lt;p&gt;Select &lt;strong&gt;Choose workspace&lt;/strong&gt;, add the relevant project directory, and select it.&lt;/p&gt;

&lt;p&gt;Use the smallest practical directory. Do not select an entire home folder or a directory containing unrelated secrets and projects.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 5: Begin with a read-only task
&lt;/h3&gt;

&lt;p&gt;A good first request is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Summarize this repository. Identify its main packages, test commands,
and likely entry points. Do not modify any files.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This lets you evaluate how the agent explores the project before allowing it to make changes.&lt;/p&gt;

&lt;p&gt;A reasonable next task is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Find one small, well-contained issue in this repository.
Explain the proposed fix and wait for approval before editing files.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only after you understand the permission flow should you try a full implementation task.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 6: Inspect the trajectory
&lt;/h3&gt;

&lt;p&gt;Do not judge the harness only by the final response.&lt;/p&gt;

&lt;p&gt;Inspect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What context was sent to the model&lt;/li&gt;
&lt;li&gt;Which tools were called&lt;/li&gt;
&lt;li&gt;Which files were accessed&lt;/li&gt;
&lt;li&gt;Whether commands required approval&lt;/li&gt;
&lt;li&gt;How tool output affected later decisions&lt;/li&gt;
&lt;li&gt;Whether the agent repeated unnecessary work&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is where a traceable harness becomes more useful than a simple chat interface.&lt;/p&gt;

&lt;h2&gt;
  
  
  A safer evaluation workflow
&lt;/h2&gt;

&lt;p&gt;For early experiments, use a disposable branch, worktree, container, or test repository.&lt;/p&gt;

&lt;p&gt;A practical workflow looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create an isolated copy of a small project.&lt;/li&gt;
&lt;li&gt;Remove production credentials and customer data.&lt;/li&gt;
&lt;li&gt;Start DeepSeek Harness from that directory.&lt;/li&gt;
&lt;li&gt;Select only that directory as the workspace.&lt;/li&gt;
&lt;li&gt;Use a read-only repository-summary task first.&lt;/li&gt;
&lt;li&gt;Ask for a plan before permitting edits.&lt;/li&gt;
&lt;li&gt;Review every requested command.&lt;/li&gt;
&lt;li&gt;Inspect the resulting diff manually.&lt;/li&gt;
&lt;li&gt;Run the project’s tests yourself.&lt;/li&gt;
&lt;li&gt;Review the trajectory for surprising behavior.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Do not let a successful demo convince you to skip these controls on the next run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building from source
&lt;/h2&gt;

&lt;p&gt;If your goal is to inspect or modify the harness itself, clone the repository and build it with its configured package manager:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/deepseek-ai/deepseek-harness.git
&lt;span class="nb"&gt;cd &lt;/span&gt;deepseek-harness
pnpm &lt;span class="nb"&gt;install
&lt;/span&gt;pnpm run build
pnpm dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use this route when you want to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Read the architecture alongside the implementation&lt;/li&gt;
&lt;li&gt;Develop or modify plugins&lt;/li&gt;
&lt;li&gt;Test changes to the runtime&lt;/li&gt;
&lt;li&gt;Contribute upstream&lt;/li&gt;
&lt;li&gt;Pin your work to a specific commit&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a first evaluation, the &lt;code&gt;npx&lt;/code&gt; command is simpler.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to evaluate before adopting it
&lt;/h2&gt;

&lt;p&gt;A successful installation only proves that the harness starts. Before using it for real work, evaluate the following.&lt;/p&gt;

&lt;h3&gt;
  
  
  Model quality
&lt;/h3&gt;

&lt;p&gt;Does your chosen model use tools reliably? Can it recover from failed commands? Does it stop when it lacks information?&lt;/p&gt;

&lt;h3&gt;
  
  
  Permission behavior
&lt;/h3&gt;

&lt;p&gt;Which actions require approval? Are writes and command execution constrained appropriately?&lt;/p&gt;

&lt;h3&gt;
  
  
  Workspace isolation
&lt;/h3&gt;

&lt;p&gt;Can the agent access files outside the intended project? Are secrets, SSH keys, cloud credentials, and production configuration isolated?&lt;/p&gt;

&lt;h3&gt;
  
  
  Trace quality
&lt;/h3&gt;

&lt;p&gt;Can you reconstruct why a change happened? Does the log contain enough information to debug failures without exposing more sensitive data than necessary?&lt;/p&gt;

&lt;h3&gt;
  
  
  Plugin trust
&lt;/h3&gt;

&lt;p&gt;A plugin can add substantial capabilities. Review its source, dependencies, permissions, maintenance status, and network behavior before installing it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Upgrade cost
&lt;/h3&gt;

&lt;p&gt;Since the project is in preview, test upgrades against pinned configurations and plugins. Do not assume a newer release will preserve every API or behavior.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cost
&lt;/h3&gt;

&lt;p&gt;The harness is open source, but model requests, hosted sandboxes, storage, and other providers may not be free. Measure token usage and infrastructure costs with realistic tasks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final perspective
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness is best understood as infrastructure for constructing and studying agents.&lt;/p&gt;

&lt;p&gt;Its value does not come from making a model magically correct. It comes from giving developers a composable way to connect models with tools, sessions, workspaces, policies, storage, orchestration, and observability.&lt;/p&gt;

&lt;p&gt;Use it when you need that control or want to experiment with agent architecture.&lt;/p&gt;

&lt;p&gt;Avoid it when a simple API call is enough, when stability is more important than extensibility, or when you cannot safely isolate what the agent can access.&lt;/p&gt;

&lt;p&gt;Most importantly, keep the model and the harness conceptually separate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The model decides.
The harness enables, constrains, executes, and records.
The developer remains responsible for the system.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Official resources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/deepseek-ai/deepseek-harness" rel="noopener noreferrer"&gt;DeepSeek Harness repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://deepseek.com/harness/en/" rel="noopener noreferrer"&gt;Official project overview&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://deepseek-harness.github.io/deepseek-harness/en/guide/quickstart" rel="noopener noreferrer"&gt;Web UI quickstart&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://deepseek-harness.github.io/deepseek-harness/en/guide/providers" rel="noopener noreferrer"&gt;Model configuration guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/docs/architecture.md" rel="noopener noreferrer"&gt;Architecture documentation&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>deepseek</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Grokbot Honest Review: Is xAI and Cursor's Computer-Use Agent Worth $200 a Month?</title>
      <dc:creator>Aditi Gupta</dc:creator>
      <pubDate>Mon, 17 Aug 2026 18:03:19 +0000</pubDate>
      <link>https://dev.to/aditi_gupta_8d81622a592aa/grokbot-honest-review-is-xai-and-cursors-computer-use-agent-worth-200-a-month-2924</link>
      <guid>https://dev.to/aditi_gupta_8d81622a592aa/grokbot-honest-review-is-xai-and-cursors-computer-use-agent-worth-200-a-month-2924</guid>
      <description>&lt;p&gt;Grokbot is xAI and Cursor's attempt at an AI agent that doesn't just chat, it uses a computer. Each bot you create gets its own machine in the cloud, with Chrome, a file manager, and its own operating system, and it keeps working after you close your laptop. I spent a week building three bots to see whether that idea holds up in practice.&lt;/p&gt;

&lt;p&gt;Prefer the quick version? I covered the three bots I built, the shared-login trap, and my full verdict in this video:&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/rUfvWFPYGMo"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What Grokbot actually is&lt;/strong&gt;&lt;br&gt;
You download it from x.ai/bot as a desktop app, not a website. You make a bot with a name and a short description of its job, and it spins up its own virtual computer to get started. Because that computer lives in the cloud, the bot runs on its own, 24/7, without touching your machine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why I tested it&lt;/strong&gt;&lt;br&gt;
I wanted to know if computer-use is real yet or still a demo. So I built three bots with very different jobs: one to find startups that raised seed funding recently and pull the founder names and amounts, one to compare flight prices from Delhi to Tokyo, and one to check the price of a phone across five stores. Same setup each time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where it provides value&lt;/strong&gt;&lt;br&gt;
The genuinely new part is that Grokbot skips the API problem. A huge chunk of the internet has no API an AI can plug into, and even when there is one it rarely lets you do everything the website can. Grokbot gets around that by using the site directly, logging in, clicking through pages, and downloading files the way a person would. It ran with my laptop shut, and I could take over the virtual computer from the mobile app. The teach-a-task feature is the standout: you record yourself doing something fiddly once, it watches the recording frame by frame, and turns it into a reusable private skill.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The main catch&lt;/strong&gt;&lt;br&gt;
Two things stood out. First, reliability. It got stuck on simple clicks, brought back half the data, and sometimes just froze on easy tasks. Second, and more important, the virtual computer and its logins seem shared across all your bots. I signed into Google on one bot and the others could suddenly reach every site tied to that login. Convenient, but you should be deliberate about which accounts you hand it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The settings and habits to change first&lt;/strong&gt;&lt;br&gt;
Always require approval before the bot sends an email, publishes anything, books a flight, buys a product, or deletes something. For anything complicated, record a demo with teach-a-task instead of hoping it figures the task out alone. And keep sensitive logins on bots you trust, given the shared-computer behavior.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it compares to Claude Code and Codex&lt;/strong&gt;&lt;br&gt;
Claude Code and Codex can drive a browser too and do more flexible work, but they usually run on your machine, you kick them off yourself, and they stop when your computer is off. Grokbot's lane is different: watching web pages, checking prices on a schedule, moving data between apps, and grinding through repetitive screen work in software you already use. It is not a replacement for a coding agent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who should and shouldn't try it&lt;/strong&gt;&lt;br&gt;
The cheapest plan is $200 a month, and that does include Cursor Ultra. If you already burn through a lot of AI tokens and you're fine switching your subscription to Cursor Ultra, the price makes more sense. If you just want the bot, it's a lot to pay for something this early. Try it on the free week if you have one repetitive task nothing else can automate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My verdict&lt;/strong&gt;&lt;br&gt;
Grokbot is a real step in the right direction. Instead of waiting for every website to build an API, it lets an AI use a computer the way we do, and if that gets reliable enough it could automate almost anything you can do on a screen. Right now the automation isn't consistent, the trial is tiny, and the price is steep, so I wouldn't tell most people to pay for it yet.&lt;/p&gt;

&lt;p&gt;Have you tried Grokbot or another computer-use agent yet? Tell me in the comments. If you like quick, honest AI-tool reviews, follow my YouTube channel for the next one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Links and sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Grokbot (xAI + Cursor): &lt;a href="https://x.ai/bot" rel="noopener noreferrer"&gt;https://x.ai/bot&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Full video review:&lt;a href="https://dev.tourl"&gt; https://youtu.be/rUfvWFPYGMo &lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;More AI tool breakdowns + weekly news: &lt;a href="https://goventure.live" rel="noopener noreferrer"&gt;https://goventure.live&lt;/a&gt; 
&lt;a href="https://youtu.be/rUfvWFPYGMo" rel="noopener noreferrer"&gt;https://youtu.be/rUfvWFPYGMo&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;My testing notes: the three bots, the shared-login behavior, and the reliability misses are all shown in the video above.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Tags: #ai #agents #cursor #aitools&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>reviews</category>
      <category>tooling</category>
    </item>
    <item>
      <title>I Tried Kimi K3 for Free in VS Code - Can It Replace Claude or GPT?</title>
      <dc:creator>Aditi Gupta</dc:creator>
      <pubDate>Wed, 12 Aug 2026 17:35:03 +0000</pubDate>
      <link>https://dev.to/aditi_gupta_8d81622a592aa/i-tried-kimi-k3-for-free-in-vs-code-can-it-replace-claude-or-gpt-6nc</link>
      <guid>https://dev.to/aditi_gupta_8d81622a592aa/i-tried-kimi-k3-for-free-in-vs-code-can-it-replace-claude-or-gpt-6nc</guid>
      <description>&lt;p&gt;My coding assistant subscription has one job: save me more time than it costs.&lt;/p&gt;

&lt;p&gt;So when Moonshot AI released a 2.8-trillion-parameter model with a one-million-token context window—and I found a way to connect it to VS Code for free—I wanted to answer one practical question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is Kimi K3 good enough to replace a paid coding assistant?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;After looking past the headline numbers, comparing the benchmark results, and testing it on a complete browser game, my short answer is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Kimi K3 is good enough to become a serious coding workhorse for many developers. But it is not a universal replacement for every paid model, and the free access comes with an important catch.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Prefer the video walkthrough? I cover the test and both setup methods here:&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/q4Da550BSMY?start=5"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  First, what exactly is Kimi K3?
&lt;/h2&gt;

&lt;p&gt;Kimi K3 is Moonshot AI's newest open-weight, native multimodal model. According to the &lt;a href="https://github.com/MoonshotAI/Kimi-K3" rel="noopener noreferrer"&gt;official release&lt;/a&gt;, it has:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;2.8 trillion total parameters&lt;/li&gt;
&lt;li&gt;104 billion active parameters per token&lt;/li&gt;
&lt;li&gt;A Mixture-of-Experts architecture that activates 16 of 896 experts&lt;/li&gt;
&lt;li&gt;A 1,048,576-token context window&lt;/li&gt;
&lt;li&gt;Native image and video understanding&lt;/li&gt;
&lt;li&gt;Kimi Delta Attention and Attention Residuals&lt;/li&gt;
&lt;li&gt;Built-in reasoning for coding, research, and agentic work&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That “104 billion active” detail matters. Kimi K3 does not use all 2.8 trillion parameters for every token. Its sparse architecture routes each token through a small subset of experts, making inference more efficient than the headline size suggests.&lt;/p&gt;

&lt;h3&gt;
  
  
  Open-weight is not the same as free to run
&lt;/h3&gt;

&lt;p&gt;You will see Kimi K3 called “open source,” but &lt;strong&gt;open-weight&lt;/strong&gt; is the more precise description. Moonshot has released the model weights under its own Kimi K3 license, so developers can inspect, deploy, and build on the model within those terms.&lt;/p&gt;

&lt;p&gt;However, downloading the weights does not make inference free. A 2.8T model is far beyond the practical local setup of most developers. Unless you have access to serious GPU infrastructure, you will use a hosted provider—and that provider pays the compute bill.&lt;/p&gt;

&lt;p&gt;That distinction becomes important when we get to the free setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  The benchmark results are genuinely competitive
&lt;/h2&gt;

&lt;p&gt;Kimi K3's coding scores are the main reason I took it seriously.&lt;/p&gt;

&lt;p&gt;Here are selected results from Moonshot's published evaluation:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Kimi K3&lt;/th&gt;
&lt;th&gt;Claude Fable 5&lt;/th&gt;
&lt;th&gt;GPT-5.6 Sol&lt;/th&gt;
&lt;th&gt;Claude Opus 4.8&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DeepSWE&lt;/td&gt;
&lt;td&gt;67.5&lt;/td&gt;
&lt;td&gt;70.0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;73.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;59.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ProgramBench&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;77.8&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;76.8&lt;/td&gt;
&lt;td&gt;77.6&lt;/td&gt;
&lt;td&gt;71.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-Bench 2.1&lt;/td&gt;
&lt;td&gt;88.3&lt;/td&gt;
&lt;td&gt;88.0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;88.8&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;84.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FrontierSWE&lt;/td&gt;
&lt;td&gt;81.2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;86.6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;71.3&lt;/td&gt;
&lt;td&gt;66.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SWE-Marathon&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;42.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;35.0&lt;/td&gt;
&lt;td&gt;39.0&lt;/td&gt;
&lt;td&gt;40.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BrowseComp&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;91.2&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;88.0&lt;/td&gt;
&lt;td&gt;90.4&lt;/td&gt;
&lt;td&gt;84.3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern is more interesting than a simple “Kimi wins” headline:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It nearly matches GPT-5.6 Sol on Terminal-Bench 2.1.&lt;/li&gt;
&lt;li&gt;It beats GPT-5.6 Sol and Opus 4.8 on FrontierSWE, although Fable 5 scores higher.&lt;/li&gt;
&lt;li&gt;It leads this comparison on ProgramBench, SWE-Marathon, and BrowseComp.&lt;/li&gt;
&lt;li&gt;It falls behind both GPT-5.6 Sol and Fable 5 on DeepSWE.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is not a clean sweep. It is something more useful: evidence that an open-weight model now belongs in the same coding conversation as the leading proprietary models.&lt;/p&gt;

&lt;h3&gt;
  
  
  One benchmark warning most reviews skip
&lt;/h3&gt;

&lt;p&gt;These numbers come from Moonshot's evaluation, and some models were tested with different agent harnesses. Kimi used Kimi Code on several tests, GPT used Codex on some, and Claude used Claude Code or other harnesses on others. All Kimi results also used maximum reasoning effort.&lt;/p&gt;

&lt;p&gt;In other words, treat the table as a strong signal—not a perfectly controlled, apples-to-apples contest. The &lt;a href="https://arxiv.org/abs/2607.24653" rel="noopener noreferrer"&gt;technical report&lt;/a&gt; and repository document the methodology if you want to inspect it.&lt;/p&gt;

&lt;h2&gt;
  
  
  My practical test: build a browser game from one prompt
&lt;/h2&gt;

&lt;p&gt;Benchmarks tell me whether a model deserves a test. They do not tell me whether I want it editing my project.&lt;/p&gt;

&lt;p&gt;So I gave Kimi K3 a long, single prompt to generate a complete browser game. This forced it to plan the interface, write the game logic, connect the files, and keep the result coherent over a longer generation.&lt;/p&gt;

&lt;p&gt;It produced a complete, playable first version in one pass.&lt;/p&gt;

&lt;p&gt;What impressed me was not one clever function. It was the model's ability to sustain a multi-part implementation without losing the original goal.&lt;/p&gt;

&lt;p&gt;But a one-shot demo has limits. A generated game can look impressive while still hiding brittle state management, accessibility problems, or code that becomes painful on the second revision. The better test is whether the model can explain its decisions, respond to bug reports, make targeted changes, and run verification without rewriting unrelated code.&lt;/p&gt;

&lt;p&gt;That is how I would evaluate Kimi K3 on a real repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to try Kimi K3 for free
&lt;/h2&gt;

&lt;p&gt;I found two practical routes: one in the browser and one inside VS Code.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Availability note:&lt;/strong&gt; The free Kimi K3 routes I tested are promotional and may be rate-limited, renamed, or removed. Check the provider's model page before following the steps.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Method 1: Use Kimi K3 in your browser
&lt;/h2&gt;

&lt;p&gt;When I tested it, GenSpark included Kimi K3 in its model menu after a free signup.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create a free GenSpark account.&lt;/li&gt;
&lt;li&gt;Open the AI chat interface.&lt;/li&gt;
&lt;li&gt;Select Kimi K3 from the model menu.&lt;/li&gt;
&lt;li&gt;Start with a real task—not “write hello world.”&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Try asking it to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Explain an unfamiliar module and identify risky dependencies.&lt;/li&gt;
&lt;li&gt;Build a small feature with clear acceptance criteria.&lt;/li&gt;
&lt;li&gt;Review a pull request and separate bugs from style preferences.&lt;/li&gt;
&lt;li&gt;Turn a screenshot into a working frontend component.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In my usage, Kimi consumed fewer GenSpark credits than the premium Claude and GPT options, so the free allowance lasted longer. Credit rules can change, so verify the current rate in the interface.&lt;/p&gt;

&lt;h2&gt;
  
  
  Method 2: Connect Kimi K3 to VS Code
&lt;/h2&gt;

&lt;p&gt;For coding, this is the more useful setup because the model can work with your files through an agent extension.&lt;/p&gt;

&lt;p&gt;You can use Kilo Code or Cline. Both support OpenAI-compatible providers, which means you normally need only three pieces of information:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A provider Base URL&lt;/li&gt;
&lt;li&gt;An API key&lt;/li&gt;
&lt;li&gt;The exact model ID&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Step 1: Install the extension
&lt;/h3&gt;

&lt;p&gt;Install &lt;strong&gt;Kilo Code&lt;/strong&gt; from the VS Code Marketplace. If its provider setup gives you trouble, install &lt;strong&gt;Cline&lt;/strong&gt; instead; &lt;a href="https://docs.cline.bot/provider-config/openai-compatible" rel="noopener noreferrer"&gt;Cline documents the same OpenAI-compatible fields&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Create a provider API key
&lt;/h3&gt;

&lt;p&gt;Use a current provider that lists a promotional Kimi K3 route. At the time of writing, the following OpenAI-compatible configuration works with ZenMux:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Provider: OpenAI Compatible
Base URL: https://zenmux.ai/api/v1
Model ID: moonshotai/kimi-k3-free
API key: &amp;lt;your provider key&amp;gt;
Reasoning: enabled
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The free model is explicitly listed as a limited-time route, so confirm that the model ID still appears in the provider's catalog before setting up the extension.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Configure Kilo Code or Cline
&lt;/h3&gt;

&lt;p&gt;Open the extension's model settings and:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Choose &lt;strong&gt;OpenAI Compatible&lt;/strong&gt; as the provider.&lt;/li&gt;
&lt;li&gt;Paste the provider's Base URL.&lt;/li&gt;
&lt;li&gt;Paste your API key.&lt;/li&gt;
&lt;li&gt;Enter the exact Kimi K3 model ID.&lt;/li&gt;
&lt;li&gt;Enable reasoning or thinking mode if the client exposes that setting.&lt;/li&gt;
&lt;li&gt;Save and send a small test prompt.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you receive “model not found,” do not keep changing random settings. Check the provider's live model list first. Promotional model IDs change frequently.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Test it safely
&lt;/h3&gt;

&lt;p&gt;Start with a disposable branch or a small personal project. Ask the agent to explain its plan before editing, review the diff after each task, and keep automatic command approval off until you trust the workflow.&lt;/p&gt;

&lt;p&gt;Also remember that a hosted endpoint can receive the prompts and code you send through it. Do not upload credentials, private customer data, or proprietary source code until you have reviewed the provider's privacy and retention policies.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real catch behind “free”
&lt;/h2&gt;

&lt;p&gt;The free access is real, but “free” can mean three different things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Free weights:&lt;/strong&gt; You can download the model under its license, but you supply the hardware.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Free web access:&lt;/strong&gt; A product absorbs the inference cost and applies its own limits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Free promotional API access:&lt;/strong&gt; A provider offers a zero-cost route temporarily, usually with rate limits and no service guarantee.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The VS Code method falls into the third category. It is excellent for testing and personal projects. I would not build a production workflow around the assumption that the endpoint will stay unlimited or free forever.&lt;/p&gt;

&lt;p&gt;There is another practical catch: the official Kimi API expects clients to preserve reasoning content across multi-turn tool calls. If your coding extension drops that state, you may see weaker follow-up behavior even when the first response looks good.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should you replace your paid coding assistant?
&lt;/h2&gt;

&lt;p&gt;Here is my honest recommendation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Kimi K3 could become your default if:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Most of your work is coding, browser automation, or agentic web tasks.&lt;/li&gt;
&lt;li&gt;You regularly need to reason across a large repository.&lt;/li&gt;
&lt;li&gt;You are a student, indie developer, or early-stage builder minimizing subscriptions.&lt;/li&gt;
&lt;li&gt;You are comfortable switching providers if a free route disappears.&lt;/li&gt;
&lt;li&gt;You review diffs and verify generated code instead of accepting it blindly.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Keep a paid model as your primary if:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Reliability and support matter more than saving the subscription fee.&lt;/li&gt;
&lt;li&gt;You work with sensitive code and need clear enterprise data controls.&lt;/li&gt;
&lt;li&gt;Your tasks are ambiguous, high-stakes, or difficult to verify.&lt;/li&gt;
&lt;li&gt;You depend on stable throughput, predictable latency, or a service-level agreement.&lt;/li&gt;
&lt;li&gt;You want one assistant for coding, writing, analysis, and specialized professional work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For many developers, the best answer is not a dramatic switch. It is a two-model workflow: use Kimi K3 as the high-context coding workhorse and keep a paid model as the fallback for difficult or high-stakes tasks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final verdict
&lt;/h2&gt;

&lt;p&gt;Kimi K3 does not make every paid coding assistant obsolete.&lt;/p&gt;

&lt;p&gt;What it does is more significant: it narrows the gap enough that “open model” no longer automatically means “second-tier coding model.” Its benchmark performance is competitive, its million-token context is genuinely useful, and the hosted free routes make it easy to test inside a real coding workflow.&lt;/p&gt;

&lt;p&gt;My verdict:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Do not cancel your paid assistant because of one leaderboard. But if you write code, Kimi K3 deserves a permanent slot in your model picker while the free access lasts.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Have you tried Kimi K3 on a real repository? Share the task, what it did well, and where it failed. That comparison is more useful than another benchmark screenshot.&lt;/p&gt;

&lt;p&gt;If this walkthrough helped, follow my YouTube channel, &lt;strong&gt;GoVenture&lt;/strong&gt;, for more practical and honest AI-tool tests.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links and sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/MoonshotAI/Kimi-K3" rel="noopener noreferrer"&gt;Kimi K3 official repository and benchmark table&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2607.24653" rel="noopener noreferrer"&gt;Kimi K3 technical report&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.kimi.com/help/kimi-api" rel="noopener noreferrer"&gt;Kimi API documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.cline.bot/provider-config/openai-compatible" rel="noopener noreferrer"&gt;Cline: configuring an OpenAI-compatible provider&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://zenmux.ai/docs/guide/quickstart" rel="noopener noreferrer"&gt;ZenMux OpenAI-compatible API quickstart&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;GenSpark (browser method): &lt;a href="http://genspark.ai/" rel="noopener noreferrer"&gt;http://genspark.ai/&lt;/a&gt; &lt;/li&gt;
&lt;li&gt;Free API provider Token Router (VS Code method) : &lt;a href="https://www.tokenrouter.com/" rel="noopener noreferrer"&gt;https://www.tokenrouter.com/&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>moonshot</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Meta Muse Spark 1.2 Honest Review: Loses the Benchmarks, Wins as an Agent</title>
      <dc:creator>Aditi Gupta</dc:creator>
      <pubDate>Sat, 08 Aug 2026 08:49:34 +0000</pubDate>
      <link>https://dev.to/aditi_gupta_8d81622a592aa/meta-muse-spark-12-honest-review-loses-the-benchmarks-wins-as-an-agent-1m8c</link>
      <guid>https://dev.to/aditi_gupta_8d81622a592aa/meta-muse-spark-12-honest-review-loses-the-benchmarks-wins-as-an-agent-1m8c</guid>
      <description>&lt;p&gt;Meta just shipped two things at once: Muse Spark 1.2, the model, and Muse Code, an agentic harness that wraps around it. The interesting part is that they pull in opposite directions. On raw coding the model is a clear underdog. As an agent it is suddenly near the top. Here is what actually held up when I tested it.&lt;/p&gt;

&lt;p&gt;Prefer the quick version? I covered the benchmark split, the broken games, the harness features, and a head to head against Qwen 3.8 in this video:&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/5DjBu90hacg"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;What Muse Spark and Muse Code are&lt;br&gt;
Muse Spark 1.2 is Meta's new coding model. Muse Code is the harness you actually run it in. The harness keeps a local log of every tool call and every edit, so if a run crashes it picks up exactly where it left off instead of starting over. It ships with built-in skills too: /plan turns a task into an approval-gated plan, /grill stress-tests that plan before you run it, and /goal keeps the agent pushing toward the objective.&lt;/p&gt;

&lt;p&gt;Where it loses&lt;br&gt;
On the coding benchmarks it never comes first. Second on Terminal Bench behind Opus 5, third on DeepSWE, and behind Opus 5 on Meta's own internal benchmark. The games I had it build showed the same weakness. They were laggy and half-finished, one dragon bounced on the spot, one character's arms were missing, another spun around when you tried to walk backwards. On the same prompts, Qwen 3.8 and Fable 5 built noticeably cleaner, more playable versions.&lt;/p&gt;

&lt;p&gt;Where it wins&lt;br&gt;
Point the same model at tools inside Muse Code and the picture flips. On agent and tool-use benchmarks it jumps to first. The crash-resume log is the standout feature for anyone running long agent jobs. To test it properly I gave both Muse Spark and Qwen 3.8 the same task: read a guide and turn it into a reusable skill. Muse Spark replied faster and produced the sharper result. It analyzed the guide, built a proper table, and got specific instead of generic.&lt;/p&gt;

&lt;p&gt;Who should run it&lt;br&gt;
If you want the best raw coding model, Opus 5 still wins and Qwen 3.8 still builds cleaner. If you care about agent workflows, tool use, and not losing progress when a long run dies, Muse Code is worth a serious look even though the model underneath loses the benchmark race.&lt;/p&gt;

&lt;p&gt;My verdict&lt;br&gt;
The model is not the story. The harness is. Muse Spark 1.2 is a reminder that in 2026 the wrapper around a model can matter as much as the weights.&lt;/p&gt;

&lt;p&gt;Have you tried Muse Spark or Muse Code yet? Tell me in the comments. If you like quick, honest AI-tool reviews, follow my YouTube channel for the next one.&lt;/p&gt;

&lt;p&gt;Tags: #ai #metaai #musespark #aiagents #coding``&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>meta</category>
    </item>
    <item>
      <title>Qwen3.8-Max Beat Claude</title>
      <dc:creator>Aditi Gupta</dc:creator>
      <pubDate>Tue, 04 Aug 2026 11:34:00 +0000</pubDate>
      <link>https://dev.to/aditi_gupta_8d81622a592aa/qwen38-max-vs-claude-what-the-16-day-coding-run-and-benchmarks-really-show-3bje</link>
      <guid>https://dev.to/aditi_gupta_8d81622a592aa/qwen38-max-vs-claude-what-the-16-day-coding-run-and-benchmarks-really-show-3bje</guid>
      <description>&lt;p&gt;Alibaba has released Qwen3.8-Max, its largest and most capable AI model to date.&lt;/p&gt;

&lt;p&gt;The headline specifications are absurd: 2.4 trillion total parameters, 95 billion activated per token, multimodal input, and a one-million-token context window.&lt;/p&gt;

&lt;p&gt;But the specification sheet is not the most interesting part.&lt;/p&gt;

&lt;p&gt;Alibaba says Qwen3.8-Max operated autonomously for roughly 16 days, starting with an empty repository and building a working software project through issues, code changes, testing, pull requests, and self-correction.&lt;/p&gt;

&lt;p&gt;That sounds impressive. It also sounds suspiciously like the sort of claim that deserves more inspection than a celebratory repost.&lt;/p&gt;

&lt;p&gt;So I examined Alibaba's announcement, the public repository, and its benchmark results to answer a more useful question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does Qwen3.8-Max actually beat Claude, or did the benchmark department simply have an excellent week?&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Prefer the two-minute version?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/oa0WWLd4uCw"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Qwen3.8-Max is a mixture-of-experts model with 2.4 trillion total parameters and 95 billion active per token.&lt;/li&gt;
&lt;li&gt;Alibaba reports a one-million-token context window and support for text, images, and agentic workflows.&lt;/li&gt;
&lt;li&gt;It beats the Claude models in Alibaba's table on Terminal Bench 2.1, PaperBench, and OSWorld-Verified.&lt;/li&gt;
&lt;li&gt;It trails Claude Fable 5 on demanding repository-level engineering benchmarks such as SWE-bench Pro and FrontierSWE.&lt;/li&gt;
&lt;li&gt;Alibaba says the model weights will be released next week. At the time of writing, the API is available, but the weights are not yet downloadable.&lt;/li&gt;
&lt;li&gt;The 16-day coding trace is public, but the results still come from Alibaba's own evaluation setup and need independent replication.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What is Qwen3.8-Max?
&lt;/h2&gt;

&lt;p&gt;Qwen3.8-Max is Alibaba's new flagship mixture-of-experts model.&lt;/p&gt;

&lt;p&gt;It contains &lt;strong&gt;2.4 trillion parameters in total&lt;/strong&gt;, with approximately &lt;strong&gt;95 billion active during each forward pass&lt;/strong&gt;. This architecture allows Alibaba to scale the model's capacity without paying the full inference cost of a dense 2.4-trillion-parameter model on every token.&lt;/p&gt;

&lt;p&gt;The model also supports a &lt;strong&gt;one-million-token context window&lt;/strong&gt;, making it suitable for large repositories, long documents, persistent agent sessions, and other tasks where context compression usually arrives carrying a shovel.&lt;/p&gt;

&lt;p&gt;Alibaba calls it the first Qwen model at Max scale that will receive an open-weight release. However, there is an important distinction:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The API is available now. The model weights are scheduled for release next week.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So calling it an "open-weight model" is reasonable when discussing Alibaba's release plan, but saying the weights are already available would be inaccurate.&lt;/p&gt;

&lt;p&gt;You can find the specifications and release details in the &lt;a href="https://qwen.ai/blog?id=qwen3.8" rel="noopener noreferrer"&gt;official Qwen3.8-Max announcement&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 16-day autonomous coding run
&lt;/h2&gt;

&lt;p&gt;Alibaba asked Qwen3.8-Max to create a project called &lt;code&gt;oh-my-cli&lt;/code&gt; from an empty repository.&lt;/p&gt;

&lt;p&gt;Instead of responding to a single prompt and stopping, the model worked through a continuous engineering loop:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Convert feedback and requirements into GitHub issues.&lt;/li&gt;
&lt;li&gt;Claim and execute individual tasks.&lt;/li&gt;
&lt;li&gt;Write and modify code.&lt;/li&gt;
&lt;li&gt;Run builds, unit tests, end-to-end tests, and lifecycle checks.&lt;/li&gt;
&lt;li&gt;Route failures back into the issue workflow.&lt;/li&gt;
&lt;li&gt;Fix the problems and verify the result.&lt;/li&gt;
&lt;li&gt;Merge completed pull requests.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;According to Alibaba, the repository had accumulated &lt;strong&gt;265 commits, 127 pull requests, and 151 issues&lt;/strong&gt; after approximately 16 days of autonomous operation.&lt;/p&gt;

&lt;p&gt;The complete project history is available in the public &lt;a href="https://github.com/qwen-code-dev-bot/oh-my-cli" rel="noopener noreferrer"&gt;&lt;code&gt;oh-my-cli&lt;/code&gt; GitHub repository&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;That transparency matters. Most autonomous-agent demonstrations give us a polished video and ask us to believe that nothing caught fire outside the frame. Here, developers can inspect the issues, commits, pull requests, tests, and failures.&lt;/p&gt;

&lt;p&gt;Still, this does not prove that the model can autonomously build any production system for 16 days. It proves that Qwen3.8-Max performed this particular task inside a structured environment with automated testing and feedback loops.&lt;/p&gt;

&lt;p&gt;That is still meaningful, just narrower than the marketing headline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Qwen3.8-Max beats Claude
&lt;/h2&gt;

&lt;p&gt;Alibaba published a large benchmark table comparing Qwen3.8-Max with Claude Opus 4.8, Claude Fable 5, GPT-5.6 Sol, and Qwen3.7-Max.&lt;/p&gt;

&lt;p&gt;Here are the most relevant results:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Qwen3.8-Max&lt;/th&gt;
&lt;th&gt;Claude Opus 4.8&lt;/th&gt;
&lt;th&gt;Claude Fable 5&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Terminal Bench 2.1&lt;/td&gt;
&lt;td&gt;86.6&lt;/td&gt;
&lt;td&gt;84.6&lt;/td&gt;
&lt;td&gt;84.6&lt;/td&gt;
&lt;td&gt;Qwen leads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PaperBench&lt;/td&gt;
&lt;td&gt;93.0&lt;/td&gt;
&lt;td&gt;80.3&lt;/td&gt;
&lt;td&gt;88.8&lt;/td&gt;
&lt;td&gt;Qwen leads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OSWorld-Verified&lt;/td&gt;
&lt;td&gt;86.1&lt;/td&gt;
&lt;td&gt;83.4&lt;/td&gt;
&lt;td&gt;85.0&lt;/td&gt;
&lt;td&gt;Qwen leads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SWE-bench Pro&lt;/td&gt;
&lt;td&gt;67.7&lt;/td&gt;
&lt;td&gt;69.2&lt;/td&gt;
&lt;td&gt;80.0&lt;/td&gt;
&lt;td&gt;Claude leads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FrontierSWE&lt;/td&gt;
&lt;td&gt;73.5&lt;/td&gt;
&lt;td&gt;70.0&lt;/td&gt;
&lt;td&gt;88.8&lt;/td&gt;
&lt;td&gt;Mixed; Fable 5 leads&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These results suggest three areas where Qwen3.8-Max looks particularly strong.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Terminal-based agent work
&lt;/h3&gt;

&lt;p&gt;Its &lt;strong&gt;86.6 score on Terminal Bench 2.1&lt;/strong&gt; puts it ahead of both Claude models in Alibaba's comparison.&lt;/p&gt;

&lt;p&gt;That makes Qwen especially interesting for command-line agents, environment setup, testing, deployment workflows, and tasks that require repeated tool use rather than a single code-generation response.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Research reproduction
&lt;/h3&gt;

&lt;p&gt;Qwen3.8-Max scored &lt;strong&gt;93.0 on PaperBench&lt;/strong&gt;, ahead of Claude Fable 5's 88.8 and Opus 4.8's 80.3.&lt;/p&gt;

&lt;p&gt;Alibaba also demonstrated a five-day research task in which the model reproduced a paper's experimental pipeline, ran 33 rounds of GPU training, and then searched for improvements to the original method.&lt;/p&gt;

&lt;p&gt;This is potentially more useful than another model becoming marginally better at generating React components nobody requested.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Computer and visual interaction
&lt;/h3&gt;

&lt;p&gt;On &lt;strong&gt;OSWorld-Verified&lt;/strong&gt;, which evaluates an agent's ability to operate computer environments, Qwen3.8-Max scored 86.1.&lt;/p&gt;

&lt;p&gt;The model uses visual output as part of its feedback loop. It can inspect an interface, identify errors, revise its plan, and try again. That matters for browser agents, desktop automation, document workflows, UI testing, and multimodal development.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Claude still wins
&lt;/h2&gt;

&lt;p&gt;The "Qwen kills Claude" headline falls apart once we examine harder repository-level engineering tasks.&lt;/p&gt;

&lt;p&gt;On &lt;strong&gt;SWE-bench Pro&lt;/strong&gt;, Qwen3.8-Max scored 67.7. Claude Fable 5 scored 80.0.&lt;/p&gt;

&lt;p&gt;On &lt;strong&gt;FrontierSWE&lt;/strong&gt;, Qwen scored 73.5 while Fable 5 reached 88.8.&lt;/p&gt;

&lt;p&gt;That is not a rounding error. It suggests Claude remains stronger when a task requires deep repository understanding, architectural judgment, and reliable changes across a complicated codebase.&lt;/p&gt;

&lt;p&gt;The more honest conclusion is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Qwen looks excellent at long-running, tool-heavy work.&lt;/li&gt;
&lt;li&gt;It performs strongly on terminal, research, multimodal, and computer-use tasks.&lt;/li&gt;
&lt;li&gt;Claude Fable 5 remains ahead on some of the hardest software-engineering benchmarks.&lt;/li&gt;
&lt;li&gt;Neither model "wins" every category, because reality rudely refuses to fit inside one thumbnail.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is another caveat: these scores come from Alibaba's evaluation table. Different benchmarks used different harnesses, time limits, context settings, and judging methods. The numbers are useful, but independent testing will matter more than launch-day charts.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to use Qwen3.8-Max with Claude Code
&lt;/h2&gt;

&lt;p&gt;QwenCloud provides an Anthropic-compatible API, allowing Claude Code to use Qwen3.8-Max without replacing the Claude Code interface.&lt;/p&gt;

&lt;p&gt;First, install Claude Code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @anthropic-ai/claude-code
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then configure it to use Qwen:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"qwen3.8-max"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_SMALL_FAST_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"qwen3.8-max"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://dashscope-intl.aliyuncs.com/apps/anthropic"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_AUTH_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"YOUR_QWEN_API_KEY"&lt;/span&gt;

claude
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You will need a QwenCloud API key. The international endpoint may differ depending on your account or deployment region, so check the current &lt;a href="https://www.qwencloud.com/" rel="noopener noreferrer"&gt;QwenCloud documentation&lt;/a&gt; before configuring it.&lt;/p&gt;

&lt;p&gt;And please do not paste your real API key into a public DEV article. Becoming an involuntary cloud-compute philanthropist is rarely part of the content strategy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can you run it locally?
&lt;/h2&gt;

&lt;p&gt;Not casually.&lt;/p&gt;

&lt;p&gt;Although only 95 billion parameters are active during each pass, the full model contains 2.4 trillion parameters. Open weights do not magically convert that into something your laptop can run between Chrome tabs.&lt;/p&gt;

&lt;p&gt;Once the weights are released, practical deployment will likely require substantial multi-GPU infrastructure, aggressive quantization, or a hosted inference provider.&lt;/p&gt;

&lt;p&gt;For most individual developers, QwenCloud will be the realistic way to use the full model.&lt;/p&gt;

&lt;h2&gt;
  
  
  The verdict
&lt;/h2&gt;

&lt;p&gt;Qwen3.8-Max does not kill Claude.&lt;/p&gt;

&lt;p&gt;It does something more consequential: it brings frontier-scale agent capabilities closer to the open-weight ecosystem.&lt;/p&gt;

&lt;p&gt;Its strongest argument is not a single benchmark score. It is the combination of long-horizon execution, terminal performance, multimodal feedback, research reproduction, and a public 16-day development trace.&lt;/p&gt;

&lt;p&gt;Based on the evidence available today:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;I would consider Qwen3.8-Max for long-running agents, terminal workflows, research automation, visual tasks, and jobs that benefit from repeated feedback.&lt;/li&gt;
&lt;li&gt;I would still prefer Claude Fable 5 for the hardest repository-level engineering work, especially when first-pass reliability matters.&lt;/li&gt;
&lt;li&gt;I would wait for independent evaluations before treating Alibaba's benchmark table as the final verdict.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Qwen3.8-Max is not the model that makes Claude irrelevant.&lt;/p&gt;

&lt;p&gt;It is the model that makes the frontier race significantly less comfortable, and that is far more interesting.&lt;/p&gt;

&lt;p&gt;Have you tested Qwen3.8-Max in QwenCloud or Claude Code? Share the task, harness, and result in the comments. "It felt smarter" is emotionally valid, but logs are sexier.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://qwen.ai/blog?id=qwen3.8" rel="noopener noreferrer"&gt;Qwen3.8-Max official announcement and benchmark table&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/qwen-code-dev-bot/oh-my-cli" rel="noopener noreferrer"&gt;Public &lt;code&gt;oh-my-cli&lt;/code&gt; autonomous coding repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.alibabacloud.com/help/en/model-studio/token-plan-harness-tool" rel="noopener noreferrer"&gt;Alibaba Cloud Model Studio integration documentation&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Claude Opus 5 Is Better at Coding and Harder to Trust</title>
      <dc:creator>Aditi Gupta</dc:creator>
      <pubDate>Wed, 29 Jul 2026 13:18:53 +0000</pubDate>
      <link>https://dev.to/aditi_gupta_8d81622a592aa/claude-opus-5-is-better-at-coding-and-harder-to-trust-4ga5</link>
      <guid>https://dev.to/aditi_gupta_8d81622a592aa/claude-opus-5-is-better-at-coding-and-harder-to-trust-4ga5</guid>
      <description>&lt;p&gt;Claude Opus 5 completed one of my coding tasks considerably faster than Opus 4.8.&lt;/p&gt;

&lt;p&gt;There was just one problem: it confidently reported that the issue was fixed when it wasn’t.&lt;/p&gt;

&lt;p&gt;That experience captures the trade-off with Anthropic’s latest Opus model. It is faster and more capable on difficult, multi-step work, but polished output can make its mistakes harder to notice.&lt;/p&gt;

&lt;p&gt;After testing it on coding and agent tasks, I changed three parts of my workflow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Start with medium reasoning effort&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;More reasoning is not automatically better. For routine coding tasks, begin with medium effort and increase it only when the problem genuinely requires deeper investigation.&lt;/p&gt;

&lt;p&gt;Higher effort can consume more tokens, expand the scope of the task, and produce a solution far more elaborate than the one you requested.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Verify outcomes, not explanations&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A convincing explanation is not evidence that the task was completed correctly.&lt;/p&gt;

&lt;p&gt;Ask for or independently run the relevant tests. Review the files that changed. Confirm the original bug no longer exists.&lt;/p&gt;

&lt;p&gt;The dangerous failure mode is not nonsense. It is an incorrect result presented like finished work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Control the scope&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Define what the model may change before it begins:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which files can be modified?&lt;/li&gt;
&lt;li&gt;What behaviour must remain unchanged?&lt;/li&gt;
&lt;li&gt;Which tests must pass?&lt;/li&gt;
&lt;li&gt;Can it create subagents or expand the task?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Opus 5 is strongest when the job requires investigation across multiple steps. For a small, clearly defined change, that same initiative can become unnecessary complexity.&lt;/p&gt;

&lt;h2&gt;
  
  
  The quick version
&lt;/h2&gt;

&lt;p&gt;I condensed my findings, the confidently wrong problem, and the three changes I recommend into this 90-second video:&lt;/p&gt;


&lt;div&gt;
    &lt;iframe src="https://www.youtube.com/embed/-NS3MOxW7EQ"&gt;
    &lt;/iframe&gt;
  &lt;/div&gt;


&lt;p&gt;My broader verdict is simple: Opus 5 is a meaningful upgrade for difficult coding and agent work, but only when verification is part of the workflow.&lt;/p&gt;

&lt;p&gt;I published the complete review, including pricing, benchmark comparisons, use cases, and switching advice, on Hashnode:&lt;/p&gt;


&lt;div class="crayons-card c-embed text-styles text-styles--secondary"&gt;
    &lt;div class="c-embed__content"&gt;
      &lt;div class="c-embed__body flex items-center justify-between"&gt;
        &lt;a href="https://hashnode.com/edit/cms61pd2q00000bj8d9gu2nx0" rel="noopener noreferrer" class="c-link fw-bold flex items-center"&gt;
          &lt;span class="mr-2"&gt;hashnode.com&lt;/span&gt;
          

        &lt;/a&gt;
      &lt;/div&gt;
    &lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;Have you tested Opus 5? Did it improve your workflow, or merely become more articulate while being wrong?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>coding</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
