<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Nancy Garg</title>
    <description>The latest articles on DEV Community by Nancy Garg (@nanseedotai).</description>
    <link>https://dev.to/nanseedotai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4164259%2Ff2d067c7-7361-4b32-9fc6-c4fdca881cc6.jpg</url>
      <title>DEV Community: Nancy Garg</title>
      <link>https://dev.to/nanseedotai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/nanseedotai"/>
    <language>en</language>
    <item>
      <title>10 Projects You Should Build with Jev</title>
      <dc:creator>Nancy Garg</dc:creator>
      <pubDate>Tue, 06 Oct 2026 19:25:08 +0000</pubDate>
      <link>https://dev.to/studio1hq/10-projects-you-should-build-with-jev-44l1</link>
      <guid>https://dev.to/studio1hq/10-projects-you-should-build-with-jev-44l1</guid>
      <description>&lt;p&gt;JEV is getting a lot of attention right now. &lt;/p&gt;

&lt;p&gt;People are excited about it, and developers are already experimenting with ways to use it inside real applications.But there is a common mistake when trying a new AI model like JEV: treating it like another general-purpose LLM.&lt;/p&gt;

&lt;p&gt;JEV becomes much more interesting when you use it for decisions rather than generation. Instead of asking it to write an email or generate a long answer, you give it some context, define the decisions it can make, and let your application act on the result.For example, a support system could use JEV to decide whether a ticket belongs to billing, engineering, or sales. A browser agent could use it to choose which button to click next. An AI application could use it to decide which model should handle a particular request.The model makes the decision, while your code handles the action.That simple pattern opens up a lot of interesting use cases. Here are 10 projects worth building with JEV.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Build a smart support ticket router
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg98gtenx38hg1nvhfq13.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg98gtenx38hg1nvhfq13.png" alt="first-usecase" width="799" height="448"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Support teams spend a lot of time reading incoming tickets and deciding where each one should go. A ticket might be a billing question, a technical issue, a feature request, or an urgent complaint from an unhappy customer.JEV can handle the first layer of that workflow.You can give it the ticket and ask which team should handle it, how urgent the issue is, whether the customer appears frustrated, and whether the ticket needs immediate human attention. Your application can then use those decisions to route the ticket automatically.A technical issue can go to engineering, a refund request can go to billing, and an urgent issue can be moved to the top of the queue.The important part is that JEV does not need to write the customer response. A generative model can handle that later if needed. JEV only needs to make the decision that determines what happens next.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Build an inbound lead scorer
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffv7y8otb52w82rqc2h90.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffv7y8otb52w82rqc2h90.png" alt="second-usecase" width="799" height="446"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The same approach works well for sales.A typical website can receive a mix of serious prospects, students, competitors, and people who are simply exploring a product. Sales teams often have to spend time sorting through these submissions before deciding which ones deserve attention.You can use JEV to score each lead based on information such as the person's job title, company, form submission, and other available context.The model can determine whether the lead fits your ideal customer profile, whether they appear ready to buy, whether they are likely to be a decision-maker, and how the lead should be prioritized.Your application can then route high-priority leads directly to sales while sending lower-priority leads into other workflows.This turns JEV into a lightweight decision layer for your sales pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Build a comment moderation system
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsiz3otcgk42levm7fq1i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsiz3otcgk42levm7fq1i.png" alt="third-usecase" width="800" height="455"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Moderation is another problem that can be expressed as a set of decisions.For a community, social platform, or Discord server, you can give JEV a comment and ask whether it is spam, harassment, unsafe content, or something that requires human review.Your application can then use those results to decide what happens to the comment. Clearly safe content can be published automatically, while potentially harmful or uncertain content can be sent to a moderator.You can also introduce confidence thresholds so that the system only automates decisions when JEV is sufficiently confident.This creates a useful balance between automation and human review instead of assuming that every AI decision should be acted on automatically.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Build an AI citation checker
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flana5yuoc08021nsxj0m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flana5yuoc08021nsxj0m.png" alt="fourth-usecase" width="800" height="449"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;AI-generated content often includes claims that sound convincing but are not actually supported by their sources.A simple JEV application could check whether a specific source supports a specific claim.Give the model the claim and the relevant source, then ask whether the source supports the claim, contradicts it, or does not contain enough information to determine the answer.For example, if an AI-generated report says that a company grew its revenue by 40% while the cited source says it grew by 12%, the application can flag the mismatch before the report reaches the user.This could be useful for research tools, education platforms, content workflows, and AI agents that work with external sources.The key is to keep the question narrow. You are not asking JEV to research the entire internet. You are asking it to make one specific judgment about a piece of evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Build a smarter RAG context picker
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxuzbxohllltiat8oostz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxuzbxohllltiat8oostz.png" alt="fifth-usecase" width="800" height="449"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;RAG systems often retrieve several documents and pass them directly into the context of a larger language model. The problem is that not every retrieved document is useful.Some documents may be irrelevant, outdated, contradictory, or contain instructions that should not be passed into the final context.JEV can sit between retrieval and generation as a filtering layer.For example, your search system could retrieve 20 documents and JEV could decide which ones are relevant to the user's question, which ones are outdated, whether any documents conflict with each other, and whether a document should be included in the final context.The main LLM can then focus on generating the answer using a cleaner set of information.In this architecture, the LLM handles generation while JEV helps decide what information the LLM should see.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Build a semantic code review bot
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fln2gaqdjhci9waia14oz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fln2gaqdjhci9waia14oz.png" alt="sixth-usecase" width="799" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Traditional linters are good at catching problems that can be expressed as explicit rules, such as formatting issues, type errors, and missing imports.Engineering teams also have rules that are much harder to express through traditional static analysis.For example, does a new API expose private customer information? Does a pull request bypass an important approval step? Does a change introduce a risky pattern? Did the developer add tests for the new behavior?These are questions that can be handled by a decision model.You can give JEV the changed code along with the relevant project rules and ask it to evaluate a small set of specific conditions. When it identifies a likely problem with high confidence, your GitHub workflow can flag the pull request or request human review.This does not replace human code review. Instead, it adds another layer that can catch potential issues before an engineer spends time reviewing the entire change.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Build an AI model router
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fznsno8vhpwka0bnj51is.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fznsno8vhpwka0bnj51is.png" alt="seventh-usecase" width="799" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Modern AI applications often have access to several models, and sending every request to the most expensive model is rarely necessary.A simple question might be handled by a small model, while a complex coding task could require a more capable model. Some requests may need retrieval, a tool, a business workflow, or even a human instead of another model.\&lt;br&gt;
\&lt;br&gt;
Give it the user's request and ask it to classify the task based on the characteristics that matter to your application. Your code can then choose the appropriate model or workflow.A simple FAQ can go to a cheaper &amp;amp; faster model like &lt;a href="https://tokenfactory.nebius.com/?modals=endpoint-details&amp;amp;model-id=zai-org/GLM-5.3-Flash" rel="noopener noreferrer"&gt;GLM 5.3 Flash&lt;/a&gt; through &lt;a href="https://x.com/@nebiustf" rel="noopener noreferrer"&gt;@nebiustf&lt;/a&gt; , a difficult coding task can go to a stronger model like &lt;a href="https://tokenfactory.nebius.com/?modals=endpoint-details&amp;amp;model-id=zai-org/GLM-5.3" rel="noopener noreferrer"&gt;GLM 5.3&lt;/a&gt;, and a refund request can be handled by your existing business logic.JEV does not need to answer the user's question. It only needs to decide where that question should go.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Build an AI agent safety firewall
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffnsjkmv270znbgpkguhd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffnsjkmv270znbgpkguhd.png" alt="eight-usecase" width="800" height="438"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;AI agents can now browse websites, send emails, modify records, call APIs, and perform other actions on behalf of users. That makes it important to evaluate an action before allowing an agent to execute it.Imagine an agent wants to delete hundreds of customer records. Before executing the action, you could use JEV to evaluate whether the action is irreversible, whether it involves sensitive information, whether it matches the user's request, and whether it should be allowed, confirmed by the user, or blocked.Your application can then enforce that decision.Low-risk actions can proceed automatically, medium-risk actions can require confirmation, and high-risk actions can be blocked.JEV should not be treated as a replacement for permissions, access controls, validation, or other security measures. It can instead provide an additional semantic layer between an AI agent and the actions it wants to perform.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Build a browser agent reflex layer
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu50j0dpkw1avgw80hawp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu50j0dpkw1avgw80hawp.png" alt="ninth-usecase" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Browser agents spend much of their time making small decisions. They need to decide which button to click, which search result to open, whether to scroll, whether to go back, or which tab to select.These decisions do not always require a large generative model.You can give JEV a compact representation of the current browser state along with a limited set of possible actions. It can then choose the action that best moves the agent toward its goal.Your browser automation system executes that action, the page changes, and JEV makes the next decision.A larger model can still handle tasks that require generation or more complex planning, such as writing an email or filling out a long form.This creates a clean architecture where a larger model handles planning, JEV handles fast decisions, and your automation code handles execution.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Build a real-time game AI
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flzts8xb6oyoftw84fwvs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flzts8xb6oyoftw84fwvs.png" alt="tenth-usecase" width="800" height="449"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One of the most interesting ways to experiment with JEV is to use it as the decision-maker for a small game.Instead of giving the model a huge screenshot and asking it to play the entire game, represent the game state in a compact format. You might tell it that an enemy is nearby, the player's health is low, an obstacle is ahead, and the available actions are attack, jump, retreat, heal, or block. JEV chooses an action, the game executes it, and the updated state is sent back for the next decision.This creates a simple loop where the application provides the state, JEV selects an action, and the game executes it.The same architecture can work for a platformer, driving game, tower defense game, or almost any environment with a relatively small action space.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bigger idea behind JEV
&lt;/h2&gt;

&lt;p&gt;Although these projects look very different, they all follow the same basic architecture. Your application has some state that needs to be interpreted. JEV makes a small, specific decision about that state, and your code takes the resulting action.That makes JEV less interesting as a replacement for a general-purpose chatbot and more interesting as a decision layer inside software.The best use cases are often the decisions that happen repeatedly throughout a product: which ticket should be escalated, which model should handle a request, which document belongs in the context, whether an agent action is safe, or which browser action should happen next.Once you start looking at software through that lens, you can find many places where a fast decision model can fit.So instead of building another chatbot with JEV, look for a workflow where information comes in, your application needs to make a decision, and that decision triggers an action.Let JEV make the decision, and let your code handle everything that happens after it.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>Gemini 4 Argon: What's New, What It Costs, and Whether to Switch</title>
      <dc:creator>Nancy Garg</dc:creator>
      <pubDate>Tue, 06 Oct 2026 19:13:43 +0000</pubDate>
      <link>https://dev.to/studio1hq/gemini-4-argon-whats-new-what-it-costs-and-whether-to-switch-1m7o</link>
      <guid>https://dev.to/studio1hq/gemini-4-argon-whats-new-what-it-costs-and-whether-to-switch-1m7o</guid>
      <description>&lt;p&gt;Google DeepMind recently launched Gemini 4 Argon, with a big benchmark table and a small guest list. Google's own numbers have it winning or tying 13 of 18 tests against Opus 5.5, Fable 5.1 and GPT-6 Astra. But for now it's only open to a group of cyber defenders.&lt;/p&gt;

&lt;p&gt;So the useful question isn't whether it's the best model. It's what you should check before you switch, and what it'll cost you when you do. Here's what we know so far.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Argon wins or ties 13 of 18 benchmarks in Google's table, mostly knowledge work, long context, and DeepSWE. It loses on FrontierSWE and Terminal-bench.&lt;/li&gt;
&lt;li&gt;The output limit jumps from 64K to 1M tokens. It also writes a lot: 62k tokens per task against 27k for GPT-6 Astra in one independent test.&lt;/li&gt;
&lt;li&gt;Launch price is $2 per million input tokens and $10 per million output. It doubles to $4 and $20 later.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What is Gemini 4 Argon?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyq1byreamlb0ie8odxoe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyq1byreamlb0ie8odxoe.png" alt="logo" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Argon is Google's new flagship model. It's the first Gemini above the Flash tier in more than seven months, and it quietly replaces the Gemini 3.5 Pro that Google teased at I/O in May and never shipped.&lt;/p&gt;

&lt;p&gt;Google built it for long, multi-step work: real-world software engineering, enterprise knowledge work like legal and finance, and cyber defense. Two details stand out for developers. The output limit is now 1M tokens, up from 64K. And it was trained to find, validate and patch software vulnerabilities on its own.&lt;/p&gt;

&lt;p&gt;Then there's access. Argon went to trusted cyber defenders in Google's Fairwind Program on September 30. Google says paid API customers and Google AI Ultra subscribers come next, then developers, enterprises and consumers, once the guardrails have been tuned with early testers. It also says it's taking part in the US government's voluntary pre-release access process while it widens the rollout.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it differs from the other models
&lt;/h2&gt;

&lt;p&gt;As compared to the previous versions, the big change is the 1M-token output limit and the cyber training. Google says Argon makes clear leaps over Gemini 3.8 Flash Cyber, its previous security-focused model.&lt;/p&gt;

&lt;p&gt;Against the rest of the field, the pattern in Google's own table is easy to read. Argon leads on knowledge work, long context, and DeepSWE. The other models keep the shell-and-terminal benchmarks.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Leads on (in Google's table)&lt;/th&gt;
&lt;th&gt;Reading&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 4 Argon&lt;/td&gt;
&lt;td&gt;Vals Index, AutomationBench, Vals Finance Agent v2, Harvey's Legal Agent, DeepSWE, Vibe Code Bench, RiemannBench, both GraphWalks tests, Agent's Last Exam, Chartography, LVBench&lt;/td&gt;
&lt;td&gt;Best on document-heavy and long-context work&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;FrontierSWE v2, Terminal-Bench Science 0.1, OSWorld-2.0&lt;/td&gt;
&lt;td&gt;Best on harder agentic and computer-use tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5.5&lt;/td&gt;
&lt;td&gt;Terminal-bench 4.0, PostTrainBench&lt;/td&gt;
&lt;td&gt;Best on shell agents and ML engineering&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5.1&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Never first in the table&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On cybersecurity, Argon and GPT-6 Astra tie at 68% on CWE-bench v1.&lt;/p&gt;

&lt;p&gt;For a neutral check, Artificial Analysis says Argon matches GPT-6 Astra on its Intelligence Index and sits one point ahead of GPT-6.1 Sol.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test scores
&lt;/h2&gt;

&lt;p&gt;Treat these as best cases. They're Google's chosen benchmarks, mostly from its own runs, as tabulated by &lt;a href="https://thenewstack.io/google-gemini-4-argon/" rel="noopener noreferrer"&gt;The New Stack&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7adw8pm54tt9oikcxvht.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7adw8pm54tt9oikcxvht.png" alt="table-of-content" width="800" height="794"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Three caveats on reading it.&lt;/p&gt;

&lt;p&gt;First, several leads are narrow. On the Vals Index and Vibe Code Bench, the gap is under two points. Vibe Code Bench is a tie in practice, since all four models score above 89%.&lt;/p&gt;

&lt;p&gt;Second, the harness matters. On CWE-bench v1, the OpenAI and Anthropic models ran in their own agent harnesses, Codex and Claude Code, so that score reflects model plus tooling. And &lt;a href="https://artificialanalysis.ai/articles/gemini-4-argon-google-top-three-labs" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt; ran Terminal Bench 4 separately and put Argon at 57%, behind Claude Sonnet 5.5 (64%), Opus 5.5 (60%), and GPT-6 Astra (59%).&lt;/p&gt;

&lt;p&gt;Third, the biggest gaps are in knowledge work and long context. The legal result looks huge next to Fable 5.1, but 19.6% still means about one task in five completed.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it will really cost you
&lt;/h2&gt;

&lt;p&gt;Price per token is the wrong number to compare. Artificial Analysis measures cost per task, and that's where Argon's chattiness shows up.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Cost per Intelligence Index task&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 4 Argon, launch price ($2 in, $10 out)&lt;/td&gt;
&lt;td&gt;$1.99&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 4 Argon, full price ($4 in, $20 out), my estimate&lt;/td&gt;
&lt;td&gt;about $3.98&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra, max effort&lt;/td&gt;
&lt;td&gt;$3.26&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first and last rows are Artificial Analysis's figures. The middle row is mine. The launch discount is a 50% promotion, so if Argon keeps using the same number of tokens, doubling the price doubles the cost. Astra's price could change too, so treat that row as a rough guide, not a forecast.&lt;/p&gt;

&lt;p&gt;A few other anchors. Opus 5.5 charges $20 per million output tokens, the same as Argon's post-launch rate. GPT-6.1 Sol's newly discounted price matches Argon's launch price. One maxed-out 1M-token response costs $10 now and $20 later. Cached input tokens get 95% off the input price, which should help if you resend the same repo-sized prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to check before you switch
&lt;/h2&gt;

&lt;p&gt;I didn't find any migration notes in Google's announcement or the coverage, so this isn't a changelog of breaking changes. It's a checklist of what the launch details imply.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You can't switch yet.&lt;/strong&gt; Access is gated, and the model ships without cyber guardrails only for Fairwind participants and Google's own teams. Google is still tuning the safeguards before a wider release, so security-heavy prompts may behave differently once you get access.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Revisit your output settings.&lt;/strong&gt; The limit went from 64K to 1M tokens. If you cap max output, set timeouts, or stream into something with size limits, those choices were made for 64K. Raise the ceiling on purpose, not by default.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expect a chattier model.&lt;/strong&gt; Artificial Analysis measured 62 K output tokens per task for Argon against 27 K for GPT-6 Astra at max effort. Watch latency and bills, not just quality.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan for the price change.&lt;/strong&gt; $2 and $10 per million tokens is an introductory rate. It becomes $4 and $20.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What Google has built with it
&lt;/h2&gt;

&lt;p&gt;Everything below comes from &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/" rel="noopener noreferrer"&gt;Google's own announcement&lt;/a&gt; and describes internal use. Read it as a demo reel, not a benchmark.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;C/C++ to Rust migrations.&lt;/strong&gt; Argon agents are porting code across Google, from tens of thousands of lines in libraries like re2 and libgav1 up to 800K+ lines for the Fuchsia Zircon kernel. Google says these rewrites are still going through automated and manual audits, emulation testing, and review before production.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;libgav1.&lt;/strong&gt; Starting from an existing Rust port, agents replaced 32K lines of SIMD code with safe Rust that the compiler vectorizes on its own. Google reports a memory-safe video decoder that runs 2.7x faster than the port, with identical video output.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data center memory.&lt;/strong&gt; A team of agents analyzed fleet-wide profiling data and applied memory optimizations, freeing over 300 TiB once rolled out. Google estimates 500 TiB to 1 PiB in total.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quantum.&lt;/strong&gt; Argon beat a published baseline by 40% on the qubits-times-gates cost of one subroutine, in a matter of minutes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security.&lt;/strong&gt; Wiz, which Google acquired in March, is using Argon in its Scan for Good program. Google says it found a critical flaw in healthcare software used by hospitals worldwide that earlier frontier models missed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The libgav1 story is the one worth stealing. The loop is profile, run an experiment, read what the compiler produced, repeat for many rounds. You can run that loop with any capable model today. Google's claim is that Argon stays on it without losing the thread.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it still loses
&lt;/h2&gt;

&lt;p&gt;This is the part the launch post skips.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;FrontierSWE v2.&lt;/strong&gt; Argon scores 55.0% to GPT-6 Astra's 65.5%, and it's last of the four models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Terminal-bench 4.0.&lt;/strong&gt; Opus 5.5 leads with 66.4% against Argon's 57.4%, a nine-point gap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Other agentic and science tests.&lt;/strong&gt; Astra beats it on Terminal-Bench Science 0.1 (68.1% to 57.6%) and OSWorld-2.0 (72.6% to 69.2%). Opus 5.5 beats it on PostTrainBench (49.3% to 45.3%).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cyber comparisons.&lt;/strong&gt; Google compares its security results only with its own Gemini 3.8 Flash Cyber: 85.8% against 71.0% on its internal vulnerability benchmark, and 70.9% against 58.2% on Wiz's penetration testing benchmark. That shows progress, not rank.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Independent testing.&lt;/strong&gt; Very few people outside Google have had hands-on time. Bloomberg reported before launch that some people inside Google worry Argon isn't as strong as the Anthropic and OpenAI models. Google disputes that and says employees have been testing versions for weeks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost per task.&lt;/strong&gt; Once the launch discount ends, Argon may not be the cheap option (see above).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One more loose end. Google's post doesn't state a context window, and one outlet reports 2M tokens. I'd wait for the docs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrapping up
&lt;/h2&gt;

&lt;p&gt;Argon looks strongest where the work is long and document-heavy: finance, legal, automation, big inputs, big outputs. It looks weaker in shell-driven agent work, where GPT-6 Astra and Opus 5.5 still hold the lead in Google's own table.&lt;/p&gt;

&lt;p&gt;The public benchmarks also disagree with each other. DeepSWE says one thing, FrontierSWE and Terminal-bench say another. So don't switch on a headline number. Pull 20 or so real tickets from your own repo, run them through whatever you use today, and run them through Argon once you have access.&lt;/p&gt;

&lt;p&gt;When you do, log two things: cost per task and whether a long output is still coherent at the end. Neither has been independently checked yet.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>gemini</category>
      <category>google</category>
      <category>llm</category>
    </item>
    <item>
      <title>How to Get Your Tool Recommended by ChatGPT, Perplexity, and Google</title>
      <dc:creator>Nancy Garg</dc:creator>
      <pubDate>Tue, 06 Oct 2026 19:10:55 +0000</pubDate>
      <link>https://dev.to/studio1hq/how-to-get-your-tool-recommended-by-chatgpt-perplexity-and-google-2dg9</link>
      <guid>https://dev.to/studio1hq/how-to-get-your-tool-recommended-by-chatgpt-perplexity-and-google-2dg9</guid>
      <description>&lt;h2&gt;
  
  
  &lt;strong&gt;TL;DR&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Getting your tool recommended by AI search isn't just about traditional SEO. You need to build the signals that ChatGPT, Perplexity, and Google use to understand, trust, and surface your product.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI recommendations depend on more than keywords. Clear positioning, authoritative content, technical depth, and consistent mentions all matter.&lt;/li&gt;
&lt;li&gt;Your documentation and technical content need to answer the exact questions developers are asking before they ever reach your website.&lt;/li&gt;
&lt;li&gt;Being mentioned across trusted websites, communities, GitHub, Reddit, and developer publications can strengthen your product's discoverability.&lt;/li&gt;
&lt;li&gt;The goal isn't to game AI search. It's to make your tool genuinely useful, easy to understand, and easy for AI systems to discover and cite.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A lot of developers nowadays ask ChatGPT before they ask Google. In &lt;a href="https://survey.stackoverflow.co/2025/ai" rel="noopener noreferrer"&gt;&lt;strong&gt;Stack Overflow's 2025 survey&lt;/strong&gt;&lt;/a&gt;, 47.1% said they use AI tools every day, and another 17.7% use them every week. That adds up to about 65%. And when someone asks which tool to use, they don't get ten blue links. They get a short list of names, with a line or two on each.&lt;/p&gt;

&lt;p&gt;So the question for any developer tool is how to land on that list. Ranking well on Google used to be most of the answer. It covers less now: &lt;a href="https://ahrefs.com/blog/ai-overview-citations-top-10/" rel="noopener noreferrer"&gt;&lt;strong&gt;Ahrefs found&lt;/strong&gt;&lt;/a&gt; that 76% of AI Overview citations came from pages in Google's top 10 in July 2025, and by January 2026 only 38% did. &lt;a href="https://developers.google.com/search/docs/fundamentals/ai-optimization-guide" rel="noopener noreferrer"&gt;&lt;strong&gt;Google says&lt;/strong&gt;&lt;/a&gt; this is still SEO at heart. But studies of what AI tools cite, like &lt;a href="https://searchengineland.com/ai-search-engines-cite-reddit-youtube-and-linkedin-most-study-473138" rel="noopener noreferrer"&gt;&lt;strong&gt;this one from Peec AI&lt;/strong&gt;&lt;/a&gt;, show that being talked about on Reddit, YouTube, and review sites matters too. This guide covers both sides.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;What is answer engine optimization (AEO)?&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Answer engine optimization, or AEO, is the work of making your product easy for AI tools to find, understand, and name when someone asks a question. You will also see it called GEO (generative engine optimization) or LLM SEO. These all mean roughly the same thing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhmmjpzt63mflc509d5pu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhmmjpzt63mflc509d5pu.png" alt="answer-by-llm" width="800" height="484"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The goal is different from classic SEO. SEO tries to get your link ranked on a results page. AEO tries to get your product's name inside the answer itself.&lt;/p&gt;

&lt;p&gt;Google sees it a little differently. Its own &lt;a href="https://developers.google.com/search/docs/fundamentals/ai-optimization-guide" rel="noopener noreferrer"&gt;&lt;strong&gt;guide to generative AI search&lt;/strong&gt;&lt;/a&gt; says that optimizing for these features is still SEO, because AI Overviews and AI Mode are built on its normal ranking and quality systems. When a question comes in, those features can run several related searches at once, which Google calls query fan-out, and then pull from the pages they find.&lt;/p&gt;

&lt;p&gt;Two things follow from that. First, basic SEO still matters, because the AI has to find your page before it can use it. Second, ranking is not the whole story. In an &lt;a href="https://ahrefs.com/blog/chatgpts-most-cited-pages/" rel="noopener noreferrer"&gt;&lt;strong&gt;Ahrefs study of the 1,000 pages ChatGPT cited most in September 2025&lt;/strong&gt;&lt;/a&gt;, 28% had no organic search keywords at all. A page can be quoted by an AI tool without ranking anywhere in Google.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;How do ChatGPT, Perplexity, and Google AI Overviews choose what to recommend?&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Each one finds pages in its own way, so your first move is a little different for each.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Engine&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;How it finds your pages&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;What studies say it leans on&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Your first move&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ChatGPT&lt;/td&gt;
&lt;td&gt;Its own search crawler, OAI-SearchBot&lt;/td&gt;
&lt;td&gt;Wikipedia, Reddit, and editorial sites like Forbes&lt;/td&gt;
&lt;td&gt;Allow OAI-SearchBot, and make your homepage say plainly what the product does&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Perplexity&lt;/td&gt;
&lt;td&gt;PerplexityBot, plus Perplexity-User when a person asks a question&lt;/td&gt;
&lt;td&gt;Reddit, LinkedIn, and G2 for business software questions&lt;/td&gt;
&lt;td&gt;Allow PerplexityBot, and earn real reviews and discussions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google AI Overviews and AI Mode&lt;/td&gt;
&lt;td&gt;Google's normal search index&lt;/td&gt;
&lt;td&gt;Pages that are indexed and can show a snippet. Google lists no extra requirements&lt;/td&gt;
&lt;td&gt;Get basic SEO right, and watch the AI report in Search Console&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sources for the table: &lt;a href="https://developers.openai.com/api/docs/bots" rel="noopener noreferrer"&gt;&lt;strong&gt;OpenAI's crawler page&lt;/strong&gt;&lt;/a&gt;, &lt;a href="https://docs.perplexity.ai/docs/resources/perplexity-crawlers" rel="noopener noreferrer"&gt;&lt;strong&gt;Perplexity's crawler page&lt;/strong&gt;&lt;/a&gt;, &lt;a href="https://developers.google.com/search/docs/appearance/ai-features" rel="noopener noreferrer"&gt;&lt;strong&gt;Google's AI features page&lt;/strong&gt;&lt;/a&gt;, and the &lt;a href="https://searchengineland.com/ai-search-engines-cite-reddit-youtube-and-linkedin-most-study-473138" rel="noopener noreferrer"&gt;&lt;strong&gt;Peec AI study of 30 million cited sources&lt;/strong&gt;&lt;/a&gt; as reported by Search Engine Land.&lt;/p&gt;

&lt;p&gt;Two patterns show up across the studies.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Community sites carry a lot of weight.&lt;/strong&gt; Peec AI's analysis found Reddit was the most cited domain across ChatGPT, Google AI Mode, Gemini, Perplexity, and AI Overviews. YouTube, LinkedIn, Wikipedia, and Forbes also made the top five. Review sites like G2 and Yelp appeared often when people asked for recommendations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Most of what ChatGPT cites is not yours to edit.&lt;/strong&gt; In Ahrefs look at ChatGPT's top 1,000 cited pages, Wikipedia made up 29.7%, homepages and landing pages 23.8%, and how-to and explainer pages 19.4%. Ahrefs counted only 32.3% as pages a company could realistically get into: explainers, reviews, news, and blog posts. Ahrefs says its sorting of page types is a rough guide, so treat the exact split loosely. The useful part is the homepage number. You cannot pitch your way onto someone else's homepage, but your own is a page you fully control.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsvz060s7j6iehe76w2zc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsvz060s7j6iehe76w2zc.png" alt="Stats" width="701" height="357"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;How to get your product recommended by AI: an 8 step playbook for developers&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Work through these in order. The first four are about your own site. The next four are about the rest of the web.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;1. Let the right crawlers in&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;If a crawler cannot read your site, nothing else in this guide matters. OpenAI says that to appear in ChatGPT search results, you should &lt;a href="https://help.openai.com/en/articles/12627856-publishers-and-developers-faq" rel="noopener noreferrer"&gt;&lt;strong&gt;allow OAI-SearchBot&lt;/strong&gt;&lt;/a&gt;. Perplexity says the same about &lt;a href="https://docs.perplexity.ai/docs/resources/perplexity-crawlers" rel="noopener noreferrer"&gt;&lt;strong&gt;PerplexityBot&lt;/strong&gt;&lt;/a&gt;. Google's &lt;a href="https://developers.google.com/search/docs/appearance/ai-features" rel="noopener noreferrer"&gt;&lt;strong&gt;AI features page&lt;/strong&gt;&lt;/a&gt; says crawling has to be allowed in robots.txt and by any CDN or hosting setup you use.&lt;/p&gt;

&lt;p&gt;A simple robots.txt that allows all three looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight jsx"&gt;&lt;code&gt;&lt;span class="nx"&gt;User&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;OAI&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;SearchBot&lt;/span&gt;
&lt;span class="nx"&gt;Allow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt;

&lt;span class="nx"&gt;User&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;PerplexityBot&lt;/span&gt;
&lt;span class="nx"&gt;Allow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt;

&lt;span class="nx"&gt;User&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Googlebot&lt;/span&gt;
&lt;span class="nx"&gt;Allow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four details trip up a lot of teams:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPTBot is a separate bot. OpenAI treats it as its own setting, used for training. You can block GPTBot and still allow OAI-SearchBot, so blocking one does not have to hide you from ChatGPT search.&lt;/li&gt;
&lt;li&gt;Your firewall can block them even when robots.txt says yes. Perplexity asks you to allow its published IP ranges and walks through the setup for Cloudflare and AWS.&lt;/li&gt;
&lt;li&gt;Changes are not instant. Perplexity says it can take up to 24 hours for a robots.txt change to show up.&lt;/li&gt;
&lt;li&gt;Docs that only appear after JavaScript runs, or that live inside images or videos, are easy to miss. Google says important content should be available as text. A quick test: open your docs with JavaScript turned off and see what is left.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;2. Put the answer in the first few lines&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Start every page with a plain answer to the question it exists for. Search Engine Land &lt;a href="https://searchengineland.com/chatgpt-citations-content-study-469483" rel="noopener noreferrer"&gt;&lt;strong&gt;reported a study&lt;/strong&gt;&lt;/a&gt; finding that 44% of ChatGPT citations come from the first third of a page's content. If your best sentence is in paragraph six, it may never get used.&lt;/p&gt;

&lt;p&gt;Here is the difference. A weak opening spends three paragraphs on how fast software is changing. A strong one looks like this (ExampleDB is made up for this guide):&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;ExampleDB is a hosted Postgres database for small teams. It costs $15 a month, works with Prisma and Drizzle, and takes about five minutes to set up.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Then write your headings as the questions people really type. "How do I connect ExampleDB to Next.js?" beats "Connecting". Keep paragraphs short, and put the answer first before the explanation.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;3. Publish the comparison pages people ask for&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Recommendation questions are comparison questions: "best tool for X", "X vs Y", "alternatives to Y". If you do not have a page for those, someone else writes it, and AI tools quote them instead.&lt;/p&gt;

&lt;p&gt;Write these pages yourself, and write them fairly. Put pricing in a table. Name the limits. Say who each tool suits, and where the other tool wins. A fair page is more useful to readers and easier to quote. If you run benchmarks, say how: the machine, the version, the date, and the code.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://developers.google.com/search/docs/fundamentals/ai-optimization-guide" rel="noopener noreferrer"&gt;&lt;strong&gt;Google's guide to generative AI search&lt;/strong&gt;&lt;/a&gt; describes the content that holds up best as non-commodity: material built on first-hand experience that goes beyond what anyone could copy from other pages. Your own tests and your own numbers are exactly that.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;4. Make your docs the best answer on the web&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;For developer tools, docs do the job that blog posts do for other products. Developers ask "how do I add X to Next.js" or paste an error message straight into a chat. Cover those questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A quickstart that works when copied, with version numbers.&lt;/li&gt;
&lt;li&gt;One page per common error, with the exact error text in the heading, the cause, and the fix.&lt;/li&gt;
&lt;li&gt;Migration guides from the tools people are leaving for yours.&lt;/li&gt;
&lt;li&gt;Pricing, limits, and supported languages written out as text, not tucked into an image or a pricing widget.&lt;/li&gt;
&lt;li&gt;Everything public. Docs behind a login cannot be read by a crawler.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;5. Get talked about where AI tools already look&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;This is the step most teams skip, and the data says it matters most. &lt;a href="https://ahrefs.com/blog/ai-overview-brand-correlation/" rel="noopener noreferrer"&gt;&lt;strong&gt;Ahrefs studied 75,000 brands&lt;/strong&gt;&lt;/a&gt; and found that how often a brand is mentioned across the web had the strongest link to showing up in Google's AI Overviews. It scored 0.664, on a scale where 1 would be a perfect match. Backlinks scored 0.218. Brands in the top quarter for web mentions averaged 169 AI Overview mentions, against 14 for the next quarter down. Brands in the bottom half barely appeared at all.&lt;/p&gt;

&lt;p&gt;Two cautions. This is correlation, and Ahrefs says so itself. It also covered established brands (domain rating above 40), so for a new tool it shows direction, not a target.&lt;/p&gt;

&lt;p&gt;What to do about it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Answer questions on Reddit and Stack Overflow when your tool is a fair answer. Say that you work on it.&lt;/li&gt;
&lt;li&gt;Make your GitHub README clear enough to quote: what it is, the install command, and a minimal example.&lt;/li&gt;
&lt;li&gt;Record tutorial videos with real transcripts. YouTube was in the top five most cited domains in the Peec AI data.&lt;/li&gt;
&lt;li&gt;Get into roundups, newsletters, and community lists that cover your category.&lt;/li&gt;
&lt;li&gt;Ask happy users for honest reviews on sites like G2, which showed up often for recommendation questions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not fake any of it. &lt;a href="https://developers.google.com/search/docs/fundamentals/ai-optimization-guide" rel="noopener noreferrer"&gt;&lt;strong&gt;Google's guide&lt;/strong&gt;&lt;/a&gt; says chasing inauthentic mentions is not as helpful as it looks, and that its systems are built to catch spam.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;6. Say the same thing everywhere&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Use the same product name, one-line description, price, and supported languages on your site, your GitHub page, your package pages (npm, PyPI, and so on), LinkedIn, and your docs. This one comes from common sense, not a study. If an AI tool reads three different prices, it may quote the wrong one or skip you.&lt;/p&gt;

&lt;p&gt;A short facts page helps: name, what it does, pricing, languages, license, and a last updated date. Every other page can then match it.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;7. Keep key pages fresh, and be honest about dates&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;AI tools lean toward newer pages. In the same &lt;a href="https://ahrefs.com/blog/chatgpts-most-cited-pages/" rel="noopener noreferrer"&gt;&lt;strong&gt;Ahrefs look at ChatGPT's top 1,000 cited pages&lt;/strong&gt;&lt;/a&gt;, Ahrefs could find an update date for just over half. About 90% of those had been updated in 2025, and with Wikipedia removed, it was still about 82%. The sample is small, but the direction is clear.&lt;/p&gt;

&lt;p&gt;So update pricing, version numbers, and screenshots when they change, and show a real last updated date. Change that date only when you changechanged the page.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;8. Schema markup and llms.txt: nice to have, not a fix&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://developers.google.com/search/docs/fundamentals/ai-optimization-guide" rel="noopener noreferrer"&gt;&lt;strong&gt;Google says&lt;/strong&gt;&lt;/a&gt; there is no special markup you need for AI Overviews, and that structured data should match the text people can see on the page. If your site setup makes FAQ or software markup easy, add it. Do not spend a sprint on it.&lt;/p&gt;

&lt;p&gt;The same goes for llms.txt. Google says its search &lt;a href="https://developers.google.com/search/docs/fundamentals/ai-optimization-guide" rel="noopener noreferrer"&gt;&lt;strong&gt;ignores those files&lt;/strong&gt;&lt;/a&gt;, so they neither help nor hurt you there. It also says it is fine to keep one for other services that use them. Treat it as optional, not a fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Can programmatic pages help your product get recommended by AI?&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Yes, but only when every page carries real information. Programmatic SEO means building many pages from one template and a set of data. For developer tools that can be a very good fit, because so many developer questions follow a pattern: "X with Next.js", "X vs Y", "error 502 in X".&lt;/p&gt;

&lt;p&gt;The risk is real, though. Google's guide warns that making separate pages for every variation of a question, mainly to manipulate rankings or AI answers, &lt;a href="https://developers.google.com/search/docs/fundamentals/ai-optimization-guide" rel="noopener noreferrer"&gt;&lt;strong&gt;violates its scaled content abuse policy&lt;/strong&gt;&lt;/a&gt;. It also says a high number of pages does not make a site better. A template with the product name swapped in is the exact thing it describes.&lt;/p&gt;

&lt;p&gt;Here are page types that work, and the real data each one needs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Page type&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Question it answers&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Real data every page needs&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Integration pages ("X with Next.js")&lt;/td&gt;
&lt;td&gt;How do I use X with this tool?&lt;/td&gt;
&lt;td&gt;Code you have tested, version numbers, known issues&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Comparison pages ("X vs Y")&lt;/td&gt;
&lt;td&gt;Which one should I pick?&lt;/td&gt;
&lt;td&gt;Pricing, limits, benchmarks with the method shown, where the other tool wins&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Error pages&lt;/td&gt;
&lt;td&gt;What does this error mean in X?&lt;/td&gt;
&lt;td&gt;The exact error text, the cause, a fix that works&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Use case pages ("X for job queues")&lt;/td&gt;
&lt;td&gt;Is X good for my situation?&lt;/td&gt;
&lt;td&gt;A real example project or customer story&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Language pages ("X in Go")&lt;/td&gt;
&lt;td&gt;How do I use X in my language?&lt;/td&gt;
&lt;td&gt;SDK version, a working snippet, language specific limits&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A simple rule: start with 10 to 20 pages you would be proud to show a customer, not 10,000. Before you publish a page, check four things. Does it answer a question someone really asks? Does it have code or numbers you verified yourself? Does it say something the other pages on the web do not? Would a developer who lands on it leave with what they came for? If any answer is no, fix the data first. Do not publish the page yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;How do you know if AI tools are recommending your product?&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;You can check this by hand in under an hour a month. Here is a simple routine:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Write down 20 questions real developers ask about your category. Pull them from support tickets, Reddit threads, and your Search Console queries.&lt;/li&gt;
&lt;li&gt;Once a month, ask each question in ChatGPT, Perplexity, and Google (including AI Mode). Use the same wording every time. Answers can change from one ask to the next, so ask the important ones a few times.&lt;/li&gt;
&lt;li&gt;For each answer, record four things: were you named, were you linked, which competitors were named, and which sites were cited.&lt;/li&gt;
&lt;li&gt;Look at the cited sites. Those are the pages shaping answers in your category, and they are your outreach list for step 5.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsg4f28zd2yuxk8m0lib3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsg4f28zd2yuxk8m0lib3.png" alt="llm" width="703" height="225"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then add the numbers your tools already give you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ChatGPT referrals.&lt;/strong&gt; ChatGPT adds &lt;strong&gt;&lt;code&gt;utm_source=chatgpt.com&lt;/code&gt;&lt;/strong&gt; to links, so you can &lt;a href="https://help.openai.com/en/articles/12627856-publishers-and-developers-faq" rel="noopener noreferrer"&gt;**filter for it in&lt;/a&gt;**&amp;nbsp;any analytics tool, like Google Analytics or &lt;a href="https://raah.dev/" rel="noopener noreferrer"&gt;raah.dev&lt;/a&gt;, which tracks both web analytics and web observability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Perplexity referrals.&lt;/strong&gt; Look for perplexity.ai in your referrer reports.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google.&lt;/strong&gt; Search Console has a &lt;a href="https://developers.google.com/search/docs/fundamentals/ai-optimization-guide" rel="noopener noreferrer"&gt;&lt;strong&gt;Generative AI performance report&lt;/strong&gt;&lt;/a&gt;. Google also warns that no outside tool has access to its internal ranking or AI systems, so be careful with tools that claim otherwise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A free check.&lt;/strong&gt; Ahrefs offers a &lt;a href="https://ahrefs.com/ai-overviews-tracker" rel="noopener noreferrer"&gt;&lt;strong&gt;free AI Overviews tracker&lt;/strong&gt;&lt;/a&gt; that shows how often AI Overviews mention your brand and which sites they cite.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Expect slow movement. Crawler access can change within a day or so. Mentions across the web take much longer to build, so judge your progress by the quarter, not the week.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;What mistakes keep developer products out of AI answers?&lt;/strong&gt;
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Blocking every AI bot.&lt;/strong&gt; Many teams block all of them to stop training and end up hiding from search too. Keep OAI-SearchBot and PerplexityBot allowed even if you block GPTBot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Publishing hundreds of thin pages.&lt;/strong&gt; Google treats this as spam, and volume does not make a site better.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Buying or faking mentions and reviews.&lt;/strong&gt; Google says it is not as helpful as it seems, and fake reviews can hurt real trust.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hiding docs.&lt;/strong&gt; Login walls, JavaScript-only pages, and text inside images all make your best material hard to read.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rewriting everything for robots.&lt;/strong&gt; Google says you do not need to chop content into tiny pieces or write in a special style for AI. Write for the developer reading it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Judging by one test.&lt;/strong&gt; One answer on one day tells you very little. Track a fixed set of questions over months.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Frequently asked questions&lt;/strong&gt;
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;How do I get my product recommended by ChatGPT?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Allow OAI-SearchBot in your robots.txt, put a clear answer at the top of your homepage and docs, and get your product mentioned on sites AI tools trust, like Reddit and review platforms. &lt;a href="https://developers.openai.com/api/docs/bots" rel="noopener noreferrer"&gt;&lt;strong&gt;OpenAI says&lt;/strong&gt;&lt;/a&gt; sites that opt out of OAI-SearchBot will not be shown in ChatGPT search answers, so access is the first thing to check.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;How long does it take to show up in AI answers?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Nobody can promise a date. &lt;a href="https://docs.perplexity.ai/docs/resources/perplexity-crawlers" rel="noopener noreferrer"&gt;&lt;strong&gt;Perplexity says&lt;/strong&gt;&lt;/a&gt; a robots.txt change can take up to 24 hours to show up, but building mentions across the web is slow. Plan in quarters, not weeks.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Do backlinks still matter for AI visibility?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;They help, but less than mentions in the &lt;a href="https://ahrefs.com/blog/ai-overview-brand-correlation/" rel="noopener noreferrer"&gt;&lt;strong&gt;Ahrefs study of 75,000 brands&lt;/strong&gt;&lt;/a&gt;: 0.218 for backlinks against 0.664 for brand mentions in AI Overviews. In &lt;a href="https://ahrefs.com/blog/chatgpts-most-cited-pages/" rel="noopener noreferrer"&gt;&lt;strong&gt;Ahrefs' look at ChatGPT's most cited pages&lt;/strong&gt;&lt;/a&gt;, the ones that did rank mostly sat on strong sites, with 65.3% on domains rated 81 or higher. Both findings are correlations, so read them as hints.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Do I need an llms.txt file?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Not for Google, which &lt;a href="https://developers.google.com/search/docs/fundamentals/ai-optimization-guide" rel="noopener noreferrer"&gt;&lt;strong&gt;says its search ignores those files&lt;/strong&gt;&lt;/a&gt;. Adding one will not hurt, and other services may use it, but it should not come before the basics in this guide.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Do I need schema markup?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://developers.google.com/search/docs/appearance/ai-features" rel="noopener noreferrer"&gt;&lt;strong&gt;Google says&lt;/strong&gt;&lt;/a&gt; no special markup is required for AI Overviews or AI Mode. If it is easy to add, keep it matched to the visible text on the page. Do not expect it to decide your results.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Should I block GPTBot?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;That is your call. &lt;a href="https://developers.openai.com/api/docs/bots" rel="noopener noreferrer"&gt;&lt;strong&gt;OpenAI treats&lt;/strong&gt;&lt;/a&gt; GPTBot, used for training, and OAI-SearchBot, used for search, as separate settings. You can block the first and still allow the second.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Is AEO different from SEO?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Mostly, it overlaps. &lt;a href="https://developers.google.com/search/docs/fundamentals/ai-optimization-guide" rel="noopener noreferrer"&gt;&lt;strong&gt;Google calls&lt;/strong&gt;&lt;/a&gt; optimizing for generative AI search a form of SEO. The extra work is off your site: getting discussed in trusted places, and making each page answer its question clearly in the first few lines.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Claude Opus 5.5: What Changed, What Breaks, and Whether to Switch</title>
      <dc:creator>Nancy Garg</dc:creator>
      <pubDate>Tue, 06 Oct 2026 19:03:33 +0000</pubDate>
      <link>https://dev.to/studio1hq/claude-opus-55-what-changed-what-breaks-and-whether-to-switch-3oa5</link>
      <guid>https://dev.to/studio1hq/claude-opus-55-what-changed-what-breaks-and-whether-to-switch-3oa5</guid>
      <description>&lt;p&gt;In September 2026, Anthropic's CEO Dario Amodei &lt;a href="https://darioamodei.com/post/we-must-pace-the-frontier" rel="noopener noreferrer"&gt;published an essay&lt;/a&gt; with this line in it: "We must slow the pace at which we improve the capabilities of AI models."&lt;/p&gt;

&lt;p&gt;A week after, on September 22, 2026, Anthropic &lt;a href="https://www.anthropic.com/claude-opus-5-5" rel="noopener noreferrer"&gt;released Claude Opus 5.5&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;It landed just three weeks after Fable 5.1, so people had jokes. But if you write code for a living, the actual questions are something else. What does it cost you, what will it break, and is it better at your actual work?&lt;/p&gt;

&lt;h3&gt;
  
  
  TL;DR
&lt;/h3&gt;

&lt;p&gt;If you want the quick take, Opus 5.5 is what Opus 5 should have shipped as, and Anthropic finally listened.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It doesn't ramble anymore.&lt;/strong&gt; Opus 5 used to pad every answer with fluff. This one gets to the point, and it's about 30% quicker to do so.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It's cheaper, and the savings actually hold up.&lt;/strong&gt; $4 in and $20 out per million tokens, down from $5 and $25. Just don't crank the thinking effort all the way up, or the savings disappear.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It comes close to Fable 5.1 on coding work,&lt;/strong&gt; which is Anthropic's priciest model, at less than half the cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A few things break if you're already using Opus 5.&lt;/strong&gt; Nothing that ruins your day, but enough that you can't just flip the switch and forget about it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're on Opus 5 right now, this is worth moving to. Try it on your own work first though, the real difference only shows up once you compare it against what you're actually building, not the benchmark charts.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is Opus 5.5?
&lt;/h3&gt;

&lt;p&gt;Claude has 3 families of Models. Haiku, Sonnet, and Opus. Then there's Fable, a separate tier above all three that costs the most.&lt;/p&gt;

&lt;p&gt;Opus 5.5 is built for agents that run for hours. It has a 1M token context window and a 128K token output limit. Its knowledge runs up to June 2026.&lt;/p&gt;

&lt;p&gt;It's the first model in a new 5.5 family. Sonnet 5.5 and Haiku 5.5 are due in the next few weeks.&lt;/p&gt;

&lt;h3&gt;
  
  
  How it differs from the other models
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Opus 5.5&lt;/th&gt;
&lt;th&gt;Opus 5&lt;/th&gt;
&lt;th&gt;Fable 5.1&lt;/th&gt;
&lt;th&gt;Sonnet 5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;API name&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;claude-opus-5-5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;claude-opus-5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;claude-fable-5-1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;claude-sonnet-5&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Input / output price&lt;/strong&gt; (per 1M tokens)&lt;/td&gt;
&lt;td&gt;$4 / $20&lt;/td&gt;
&lt;td&gt;$5 / $25&lt;/td&gt;
&lt;td&gt;$10 / $50&lt;/td&gt;
&lt;td&gt;$2 / $10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Cache reads&lt;/strong&gt; (per 1M tokens)&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;td&gt;not listed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Speed&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;About 30% slower than 5.5&lt;/td&gt;
&lt;td&gt;Slower&lt;/td&gt;
&lt;td&gt;Fast&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Context / max output&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1M / 128K&lt;/td&gt;
&lt;td&gt;1M / not listed&lt;/td&gt;
&lt;td&gt;1M / 128K&lt;/td&gt;
&lt;td&gt;1M / 128K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Safety filters&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Cyber and biology&lt;/td&gt;
&lt;td&gt;Cyber&lt;/td&gt;
&lt;td&gt;Cyber and biology&lt;/td&gt;
&lt;td&gt;not listed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Terminal-Bench 4.0&lt;/strong&gt; (coding)&lt;/td&gt;
&lt;td&gt;66.4%&lt;/td&gt;
&lt;td&gt;52.3%&lt;/td&gt;
&lt;td&gt;55.8%&lt;/td&gt;
&lt;td&gt;not tested&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Best for&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Long coding jobs, big refactors, reports&lt;/td&gt;
&lt;td&gt;Same jobs, wordier&lt;/td&gt;
&lt;td&gt;The hardest reasoning&lt;/td&gt;
&lt;td&gt;Quick, simple, high-volume work&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Compared with Opus 5&lt;/strong&gt;, the biggest change is how it talks. Opus 5 wrote long, twisty answers and did extra work nobody asked for. Opus 5.5 puts the main point first, writes about 30% faster, and costs less, with cache reads down from $0.50 to $0.20. That last part adds up when an agent re-reads the same files all day.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Against Fable 5.1&lt;/strong&gt;, it holds its own for a lot less money. It costs 60% less and beats Fable on most coding tests, including Terminal-Bench 4.0 (66.4% to 55.8%). Fable is still the one to move up to when the hardest problems get past Opus.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compared to Sonnet 5&lt;/strong&gt;, it's the heavier tool. Sonnet is half the price and faster, so it's the better pick for small edits, summaries, and routing. Reach for Opus 5.5 when the job is long or tricky.&lt;/p&gt;

&lt;h3&gt;
  
  
  What breaks when you switch
&lt;/h3&gt;

&lt;p&gt;Some requests that worked on Opus 5 now return errors. The &lt;a href="https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5" rel="noopener noreferrer"&gt;full list is in the docs&lt;/a&gt;. These are the four you'll actually hit:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Thinking can't be turned off.&lt;/strong&gt; Sending &lt;code&gt;thinking: {"type": "disabled"}&lt;/code&gt; or a manual token budget returns a 400. Leave it out, or use adaptive, and control depth with &lt;code&gt;effort&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Forced tool use is gone.&lt;/strong&gt; &lt;code&gt;tool_choice&lt;/code&gt; set to &lt;code&gt;any&lt;/code&gt; or a named tool returns a 400. Use &lt;code&gt;auto&lt;/code&gt; with strict tool use, and say in the prompt when the tool should run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Progress messages can go quiet.&lt;/strong&gt; The short notes the model writes between tool calls now arrive as thinking blocks, which are empty by default. A UI that streams them will look frozen, with no error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The old computer use tool is rejected.&lt;/strong&gt; On the Claude API and Google Cloud, &lt;code&gt;computer_20251124&lt;/code&gt; returns a 400. Use &lt;code&gt;computer_toolset_20260801&lt;/code&gt;. Bedrock still accepts the old one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There's one more, and it only affects newer accounts. Thinking blocks are tied to the conversation, so editing anything earlier in the conversation before replaying them returns an error. Keep conversations append-only.&lt;/p&gt;

&lt;h3&gt;
  
  
  What it will really cost you
&lt;/h3&gt;

&lt;p&gt;Anthropic says the average job costs about &lt;a href="https://www.anthropic.com/claude-opus-5-5" rel="noopener noreferrer"&gt;40% less than on Opus 5&lt;/a&gt;. That holds at the default effort, which is now medium.&lt;/p&gt;

&lt;p&gt;At max effort it doesn't. An independent test found the model writes about 119,000 tokens per task there, against about 73,000 for Opus 5. The price cuts cancel that out, so the cost per task came out flat at roughly $6.&lt;/p&gt;

&lt;p&gt;It also thinks harder at each effort level than Opus 5 did. So don't copy your old setting over. Re-run a sweep on your own tasks and compare cost per finished task, not price per token.&lt;/p&gt;

&lt;p&gt;A few other prices, if you're budgeting:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Batch:&lt;/strong&gt; half price, $2 in and $10 out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fast mode:&lt;/strong&gt; up to 2.5x speed at $8 in and $40 out. Claude API only, and still a research preview.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI:&lt;/strong&gt; GPT-6 Astra is reported at $10 and $50, the same as Fable. The cheaper GPT-6 Sol costs far less per task but scored well below Opus 5.5 in the same test.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Test Scores
&lt;/h3&gt;

&lt;p&gt;Benchmarks are standard exams for models. These are the ones that matter most for dev work (best score in bold):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;th&gt;What it checks&lt;/th&gt;
&lt;th&gt;Opus 5.5&lt;/th&gt;
&lt;th&gt;Fable 5.1&lt;/th&gt;
&lt;th&gt;Opus 5&lt;/th&gt;
&lt;th&gt;GPT-6 Astra&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-Bench 4.0&lt;/td&gt;
&lt;td&gt;Multi-step tasks in a command line&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;66.4%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;55.8%&lt;/td&gt;
&lt;td&gt;52.3%&lt;/td&gt;
&lt;td&gt;57.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FrontierCode&lt;/td&gt;
&lt;td&gt;Would the code change get merged&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;54.4%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;50.3%&lt;/td&gt;
&lt;td&gt;48.0%&lt;/td&gt;
&lt;td&gt;53.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CursorBench 4.0&lt;/td&gt;
&lt;td&gt;Real tasks from Cursor users&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;57.8%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;51.8%&lt;/td&gt;
&lt;td&gt;46.6%&lt;/td&gt;
&lt;td&gt;not reported&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OSWorld 2.0&lt;/td&gt;
&lt;td&gt;Using a computer by clicking and typing&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;81.8%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;80.7%&lt;/td&gt;
&lt;td&gt;74.0%&lt;/td&gt;
&lt;td&gt;not reported&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AutomationBench&lt;/td&gt;
&lt;td&gt;Business tasks across many apps&lt;/td&gt;
&lt;td&gt;40.0%&lt;/td&gt;
&lt;td&gt;31.4%&lt;/td&gt;
&lt;td&gt;26.9%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;41.4%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-Bench-Science&lt;/td&gt;
&lt;td&gt;Science research tasks&lt;/td&gt;
&lt;td&gt;58.7%&lt;/td&gt;
&lt;td&gt;52.6%&lt;/td&gt;
&lt;td&gt;29.0%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;64.6%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are best-case scores, mostly from the maker's own runs. Even Anthropic says the real gap to Fable is smaller than the table looks.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fov69o75a30ok2b1ghdqg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fov69o75a30ok2b1ghdqg.png" alt="Article-preview" width="800" height="585"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What people have built with it
&lt;/h3&gt;

&lt;p&gt;These are some great use cases built by people who got their hands on Opus 5.5 in the first two days. Most came from a single prompt.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sketch to simulator.&lt;/strong&gt; &lt;a href="https://twitter.com/poolio/status/2102445641205248145" rel="noopener noreferrer"&gt;Ben Poole&lt;/a&gt;, formerly of Google Brain and DeepMind, drew a trebuchet on paper and sent Opus 5.5 a photo. It gave him a working 3D simulator where you change the weight and angle, fire at a stack of blocks, and replay the shot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Minecraft in a browser tab.&lt;/strong&gt; &lt;a href="https://twitter.com/noahwachnik/status/2102470200415166699" rel="noopener noreferrer"&gt;Noah Wachnik&lt;/a&gt; asked for a playable Minecraft with fancy shaders. At max effort it built a game called Lumen Vale in 1 hour 37 minutes. You can walk around, break and place blocks, and the water ripples while the lighting changes with the time of day.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pixel art wizard from a strict spec.&lt;/strong&gt; &lt;a href="https://twitter.com/majidmanzarpour/status/2102476258948927543" rel="noopener noreferrer"&gt;Majid&lt;/a&gt; wrote a tight prompt: one HTML file, 128x96 resolution, 24 colors, particle effects, and a state machine for idle, charge and cast. It followed the whole spec and runs at 60 fps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An 80-second film with no libraries.&lt;/strong&gt; &lt;a href="https://twitter.com/LCSlates/status/2102503027340988559" rel="noopener noreferrer"&gt;Chris Riley&lt;/a&gt; wanted a film in one HTML file using only WebGL2 and plain JavaScript. No libraries, no images, no audio files. The result is a glass mosaic where the fish and birds are made of moving tiles.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A Blender castle, with the bill.&lt;/strong&gt; &lt;a href="https://twitter.com/Stefan_3D_AI/status/2102471841046786153" rel="noopener noreferrer"&gt;Stefan Vaskevich&lt;/a&gt; had Opus 5.5 and GPT-6 Astra each build a 10-second Blender animation of a castle from one prompt. Opus took 35 minutes and cost about $13.30. Astra was faster at 28 minutes but cost $14.50. Stefan felt Opus handled more complexity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are demos people chose to post, so treat them as best cases. &lt;/p&gt;

&lt;h3&gt;
  
  
  Where it still loses
&lt;/h3&gt;

&lt;p&gt;Not everything's better. A few real gaps before you switch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Coding.&lt;/strong&gt; Sonar found it writes about 27% less code, with 42% fewer issues overall. But bugs per line went up 12%, and problems in concurrent code went up 44%. If your codebase leans on threads, locks, or async, don't skip the review.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Motion graphics and creative work.&lt;/strong&gt; Opus 5.5 is also being used for motion graphics and creative coding, with people generating animations from simple prompts. We have created &lt;a href="https://x.com/Astrodevil_/status/2103600734017392708?s=20" rel="noopener noreferrer"&gt;one ourselves&lt;/a&gt; using a single prompt. The model goes beyond code generation into actually building visual experiences.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Business tasks and science.&lt;/strong&gt; OpenAI's Astra still wins both. It beats Opus 5.5 on business tasks, and by a wider margin, on science.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hard reasoning.&lt;/strong&gt; Fable still wins here. If your evals fail at high effort, that's your next move, not a higher setting on Opus.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Speed.&lt;/strong&gt; It's slow to start. One test clocked about 22 seconds before the first word at medium effort. Fine for a background job, bad for a chat window someone's watching.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security and biology.&lt;/strong&gt; Some work gets rerouted without asking. Most security tasks, like finding exploits, go to the older Opus 4.8. Flagged biology requests go to Opus 5. The filters read your files and search results too, so something already sitting in your repo can trigger a reroute. Worth checking if you build anything security or bio adjacent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Simple, everyday tasks.&lt;/strong&gt; Sonnet 5 is still the better call here. Cheaper, faster, no reason to reach for Opus.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Wrapping up
&lt;/h3&gt;

&lt;p&gt;Opus 5.5 is basically Opus 5 with some parts fixed. Anthropic heard the complaints about Opus 5 and cheaper, faster, less bloated is the result.&lt;/p&gt;

&lt;p&gt;It's not for everyone. Security, biology, or low-latency work still needs a second look, and a few breaking changes mean this isn't a plug-and-play swap.&lt;/p&gt;

&lt;p&gt;Test it on your own tasks before you switch. That's the only benchmark that matters.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
      <category>programming</category>
    </item>
    <item>
      <title>What is Jev? In 5 minutes</title>
      <dc:creator>Nancy Garg</dc:creator>
      <pubDate>Tue, 06 Oct 2026 18:59:13 +0000</pubDate>
      <link>https://dev.to/studio1hq/what-is-jev-in-5-minutes-3jij</link>
      <guid>https://dev.to/studio1hq/what-is-jev-in-5-minutes-3jij</guid>
      <description>&lt;p&gt;&lt;a href="https://typesafe.ai/" rel="noopener noreferrer"&gt;TypeSafe&lt;/a&gt; AI released a new model called &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" rel="noopener noreferrer"&gt;Jev on September 15, 2026&lt;/a&gt;, and it works differently from most AI models you've probably used.&lt;/p&gt;

&lt;p&gt;Instead of writing you an answer, it picks one from a list you give it, and tells you how sure it is about that pick. TypeSafe has named it &lt;a href="https://docs.typesafe.ai/concepts/system-one" rel="noopener noreferrer"&gt;System One&lt;/a&gt; model.&lt;/p&gt;

&lt;p&gt;Most of the time when software uses AI to make a decision, like sorting a support ticket or flagging a transaction, it doesn't actually need a written answer, it need few outcomes.&lt;/p&gt;

&lt;p&gt;But an LLM throws a full paragraph and then you have to write code to pull out the actual decision. But Jev skips all of that and gives you the decision directly. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fttq88zgsfw066yeo211d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fttq88zgsfw066yeo211d.png" alt=" " width="800" height="165"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  TL;DR
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Jev takes a piece of text (&lt;code&gt;state&lt;/code&gt;) plus one or more typed questions, and returns a typed answer with calibrated probabilities. Never free-form text.&lt;/li&gt;
&lt;li&gt;Three answer types: &lt;strong&gt;Choice&lt;/strong&gt; (pick one of N options), &lt;strong&gt;Score&lt;/strong&gt; (rate on an ordered scale), &lt;strong&gt;Noul&lt;/strong&gt; (probability that a yes/no statement is true).&lt;/li&gt;
&lt;li&gt;It's fast (well under a second in most cases) and cheaper $0.042 per million input tokens, output is free.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What exactly is Jev?
&lt;/h2&gt;

&lt;p&gt;Jev is built by TypeSafe AI, started by &lt;a href="https://x.com/CompleteSkeptic/status/2099925682726002904?s=20" rel="noopener noreferrer"&gt;Diogo Almeida&lt;/a&gt;, the person who worked on ChatGPT and RLHF at OpenAI, along with Erik Gafni and Sasha Sheng.&lt;/p&gt;

&lt;p&gt;The way it works is pretty simple once you see it. You give it a piece of text, called the &lt;a href="https://docs.typesafe.ai/concepts/state" rel="noopener noreferrer"&gt;state&lt;/a&gt;, and then you ask it a question that has a fixed set of possible answers. It reads the text and answers the question, along with a number that says how confident it is. It gives a answer in 70 to 500 milliseconds. &lt;/p&gt;

&lt;p&gt;There are three &lt;a href="https://docs.typesafe.ai/primitives" rel="noopener noreferrer"&gt;kinds of questions it can answer&lt;/a&gt;. You can ask it to pick one option out of a list, give a score on a scale, or answer yes or no as a probability. And you can send it several questions about the same piece of text, and it answers all of them together in one go.&lt;/p&gt;

&lt;p&gt;A real example: asking it whether a support ticket needs a refund on a short message cost about $0.00002. Run that on a million similar tickets, and the total comes to around $19. &lt;/p&gt;

&lt;h2&gt;
  
  
  How Jev Is different
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://docs.typesafe.ai/models" rel="noopener noreferrer"&gt;Jev&lt;/a&gt; doesn’t replace them. They do different jobs. Here is a quick comparison between Jev and standard LLM&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;GPT / Claude&lt;/th&gt;
&lt;th&gt;Jev&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What it gives you&lt;/td&gt;
&lt;td&gt;Written text&lt;/td&gt;
&lt;td&gt;A picked answer with a confidence score&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;How fast&lt;/td&gt;
&lt;td&gt;A few seconds&lt;/td&gt;
&lt;td&gt;Under a second&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input cost&lt;/td&gt;
&lt;td&gt;Around $10 per million tokens&lt;/td&gt;
&lt;td&gt;$0.042 per million tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output cost&lt;/td&gt;
&lt;td&gt;You pay for every word&lt;/td&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Can it write essays or code&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Can it have a conversation&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Can it be wrong&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes, but it can never give an answer outside the options you gave it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last row is worth pausing on. Jev being unable to go outside the list of answers you gave it doesn't mean it's always right, it just means it can't get creative in a way that breaks your code. If you told it the answer has to be low, medium, or high, it will never hand you back something else, but it can still confidently pick medium when the real answer was low.&lt;/p&gt;

&lt;p&gt;TypeSafe's own number is 193x faster and 444x cheaper on their internal benchmark, self-run and openly flagged as the high end. &lt;/p&gt;

&lt;p&gt;One more practical detail: it can only read text, so no images or audio, and it can handle roughly 32,000 tokens per question, or about 64,000 tokens total once you count the input plus the question itself.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhoeukg6mkiktapog3ent.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhoeukg6mkiktapog3ent.png" alt=" " width="800" height="439"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: Claude thinks, Jev decides
&lt;/h2&gt;

&lt;p&gt;Once you actually sit with how Jev is &lt;a href="https://forkast.news/typesafe-ais-jev-is-not-an-llm-and-that-may-be-the-point/" rel="noopener noreferrer"&gt;meant to be used&lt;/a&gt;, a pattern shows up everywhere. You keep a model like Claude or GPT around for the parts that need thinking or writing, and you hand Jev the small decision sitting in between.&lt;/p&gt;

&lt;p&gt;A task comes in. Claude or GPT reads it and works out what needs to happen, maybe drafts a reply, maybe explains something. Jev sits next to that, and every time there's a small decision in the road, which row is urgent, whether to brake, whether a tool call is even needed, Jev picks the answer in a fraction of a second, and the app acts on it.&lt;/p&gt;

&lt;p&gt;The rule that falls out of this is simple: if you're making a model write out a full explanation just to arrive at one of three or four possible outcomes, you're paying for an essay when all you needed was a checkbox.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiihvmbdxjix0bptgkg4c.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiihvmbdxjix0bptgkg4c.png" alt=" " width="428" height="291"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Usecase
&lt;/h2&gt;

&lt;p&gt;Here are a &lt;a href="https://docs.typesafe.ai/concepts/use-case-map" rel="noopener noreferrer"&gt;few usecases&lt;/a&gt; that fit this shape well, the kind you'd normally build by prompting a chat model and parsing its reply.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sorting leads.&lt;/strong&gt; Feed it a form submission or a call summary, and let it sort into hot, warm, or cold. Nobody has to read through every lead by hand, and the hot ones get called first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routing support messages.&lt;/strong&gt; Give it an incoming email or chat message, and have it decide whether it goes to billing, technical, refunds, or gets flagged urgent. The right team sees it immediately instead of after someone manually sorts the queue.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Flagging invoices for review.&lt;/strong&gt; Pass in the invoice details and let it mark each one clean, suspicious, or needs review. Your team only has to actually look at the handful that came back flagged.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Triaging incoming customer messages.&lt;/strong&gt; For a business getting messages through chat or WhatsApp, sort each one into booking, pricing question, complaint, or spam, so replies go out to the ones that matter first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Choosing which model an agent should use.&lt;/strong&gt; Instead of burning tokens on a big model deciding whether to search, calculate, query a database, or hand off to a human, let Jev make that pick and save the expensive model for the actual task.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tagging comments and content.&lt;/strong&gt; Sort incoming comments or reviews into positive, negative, question, or needs a reply, so nothing important gets buried in a pile you were never going to read in full anyway.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A quick way to check if your use-case need Jev : Can the answer be one item picked from a short list? If yes, it fits. If the answer actually needs a sentence or two of explanation, that's still a job for LLM. &lt;/p&gt;

&lt;h2&gt;
  
  
  How to start?
&lt;/h2&gt;

&lt;p&gt;If you want to actually try it, the path is short:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Try Jev at &lt;a href="https://typesafe.ai/" rel="noopener noreferrer"&gt;typesafe.ai&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Try it in the playground first, give it a situation and three or four options, and see what it picks&lt;/li&gt;
&lt;li&gt;Move to the API once you've seen one decision work the way you expect&lt;/li&gt;
&lt;li&gt;Start with something small and repetitive, not your hardest decision. Pick the boring one you make fifty times a day&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;At $0.042 per million input tokens with free output, testing this costs close to nothing.&lt;/p&gt;

&lt;p&gt;For reference you can check what people are already building using Jev &lt;a href="https://madewithjev.com/" rel="noopener noreferrer"&gt;here.&lt;/a&gt; &lt;/p&gt;

&lt;h2&gt;
  
  
  When not to use Jev?
&lt;/h2&gt;

&lt;p&gt;A few situations where you still have to use LLM:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You need actual text back.&lt;/strong&gt; Emails, summaries, explanations, and other generated content are still better suited to LLMs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You can't define the possible outputs.&lt;/strong&gt; Jev works with predefined decisions such as yes/no, choices, or scores, so the decision needs to be structured ahead of time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The decision is genuinely high stakes.&lt;/strong&gt; For money, health, or legal outcomes, use Jev to sort, score, or prioritize, but keep a person involved in the final decision.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You haven't tested it on your own data.&lt;/strong&gt; Before relying on Jev in production, test it against real examples from your use case and see where it gets decisions right or wrong.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Jev isn't trying to replace GPT or Claude. LLMs are still the better fit when you need writing, explanation, or open-ended reasoning.&lt;/p&gt;

&lt;p&gt;Jev targets a different problem: &lt;strong&gt;decisions that software needs to make repeatedly.&lt;/strong&gt; For true/false checks, classification, routing, scoring, and verification, returning a typed decision and confidence score can be more useful than generating text that your code then has to parse.&lt;/p&gt;

&lt;p&gt;It won't outperform LLMs at every decision task, but its combination of structured outputs, speed, cost, and calibrated confidence makes it an interesting new primitive for building AI-powered software.&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
