<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: albe_sf</title>
    <description>The latest articles on DEV Community by albe_sf (@albertomontagnese).</description>
    <link>https://dev.to/albertomontagnese</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3928059%2F8788e7f6-c941-4959-b1cf-18686efc9034.jpg</url>
      <title>DEV Community: albe_sf</title>
      <link>https://dev.to/albertomontagnese</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/albertomontagnese"/>
    <language>en</language>
    <item>
      <title>The Assistants API is Gone. Here's How to Build Agents Now.</title>
      <dc:creator>albe_sf</dc:creator>
      <pubDate>Fri, 02 Oct 2026 15:04:00 +0000</pubDate>
      <link>https://dev.to/albertomontagnese/the-assistants-api-is-gone-heres-how-to-build-agents-now-3428</link>
      <guid>https://dev.to/albertomontagnese/the-assistants-api-is-gone-heres-how-to-build-agents-now-3428</guid>
      <description>&lt;p&gt;The OpenAI Assistants API was officially shut down on August 26, 2026. This deprecation forces a fundamental shift in how we build stateful agents, moving from a managed, persistent-thread model to a more direct, stateless approach with the new Responses and Conversations APIs. This change simplifies some parts of the stack but places the responsibility for state management squarely back on you, the developer.&lt;/p&gt;

&lt;h2&gt;
  
  
  what changed: from assistants to responses
&lt;/h2&gt;

&lt;p&gt;The original Assistants API was an attempt to abstract away the complexity of building conversational AI. It provided a stateful environment with persistent "threads," allowing developers to build agents that could maintain context over long interactions without managing the conversation history themselves. It also bundled powerful tools like Code Interpreter and a retrieval system.&lt;/p&gt;

&lt;p&gt;However, this abstraction came with trade-offs. The API could feel clunky, and its stateful nature led to unpredictable costs, as the entire conversation thread might be re-processed on every turn. Performance was also a concern, as developers had to poll for updates rather than using a real-time stream.&lt;/p&gt;

&lt;p&gt;The new model, centered on the Responses API, is a return to a more primitive, stateless paradigm. The core idea is a direct request/response flow. Persistent threads are replaced by &lt;code&gt;Conversation&lt;/code&gt; objects, which you must create and manage. The responsibility for maintaining context between turns now falls to your application code. This is a significant architectural change, but one that offers more control, better performance, and more predictable costs.&lt;/p&gt;

&lt;h2&gt;
  
  
  the new core of RAG: file_search
&lt;/h2&gt;

&lt;p&gt;For developers building Retrieval-Augmented Generation (RAG) systems, the most important change is the introduction of the &lt;code&gt;file_search&lt;/code&gt; tool within the Responses API. This is the successor to the old &lt;code&gt;Retrieval&lt;/code&gt; tool and is now the standard way to have models access knowledge from your private documents.&lt;/p&gt;

&lt;p&gt;The workflow is straightforward: you create a &lt;code&gt;vector_store&lt;/code&gt;, upload your files to it, and OpenAI's backend handles the entire pipeline of chunking, embedding, and indexing. You no longer have to build and manage your own embedding and retrieval logic.&lt;/p&gt;

&lt;p&gt;When you make a call to the Responses API, you can make the &lt;code&gt;file_search&lt;/code&gt; tool available. The model then intelligently decides when to use it based on the user's query. It performs a semantic search against your vector store, retrieves the most relevant passages, and incorporates them into its response, complete with citations.&lt;/p&gt;

&lt;p&gt;Here is a conceptual look at what an API call might look like in Python, enabling the tool for a specific vector store:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;
&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# Assuming a vector_store_id has been created and files uploaded
&lt;/span&gt;&lt;span class="n"&gt;vector_store_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;vs_123abc&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="c1"&gt;# Create a conversation and add a user message
&lt;/span&gt;&lt;span class="n"&gt;conversation&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;conversations&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;conversations&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;conversation_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;role&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What were the key findings in the Q3 financial report?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Create a response, enabling the file_search tool
&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-6.1-sol&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;conversation_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;file_search&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;file_search&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;vector_store_ids&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;vector_store_id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# The response will contain the model's answer,
# potentially using information retrieved from the file search.
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  what you gain and what you lose
&lt;/h2&gt;

&lt;p&gt;The primary gain from this migration is control. The move to a stateless API gives you direct authority over conversation history and state management. This means more predictable performance and costs, as you are no longer subject to the black box of the old persistent-thread system. The Responses API also consolidates complex workflows into a single API call, simplifying the overall interaction model.&lt;/p&gt;

&lt;p&gt;The most significant loss is convenience. The hand-holding of the managed, infinite-context thread is gone. If your application was architected around the assumption that OpenAI would manage the state, you now have a non-trivial migration project ahead of you. OpenAI does not provide an automatic tool for migrating old Threads to new Conversations.&lt;/p&gt;

&lt;p&gt;There is also a key technical trade-off with the managed &lt;code&gt;file_search&lt;/code&gt; tool. While it removes the burden of building a RAG pipeline, it also takes away your control over the chunking strategy. For highly structured or complex documents, the automated chunking might not be optimal, which can impact the quality of retrieval. This is a critical limitation to be aware of when deciding whether to use the built-in tool or roll your own retrieval system.&lt;/p&gt;

&lt;h3&gt;
  
  
  the final word
&lt;/h3&gt;

&lt;p&gt;This shift marks a maturation of the AI developer stack. The initial, heavily-abstracted Assistants API has been replaced by more powerful and primitive building blocks. It’s a move away from providing a magic box and toward giving engineers direct access to the core components of RAG and agentic systems. For builders, this means more responsibility, but also more power. The era of the fully managed agent is over; the era of building your own is here.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://platform.openai.com/docs/assistants/migration" rel="noopener noreferrer"&gt;Assistants migration guide | OpenAI API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.openai.com/docs/guides/tools-file-search" rel="noopener noreferrer"&gt;File search | OpenAI API&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>python</category>
    </item>
    <item>
      <title>Anthropic's Sonnet 5.5 release is about cost-per-task, not peak performance</title>
      <dc:creator>albe_sf</dc:creator>
      <pubDate>Wed, 30 Sep 2026 15:04:14 +0000</pubDate>
      <link>https://dev.to/albertomontagnese/anthropics-sonnet-55-release-is-about-cost-per-task-not-peak-performance-27i6</link>
      <guid>https://dev.to/albertomontagnese/anthropics-sonnet-55-release-is-about-cost-per-task-not-peak-performance-27i6</guid>
      <description>&lt;p&gt;Anthropic released Claude Sonnet 5.5 this week, and the main takeaway isn't a new state-of-the-art benchmark. The real story is the focus on cost-per-task for the bulk of everyday engineering work. This isn't a model for moonshots; it's a workhorse for bug fixes, documentation, and routine code generation, where speed and efficiency matter more than cutting-edge reasoning.&lt;/p&gt;

&lt;h2&gt;
  
  
  what changed
&lt;/h2&gt;

&lt;p&gt;Sonnet 5.5 is the second model in Anthropic's 5.5 family, positioned as a faster, more economical alternative to the top-tier Claude Opus 5.5. Where Opus is designed for complex tasks requiring careful judgment, Sonnet is aimed at well-defined, everyday work.&lt;/p&gt;

&lt;p&gt;The key improvements over its predecessor, Sonnet 5, are in efficiency. Sonnet 5.5 generates output over 30% faster. While the per-token pricing remains the same, Anthropic claims it costs up to 30% less for most tasks because it requires fewer tokens to complete the same work.&lt;/p&gt;

&lt;p&gt;The performance jump on agentic coding is significant. On the Terminal-Bench 4.0 evaluation, Sonnet 5.5 scored 70.6%, a massive leap from Sonnet 5's 10.3%. This suggests it's far more capable for tasks that require tool use and autonomous operation within a terminal environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  why it matters for builders
&lt;/h2&gt;

&lt;p&gt;The most expensive model is rarely the right tool for every job. The release of Sonnet 5.5, alongside recent models like OpenAI's GPT-6 Sol and Luna, shows the market maturing. Labs are now competing on the cost-performance curve, not just on leaderboards. For a significant portion of a developer's workflow—writing unit tests, refactoring a function, generating boilerplate, summarizing a pull request—the raw intelligence of a frontier model is overkill. Latency and cost become the primary constraints.&lt;/p&gt;

&lt;p&gt;A model that is 30% faster and cheaper for 80% of your daily tasks is a material change to your workflow. It makes continuous, ambient use of AI more practical. When calling the API, you are simply targeting the new model name.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="c1"&gt;# api_key="my_api_key"
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-sonnet-5-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Write a Python function to calculate the Fibonacci sequence and include docstrings.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This shift means you can afford to integrate AI into more granular parts of your development loop without worrying about the bill. The gains in agentic coding also suggest that mid-tier models are becoming viable for more complex automations that were previously the domain of top-tier models.&lt;/p&gt;

&lt;h2&gt;
  
  
  some gotchas
&lt;/h2&gt;

&lt;p&gt;This is not a frontier model, and Anthropic is clear about that. The company stated that Sonnet 5.5 does not advance the frontier of its models' capabilities. For high-stakes, complex reasoning, Opus 5.5 remains the recommended choice.&lt;/p&gt;

&lt;p&gt;Interestingly, because Sonnet 5.5's cybersecurity capabilities are a significant improvement over Sonnet 5, it is the first Sonnet model to ship with the kind of cyber safeguards previously reserved for top-tier models. For certain high-risk cybersecurity requests, the model will visibly fall back to Sonnet 5.&lt;/p&gt;

&lt;h2&gt;
  
  
  the so-what
&lt;/h2&gt;

&lt;p&gt;The era of just chasing the highest benchmark score is giving way to a more pragmatic focus on the right tool for the job. Sonnet 5.5 is a strong signal that the major labs see a huge market in providing capable, efficient, and economically viable models for the vast majority of software development tasks. For builders, this means more choice and better tools for the everyday grind. The most important model in your stack might not be the most powerful one, but the one that delivers reliable results with the best balance of speed and cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/" rel="noopener noreferrer"&gt;Introducing Claude Sonnet 5.5&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>claude</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>DBRX and the Rise of Fine-Grained MoE</title>
      <dc:creator>albe_sf</dc:creator>
      <pubDate>Mon, 28 Sep 2026 15:04:18 +0000</pubDate>
      <link>https://dev.to/albertomontagnese/dbrx-and-the-rise-of-fine-grained-moe-52bc</link>
      <guid>https://dev.to/albertomontagnese/dbrx-and-the-rise-of-fine-grained-moe-52bc</guid>
      <description>&lt;p&gt;A new state-of-the-art open language model just arrived, but the real story isn't just its benchmark scores. The release of DBRX provides a clear signal about where efficient model architecture is heading: fine-grained Mixture-of-Experts (MoE). For builders, this approach is the critical takeaway, as it directly impacts inference speed, serving costs, and the viability of deploying powerful custom models.&lt;/p&gt;

&lt;h2&gt;
  
  
  a new benchmark for open models
&lt;/h2&gt;

&lt;p&gt;Databricks released DBRX as a general-purpose large language model that outperforms other established open-source models on a variety of standard benchmarks, including language understanding, programming, and math. It was developed to provide enterprises with a platform to build their own custom, high-performance AI systems without relying on a few closed-source providers.&lt;/p&gt;

&lt;p&gt;The model itself is a decoder-only transformer, but its architecture is what sets it apart. This design choice is a deliberate move toward efficiency, aiming to deliver top-tier performance while managing the computational costs associated with massive models.&lt;/p&gt;

&lt;h2&gt;
  
  
  the fine-grained mixture-of-experts architecture
&lt;/h2&gt;

&lt;p&gt;The core innovation in DBRX is its fine-grained MoE architecture. Instead of a single, dense network, an MoE model comprises multiple specialized "expert" networks and a router that selects which experts to engage for a given input. While other models like Mixtral-8x7B use this approach, DBRX implements a more granular strategy.&lt;/p&gt;

&lt;p&gt;DBRX has 16 total experts and selects 4 of them for any given input. This is in contrast to models like Mixtral or Grok-1, which use 8 experts and select 2. This finer-grained approach provides a vastly larger number of possible expert combinations, which improves overall model quality. While the model has a total of 132 billion parameters, only 36 billion are active during inference on any single input. This makes the model significantly faster and more cost-effective to serve than a dense model of a similar size.&lt;/p&gt;

&lt;p&gt;This architecture is built on open-source projects, making it a design pattern that other builders can adopt. The combination of high active parameter count and a large number of fine-grained experts appears to be a key recipe for its performance.&lt;/p&gt;

&lt;h2&gt;
  
  
  running dbrx locally
&lt;/h2&gt;

&lt;p&gt;For engineers who want to experiment with the model, DBRX is available on Hugging Face. Getting it running requires the &lt;code&gt;transformers&lt;/code&gt; library and a significant amount of VRAM, but the process itself is straightforward. The model uses the GPT-4 tokenizer.&lt;/p&gt;

&lt;p&gt;Here is a basic example of how you might load the instruction-tuned model and run inference:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;AutoModelForCausalLM&lt;/span&gt;

&lt;span class="c1"&gt;# Ensure you have torch and transformers installed
# pip install torch transformers sentencepiece
&lt;/span&gt;
&lt;span class="n"&gt;tokenizer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;databricks/dbrx-instruct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# Note: This requires substantial memory. Use device_map="auto" for multi-GPU.
&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AutoModelForCausalLM&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;databricks/dbrx-instruct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; 
    &lt;span class="n"&gt;torch_dtype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bfloat16&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; 
    &lt;span class="n"&gt;device_map&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;auto&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;trust_remote_code&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# The prompt format should follow the model's chat template
&lt;/span&gt;&lt;span class="n"&gt;user_prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Explain the concept of a Mixture-of-Experts (MoE) model in a few sentences.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;user_prompt&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;input_ids&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tokenizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;apply_chat_template&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;return_tensors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;to&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;device&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;outputs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;input_ids&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; 
    &lt;span class="n"&gt;max_new_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; 
    &lt;span class="n"&gt;do_sample&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; 
    &lt;span class="n"&gt;top_k&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;top_p&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.95&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tokenizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;outputs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;skip_special_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This snippet demonstrates loading the model and tokenizer, formatting a prompt, and generating a response. The key takeaway for builders is the accessibility of a model with this architecture through standard open-source tooling.&lt;/p&gt;

&lt;h2&gt;
  
  
  why this matters now
&lt;/h2&gt;

&lt;p&gt;The release of a powerful, open model with a fine-grained MoE architecture is not just an incremental update. It's a clear indicator that the frontier of AI development is increasingly focused on architectural efficiency, not just scaling parameter counts. For engineering teams, this trend is a welcome one. Efficient models like DBRX lower the barrier to entry for building and deploying custom AI applications.&lt;/p&gt;

&lt;p&gt;This shift allows more organizations to move from proprietary, closed models to open-source alternatives that they can fine-tune and control. As builders, we should be paying close attention to these architectural patterns. They represent the next step in democratizing access to state-of-the-art AI, making powerful systems more practical and affordable to build and serve.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.databricks.com/blog/introducing-dbrx-new-state-art-open-llm" rel="noopener noreferrer"&gt;Introducing DBRX: A New State-of-the-Art Open LLM&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Gemma 2 is here. It's time to pay attention to open models.</title>
      <dc:creator>albe_sf</dc:creator>
      <pubDate>Fri, 25 Sep 2026 15:03:35 +0000</pubDate>
      <link>https://dev.to/albertomontagnese/gemma-2-is-here-its-time-to-pay-attention-to-open-models-5538</link>
      <guid>https://dev.to/albertomontagnese/gemma-2-is-here-its-time-to-pay-attention-to-open-models-5538</guid>
      <description>&lt;p&gt;Google has released Gemma 2, the next generation of its open models, and it's a release that warrants a closer look from builders. The key takeaway is this: the 27B parameter model delivers performance competitive with models more than twice its size, while running on a single GPU. This isn't just an incremental update; it's a shift in the cost-to-performance ratio for high-quality open models.&lt;/p&gt;

&lt;h2&gt;
  
  
  what's new under the hood
&lt;/h2&gt;

&lt;p&gt;Gemma 2 comes in 9 billion and 27 billion parameter sizes, with a smaller 2B version also available. The architecture introduces some notable changes from the first generation. It now uses a combination of local and global attention, allowing it to focus on both immediate context and the broader meaning of a text. It also implements Grouped Query Attention (GQA) to improve inference speed and parameter efficiency.&lt;/p&gt;

&lt;p&gt;The training data for the 27B model consisted of 13 trillion tokens, while the 9B model was trained on 8 trillion tokens, sourced from a diverse mix of web documents and code. For the smaller 2B and 9B models, Google used knowledge distillation from the larger 27B model, which they report leads to significant performance gains compared to training from scratch with the same token count.&lt;/p&gt;

&lt;h2&gt;
  
  
  performance and efficiency claims
&lt;/h2&gt;

&lt;p&gt;The headline claim is that the 27B model provides a competitive alternative to much larger, closed-source models. Google states that the 9B model also delivers class-leading performance, outperforming other open models in its size category like Llama 3 8B.&lt;/p&gt;

&lt;p&gt;Crucially, the 27B model is designed for efficient inference on a single NVIDIA H100 or even an A100 80GB GPU. This significantly lowers the barrier to entry for deploying a high-performance model, moving it out of the realm of large-scale clusters and onto a single machine. This focus on efficiency makes self-hosting and running local inference for product development much more feasible.&lt;/p&gt;

&lt;h2&gt;
  
  
  how to get started
&lt;/h2&gt;

&lt;p&gt;Getting up and running with Gemma 2 is straightforward. The model weights are available on Hugging Face, Kaggle, and through Google AI Studio. For local development, tools like Ollama provide a simple way to pull and run the models with a single command.&lt;/p&gt;

&lt;p&gt;Here's a basic example of how you might run the 9B instruction-tuned model locally after installing Ollama:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Pull the Gemma 2 9B model&lt;/span&gt;
ollama run gemma2:9b

&lt;span class="c"&gt;# After pulling, you can interact with it directly&lt;/span&gt;
&lt;span class="c"&gt;# or use it via the Ollama API&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For more integrated use cases, you can use libraries like &lt;code&gt;transformers&lt;/code&gt; from Hugging Face. Here's a Python snippet to load the model and tokenizer, using 4-bit quantization to manage memory usage on consumer hardware.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;AutoModelForCausalLM&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;BitsAndBytesConfig&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;

&lt;span class="n"&gt;model_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;google/gemma-2-9b-it&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="c1"&gt;# Configure 4-bit quantization
&lt;/span&gt;&lt;span class="n"&gt;quantization_config&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BitsAndBytesConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;load_in_4bit&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;bnb_4bit_quant_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nf4&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;bnb_4bit_compute_dtype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bfloat16&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;tokenizer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AutoModelForCausalLM&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;quantization_config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;quantization_config&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;device_map&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;auto&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;input_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Explain the key architectural changes in Gemma 2.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;input_ids&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;tokenizer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;input_text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;return_tensors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;to&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;device&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;outputs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;input_ids&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_new_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;150&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tokenizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;outputs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This approach makes it possible to run a powerful model on a developer's laptop or a single cloud GPU, which is a game-changer for rapid prototyping and building privacy-sensitive applications.&lt;/p&gt;

&lt;h2&gt;
  
  
  the so-what for builders
&lt;/h2&gt;

&lt;p&gt;The release of Gemma 2 is another sign that the performance gap between state-of-the-art proprietary models and open models is closing, especially when you factor in efficiency. For builders, this means more freedom. You can fine-tune and deploy a capable model without being locked into a specific API provider or facing unpredictable usage bills.&lt;/p&gt;

&lt;p&gt;The ability to run a 27B parameter model that competes with 50B+ models on a single GPU is the real story here. It changes the economics of building AI features and products. It makes local, private, and cost-effective AI applications a more realistic goal for a wider range of developers and organizations. This is a good week to re-evaluate your stack and see where an open model might fit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://blog.google/" rel="noopener noreferrer"&gt;https://blog.google/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>MiMo-V2.6 and the bet on open reinforcement learning</title>
      <dc:creator>albe_sf</dc:creator>
      <pubDate>Wed, 23 Sep 2026 15:08:28 +0000</pubDate>
      <link>https://dev.to/albertomontagnese/mimo-v26-and-the-bet-on-open-reinforcement-learning-3mn5</link>
      <guid>https://dev.to/albertomontagnese/mimo-v26-and-the-bet-on-open-reinforcement-learning-3mn5</guid>
      <description>&lt;p&gt;Xiaomi just released and open-sourced its MiMo-V2.6 series, with models positioned to compete with leading closed-source offerings on agentic tasks. But the more significant release is the methodology. The accompanying technical report details a scaled-up reinforcement learning (RL) strategy, making a public case that the path forward is not just bigger base models, but better, continuous improvement through RL at scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  what changed: scaling rl in public
&lt;/h2&gt;

&lt;p&gt;The core idea behind the MiMo-V2.6 series is scaling three components of reinforcement learning in tandem: RL compute, the diversity of task environments, and grader compute. Instead of discrete training runs for different domains, the approach uses one mixed RL run across coding, general agent tasks, visual tasks, and cybersecurity. This produced two main omni-modal models: MiMo-V2.6-Pro, the flagship, and MiMo-V2.6-Flash, a smaller model optimized for cost and speed.&lt;/p&gt;

&lt;p&gt;The Pro model, a sparse mixture-of-experts architecture with over a trillion total parameters, now stands as one of the most powerful open-source models available, particularly on benchmarks that measure long-horizon engineering tasks. The company has fully open-sourced the weights and technical report for both the Pro and Flash models. They also released over 7,000 high-quality RL task environments covering software engineering, web design, and knowledge work, providing a valuable resource for others building agentic systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  the engineering details that matter
&lt;/h2&gt;

&lt;p&gt;The technical report provides specific details about the training process. The RL phase consumed massive compute, processing 1,568 prompts per training step, with each prompt generating 16 attempts. This high-throughput, asynchronous training pipeline allowed the model to learn from billions of tokens at each step within a context length of up to one million tokens.&lt;/p&gt;

&lt;p&gt;A key improvement was in the quality of the feedback signal. The team moved beyond simple correctness verifiers to a more sophisticated "groupwise agentic grading" system. This grader analyzes execution traces to provide more accurate reward signals for long-running tasks, steering the model toward more token-efficient solutions. This investment in grader compute represented a significant portion of the total RL cost but was critical for keeping the training process stable and preventing reward hacking.&lt;/p&gt;

&lt;p&gt;For builders who want to run the model directly, the Hugging Face page provides configuration details. Serving the Pro model can be done using a containerized environment with a library like vLLM.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Example command to serve the model using a pre-built container&lt;/span&gt;
docker pull vllm/vllm-openai:mimov25-cu129

vllm serve XiaomiMiMo/MiMo-V2.6-Pro-RL &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--tensor-parallel-size&lt;/span&gt; 8 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--trust-remote-code&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--gpu-memory-utilization&lt;/span&gt; 0.95 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--max-model-len&lt;/span&gt; auto &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--reasoning-parser&lt;/span&gt; mimo &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--tool-call-parser&lt;/span&gt; mimo &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--enable-auto-tool-choice&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--generation-config&lt;/span&gt; vllm
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This configuration highlights the model's native support for features like automatic tool choice, which is essential for agentic workflows.&lt;/p&gt;

&lt;h2&gt;
  
  
  the so-what for builders
&lt;/h2&gt;

&lt;p&gt;The release of MiMo-V2.6 pushes the performance of open-source models further, particularly for complex, multi-step tasks that require an agent to interact with tools and environments. On some benchmarks for long-horizon software engineering, the Pro model's performance improved significantly over the course of its RL training.&lt;/p&gt;

&lt;p&gt;While the closed-source labs are currently focused on a price war for high-volume inference, this release signals a different competitive front for open-source: the training methodology itself. By open-sourcing not just the model but also the RL environments, the release provides a toolkit for teams building their own specialized agents. It allows builders to replicate, fine-tune, and extend the RL process on their own data and tasks. This is less about providing a drop-in replacement for a commercial API and more about providing the foundation for building custom, self-improving systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://alphaxiv.org/abs/2409.15832" rel="noopener noreferrer"&gt;MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement (Technical Report)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL" rel="noopener noreferrer"&gt;MiMo-V2.6-Pro-RL on Hugging Face&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://mimo.xiaomi.com/" rel="noopener noreferrer"&gt;Introducing MiMo-V2.6 series (Official Announcement)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Real-Time Translation Latency Just Dropped. Here's What Changed.</title>
      <dc:creator>albe_sf</dc:creator>
      <pubDate>Mon, 21 Sep 2026 15:03:44 +0000</pubDate>
      <link>https://dev.to/albertomontagnese/real-time-translation-latency-just-dropped-heres-what-changed-42g6</link>
      <guid>https://dev.to/albertomontagnese/real-time-translation-latency-just-dropped-heres-what-changed-42g6</guid>
      <description>&lt;p&gt;Simultaneous interpretation has always been a trade-off between speed and accuracy. Waiting for more context yields a better translation, but speaking sooner reduces the awkward silence for the listener. Qwen's release of Qwen3.8-LiveTranslate suggests a change in the underlying architecture that meaningfully moves the needle, cutting average lag time to 2.3 seconds.&lt;/p&gt;

&lt;p&gt;This isn't just an incremental improvement. It's a structural change that makes real-time, multi-speaker translation practical for production applications.&lt;/p&gt;

&lt;h2&gt;
  
  
  what just shipped
&lt;/h2&gt;

&lt;p&gt;The Qwen team at Alibaba released Qwen3.8-LiveTranslate, a model for real-time, simultaneous translation of audio and video streams. It's available as a hosted API on Alibaba Cloud Model Studio and QwenCloud, accessible over a WebSocket connection.&lt;/p&gt;

&lt;p&gt;The key metric is a drop in Length-Adaptive Average Lagging (LAAL) from 2.8 to 2.3 seconds, an 18% reduction in the average time the translation trails the source speech. The model understands 60 languages and can generate speech in 29 of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  the interleave architecture
&lt;/h2&gt;

&lt;p&gt;The performance gain comes from a new design Qwen calls the Interleave architecture. Previous systems often treated speech recognition, translation, and speech synthesis as separate, sequential steps. This new model reframes the problem by treating the incoming audio and the outgoing translated text as a single, interleaved data stream.&lt;/p&gt;

&lt;p&gt;Under the hood is a two-module design described as a "Thinker" and a "Talker". The Thinker module arranges the audio, source text, and translation into one causal sequence. This allows the model to process audio and generate translated text within a single, unified process, which improves both quality and latency. The Talker module then handles speech synthesis, preserving the original speaker's voice.&lt;/p&gt;

&lt;h2&gt;
  
  
  new capabilities for builders
&lt;/h2&gt;

&lt;p&gt;Beyond the latency reduction, the release includes two features that address common pain points in building real-world translation apps.&lt;/p&gt;

&lt;p&gt;First is real-time speaker diarization. The model can distinguish between different speakers in a multi-party conversation and attribute the translation correctly. The API also exposes voice cloning capabilities to maintain a more stable voice for each speaker throughout a session.&lt;/p&gt;

&lt;p&gt;Second is a synchronized bilingual display. The API can stream both the source transcription and the translation as separate, aligned events. This allows you to build UIs that show both languages simultaneously, which is critical for applications like subtitling or meeting summaries where users might need to reference the original text.&lt;/p&gt;

&lt;p&gt;Here is a conceptual look at how you might handle the WebSocket stream in Python.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;asyncio&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;websockets&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;

&lt;span class="c1"&gt;# Note: This is a conceptual example. 
# Refer to official Alibaba Cloud documentation for the actual API endpoint and auth.
&lt;/span&gt;
&lt;span class="n"&gt;WEBSOCKET_URI&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;wss://api.qwen.ai/v1/translate/qwen3.8-livetranslate-flash-realtime&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;stream_audio_for_translation&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio_chunk_iterator&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;websockets&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;WEBSOCKET_URI&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;websocket&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Connection established.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;send_audio&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;audio_chunk_iterator&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;websocket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;send&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;asyncio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;0.1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# Simulate real-time streaming
&lt;/span&gt;            &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;websocket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;send&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;stream_end&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}))&lt;/span&gt;

        &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;receive_translation&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
            &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;websocket&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;event&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;transcription&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SOURCE: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;event&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;translation&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TRANSLATION: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;event&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;error&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Error: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                    &lt;span class="k"&gt;break&lt;/span&gt;

        &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;asyncio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;gather&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;send_audio&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="nf"&gt;receive_translation&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;

&lt;span class="c1"&gt;# Example usage:
# async def get_audio_chunks():
#     # Your logic to get real-time audio chunks from a mic or stream
#     for i in range(10):
#         yield f"audio_chunk_{i}".encode('utf-8')
#
# asyncio.run(stream_audio_for_translation(get_audio_chunks()))
&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  so what
&lt;/h2&gt;

&lt;p&gt;For engineers who have previously dismissed simultaneous translation as too slow or inaccurate for interactive use cases, this release is a signal to re-evaluate. The architectural shift from a sequential pipeline to an interleaved stream is a meaningful change. When combined with practical features like speaker separation, it makes building robust, multilingual, real-time voice applications significantly more feasible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.alibabacloud.com/blog/qwen3-8-livetranslate-names-the-speaker-carries-the-meaning_602166" rel="noopener noreferrer"&gt;Qwen3.8-LiveTranslate: Names the Speaker. Carries the Meaning.&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://modelstudio.aliyuncs.com/models/qwen/qwen3.8-livetranslate-flash-realtime/summary" rel="noopener noreferrer"&gt;Alibaba Cloud Model Studio: qwen3.8-livetranslate-flash-realtime&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>python</category>
      <category>news</category>
    </item>
    <item>
      <title>Mistral just shipped a vision model and cut flagship prices</title>
      <dc:creator>albe_sf</dc:creator>
      <pubDate>Fri, 18 Sep 2026 15:02:31 +0000</pubDate>
      <link>https://dev.to/albertomontagnese/mistral-just-shipped-a-vision-model-and-cut-flagship-prices-3o6i</link>
      <guid>https://dev.to/albertomontagnese/mistral-just-shipped-a-vision-model-and-cut-flagship-prices-3o6i</guid>
      <description>&lt;p&gt;Mistral just made two significant moves that position it as an increasingly practical alternative to the major closed-source labs. The company released Pixtral, a new vision model, and simultaneously announced a price reduction for its flagship model, Mistral Large 2. This isn't just a routine update; it's a clear signal about their strategy: compete directly on multimodal features while aggressively pushing down cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  what just shipped
&lt;/h2&gt;

&lt;p&gt;The update, which appeared on their official changelog on September 17, included three key changes.&lt;/p&gt;

&lt;p&gt;First, the release of &lt;code&gt;pixtral-12b-2409&lt;/code&gt;, their new vision model. The naming suggests a 12-billion parameter model, a deliberate choice to offer strong performance in a cost-effective, easily deployable size. This follows a broader industry trend toward multimodality, but Mistral's commitment to open and efficient models makes this release particularly notable for builders.&lt;/p&gt;

&lt;p&gt;Second, they released an updated version of their small model, &lt;code&gt;mistral-small-2409&lt;/code&gt;. While less flashy than a new vision model, continuous improvement of smaller, efficient models is critical for production use cases where latency and cost are primary constraints.&lt;/p&gt;

&lt;p&gt;Third, Mistral cut the price on their most capable model, Mistral Large 2, and introduced a free API tier on their platform, La Plateforme. This directly addresses one of the biggest barriers to adoption for smaller teams and individual developers: cost. Lowering the price of your top-tier model is a confident move designed to capture more production workloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  competing on cost and capability
&lt;/h2&gt;

&lt;p&gt;The release of Pixtral is a direct answer to the multimodal models from larger labs. For developers building applications that need to understand or process images, this provides a new, potentially more open and efficient option. While benchmarks are not yet available, a 12B parameter model is large enough for serious tasks without incurring the inference costs of massive frontier models.&lt;/p&gt;

&lt;p&gt;Integrating a model like this into a workflow is straightforward. You can expect to interact with it through the standard API, likely with a modified payload to handle image inputs, such as a base64-encoded string or a URL.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;

&lt;span class="n"&gt;API_KEY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_MISTRAL_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;MODEL_NAME&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pixtral-12b-2409&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="c1"&gt;# Encode a local image file
&lt;/span&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;path/to/your/image.jpg&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;image_file&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;encoded_string&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;b64encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image_file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;headers&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;API_KEY&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;MODEL_NAME&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;messages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
                &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Describe the contents of this image in detail.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
                &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data:image/jpeg;base64,&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;encoded_string&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;
            &lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.mistral.ai/v1/chat/completions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is more than just adding a feature. It's about providing the core building blocks that engineers need. The simultaneous price cut on Mistral Large 2 reinforces this. It makes the entire stack, from small and efficient to large and powerful, more economically viable. For teams running systems at scale, these cost differences add up quickly.&lt;/p&gt;

&lt;h2&gt;
  
  
  the bigger picture: open distribution
&lt;/h2&gt;

&lt;p&gt;Mistral's strategy appears to extend beyond just models and APIs. Their recent partnership with Mozilla to integrate Mistral models into the Firefox browser is a move to control distribution and reach users outside of the typical developer-focused cloud platforms. By building a presence directly in the browser, they are creating a new channel for their technology, one that is built on a foundation of open technology.&lt;/p&gt;

&lt;p&gt;This matters. As AI becomes more deeply embedded in our daily tools, the question of who controls the underlying models becomes critical. By partnering with organizations like Mozilla and continuing to release open-weight models, Mistral is providing a real alternative to the closed ecosystems of Big Tech.&lt;/p&gt;

&lt;p&gt;For builders, this is a positive development. It means more choice, better pricing, and the ability to build on platforms that align with the open principles of the web. The latest releases are not just new tools, but a continued investment in an ecosystem that offers more than one way to build.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.mistral.ai/changelog" rel="noopener noreferrer"&gt;Mistral AI Docs: Changelog&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Google's New Gemini Voice Models Change The Rules For Agent State Management</title>
      <dc:creator>albe_sf</dc:creator>
      <pubDate>Wed, 16 Sep 2026 15:04:00 +0000</pubDate>
      <link>https://dev.to/albertomontagnese/googles-new-gemini-voice-models-change-the-rules-for-agent-state-management-e2m</link>
      <guid>https://dev.to/albertomontagnese/googles-new-gemini-voice-models-change-the-rules-for-agent-state-management-e2m</guid>
      <description>&lt;p&gt;Google released two new live dialogue models on September 15, 2026: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. While incremental model numbers are common, the architectural shift in the Extended Thinking variant is not. It introduces parallel reasoning, allowing the model to process tasks in the background while streaming an audio response, fundamentally changing the interaction pattern for building complex voice agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  what just shipped
&lt;/h2&gt;

&lt;p&gt;Google has positioned the two models for different use cases. Gemini 3.8 Live is framed for scale, cost efficiency, and fluid dialogue. The more interesting model for developers building complex agents is Gemini 3.8 Live Extended Thinking. This version is designed for high-complexity tasks that require multi-step reasoning.&lt;/p&gt;

&lt;p&gt;Both models are being rolled out across the Gemini API, Google AI Studio, and into products like Google Workspace and Search. The core capability upgrade is a move toward more natural and fluid voice interactions, but the mechanism for achieving this has direct implications for developers using the API.&lt;/p&gt;

&lt;h2&gt;
  
  
  asynchronous reasoning is the new default
&lt;/h2&gt;

&lt;p&gt;The most significant change is how the Extended Thinking model handles work. It can process background tasks, like asynchronous tool calls, while simultaneously generating and streaming a continuous audio response. This is a departure from the traditional, blocking request-response cycle where the user waits in silence while the model completes its entire thought process.&lt;/p&gt;

&lt;p&gt;For anyone who has built a voice agent, the benefit is obvious: the user gets immediate verbal feedback while the agent continues to work on a complex query. This makes the interaction feel less robotic.&lt;/p&gt;

&lt;p&gt;However, it imposes a new burden on the client application. The developer documentation notes that when using asynchronous reasoning, a &lt;code&gt;turnComplete: true&lt;/code&gt; signal no longer indicates that the model is idle. The server may still be processing tool calls or other background reasoning. Your client must continue to listen for subsequent server messages even after the initial audio stream for a given "turn" has finished.&lt;/p&gt;

&lt;h2&gt;
  
  
  updating your client-side logic
&lt;/h2&gt;

&lt;p&gt;This new interaction model requires a shift in state management. A simple &lt;code&gt;while&lt;/code&gt; loop that terminates on a completion flag is no longer sufficient. Your application needs to handle a more persistent connection and be prepared for out-of-band messages from the server.&lt;/p&gt;

&lt;p&gt;Consider a conceptual Python client. Previously, you might have done something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Old pattern: simple request-response loop
&lt;/span&gt;&lt;span class="n"&gt;response_stream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate_content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio_chunks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response_stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;play_audio&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;turn_complete&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt; &lt;span class="c1"&gt;# The turn is over, we can stop listening.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The new model requires logic that persists and handles different message types after a turn appears to be complete.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# New pattern: persistent listening for async events
&lt;/span&gt;&lt;span class="n"&gt;response_stream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate_content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio_chunks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Assume a persistent connection or long-lived stream
&lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response_stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has_audio_chunk&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="nf"&gt;play_audio&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;audio_chunk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has_tool_call&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="c1"&gt;# Even if audio is playing or has finished for this turn,
&lt;/span&gt;        &lt;span class="c1"&gt;# we receive and dispatch a tool call here.
&lt;/span&gt;        &lt;span class="nf"&gt;dispatch_tool_call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool_call&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;turn_complete&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# This no longer means the model is idle.
&lt;/span&gt;        &lt;span class="c1"&gt;# It's just the end of this speech segment.
&lt;/span&gt;        &lt;span class="c1"&gt;# The loop continues, listening for more messages like tool results.
&lt;/span&gt;        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;INFO: Turn complete, but continuing to listen for background tasks.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# The connection would remain open to receive further updates
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a more complex, event-driven approach. The client has to be able to play audio while simultaneously listening for and processing other events, like function call requests from the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  the so-what for builders
&lt;/h2&gt;

&lt;p&gt;The introduction of asynchronous reasoning is a significant step toward more capable voice agents that can tackle complex, multi-step tasks without creating an awkward, silent user experience. It moves the interaction closer to a human-like collaboration where speaking and thinking can happen in parallel. For any team building production-ready voice agents, this is a pattern worth paying attention to.&lt;/p&gt;

&lt;p&gt;However, this capability comes with the engineering overhead of more complex client-side logic. It also introduces an open question on cost. While competitors often publish per-minute pricing for their real-time voice APIs, Google has not yet done so for these new models, leaving a critical variable unknown for teams evaluating them for production use.&lt;/p&gt;

&lt;p&gt;Ultimately, the release of Gemini 3.8 Live Extended Thinking signals that the frontier of voice AI is moving beyond simple, turn-based interactions. The next challenge for builders is to create applications that can gracefully manage the state of an agent that thinks and speaks at the same time.&lt;/p&gt;

&lt;h2&gt;
  
  
  sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://blog.google/" rel="noopener noreferrer"&gt;Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ai.google/gemini-api/docs/models/gemini-3.8-live-extended-thinking" rel="noopener noreferrer"&gt;Gemini 3.8 Live Extended Thinking - Google AI for Developers&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Claude 3.5 Sonnet is more than an upgrade. It’s a new workflow.</title>
      <dc:creator>albe_sf</dc:creator>
      <pubDate>Mon, 14 Sep 2026 15:04:51 +0000</pubDate>
      <link>https://dev.to/albertomontagnese/claude-35-sonnet-is-more-than-an-upgrade-its-a-new-workflow-2opk</link>
      <guid>https://dev.to/albertomontagnese/claude-35-sonnet-is-more-than-an-upgrade-its-a-new-workflow-2opk</guid>
      <description>&lt;p&gt;Anthropic just shipped Claude 3.5 Sonnet, its first model in the new 3.5 series. It’s faster and cheaper than Opus, and outperforms it on several key benchmarks, but the bigger story for builders is a new feature called Artifacts. This addition moves Claude from a simple conversational AI into a collaborative work environment, changing the core loop of how we build with these models.&lt;/p&gt;

&lt;h2&gt;
  
  
  from chat loop to live workspace
&lt;/h2&gt;

&lt;p&gt;Historically, the workflow for using an LLM to code or create has been a turn-based conversation. You ask for a code snippet, you copy it, paste it into your local editor, run it, find an error, and then return to the chat to describe the problem. This cycle repeats, with the context split between your IDE and the chat window.&lt;/p&gt;

&lt;p&gt;The new Artifacts feature in the claude.ai interface changes this. When you ask for something like a React component or an SVG graphic, Claude now generates it in a dedicated panel next to the conversation. This panel is a live workspace where you can see the rendered output, edit the code, and iterate in real-time without leaving the browser.&lt;/p&gt;

&lt;p&gt;This creates a tighter feedback loop. Instead of copying and pasting, you can prompt for a change—"add a login feature"—and see the code and the rendered preview update in the same window. It’s a move toward making the AI an active participant in a shared workspace, not just a passive respondent in a chat.&lt;/p&gt;

&lt;h2&gt;
  
  
  agentic coding gets a boost
&lt;/h2&gt;

&lt;p&gt;Beyond the UI, the model itself has significant improvements for developers. In an internal agentic coding evaluation, Claude 3.5 Sonnet solved 64% of problems, a marked improvement over the 38% solved by Claude 3 Opus. This evaluation tests the model's ability to fix bugs or add functionality to open-source codebases based on a natural language description.&lt;/p&gt;

&lt;p&gt;This suggests the model is more capable of complex reasoning, troubleshooting, and independent code execution when given the right tools. For teams working on migrating legacy applications or automating codebase maintenance, this is a meaningful step up. The model handles code translations and updates with more proficiency.&lt;/p&gt;

&lt;p&gt;Accessing the new model is straightforward. It's available through the Anthropic API, as well as on Amazon Bedrock and Google Cloud's Vertex AI.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="c1"&gt;# defaults to os.environ.get("ANTHROPIC_API_KEY")
&lt;/span&gt;    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;my_api_key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-3-5-sonnet-20240620&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Write a python function to check if a number is prime.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The pricing is set at $3 per million input tokens and $15 per million output tokens, with a 200K token context window. This makes it significantly more cost-effective than Opus for tasks that require high intelligence but also demand speed, like orchestrating multi-step agentic workflows.&lt;/p&gt;

&lt;h2&gt;
  
  
  what this means for builders
&lt;/h2&gt;

&lt;p&gt;The release of Claude 3.5 Sonnet and the Artifacts feature isn't just another incremental model bump. It's a signal of where the developer experience is heading. The friction of context-switching between a chatbot and an IDE is being actively designed out of the process.&lt;/p&gt;

&lt;p&gt;This evolution from conversational partner to collaborative tool is the key takeaway. For builders, it means we can start designing workflows that assume the AI is in the editor with us, not just in another window. This has implications for prototyping, debugging, and even how we onboard new developers to a complex codebase. The model is becoming an on-demand teammate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/news/claude-3-5-sonnet" rel="noopener noreferrer"&gt;https://www.anthropic.com/news/claude-3-5-sonnet&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>claude</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Cohere's 218B Parameter MoE Model for Translation Dropped Quietly</title>
      <dc:creator>albe_sf</dc:creator>
      <pubDate>Fri, 11 Sep 2026 15:06:58 +0000</pubDate>
      <link>https://dev.to/albertomontagnese/coheres-218b-parameter-moe-model-for-translation-dropped-quietly-5adn</link>
      <guid>https://dev.to/albertomontagnese/coheres-218b-parameter-moe-model-for-translation-dropped-quietly-5adn</guid>
      <description>&lt;p&gt;Cohere released a 218-billion-parameter translation model, and it barely made a sound. North Small Translate is a massive sparse model that sets a new performance benchmark for its domain. Its architecture and release strategy show where production-grade specialized models are heading: massive scale, focused on a single task, with efficiency coming from sparsity.&lt;/p&gt;

&lt;h2&gt;
  
  
  what shipped
&lt;/h2&gt;

&lt;p&gt;On September 9, 2026, Cohere published release notes for North Small Translate, an open-weight Mixture-of-Experts (MoE) model built specifically for machine translation. The model has 218 billion total parameters, with 25 billion active for any given token. It's a sparse architecture with 128 experts, activating 8 per token.&lt;/p&gt;

&lt;p&gt;This isn't a general-purpose chat model. It's a specialist, supporting translation across 50 languages. The weights are available on Hugging Face for research and non-commercial use under a CC BY-NC 4.0 license. For production use, Cohere routes you to a commercial license and their Model Vault deployment.&lt;/p&gt;

&lt;p&gt;Performance-wise, Cohere reports a WMT26 score of 83.60 across all evaluated languages. They also note this can be pushed to 84.36 using an agentic multi-pass workflow where the model refines its own output.&lt;/p&gt;

&lt;h2&gt;
  
  
  why it matters for builders
&lt;/h2&gt;

&lt;p&gt;The most significant takeaway is the hardware footprint versus the parameter count. Because it's a sparse MoE model, you aren't loading all 218B parameters for every inference. Cohere provides clear hardware minimums for different quantization levels. A 4-bit quantized version can run on a single NVIDIA B200 or two H100 GPUs. The full BF16 precision requires four B200s or eight H100s. This is still substantial, but it puts a model of this scale within reach for self-hosting, which is not the case for dense models of a similar size.&lt;/p&gt;

&lt;p&gt;The release strategy itself is also notable. This was a quiet drop, first appearing on Hugging Face weeks before the official release note. It's a move towards treating large models less like blockbuster events and more like industrial components. You have a specific, high-value problem like translation at enterprise scale. You deploy a specialized, high-performance component to solve it.&lt;/p&gt;

&lt;p&gt;For teams working with multilingual systems, this model represents a new frontier for quality, especially for the 32 high-resource languages it covers well.&lt;/p&gt;

&lt;h2&gt;
  
  
  using specialized translation models
&lt;/h2&gt;

&lt;p&gt;While you can download the weights for evaluation, most production use will be via an API. Interacting with a dedicated translation model is more direct than prompting a general-purpose model. You're not engineering a complex prompt with few-shot examples; you're calling a function.&lt;/p&gt;

&lt;p&gt;Here’s a hypothetical Python snippet of what an SDK interaction might look like. Note that this is a representative example, not a direct copy of Cohere's current SDK.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;cohere&lt;/span&gt;

&lt;span class="c1"&gt;# Assuming API key is configured in environment variables
&lt;/span&gt;&lt;span class="n"&gt;co&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cohere&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Client&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# The model ID would point to the specialized translation model
&lt;/span&gt;&lt;span class="n"&gt;model_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;north-small-translate-1-0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;source_texts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;To build great AI products, focus on the user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s workflow.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;La arquitectura de transformadores es la base de los modelos lingüísticos modernos.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;target_language&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;de&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="c1"&gt;# German
&lt;/span&gt;
&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;co&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;translate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;texts&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;source_texts&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;target_language&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;target_language&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;translation&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;translations&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Original: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;translation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;source_text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Translation: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;translation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key is the shift from conversational prompting to a more structured, tool-like interaction. The model expects a specific input (text and a target language) and provides a specific output (the translated text). The complexity is in the model's architecture, not in your prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  the takeaway
&lt;/h2&gt;

&lt;p&gt;North Small Translate is a signal of maturity in the AI space. We are moving past the era where every new model had to be a better generalist. Instead, we are seeing the rise of massive, hyper-specialized models that are state-of-the-art at a single, commercially valuable task. For builders, this means having more powerful and efficient tools for specific jobs, even if it requires significant hardware to run them yourself. It pays to watch the specialists, not just the chatbots.&lt;/p&gt;

&lt;h2&gt;
  
  
  sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.cohere.com/" rel="noopener noreferrer"&gt;Cohere Release Notes&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.unite.ai/cohere-debuts-open-weight-218b-mixture-of-experts-machine-translation-model/" rel="noopener noreferrer"&gt;Unite.AI: Cohere Debuts Open-Weight 218B Mixture-of-Experts Machine Translation Model&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://huggingface.co/CohereLabs/North-Small-Translate-1.0" rel="noopener noreferrer"&gt;Hugging Face Model Card: North-Small-Translate-1.0&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Gemini 3.8 Flash is not a routine update</title>
      <dc:creator>albe_sf</dc:creator>
      <pubDate>Wed, 09 Sep 2026 15:04:52 +0000</pubDate>
      <link>https://dev.to/albertomontagnese/gemini-38-flash-is-not-a-routine-update-3j71</link>
      <guid>https://dev.to/albertomontagnese/gemini-38-flash-is-not-a-routine-update-3j71</guid>
      <description>&lt;p&gt;Google's release of Gemini 3.8 Flash is more than an incremental version bump. It represents a deliberate focus on the workflows that builders are actually shipping: long-running agents, multi-step reasoning, and software engineering tasks. This isn't about chasing chatbot benchmarks; it's about providing more effective tools for complex, automated systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  what changed with 3.8 flash
&lt;/h2&gt;

&lt;p&gt;The key advancements in Gemini 3.8 Flash are centered on performance for software engineering and agentic knowledge workflows. This is a direct response to how developers are using these models in production. While general capability improvements are always welcome, targeted enhancements for code generation, debugging, and orchestrating complex tasks are what move the needle on a day-to-day basis.&lt;/p&gt;

&lt;p&gt;For teams already using the Gemini 3.x series, the transition is straightforward. The introductory API pricing for 3.8 Flash remains the same as it was for 3.7 Flash, though Google has indicated this pricing will change in January. This provides a window for developers to integrate and test the new model's capabilities without an immediate cost increase.&lt;/p&gt;

&lt;p&gt;The model continues to support customizable effort levels, allowing a trade-off between quality, cost, and latency. This is a critical feature for production systems where you might want to use a faster, cheaper response for one task and a slower, higher-quality one for another.&lt;/p&gt;

&lt;h2&gt;
  
  
  a dedicated model for cyber
&lt;/h2&gt;

&lt;p&gt;The most significant part of this release is the introduction of Gemini 3.8 Flash Cyber. This is a specialized variant of the model fine-tuned for cybersecurity use cases, specifically for vulnerability discovery and automated patching.&lt;/p&gt;

&lt;p&gt;Access to this model is not public. It's being made available to trusted defenders through a new channel called Google's Fairwind Program. This gated approach is becoming a pattern for frontier models with sensitive capabilities. By controlling the release, providers aim to mitigate misuse while still getting the tool into the hands of security professionals who can use it for defense.&lt;/p&gt;

&lt;p&gt;This move signals a broader industry trend. As models become more powerful, we will see more of these specialized, access-controlled variants for high-stakes domains. Expect to see similar models for finance, medicine, and critical infrastructure in the near future. For builders, this means the most powerful tools may require a verification process, not just an API key.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"task"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"scan_and_patch"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"target"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"repository"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"github.com/example/repo"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"branch"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"main"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model_config"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"google"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gemini-3.8-flash-cyber"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"credentials_secret"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"GOOGLE_FAIRWIND_TOKEN"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"parameters"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"vulnerability_types"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"sql_injection"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"cross_site_scripting"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"dependency_confusion"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"propose_pull_request"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"notify_channel"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"#security-alerts"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  the so-what for builders
&lt;/h2&gt;

&lt;p&gt;The release of Gemini 3.8 Flash and its Cyber variant confirms that the next phase of AI development is specialization. Foundational, general-purpose models are becoming a commodity. The real value is in models that are expertly tuned for specific, high-value vertical tasks like software engineering and cybersecurity.&lt;/p&gt;

&lt;p&gt;This shift has direct implications for how you build. It means that simply calling a generic model API is no longer the optimal approach. Instead, you should be evaluating a portfolio of models, including specialized ones, and routing tasks to the tool best suited for the job. The future of building with AI is less about having one all-powerful model and more about orchestrating a fleet of specialized agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://deepmind.google/news/" rel="noopener noreferrer"&gt;https://deepmind.google/news/&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Anthropic's Fable and Mythos 5.1: More Than a Model Update</title>
      <dc:creator>albe_sf</dc:creator>
      <pubDate>Mon, 07 Sep 2026 15:05:43 +0000</pubDate>
      <link>https://dev.to/albertomontagnese/anthropics-fable-and-mythos-51-more-than-a-model-update-4h5p</link>
      <guid>https://dev.to/albertomontagnese/anthropics-fable-and-mythos-51-more-than-a-model-update-4h5p</guid>
      <description>&lt;p&gt;Anthropic's release of Claude Fable 5.1 and Mythos 5.1 is more than an incremental update. It marks a strategic shift in how frontier models are productized, splitting a single underlying architecture into two distinct offerings tailored for different risk profiles and use cases. For builders, this means a new state-of-the-art model for coding and knowledge work that is also cheaper to run, coupled with a clearer framework for how the most powerful capabilities will be gated.&lt;/p&gt;

&lt;h2&gt;
  
  
  what changed: fable vs. mythos
&lt;/h2&gt;

&lt;p&gt;Fable 5.1 is the new flagship model for general availability, setting a new performance standard for coding, knowledge work, and complex problem-solving. It replaces its predecessor as the state-of-the-art option for most developers building on the platform. The key change is that Fable 5.1 is one of two new models. The other, Mythos 5.1, is the same underlying model but with different safeguards.&lt;/p&gt;

&lt;p&gt;Mythos 5.1 is designed specifically for high-stakes research in sensitive fields like cybersecurity and biology. Access is restricted to a small number of vetted organizations through trusted access programs. This bifurcation is the main story: instead of a single model with one-size-fits-all safety controls, Anthropic is creating distinct products from the same core intelligence. Fable 5.1 gets more precise safeguards that are less likely to intervene on benign requests, while Mythos 5.1 provides more specialized capabilities for trusted partners.&lt;/p&gt;

&lt;h2&gt;
  
  
  pragmatic shifts for builders
&lt;/h2&gt;

&lt;p&gt;For engineers shipping products, the most significant changes are economic and practical. Fable 5.1 is estimated to be 25% less expensive than Fable 5 for typical workloads, a meaningful reduction for production systems. This cost reduction makes it more feasible to use a frontier-class model for tasks that might have previously been relegated to smaller, less capable models.&lt;/p&gt;

&lt;p&gt;The updated safety mechanisms in Fable 5.1 are also a practical benefit. The new safeguards are more precise, with interventions on benign biology-related requests reportedly reduced by 85%. The model can now be used to identify software vulnerabilities in source code, a task that was previously more restricted. This fine-tuning of the safety layer means fewer false positives and a more reliable experience for developers working on legitimate but potentially sensitive applications.&lt;/p&gt;

&lt;p&gt;When using the API, the model name is the primary change, but developers should also be aware of new cost-saving mechanics like improved caching.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="c1"&gt;# api_key="my_api_key",
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# A typical call to the new Fable 5.1 model
&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-fable-5-1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;4096&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Review this Python code for potential vulnerabilities and suggest improvements.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This approach allows developers to access state-of-the-art performance for general coding and analysis tasks without needing to apply for specialized access programs.&lt;/p&gt;

&lt;h2&gt;
  
  
  the new playbook for frontier model deployment
&lt;/h2&gt;

&lt;p&gt;The Fable/Mythos split is a clear signal of where the industry is heading. As model capabilities increase, especially in scientifically sensitive areas, a single safety policy becomes untenable. A blanket approach either stifles legitimate research or fails to adequately contain risk.&lt;/p&gt;

&lt;p&gt;By creating a tiered system, labs can offer a powerful, general-purpose model like Fable 5.1 to a broad audience while reserving the most potent, potentially dual-use capabilities for partners who have undergone a vetting process. This allows them to continue pushing the research frontier with Mythos 5.1 while providing a more stable and predictable product for the majority of their customers.&lt;/p&gt;

&lt;p&gt;For builders, this trend is worth watching closely. It suggests that future model access will be less about a single API endpoint and more about a portfolio of models, each with specific capabilities, safeguards, and access requirements. Understanding this structure will be as important as understanding the model's performance on benchmarks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/" rel="noopener noreferrer"&gt;Introducing Claude Fable 5.1 and Claude Mythos 5.1&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/claude-mythos" rel="noopener noreferrer"&gt;Claude Mythos&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>claude</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
