<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Armin Burger</title>
    <description>The latest articles on DEV Community by Armin Burger (@armin_burger_ab136b2f8bb1).</description>
    <link>https://dev.to/armin_burger_ab136b2f8bb1</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2642895%2F737f3d5a-6e7a-4a79-a423-4f3944b3bbbc.png</url>
      <title>DEV Community: Armin Burger</title>
      <link>https://dev.to/armin_burger_ab136b2f8bb1</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/armin_burger_ab136b2f8bb1"/>
    <language>en</language>
    <item>
      <title>Beyond the ReAct Loop: Why Your AI Agent Breaks in Production</title>
      <dc:creator>Armin Burger</dc:creator>
      <pubDate>Fri, 09 Oct 2026 09:55:31 +0000</pubDate>
      <link>https://dev.to/armin_burger_ab136b2f8bb1/beyond-the-react-loop-why-your-ai-agent-breaks-in-production-45he</link>
      <guid>https://dev.to/armin_burger_ab136b2f8bb1/beyond-the-react-loop-why-your-ai-agent-breaks-in-production-45he</guid>
      <description>&lt;p&gt;The ReAct (Reasoning and Acting) pattern is deceptively simple. In a tutorial, you define an agent, give it a few tools, and watch it solve complex queries. It looks magical. But moving that same loop into production reveals a harsh reality: the complexity isn’t in the reasoning logic itself, but in the infrastructure surrounding it. Most "agent failures" attributed to poor LLM reasoning are actually symptoms of inadequate tool design, missing cost controls, or security oversights.&lt;/p&gt;

&lt;h2&gt;
  
  
  Iteration Limits Are Cost Controls, Not Just Safety Nets
&lt;/h2&gt;

&lt;p&gt;A common configuration mistake is treating &lt;code&gt;max_iterations&lt;/code&gt; solely as a safeguard against infinite loops. While it prevents runaway processes, its primary function in production is cost management. Every iteration consumes tokens for both the input context and the generated output. Without strict limits derived from your specific cost model, a single user query can drain resources rapidly.&lt;/p&gt;

&lt;p&gt;For example, setting &lt;code&gt;max_iterations=10&lt;/code&gt; in a LangChain &lt;code&gt;AgentExecutor&lt;/code&gt; isn't just about stopping a stuck bot; it's about capping the financial exposure per request. You must also handle parsing errors gracefully. Instead of raising exceptions when the model outputs malformed JSON, feed the error back as an observation. This allows the agent to self-correct without terminating the session prematurely.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;langchain.agents&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AgentExecutor&lt;/span&gt;

&lt;span class="n"&gt;executor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;AgentExecutor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; 
    &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; 
    &lt;span class="n"&gt;max_iterations&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; 
    &lt;span class="n"&gt;handle_parsing_errors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Tool Design Determines Quality More Than Prompts
&lt;/h2&gt;

&lt;p&gt;There is a misconception that better prompts lead to better agents. In practice, agent quality depends heavily on well-scoped, clearly described tools. Ambiguous descriptions or poorly truncated results cause what appears to be model failure but is actually an interface problem.&lt;/p&gt;

&lt;p&gt;Consider a web search tool. If it returns 50 raw HTML snippets, you flood the context window, increasing costs and confusing the model. You must enforce truncation at the tool boundary. A robust implementation might limit results explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;langchain.tools&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Tool&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;duckduckgo_search&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;search_web&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;ddgs&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_results&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# Strict limit
&lt;/span&gt;
&lt;span class="n"&gt;web_search&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;WebSearch&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;search_web&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Useful for searching current events. Returns top 5 results.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The description matters as much as the code. Vague instructions lead to hallucinated tool usage. Precise scoping ensures the model knows exactly when and how to use the tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security: Treat Tool Results as Untrusted Input
&lt;/h2&gt;

&lt;p&gt;This is often overlooked. In a standard chatbot, user input is untrusted. In an agent system, &lt;strong&gt;tool outputs&lt;/strong&gt; are also untrusted. If your agent fetches data from a public API or scrapes a website, that content could contain prompt injection attacks. Malicious text embedded in a search result could instruct the agent to execute dangerous commands or leak data. Because agents have side effects (executing code, sending emails), these injections are far more dangerous than in static chats.&lt;/p&gt;

&lt;h2&gt;
  
  
  Streaming Requires Intermediate Events
&lt;/h2&gt;

&lt;p&gt;Finally, user experience suffers if you only stream the final answer. An agent takes time to think, select tools, and process results. To maintain engagement and provide transparency, you must stream intermediate step events—such as "thinking," "tool selection," and "execution status"—rather than just waiting for the final token. This requires a custom streaming protocol that exposes the internal state of the ReAct loop.&lt;/p&gt;

&lt;p&gt;Production agents aren't built by tweaking prompts. They are engineered through rigorous cost modeling, secure tool boundaries, and transparent execution flows.&lt;/p&gt;

&lt;p&gt;Repo: &lt;code&gt;github.com/armbur19-collab/chimerai-app&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://chimerai.dev/blog?utm_campaign=blog-promo&amp;amp;utm_medium=social&amp;amp;utm_source=devto-08-ai-agents-in-production-what-the-react-loop-does-n_devto" rel="noopener noreferrer"&gt;ChimerAI Blog&lt;/a&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>production</category>
    </item>
    <item>
      <title>Stop Outsourcing Your Agent’s Brain: Why DIY Often Beats Frameworks in Production</title>
      <dc:creator>Armin Burger</dc:creator>
      <pubDate>Tue, 06 Oct 2026 16:46:50 +0000</pubDate>
      <link>https://dev.to/armin_burger_ab136b2f8bb1/stop-outsourcing-your-agents-brain-why-diy-often-beats-frameworks-in-production-4mdg</link>
      <guid>https://dev.to/armin_burger_ab136b2f8bb1/stop-outsourcing-your-agents-brain-why-diy-often-beats-frameworks-in-production-4mdg</guid>
      <description>&lt;p&gt;The debate around AI agent architecture often frames the choice between using a hosted platform, adopting an open-source framework, or building from scratch as a question of engineering maturity. This is a misconception. The decision is strictly a matter of fit, determined by two critical factors: the complexity of your control flow and who owns the failure modes when things inevitably break.&lt;/p&gt;

&lt;h3&gt;
  
  
  When to Use a Platform
&lt;/h3&gt;

&lt;p&gt;Hosted platforms are excellent for specific scenarios, primarily when the agent is a secondary feature within a larger product rather than the core value proposition itself. If you are adding a chatbot to a SaaS dashboard where the primary business logic resides elsewhere, outsourcing the agent infrastructure makes sense. However, this convenience comes with a significant trade-off: moving agent operations to a platform shifts core behavior outside your codebase. This limits your ability to locally reproduce failures and patch issues independently. You become dependent on third-party release schedules for fixes that affect your user experience.&lt;/p&gt;

&lt;h3&gt;
  
  
  When to Use a Framework
&lt;/h3&gt;

&lt;p&gt;Frameworks occupy the middle ground. They are optimal when your agent’s logic aligns with standard multi-step reasoning patterns and common tool interfaces. These tools handle the tedious plumbing, such as the tool-calling protocol, message history formatting, retry logic for parse errors, and streaming interfaces. For many developers, this abstraction saves time. But beware: working around a framework’s abstraction can sometimes be harder than implementing the functionality without it. If your requirements deviate from the "happy path" supported by the framework, you may find yourself fighting the library rather than leveraging it.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Case for Building It Yourself (DIY)
&lt;/h3&gt;

&lt;p&gt;Building an agent loop from scratch is justified when specific requirements make framework abstractions obstructive. A prime example is exact cost accounting or unusual control flow logic. A DIY implementation typically involves calling provider APIs directly within a simple &lt;code&gt;while&lt;/code&gt; loop with a clear stop condition. While this requires more initial effort, it grants complete visibility into every step of the execution.&lt;/p&gt;

&lt;p&gt;If the agent &lt;em&gt;is&lt;/em&gt; your product, its loop, error handling, and cost model constitute core competencies that should not be outsourced. Relying on external libraries means you cannot fully control how prompt injection vulnerabilities are handled or how tool results are sanitized before being returned to the model context.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Non-Negotiable: Cost and Observability
&lt;/h3&gt;

&lt;p&gt;Regardless of the strategy chosen—platform, framework, or DIY—two technical realities remain mandatory. First, unbounded agent execution equates to unbounded financial liability. Every agent loop must have an iteration limit derived from a strict cost model to prevent runaway spending. Second, debugging capability is more critical than feature richness. Observability requires tracing every tool call, argument, and result. If a solution does not provide full trace visibility, it is fundamentally flawed for production use because you cannot diagnose why an agent failed.&lt;/p&gt;

&lt;h3&gt;
  
  
  Conclusion
&lt;/h3&gt;

&lt;p&gt;Choose your architecture based on ownership of failure modes. If you need independent patching capabilities and precise cost controls, build it yourself. If you are shipping a minor feature quickly, use a platform. Only choose a framework if your logic perfectly matches its abstractions. Don’t let convenience compromise your ability to debug your own product.&lt;/p&gt;

&lt;p&gt;Repo: &lt;code&gt;github.com/armbur19-collab/chimerai-app&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://chimerai.dev/blog?utm_campaign=blog-promo&amp;amp;utm_medium=social&amp;amp;utm_source=devto-07-ai-agent-platforms-framework-platform-or-just-buil_devto" rel="noopener noreferrer"&gt;ChimerAI Blog&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>productivity</category>
      <category>aiops</category>
    </item>
    <item>
      <title>Your prompts are code. Manage them like data — but know what you lose when you do</title>
      <dc:creator>Armin Burger</dc:creator>
      <pubDate>Wed, 30 Sep 2026 11:43:16 +0000</pubDate>
      <link>https://dev.to/armin_burger_ab136b2f8bb1/your-prompts-are-code-manage-them-like-data-but-know-what-you-lose-when-you-do-4i3j</link>
      <guid>https://dev.to/armin_burger_ab136b2f8bb1/your-prompts-are-code-manage-them-like-data-but-know-what-you-lose-when-you-do-4i3j</guid>
      <description>&lt;p&gt;The three things that separate an AI demo from an AI product aren't models. They're: being able&lt;br&gt;
to change a prompt without a deploy, knowing what each request cost you, and not leaking a&lt;br&gt;
user's credit card into logs. None of these are exciting. All three are in the ChimerAI stack,&lt;br&gt;
so here's what they do and — more usefully — the tradeoffs they hide.&lt;/p&gt;
&lt;h2&gt;
  
  
  1. Prompt templates in the database
&lt;/h2&gt;

&lt;p&gt;Instead of prompt strings scattered across &lt;code&gt;.py&lt;/code&gt; files, there's a &lt;code&gt;PromptTemplate&lt;/code&gt; model with&lt;br&gt;
Mustache-style variables:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;model PromptTemplate {
  id        String   @id @default(cuid())
  name      String   @unique
  category  String   // "system" | "user" | "rag" | "agent" | "chat" | "custom"
  content   String   @db.Text
  variables String[] // extracted from {{placeholders}} at save time
  language  String   @default("en")
  version   Int      @default(1)
  isDefault Boolean  @default(false)
  isActive  Boolean  @default(true)
  tags      String[]
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;variables&lt;/code&gt; array is derived from the content, so the app can validate that a caller&lt;br&gt;
supplied every &lt;code&gt;{{context}}&lt;/code&gt; / &lt;code&gt;{{query}}&lt;/code&gt; before rendering — a missing variable becomes a&lt;br&gt;
400 at the API boundary, not a prompt that silently ships a literal &lt;code&gt;{{context}}&lt;/code&gt; to the model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST /api/prompts
{ "name": "RAG System Prompt",
  "category": "rag",
  "content": "Use {{context}} to answer {{query}}",
  "isDefault": true }
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Exactly one default per category, multi-language (EN/DE/FR/ES/IT), and the whole thing is&lt;br&gt;
behind an &lt;code&gt;manage_prompts&lt;/code&gt; permission so non-admins can't edit production prompts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The tradeoff nobody tells you about:&lt;/strong&gt; the moment a prompt lives in the DB instead of git, you&lt;br&gt;
lose your diff, your code review, and your "what shipped on Tuesday" answer. Versioning here is&lt;br&gt;
an &lt;code&gt;Int&lt;/code&gt; counter, not a history — it tells you &lt;em&gt;this is v7&lt;/em&gt;, not &lt;em&gt;what v5 said&lt;/em&gt;. If prompts&lt;br&gt;
matter to you (and for a RAG system they do), either export template changes to your audit log&lt;br&gt;
or keep the canonical copy in git and treat the DB as a runtime override. I lean toward the&lt;br&gt;
latter and haven't fully committed to either, which is honest.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. Per-request cost tracking
&lt;/h2&gt;

&lt;p&gt;Every chat call runs through &lt;code&gt;trackApiUsage&lt;/code&gt;, which records the endpoint, model, token counts,&lt;br&gt;
success/failure, and status code. The provider's own token counts are used where available; the&lt;br&gt;
streaming path accumulates them from chunks and reports once at the end:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;provider_id&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;total_prompt_tokens&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;total_completion_tokens&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;provider_client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;report_usage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;provider_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;provider_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;prompt_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;total_prompt_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;completion_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;total_completion_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;endpoint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/api/chat/stream&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is what lets the credit check in the request path exist at all, and what turns "our AI&lt;br&gt;
bill doubled" from a mystery into a query. The model is chosen at runtime and can be swapped&lt;br&gt;
per provider without a deploy, so cost is tracked against the model that actually answered, not&lt;br&gt;
a config constant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Caveat:&lt;/strong&gt; the &lt;em&gt;pre-request&lt;/em&gt; budget check uses a character-based token estimate&lt;br&gt;
(&lt;code&gt;chars / 4 × 1.2&lt;/code&gt;), not the real tokenizer. It's a soft gate to stop runaway spend, not a&lt;br&gt;
billing system. The recorded usage is the accurate number; the estimate is a heuristic that&lt;br&gt;
will occasionally reject a request slightly early. Don't invoice off the estimate.&lt;/p&gt;
&lt;h2&gt;
  
  
  3. Guardrails: moderation, PII, injection
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;chimerai add guardrails&lt;/code&gt; gives you four endpoints. The PII detector is regex-based — which is&lt;br&gt;
both its strength (fast, no network, deterministic) and its ceiling (it catches patterns, not&lt;br&gt;
meaning):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;       &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;phone&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;       &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;\b\d{3}[-.]?\d{3}[-.]?\d{4}\b&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ssn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;         &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;\b\d{3}-\d{2}-\d{4}\b&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;credit_card&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;\b\d{4}[- ]?\d{4}[- ]?\d{4}[- ]?\d{4}\b&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Endpoints:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Path&lt;/th&gt;
&lt;th&gt;What&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;POST&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/api/guardrails/pii/detect&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;find email/phone/SSN/CC/IP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;POST&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/api/guardrails/pii/redact&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;mask the above&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;POST&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/api/guardrails/toxicity&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;keyword score 0–1 + risk level&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;POST&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/api/guardrails/injection&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;prompt-injection patterns&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Toxicity and injection detection are keyword/pattern-based with a confidence score and a risk&lt;br&gt;
classification. That's enough to catch the obvious stuff and to add a cheap first line of&lt;br&gt;
defense before a request hits your paid model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where it stops being sufficient:&lt;/strong&gt; US phone/SSN/CC formats only (a German project will want&lt;br&gt;
different patterns — worth noting since this kit is partly German-authored); "toxicity" by&lt;br&gt;
keyword misses rephrased hostility entirely; regex injection detection is a tripwire, not a&lt;br&gt;
classifier. If compliance demands real content moderation, put a dedicated model in front and&lt;br&gt;
keep these as a fast pre-filter. &lt;code&gt;detect&lt;/code&gt; vs &lt;code&gt;redact&lt;/code&gt; as separate calls is deliberate — you&lt;br&gt;
often want to log what was found even when you forward a masked version.&lt;/p&gt;

&lt;h2&gt;
  
  
  How they fit together
&lt;/h2&gt;

&lt;p&gt;The RAG pipeline in the same stack reads its system prompt from category &lt;code&gt;rag&lt;/code&gt; (feature 1),&lt;br&gt;
runs the user query through guardrails before embedding (feature 3), and reports token usage&lt;br&gt;
after the LLM call (feature 2). Three boring features, one grounded-and-metered request.&lt;/p&gt;

&lt;p&gt;None of this is the part that demos well. It's the part that lets you sleep once users find&lt;br&gt;
your app.&lt;/p&gt;

&lt;p&gt;Repo: &lt;code&gt;github.com/armbur19-collab/chimerai-app&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://chimerai.dev/blog?utm_campaign=blog-promo&amp;amp;utm_medium=content&amp;amp;utm_source=devto-06-prompts-cost-tracking-guardrails" rel="noopener noreferrer"&gt;ChimerAI Blog&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>promptengineering</category>
      <category>productivity</category>
      <category>architecture</category>
    </item>
    <item>
      <title>A RAG system in one command — what chimerai add rag actually wires up</title>
      <dc:creator>Armin Burger</dc:creator>
      <pubDate>Tue, 29 Sep 2026 12:24:15 +0000</pubDate>
      <link>https://dev.to/armin_burger_ab136b2f8bb1/a-rag-system-in-one-command-what-chimerai-add-rag-actually-wires-up-59ff</link>
      <guid>https://dev.to/armin_burger_ab136b2f8bb1/a-rag-system-in-one-command-what-chimerai-add-rag-actually-wires-up-59ff</guid>
      <description>&lt;p&gt;RAG is one of those things where the tutorial is 30 lines and the production system is a&lt;br&gt;
vector-database procurement process. This post is about the middle ground: a working retrieval&lt;br&gt;
pipeline you can drop into a project in one command, what it actually installs, and the exact&lt;br&gt;
point where you'd want to replace it.&lt;/p&gt;
&lt;h2&gt;
  
  
  The command
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;my-nextjs-app
npx chimerai add rag
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;code&gt;rag&lt;/code&gt; depends on the chat module, so if &lt;code&gt;ai-chat&lt;/code&gt; isn't installed the CLI adds it first. What&lt;br&gt;
you get is a self-contained Python AI service (FastAPI + LiteLLM) under &lt;code&gt;services/ai/&lt;/code&gt; with the&lt;br&gt;
RAG module switched on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;services/ai/
├── config.py                    # pydantic settings, env loading
├── provider_client.py           # LiteLLM — multi-provider routing
├── main.py                      # FastAPI entry (generated from manifest)
├── services/
│   ├── rag_service.py           # ingest → retrieve → answer pipeline
│   ├── vector_store.py          # FAISS index + persistence
│   └── embedding_service.py     # text embeddings via LiteLLM
├── routes/rag_routes.py         # /api/rag/upload, /query, /stats
└── data/                        # FAISS index on disk
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Next.js side gets proxy routes that forward to &lt;code&gt;AI_SERVICE_URL&lt;/code&gt; (default&lt;br&gt;
&lt;code&gt;http://localhost:8002&lt;/code&gt;), so your frontend calls &lt;code&gt;/api/rag/...&lt;/code&gt; and never talks to Python&lt;br&gt;
directly. &lt;code&gt;chimerai dev&lt;/code&gt; starts both.&lt;/p&gt;
&lt;h2&gt;
  
  
  The pipeline, in the order it runs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Ingest&lt;/strong&gt; — you hand it raw text, it chunks and embeds. Chunking is&lt;br&gt;
&lt;code&gt;RecursiveCharacterTextSplitter&lt;/code&gt; with defaults you can see rather than guess:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text_splitter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RecursiveCharacterTextSplitter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;chunk_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;chunk_overlap&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;length_function&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;len&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;separators&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;chunk_size=1000&lt;/code&gt; is in &lt;em&gt;characters&lt;/em&gt;, not tokens (&lt;code&gt;length_function=len&lt;/code&gt;) — a detail that&lt;br&gt;
matters if you're used to token-based splitters; 1000 chars is roughly 250 tokens of English.&lt;br&gt;
The 200-char overlap is there so a fact split across a boundary still shows up whole in at&lt;br&gt;
least one chunk. Each chunk keeps a pointer back to its source:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;all_chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;all_metadatas&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;_chunk_index&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;_chunk_total&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;_source_doc_index&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those &lt;code&gt;_chunk_*&lt;/code&gt; fields are what let you show "source: page 3" in the UI instead of a bare&lt;br&gt;
similarity score.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Store&lt;/strong&gt; — a &lt;code&gt;faiss.IndexFlatL2&lt;/code&gt; on 1536-dim vectors (OpenAI &lt;code&gt;text-embedding-ada-002&lt;/code&gt;). Flat&lt;br&gt;
L2 means &lt;em&gt;exact&lt;/em&gt; brute-force search: no approximation, so recall is perfect and the index just&lt;br&gt;
scans everything. That's fine to tens of thousands of vectors and gets slow past that — the&lt;br&gt;
code comment says as much:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Use IndexFlatL2 for exact search (can be changed to IndexIVFFlat for large datasets)
&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;faiss&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;IndexFlatL2&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;dimension&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Retrieve + answer&lt;/strong&gt; — the RAG chat is a plain "stuff context into the system prompt" flow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;relevant_docs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;vector_store&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;similarity_search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;context_parts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;relevant_docs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;context_parts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[Document &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;]&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context_parts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;context_parts&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;No relevant documents found.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;system_message&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;You are a helpful assistant. Use the following context to answer &lt;/span&gt;&lt;span class="se"&gt;\
&lt;/span&gt;&lt;span class="s"&gt;the user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s question. If the context doesn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;t contain relevant information, say so and &lt;/span&gt;&lt;span class="se"&gt;\
&lt;/span&gt;&lt;span class="s"&gt;provide a general answer.

Context:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;[Document i]&lt;/code&gt; labels exist so the model can cite, and the "if the context doesn't contain&lt;br&gt;
the answer, say so" line is the single most important prompt sentence in the whole file —&lt;br&gt;
without it the model answers from parametric memory and your eval numbers are fiction.&lt;/p&gt;
&lt;h2&gt;
  
  
  Using it
&lt;/h2&gt;

&lt;p&gt;Add documents:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;POST http://localhost:8002/api/rag/documents
&lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="s2"&gt;"documents"&lt;/span&gt;: &lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"FastAPI is a modern web framework for Python."&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;,
  &lt;span class="s2"&gt;"metadatas"&lt;/span&gt;: &lt;span class="o"&gt;[{&lt;/span&gt; &lt;span class="s2"&gt;"source"&lt;/span&gt;: &lt;span class="s2"&gt;"docs"&lt;/span&gt;, &lt;span class="s2"&gt;"page"&lt;/span&gt;: 1 &lt;span class="o"&gt;}]&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ask a grounded question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;POST http://localhost:8002/api/rag/chat
&lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="s2"&gt;"query"&lt;/span&gt;: &lt;span class="s2"&gt;"Tell me about FastAPI"&lt;/span&gt;, &lt;span class="s2"&gt;"model"&lt;/span&gt;: &lt;span class="s2"&gt;"gpt-3.5-turbo"&lt;/span&gt;, &lt;span class="s2"&gt;"k"&lt;/span&gt;: 3 &lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response is a normal OpenAI-shaped completion plus a &lt;code&gt;rag_metadata&lt;/code&gt; block that lists the&lt;br&gt;
retrieved chunks and their scores — so you can render citations and debug "why did it answer&lt;br&gt;
that" without a separate search call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"rag_metadata"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"retrieved_documents"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"documents"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"FastAPI is a modern web framework..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                  &lt;/span&gt;&lt;span class="nl"&gt;"score"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.123&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"metadata"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"source"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"docs"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There's also &lt;code&gt;/api/rag/search&lt;/code&gt; (just retrieval, no LLM), &lt;code&gt;/api/rag/stats&lt;/code&gt;, and&lt;br&gt;
&lt;code&gt;/api/rag/clear&lt;/code&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  The FAISS import is guarded — and that's the point
&lt;/h2&gt;

&lt;p&gt;On Python 3.13 / some Windows setups, &lt;code&gt;faiss-cpu&lt;/code&gt; or &lt;code&gt;numpy&lt;/code&gt; won't install cleanly. Rather than&lt;br&gt;
crash the whole service at import time, the module degrades:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;faiss&lt;/span&gt;
    &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;
    &lt;span class="n"&gt;FAISS_AVAILABLE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;FAISS_AVAILABLE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Warning: FAISS/Numpy not available: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every RAG entry point calls &lt;code&gt;_check_availability()&lt;/code&gt; first and raises a clear&lt;br&gt;
"FAISS vector store is not available" instead of an &lt;code&gt;ImportError&lt;/code&gt; three frames deep. Chat,&lt;br&gt;
guardrails, and tools keep working; only retrieval is off. Small design choice, but it's the&lt;br&gt;
difference between "my RAG doesn't work" and "my whole app won't boot."&lt;/p&gt;

&lt;h2&gt;
  
  
  Persistence
&lt;/h2&gt;

&lt;p&gt;The index is saved to &lt;code&gt;data/faiss_index&lt;/code&gt; and metadata to a matching &lt;code&gt;.pkl&lt;/code&gt; — loaded on startup,&lt;br&gt;
written after ingest. It's a single-process, on-disk store with no locking. Which is exactly the&lt;br&gt;
right scope for a starter and exactly the wrong scope for multi-instance:&lt;/p&gt;

&lt;h2&gt;
  
  
  When you outgrow it
&lt;/h2&gt;

&lt;p&gt;Honest boundaries, so you know what you're buying:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Flat index&lt;/strong&gt; → fine to ~tens of thousands of chunks, then switch to &lt;code&gt;IndexIVFFlat&lt;/code&gt;/HNSW or
a real vector DB (pgvector, Qdrant, Weaviate).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One global index&lt;/strong&gt; → no per-tenant / per-user namespaces. Multi-tenant retrieval needs
metadata filtering you'd add yourself, or a DB that does it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local disk, single writer&lt;/strong&gt; → no horizontal scaling; two instances = two divergent indexes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;k&lt;/code&gt; is a fixed int&lt;/strong&gt; → no hybrid (BM25 + dense) search, no re-ranking, no MMR
diversification. Those are the usual next levers when retrieval quality plateaus.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that is a criticism of a scaffold — it's the list of what a "RAG in one command" is&lt;br&gt;
&lt;em&gt;not&lt;/em&gt;, so you replace the right thing at the right time instead of rewriting from scratch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;chimerai create my-rag-app &lt;span class="nt"&gt;--sqlite&lt;/span&gt; &lt;span class="nt"&gt;--yes&lt;/span&gt;
&lt;span class="nb"&gt;cd &lt;/span&gt;my-rag-app
chimerai add rag
chimerai dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--sqlite&lt;/code&gt; skips Docker for the app DB; the AI service still runs as its own Python process.&lt;/p&gt;

&lt;p&gt;Repo: &lt;code&gt;github.com/armbur19-collab/chimerai-kickstart&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://chimerai.dev/blog?utm_campaign=blog-promo&amp;amp;utm_medium=content&amp;amp;utm_source=devto-03-rag-mit-chimerai-add-rag-und-faiss" rel="noopener noreferrer"&gt;ChimerAI Blog&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>rag</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Streaming an AI chat from Python FastAPI through a Next.js proxy: the four annoying details</title>
      <dc:creator>Armin Burger</dc:creator>
      <pubDate>Mon, 28 Sep 2026 14:03:53 +0000</pubDate>
      <link>https://dev.to/armin_burger_ab136b2f8bb1/streaming-an-ai-chat-from-python-fastapi-through-a-nextjs-proxy-the-four-annoying-details-1eni</link>
      <guid>https://dev.to/armin_burger_ab136b2f8bb1/streaming-an-ai-chat-from-python-fastapi-through-a-nextjs-proxy-the-four-annoying-details-1eni</guid>
      <description>&lt;p&gt;Streaming chat looks trivial in tutorials: &lt;code&gt;for chunk in response: print(chunk.text)&lt;/code&gt;. It stops&lt;br&gt;
looking trivial the moment you add the things a real product needs — a logged-in user, a&lt;br&gt;
provider that can be swapped at runtime, a cost record per request, and an error that happens&lt;br&gt;
&lt;em&gt;after&lt;/em&gt; you already sent &lt;code&gt;200 OK&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Here's the path a message takes in the ChimerAI stack, with the parts that cost me time.&lt;/p&gt;
&lt;h2&gt;
  
  
  The shape
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ChatWindow (React)
  → POST /api/v1/chat/stream        Next.js API route, TypeScript
      auth + credits + persist user message
  → POST /api/chat/stream           FastAPI, Python + LiteLLM
      async generator, yields SSE chunks
  ← data: {...}\n\n  …  data: [DONE]\n\n
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Two layers, because the LLM ecosystem is Python and the UI is TypeScript. The Next.js route is&lt;br&gt;
not a pass-through — it's where auth, persistence, and metering happen, so the Python service&lt;br&gt;
stays a stateless completion endpoint you could put behind any other frontend.&lt;/p&gt;
&lt;h2&gt;
  
  
  1. The frontend is a component, not a page
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;ChatWindow&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@chimerai/chat-ui&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;ChatPage&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;ChatWindow&lt;/span&gt;
      &lt;span class="na"&gt;apiEndpoint&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"/api/v1/chat/stream"&lt;/span&gt;
      &lt;span class="na"&gt;placeholder&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"Ask me anything..."&lt;/span&gt;
      &lt;span class="na"&gt;showModelSelector&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
      &lt;span class="na"&gt;models&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-4&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;claude-3&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;ollama/llama3&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The component handles markdown rendering, code blocks with a copy button, auto-scroll, and the&lt;br&gt;
streaming reader. If you're embedding into someone else's site rather than shipping your own&lt;br&gt;
app, there's a separate &lt;code&gt;chat-widget&lt;/code&gt; — a self-contained Web Component with Shadow DOM, mounted&lt;br&gt;
per API key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;script &lt;/span&gt;&lt;span class="na"&gt;src=&lt;/span&gt;&lt;span class="s"&gt;"https://your-app.com/widget/chat.js"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&amp;lt;/script&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;div&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;"chat"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&amp;lt;/div&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;script&amp;gt;&lt;/span&gt;&lt;span class="nx"&gt;ChimerAI&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mount&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;#chat&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;sk_live_...&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;theme&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;auto&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;&lt;span class="nt"&gt;&amp;lt;/script&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  2. Check what you can't verify: estimate tokens before the call
&lt;/h2&gt;

&lt;p&gt;You don't know the completion length before the model generates it, so a credit check can't be&lt;br&gt;
exact. The route estimates from the prompt and treats it as a precondition, then reconciles&lt;br&gt;
afterwards:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;estimateTokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;any&lt;/span&gt;&lt;span class="p"&gt;[]):&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nb"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isArray&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;totalChars&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reduce&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;sum&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ceil&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;totalChars&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;1.2&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;estimatedTokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;estimateTokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;creditResult&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;requireCredits&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;estimatedTokens&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;creditResult&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;authorized&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;createErrorResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;creditResult&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;402&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;chars / 4&lt;/code&gt; is the usual rough tokenizer rule; the &lt;code&gt;1.2&lt;/code&gt; is headroom. It's a heuristic and it&lt;br&gt;
fails in both directions — a short prompt with a long completion overshoots the balance, a&lt;br&gt;
user with exactly enough credits can get cut off mid-stream. That's acceptable for a soft&lt;br&gt;
limit; it is &lt;em&gt;not&lt;/em&gt; acceptable as billing. Actual usage is recorded from the provider's own&lt;br&gt;
token counts at the end of the stream, and that's the number that goes into &lt;code&gt;ApiUsage&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Note the ordering: permission check → credit check → &lt;em&gt;then&lt;/em&gt; create the conversation row and&lt;br&gt;
persist the user message. Doing the DB writes first means a rejected request still costs you a&lt;br&gt;
conversation with an orphan message in it.&lt;/p&gt;
&lt;h2&gt;
  
  
  3. Errors after &lt;code&gt;200 OK&lt;/code&gt; must travel inside the stream
&lt;/h2&gt;

&lt;p&gt;This is the one that bites everyone. &lt;code&gt;StreamingResponse&lt;/code&gt; sends headers immediately, so a Python&lt;br&gt;
&lt;code&gt;try/except&lt;/code&gt; around the generator body is useless — by the time the provider call fails, you've&lt;br&gt;
already committed to a 200:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@router.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/chat/stream&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;create_chat_completion_stream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ChatCompletionRequest&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# No try/except here — StreamingResponse returns 200 immediately,
&lt;/span&gt;    &lt;span class="c1"&gt;# the generator runs lazily. Errors are handled inside the generator
&lt;/span&gt;    &lt;span class="c1"&gt;# as SSE error events.
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;StreamingResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;chat_service&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create_streaming_completion&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;provider_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;...,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;...),&lt;/span&gt;
        &lt;span class="n"&gt;media_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text/event-stream&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the error handling lives &lt;em&gt;inside&lt;/em&gt; the generator and speaks the protocol:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;        &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;streaming_completion_error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="c1"&gt;# Yield error as SSE event so the client receives it
&lt;/span&gt;            &lt;span class="c1"&gt;# (raising here would be swallowed — 200 OK was already sent)
&lt;/span&gt;            &lt;span class="n"&gt;error_payload&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;server_error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}})&lt;/span&gt;
            &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;error_payload&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data: [DONE]&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The client therefore has to handle an &lt;code&gt;error&lt;/code&gt; key in the middle of an otherwise healthy stream,&lt;br&gt;
and still terminate on &lt;code&gt;[DONE]&lt;/code&gt;. If your reader assumes "chunks only contain content", a&lt;br&gt;
provider outage renders as a silently truncated answer instead of an error message — which is&lt;br&gt;
the worst possible failure mode, because it looks like the model just stopped talking.&lt;/p&gt;
&lt;h2&gt;
  
  
  4. Normalize chunks into your own schema
&lt;/h2&gt;

&lt;p&gt;The service doesn't forward provider chunks verbatim; it rebuilds each one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;            &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="k"&gt;continue&lt;/span&gt;
                &lt;span class="n"&gt;choice&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;hasattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;usage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="n"&gt;total_prompt_tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;getattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
                    &lt;span class="n"&gt;total_completion_tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;getattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;completion_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;

                &lt;span class="n"&gt;chunk_response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ChatCompletionChunk&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                    &lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chatcmpl-&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uuid4&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nb"&gt;hex&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="nb"&gt;object&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chat.completion.chunk&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="n"&gt;created&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()),&lt;/span&gt;
                    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nc"&gt;ChatCompletionChunkChoice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                        &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;choice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                        &lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nc"&gt;DeltaMessage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                            &lt;span class="n"&gt;role&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;getattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;choice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                            &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;getattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;choice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                        &lt;span class="p"&gt;),&lt;/span&gt;
                        &lt;span class="n"&gt;finish_reason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;choice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;finish_reason&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="p"&gt;)],&lt;/span&gt;
                &lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;chunk_response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;model_dump_json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two reasons this indirection is worth the boilerplate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;if not chunk.choices: continue&lt;/code&gt; — providers emit usage-only or role-only frames. Anthropic
and Ollama don't frame deltas the way OpenAI does; LiteLLM gets you most of the way, and
these edge cases are the rest.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;id=chunk.id or f"chatcmpl-{uuid4().hex[:8]}"&lt;/code&gt; — some providers omit the id on chunks. The
client uses it to correlate, so a missing id becomes a dropped message.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Usage is accumulated from whichever chunks carry it and reported once, after the stream closes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;provider_id&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;total_prompt_tokens&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;total_completion_tokens&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;provider_client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;report_usage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                    &lt;span class="n"&gt;provider_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;provider_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="n"&gt;prompt_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;total_prompt_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="n"&gt;completion_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;total_completion_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="n"&gt;endpoint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/api/chat/stream&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Swapping providers without touching the chat code
&lt;/h2&gt;

&lt;p&gt;The model isn't hardcoded in the route. The request may omit it, in which case it falls back to&lt;br&gt;
the active provider's &lt;code&gt;config.defaultModel&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// Optional — resolved after provider loading&lt;/span&gt;
&lt;span class="c1"&gt;// Model validation is deferred until after provider loading,&lt;/span&gt;
&lt;span class="c1"&gt;// so we can fall back to provider.config.defaultModel&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's why the permission check has a deferred branch (&lt;code&gt;requireModelPermission(request,&lt;br&gt;
'__deferred__')&lt;/code&gt;) — you can't validate a model you haven't resolved yet. Providers are DB rows&lt;br&gt;
with AES-256-encrypted keys, so adding Claude or a local Ollama instance is an admin-UI action,&lt;br&gt;
not a deploy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/api/providers&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Local Llama&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;ollama&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                         &lt;span class="na"&gt;baseUrl&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;http://localhost:11434&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;isActive&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;p&gt;The token estimator is the weakest link — it's a guess in the request path. Cleaner would be a&lt;br&gt;
queue/worker split: accept the request, enqueue it, stream from the worker, and enforce budget&lt;br&gt;
per user per window instead of per request. I haven't done that because at current traffic the&lt;br&gt;
heuristic's failure mode (a request rejected slightly early) is cheaper than the complexity.&lt;/p&gt;

&lt;p&gt;Also: &lt;code&gt;1.2&lt;/code&gt; is a magic number I chose. If you have real prompt/completion distributions,&lt;br&gt;
replace it with a p95 multiplier from your own logs.&lt;/p&gt;

&lt;p&gt;Repo: &lt;code&gt;github.com/armbur19-collab/chimerai-kickstart&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://chimerai.dev/blog?utm_campaign=blog-promo&amp;amp;utm_medium=content&amp;amp;utm_source=devto-02-streaming-chat-von-fastapi-durch-nextjs" rel="noopener noreferrer"&gt;ChimerAI Blog&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>nextjs</category>
      <category>python</category>
      <category>webdev</category>
    </item>
    <item>
      <title>I built a CLI that scaffolds the boring parts of an AI SaaS — here's what it actually generates</title>
      <dc:creator>Armin Burger</dc:creator>
      <pubDate>Thu, 24 Sep 2026 09:56:07 +0000</pubDate>
      <link>https://dev.to/armin_burger_ab136b2f8bb1/i-built-a-cli-that-scaffolds-the-boring-parts-of-an-ai-saas-heres-what-it-actually-generates-248f</link>
      <guid>https://dev.to/armin_burger_ab136b2f8bb1/i-built-a-cli-that-scaffolds-the-boring-parts-of-an-ai-saas-heres-what-it-actually-generates-248f</guid>
      <description>&lt;p&gt;Every AI SaaS project starts with the same two days of nothing-interesting: auth, a users&lt;br&gt;
table, an encrypted place to store provider API keys, a Prisma schema. Then you finally get to&lt;br&gt;
the part you actually wanted to build.&lt;/p&gt;

&lt;p&gt;I've been working on a CLI (&lt;code&gt;@chimerai/cli&lt;/code&gt;) that scaffolds exactly that prefix. This post is&lt;br&gt;
about what it &lt;em&gt;actually writes to disk&lt;/em&gt; — not what a landing page would claim — because that's&lt;br&gt;
the only thing that matters when you decide whether to adopt a generator.&lt;/p&gt;
&lt;h2&gt;
  
  
  The one command
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @chimerai/cli create my-ai-app
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Interactive feature selector. Defaults are auth + RBAC + admin dashboard + analytics; AI&lt;br&gt;
features are opt-in. Flags worth knowing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;chimerai create my-ai-app &lt;span class="nt"&gt;--yes&lt;/span&gt;              &lt;span class="c"&gt;# no prompts, defaults&lt;/span&gt;
chimerai create my-ai-app &lt;span class="nt"&gt;--sqlite&lt;/span&gt;           &lt;span class="c"&gt;# no Docker needed, DATABASE_URL=file:./dev.db&lt;/span&gt;
chimerai create my-ai-app &lt;span class="nt"&gt;--yes&lt;/span&gt; &lt;span class="nt"&gt;--install&lt;/span&gt;    &lt;span class="c"&gt;# + npm install&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--sqlite&lt;/code&gt; is the one I'd start with if you just want to look at the code. Without it you get&lt;br&gt;
a &lt;code&gt;docker-compose.yml&lt;/code&gt; for PostgreSQL 16 + Redis 7.&lt;/p&gt;
&lt;h2&gt;
  
  
  What's in the output
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;my-ai-app/
├── app/
│   ├── layout.tsx
│   ├── page.tsx
│   ├── api/auth/          # NextAuth routes (if auth selected)
│   └── admin/             # admin pages (if selected)
├── components/ui/         # shadcn/ui components
├── lib/
│   ├── prisma.ts          # PrismaClient singleton
│   ├── auth.ts            # auth config
│   └── encryption.ts      # API key encryption
├── prisma/schema.prisma   # only models for selected features
├── .env
├── docker-compose.yml
└── package.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The generated Prisma schema is additive per feature, which is the part I check first in any&lt;br&gt;
scaffolding tool — if it gives me six models when I asked for two, I stop trusting it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Selected&lt;/th&gt;
&lt;th&gt;Models added&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Auth&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;User&lt;/code&gt;, &lt;code&gt;Account&lt;/code&gt;, &lt;code&gt;Session&lt;/code&gt;, &lt;code&gt;VerificationToken&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RBAC&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;Role&lt;/code&gt;, &lt;code&gt;Permission&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Providers&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ModelProvider&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompts&lt;/td&gt;
&lt;td&gt;&lt;code&gt;PromptTemplate&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Analytics&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ApiUsage&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Feature dependencies are resolved by the CLI rather than left to you: &lt;code&gt;admin-dashboard&lt;/code&gt;&lt;br&gt;
requires RBAC, &lt;code&gt;model-providers&lt;/code&gt; requires auth (the encryption key lives in the auth config),&lt;br&gt;
&lt;code&gt;chat-ui&lt;/code&gt; requires &lt;code&gt;model-providers&lt;/code&gt;, &lt;code&gt;rag&lt;/code&gt; requires &lt;code&gt;model-providers&lt;/code&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  Stack, stated plainly
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Next.js 14/15 App Router, API routes serve as the backend — no separate Node service&lt;/li&gt;
&lt;li&gt;TypeScript everywhere except the AI service, which is Python (FastAPI + LiteLLM)&lt;/li&gt;
&lt;li&gt;Prisma on PostgreSQL, NextAuth for credentials + OAuth&lt;/li&gt;
&lt;li&gt;Tailwind + shadcn/ui&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The split is deliberate: the LLM orchestration layer is Python because that's where the&lt;br&gt;
ecosystem lives (LangChain, FAISS, spaCy), and everything type-safe and user-facing stays&lt;br&gt;
TypeScript. The Next.js side proxies to &lt;code&gt;AI_SERVICE_URL&lt;/code&gt; (default &lt;code&gt;http://localhost:8002&lt;/code&gt;).&lt;/p&gt;
&lt;h2&gt;
  
  
  The bit I'd point at as "not boilerplate"
&lt;/h2&gt;

&lt;p&gt;Provider management with encrypted keys, which is the thing everyone re-implements badly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/api/providers&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;OpenAI GPT-4&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;openai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;      &lt;span class="c1"&gt;// | 'anthropic' | 'ollama' | 'custom'&lt;/span&gt;
    &lt;span class="na"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;sk-...&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;isActive&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;}),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keys are stored AES-256 encrypted at rest using &lt;code&gt;ENCRYPTION_KEY&lt;/code&gt; from &lt;code&gt;.env&lt;/code&gt;, never in the&lt;br&gt;
database in plaintext. &lt;code&gt;provider: 'custom'&lt;/code&gt; takes any OpenAI-compatible &lt;code&gt;baseURL&lt;/code&gt;, which is how&lt;br&gt;
you wire up a self-hosted vLLM or an OpenRouter-style gateway without the core knowing about it.&lt;br&gt;
You can also add providers through the admin UI and test the connection before using it — small&lt;br&gt;
thing, saves an annoying debugging loop.&lt;/p&gt;

&lt;p&gt;RBAC uses &lt;code&gt;resource:action&lt;/code&gt; permission strings, checked in API routes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;requirePermission&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@/lib/auth/require-permission&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;GET&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;NextRequest&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;permissionError&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;requirePermission&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;posts:read&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;permissionError&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;permissionError&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="c1"&gt;// ...&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;admin:*&lt;/code&gt; is a wildcard. Roles carry permission lists; there's a seeded admin&lt;br&gt;
(&lt;code&gt;admin@example.com&lt;/code&gt; / &lt;code&gt;admin123&lt;/code&gt; — remove it, obviously).&lt;/p&gt;
&lt;h2&gt;
  
  
  Honest limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Packages aren't on npm yet.&lt;/strong&gt; &lt;code&gt;@chimerai/model-providers&lt;/code&gt;, &lt;code&gt;@chimerai/admin-ui&lt;/code&gt; etc. exist
in the monorepo. &lt;code&gt;chimerai create&lt;/code&gt; therefore generates &lt;em&gt;standalone&lt;/em&gt; projects with the selected
features inlined as code templates — you get readable code you own, not a dependency you
can't inspect. For the full monorepo with shared packages you'd clone the repo instead.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;chimerai add&lt;/code&gt; requires Next.js 15+ (generated route handlers use the async &lt;code&gt;params&lt;/code&gt; API).
The CLI detects the version and warns below 15. Standalone &lt;code&gt;create&lt;/code&gt; output runs on port 3001
by default to avoid colliding with another dev server.&lt;/li&gt;
&lt;li&gt;The feature set is opinionated toward the SaaS shape (auth, providers, prompts, chat, RAG,
admin). If your app doesn't need that shape, a plain &lt;code&gt;create-next-app&lt;/code&gt; plus a library like
LangChain is a perfectly reasonable starting point — pick the generator that matches the
problem.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Adding to an existing project instead
&lt;/h2&gt;

&lt;p&gt;You don't have to start from a generated project:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;my-nextjs-app
npx chimerai add auth
npx chimerai add model-providers
npx chimerai add chat-ui
pnpm &lt;span class="nb"&gt;install
&lt;/span&gt;npx prisma db push &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; npx prisma db seed
pnpm dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The CLI finds the project root by walking up for a &lt;code&gt;.chimerai&lt;/code&gt; marker, then falls back to a&lt;br&gt;
registry in &lt;code&gt;~/.chimerai/projects.json&lt;/code&gt;, then &lt;code&gt;--dir&lt;/code&gt;. In a monorepo, point &lt;code&gt;--dir&lt;/code&gt; at the&lt;br&gt;
Next.js app (&lt;code&gt;apps/frontend&lt;/code&gt;), not the workspace root — the root has no &lt;code&gt;app/&lt;/code&gt; and no &lt;code&gt;prisma/&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;chimerai doctor&lt;/code&gt; runs health checks on env vars, DB connectivity, and installed components,&lt;br&gt;
which is mostly there because I forgot to set &lt;code&gt;ENCRYPTION_KEY&lt;/code&gt; often enough to be embarrassing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next in this series
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Streaming chat: SSE from FastAPI through a Next.js proxy route&lt;/li&gt;
&lt;li&gt;RAG with &lt;code&gt;chimerai add rag&lt;/code&gt; — FAISS index, chunking, and the &lt;code&gt;/api/rag/query&lt;/code&gt; round trip&lt;/li&gt;
&lt;li&gt;Prompt templates in the DB instead of in git, and what breaks when you do that&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Repo: &lt;code&gt;github.com/armbur19-collab/chimerai-kickstart&lt;/code&gt; · site: chimerai.dev&lt;/p&gt;

&lt;p&gt;&lt;a href="https://chimerai.dev/blog?utm_campaign=blog-promo&amp;amp;utm_medium=content&amp;amp;utm_source=devto-01-chimerai-create-was-die-cli-wirklich-generiert" rel="noopener noreferrer"&gt;ChimerAI Blog&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>typescript</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Stop Counting Requests: The Case for Token-Based Quotas in LLM SaaS</title>
      <dc:creator>Armin Burger</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:42:23 +0000</pubDate>
      <link>https://dev.to/armin_burger_ab136b2f8bb1/stop-counting-requests-the-case-for-token-based-quotas-in-llm-saas-3gb9</link>
      <guid>https://dev.to/armin_burger_ab136b2f8bb1/stop-counting-requests-the-case-for-token-based-quotas-in-llm-saas-3gb9</guid>
      <description>&lt;p&gt;If you are building a multi-tenant AI application, your current rate-limiting strategy is likely broken. Traditional "requests per minute" (RPM) metrics, inherited from standard REST APIs, fail catastrophically when applied to Large Language Models. This isn't just an imprecision; it is a fundamental architectural error that systematically throttles efficient short queries while allowing high-cost massive contexts to bypass limits undetected.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why RPM Metrics Fail for LLMs
&lt;/h3&gt;

&lt;p&gt;The core issue is cost variance. In traditional APIs, the computational cost of a request is relatively uniform. In LLMs, cost is driven by token count, not frequency. A single request with a 100k-token context window can cost hundreds of times more than a simple 50-token query. If you limit users by request count, you inadvertently penalize lightweight users while giving heavy users free rein to drain resources until the provider’s account-level quota is exhausted.&lt;/p&gt;

&lt;p&gt;This leads to the "noisy neighbor" effect. Most LLM providers enforce rate limits at the organization or account level, not per user. Without internal controls, one tenant’s batch processing spike can throttle API access for every other customer on your platform, causing widespread service degradation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tokens ≠ Costs: The Complexity Layer
&lt;/h3&gt;

&lt;p&gt;Even shifting to token counts introduces inaccuracies because tokens do not map linearly to dollars. Three factors complicate this relationship:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Input vs. Output Pricing:&lt;/strong&gt; Output tokens are often 3 to 5 times more expensive than input tokens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model Variance:&lt;/strong&gt; Tenants using premium models can incur 3 to 4 times the costs of those using cheaper models, even with identical usage patterns.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompt Caching:&lt;/strong&gt; Repeated context blocks may be charged at a fraction of the regular price upon cache hits, meaning two requests with the same token count can have vastly different costs.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Therefore, financial control requires managing &lt;strong&gt;token credits&lt;/strong&gt; rather than raw token counts. You must normalize costs across different models and pricing tiers into a common currency for your internal ledger.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Reservation Pattern: Solving the Temporal Problem
&lt;/h3&gt;

&lt;p&gt;A critical challenge in LLM billing is temporal uncertainty. You only know the exact cost after generation completes. If you wait for completion to charge, a user could trigger a long-running stream that exceeds their budget, resulting in negative balances.&lt;/p&gt;

&lt;p&gt;The solution is a reservation pattern, similar to credit card pre-authorizations. Before calling the provider, your system should:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Estimate the maximum possible cost using &lt;code&gt;max_tokens&lt;/code&gt; as a conservative upper bound.&lt;/li&gt;
&lt;li&gt;Reserve these funds in an internal credit ledger (&lt;code&gt;creditLedger.reserve&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Execute the API call.&lt;/li&gt;
&lt;li&gt;Settle the exact cost based on the provider’s actual usage report, releasing any excess reservation.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If the call fails, the reservation is released immediately. This ensures that no tenant can spend more than they have authorized, even during streaming processes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Implementation Strategy: Two-Tier Limits
&lt;/h3&gt;

&lt;p&gt;To maintain stability, implement a two-tier limiting structure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Internal Per-Tenant Sub-Limits:&lt;/strong&gt; Use a token bucket algorithm tied to the user’s credit balance to prevent individual overspending.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Global Provider Quota Checks:&lt;/strong&gt; Monitor the aggregate usage against the provider’s account limits to prevent the noisy neighbor effect.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Crucially, always base final cost tracking on the actual usage data returned in the provider’s API response. Internal estimates are useful for reservations, but discrepancies between your bookkeeping and provider invoices signal bugs in your cost logic, not mere estimation errors. Real-time monitoring of usage fidelity is essential for accurate billing.&lt;/p&gt;

&lt;p&gt;By shifting from request counting to token-based credit management with reservation patterns, you align your architecture with the economic reality of LLMs, ensuring both financial accuracy and service reliability for all tenants.&lt;/p&gt;

&lt;p&gt;Repo: &lt;code&gt;github.com/armbur19-collab/chimerai-app&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://chimerai.dev/blog?utm_campaign=blog-promo&amp;amp;utm_medium=content&amp;amp;utm_source=devto-05-token-costs-per-tenant-why-rate-limiting-works-dif_devto" rel="noopener noreferrer"&gt;ChimerAI Blog&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>sass</category>
    </item>
    <item>
      <title>Beyond `tenant_id`: Why Classical Multi-Tenancy Fails for RAG Systems</title>
      <dc:creator>Armin Burger</dc:creator>
      <pubDate>Fri, 18 Sep 2026 12:09:46 +0000</pubDate>
      <link>https://dev.to/armin_burger_ab136b2f8bb1/beyond-tenantid-why-classical-multi-tenancy-fails-for-rag-systems-1off</link>
      <guid>https://dev.to/armin_burger_ab136b2f8bb1/beyond-tenantid-why-classical-multi-tenancy-fails-for-rag-systems-1off</guid>
      <description>&lt;p&gt;Most engineering teams assume that because they have implemented row-level security (RLS) and a &lt;code&gt;tenant_id&lt;/code&gt; column in their relational database, their application is securely multi-tenant. This assumption holds true for traditional CRUD applications. However, when you integrate Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) pipelines, classical isolation patterns become dangerously insufficient.&lt;/p&gt;

&lt;p&gt;The core issue is that AI systems introduce new attack surfaces and failure modes that do not exist in standard SaaS architectures. Here are four critical frontiers where traditional multi-tenancy breaks down, along with the architectural adjustments required to secure them.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Vector Store Isolation: The Silent Leak Risk
&lt;/h2&gt;

&lt;p&gt;Approximate Nearest Neighbor (ANN) indexes, which power most vector databases, lack default access controls comparable to SQL. If you rely solely on post-filtering results by &lt;code&gt;tenant_id&lt;/code&gt;, you risk exposing data during the search phase or leaking metadata through similarity scores.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Solution:&lt;/strong&gt; You must enforce isolation at the index level. Use namespace models or explicit pre-filtering strategies within the vector store itself. Treating the tenant boundary as an absolute filter component—rather than just a metadata tag—is essential to prevent cross-tenant data leaks.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Semantic Caching: Keys Must Include Tenant ID
&lt;/h2&gt;

&lt;p&gt;Caching is often the most overlooked frontier for security breaches in AI apps. Semantic caches retrieve answers based on question similarity rather than exact matches. If your cache key does not include the &lt;code&gt;tenant_id&lt;/code&gt; as a hard component, a query from Tenant A might match a cached response generated for Tenant B.&lt;/p&gt;

&lt;p&gt;This leads to silent failures: users receive plausible-sounding but incorrect answers derived from another company’s data. There are no error codes or crashes; the system simply returns wrong information.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Solution:&lt;/strong&gt; Ensure the tenant ID is a mandatory part of every semantic cache key. Maintain per-tenant cache spaces instead of using a single global cache where tenant identity is treated merely as a similarity signal.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Token Cost Management: Beyond Request Rate Limiting
&lt;/h2&gt;

&lt;p&gt;Traditional rate limiting counts requests per minute. In RAG products, this is inadequate because token costs vary significantly by provider, model, and direction (input vs. output). A "noisy neighbor" can consume disproportionate resources with a few complex prompts, impacting latency and cost for all other tenants.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Solution:&lt;/strong&gt; Shift from request counting to token usage tracking. Implement reservation patterns to manage resource allocation effectively, ensuring that one tenant’s heavy usage does not degrade service quality for others.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Context Assembly and Guardrails
&lt;/h2&gt;

&lt;p&gt;Hardcoding global middleware for guardrails fails when different tenants have varying compliance needs and latency tolerances. Furthermore, concurrency issues in prompt assembly can lead to catastrophic mixing of sensitive data if module-global state is used.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Solution:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Dynamic Guardrails:&lt;/strong&gt; Configure PII filters and prompt-injection defenses per tenant.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Strict Scoping:&lt;/strong&gt; Avoid object sharing between parallel requests. Place base system instructions before the cache boundary, and keep all tenant-specific context (names, configs, RAG data) strictly after it.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Granular Observability:&lt;/strong&gt; Global average metrics mask individual tenant failures. Measure evaluation metrics like RAGAS faithfulness at the tenant granularity to detect subtle quality degradation early.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Security architecture for AI requires moving beyond database-level isolation. It must encompass vector indices, prompt construction logic, and API-level caching mechanisms. Thinking through each dimension individually during design beats bundling them under a generic "multi-tenancy" label, because failure modes in AI systems are silent and subtle, not obvious crashes.&lt;/p&gt;

&lt;p&gt;Repo: &lt;code&gt;github.com/armbur19-collab/chimerai-app&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://chimerai.dev/blog?utm_campaign=blog-promo&amp;amp;utm_medium=content&amp;amp;utm_source=devto-04-multi-tenancy-for-ai-rag-systems-the-landscape-beh_devto" rel="noopener noreferrer"&gt;ChimerAI Blog&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>rag</category>
      <category>security</category>
    </item>
    <item>
      <title>Stop Hardcoding Your AI Guardrails: A Multi-Tenant Configuration Strategy</title>
      <dc:creator>Armin Burger</dc:creator>
      <pubDate>Tue, 15 Sep 2026 15:36:15 +0000</pubDate>
      <link>https://dev.to/armin_burger_ab136b2f8bb1/stop-hardcoding-your-ai-guardrails-a-multi-tenant-configuration-strategy-5bha</link>
      <guid>https://dev.to/armin_burger_ab136b2f8bb1/stop-hardcoding-your-ai-guardrails-a-multi-tenant-configuration-strategy-5bha</guid>
      <description>&lt;p&gt;Most early-stage AI products make a critical architectural mistake: they implement security guardrails as global, hardcoded middleware. While this works for a single use case, it becomes a liability the moment you onboard a second enterprise customer with different regulatory needs.&lt;/p&gt;

&lt;p&gt;The core problem is that "one-size-fits-all" security fails in multi-tenant environments. A healthcare tenant requires strict PII detection and high latency tolerance, while a developer tools tenant prioritizes low-latency responses and minimal false positives. When these conflicting requirements meet a static codebase, you face risky rewrites instead of simple config changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Pipeline Model
&lt;/h2&gt;

&lt;p&gt;Instead of treating guardrails as binary on/off feature flags, model them as an ordered, parameterized pipeline. This allows each tenant to define their own sequence of checks for both input and output streams.&lt;/p&gt;

&lt;p&gt;A robust configuration structure looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"input"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pii_detection"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"sensitivity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"high"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"redact"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"toxicity_check"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"sensitivity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"medium"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"flag"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"harmful_content"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"sensitivity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"low"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"block"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;By loading this JSON from a database per tenant, you enable adjustments without code deployments. Actions like &lt;code&gt;block&lt;/code&gt;, &lt;code&gt;flag&lt;/code&gt;, or &lt;code&gt;redact&lt;/code&gt; provide granular control over how violations are handled.&lt;/p&gt;

&lt;h2&gt;
  
  
  Handling Streaming Output
&lt;/h2&gt;

&lt;p&gt;Input guardrails are relatively straightforward because you can analyze the full prompt before processing. Output guardrails present a unique architectural conflict: real-time streaming UX versus holistic text analysis.&lt;/p&gt;

&lt;p&gt;Waiting for the entire response to finish defeats the purpose of streaming, but checking every token individually misses context-dependent toxicity. The solution is a middle-ground approach: process output in sentences or chunks. If a chunk triggers a rule, abort the stream immediately. This balances user experience with security compliance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Audit Trails and Compliance
&lt;/h2&gt;

&lt;p&gt;Guardrails are not just blockers; they are evidence generators. Without granular audit logs containing confidence scores and specific triggered rules, you cannot prove compliance to auditors—you can only claim it exists.&lt;/p&gt;

&lt;p&gt;Implementing tenant-specific "golden sets" for regression testing ensures that configuration changes do not inadvertently break security policies. Furthermore, sensitivity thresholds should be treated as business risk decisions configurable by the tenant, not fixed by the platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Two-Tier Configuration Model
&lt;/h2&gt;

&lt;p&gt;To balance usability with legal safety, adopt a two-tier model:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Coarse Presets:&lt;/strong&gt; Allow self-service tenants to choose from predefined security profiles (e.g., "Standard," "Strict").&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Granular Overrides:&lt;/strong&gt; Enable Enterprise customers to tweak specific parameters via approval workflows.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Hardcoding guardrails globally is not just bad practice; it is a guaranteed path to losing enterprise contracts. By shifting to a configurable, per-tenant pipeline, you transform security from a rigid constraint into a flexible product feature.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>security</category>
      <category>systemdesign</category>
    </item>
  </channel>
</rss>
