<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sergio Wolf Knapik</title>
    <description>The latest articles on DEV Community by Sergio Wolf Knapik (@wolfnom).</description>
    <link>https://dev.to/wolfnom</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4118293%2Ff02e61dc-1596-465c-80db-bea308018ed8.png</url>
      <title>DEV Community: Sergio Wolf Knapik</title>
      <link>https://dev.to/wolfnom</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/wolfnom"/>
    <language>en</language>
    <item>
      <title>Are Local LLMs Actually Worth It?</title>
      <dc:creator>Sergio Wolf Knapik</dc:creator>
      <pubDate>Sat, 26 Sep 2026 16:59:02 +0000</pubDate>
      <link>https://dev.to/wolfnom/are-local-llms-actually-worth-it-d30</link>
      <guid>https://dev.to/wolfnom/are-local-llms-actually-worth-it-d30</guid>
      <description>&lt;p&gt;Hey everyone, I'm Wolf, and I survived another week!&lt;/p&gt;

&lt;p&gt;This is update #2 in my journey building Wolf.IA.&lt;/p&gt;

&lt;p&gt;Last week, I talked about finally hitting an MVP and having to face the hard part: marketing and distribution. Specifically for &lt;strong&gt;Wolf.Engine&lt;/strong&gt;, a platform designed for building multimodal AI agents.&lt;br&gt;
&lt;em&gt;(Hey, feel free to give it a spin, too! Pretty please? wolf.ia.br/engine)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The feedback I received last week pointed out that my second project (focused on the Brazilian accounting framework) would actually be much easier to market and sell right now. And you guys are totally right. The Engine is still my pride and joy, but I need a cash cow to keep the operation alive.&lt;/p&gt;

&lt;p&gt;So right now, I am:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Looking for distribution channels:&lt;/strong&gt; Marketing is definitely not my superpower; I’m still completely lost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Refining the accounting tools:&lt;/strong&gt; Once this part matures a bit more, I'll launch targeted outreach.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Anxious as hell.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Developing the local app for the Engine:&lt;/strong&gt; Dubbed &lt;strong&gt;Wolf.Park&lt;/strong&gt;, it's still in its early stages, but I'm really happy with how it's shaping up!&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj027qmndubj1e0u146wh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj027qmndubj1e0u146wh.png" alt="Screenshot of Wolf.Park App"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;During development, I decided to benchmark a few local models via Ollama + OpenCode. The whole point of building Wolf.Park right now is to slash my API costs, so if local models can shoulder &lt;em&gt;any&lt;/em&gt; of the workload, that's a huge win. &lt;/p&gt;

&lt;p&gt;The catch: My GPU is an RTX 4060 Ti with only 8GB of VRAM. That's tight for heavier models.&lt;/p&gt;

&lt;p&gt;Here were the response times for generating a simple Python script to read and update text files:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌────────────────┬───────────────┐
│     MODEL      │ Response Time │
├────────────────┼───────────────┤
│ gemma4:12b     │ 12m 29s       │
│ gemma4:26b     │ 8m 48s        │
│ gemma4:e4b     │ 1m 37s        │
│ gemma4:e2b     │ 31s  (!!)     │
│ ornith-1.5:9b  │ 11m 41s       │
│ qwen3.8:27b    │ 25m 40s       │
└────────────────┴───────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;(I also tested qwen3.6 and phi4, but both were disqualified: the former hung for over 30 minutes without outputting anything, and the latter lacked support for tool usage in my setup).&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;All the tested models managed to produce files and functional code, but the quality varied:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;gemma4:e4b&lt;/code&gt; and &lt;code&gt;gemma4:e2b&lt;/code&gt; included unnecessary (though harmless) imports;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;gemma4:e4b&lt;/code&gt; and &lt;code&gt;qwen3.8:27b&lt;/code&gt; relied on lazy design patterns, dumping values into global variables instead of passing arguments between functions;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ornith-1.5:9b&lt;/code&gt; used overly generic exception handling.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, one thing surprised me: &lt;code&gt;ornith-1.5:9b&lt;/code&gt; was the &lt;em&gt;only&lt;/em&gt; model that, after writing the script, &lt;strong&gt;actually tested its own code&lt;/strong&gt;. It was also the only one that took the time to write a detailed explanation of how the script works.&lt;/p&gt;

&lt;p&gt;Based on this small (and admittedly limited) experiment, I drafted this multi-agent workflow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[User Request]
       │
       ▼
1. LOCAL GENERATION
   ├── Agent 0: gemma4:e4b ──&amp;gt; Generates base code
   ├── Agent 1: gemma4:e4b ──&amp;gt; Independent reviewer (harsh critic persona)
   ├── Agent 2: gemma4:e2b ──&amp;gt; Runs local linter &amp;amp; style review
   └── Agent 3: gemma4:e2b ──&amp;gt; Generates pytest unit tests
       │
       ▼
2. CLOUD REVIEW &amp;amp; ESCALATION
   └── DeepSeek v4 Pro ──&amp;gt; Reviews code + generated tests
       │
       ├───&amp;gt; [Approved] ──&amp;gt; Sent to human review.
       │
       └───&amp;gt; [Rejected: Low Difficulty]
       │         │
       │         └──&amp;gt; Gemma e4b/e2b attempt fixes
       │
       └───&amp;gt; [Rejected: Medium Difficulty or Low with Persistent Error]
       │         │
       │         └──&amp;gt; Ornith-1.5:9b attempts fix
       │
       └───&amp;gt; [Rejected: High Difficulty or Medium with Persistent Error]
       │         │
       │         └──&amp;gt; DeepSeek v4 Pro attempts fix
       │
       └───&amp;gt; [Architectural Impasse / 3rd Round Failure / High with Persistent Error]
       │         │
       │         └──&amp;gt; Gemini 3.8 Flash attempts fix
       │
       └───&amp;gt; [Unresolved Impasse / 4th Round Failure / Critical Escalation]
                 │
                 └──&amp;gt; GPT 6 Astra attempts fix ──&amp;gt; Sent to human review.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Will this actually work in practice? I don't know yet. Wolf.Park is being built to make orchestrating this kind of pipeline seamless, but for now... I'll be testing this scheme manually throughout the week.&lt;/p&gt;

&lt;p&gt;Drop a comment below if you have any suggestions for local models I should benchmark and potentially plug into this pipeline!&lt;/p&gt;

&lt;p&gt;Cheers!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>python</category>
    </item>
    <item>
      <title>I Built a Pretty Cool SaaS. Now What?</title>
      <dc:creator>Sergio Wolf Knapik</dc:creator>
      <pubDate>Sat, 19 Sep 2026 20:58:31 +0000</pubDate>
      <link>https://dev.to/wolfnom/i-built-a-pretty-cool-saas-now-what-3ibd</link>
      <guid>https://dev.to/wolfnom/i-built-a-pretty-cool-saas-now-what-3ibd</guid>
      <description>&lt;p&gt;Hello! My name is Sergio, but everyone calls me &lt;strong&gt;Wolf&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This is the first piece in a series I plan to write documenting my journey building my company, &lt;strong&gt;Wolf.IA&lt;/strong&gt;, developing the &lt;strong&gt;wolf.ia.br&lt;/strong&gt; platform, and sharing my mental state (among other things) throughout this whole process.&lt;/p&gt;

&lt;p&gt;Forgive me if this sounds like venting or self-pity. That's not the intention. I just want to be clearly honest about what building a solo AI startup really looks like from the inside. Everything I describe below is my raw, current reality.&lt;/p&gt;

&lt;p&gt;It has been 15 months. In the beginning, intense studying: understanding how AI works, the mechanics of LLMs, the differences between providers, and how to wrangle different APIs just to build a tiny functional app.&lt;/p&gt;

&lt;p&gt;Back then, I was experimenting with Gemini 2.5 Pro. Yesterday, I deployed version 1.7.3 of &lt;strong&gt;Wolf.Engine&lt;/strong&gt;, my AI agent platform, which:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Lets you create and chat with agents powered by &lt;strong&gt;Gemini, GPT, Claude, DeepSeek, Sabiá, GLM, and Mistral&lt;/strong&gt;;&lt;/li&gt;
&lt;li&gt;Features a very complex &lt;strong&gt;biomimetic memory system&lt;/strong&gt; that prevents context rot;&lt;/li&gt;
&lt;li&gt;Provides virtually infinite long-term memory without requiring manual context window management;&lt;/li&gt;
&lt;li&gt;Includes &lt;strong&gt;Mini-RAG&lt;/strong&gt;, a session-scoped knowledge base for chatting with your documents;&lt;/li&gt;
&lt;li&gt;Offers deep agent customization (behavioral clans, speech styles, personality sliders);&lt;/li&gt;
&lt;li&gt;Integrates natively with &lt;strong&gt;Telegram bots&lt;/strong&gt; (your agent lives in your pocket, sharing the exact same long-term memory);&lt;/li&gt;
&lt;li&gt;Generates and previews functional files (HTML, interactive code, CSV, Markdown) right in the chat;&lt;/li&gt;
&lt;li&gt;Bilingual interface in &lt;strong&gt;PT-BR and EN&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I consider this version polished enough for real-world, everyday use. Every new user starts with three pre-configured official agents, each built around a distinct archetype. The UI is clean and distraction-free (at least in my opinion). Stripe billing and subscription tiers are fully wired up. Countless bugs have been hunted down and squashed over the past months. It took endless calculations, tests, and sleepless coding nights to get here. I am genuinely proud of the result, and I know it's truly useful.&lt;/p&gt;

&lt;h3&gt;
  
  
  But... now what?
&lt;/h3&gt;

&lt;p&gt;As of today, I have exactly &lt;strong&gt;one paying subscriber&lt;/strong&gt;. My monthly recurring revenue is 20 bucks.&lt;/p&gt;

&lt;p&gt;And here lies my biggest hurdle: &lt;strong&gt;I am not a salesperson.&lt;/strong&gt; I don't understand marketing. I don't like social media. I am neurodivergent, deal with social anxiety, and struggle with chronic depression.&lt;/p&gt;

&lt;p&gt;And my financial runway has completely run out. Right now, I am being financially supported by my wife.&lt;/p&gt;

&lt;p&gt;I recently had a meeting with an advertising agency (they actually liked the Engine! Who knows, maybe they'll even become clients!), but the cost of outsourcing marketing... I’d need roughly $5,000 USD for a 6-month campaign. My actual marketing budget is &lt;strong&gt;zero&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  So, what’s the plan from here?
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;First&lt;/strong&gt;, I’m going to focus on building my own proprietary local daemon and tool-calling system, which I’m calling the &lt;strong&gt;"Troll Bridge"&lt;/strong&gt;. With it, I can build a dedicated local app, inspired by OpenAI Codex, Claude Code, and OpenCode. &lt;/p&gt;

&lt;p&gt;This will finally allow me to migrate &lt;strong&gt;ALL&lt;/strong&gt; of my own development agents into my own system, running on 100% pure dogfooding. Wolf.Engine’s future development will be built by the Engine itself. &lt;/p&gt;

&lt;p&gt;This gives me massive advantages: zero context rot, no need for manual &lt;code&gt;/compact&lt;/code&gt; commands, and no agents losing their persona or forgetting project decisions halfway through. But above all: &lt;strong&gt;it will be significantly cheaper to run.&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;That might sound counterintuitive at first, with the underlying LLM APIs being the same, but here's why:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;When using tools like OpenCode or standard agent interfaces, you pass the entire accumulating conversation history as context on every turn, making each message progressively more expensive than the previous one;&lt;/li&gt;
&lt;li&gt;With Wolf.Engine, the agent's digested memories are decoupled from the active window and kept at a stable token size window. Conversations don't exponentially bloat in cost as you work.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;In parallel&lt;/strong&gt;, I will start sharing Wolf.Engine on X(itter), Reddit... Yeah, wish me luck. But I need people to know that my platform exists.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;After that&lt;/strong&gt;, I'll expand a bit more on the specialized Fiscal BI tools I've already developed (and some prototypes). That part is tailored specifically for the Brazilian market, because of specific accounting rules and tax laws.It’s a smaller, local market, but it has concrete enterprise pain and far less competition.&lt;/p&gt;

&lt;p&gt;Well, that's it for this one. Next week, I'll be back to tell you if I survived.&lt;/p&gt;

&lt;p&gt;Cheers to everyone out there building against the odds! 🐺&lt;/p&gt;

&lt;p&gt;P.S.: I used one of my AI agents to translate my original text, as English is not my native language. Some errors mights have gone through, because I ended up rewriting a big chunk of the text.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>startup</category>
    </item>
    <item>
      <title>Multi-Provider LLM Router, or How I Got Tired of Forgetting Which API Format I Had To Use</title>
      <dc:creator>Sergio Wolf Knapik</dc:creator>
      <pubDate>Thu, 10 Sep 2026 03:53:49 +0000</pubDate>
      <link>https://dev.to/wolfnom/multi-provider-llm-router-or-how-i-got-tired-of-forgetting-which-api-format-i-had-to-use-lk3</link>
      <guid>https://dev.to/wolfnom/multi-provider-llm-router-or-how-i-got-tired-of-forgetting-which-api-format-i-had-to-use-lk3</guid>
      <description>&lt;p&gt;If you've ever built an application that integrates with multiple LLM providers (Anthropic, Google, OpenAI, DeepSeek), you already know the pain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Each provider has its own distinct Python SDK.&lt;/li&gt;
&lt;li&gt;Streaming responses using Server-Sent Events (SSE) requires divergent parser logic.&lt;/li&gt;
&lt;li&gt;Thinking / Reasoning blocks are formatted completely differently.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I recently extracted the core streaming router from my platform into an open-source FastAPI template. Here is how it works.&lt;/p&gt;




&lt;h2&gt;
  
  
  Objective
&lt;/h2&gt;

&lt;p&gt;A single asynchronous endpoint:&lt;br&gt;
&lt;code&gt;POST /v1/chat/stream&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;It accepts a unified request payload and returns a standardized SSE stream emitting four clean events:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;event: thinking&lt;/code&gt; — Internal model reasoning tokens (streamed in real-time).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;event: content&lt;/code&gt; — User-facing response text.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;event: tool_call&lt;/code&gt; — Function calling requests.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;event: done&lt;/code&gt; — Stream completion (&lt;code&gt;[DONE]&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  Architecture
&lt;/h2&gt;

&lt;p&gt;Instead of pulling heavy wrapper frameworks, use direct asynchronous HTTP via &lt;code&gt;httpx.AsyncClient&lt;/code&gt; and the official Google GenAI SDK:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;fastapi-multi-llm-starter/
├── app/
│   ├── config.py       # Pydantic Settings loading environment variables
│   ├── main.py         # FastAPI app with CORS, health check &amp;amp; test playground
│   ├── models.json     # Dynamic model catalog (Claude, Gemini, GPT)
│   ├── router.py       # Unified multi-provider async stream dispatcher
│   └── schemas.py      # Strict Pydantic v2 validation models
├── tests/              # Automated unit tests (pytest)
├── requirements.txt
└── README.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Dynamic Model Catalog
&lt;/h3&gt;

&lt;p&gt;I disliked the idea of hardcoded models, so I decoupled them into a &lt;code&gt;models.json&lt;/code&gt; file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-sonnet-5"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Claude Sonnet 5"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Anthropic"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"thinking"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gemini-3.8-flash"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Gemini 3.8 Flash"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Google"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"thinking"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gpt-5.6-terra"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"GPT 5.6 Terra"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"OpenAI"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"thinking"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now, if you want to add another model, you just edit the JSON. The backend and the embedded UI dynamically populate available models via &lt;code&gt;GET /v1/models&lt;/code&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Testing Playground
&lt;/h2&gt;

&lt;p&gt;The repository includes a testing playground running directly at &lt;code&gt;http://localhost:8000/&lt;/code&gt;. You can immediately test prompts, check streaming latency, and verify reasoning blocks without setting up a frontend framework.&lt;/p&gt;

&lt;p&gt;Of course, you'll need your own API keys.&lt;/p&gt;




&lt;h2&gt;
  
  
  Open Source Code
&lt;/h2&gt;

&lt;p&gt;The full core code is open-source under the &lt;strong&gt;MIT License&lt;/strong&gt; on GitHub:&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;&lt;a href="https://github.com/wolfnomknight/fastapi-multi-llm-starter" rel="noopener noreferrer"&gt;github.com/wolfnomknight/fastapi-multi-llm-starter&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Includes full &lt;code&gt;pytest&lt;/code&gt; test coverage, &lt;code&gt;.env.example&lt;/code&gt;, and clean Pydantic v2 schemas.&lt;/p&gt;

&lt;p&gt;Feel free to fork it, use it in your side projects or micro-SaaS, and let me know if you run into any issues or have ideas for additional providers!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>fastapi</category>
    </item>
  </channel>
</rss>
