<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Jyotiraditya Singh</title>
    <description>The latest articles on DEV Community by Jyotiraditya Singh (@jyotir07).</description>
    <link>https://dev.to/jyotir07</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4007823%2Ff21112a5-ef6e-4518-b16f-1563ef4a92e2.jpg</url>
      <title>DEV Community: Jyotiraditya Singh</title>
      <link>https://dev.to/jyotir07</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jyotir07"/>
    <language>en</language>
    <item>
      <title>SECONDO: An Offline Baking Planner for My Friend's Café, Built on Gemma and TabPFN</title>
      <dc:creator>Jyotiraditya Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 00:02:52 +0000</pubDate>
      <link>https://dev.to/jyotir07/secondo-an-offline-baking-planner-for-my-friends-cafe-built-on-gemma-and-tabpfn-517g</link>
      <guid>https://dev.to/jyotir07/secondo-an-offline-baking-planner-for-my-friends-cafe-built-on-gemma-and-tabpfn-517g</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;SECONDO&lt;/strong&gt; is a local-first planning tool for very small food businesses: home bakers, tiffin services, one-counter cafés. It takes sales history and messy WhatsApp orders and turns them into a &lt;strong&gt;daily baking plan&lt;/strong&gt;. The owner reviews it, edits it and approves it. Every number in the plan shows the reason behind it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who it's for.&lt;/strong&gt; My friend Sidd runs a small café in Bangalore. Most pre-orders arrive on WhatsApp, and every evening Sidd has to decide how much of each item to prepare for the next day. In Sidd's words: &lt;em&gt;"I usually just look at the orders and estimate how much I need to prepare."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That estimate depends on information scattered across several places:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pre-orders are buried in WhatsApp.&lt;/strong&gt; A message like &lt;em&gt;"2 sourdough and half a dozen cinnamon rolls for Saturday, eggless please"&lt;/em&gt; sits between supplier chats and family groups, and &lt;em&gt;"sometimes people change things at the last minute."&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Walk-in demand is a gut feeling.&lt;/strong&gt; Weekdays differ from weekends, and some days are just busier. The owner knows these patterns only roughly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A wrong guess costs money either way.&lt;/strong&gt; Bake too much and butter, flour and hours of work go to waste. Bake too little and a regular walks out without their croissant.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Proper forecasting tools are built for chains. They're expensive, they're made for analysts, and they want a café's sales and its customers' names in someone else's cloud. What Sidd needs is a sensible number for each tray and the reason behind that number.&lt;/p&gt;

&lt;p&gt;SECONDO works in four steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Understand.&lt;/strong&gt; Import past sales from a CSV, with row-level validation and duplicate detection. Paste a WhatsApp order and &lt;strong&gt;Gemma, running locally&lt;/strong&gt;, pulls out the products, quantities, date, customer and dietary notes. The owner checks the result before anything is saved.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Predict.&lt;/strong&gt; &lt;strong&gt;TabPFN&lt;/strong&gt; forecasts tomorrow's demand for each product, with an 80% range. It is compared against simple baselines in a backtest where no model ever sees the future.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide.&lt;/strong&gt; A kitchen plan lists a quantity per item with its reasoning: the forecast, the range, recent same-weekday sales and pre-orders. It also includes warnings, dietary requests and ingredient totals.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Act.&lt;/strong&gt; The owner changes any number they disagree with, then approves or rejects the plan. &lt;strong&gt;Nothing is automatic.&lt;/strong&gt; The plan stays a draft until a person approves it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwwlpwjtljejqerwjjas2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwwlpwjtljejqerwjjas2.png" alt="Kitchen plan" width="800" height="556"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What Sidd said
&lt;/h3&gt;

&lt;p&gt;I showed SECONDO to Sidd and asked four questions.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What do you think of SECONDO?&lt;/strong&gt;&lt;br&gt;
"Honestly, this could be pretty useful. I usually just look at the orders and estimate how much I need to prepare. Having something organise that for me would save time, especially on busy days."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Would you actually use it?&lt;/strong&gt;&lt;br&gt;
"Yeah, if it's simple enough. I don't want to spend 20 minutes entering orders just to get a list of what I need to bake. If it can take my existing orders and make sense of them, that's where I'd find it useful."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's the biggest issue?&lt;/strong&gt;&lt;br&gt;
"Probably getting the orders in. Most of them come through WhatsApp, and sometimes people change things at the last minute. I'd want to be able to correct something quickly."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What would make you trust its recommendations?&lt;/strong&gt;&lt;br&gt;
"If it tells me to make 15 croissants, I'd want to know why 15. Is it based on previous orders or just guessing? I'd rather have a number I can adjust than something that acts like it knows everything."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Some of these answers line up with design decisions SECONDO already makes. Others point to things it doesn't do yet:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;"I'd want to know why 15."&lt;/strong&gt; Every line in the plan shows its reasoning: the forecast, its range, what sold on recent same weekdays and what's already pre-ordered. Sidd's question, "previous orders or just guessing?", is answered right next to the number.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"A number I can adjust."&lt;/strong&gt; Plans are drafts. The owner edits quantities and approves the plan, and the model never decides on its own.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Take my existing orders and make sense of them."&lt;/strong&gt; That's the paste-a-WhatsApp-message flow. Gemma reads the message, and the owner confirms the result instead of typing it in.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Correct something quickly" is a real gap.&lt;/strong&gt; Today an extracted order can be fixed before it's saved, and plan quantities can be changed. A &lt;em&gt;saved&lt;/em&gt; order can't be edited or cancelled yet, so last-minute changes aren't handled well. That's the first thing on my list.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"I don't want to spend 20 minutes entering orders."&lt;/strong&gt; This sets the bar for the order flow. I haven't timed it with Sidd's real orders yet, and that's the next test.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;🔗 &lt;strong&gt;Live demo (synthetic data):&lt;/strong&gt; &lt;a href="https://secondo-web.onrender.com" rel="noopener noreferrer"&gt;https://secondo-web.onrender.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Click &lt;strong&gt;Load synthetic sample sales&lt;/strong&gt;, then try Orders → Forecast → Kitchen plan. The free Render instance sleeps when idle, so the first load can take up to a minute.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Free hosting has 512 MB of RAM, which can't hold PyTorch or a 4B model. The hosted demo therefore uses the &lt;strong&gt;weekday-average forecaster&lt;/strong&gt; and the &lt;strong&gt;rule-based message parser&lt;/strong&gt; as fallbacks, and the UI labels which one is active. Gemma and TabPFN run in the local setup, which is the real product (see below).&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Orders&lt;/th&gt;
&lt;th&gt;Forecast&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5wjvwsuvkozm4gahcyyt.png" alt="Orders" width="800" height="611"&gt;&lt;/td&gt;
&lt;td&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwkqru3hpc1yl04q4j5zj.png" alt="Forecast" width="799" height="694"&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/jyotir07" rel="noopener noreferrer"&gt;
        jyotir07
      &lt;/a&gt; / &lt;a href="https://github.com/jyotir07/secondo" rel="noopener noreferrer"&gt;
        secondo
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      SECONDO turns sales history and messy customer messages into a daily baking plan that the owner reviews and approves.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;SECONDO&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Your business. Your data. Your intelligence.&lt;/strong&gt;
SECONDO is a local-first operations copilot for small food businesses: home bakers, tiffin
services and other one- or two-person kitchens. It turns sales history and messy customer
messages into a daily baking plan that the owner reviews and approves. Every number in the plan
comes with the reason behind it.&lt;/p&gt;
&lt;p&gt;It runs on the owner's own computer, with open-weight models and no API keys.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Live demo (synthetic data, hosted on Render):&lt;/strong&gt; &lt;a href="https://secondo-web.onrender.com" rel="nofollow noopener noreferrer"&gt;https://secondo-web.onrender.com&lt;/a&gt;. The free instance sleeps when idle, so the first load can take up to a minute. The hosted version uses the weekday average and the rule-based parser; TabPFN and Gemma run in the local setup.&lt;/p&gt;
&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/jyotir07/secondo/docs/screenshots/overview.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fjyotir07%2Fsecondo%2FHEAD%2Fdocs%2Fscreenshots%2Foverview.png" alt="Overview"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;About the data in this repository.&lt;/strong&gt; All sales data bundled here is &lt;strong&gt;synthetic&lt;/strong&gt;: it was
generated by &lt;code&gt;backend/scripts/generate_sample_data.py&lt;/code&gt; with a fixed seed. The app shows a
"Synthetic sample data" badge whenever it is loaded…&lt;/p&gt;
&lt;/blockquote&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/jyotir07/secondo" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;Running it locally takes about five minutes and needs no API keys:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Backend&lt;/span&gt;
&lt;span class="nb"&gt;cd &lt;/span&gt;backend
uv &lt;span class="nb"&gt;sync&lt;/span&gt; &lt;span class="nt"&gt;--extra&lt;/span&gt; tabpfn
uv run uvicorn app.main:create_app &lt;span class="nt"&gt;--factory&lt;/span&gt; &lt;span class="nt"&gt;--port&lt;/span&gt; 8000

&lt;span class="c"&gt;# Frontend (second terminal)&lt;/span&gt;
&lt;span class="nb"&gt;cd &lt;/span&gt;frontend
npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; npm run dev        &lt;span class="c"&gt;# http://localhost:3000&lt;/span&gt;

&lt;span class="c"&gt;# Optional: local Gemma for reading messages&lt;/span&gt;
ollama pull gemma3:4b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;SECONDO is a &lt;strong&gt;modular monolith&lt;/strong&gt;: one FastAPI service and one Next.js dashboard. It has no queues and no microservices, because one café doesn't need them. In the local setup everything lives in a single SQLite file on the owner's laptop.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Browser ─► Next.js (:3000) ─/api/*─► FastAPI (:8000)
                                        │
   Ingestion ──── Extraction ──── Demand ──── Planning
   CSV + dedupe   Gemma/Ollama    TabPFN      explanations,
                  ↳ rules fallback ↳ baselines approve/reject
                                        │
                                  SQLite file
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The open-source AI it uses:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemma 3 4B via Ollama&lt;/strong&gt; reads customer messages. It runs locally and has to return JSON that matches a schema. The output is checked with Pydantic, and if the JSON is malformed the call is retried. Gemma never gets the final say:

&lt;ul&gt;
&lt;li&gt;Products are matched against the real menu by &lt;strong&gt;deterministic code&lt;/strong&gt;, so the model can't book something the café doesn't sell.&lt;/li&gt;
&lt;li&gt;Dates are worked out in code from the phrase Gemma quotes ("Saturday"), because small models are unreliable at weekday arithmetic.&lt;/li&gt;
&lt;li&gt;If Ollama isn't running, a rule-based parser takes over and the UI says so.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TabPFN v2&lt;/strong&gt;, an open-weight tabular foundation model, does the forecasting. A café's history is a &lt;em&gt;tiny&lt;/em&gt; table: a few months and a handful of products. That's too little data for most ML to do well, and it's exactly the case TabPFN is built for. It needs no training pipeline for each business. It gets one row per (day, product) with lag and weekday features, and &lt;strong&gt;every feature is computed only from earlier days&lt;/strong&gt;. Tests check that no row can see its own sales.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Honest evaluation.&lt;/strong&gt; The evaluation is a rolling-origin backtest over the last 14 open days, using only the bundled &lt;strong&gt;synthetic&lt;/strong&gt; data:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;MAE (units/product/day)&lt;/th&gt;
&lt;th&gt;WAPE&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Historical average (baseline)&lt;/td&gt;
&lt;td&gt;5.95&lt;/td&gt;
&lt;td&gt;29.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Weekday average&lt;/td&gt;
&lt;td&gt;3.63&lt;/td&gt;
&lt;td&gt;18.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;TabPFN v2&lt;/strong&gt; (CPU)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3.42&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;17.0%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;TabPFN clearly beats the naive baseline and only slightly beats a weekday average. The synthetic data was &lt;em&gt;generated&lt;/em&gt; from weekday patterns, so these results show the pipeline works and doesn't leak future data. They say nothing yet about how a real café would do. On a laptop CPU (i5-12500H), a cold backtest takes about 30 seconds and a next-day forecast about 10.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fallbacks are visible.&lt;/strong&gt; If Ollama is down, extraction drops to rules. If TabPFN is missing or crashes, forecasting drops to the weekday average. The API and the UI always say which path ran and why. There are 54 backend tests, covering the Gemma path against a mocked Ollama (malformed JSON, schema violations, timeouts, a missing model) and leakage checks on the backtest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agent tracing with Sentry.&lt;/strong&gt; Each AI step is wrapped in a custom span: &lt;code&gt;secondo.extract&lt;/code&gt; contains a &lt;code&gt;gen_ai.request&lt;/code&gt; span for the Gemma call, recording the model, token counts and attempt number. Plan generation, the forecast backtest and each TabPFN run (&lt;code&gt;secondo.model.tabpfn&lt;/code&gt;) get their own spans too. Gemma failures while Ollama is up, and TabPFN crashes, become Sentry issues instead of silent fallbacks. Sentry never receives the message text or customer names, and &lt;code&gt;send_default_pii&lt;/code&gt; is off. The traces showed where the time goes: one real trace had plan generation at &lt;strong&gt;22.3 s&lt;/strong&gt;, almost all of it three ~7 s TabPFN runs. Everything else in that plan generation took about a second, so TabPFN is clearly the thing to optimise.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2y016gvb86325wu3dpfq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2y016gvb86325wu3dpfq.png" alt="Plan-generation spans in Sentry" width="800" height="415"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Every plan generation is recorded as a &lt;code&gt;secondo.plan.generate&lt;/code&gt; span in Sentry.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Voice briefing with ElevenLabs.&lt;/strong&gt; An approved plan gets a &lt;strong&gt;Play briefing&lt;/strong&gt; button, so the owner can hear the plan read aloud. The backend calls ElevenLabs, and the API key never reaches the browser. The text it sends contains only quantities, dietary labels and warnings, never customer names or notes, and the same text is shown on screen as a transcript. Audio is cached per plan. In testing, a 244-character briefing became about 23 seconds of MP3 in 8.7 s, and the cached replay took 0.05 s.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F16o228oidhqtwwurrjt1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F16o228oidhqtwwurrjt1.png" alt="Voice briefing" width="800" height="944"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MongoDB Atlas for the hosted demo.&lt;/strong&gt; Free Render disks are wiped on restart, so the hosted demo can't keep a SQLite file. &lt;code&gt;repository_mongo.py&lt;/code&gt; implements the same &lt;code&gt;Repository&lt;/code&gt; interface as the SQLite version and is used whenever &lt;code&gt;MONGODB_URI&lt;/code&gt; is set, so no other code changes. Duplicate detection is enforced in the database itself by a unique index on &lt;code&gt;(business_id, dedupe_key)&lt;/code&gt;. The full test suite also runs against a real Atlas cluster (56/56 passing), and an approved plan survived a backend restart.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stack:&lt;/strong&gt; Python 3.11, FastAPI, Pydantic, pandas, SQLite (uv) · Next.js 16, TypeScript, Tailwind 4, Recharts · Ollama + Gemma 3 4B · TabPFN v2 (CPU PyTorch). The hosted demo adds Render, MongoDB Atlas, Sentry and ElevenLabs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;For a one-counter café, open models are what make the product possible at all:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Customer data stays on the café's laptop.&lt;/strong&gt; Names, orders and dietary needs ("eggless", "nut allergy") are personal data about the owner's regulars. Gemma runs through a local Ollama server, so a customer's WhatsApp message is never sent to a cloud LLM. With a closed API, every order would leave the building.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It runs offline.&lt;/strong&gt; After the one-time model downloads, everything runs on a laptop CPU. Making the evening plan doesn't need an internet connection or a working API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It costs nothing to run.&lt;/strong&gt; There are no API keys, no subscription and no per-message billing. A tiny business can use it indefinitely for free, which is the only price that works at this scale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The right model for the data.&lt;/strong&gt; A closed general-purpose LLM is the wrong tool for forecasting four months of sales for six products. An open tabular foundation model like TabPFN fits that data, and I could inspect it, benchmark it honestly against baselines and run it locally.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Swappable and controllable.&lt;/strong&gt; The extraction model is a config value (&lt;code&gt;OLLAMA_MODEL&lt;/code&gt;), and so is the forecaster (&lt;code&gt;FORECAST_PROVIDER&lt;/code&gt;). If a better small open model comes out next month, switching is a one-line change. No vendor can deprecate it out from under the café.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ownership.&lt;/strong&gt; All data lives in one SQLite file, and there's an &lt;strong&gt;Export all my data&lt;/strong&gt; button. If SECONDO disappeared tomorrow, the café's records wouldn't go with it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  My Agent Session
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;p&gt;Best Use of Gemma · Best Use of TabPFN (Prior Labs) · Best Use of Render · Best Use of ElevenLabs · Best Use of MongoDB Atlas · Best Use of Sentry Agent Tracing&lt;/p&gt;

</description>
      <category>hacktoberfest</category>
      <category>opensource</category>
      <category>ai</category>
      <category>python</category>
    </item>
    <item>
      <title>Stop Scattering LLM Code Across Your Codebase</title>
      <dc:creator>Jyotiraditya Singh</dc:creator>
      <pubDate>Thu, 02 Jul 2026 15:19:10 +0000</pubDate>
      <link>https://dev.to/jyotir07/stop-scattering-llm-code-across-your-codebase-5027</link>
      <guid>https://dev.to/jyotir07/stop-scattering-llm-code-across-your-codebase-5027</guid>
      <description>&lt;p&gt;Every AI project starts the same way.&lt;/p&gt;

&lt;p&gt;You pick a model, install its SDK, and write something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-4.1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Explain recursion.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Everything works.&lt;/p&gt;

&lt;p&gt;Then a week later someone asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Can we try Claude?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So you install another SDK.&lt;/p&gt;

&lt;p&gt;Now your project has OpenAI code in one file, Anthropic code in another, Gemini somewhere else, and every agent knows exactly which provider it's talking to.&lt;/p&gt;

&lt;p&gt;A month later, switching models means touching dozens of files.&lt;/p&gt;

&lt;p&gt;I've made this mistake more than once.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Problem Isn't Which Model You Use
&lt;/h2&gt;

&lt;p&gt;It's that your application is tightly coupled to a provider.&lt;/p&gt;

&lt;p&gt;Instead of your business logic saying:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;it says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your application now depends on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenAI's SDK&lt;/li&gt;
&lt;li&gt;OpenAI's response format&lt;/li&gt;
&lt;li&gt;OpenAI's streaming API&lt;/li&gt;
&lt;li&gt;OpenAI's error handling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Switch providers and all of that changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Clean AI Architecture Looks Like
&lt;/h2&gt;

&lt;p&gt;Instead of every component talking directly to an SDK, give your application a single interface.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agents
 ↓
Services
 ↓
LLM Interface
 ↓
Provider
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your agents don't know whether the response came from OpenAI, Claude, Gemini, or something else.&lt;/p&gt;

&lt;p&gt;They just ask for text.&lt;/p&gt;

&lt;p&gt;That's the only thing they should care about.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Simple Example
&lt;/h2&gt;

&lt;p&gt;Instead of this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;your code becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;provider&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;openai&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-4.1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;modality&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Explain recursion.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rest of your application stays exactly the same.&lt;/p&gt;

&lt;p&gt;Changing providers doesn't require rewriting your business logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Matters
&lt;/h2&gt;

&lt;p&gt;As AI projects grow, you'll probably want to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;compare providers&lt;/li&gt;
&lt;li&gt;benchmark models&lt;/li&gt;
&lt;li&gt;add fallbacks&lt;/li&gt;
&lt;li&gt;route requests dynamically&lt;/li&gt;
&lt;li&gt;experiment with new releases&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If provider-specific code is scattered throughout your project, each of those becomes a refactoring exercise.&lt;/p&gt;

&lt;p&gt;If it's isolated behind one interface, they're configuration changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Pattern Scales
&lt;/h2&gt;

&lt;p&gt;This approach also makes it much easier to add things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;retries&lt;/li&gt;
&lt;li&gt;timeout handling&lt;/li&gt;
&lt;li&gt;logging&lt;/li&gt;
&lt;li&gt;usage tracking&lt;/li&gt;
&lt;li&gt;streaming&lt;/li&gt;
&lt;li&gt;structured outputs&lt;/li&gt;
&lt;li&gt;provider fallbacks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;without changing every agent or API route.&lt;/p&gt;

&lt;p&gt;Instead, those concerns live in one place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I Built Loom
&lt;/h2&gt;

&lt;p&gt;After repeating this pattern across multiple projects, I wanted a single abstraction that kept provider-specific code isolated.&lt;/p&gt;

&lt;p&gt;That's why I built &lt;strong&gt;Loom&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The idea is simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Your application shouldn't know which LLM provider it's using.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Only one file should interact with provider SDKs.&lt;/p&gt;

&lt;p&gt;Everything else should depend on a clean interface.&lt;/p&gt;

&lt;p&gt;Loom currently supports multiple providers behind the same API so switching providers doesn't require rewriting your application logic.&lt;/p&gt;

&lt;p&gt;If you're interested, you can explore the project here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt; Website: &lt;a href="https://loom-weaves.vercel.app" rel="noopener noreferrer"&gt;https://loom-weaves.vercel.app&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;GitHub: &lt;a href="https://github.com/jyotir07/Loom" rel="noopener noreferrer"&gt;https://github.com/jyotir07/Loom&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'm also planning to write more about the architecture behind provider abstraction, routing strategies, streaming, and production AI systems.&lt;/p&gt;

&lt;p&gt;I'd love to hear how you're structuring LLM integrations in your own projects. Are you calling provider SDKs directly, or have you built an abstraction layer too?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>llm</category>
      <category>python</category>
    </item>
  </channel>
</rss>
