<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Shaw Sha</title>
    <description>The latest articles on DEV Community by Shaw Sha (@shadie_ai).</description>
    <link>https://dev.to/shadie_ai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3958538%2Fb37de443-b097-419e-8e05-2f83abbbbcec.png</url>
      <title>DEV Community: Shaw Sha</title>
      <link>https://dev.to/shadie_ai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shadie_ai"/>
    <language>en</language>
    <item>
      <title>Building an AI Side Project That Actually Ships — Lessons from Shipping 3 MVPs</title>
      <dc:creator>Shaw Sha</dc:creator>
      <pubDate>Fri, 02 Oct 2026 00:55:56 +0000</pubDate>
      <link>https://dev.to/shadie_ai/building-an-ai-side-project-that-actually-ships-lessons-from-shipping-3-mvps-4bkh</link>
      <guid>https://dev.to/shadie_ai/building-an-ai-side-project-that-actually-ships-lessons-from-shipping-3-mvps-4bkh</guid>
      <description>&lt;p&gt;I’ve lost count of how many times I’ve started an AI project with a burst of excitement, only to abandon it two weekends later. The pattern was always the same: a brilliant idea, a deep dive into fine-tuning a model, a week of wrestling with CUDA drivers, and then—silence. The repo goes cold, the README stays half-written, and I’m back to scrolling Twitter.&lt;/p&gt;

&lt;p&gt;But something shifted for me over the last two months. I shipped three AI MVPs, got real users on two of them, and actually learned what it takes to cross the finish line. It wasn’t about smarter engineering or better prompts. It was about brutally cutting scope and embracing the unglamorous parts of building.&lt;/p&gt;

&lt;p&gt;Here’s what that journey looked like, warts and all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Wake-Up Call: My First "AI" Project Was a Disaster
&lt;/h2&gt;

&lt;p&gt;Let me rewind to January. I had this idea for an AI-powered meeting summarizer that would integrate with Slack. The vision was grand: it would identify action items, assign owners, and even draft follow-up emails. I spent three weeks building a custom pipeline using LangChain, vector embeddings, and a local fine-tuned model because I thought that was the "real" way to do it.&lt;/p&gt;

&lt;p&gt;The result? It worked—sort of. The summaries were 60% accurate, the latency was terrible (15 seconds per call), and I spent more time debugging the infrastructure than actually using the tool. I showed it to two friends, both nodded politely, and then never opened it again.&lt;/p&gt;

&lt;p&gt;I had fallen into the classic trap: I was building an &lt;em&gt;AI system&lt;/em&gt;, not a &lt;em&gt;product&lt;/em&gt;. The user didn't care about my custom tokenizer. They cared about not having to read a 45-minute meeting transcript.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Pivot: Shipping Small and Dumb
&lt;/h2&gt;

&lt;p&gt;After that failure, I set a new rule for myself: &lt;strong&gt;The MVP must be buildable in a single weekend.&lt;/strong&gt; No exceptions. If I can't demo it by Sunday night, I'm not building it.&lt;/p&gt;

&lt;p&gt;This forced me to think differently. Instead of asking "What can I build with AI?", I started asking "What's the smallest annoying problem I can solve with one API call?"&lt;/p&gt;

&lt;h3&gt;
  
  
  MVP #1: The Resume Screener (Weekend 1)
&lt;/h3&gt;

&lt;p&gt;A friend of mine was drowning in resumes for a junior developer role. She had 200 applications and no time. My idea was dead simple: a web app where she pastes the job description, uploads PDFs, and gets a ranked list based on keyword and semantic similarity.&lt;/p&gt;

&lt;p&gt;Here's the core logic—it's embarrassingly simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pandas&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;rank_resumes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;job_description&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;resumes&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;scores&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;res&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;resumes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;embeddings&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text-embedding-3-small&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;job_description&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;res&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;emb1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;emb2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt;
        &lt;span class="n"&gt;similarity&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;emb1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;emb2&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;similarity&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;scores&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it. No vector database, no LangChain, no fine-tuning. Just cosine similarity on embeddings. It took me four hours to build the backend and three hours to throw together a basic HTML form. Total time: 7 hours.&lt;/p&gt;

&lt;p&gt;She used it to shortlist 20 candidates in an afternoon. It wasn't perfect, but it saved her a day of work. I didn't monetize it, but I learned something crucial: &lt;strong&gt;the "dumb" solution is often the best one&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  MVP #2: The Changelog Generator (Weekend 3)
&lt;/h3&gt;

&lt;p&gt;After the first success, I wanted to push a bit further. I was working on a small SaaS product and hated writing changelog entries. So I built a tool that takes a GitHub diff and generates a human-readable summary.&lt;/p&gt;

&lt;p&gt;The trick was to stop thinking about "understanding the code" and instead just feed the diff text into a prompt. The key was in the system message:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a technical writer. Given the following git diff, write a concise changelog entry for a non-technical audience. Focus on user-visible changes. Ignore formatting and refactoring. Use bullet points.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I used the &lt;code&gt;gpt-4o-mini&lt;/code&gt; model for this because it's cheap and fast. The whole thing runs in a single Python script that I wrapped in a Flask app. The hardest part wasn't the AI—it was parsing the &lt;code&gt;git diff&lt;/code&gt; output into a clean string.&lt;/p&gt;

&lt;p&gt;This one took me about 6 hours across two evenings. A few friends started using it, and one even said it saved them from writing boring release notes for their client work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Reality Check: What Actually Mattered
&lt;/h2&gt;

&lt;p&gt;Here's the thing that surprised me. In all three projects, the AI part was maybe 20% of the work. The other 80% was:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Handling edge cases&lt;/strong&gt;: What if the PDF is scanned? What if the diff is 10,000 lines?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Making it fast&lt;/strong&gt;: No one cares about a smart AI if it takes 30 seconds to respond.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Simple UI&lt;/strong&gt;: I used plain HTML and a tiny bit of CSS. Nobody complimented my frontend skills, but they didn't complain either.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The biggest lesson? &lt;strong&gt;Prompt engineering is more important than model choice.&lt;/strong&gt; I spent hours comparing &lt;code&gt;gpt-4&lt;/code&gt; vs &lt;code&gt;claude-3&lt;/code&gt; vs &lt;code&gt;llama-3&lt;/code&gt; on benchmarks. In the end, the winning factor was always how I structured the prompt, not which model I used. For 95% of use cases, a cheap model with a great prompt beats an expensive model with a lazy prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Infrastructure Trap I Almost Fell Into
&lt;/h2&gt;

&lt;p&gt;On MVP #3—a content repurposing tool that turns blog posts into Twitter threads—I hit a wall. The API costs were creeping up, and I thought, "I should just self-host a small model to cut costs."&lt;/p&gt;

&lt;p&gt;I spent an entire weekend trying to get a quantized model running on a rented GPU. I installed Docker, fiddled with &lt;code&gt;vLLM&lt;/code&gt;, battled with CUDA version mismatches. I even got it working—once. Then the GPU instance rebooted and I lost all my environment setup. I rage-quit and went back to the API.&lt;/p&gt;

&lt;p&gt;That weekend taught me a valuable lesson about opportunity cost. I could have spent that time marketing the product or improving the prompt. Instead, I was fighting infrastructure battles that had nothing to do with my users' problems.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Numbers That Keep Me Honest
&lt;/h2&gt;

&lt;p&gt;Here are the raw numbers from my two-month experiment:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Projects started:&lt;/strong&gt; 5&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Projects shipped:&lt;/strong&gt; 3&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Projects with external users:&lt;/strong&gt; 2&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Total API spend:&lt;/strong&gt; $47.32&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Time invested per project:&lt;/strong&gt; 6-10 hours&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lines of code per project:&lt;/strong&gt; 200-400&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The most successful project wasn't the one with the most complex AI. It was the one that solved a specific, painful problem for a specific person. The resume screener got 3 users in a week. The changelog generator got 15 users because I posted it on a niche Slack group.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd Tell My Past Self
&lt;/h2&gt;

&lt;p&gt;If I could go back to January, I'd tell myself:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Start with the API, not the model.&lt;/strong&gt; Don't even think about self-hosting until you have paying users. And even then, question it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Time-box the AI part.&lt;/strong&gt; If you can't get the core AI feature working in 2 hours, you're overcomplicating it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The prompt is the product.&lt;/strong&gt; Spend 30 minutes crafting your system prompt before you even write a line of code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Show it to one person early.&lt;/strong&gt; The resume screener was built for my friend, not for "the market." That clarity changed everything.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Practical Side: What I Use Now
&lt;/h2&gt;

&lt;p&gt;Since that GPU nightmare, I've been a firm believer in using managed, pay-as-you-go APIs. I don't want to think about provisioning, scaling, or keeping models warm. I want to spend my weekends building features, not debugging &lt;code&gt;nvidia-smi&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;These days, I default to using a single aggregation endpoint that gives me access to multiple models without the overhead of managing separate API keys and billing. It's been a lifesaver for rapid prototyping. When I want to test if a cheaper model works for a task, I can just swap it in with a config change instead of rewriting code.&lt;/p&gt;

&lt;p&gt;I've been using &lt;strong&gt;tai.shadie-oneapi.com&lt;/strong&gt; for this. It's not the flashiest tool, but it does exactly what I need: one API key, multiple models, and no surprise bills. For someone who ships MVPs on a weekend, that's worth more than any fancy feature.&lt;/p&gt;

&lt;p&gt;The bottom line is this: AI projects don't fail because the AI isn't good enough. They fail because we get lost in the technology and forget we're building a tool for humans. Ship something small, get it in front of someone, and iterate. The rest is just noise.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>beginners</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How I Cut My LLM API Costs by 70% Without Touching My Code</title>
      <dc:creator>Shaw Sha</dc:creator>
      <pubDate>Thu, 01 Oct 2026 00:55:25 +0000</pubDate>
      <link>https://dev.to/shadie_ai/how-i-cut-my-llm-api-costs-by-70-without-touching-my-code-2811</link>
      <guid>https://dev.to/shadie_ai/how-i-cut-my-llm-api-costs-by-70-without-touching-my-code-2811</guid>
      <description>&lt;p&gt;I was spending $200 a month on AI APIs. Now I'm down to $60, and my application works exactly the same. Same latency, same response quality, same user experience. The only thing that changed was how I route my requests.&lt;/p&gt;

&lt;p&gt;Let me walk you through what I did, because it took me about three hours to set up, and it's been saving me money every single month since.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Starting Point
&lt;/h2&gt;

&lt;p&gt;My application uses GPT-4o for chat completions, Claude for long-form content generation, and a custom fine-tuned model for classification tasks. When I first built this, I picked the models based on quality and just wired them up directly. Every call went straight to whatever provider I chose at build time.&lt;/p&gt;

&lt;p&gt;The problem is that direct integration locks you into one provider's pricing. And the pricing differences between providers for similar quality outputs are honestly wild if you actually sit down and compare them.&lt;/p&gt;

&lt;p&gt;I was running about 500,000 API calls per month. Not huge, but not trivial either. At roughly $0.40 per 1M tokens average across my workloads, with most of that being output tokens which are always more expensive, I was looking at a $200 average monthly bill. When my usage spiked to 700,000 calls, that hit $280.&lt;/p&gt;

&lt;p&gt;The pain point wasn't just the absolute number. It was watching that number grow while knowing I couldn't do much about it without rewriting large chunks of my integration code.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Realization
&lt;/h2&gt;

&lt;p&gt;Here's what I discovered: model prices change constantly, and the "best" model for a task shifts over time. But my code had hardcoded endpoints and model names everywhere. Switching from one provider to another meant touching dozens of files.&lt;/p&gt;

&lt;p&gt;I remember sitting there at 11 PM, looking at a Git diff that was essentially a global find-and-replace of API endpoints, thinking there has to be a better way.&lt;/p&gt;

&lt;p&gt;There is. It's called an AI gateway.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Worked
&lt;/h2&gt;

&lt;p&gt;Instead of calling OpenAI, Anthropic, or any provider directly, my code now calls one endpoint. That endpoint handles the routing, load balancing, and fallbacks. My application code doesn't know or care which provider is actually serving the request.&lt;/p&gt;

&lt;p&gt;The setup looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;

&lt;span class="c1"&gt;# Before: direct call to OpenAI
# response = openai.chat.completions.create(model="gpt-4o", messages=messages)
&lt;/span&gt;
&lt;span class="c1"&gt;# After: gateway call, same interface
&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://my-gateway.example.com/v1/chat/completions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer my-gateway-key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-4o&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# still specify what you want
&lt;/span&gt;        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;messages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The gateway takes my request, checks which provider currently offers the best price for that model category, and routes accordingly. It also handles rate limits, retries, and load balancing.&lt;/p&gt;

&lt;p&gt;The key insight is that the interface matches the provider format. I didn't have to change any of my existing logic. The gateway presents the exact same API convention that the big providers use, so my code, my prompts, my logging all stayed intact.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Numbers That Matter
&lt;/h2&gt;

&lt;p&gt;Here's the breakdown of where the savings actually came from:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Dynamic Provider Selection
&lt;/h3&gt;

&lt;p&gt;The biggest win was letting the gateway pick which provider serves the request. For casual chat completion, the price difference between providers for comparable models was around 30-40%. By routing to the cheapest provider with adequate quality at any given moment, I cut that cost instantly.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Model Tiering
&lt;/h3&gt;

&lt;p&gt;I had been using the same high-end model for everything. Turns out, a lot of my requests didn't need the smartest model available. Simple classification, basic extraction, or quick summarization tasks work perfectly fine on smaller, faster models that cost about 70% less per token.&lt;/p&gt;

&lt;p&gt;The gateway let me set rules: if the request is classified as "simple," route it to a cheaper model. Only complex reasoning tasks go to the top-tier models. My quality metrics didn't budge, but the cost did.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Batch Processing and Caching
&lt;/h3&gt;

&lt;p&gt;This is where I got a bit nerdy. The gateway caches identical requests, which I didn't even realize I was making so many of. When the same request comes in twice, it serves the cached response. My cache hit rate is around 18% now, which sounds low, but those are 18% of calls I'm paying exactly zero for.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Pay-As-You-Go Over Subscriptions
&lt;/h3&gt;

&lt;p&gt;This was the biggest mindset shift. I was holding standing API credits with multiple providers, paying monthly minimums, and letting that money just sit there. When I moved to a pay-as-you-go routing model, I stopped paying for unused capacity.&lt;/p&gt;

&lt;p&gt;The cost comparison after a full month:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cost Category&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Provider fees&lt;/td&gt;
&lt;td&gt;$185&lt;/td&gt;
&lt;td&gt;$52&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gateway usage&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;$8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$185&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$60&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's a 68% reduction, and it's been consistent for three months now.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Implementation Details
&lt;/h2&gt;

&lt;p&gt;If you're thinking about doing this, here's the honest breakdown of what it takes:&lt;/p&gt;

&lt;h3&gt;
  
  
  What You Need
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;A gateway service (open-source options exist)&lt;/li&gt;
&lt;li&gt;Read access to provider pricing (check their published rate cards)&lt;/li&gt;
&lt;li&gt;Your existing API keys&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Core Config
&lt;/h3&gt;

&lt;p&gt;Most gateways are configured with YAML or JSON rules. My config has two critical sections:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Provider pool&lt;/strong&gt;: which providers are available and their base URLs&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routing rules&lt;/strong&gt;: when to use which provider/model
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;providers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openai&lt;/span&gt;
    &lt;span class="na"&gt;base_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://api.openai.com/v1&lt;/span&gt;
    &lt;span class="na"&gt;api_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;env.OPENAI_KEY&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;anthropic&lt;/span&gt;
    &lt;span class="na"&gt;base_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://api.anthropic.com/v1&lt;/span&gt;
    &lt;span class="na"&gt;api_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;env.ANTHROPIC_KEY&lt;/span&gt;

&lt;span class="na"&gt;routes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;classification"&lt;/span&gt;
    &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openai&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gpt-3.5-turbo&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;long-form"&lt;/span&gt;
    &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;anthropic&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;claude-3-5-sonnet&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;general"&lt;/span&gt;
    &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;cheapest_available&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;auto&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The Migration
&lt;/h3&gt;

&lt;p&gt;Here's the part that surprised me: migration took me about three hours for an application with 60+ integration points.&lt;/p&gt;

&lt;p&gt;The steps were:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Spin up the gateway locally&lt;/li&gt;
&lt;li&gt;Point a test environment at it&lt;/li&gt;
&lt;li&gt;Run my existing test suite (which mocked API calls) against the gateway&lt;/li&gt;
&lt;li&gt;Swap the base URL in production config&lt;/li&gt;
&lt;li&gt;Watch logs for a week&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;No code changes. No prompt changes. The switch was literally a configuration change.&lt;/p&gt;

&lt;h2&gt;
  
  
  Edge Cases I Hit
&lt;/h2&gt;

&lt;p&gt;I want to be upfront about the complications, because they're real:&lt;/p&gt;

&lt;h3&gt;
  
  
  Token Counting Difference
&lt;/h3&gt;

&lt;p&gt;Different providers count tokens slightly differently for billing. What's a "token" to Anthropic isn't always identical to OpenAI's definition. Your billing metering might show different usage numbers than what the gateway reports. I had to build a normalization layer for my internal cost tracking.&lt;/p&gt;

&lt;h3&gt;
  
  
  Rate Limit Headaches
&lt;/h3&gt;

&lt;p&gt;Free or cheaper tiers often have rate limits that premium tiers don't. When I routed more traffic to budget providers, I hit rate limits on a few high-traffic days. The gateway's retry logic handled it, but I had to tune the retry backoff to avoid spamming.&lt;/p&gt;

&lt;h3&gt;
  
  
  Quality Consistency
&lt;/h3&gt;

&lt;p&gt;No two models are identical, even if they score similarly on benchmarks. My classification task (which uses GPT-3.5-turbo on the budget side) gave slightly different confidence scores than the premium model. I had to adjust my confidence threshold by about 5% to maintain the same precision.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm Doing Now
&lt;/h2&gt;

&lt;p&gt;Since setting this up, I've built a habit of checking my gateway analytics weekly. Some providers have dropped prices by 20-30% in the months since. The gateway picks that up automatically.&lt;/p&gt;

&lt;p&gt;I'm also experimenting with speculative routing, where the gateway sends the first few tokens to a cheap model and evaluates whether to upgrade mid-stream for complex responses. Early results suggest another 10-15% savings on long generation tasks, though the infrastructure is a bit more involved.&lt;/p&gt;

&lt;p&gt;One more thing: I moved to a pay-as-you-go gateway service instead of running my own infrastructure for this. I don't want to maintain a load balancer, handle failover, and monitor uptime on top of my actual application. The per-request fee is small enough that it's worth it.&lt;/p&gt;

&lt;p&gt;By the way, the gateway I moved to is &lt;strong&gt;&lt;a href="https://tai.shadie-oneapi.com" rel="noopener noreferrer"&gt;tai.shadie-oneapi.com&lt;/a&gt;&lt;/strong&gt; — it's a pay-as-you-go aggregator that doesn't lock you into a monthly subscription. You only pay for the tokens you actually use across whatever providers it routes to. I found it while comparing aggregate API pricing, and the no-vendor-lock-in model was what sold me.&lt;/p&gt;

&lt;p&gt;The whole experience has shifted how I think about API costs. I used to treat them as a fixed overhead. Now they're an optimizable variable, and it's been a significant chunk of money back every month. If you're spending over $100/month on API calls, I'd genuinely recommend spending an afternoon setting up routing. The math works out.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>I Spent 10x Longer Debugging AI Code Than Writing It — Here's What Changed</title>
      <dc:creator>Shaw Sha</dc:creator>
      <pubDate>Wed, 30 Sep 2026 00:55:59 +0000</pubDate>
      <link>https://dev.to/shadie_ai/i-spent-10x-longer-debugging-ai-code-than-writing-it-heres-what-changed-10k6</link>
      <guid>https://dev.to/shadie_ai/i-spent-10x-longer-debugging-ai-code-than-writing-it-heres-what-changed-10k6</guid>
      <description>&lt;p&gt;Everyone talks about how AI speeds up coding. I've read a hundred posts about generating entire functions in seconds, scaffolding CRUD apps before your coffee gets cold, shipping features at 3x velocity. Nobody talks about the debugging. Nobody warns you that the real time sink isn't the generation — it's the 45 minutes you spend figuring out why the AI's "perfect" solution silently fails on edge case #17.&lt;/p&gt;

&lt;p&gt;I learned this the hard way. Over the past three months, I tracked my time on a side project — a data pipeline that processes real-time stock feeds. I logged every hour: writing code, debugging code, and &lt;em&gt;debugging AI-written code&lt;/em&gt;. The numbers were brutal. I spent roughly 8 hours writing my own code, and 62 hours debugging the AI's output. That's not a typo. 62 hours. Almost 10x the time I spent typing my own logic. And that's not even counting the hours I spent rewriting the AI's "optimizations" that were actually just slower, more convoluted versions of what I'd already written.&lt;/p&gt;

&lt;p&gt;Here's what I learned, the hard way, so you don't have to.&lt;/p&gt;

&lt;h2&gt;
  
  
  The False Confidence Trap
&lt;/h2&gt;

&lt;p&gt;The problem isn't that AI generates bad code. It's that AI generates &lt;em&gt;confident&lt;/em&gt; code. It spits out a function with perfect syntax, reasonable variable names, and comments that explain exactly what it's &lt;em&gt;supposed&lt;/em&gt; to do. And you read it, and it looks right. So you paste it in, run your tests, and they pass. You ship it.&lt;/p&gt;

&lt;p&gt;Then three days later, a user reports that the timestamp on their invoice is off by exactly 4 hours. You dig in. The AI's code handles UTC conversion perfectly — except for one path where it uses &lt;code&gt;.toISOString()&lt;/code&gt; on a date that's already a string, which coerces to &lt;code&gt;NaN&lt;/code&gt;, which gets caught by a fallback that defaults to &lt;code&gt;new Date()&lt;/code&gt;, which uses the &lt;em&gt;local&lt;/em&gt; timezone instead of UTC. And the AI wrote a comment saying "// normalize to UTC" right above it.&lt;/p&gt;

&lt;p&gt;I've seen this exact pattern play out a dozen times. The AI doesn't know it's wrong. It's predicting the next token, not reasoning about your data. So it writes code that &lt;em&gt;looks&lt;/em&gt; correct, passes the happy path, and breaks in the exact place where your domain logic gets weird.&lt;/p&gt;

&lt;h2&gt;
  
  
  The "It Works in My Tests" Fallacy
&lt;/h2&gt;

&lt;p&gt;Let me show you a concrete example. I asked an AI to write a function that deduplicates a list of user objects by email, case-insensitively.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;deduplicate_users&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;users&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;seen&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;users&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Looks fine, right? I thought so too. My tests passed. Then I ran it on real data and found that two users with the same email but different casing — &lt;code&gt;User@Example.com&lt;/code&gt; and &lt;code&gt;user@example.com&lt;/code&gt; — were being deduplicated, when they were actually &lt;em&gt;different&lt;/em&gt; accounts with different permissions. The AI's code was technically correct for the letter of my prompt, but completely wrong for the spirit of my domain.&lt;/p&gt;

&lt;p&gt;The fix took me two hours. I had to write a custom key function that only normalizes for comparison, but preserves the original for output. I had to handle the edge case where two users have the same normalized email but different IDs. I had to write a regression test. And I had to explain to my PM why a "5-minute AI task" took half a day.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Changed
&lt;/h2&gt;

&lt;p&gt;After that experience, I didn't stop using AI. That would be throwing the baby out with the bathwater. But I completely changed &lt;em&gt;how&lt;/em&gt; I use it. Here are the three rules that cut my debugging time from 10x down to maybe 1.5x.&lt;/p&gt;

&lt;h3&gt;
  
  
  Rule 1: AI for Structure, Not for Logic
&lt;/h3&gt;

&lt;p&gt;I stopped asking AI to write the actual business logic. Instead, I ask it to write the scaffolding — the boilerplate, the data transformations, the glue code. When I need a function that loops through a list and applies a transformation, I let the AI write the loop. When I need to decide &lt;em&gt;what&lt;/em&gt; transformation to apply, that's on me.&lt;/p&gt;

&lt;p&gt;For example, instead of asking "write a function that validates user input," I ask "write a Pydantic model for a user registration form with these fields." The AI is great at that. It's terrible at understanding &lt;em&gt;why&lt;/em&gt; certain validation rules matter in my specific context.&lt;/p&gt;

&lt;h3&gt;
  
  
  Rule 2: The 15-Minute Rule
&lt;/h3&gt;

&lt;p&gt;I now have a hard rule: if I can't understand what the AI's code does within 15 minutes, I delete it and write it myself. Every single time I've broken this rule, I've regretted it. The AI's clever one-liner that uses three nested list comprehensions and a generator expression? It might be elegant, but if I can't trace through it quickly, I can't debug it quickly either. And the debugging is where the time goes.&lt;/p&gt;

&lt;p&gt;This rule has saved me more hours than any other single change. It forces me to treat AI output as a draft, not a final answer. I read it, understand it, and then often rewrite it in a way that's less clever but more maintainable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Rule 3: I Write the Tests First
&lt;/h3&gt;

&lt;p&gt;This one sounds obvious, but I never did it with AI code. I used to just ask for the implementation, run it, and trust the output. Now I write the test cases &lt;em&gt;before&lt;/em&gt; I even look at the AI's solution. I think about my edge cases — empty inputs, duplicate data, timezone boundaries, Unicode characters — and I write tests for all of them.&lt;/p&gt;

&lt;p&gt;Then I ask the AI to implement the function. When it inevitably fails on my edge case tests, I can see exactly where it breaks, and I can either fix it or write a new prompt that addresses the failure. This turns debugging from a hunt through unfamiliar code into a targeted exercise.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Consistency Problem
&lt;/h2&gt;

&lt;p&gt;The other thing that changed was my tooling. I realized that a huge source of my debugging time was actually &lt;em&gt;inconsistency&lt;/em&gt;. I'd get one answer from one model, then a slightly different answer from another model, and they'd conflict in subtle ways. One would handle timezone DST correctly, the other wouldn't. One would use &lt;code&gt;O(n)&lt;/code&gt; space, the other &lt;code&gt;O(1)&lt;/code&gt;. And I'd spend an hour reconciling their behavior.&lt;/p&gt;

&lt;p&gt;That's when I started being more deliberate about which API I used. I needed something stable, consistent, and — honestly — affordable, because I was burning tokens on all these debugging iterations. After trying a few options, I landed on a pay-as-you-go API gateway that gives me access to multiple models through a single, stable endpoint. It's been a game-changer because I can stick with one model configuration long enough to learn its quirks, rather than chasing a moving target.&lt;/p&gt;

&lt;p&gt;I'm not saying you need to use a specific tool — but I do think the consistency of your model access matters more than you'd expect. If you're constantly switching between models or hitting rate limits, you're adding a whole other layer of unpredictability to an already unpredictable process.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Metric
&lt;/h2&gt;

&lt;p&gt;Here's the thing I wish someone had told me: the metric that matters isn't "lines of code written per hour." It's "features shipped per week." And that metric only improves if you're not spending all your time debugging code you didn't write.&lt;/p&gt;

&lt;p&gt;AI has made me faster, but not in the way I expected. It's made me faster because I'm more disciplined about &lt;em&gt;what&lt;/em&gt; I delegate and &lt;em&gt;how&lt;/em&gt; I verify the results. I still write the core logic myself. I still write the tests first. I still spend time understanding every line that goes into production — whether I wrote it or not.&lt;/p&gt;

&lt;p&gt;The 10x debugging problem was real, but it wasn't inevitable. It was a symptom of treating AI as a colleague who could be trusted to get it right, rather than a tool that needs careful supervision. Once I shifted my mindset, the time I spent debugging dropped dramatically. I still use AI every day, but now I use it as a junior developer who needs explicit instructions and thorough review — not as a senior who knows what they're doing.&lt;/p&gt;

&lt;p&gt;And honestly? That's the right way to think about it. AI is a powerful tool, but it's still a tool. It doesn't understand your domain, your users, or your constraints. It's really good at generating plausible code, and really bad at knowing when that code is wrong. The sooner you internalize that, the sooner you'll stop spending 10x longer debugging than writing.&lt;/p&gt;

&lt;p&gt;If you're just starting to integrate AI into your workflow, or if you've been frustrated by inconsistent model behavior, I'd recommend finding a stable, reliable API endpoint that doesn't make you think about quotas or rate limits — something like shadie-oneapi.com, which lets you pay as you go without the overhead of managing multiple subscriptions. That consistency, combined with the discipline of writing tests first and reviewing everything, has made the difference between AI being a time-saver and a time-sink for me.&lt;/p&gt;

&lt;p&gt;The code you write is yours. The bugs you fix are yours too. Make sure you understand both — whether you wrote them or the AI did.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Why I Stopped Self-Hosting AI Models (And You Probably Should Too)</title>
      <dc:creator>Shaw Sha</dc:creator>
      <pubDate>Tue, 29 Sep 2026 00:55:52 +0000</pubDate>
      <link>https://dev.to/shadie_ai/why-i-stopped-self-hosting-ai-models-and-you-probably-should-too-5fig</link>
      <guid>https://dev.to/shadie_ai/why-i-stopped-self-hosting-ai-models-and-you-probably-should-too-5fig</guid>
      <description>&lt;p&gt;I’m not going to lie—I was that guy. The one who bought a used RTX 3090 off eBay, maxed out my home’s breaker panel, and spent three months convincing myself that self-hosting a 7B parameter LLM was the future of my side projects. I had a blog post drafted in my head: "How I Built a Private ChatGPT for Under $500." It was going to be epic.&lt;/p&gt;

&lt;p&gt;It wasn’t.&lt;/p&gt;

&lt;p&gt;After burning through $500 in hardware, countless weekends, and a small fortune in electricity bills, I finally pulled the plug and switched to a $1 API. Here’s the story of my descent into self-hosting madness, and why I think 99% of developers should just use an API.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Siren Call of Open Weights
&lt;/h2&gt;

&lt;p&gt;It started innocently enough. I read about Llama 2, then Mistral, then the explosion of fine-tuned models on Hugging Face. The open-source community was doing incredible things. I remember thinking, &lt;em&gt;"If these models are 'good enough' and 'free,' why would I pay OpenAI or Anthropic a monthly subscription?"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The appeal was threefold:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Privacy:&lt;/strong&gt; My data stays on my machine. No one can peek at my prompts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Control:&lt;/strong&gt; I can fine-tune the model to my exact use case.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost:&lt;/strong&gt; No per-token fees. Just electricity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I was sold. I dove headfirst into the rabbit hole.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Build: A Tale of Woe
&lt;/h2&gt;

&lt;p&gt;I’m a competent DevOps engineer. I’ve containerized microservices, orchestrated Kubernetes clusters, and debugged network latency issues that would make seasoned sysadmins weep. I figured hosting a model would be a walk in the park.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hardware Haul:&lt;/strong&gt; I bought a used RTX 3090 (24GB VRAM) for $450. My power supply wasn’t beefy enough, so that was another $120. Let’s call it $570 total.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Setup:&lt;/strong&gt; I chose Ollama for simplicity. It’s a fantastic tool, honestly. &lt;code&gt;ollama run llama2&lt;/code&gt; and you’re off to the races. The initial test was mind-blowing. The model responded faster than I expected, and the quality was... decent.&lt;/p&gt;

&lt;p&gt;But "decent" in a controlled demo is very different from "production-ready" in a real application.&lt;/p&gt;

&lt;p&gt;Here’s where the cracks started to show.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Memory Wall
&lt;/h3&gt;

&lt;p&gt;My 24GB VRAM was fine for a 7B model with 4-bit quantization. But I wanted to run a 13B model for better code generation. That immediately pushed me into the territory of offloading layers to system RAM. The inference speed dropped from "snappy" to "watching paint dry."&lt;/p&gt;

&lt;p&gt;I remember a specific instance: I was building a small agent that needed to summarize emails. A simple task. With the 13B model over CPU offload, a single 200-word email took &lt;strong&gt;45 seconds&lt;/strong&gt; to summarize. That’s not an agent; that’s a time machine to 1998.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Concurrency Problem
&lt;/h3&gt;

&lt;p&gt;The real killer wasn't speed—it was concurrency.&lt;/p&gt;

&lt;p&gt;My API server (FastAPI + Uvicorn) could handle dozens of requests simultaneously. But my GPU? Not so much. With a single GPU, you can process one batch at a time effectively. When two requests hit at once, the second one queues. When five hit, the queue backs up.&lt;/p&gt;

&lt;p&gt;I stress-tested it once. I sent 10 concurrent requests to a simple text-generation endpoint.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Self-hosted:&lt;/strong&gt; Average latency 12 seconds, p99 latency 38 seconds. Half the requests timed out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API (GPT-4o-mini):&lt;/strong&gt; Average latency 0.8 seconds, p99 1.5 seconds. Zero failures.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The difference wasn't incremental. It was a chasm.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Hidden Costs You Don't Think About
&lt;/h2&gt;

&lt;p&gt;People talk about the "cost of GPUs," but that’s only the tip of the iceberg. Let’s break down the numbers from my three-month experiment:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cost Category&lt;/th&gt;
&lt;th&gt;Monthly Expense&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Electricity (GPU at load)&lt;/td&gt;
&lt;td&gt;~$40 - $60&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud backup (for model weights)&lt;/td&gt;
&lt;td&gt;~$10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time spent debugging (avg 5 hrs/week)&lt;/td&gt;
&lt;td&gt;Priceless (but let's say $250 at freelance rates)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$300 - $320 / month&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And that’s not even counting the initial $570 hardware investment.&lt;/p&gt;

&lt;p&gt;Now, let’s look at what I actually use today. I subscribe to a small API plan that costs me &lt;strong&gt;$1&lt;/strong&gt; a month for my low-traffic hobby projects. That’s $300+ vs $1. The math isn’t just lopsided—it’s insulting.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Maintenance Nightmare
&lt;/h2&gt;

&lt;p&gt;The final straw for me was the update cycle.&lt;/p&gt;

&lt;p&gt;One Tuesday, I decided to update my model from Llama 2 to Llama 3. I thought it would be a simple swap. I was wrong.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Step 1:&lt;/strong&gt; Download the new weights. (20 minutes)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Step 2:&lt;/strong&gt; Re-format the Ollama model file. (I forgot the syntax, so 30 minutes of Googling.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Step 3:&lt;/strong&gt; Realize the new model needs a different prompt template to work with my code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Step 4:&lt;/strong&gt; Rewrite my backend logic to accommodate the new output format.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Step 5:&lt;/strong&gt; Re-run my entire test suite because the model's JSON output changed slightly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That took me an entire evening. An entire evening I could have spent building features.&lt;/p&gt;

&lt;p&gt;When I use an API, I just change the model name from &lt;code&gt;gpt-4o-mini&lt;/code&gt; to &lt;code&gt;claude-3-5-sonnet&lt;/code&gt; and adjust the prompt slightly. Done. It took me 5 minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Self-Hosting &lt;em&gt;Does&lt;/em&gt; Make Sense
&lt;/h2&gt;

&lt;p&gt;I want to be fair here. There are cases where self-hosting is the right call.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Regulated Industries:&lt;/strong&gt; If you’re handling PHI (Protected Health Information) or financial data with strict compliance rules, you might &lt;em&gt;need&lt;/em&gt; on-prem inference.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero-Data-Retention:&lt;/strong&gt; If the API provider doesn't offer a zero-data-retention agreement, and you have strict privacy needs, self-hosting might be the only option.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scale:&lt;/strong&gt; If you’re processing billions of tokens a day, the marginal cost of GPU time might beat API per-token costs. But you need to be at &lt;em&gt;enormous&lt;/em&gt; scale for this to kick in.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But for the other 99% of us—the indie hackers, the startup devs, the side project enthusiasts—the API is the better engineering decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Pragmatic Shift
&lt;/h2&gt;

&lt;p&gt;I’m not saying APIs are perfect. They have their own issues: vendor lock-in, rate limits, and the occasional API outage. But they solve the &lt;em&gt;hard&lt;/em&gt; problems for you.&lt;/p&gt;

&lt;p&gt;Here’s what I realized: I don’t want to be a GPU whisperer. I want to build products. I want to focus on the application logic, the UX, and the business value. Managing CUDA drivers, VRAM allocation, and quantization levels is a job—a full-time one.&lt;/p&gt;

&lt;p&gt;When I switched to an API, my productivity skyrocketed. I went from spending 20% of my time on model serving to 0%. I could iterate faster, test more ideas, and ship features that actually matter.&lt;/p&gt;

&lt;p&gt;I also stopped cringing when I opened my electricity bill.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Current Setup
&lt;/h2&gt;

&lt;p&gt;Today, my stack is boring and efficient. I use a managed API for most heavy lifting (chat, summarization, embeddings). For my truly quick-and-dirty experiments, I still have that RTX 3090 sitting in a drawer. But it's a paperweight now.&lt;/p&gt;

&lt;p&gt;If you’re on the fence, do the math. Calculate your hourly rate, estimate the time you'll spend on maintenance, and add up your electricity and hardware costs. Then compare that to the cost of an API.&lt;/p&gt;

&lt;p&gt;For me, the answer was clear: &lt;strong&gt;I value my time more than my GPU.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;Look, I know the open-source purists will hate this post. I get it—I used to be one of them. The idea of running a model locally, free from external dependencies, is incredibly appealing. But the reality is that the infrastructure around these models (the serving layer, the scaling, the security patches) is a full-time job.&lt;/p&gt;

&lt;p&gt;If you’re just starting out, or even if you’re a seasoned dev, don’t make my mistake. Start with an API. Get your product working. Then, if you &lt;em&gt;really&lt;/em&gt; need to self-host, you'll have the revenue and the user base to justify the complexity.&lt;/p&gt;

&lt;p&gt;And if you’re looking for a solid middle ground, I’ve found that a simple API gateway can save you a ton of headaches. I use &lt;strong&gt;tai.shadie-oneapi.com&lt;/strong&gt; for my personal projects—it gives me a unified way to access multiple models without managing the underlying infrastructure. For a hobbyist, it’s been a game-changer. It's not a magic bullet, but it's a hell of a lot easier than babysitting a GPU in your closet.&lt;/p&gt;

&lt;p&gt;Your time is better spent building the future, not fixing your CUDA install. Trust me on that one.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>From Curious to Confident: How I Use AI APIs Without Being a Machine Learning Expert</title>
      <dc:creator>Shaw Sha</dc:creator>
      <pubDate>Mon, 28 Sep 2026 00:55:53 +0000</pubDate>
      <link>https://dev.to/shadie_ai/from-curious-to-confident-how-i-use-ai-apis-without-being-a-machine-learning-expert-2ngk</link>
      <guid>https://dev.to/shadie_ai/from-curious-to-confident-how-i-use-ai-apis-without-being-a-machine-learning-expert-2ngk</guid>
      <description>&lt;p&gt;I remember the exact moment I almost gave up on AI. It was 2 AM, I had three browser tabs open—one explaining what a transformer is, another comparing tokenization methods, and a third that looked like it was written in ancient Greek. I was trying to build a simple chatbot for my portfolio, but I kept hitting walls. Every tutorial assumed I knew what "fine-tuning" meant. Every forum post made me feel like I'd walked into a party where everyone knew the secret handshake.&lt;/p&gt;

&lt;p&gt;Fast forward six months, and I've shipped three separate projects using AI APIs. I still can't explain the math behind attention mechanisms, and honestly, I don't need to. Here's how I went from being paralyzed by the complexity to feeling genuinely confident—and how you can too.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Wake-Up Call: You're Not Building the Model, You're Using It
&lt;/h2&gt;

&lt;p&gt;My breakthrough moment came when I realized something obvious: &lt;strong&gt;I don't need to know how a car engine works to drive to the grocery store.&lt;/strong&gt; The same logic applies to AI. When I use the OpenAI API, Anthropic's Claude, or any of the other major providers, I'm not building a model from scratch. I'm making a request to a service that someone else spent millions of dollars training.&lt;/p&gt;

&lt;p&gt;The first time I actually got a response back from an API call, it was almost anticlimactic. Twenty lines of Python, and I had a program that could write a haiku about my cat. That's it. That's the whole magic trick.&lt;/p&gt;

&lt;p&gt;Here's what that first successful call looked like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://tai.shadie-oneapi.com/v1/chat/completions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer YOUR_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-4o-mini&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;messages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Write a haiku about a cat who loves debugging code&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;choices&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it. No machine learning degree required. No math beyond what you learned in high school. Just a POST request with a JSON body—the same thing you've probably done a hundred times with other APIs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 80/20 Rule of AI APIs
&lt;/h2&gt;

&lt;p&gt;After building those three projects, I've noticed that 80% of what you'll ever do with AI APIs falls into four simple patterns:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Text generation&lt;/strong&gt; - "Write me a summary of this article"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Classification&lt;/strong&gt; - "Is this email a complaint or a compliment?"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Extraction&lt;/strong&gt; - "Pull out all the dates from this text"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transformation&lt;/strong&gt; - "Rewrite this in a more professional tone"&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's it. Everything else—agents, RAG, chain-of-thought prompting—is just clever combinations of these four basics. Once I internalized that, the anxiety melted away.&lt;/p&gt;

&lt;h2&gt;
  
  
  My First Real Project: A Support Ticket Classifier
&lt;/h2&gt;

&lt;p&gt;Let me walk you through my first actual use case, because it's the perfect example of how simple this can be. My friend runs a small e-commerce store, and she was drowning in customer emails. She asked if I could build something to sort them into categories: "problem," "question," "praise," or "return request."&lt;/p&gt;

&lt;p&gt;My first instinct was panic. I started thinking about natural language processing, sentiment analysis algorithms, training data... but then I stopped myself. I just needed to ask the API to do the classification for me.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;axios&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;axios&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;classifyTicket&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;axios&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://tai.shadie-oneapi.com/v1/chat/completions&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-4o-mini&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;system&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;You are a customer service classifier. Respond with ONLY one word: problem, question, praise, or return.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
      &lt;span class="p"&gt;],&lt;/span&gt;
      &lt;span class="na"&gt;temperature&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.1&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Authorization&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Bearer YOUR_API_KEY&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Usage&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;email&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Hey, my order #4829 arrived but the box was crushed. Can I send it back?&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;category&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;classifyTicket&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;email&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;category&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// "problem"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The whole thing took me about 45 minutes to write. My friend has been using it for three months now, and it's caught over 1,200 emails, correctly categorizing them about 94% of the time. That's better than any manual solution we could have built.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Hidden Costs Nobody Tells You About
&lt;/h2&gt;

&lt;p&gt;Here's what I wish someone had told me before I started: the API costs money, but not in the way you think. The token system can be sneaky. I once burned through $15 in a weekend because I was sending giant blocks of text without realizing that both input &lt;strong&gt;and&lt;/strong&gt; output are charged.&lt;/p&gt;

&lt;p&gt;A quick tip: always set &lt;code&gt;max_tokens&lt;/code&gt; on your requests. Otherwise, you might get a 4,000-token response when you only needed 100. That's like ordering a single taco and being charged for the entire catering service.&lt;/p&gt;

&lt;p&gt;Also, &lt;strong&gt;don't use GPT-4 for everything&lt;/strong&gt;. For simple tasks like classification, the smaller models (like &lt;code&gt;gpt-4o-mini&lt;/code&gt; or &lt;code&gt;claude-3-haiku&lt;/code&gt;) are 10x cheaper and just as accurate. I was blowing through credits for months before I realized I was using a sledgehammer to crack a walnut.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Prompt Engineering Myth
&lt;/h2&gt;

&lt;p&gt;Everyone talks about "prompt engineering" like it's a mystical art form. It's not. It's just being clear about what you want. Here are the three rules I actually follow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Be specific about the output format&lt;/strong&gt; - "Respond with JSON" is better than "Give me the info"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Provide context&lt;/strong&gt; - The system message is your friend&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Show examples&lt;/strong&gt; - One example is worth a thousand words of explanation&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's it. You don't need to learn about temperature curves or sampling strategies. Just be the kind of person who writes clear instructions.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Things Go Wrong (They Will)
&lt;/h2&gt;

&lt;p&gt;I had a moment last month where my classification bot suddenly started responding with "I cannot assist with that request" for every email. Turns out, someone on the e-commerce site had sent an email that triggered the content filter, and I hadn't handled that response type in my code.&lt;/p&gt;

&lt;p&gt;Here's the lesson: &lt;strong&gt;AI APIs are not deterministic.&lt;/strong&gt; The same input can give you different outputs 5% of the time. You need to handle errors, edge cases, and weird responses. My code now always has a fallback:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choices&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;uncertain&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This ten lines of defensive programming has saved me more headaches than any amount of theoretical knowledge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building Confidence Through Iteration
&lt;/h2&gt;

&lt;p&gt;The most important shift in my mindset was going from "I need to understand everything first" to "I'll build something small and iterate." My first project was a joke generator that only worked 70% of the time. My second was a summary tool that was just okay. By the third project, I felt like I actually knew what I was doing.&lt;/p&gt;

&lt;p&gt;The confidence didn't come from mastering the technology. It came from shipping things and fixing problems as they appeared. Just like learning any other skill, the reps matter more than the theory.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Setup I Use Now
&lt;/h2&gt;

&lt;p&gt;If you're curious about getting started, here's a practical recommendation: find an API aggregator or gateway that gives you access to multiple models with one key. I use &lt;strong&gt;tai.shadie-oneapi.com&lt;/strong&gt; as my endpoint for most of my projects. It's not because I'm an expert—it's because it removes the friction of managing multiple accounts and billing cycles. One key, one endpoint, and I can switch between different models depending on the task.&lt;/p&gt;

&lt;p&gt;The URL I've been using in the code examples above is from that service. It works with the standard OpenAI SDK, so you don't need to learn any new tools. Just swap out the base URL and you're good to go.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your First Step
&lt;/h2&gt;

&lt;p&gt;Don't start by reading a textbook on neural networks. Start by making a single API call. Literally, copy the first code block in this article, paste it into your editor, and run it. Then change the prompt. Then add a loop. Then build something that solves a problem you actually have.&lt;/p&gt;

&lt;p&gt;The barrier to entry for AI development isn't knowledge anymore—it's confidence. And confidence comes from doing, not studying.&lt;/p&gt;

&lt;p&gt;I'm still not an AI expert. I couldn't tell you what a "multi-head attention mechanism" actually does without Googling it. But I've built tools that people use every day, and that's a pretty good definition of "good enough." Now go build something.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
      <category>tutorial</category>
      <category>javascript</category>
    </item>
    <item>
      <title>The Silent Costs of AI APIs Nobody Warns You About</title>
      <dc:creator>Shaw Sha</dc:creator>
      <pubDate>Sun, 27 Sep 2026 00:55:53 +0000</pubDate>
      <link>https://dev.to/shadie_ai/the-silent-costs-of-ai-apis-nobody-warns-you-about-41e1</link>
      <guid>https://dev.to/shadie_ai/the-silent-costs-of-ai-apis-nobody-warns-you-about-41e1</guid>
      <description>&lt;p&gt;I remember the day I got my first API bill from a major AI provider. I'd been building a prototype chatbot for a client, meticulously tracking my token usage, convinced I had a handle on costs. The estimate I'd given was $200 a month. The bill was $1,400.&lt;/p&gt;

&lt;p&gt;I stared at the invoice for a solid five minutes, convinced it was a glitch. It wasn't.&lt;/p&gt;

&lt;p&gt;That was the moment I realized that AI API pricing is a bit like buying a car—the sticker price is just the beginning. The real costs are hidden in the fine print, in the infrastructure you didn't plan for, and in the code you'll have to rewrite when the ground shifts beneath you.&lt;/p&gt;

&lt;p&gt;Here’s what nobody tells you about the silent costs of AI APIs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Token Trap
&lt;/h2&gt;

&lt;p&gt;The first trap is the most obvious, yet the most commonly miscalculated: token usage is not linear. We all know the formula: input tokens cost less than output tokens. But the way we &lt;em&gt;estimate&lt;/em&gt; those tokens in the planning phase is almost always wrong.&lt;/p&gt;

&lt;p&gt;I built a summarization tool that processes support tickets. The tickets were short, maybe 50 words each. I estimated that a single API call would use about 200 tokens. Easy.&lt;/p&gt;

&lt;p&gt;But then I added context. To make the summary accurate, I needed to feed the AI the previous conversation thread, the customer’s history, and the product details. Suddenly, that 200-token request turned into a 2,000-token request. The output, which I wanted to be a concise summary, often came back as a verbose paragraph.&lt;/p&gt;

&lt;p&gt;The ratio of input to output matters. If you’re building a chain-of-thought system or a multi-step agent, you aren't just paying for the final answer. You’re paying for every intermediary thought, every failed attempt, and every retry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; Log every single request with actual token counts from the response headers. Don’t trust your estimates. Build a dashboard that shows you the cost per successful &lt;em&gt;user action&lt;/em&gt;, not per API call. I found that one "simple" user query often triggers 15-20 API calls in the background.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Latency Tax
&lt;/h2&gt;

&lt;p&gt;The next silent cost is latency. Not in terms of user experience (though that matters), but in terms of &lt;em&gt;infrastructure&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;To make an AI-powered feature feel responsive, you can't just call the API synchronously and wait. You need background workers, queues, and async processing. That means you’re now paying for a message queue service, a worker server (or Lambda functions), and the cold-start time of your containers.&lt;/p&gt;

&lt;p&gt;I once built a "smart search" feature that embedded user queries and compared them against a vector database. The API cost was negligible. But the vector database? That was $70 a month. The GPU instance I thought I needed to keep latency low? Another $50. The Redis cache to store common embeddings? Twenty bucks.&lt;/p&gt;

&lt;p&gt;The AI API was the cheapest part of the entire stack.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The reality:&lt;/strong&gt; When you calculate the ROI of an AI feature, you have to add 50-100% on top of the API cost just for the plumbing. The "serverless" promise often breaks down when you have a long-running streaming response that holds a connection open.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Rate Limit Whiplash
&lt;/h2&gt;

&lt;p&gt;Nothing breaks a production app faster than a 429 status code. Rate limits are the hidden tax on your development time.&lt;/p&gt;

&lt;p&gt;The providers give you these nice tiers: RPM (requests per minute), TPM (tokens per minute), and IPM (images per minute). They look generous on paper. But they are calculated based on &lt;em&gt;their&lt;/em&gt; infrastructure, not your traffic patterns.&lt;/p&gt;

&lt;p&gt;I moved a feature to production and immediately hit a wall. My traffic spiked at 9:00 AM when users logged in. The provider’s limit was 3,000 RPM, but I was bursting at 2,500 and getting throttled hard. Why? Because the limits are shared across a pool, and if you have a burst that exceeds the &lt;em&gt;tier&lt;/em&gt; average, you get blocked.&lt;/p&gt;

&lt;p&gt;The workaround is to implement exponential backoff and retries. But that means your code needs to be more resilient, which means more engineering time. And let me tell you, debugging a system that fails only when an external service says "slow down" is a nightmare.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The trick:&lt;/strong&gt; Build your own internal rate limiter that sits at 70% of your provider’s limit. This gives you headroom. But more importantly, architect your system to handle failure gracefully. If the API is down, your app should still function, just with reduced features. That's a fallback layer you have to build, test, and maintain.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Token Eater: System Prompts
&lt;/h2&gt;

&lt;p&gt;Here is the cost that nobody mentions in the marketing blogs: the system prompt.&lt;/p&gt;

&lt;p&gt;We all write these massive, detailed system prompts to get the AI to behave correctly. I have one that is about 1,500 tokens long. That’s fine for a single call.&lt;/p&gt;

&lt;p&gt;But if you’re building an agent that calls tools, you have to include the tool schemas. Those schemas are verbose JSON. Add another 800 tokens. Now, every single interaction with the model has a 2,300 token overhead &lt;em&gt;before the user even types a single character&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;If you have 10,000 users a day, and each of them sends 5 messages, that’s 50,000 calls. Multiply that by the overhead tokens, and you’re burning through 115 million tokens &lt;em&gt;just on the system prompt&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;That’s not the cost of the AI; that’s the cost of the software engineering around it. I started using dynamic prompt compression—only including the relevant tool schemas for the current step. It cut my costs by 40% immediately.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Vendor Lock-In Paradox
&lt;/h2&gt;

&lt;p&gt;The most expensive cost of all, though, is the switching cost.&lt;/p&gt;

&lt;p&gt;You spend three months building a pipeline that works perfectly with OpenAI’s function calling. You optimize for their tokenizer. You use their specific parameters for temperature and top_p. You build a feedback loop that relies on their logprobs.&lt;/p&gt;

&lt;p&gt;Then, their pricing changes. Or they deprecate a model. Or the CFO looks at the bill and says, "Can we switch to the open-source model running on our own hardware?"&lt;/p&gt;

&lt;p&gt;You say "yes," and then you realize that the data you have stored in their format needs to be converted. Your prompts need to be rewritten because the new model doesn't follow instructions the same way. Your evaluation suite needs to be recalibrated because the output distribution is different.&lt;/p&gt;

&lt;p&gt;I spent two weeks converting a pipeline to a different provider last year. The code was abstracted (I had a wrapper layer), but the &lt;em&gt;behavior&lt;/em&gt; was not. The new model was smarter, but it gave me different formatting. My regex to parse the output broke. My validation logic failed.&lt;/p&gt;

&lt;p&gt;The cost of that migration wasn't in API fees. It was in the opportunity cost of two weeks of my salary that I couldn't spend on new features.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Current Approach
&lt;/h2&gt;

&lt;p&gt;I've learned to stop treating AI APIs as a commodity. They are specialized services, and you have to plan for their quirks.&lt;/p&gt;

&lt;p&gt;Today, I look for three things before I commit to a provider:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Transparency:&lt;/strong&gt; Can I see a live cost calculator? Or do they hide the pricing behind a "contact sales" link?&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Simplicity:&lt;/strong&gt; Are the rate limits clear? Is the token counting straightforward?&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Stability:&lt;/strong&gt; Do they have a track record of keeping models available, or do they retire them every six months?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Recently, I’ve moved a few of my hobby projects over to a service that offers a pay-as-you-go model without the aggressive rate limiting I was facing. It’s not a huge enterprise solution, but for prototyping and mid-scale production, it takes away the "surprise bill" anxiety.&lt;/p&gt;

&lt;p&gt;If you’re tired of the spreadsheet gymnastics required to forecast your monthly spend, it’s worth checking out a provider that keeps the pricing model simple and transparent. I use &lt;strong&gt;tai.shadie-oneapi.com&lt;/strong&gt; for a few side projects now—it aggregates several models and bills me a flat rate per token with no hidden infrastructure costs. It’s not the only option out there, but it’s the one that finally let me sleep at night without worrying about a $1,400 invoice.&lt;/p&gt;

&lt;p&gt;The bottom line is this: AI APIs are powerful, but they are not cheap. The cost isn't the API call; it's the system you have to build around it to make it reliable, fast, and maintainable. Plan for the hidden costs, and you won't get burned. Ignore them, and you'll be staring at a spreadsheet wondering where your budget went.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
    <item>
      <title>AI APIs in 2026: The Honest Developer's Guide to Choosing One</title>
      <dc:creator>Shaw Sha</dc:creator>
      <pubDate>Sat, 26 Sep 2026 00:55:24 +0000</pubDate>
      <link>https://dev.to/shadie_ai/ai-apis-in-2026-the-honest-developers-guide-to-choosing-one-4mhe</link>
      <guid>https://dev.to/shadie_ai/ai-apis-in-2026-the-honest-developers-guide-to-choosing-one-4mhe</guid>
      <description>&lt;p&gt;Choosing an AI API in 2026 isn't about picking the "best" model—it's about picking the right tradeoff. I've spent the last three years building everything from internal chat bots to production-grade document parsers, and I've been burned more times than I'd like to admit.&lt;/p&gt;

&lt;p&gt;The reality is that the landscape has shifted dramatically. In 2023, we had a handful of serious players. Today, the market is fragmented into dozens of providers, each with their own quirks, pricing models, and hidden gotchas. I've watched developers get paralyzed by choice, and I've watched others lock themselves into a single vendor and regret it six months later.&lt;/p&gt;

&lt;p&gt;This guide is the result of my own trial-and-error—the things I wish someone had told me before I wasted $2,000 on API calls that could have been done with a regex and a prayer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three axes that actually matter
&lt;/h2&gt;

&lt;p&gt;Before you even look at benchmark scores, you need to understand that every AI API decision boils down to three axes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Quality&lt;/strong&gt; — how well the model handles your specific use case&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost&lt;/strong&gt; — both per-token price and the infrastructure overhead of using it&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reliability&lt;/strong&gt; — uptime, rate limits, and consistency of responses&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Most developers obsess over quality first. That's a mistake. I've seen teams spend hours comparing GPT-4.5 vs Claude 4 vs Gemini 2.5 benchmark scores, only to discover that their actual bottleneck was response latency during peak hours—something no benchmark will tell you.&lt;/p&gt;

&lt;h3&gt;
  
  
  The quality trap
&lt;/h3&gt;

&lt;p&gt;Here's what I've learned: benchmark scores are only marginally useful for real-world applications. A model that scores 92% on MMLU might completely fail at extracting structured data from messy invoices. I had a project where Claude consistently outperformed GPT on code generation, but the gap was almost invisible in day-to-day use.&lt;/p&gt;

&lt;p&gt;The smarter approach is to build a small evaluation set that represents &lt;em&gt;your&lt;/em&gt; actual workload—maybe 50 to 100 examples—and run every candidate model through it. I've built a simple script that automates this, and it's saved me from making expensive mistakes more than once.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I actually use in production
&lt;/h2&gt;

&lt;p&gt;Let me break down my current stack, because I think it's fairly representative of where the industry is heading.&lt;/p&gt;

&lt;h3&gt;
  
  
  OpenAI (GPT-4.5 and beyond)
&lt;/h3&gt;

&lt;p&gt;OpenAI is still the default for most developers, and for good reason. The API is clean, the documentation is excellent, and the ecosystem around it—from LangChain integrations to fine-tuning tools—is unmatched. But the cost has been creeping up. For a typical chat completion with moderate context, I'm looking at roughly $0.01 to $0.03 per query on the standard tier.&lt;/p&gt;

&lt;p&gt;The biggest annoyance? Rate limits. At the tier I'm on, I've hit throttling during peak usage more times than I can count. It's the classic "works fine in dev, falls over in production" scenario.&lt;/p&gt;

&lt;h3&gt;
  
  
  Anthropic (Claude 4 and beyond)
&lt;/h3&gt;

&lt;p&gt;I'll be honest—I was late to the Claude bandwagon. I dismissed it as the "safety-focused" alternative that couldn't compete on raw capability. I was wrong.&lt;/p&gt;

&lt;p&gt;For code generation, Claude has been consistently better in my testing. The contextual understanding of multi-file changes is eerie. I had a refactoring task that involved renaming a class across 20+ files, and Claude handled it with maybe 8% hallucination rate compared to GPT's 15%. That's not a small difference when you're dealing with a 10,000-line codebase.&lt;/p&gt;

&lt;p&gt;The downside is less straightforward integration with existing tooling. If you're deeply embedded in the OpenAI ecosystem, switching takes effort.&lt;/p&gt;

&lt;h3&gt;
  
  
  Google (Gemini 2.5 and Pro)
&lt;/h3&gt;

&lt;p&gt;Gemini has gotten legitimately good, but I still find it awkward for most developer workflows. The API is solid, but the documentation feels like it was written by a product manager who never actually built anything. The context window is impressive—I've fed it entire codebases for analysis—but the response quality degrades noticeably at the edges of that context.&lt;/p&gt;

&lt;p&gt;Where Gemini shines is multimodal processing. If you're doing heavy image or video analysis, it's worth considering.&lt;/p&gt;

&lt;h2&gt;
  
  
  A working example: parsing messy data
&lt;/h2&gt;

&lt;p&gt;Let me show you something real. I recently built a tool that extracts order information from raw email threads. Here's the core function:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;extract_order_details&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;email_thread&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;system_prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    You are an order extraction assistant. Parse the email thread 
    and extract order details in JSON format with fields:
    - order_id
    - customer_name
    - items (array of {item_name, quantity, unit_price})
    - total_amount
    - shipping_address

    Handle inconsistencies in formatting. If data is missing, 
    omit the field entirely—do not hallucinate.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-4-turbo&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# or gpt-4.5 if available
&lt;/span&gt;        &lt;span class="n"&gt;response_format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;json_object&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;system_prompt&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;email_thread&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The interesting thing is that I tested this exact function across three providers. GPT-4-turbo had a 94% success rate on my eval dataset. Claude 4 hit 97%. Gemini 2.5 was at 91%. But none of that mattered as much as the failure modes.&lt;/p&gt;

&lt;p&gt;When GPT failed, it was often due to merging two different customers in the same thread. Claude's failures were more about missing edge-case fields. Gemini just gave up more often on long threads and returned empty responses. Understanding these patterns made the choice much easier than chasing benchmarks.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reliability gamble
&lt;/h2&gt;

&lt;p&gt;This is the part that nobody talks about enough. I've had API providers go down at 3 AM on a Saturday. I've had response times balloon from 2 seconds to 30 seconds without warning. I've had models get deprecated with two weeks' notice, breaking my production code.&lt;/p&gt;

&lt;p&gt;My advice: &lt;strong&gt;design for failure from day one.&lt;/strong&gt; Build a fallback chain. If your primary provider is down, route to a secondary one. If you're using a large context window, prepare a shortened prompt as a backup. The extra work upfront is nothing compared to the cost of a production outage.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pricing puzzle
&lt;/h2&gt;

&lt;p&gt;The per-token pricing is only half the story. What really matters is the total cost of ownership. Here's what I mean:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Compute overhead&lt;/strong&gt;: Some APIs have better caching, which can cut your effective cost by 60-70%&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retry logic&lt;/strong&gt;: If a provider has flaky reliability, you'll burn tokens on retries&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Response formatting&lt;/strong&gt;: If the API forces you to make multiple calls to get structured output, that's 3x the cost&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When I crunched my actual invoices across a three-month period, I found that my effective cost per successful request was between 2x and 4x the listed price for some providers. That's a huge difference.&lt;/p&gt;

&lt;h2&gt;
  
  
  A comparison table that actually helps
&lt;/h2&gt;

&lt;p&gt;I've assembled a practical comparison based on my production experience, not marketing pages:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Provider&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;th&gt;Typical Cost (per 1K tokens)&lt;/th&gt;
&lt;th&gt;Gotcha&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;General purpose, ecosystem&lt;/td&gt;
&lt;td&gt;$0.02-$0.06&lt;/td&gt;
&lt;td&gt;Rate limits, rising costs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;Code generation, long context&lt;/td&gt;
&lt;td&gt;$0.015-$0.05&lt;/td&gt;
&lt;td&gt;Tooling integration is rough&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google Gemini&lt;/td&gt;
&lt;td&gt;Multimodal, large context&lt;/td&gt;
&lt;td&gt;$0.01-$0.04&lt;/td&gt;
&lt;td&gt;Inconsistent quality at context edges&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;shadie-oneapi&lt;/td&gt;
&lt;td&gt;Multi-provider access on demand&lt;/td&gt;
&lt;td&gt;Varies by usage&lt;/td&gt;
&lt;td&gt;Simpler integration, no monthly commitments&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Let me talk about that last one for a second. I've been using shadie-oneapi for a few months now, and it's solved the biggest headache I had: switching providers without rewriting code. It gives me instant access to multiple models from a single API key, and there's no monthly fee—I pay for what I actually use. For someone like me who runs experiments across different models weekly, that's a game-changer.&lt;/p&gt;

&lt;h2&gt;
  
  
  My actual recommendation
&lt;/h2&gt;

&lt;p&gt;Here's the honest truth: there's no single best AI API. It depends on what you're building.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;If you're prototyping&lt;/strong&gt;: Start with OpenAI. The ecosystem is the smoothest, and you'll get to a working demo fastest.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you're doing serious code generation&lt;/strong&gt;: Invest in setting up Claude. The quality difference in code is worth the integration friction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you're dealing with documents or images&lt;/strong&gt;: Give Gemini a serious look, but build thorough eval tests first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you're building something you want to maintain long-term&lt;/strong&gt;: Abstract your provider calls behind an interface. Make it so you can swap providers in a configuration file, not in your core logic. This is where I've saved the most time and money.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The bottom line
&lt;/h2&gt;

&lt;p&gt;Don't chase benchmarks. Don't lock yourself into a single vendor. Design for flexibility, measure on your own data, and always keep a fallback plan.&lt;/p&gt;

&lt;p&gt;The AI API space is moving fast, and what's "best" today might be obsolete in six months. Build your systems to adapt, not to commit. And if you want to simplify your life by accessing multiple providers without the overhead, check out shadie-oneapi.com—I've found it to be a genuinely practical addition to my stack.&lt;/p&gt;

&lt;p&gt;In the end, the right AI API is the one that handles your specific problem, at a cost you can tolerate, with reliability you can trust. Everything else is just noise.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>tutorial</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Building an AI Side Project That Actually Ships — Lessons from Shipping 3 MVPs</title>
      <dc:creator>Shaw Sha</dc:creator>
      <pubDate>Fri, 25 Sep 2026 00:56:01 +0000</pubDate>
      <link>https://dev.to/shadie_ai/building-an-ai-side-project-that-actually-ships-lessons-from-shipping-3-mvps-52ni</link>
      <guid>https://dev.to/shadie_ai/building-an-ai-side-project-that-actually-ships-lessons-from-shipping-3-mvps-52ni</guid>
      <description>&lt;p&gt;There’s a graveyard of half-finished AI projects on my hard drive. I’m not ashamed to admit it. There’s the chatbot that was supposed to summarize my emails (it died because I realized I hate reading emails anyway). There’s the "smart" habit tracker that required more effort to log data than the habit itself. And there’s the one that got 5,000 stars on GitHub before I realized I had no idea how to monetize it (we don't talk about that one).&lt;/p&gt;

&lt;p&gt;But then, something clicked. In the last two months, I shipped three AI-powered side projects that are actually being used by real people. Not millions, but hundreds. They’re not unicorns, but they’re alive. They don't have massive cloud bills, and they don't take up my weekends.&lt;/p&gt;

&lt;p&gt;This isn't a story about a magical framework or a genius idea. It’s about ditching the "perfect" engineering mindset and embracing the ugly, pragmatic reality of shipping. Here are the lessons I learned the hard way, so you don't have to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 1: The Model is a Commodity, Not the Product
&lt;/h2&gt;

&lt;p&gt;The biggest mental shift was realizing that the LLM itself is the least interesting part of the product. The magic isn't in the API call; it's in the orchestration, the prompt engineering, and the user experience around it.&lt;/p&gt;

&lt;p&gt;My first failed project, "EmailSage," tried to fine-tune a model on my personal writing style. I spent two weeks collecting data, cleaning it, and messing with LoRA adapters. It was a technical nightmare. The result was a model that sounded vaguely like me but took 10 seconds to respond and cost a fortune to run.&lt;/p&gt;

&lt;p&gt;For my next project, &lt;strong&gt;"RecipeRehab"&lt;/strong&gt; (an app that tells you what to cook based on the random ingredients in your fridge), I did the opposite. I used a standard, off-the-shelf model. No fine-tuning. No custom hosting. Just a smart system prompt that included the list of ingredients and a JSON output schema.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_recipe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ingredients&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]):&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-4o-mini&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;# The cheap, fast one
&lt;/span&gt;        &lt;span class="n"&gt;response_format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;json_object&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
            You are a culinary genius. Given a list of ingredients, suggest one recipe.
            You MUST respond in JSON format with the following structure:
            {
              &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Recipe Name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;,
              &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;time_to_cook&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;15 mins&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;,
              &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;instructions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;: [&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;step1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;step2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;],
              &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;missing_ingredients&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;: [&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;salt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pepper&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;]
            }
            &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()},&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Here are my ingredients: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ingredients&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That was it. The entire "AI" part was about 20 lines of code. The real work—and the real value—was in the front-end design that made it dead simple to type in "chicken, rice, soy sauce" and get a delicious meal in under 3 seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Takeaway:&lt;/strong&gt; Stop trying to build your own model. Stop trying to fine-tune one. The prompt is your product. The UX is your moat. If you're spending more time on the ML pipeline than the user flow, you're building the wrong thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 2: The 80/20 Rule of Infrastructure (or, Why I Stopped Hosting)
&lt;/h2&gt;

&lt;p&gt;I used to be a "self-host everything" guy. I have a beefy server in my closet at home. I ran a Kubernetes cluster for a project that got maybe 20 hits a day. It was absurd. I was spending more time patching the cluster than building the app.&lt;/p&gt;

&lt;p&gt;For my second project, &lt;strong&gt;"ChillBeats"&lt;/strong&gt; (an AI that generates ambient soundscapes based on your current mood), I hit a wall. I initially tried to host an open-source music generation model locally. The latency was terrible, the CPU was maxed out, and my cat was getting angry at the fan noise.&lt;/p&gt;

&lt;p&gt;I pivoted. Instead of hosting a model, I used a hosted API. This changed everything.&lt;/p&gt;

&lt;p&gt;The speed to market was insane. What took me a week of fighting with CUDA drivers and Python environments took me an afternoon to integrate via an API. I didn't have to worry about scaling, uptime, or GPU costs. I just paid for what I used.&lt;/p&gt;

&lt;p&gt;This is where I have to be completely honest. The biggest game-changer for me was moving to a &lt;strong&gt;pay-as-you-go API&lt;/strong&gt; model rather than a subscription. With a monthly plan, I was constantly worried about hitting my token limit. With pay-as-you-go, I just... stopped caring. I could experiment without the anxiety of a cap.&lt;/p&gt;

&lt;p&gt;For &lt;code&gt;ChillBeats&lt;/code&gt;, I ended up building a simple wrapper around a few different providers, switching based on the task. For the sound generation, I used one API; for the mood analysis of the user's text, I used another. This modular approach meant if one provider had a hiccup, the other kept working.&lt;/p&gt;

&lt;p&gt;The infrastructure stack for all three projects is now embarrassingly simple:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Frontend:&lt;/strong&gt; Vercel for static hosting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backend:&lt;/strong&gt; A single serverless function (Node.js).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Database:&lt;/strong&gt; A simple KV store (Redis).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI:&lt;/strong&gt; A mix of hosted APIs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's it. No Docker Compose files. No &lt;code&gt;docker-compose.yml&lt;/code&gt; nightmares. No &lt;code&gt;nginx&lt;/code&gt; configs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 3: The "Good Enough" Launch
&lt;/h2&gt;

&lt;p&gt;My third project, &lt;strong&gt;"StudyMate"&lt;/strong&gt;, is a flashcard app that generates quizzes from your class notes. I had this grand vision of a collaborative learning platform with sharing, leaderboards, and gamification.&lt;/p&gt;

&lt;p&gt;I launched with none of that.&lt;/p&gt;

&lt;p&gt;I launched with a text box, a "Generate" button, and a list of flashcards. It was ugly. The CSS was a single file, and it looked like it was made in 2005. But it worked.&lt;/p&gt;

&lt;p&gt;And you know what? People used it. They didn't care about the missing leaderboards. They cared that they could paste 5,000 words of dense neuroscience notes and get a quiz in 10 seconds.&lt;/p&gt;

&lt;p&gt;I'm a firm believer in the "ugly launch" now. If the core value proposition is strong, people will forgive a poor UI. But if the UI is beautiful and the AI is slow or inaccurate, they will leave immediately.&lt;/p&gt;

&lt;p&gt;Here’s a graph I wish I had seen before I started:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Week 1:&lt;/strong&gt; Build the core feature (the API call).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Week 2:&lt;/strong&gt; Make the UI functional but not pretty.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Week 3:&lt;/strong&gt; Tell 10 people about it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Week 4:&lt;/strong&gt; Iterate based on feedback.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I spent my first month on &lt;code&gt;EmailSage&lt;/code&gt; doing Week 1 for three weeks. I was polishing the "prompt chain" for a feature that nobody had asked for.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Financial Reality
&lt;/h2&gt;

&lt;p&gt;Let's talk numbers, because it's the thing everyone is scared to talk about.&lt;/p&gt;

&lt;p&gt;My monthly bill for hosting all three projects last month was &lt;strong&gt;$3.41&lt;/strong&gt;. That's it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Vercel:&lt;/strong&gt; $0 (Hobby tier)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Redis:&lt;/strong&gt; $0 (Free tier)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API Usage:&lt;/strong&gt; $3.41 (The vast majority of this was me testing and playing with the prompts).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I have one project that has 200 active users. At the current usage rate, it costs me about $0.005 per user per month. That is insanely cheap.&lt;/p&gt;

&lt;p&gt;This is why the "AI gold rush" feels different. The infrastructure cost is a rounding error. The cost is your time. So, optimize for your time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Pragmatic Path Forward
&lt;/h2&gt;

&lt;p&gt;If you're looking to start an AI side project tomorrow, here is my brutally honest advice:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Don't&lt;/strong&gt; buy a GPU.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Don't&lt;/strong&gt; set up a vector database until you have 10,000 documents to search.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Don't&lt;/strong&gt; write a full spec document. Write a paragraph.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Do&lt;/strong&gt; use a hosted model that gives you a simple &lt;code&gt;curl&lt;/code&gt; command.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Do&lt;/strong&gt; focus on the input/output flow. What does the user type? What do they see?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And most importantly, &lt;strong&gt;ship it on day one&lt;/strong&gt;. Not day 30. Day one. Get the ugly version out there.&lt;/p&gt;

&lt;p&gt;For my infrastructure, I've settled into a comfortable rhythm. I use a mix of services, but I keep coming back to the ones that let me pay as I go. I don't want a contract; I don't want a subscription that I forget to cancel.&lt;/p&gt;

&lt;p&gt;That's why I've been leaning heavily on services like &lt;strong&gt;tai.shadie-oneapi.com&lt;/strong&gt; for my AI calls. It aggregates a bunch of models behind a single, simple API, and the pay-as-you-go model means I'm never locked into a specific vendor or a monthly quota. It just works, and I only pay for the tokens I actually burn. It’s the kind of "set and forget" infrastructure that lets me focus on the app, not the plumbing.&lt;/p&gt;

&lt;p&gt;The dream isn't to build the next ChatGPT. The dream is to build a tiny, useful tool that solves a specific problem for a niche group of people. And you don't need to be an AI researcher to do that. You just need to be a developer who can glue two APIs together and ship.&lt;/p&gt;

&lt;p&gt;Stop polishing. Start shipping.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>beginners</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How I Cut My LLM API Costs by 70% Without Touching My Code</title>
      <dc:creator>Shaw Sha</dc:creator>
      <pubDate>Thu, 24 Sep 2026 00:57:48 +0000</pubDate>
      <link>https://dev.to/shadie_ai/how-i-cut-my-llm-api-costs-by-70-without-touching-my-code-4m2d</link>
      <guid>https://dev.to/shadie_ai/how-i-cut-my-llm-api-costs-by-70-without-touching-my-code-4m2d</guid>
      <description>&lt;p&gt;I was staring at my credit card statement, and it wasn't pretty. $217.43 on AI APIs in a single month. For a solo developer building a side project that wasn't even generating revenue yet, that was painful.&lt;/p&gt;

&lt;p&gt;The worst part? I knew I was being wasteful. I just didn't know how wasteful until I actually dug into the numbers.&lt;/p&gt;

&lt;p&gt;So I spent a weekend doing what any reasonable developer would do: I treated my AI API usage like a performance bug. I profiled it, broke it down, and rebuilt the pipeline piece by piece. By Monday, my monthly bill was projected at $63. Same output quality. Same code. Just a fundamentally different approach to how I talk to these models.&lt;/p&gt;

&lt;p&gt;Here's exactly how I did it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 1: I Stopped Overprovisioning Every Request
&lt;/h2&gt;

&lt;p&gt;The first thing I discovered was embarrassing. I was sending max tokens of 4096 on every single call. For everything. Even simple classification tasks that needed a yes/no answer.&lt;/p&gt;

&lt;p&gt;Here's a snapshot of what my calls looked like before:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-4o&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are a helpful assistant that categorizes emails.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Categorize this email: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;email_text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;4096&lt;/span&gt;  &lt;span class="c1"&gt;# Way overkill for this task
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I was paying for a truck when I needed a bicycle. When I checked the actual usage, most of my responses were coming back with 80-200 tokens. So I cut &lt;code&gt;max_tokens&lt;/code&gt; down to 512 for classification tasks and 2048 for generation tasks. That single change shaved off about 15% of my costs right away.&lt;/p&gt;

&lt;p&gt;But that was just the warm-up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: I Realized I Was Asking Stupid Questions
&lt;/h2&gt;

&lt;p&gt;Here's the thing about LLM APIs: you're not paying for the question. You're paying for the thinking. And I was making every model do a PhD-level analysis before answering basic questions.&lt;/p&gt;

&lt;h3&gt;
  
  
  The fix: cascading model tiers
&lt;/h3&gt;

&lt;p&gt;I built a simple router that uses small models for small problems. My rule of thumb is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GPT-4o or Claude 3.5 Sonnet&lt;/strong&gt;: Complex reasoning, multi-step planning, code generation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPT-4o-mini or Claude 3 Haiku&lt;/strong&gt;: Summarization, extraction, classification
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPT-3.5 Turbo or Llama 3.1 8B&lt;/strong&gt;: Trivial tasks where a wrong answer costs nothing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here's the logic that saved me the most:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;smartCompletion&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;complexity&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;estimateComplexity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;complexity&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;simple&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;callModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-4o-mini&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;  &lt;span class="c1"&gt;// ~2.5x cheaper per token&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;complexity&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;medium&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;callModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-4o-mini&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;max_tokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2048&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;callModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-4o&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;estimateComplexity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;extract&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;categorize&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;simple&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;summarize&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;rewrite&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;medium&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;complex&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You'd be surprised how many "complex" tasks are actually simple. When I started logging prompts and checking their actual output, about 70% of my calls were handled by the mini tier without any quality drop.&lt;/p&gt;

&lt;p&gt;That moved the needle. Hard.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Caching Was My Biggest Blindspot
&lt;/h2&gt;

&lt;p&gt;Look, I know "add caching" sounds like the most boring advice in the world. But let me show you the numbers from my own logs:&lt;/p&gt;

&lt;p&gt;In a typical week, I was sending &lt;strong&gt;11,847 identical requests&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Not similar.&lt;/p&gt;

&lt;p&gt;Identical. Byte-for-byte the same prompt.&lt;/p&gt;

&lt;p&gt;A lot of this came from my cron jobs — I was re-processing the same data every hour, even when nothing had changed since the last run. That's not intelligence. That's insanity.&lt;/p&gt;

&lt;p&gt;I built a simple in-memory cache (and later Redis) keyed by a hash of the prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;redis&lt;/span&gt;

&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Redis&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;localhost&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;port&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;6379&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;decode_responses&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;cached_completion&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;cache_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sha256&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;hexdigest&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;cached&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cache_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;cached&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cached&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_llm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cache_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;ex&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3600&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="mi"&gt;24&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# 24hr TTL
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That cut my request count dramatically. But here's the more interesting part: caching isn't just about deduplication. It's also about being smart about &lt;em&gt;when&lt;/em&gt; to call.&lt;/p&gt;

&lt;p&gt;For tasks that weren't time-sensitive (like batch processing yesterday's support tickets), I queued them up and ran them at off-peak hours. Some providers charge different rates at different times, but more importantly, I could batch multiple tasks into a single call with structured prompting. Batching 5 tasks into one call with a JSON output block cut my per-task cost by about 40%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Context Engineering Beat Prompt Engineering
&lt;/h2&gt;

&lt;p&gt;Here's the mistake that cost me the most money without me knowing it: I was sending massive system prompts and entire conversation histories with every call.&lt;/p&gt;

&lt;p&gt;I had one particular function that kept a 10,000-token conversation history alive so the model could "remember context." That's pure waste. For one call that needed 100 tokens of output, I was paying for 10,000 tokens of input context.&lt;/p&gt;

&lt;p&gt;I switched to what I call "flash context" — strip everything down to just the essential inputs:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;FULL_INSTRUCTIONS_BOOKLET&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;full_conversation_history&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;assistant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;last_response&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;new_request&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;After:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You extract action items from support tickets. Return JSON only.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ticket_text&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="mi"&gt;2000&lt;/span&gt;&lt;span class="p"&gt;]}&lt;/span&gt;  &lt;span class="c1"&gt;# Truncate to relevant portion
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I tested this against my old approach with 50 real tickets. The extraction quality was statistically identical (97% overlap in extracted items). But the token input dropped by 80%.&lt;/p&gt;

&lt;h3&gt;
  
  
  Context compression technique
&lt;/h3&gt;

&lt;p&gt;Here's a trick worth stealing: I built a "context preprocessor" that summarizes long inputs before sending them to the model. If the input document is over 3,000 characters, I run a cheap compression pass first.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;compress_context&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;3000&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;
    &lt;span class="c1"&gt;# Use a small, cheap model to summarize to 500 words
&lt;/span&gt;    &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_small_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Summarize: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I benchmarked this against raw contexts. The results were within 2-3% accuracy for most tasks. But the cost savings were massive — because with GPT-4o you're paying roughly 4x more for input than output per token.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: I Got Rid of My "Always GPT-4o" Default
&lt;/h2&gt;

&lt;p&gt;I'm embarrassed to admit this, but I was just defaulting to the biggest model because I didn't want to think about it. That's lazy engineering.&lt;/p&gt;

&lt;p&gt;When I did the comparison test — same prompts, same temperature, same everything — here's what I found:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;For code generation tasks&lt;/strong&gt;: GPT-4o was measurably better, but GPT-4o-mini was &lt;em&gt;acceptable&lt;/em&gt; for 60% of cases&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For text classification&lt;/strong&gt;: GPT-3.5 Turbo matched GPT-4o 94% of the time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For data extraction&lt;/strong&gt;: All three models were within noise of each other on accuracy&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The gap between flagship and mid-tier models is shrinking. If your application doesn't have a critical failure mode, you can get away with considerably cheaper models with better prompt design.&lt;/p&gt;

&lt;p&gt;At the end of that test weekend, I had a matrix of which model to use for which task type. And I let go of the idea that "big model = better output." Sometimes, it just means "bigger bill."&lt;/p&gt;




&lt;h2&gt;
  
  
  What This Actually Saved Me
&lt;/h2&gt;

&lt;p&gt;Let me break down the real numbers:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model tiering&lt;/td&gt;
&lt;td&gt;$112&lt;/td&gt;
&lt;td&gt;$31&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Token waste (max_tokens)&lt;/td&gt;
&lt;td&gt;$28&lt;/td&gt;
&lt;td&gt;$8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Caching &amp;amp; deduplication&lt;/td&gt;
&lt;td&gt;$45&lt;/td&gt;
&lt;td&gt;$12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context compression&lt;/td&gt;
&lt;td&gt;$32&lt;/td&gt;
&lt;td&gt;$12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$217&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$63&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's a 71% reduction, and I didn't touch a single line of application logic. The output quality stayed the same because every change happened at the infrastructure level, not the prompt level.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Extra 10%: Rate Limits and Batching
&lt;/h2&gt;

&lt;p&gt;Beyond the big wins, there are smaller optimizations that add up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Batch low-priority jobs&lt;/strong&gt; into single calls where possible&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use streaming for interactive responses&lt;/strong&gt; — you're only charged for tokens you actually receive, and streaming lets you cut off early&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Monitor token usage per user&lt;/strong&gt; — I found that 3 heavy users were consuming 40% of my API budget, so I implemented per-user quotas&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I Use Now
&lt;/h2&gt;

&lt;p&gt;I ended up consolidating on a setup that gives me flexible routing across multiple providers without being locked into one vendor's pricing. One thing that made a real difference was using a gateway that lets me switch between model providers based on current costs and availability — because honestly, the "best" model changes month to month now.&lt;/p&gt;

&lt;p&gt;These days I'm running through a unified API endpoint that handles the routing, caching, and fallbacks for me. If you're in the same boat and want to avoid vendor lock-in without building all that infra yourself, there's a service I've been using daily called &lt;strong&gt;tai.shadie-oneapi.com&lt;/strong&gt;. It's a pay-as-you-go aggregator that plugs into multiple LLM providers and handles the cost-tiering in the background. No monthly minimums, no enterprise sales call — you just top up and go.&lt;/p&gt;

&lt;p&gt;Look, I'm not saying it's magic. But if you're tired of watching your AWS bill creep up every month just because you wanted to call a chatbot API, it's worth a look.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Takeaway
&lt;/h2&gt;

&lt;p&gt;The biggest lesson from this entire exercise? I was paying for intelligence I didn't need. The models are powerful, sure. But most of my workload is boring, repetitive, deterministic work that doesn't need the full firepower of a frontier model.&lt;/p&gt;

&lt;p&gt;Cutting costs wasn't about compromising quality. It was about matching the tool to the problem. That's always been good engineering. The AI API ecosystem just made it easy to forget that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your move&lt;/strong&gt;: go check your logs. Look at the actual token usage and prompt lengths from the last week. I bet you'll find the same waste I did — the same 70% sitting there waiting to be reclaimed.&lt;/p&gt;

&lt;p&gt;Happy coding. And may your API bills be small and your test coverage be large.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>I Spent 10x Longer Debugging AI Code Than Writing It — Here's What Changed</title>
      <dc:creator>Shaw Sha</dc:creator>
      <pubDate>Wed, 23 Sep 2026 00:56:19 +0000</pubDate>
      <link>https://dev.to/shadie_ai/i-spent-10x-longer-debugging-ai-code-than-writing-it-heres-what-changed-3g9a</link>
      <guid>https://dev.to/shadie_ai/i-spent-10x-longer-debugging-ai-code-than-writing-it-heres-what-changed-3g9a</guid>
      <description>&lt;p&gt;Everyone talks about AI speeding up coding. Nobody talks about debugging AI-generated code. I learned this the hard way — and it cost me three weekends and a lot of sleep.&lt;/p&gt;

&lt;p&gt;It started innocently enough. I had a feature to build: a real-time dashboard that pulls data from three APIs, merges it, and visualizes it with some custom charts. Nothing crazy. I'd normally budget two days for it. With AI, I figured, I could knock it out in an afternoon.&lt;/p&gt;

&lt;p&gt;I was wrong. Spectacularly wrong.&lt;/p&gt;




&lt;h2&gt;
  
  
  The First Hour: Pure Magic
&lt;/h2&gt;

&lt;p&gt;Let me set the scene. It's a Friday afternoon. I open up my editor, pull up Claude, and start prompting. The first response is beautiful. Clean TypeScript, proper error handling, even comments explaining the tricky parts. I copy-paste it in, run it, and... it works. First try.&lt;/p&gt;

&lt;p&gt;I'm grinning. This is the future. I start chaining prompts: "Now add websocket support," "Make the charts responsive," "Add retry logic for the API calls." Each response is more impressive than the last. By hour two, I have what looks like a complete feature. 800 lines of code, all generated.&lt;/p&gt;

&lt;p&gt;I commit it, push it, and go get coffee feeling like a genius.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Second Hour: The Cracks Appear
&lt;/h2&gt;

&lt;p&gt;The first bug shows up during testing. One of the API endpoints returns data in a slightly different shape than the AI assumed. No problem, I think. I'll just prompt it to fix that specific function.&lt;/p&gt;

&lt;p&gt;But here's the thing I didn't realize yet: every prompt I give the AI to fix something doesn't just modify the code — it sometimes rewrites entire adjacent functions, changes variable names, or introduces a new pattern that clashes with what's already there. The AI doesn't have a mental model of the whole codebase. It's just pattern-matching on my latest prompt.&lt;/p&gt;

&lt;p&gt;So I ask it to fix the API response shape. It does. But now the chart rendering breaks because the AI decided to rename a data transformation function and update half the call sites. The other half it didn't touch.&lt;/p&gt;

&lt;p&gt;Now I'm debugging. Not the original bug — the &lt;em&gt;new&lt;/em&gt; bugs introduced while fixing the first bug.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Debugging Hell Timeline
&lt;/h2&gt;

&lt;p&gt;Here's what the next two weeks looked like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Day 1-2&lt;/strong&gt;: I try prompting my way out. Every fix introduces two new issues. It's whack-a-mole with a jackhammer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Day 3&lt;/strong&gt;: I give up on prompting and manually read through all 800 lines. I find the real problems: the AI generated code that looks correct but has subtle logic errors. Off-by-one errors in loops. Async functions that don't actually await. State updates that happen in the wrong order.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Day 4-5&lt;/strong&gt;: I rewrite about 60% of the code by hand. I keep the parts that work, but I've spent more time understanding the AI's code than I would have writing my own.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Week 2&lt;/strong&gt;: I'm still finding edge cases the AI didn't handle. Timezone issues. Null pointer exceptions from data that was "guaranteed" to exist. Race conditions in the websocket handler.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Total time spent: roughly 60 hours. Total time I would have spent writing it myself: maybe 12-15 hours. I spent &lt;strong&gt;4x longer&lt;/strong&gt;, not 10x — but for some of my colleagues, it's been worse.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why AI Code Is So Hard to Debug
&lt;/h2&gt;

&lt;p&gt;After that experience, I started paying attention to &lt;em&gt;why&lt;/em&gt; AI-generated code breaks in ways that human-written code doesn't. Here's what I've found:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. It's Confidently Wrong
&lt;/h3&gt;

&lt;p&gt;When a human writes code, they usually know where they're uncertain. They'll add a TODO, write a comment, or flag a risk. AI doesn't do that. It generates code with the same confidence for a well-tested pattern and a wild guess at an API you've never heard of.&lt;/p&gt;

&lt;p&gt;I found a function that called &lt;code&gt;fs.readFileSync&lt;/code&gt; with a path that was built from user input. No validation. No error handling. It would crash the server if the file didn't exist. The AI just... didn't think about that case.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. It Has No Memory of Your Codebase
&lt;/h3&gt;

&lt;p&gt;The AI doesn't know that you've established a pattern of using &lt;code&gt;useCallback&lt;/code&gt; for all event handlers, or that you have a utility function for date formatting, or that your team wraps all external API calls in a specific error boundary.&lt;/p&gt;

&lt;p&gt;So it generates code that &lt;em&gt;could&lt;/em&gt; work in isolation but doesn't integrate with your existing patterns. You end up with three different ways to format dates, two different error handling strategies, and a mix of Promises and callbacks in the same file.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. The Invisible Dependencies
&lt;/h3&gt;

&lt;p&gt;The worst bugs are the ones where the AI's code works &lt;em&gt;most&lt;/em&gt; of the time. It handles the happy path perfectly. But there's a hidden dependency — a global state that should be reset, a cache that should be invalidated, a listener that should be removed — that the AI didn't know about.&lt;/p&gt;

&lt;p&gt;I had a function that set up a timer to refresh data. The AI didn't clean it up on component unmount. Memory leak. Subtle. Wouldn't show up in testing until the app had been running for hours.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Actually Changed
&lt;/h2&gt;

&lt;p&gt;After that project, I didn't stop using AI — I just completely changed how I use it. Here's what works:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Use AI for the Parts You Understand
&lt;/h3&gt;

&lt;p&gt;Now I only use AI for code that I could write myself in a few minutes but that's tedious. Boilerplate. Config files. Regex patterns. Basic CRUD operations. Things where I can spot-check the output in 30 seconds.&lt;/p&gt;

&lt;p&gt;I don't use it for anything architecturally complex, or anything where I don't have a clear mental model of what "correct" looks like. If I can't immediately tell if the output is right, I don't use it.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Always Write the Tests First
&lt;/h3&gt;

&lt;p&gt;This is the biggest game-changer. Before I even look at the AI's output, I write tests that define what the code should do. Then I run the AI-generated code against those tests.&lt;/p&gt;

&lt;p&gt;It's not about catching bugs (though it does). It's about having a safety net so that when the AI's fix introduces a new bug, I find out immediately instead of two days later.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Before: I'd ask AI to generate a function and trust it
# After: I define the contract first
&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;unittest&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_merge_user_data&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;The contract: merge two API responses, deduplicate by user_id,
    keep the most recent data for conflicts.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;merge_user_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;api_a&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Alice&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;updated&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2024-01-01&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
        &lt;span class="n"&gt;api_b&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Alice B.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;updated&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2024-01-05&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Alice B.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now when I use AI, I'm not asking it to "write a function." I'm asking it to "implement this contract" and I've got a way to verify it.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Treat AI Like a Junior Developer
&lt;/h3&gt;

&lt;p&gt;I know this sounds patronizing, but it's helped me the most. I don't let a junior dev push directly to main without review. I don't trust their code until I've read it and tested it.&lt;/p&gt;

&lt;p&gt;Same with AI. I read every single line it generates. I check for the edge cases it probably missed. I refactor to fit our patterns. I add the error handling that it skipped.&lt;/p&gt;

&lt;p&gt;The difference is that a junior dev gets better with feedback. AI doesn't — at least not in a way that persists across sessions. So the review burden is on me, every time.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Keep the AI Context Small
&lt;/h3&gt;

&lt;p&gt;This is counterintuitive, but I've found that giving the AI more context makes things worse. If I paste in my entire codebase and ask for a change, it's more likely to produce a mess.&lt;/p&gt;

&lt;p&gt;Instead, I isolate: "Here's this one function. Here's the data shape it receives. Here's what I want it to return." Small, focused prompts produce more predictable output.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Numbers That Convinced Me
&lt;/h2&gt;

&lt;p&gt;I track my time now. Here's what the last three months look like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Before the change&lt;/strong&gt;: ~15 hours of AI coding per week, ~10 hours of debugging AI code&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;After the change&lt;/strong&gt;: ~8 hours of AI coding per week, ~2 hours of debugging AI code&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'm producing less AI code, but I'm shipping &lt;em&gt;more&lt;/em&gt; working features. The debugging time savings more than makes up for the reduction in raw generation.&lt;/p&gt;

&lt;p&gt;My team's story is similar. We did a retrospective where we found that the most productive sprint in the last quarter was the one where we used AI the &lt;em&gt;least&lt;/em&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Consistency Problem Nobody Mentions
&lt;/h2&gt;

&lt;p&gt;There's one more thing that tripped me up: AI model outputs aren't stable. I'd generate a function one day, come back the next, and the updated model would generate something completely different for the same prompt. Different naming conventions. Different structure. Different error handling.&lt;/p&gt;

&lt;p&gt;That's when I started caring about the API layer. If I'm going to rely on AI at all, I need consistent access to the same model, with predictable pricing. Nothing kills a workflow faster than a rate limit hit in the middle of a debugging session.&lt;/p&gt;

&lt;p&gt;That's part of why I've settled on an aggregator for my AI API calls. I use &lt;a href="https://shadie-oneapi.com" rel="noopener noreferrer"&gt;shadie-oneapi.com&lt;/a&gt; because it gives me pay-as-you-go access to multiple models without worrying about quotas or sudden price spikes. The output is consistent because I'm hitting the same endpoints every time, and I don't have to juggle five different API keys.&lt;/p&gt;

&lt;p&gt;I'm not saying you need to use that specific service — but pay attention to how you're accessing AI. The tool matters less than the reliability. If your AI provider changes behavior mid-project, you're going to have a bad time.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;AI code generation is a multiplier, not a replacement. It amplifies whatever you already have — including your debugging skills. If you're a strong developer who can review and refactor code quickly, AI makes you faster. If you're relying on it to write code you couldn't write yourself, you're going to spend a lot of time debugging code you don't understand.&lt;/p&gt;

&lt;p&gt;The shift for me was moving from "let the AI write it and I'll fix it if it breaks" to "let the AI write it, and I'll understand it before it ships."&lt;/p&gt;

&lt;p&gt;That single change cut my debugging time by 80%. And honestly? I sleep better now too.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Why I Stopped Self-Hosting AI Models (And You Probably Should Too)</title>
      <dc:creator>Shaw Sha</dc:creator>
      <pubDate>Tue, 22 Sep 2026 00:55:26 +0000</pubDate>
      <link>https://dev.to/shadie_ai/why-i-stopped-self-hosting-ai-models-and-you-probably-should-too-4oo2</link>
      <guid>https://dev.to/shadie_ai/why-i-stopped-self-hosting-ai-models-and-you-probably-should-too-4oo2</guid>
      <description>&lt;p&gt;I spent three months and roughly $500 of my own money trying to get a self-hosted LLM to work reliably. I had the GPUs, the Docker containers, the whole nine yards. And I ultimately tore it all down in a weekend.&lt;/p&gt;

&lt;p&gt;Let me tell you exactly why.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Allure of the Self-Hosted Setup
&lt;/h2&gt;

&lt;p&gt;It started innocently enough. I was building a tool to summarize internal support tickets. The data isn't top-secret, but it's not something I wanted floating around public API logs either. The open-source community had just dropped some impressive quantized models, and the logic was simple: I own the hardware, I control the data, and I don't pay per token. It sounded like a win-win.&lt;/p&gt;

&lt;p&gt;I scored a used NVIDIA RTX 3090 (24GB VRAM) for a decent price on eBay. My rig already had decent specs, so I figured I was set.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Hardware Shuffle
&lt;/h3&gt;

&lt;p&gt;The first week was all about the hardware. I spent hours fiddling with CUDA versions, driver updates, and the eternal struggle of &lt;code&gt;nvidia-smi&lt;/code&gt; not showing up after a kernel update. It brought back flashbacks to my early Linux days.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# My first test script (obviously) was to just check if the GPU was alive
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cuda&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;is_available&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;  &lt;span class="c1"&gt;# This was True
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cuda&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_device_name&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;  &lt;span class="c1"&gt;# This said "NVIDIA GeForce RTX 3090"
&lt;/span&gt;
&lt;span class="c1"&gt;# Then the actual test
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;transformers&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AutoModelForCausalLM&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;

&lt;span class="n"&gt;model_name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TinyLlama/TinyLlama-1.1B-Chat-v1.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;tokenizer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AutoModelForCausalLM&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;torch_dtype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;float16&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;device_map&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;auto&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;to&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cuda&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Success! It runs. But this is a small model...
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That worked. But the moment I tried something with real reasoning capabilities—a 7B or 13B parameter model—the cracks started to show.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Memory Wall
&lt;/h3&gt;

&lt;p&gt;I realized quickly that VRAM is the currency of the AI world. My 24GB card seemed huge until I looked at the requirements for running a genuinely capable model with a decent context window.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;TinyLlama (1.1B)&lt;/strong&gt;: Fast, but dumb. Good for testing, useless for production.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Llama 2 7B&lt;/strong&gt;: Ran okay with 4-bit quantization, but the quality was mediocre.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mistral 7B&lt;/strong&gt;: Better, but only held a small context window before the GPU exploded.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The second I inched toward a larger context window (like handling a 2,000-word email thread), I hit the memory ceiling. I started researching cloud GPU rentals to run it locally, which defeats the entire purpose of "self-hosting."&lt;/p&gt;

&lt;h2&gt;
  
  
  The Latency Disaster
&lt;/h2&gt;

&lt;p&gt;Once I got a model serving stably, I realized the real problem: speed.&lt;/p&gt;

&lt;p&gt;Inference on a 3090 for a 7B model is &lt;em&gt;okay&lt;/em&gt;. You're looking at maybe 20-40 tokens per second, depending on the quantization. But here’s the thing—when you're using OpenAI or Anthropic APIs, the "time to first byte" is almost instant. The server-side batching is massive.&lt;/p&gt;

&lt;p&gt;When I self-hosted, every single request I made was a cold start or a queue. If two users hit the endpoint simultaneously, the response time went from 2 seconds to 45 seconds. That lag makes the entire application feel broken.&lt;/p&gt;

&lt;p&gt;I spent a week trying to set up vLLM or Text Generation Inference to fix the queuing. It worked, but it consumed even more RAM and required a lot of maintenance. I was becoming a DevOps engineer for a project that was supposed to be a simple utility.&lt;/p&gt;

&lt;h3&gt;
  
  
  The MLOps Slippery Slope
&lt;/h3&gt;

&lt;p&gt;This is where the project really started to unravel for me. I was not just writing code anymore; I was doing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Watching GPU temps&lt;/strong&gt;: I bought a thermal probe and a fan controller so the PC wouldn't sound like a jet engine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dealing with crashes&lt;/strong&gt;: One weekend, the power flickered in my apartment. My model wasn't set up to auto-restart, so the entire system was down for 6 hours until I got home.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Updating dependencies&lt;/strong&gt;: Every time &lt;code&gt;transformers&lt;/code&gt; or &lt;code&gt;torch&lt;/code&gt; released a new version, I had to test it to make sure my serving script didn't break.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I remember calculating the "opportunity cost" during a particularly tedious debugging session. I was paying roughly $0.15/kWh for electricity. Running the GPU at 350W 24/7 was about $38 a month. Add that to the amortized cost of the card, and I was paying about $40-50 a month just to have a worse experience than a $20 API subscription.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Realization: "Cheap" Isn't Always Cheap
&lt;/h2&gt;

&lt;p&gt;Here is the math I finally did. It was sobering.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Hardware&lt;/strong&gt;: $500 for the card (plus a new PSU, because my old one couldn't handle the power draw—another $150).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Time&lt;/strong&gt;: ~2-3 hours a week maintaining it. Across twelve weeks, that's about 30 hours.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Electricity&lt;/strong&gt;: Roughly $120 total over those three months.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Conversely, I tested an API-based solution. It cost me &lt;strong&gt;$1&lt;/strong&gt; in that same period because I was only running a few thousand requests a month. But the real kicker was the speed and reliability. The API never had a power outage. The API never needed a driver update.&lt;/p&gt;

&lt;p&gt;I switched my codebase in one afternoon. It was a moment of ridiculous clarity.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# The exact same "logic" I had on my GPU, but now via an Async client
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;asyncio&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AsyncOpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;AsyncOpenAI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;summarize_ticket&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ticket_text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-4o-mini&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# The cost-per-token is insanely low here
&lt;/span&gt;        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Summarize this support ticket.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ticket_text&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;150&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;

&lt;span class="c1"&gt;# Run it
&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;asyncio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;summarize_ticket&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;My laptop is on fire...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No CUDA checks. No VRAM monitoring. Just... a network call.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Self-Hosting Still Makes Sense
&lt;/h2&gt;

&lt;p&gt;I have to be fair. Self-hosting isn't dead. It's still absolutely essential if:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;You have enterprise-scale security requirements&lt;/strong&gt; where data literally can't leave the VPC.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;You are a genuine ML researcher&lt;/strong&gt; who needs to fine-tune on private datasets.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;You run massive, sustained workloads&lt;/strong&gt; (like processing millions of tokens per day) where the cost curves actually intersect.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;But for the average developer building a feature? For the side project? For the startup MVP where engineering hours are your most precious asset?&lt;/p&gt;

&lt;p&gt;It's a terrible trade-off.&lt;/p&gt;

&lt;h2&gt;
  
  
  The "Hidden" Infrastructure Costs
&lt;/h2&gt;

&lt;p&gt;The thing nobody tells you about self-hosting is the monitoring. When I used an API, I got a dashboard showing me token usage, latency, and error rates. I could see &lt;em&gt;everything&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;When I self-hosted, I had to set up Grafana, export Prometheus metrics, and build my own dashboards just to answer the basic question: "Is it working?"&lt;/p&gt;

&lt;p&gt;One time, my model silently corrupted its weights due to a bad RAM stick in my machine. It didn't crash; it just started producing nonsense outputs. I spent an entire evening testing my code logic before I ran a checksum on the model files. That was the day I realized I was spending more time on infrastructure than application logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Pragmatic Middle Ground
&lt;/h2&gt;

&lt;p&gt;Right now, I run a hybrid approach. I use local models via Ollama for quick, one-off context-free snippets when I'm prototyping ideas. It's great for a brainstorming chat where I don't care about the latency or the data.&lt;/p&gt;

&lt;p&gt;But the moment it hits a production endpoint or a user-facing feature, I route it through an API. I found a solid aggregation service called tai.shadie-oneapi.com that acts as a unified gateway. It lets me buy tokens at rock-bottom prices without maintaining multiple API keys for different providers. It's basically what I wanted my self-hosted setup to be: a single endpoint that just works.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Verdict
&lt;/h2&gt;

&lt;p&gt;Stopping my self-hosting journey was a relief. I sold the 3090 (got $450 for it, so not a total loss) and went back to writing business logic instead of systemd service files.&lt;/p&gt;

&lt;p&gt;The "cool factor" of seeing a GPU at 100% utilization is high. But the maintenance tax is real. I’ve learned that, for 99% of developers, the API route is the viable one. The providers have teams of engineers optimizing kernels and uptime. I want to leverage that expertise, not replicate it.&lt;/p&gt;

&lt;p&gt;Stop fighting your hardware. Spend that time writing code that actually matters to your users. Trust me, your GPU does not need a "project" to feel useful.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>From Curious to Confident: How I Use AI APIs Without Being a Machine Learning Expert</title>
      <dc:creator>Shaw Sha</dc:creator>
      <pubDate>Mon, 21 Sep 2026 00:55:34 +0000</pubDate>
      <link>https://dev.to/shadie_ai/from-curious-to-confident-how-i-use-ai-apis-without-being-a-machine-learning-expert-4bjf</link>
      <guid>https://dev.to/shadie_ai/from-curious-to-confident-how-i-use-ai-apis-without-being-a-machine-learning-expert-4bjf</guid>
      <description>&lt;p&gt;I remember staring at a machine learning research paper three years ago, feeling my eyes glaze over by the second paragraph. The math was dense, the terminology was alien, and I was convinced that building anything with AI required a PhD in neural networks. So I shelved the idea entirely — for about six months, until a side project forced my hand.&lt;/p&gt;

&lt;p&gt;Here's what I learned in the process: you don't need to understand gradient descent, transformer architectures, or any of the heavy theoretical stuff to ship real products with AI. You need the right API key and, in my case, about ten lines of code.&lt;/p&gt;

&lt;h2&gt;
  
  
  The myth of the ML prerequisite
&lt;/h2&gt;

&lt;p&gt;My "aha moment" happened in an unlikely place — a weekend hackathon. I was paired with a guy who worked at a fintech company, and he was building a chatbot to parse banking emails and extract refund information. I asked him how long he'd studied machine learning. He laughed.&lt;/p&gt;

&lt;p&gt;"I'm a frontend developer," he said. "I just call OpenAI's API and parse the JSON."&lt;/p&gt;

&lt;p&gt;That was it. That was the entire secret. He wasn't training models, wasn't tuning hyperparameters, wasn't even writing any ML code. He was just sending HTTP requests with a &lt;code&gt;prompt&lt;/code&gt; field, and handling what came back. The hardest part, he told me, was figuring out which plan to pay for.&lt;/p&gt;

&lt;p&gt;That conversation completely changed my career trajectory. In the months that followed, I built four or five small AI-powered tools without ever once touching a training dataset.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an AI API actually is (for the curious)
&lt;/h2&gt;

&lt;p&gt;Here's the mental model that finally made everything click for me. An AI API is a black box with two holes — one you put text in, one you get text out. The complexity, the training, the billions of parameters — all of that lives entirely inside the box, maintained by engineers who &lt;em&gt;do&lt;/em&gt; have those PhDs.&lt;/p&gt;

&lt;p&gt;My job, as the person building the application, is just to communicate clearly. That's it.&lt;/p&gt;

&lt;p&gt;The whole interaction boils down to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The endpoint&lt;/strong&gt;: the URL you hit&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The headers&lt;/strong&gt;: your authentication key and content type&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The payload&lt;/strong&gt;: your prompt, model choice, and a few parameters like temperature&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The response&lt;/strong&gt;: usually some JSON structure containing the generated text&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Take something I built recently — a simple expense categorizer. I spend way too much on food delivery, but I also mix up work lunches with personal dinners, and my accounting is a disaster. Instead of manually sorting fifty transactions per week, I wrote a script that reads my bank CSV and sends each transaction description to an AI API.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Simple expense categorizer using an AI API&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://tai.shadie-oneapi.com/v1/chat/completions&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Authorization&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Bearer &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;API_KEY&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt; &lt;span class="c1"&gt;// your key goes here&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-3.5-turbo&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;system&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;You are a helpful expense categorization assistant.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Categorize the following transaction into one word: "Chicken &amp;amp; Waffle House". 
        Reply with only the category name, nothing else.`&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="na"&gt;temperature&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;max_tokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;
  &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;category&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Category:&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;category&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="c1"&gt;// Output: Category: Dining&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's seventeen lines of code — including blank space. I didn't train anything. I didn't fine-tune a model. I just wrote instructions in plain English, and the API understood what I needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The skills that actually matter
&lt;/h2&gt;

&lt;p&gt;Once I stopped worrying about the ML math, I realized I already had the skills that mattered for almost every AI use case I encountered:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Prompt engineering is just clear writing.&lt;/strong&gt; The better I articulated what I wanted, the better the results. It's like giving instructions to a smart intern who is &lt;em&gt;literal to a fault&lt;/em&gt;. If you don't say "reply with only the category name," you'll get a paragraph explaining why your waffle purchase is a dining expense — and also a recommendation for their waffle sauce.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;JSON handling is familiar territory.&lt;/strong&gt; Every AI API returns structured data. I've been parsing JSON for years. The new part was just understanding where the response lived in the object (usually &lt;code&gt;choices[0].message.content&lt;/code&gt; for chat-style endpoints).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Error handling became more important.&lt;/strong&gt; I learned that AI APIs return a status code &lt;code&gt;429&lt;/code&gt; when you hit rate limits, and &lt;code&gt;401&lt;/code&gt; when your key is misconfigured. Once I treated these like any other HTTP errors — adding retry logic, exponential backoff, and decent user-facing error messages — my tools stopped breaking randomly.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The actual ML stuff — tokenization, embeddings, attention mechanisms — stayed as background curiosity, not a blocker.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real numbers from my experiments
&lt;/h2&gt;

&lt;p&gt;To give you a sense of scale, let me share some actual figures from my projects:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;My expense categorizer processes about 40 transactions per call batch. At roughly 1,000 tokens per batch, I ran through 30,000 transactions in a month for about &lt;strong&gt;$1.80&lt;/strong&gt; in API costs. That's less than what I waste on late coffee runs.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;A summarization tool I built for a client newsletter — feeding it 5,000-word articles and asking for a 150-word summary — uses about 3,500 tokens per article. At GPT-3.5 pricing, that's roughly &lt;strong&gt;$0.007 per article&lt;/strong&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The point isn't that it's cheap (though it is). The point is that the entry barrier for cost is microscopic. You can build, test, and break things for the price of a vending machine snack, which is a radically different situation from buying GPUs or paying for enterprise ML training runs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common beginner mistakes (I made all of these)
&lt;/h2&gt;

&lt;p&gt;If you're starting out, let me save you some pain. Here's what I got wrong initially:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Over-specifying the prompt.&lt;/strong&gt; I wrote paragraphs of constraints and got worse results than with four clear sentences. The model doesn't reward word count — it rewards precision.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Ignoring temperature.&lt;/strong&gt; I left it at default, which for most chat models is 0.7. My categorization script kept giving varied outputs — sometimes "Dining", sometimes "Restaurant", sometimes "Food". Cranked it down to 0.2, and suddenly everything was consistent. A single parameter fixed my data quality problem.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Using big models for small tasks.&lt;/strong&gt; For a while, I used GPT-4 for everything. Then I tested my scripts with GPT-3.5 and realized the outputs were 95% the same quality — and the cost was about 20x lower.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Not handling long outputs.&lt;/strong&gt; My newsletter summarizer once returned a 3,000-word "summary" because I didn't set &lt;code&gt;max_tokens&lt;/code&gt;. A tiny parameter, a big difference.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Building confidence through iteration
&lt;/h2&gt;

&lt;p&gt;The confidence I now have with AI APIs didn't come from a bootcamp or a certification. It came from shipping.&lt;/p&gt;

&lt;p&gt;My first script took me an evening to write. My second took an hour. The third — which now runs automatically every Sunday morning to generate a grocery list from my weekly meal plan — took about twenty minutes, because I had a working pattern in my codebase and I just copied and adapted it.&lt;/p&gt;

&lt;p&gt;That's the real unlock. Once you wrap your head around the abstract interface of an AI API, every subsequent project feels the same: define your input, write clear instructions, parse the output, and handle the edge cases.&lt;/p&gt;

&lt;p&gt;You don't need to understand what's inside the black box to benefit from it, any more than you need to understand combustion engines to drive a car. And honestly, the more I use these APIs, the more I appreciate the people who built the internals — because their entire job is to make the box so good that I can stay happily ignorant and just build applications on top.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I landed
&lt;/h2&gt;

&lt;p&gt;These days, I have a small toolkit of scripts and utilities that run my life's admin — email drafting, receipt categorization, meeting note summaries. None of them required me to become a machine learning engineer. All of them required me to become a better API consumer.&lt;/p&gt;

&lt;p&gt;If you're curious and you haven't started yet, my recommendation is to pick a boring, repetitive task you hate, and try to automate it with one of these APIs. Start small, budget a couple of dollars, and don't get caught up in model theory. The API will do the heavy lifting.&lt;/p&gt;

&lt;p&gt;And if you're looking for a practical gateway, I've been using one endpoint that aggregates access to multiple AI models under a single key — &lt;strong&gt;tai.shadie-oneapi.com&lt;/strong&gt; — which has saved me from managing separate accounts and billing quirks for each service. I plugged that URL into the same fetch code pattern above, and my whole toolkit works against one consolidated interface. It's a small convenience, but it's exactly the kind of friction that would have stopped me early on.&lt;/p&gt;

&lt;p&gt;Now, when I tell my friends I've been "working with AI," they assume I spent months in a dark room surrounded by textbooks. The reality is that I spent a Saturday writing &lt;code&gt;fetch&lt;/code&gt; requests, and the rest of it was just iterative learning.&lt;/p&gt;

&lt;p&gt;You don't need the PhD. You need curiosity and a willingness to read a few error messages.&lt;/p&gt;

&lt;p&gt;That's not a bar. That's a doorway.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
      <category>tutorial</category>
      <category>javascript</category>
    </item>
  </channel>
</rss>
