<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: LearnAI Resource</title>
    <description>The latest articles on DEV Community by LearnAI Resource (@learnairesource).</description>
    <link>https://dev.to/learnairesource</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3896826%2Fc2131d9a-7159-4fb1-b639-acbdf2d65ecc.png</url>
      <title>DEV Community: LearnAI Resource</title>
      <link>https://dev.to/learnairesource</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/learnairesource"/>
    <language>en</language>
    <item>
      <title>Why Your AI Project Is Bleeding Money (And How to Stop It)</title>
      <dc:creator>LearnAI Resource</dc:creator>
      <pubDate>Mon, 24 Aug 2026 15:00:57 +0000</pubDate>
      <link>https://dev.to/learnairesource/why-your-ai-project-is-bleeding-money-and-how-to-stop-it-4i4o</link>
      <guid>https://dev.to/learnairesource/why-your-ai-project-is-bleeding-money-and-how-to-stop-it-4i4o</guid>
      <description>&lt;h1&gt;
  
  
  Why Your AI Project Is Bleeding Money (And How to Stop It)
&lt;/h1&gt;

&lt;p&gt;You built something cool with Claude, ChatGPT, or Gemini. It works great. Then the bill hits and you're asking whether you accidentally funded a data center.&lt;/p&gt;

&lt;p&gt;Here's the reality: most developers throw money at AI APIs without thinking about it. You're not stupid. You're just not trained to think about token efficiency the way you think about SQL queries.&lt;/p&gt;

&lt;p&gt;Let's fix that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Hidden Cost Structure
&lt;/h2&gt;

&lt;p&gt;Every API call has three sneaky costs:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Input tokens (what you send)&lt;/strong&gt;&lt;br&gt;
You're paying per token. Your 50KB context window? That's like 12,500 tokens. Do that ten times a day, it adds up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Output tokens (what you get back)&lt;/strong&gt;&lt;br&gt;
Usually 2-3x more expensive than input. Asking for detailed responses is literally more expensive than terse ones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Model selection&lt;/strong&gt;&lt;br&gt;
GPT-4 is 60x more expensive than GPT-3.5. Claude 3.5 Sonnet is cheaper than Opus. If you're using the wrong model, you're leaving money on the table.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real Numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Analyzing 100 customer support emails with GPT-4: ~\$15&lt;/li&gt;
&lt;li&gt;Same task with Claude 3.5 Haiku: ~\$0.20&lt;/li&gt;
&lt;li&gt;Using 100KB context per request instead of 20KB: 5x cost increase&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Your project isn't broken. You're just using a Ferrari to go to the grocery store.&lt;/p&gt;

&lt;h2&gt;
  
  
  Optimization Tactics (That Actually Work)
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Batch Your Requests
&lt;/h3&gt;

&lt;p&gt;Don't ask the AI to analyze one thing at a time.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;\&lt;/code&gt;`python&lt;/p&gt;

&lt;h1&gt;
  
  
  ❌ Expensive (10 API calls)
&lt;/h1&gt;

&lt;p&gt;for email in emails:&lt;br&gt;
    response = client.messages.create(&lt;br&gt;
        model="claude-3-5-sonnet-20241022",&lt;br&gt;
        messages=[{"role": "user", "content": f"Categorize: {email}"}]&lt;br&gt;
    )&lt;/p&gt;

&lt;h1&gt;
  
  
  ✅ Cheap (1 API call)
&lt;/h1&gt;

&lt;p&gt;response = client.messages.create(&lt;br&gt;
    model="claude-3-5-sonnet-20241022",&lt;br&gt;
    messages=[{"role": "user", "content": f"""&lt;br&gt;
    Categorize these 10 emails:&lt;br&gt;
    1. {emails[0]}&lt;br&gt;
    2. {emails[1]}&lt;br&gt;
    ...&lt;br&gt;
    10. {emails[9]}&lt;br&gt;
    """}]&lt;br&gt;
)&lt;br&gt;
`&lt;code&gt;\&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;One call with 10 items costs roughly the same as one call with 1 item (you pay for tokens, not requests). But you're making 90% fewer requests.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Right-Size Your Model
&lt;/h3&gt;

&lt;p&gt;Stop using your Cadillac for everything.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude 3.5 Haiku&lt;/strong&gt;: \$0.80/\$2.40 per million tokens. Perfect for routing, classification, simple transformations&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude 3.5 Sonnet&lt;/strong&gt;: \$3/\$15. Sweet spot for most work&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude 3 Opus&lt;/strong&gt;: \$15/\$45. Use when you actually need it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here's the rule: start with Haiku. If it fails, upgrade. Don't start with Opus.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;\&lt;/code&gt;`python&lt;/p&gt;

&lt;h1&gt;
  
  
  Route to the right model based on task complexity
&lt;/h1&gt;

&lt;p&gt;if task == "classification":&lt;br&gt;
    model = "claude-3-5-haiku-20241022"  # \$0.80/1M input&lt;br&gt;
elif task == "code_review":&lt;br&gt;
    model = "claude-3-5-sonnet-20241022"  # \$3/1M input&lt;br&gt;
else:&lt;br&gt;
    model = "claude-3-opus-20240229"  # \$15/1M input&lt;br&gt;
`&lt;code&gt;\&lt;/code&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Compress Your Context
&lt;/h3&gt;

&lt;p&gt;You don't need to send the entire file. Extract what matters.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;\&lt;/code&gt;`python&lt;/p&gt;

&lt;h1&gt;
  
  
  ❌ Send whole document (12KB = 3000 tokens)
&lt;/h1&gt;

&lt;p&gt;prompt = f"Summarize this:\n\n{document}"&lt;/p&gt;

&lt;h1&gt;
  
  
  ✅ Extract key sections first
&lt;/h1&gt;

&lt;p&gt;key_sections = extract_sections(document, ["intro", "methodology", "results"])&lt;br&gt;
prompt = f"Summarize this:\n\n{key_sections}"&lt;br&gt;
`&lt;code&gt;\&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;50% less context = 50% cheaper, often with better results because the signal-to-noise ratio improved.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Cache Frequently Used Prompts
&lt;/h3&gt;

&lt;p&gt;If you're running the same analysis 100 times with different data, use prompt caching.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;\&lt;/code&gt;&lt;code&gt;python&lt;br&gt;
response = client.messages.create(&lt;br&gt;
    model="claude-3-5-sonnet-20241022",&lt;br&gt;
    system=[&lt;br&gt;
        {&lt;br&gt;
            "type": "text",&lt;br&gt;
            "text": "You are a code reviewer. Check for security issues, performance problems, and code style violations.",&lt;br&gt;
            "cache_control": {"type": "ephemeral"}&lt;br&gt;
        }&lt;br&gt;
    ],&lt;br&gt;
    messages=[{"role": "user", "content": new_code}]&lt;br&gt;
)&lt;br&gt;
\&lt;/code&gt;&lt;code&gt;\&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;After the first request, Claude caches that system prompt. Subsequent requests cost 90% less for that context.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Use Local Models for High-Volume Work
&lt;/h3&gt;

&lt;p&gt;For simple tasks (embedding, classification, summarization), Ollama or similar local LLMs cost literally nothing after setup.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;\&lt;/code&gt;&lt;code&gt;bash&lt;br&gt;
ollama pull mistral&lt;br&gt;
\&lt;/code&gt;&lt;code&gt;\&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Now you've got a fast, free classifier for your background jobs. No API costs. No latency. No rate limits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real Example: Customer Support Triage
&lt;/h2&gt;

&lt;p&gt;Let's say you process 500 support tickets daily.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before optimization:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPT-4: 500 × 3000 tokens × (\$0.03/1K input) = \$45/day&lt;/li&gt;
&lt;li&gt;Costs: \$1,350/month&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;After optimization:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Route with Haiku: 500 × 300 tokens × (\$0.0008/1K input) = \$0.12/day&lt;/li&gt;
&lt;li&gt;Complex ones to Sonnet: 100 × 1500 tokens × (\$0.003/1K input) = \$0.45/day&lt;/li&gt;
&lt;li&gt;Costs: \$17/month&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same functionality. 98% cheaper.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Question
&lt;/h2&gt;

&lt;p&gt;Not "can we use AI for this?" but "what's the cheapest way to do this well?"&lt;/p&gt;

&lt;p&gt;Start small. Measure. Optimize. Most teams could cut costs 70-80% by moving to the right model and batching requests.&lt;/p&gt;

&lt;p&gt;Your project isn't too expensive because AI is expensive. It's expensive because you haven't thought about efficiency yet.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Want practical guidance on building with AI without breaking the bank?&lt;/strong&gt; Check out &lt;a href="https://learnairesource.com/newsletter" rel="noopener noreferrer"&gt;LearnAI Weekly&lt;/a&gt; — fresh insights on AI tools, cost optimization, and production workflows every week.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>optimization</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Stop Waiting for Code Reviews: Use AI as Your First Line of Defense</title>
      <dc:creator>LearnAI Resource</dc:creator>
      <pubDate>Sun, 23 Aug 2026 15:00:43 +0000</pubDate>
      <link>https://dev.to/learnairesource/stop-waiting-for-code-reviews-use-ai-as-your-first-line-of-defense-b84</link>
      <guid>https://dev.to/learnairesource/stop-waiting-for-code-reviews-use-ai-as-your-first-line-of-defense-b84</guid>
      <description>&lt;h1&gt;
  
  
  Stop Waiting for Code Reviews: Use AI as Your First Line of Defense
&lt;/h1&gt;

&lt;p&gt;We've all been there—you push a PR, grab coffee, and wait hours for someone to review 400 lines of code. Or worse, the feedback comes back and it's things a linter should've caught.&lt;/p&gt;

&lt;p&gt;I started using Claude and ChatGPT as a pre-review layer, and it's genuinely changed how fast I move. Here's what actually works.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Setup (2 Minutes)
&lt;/h2&gt;

&lt;p&gt;Most editors support Claude through extensions now. VS Code has the official Anthropic extension. Drop your API key in, select code, and ask.&lt;/p&gt;

&lt;p&gt;If you're command-line focused, curl works fine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.anthropic.com/v1/messages &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"x-api-key: &lt;/span&gt;&lt;span class="nv"&gt;$ANTHROPIC_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"content-type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; @- &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt; &lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt;
{
  "model": "claude-3-5-sonnet-20241022",
  "max_tokens": 2048,
  "messages": [
    {
      "role": "user",
      "content": "Review this code for logic errors, performance issues, and readability: [your code]"
    }
  ]
}
&lt;/span&gt;&lt;span class="no"&gt;EOF
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Dead simple. Runs locally in your editor or terminal.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Works
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Logic errors:&lt;/strong&gt; AI catches the obvious ones you missed at 11 PM. A few weeks back, I had a loop that would've deleted user records instead of archiving them. Claude spotted it immediately.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Performance:&lt;/strong&gt; Ask specifically about O(n²) operations, unnecessary loops, or API calls in loops. It's especially useful if you're not a performance expert in that language.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Security issues:&lt;/strong&gt; SQL injection patterns, hardcoded secrets, auth bypass logic. Not perfect, but way better than nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Style consistency:&lt;/strong&gt; Does your code match your team's style guide? Ask it to check. Saves your actual reviewer from nitpicking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real Example
&lt;/h2&gt;

&lt;p&gt;I had a Node function that was checking permissions wrong:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;deletePost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;postId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;post&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;Post&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;postId&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;post&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Not authorized&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;Post&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;deleteOne&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;postId&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Seems fine, right? I asked Claude: "Does this have any security issues?" It flagged a race condition—what if the post gets deleted between the check and the delete? What if the user object gets updated? It suggested:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;deletePost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;postId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;Post&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;deleteOne&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; 
    &lt;span class="na"&gt;_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;postId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;userId&lt;/span&gt;  &lt;span class="c1"&gt;// Permission check in the query itself&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;deletedCount&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Post not found or not authorized&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Much better. I wouldn't have thought of that without a prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Limitations (Be Real)
&lt;/h2&gt;

&lt;p&gt;AI code review is NOT a replacement for human review. It'll miss:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Architectural decisions that don't make sense for your system&lt;/li&gt;
&lt;li&gt;Dead code that should stay for backward compatibility&lt;/li&gt;
&lt;li&gt;Business logic that contradicts requirements&lt;/li&gt;
&lt;li&gt;Context from other parts of your codebase&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use it as a filter. Let it catch the dumb stuff (off-by-one errors, missing error handling, typos in variable names). Then have humans review the important bits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pro Tips
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Ask specific questions.&lt;/strong&gt; Don't say "review this." Say "check if this handles null values correctly" or "look for database performance issues."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Feed it context.&lt;/strong&gt; Paste the relevant parts of your codebase that the code interacts with. More context = better feedback.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Iterate.&lt;/strong&gt; If it suggests something you don't understand, ask it to explain. If you disagree, ask why. You're using it as a thinking partner, not gospel.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Set it up in CI.&lt;/strong&gt; Some teams run AI review as part of their CI pipeline. It doesn't block PRs, but it leaves comments automatically. Nice friction-free layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Win
&lt;/h2&gt;

&lt;p&gt;I'm not faster because the AI is smarter. I'm faster because I catch my own mistakes before they go to review, and reviewers spend time on what actually matters—does this fit our architecture? Is the approach right?&lt;/p&gt;

&lt;p&gt;That's worth an API key.&lt;/p&gt;

&lt;p&gt;Check out more practical AI workflows in the &lt;a href="https://learnairesource.com/newsletter" rel="noopener noreferrer"&gt;LearnAI Weekly newsletter&lt;/a&gt; if you want to stay on top of tools that actually move the needle.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>coding</category>
      <category>productivity</category>
      <category>webdev</category>
    </item>
    <item>
      <title>5 Claude API Patterns That Actually Ship (No Prompt Engineering Rabbit Holes)</title>
      <dc:creator>LearnAI Resource</dc:creator>
      <pubDate>Sat, 22 Aug 2026 15:00:39 +0000</pubDate>
      <link>https://dev.to/learnairesource/5-claude-api-patterns-that-actually-ship-no-prompt-engineering-rabbit-holes-3nid</link>
      <guid>https://dev.to/learnairesource/5-claude-api-patterns-that-actually-ship-no-prompt-engineering-rabbit-holes-3nid</guid>
      <description>&lt;p&gt;So you've got a Claude API key and a vague idea that you want to use it for something. Maybe a chatbot. Maybe parsing. Maybe automation. The problem is most tutorials are either "hello world" garbage or they dive into prompt engineering theory that doesn't help you ship anything.&lt;/p&gt;

&lt;p&gt;Here's what actually works. Five patterns I've seen developers successfully throw into production without losing their minds.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Structured Output for Data Extraction
&lt;/h2&gt;

&lt;p&gt;Stop writing regex. Stop parsing XML. Use the API's built-in structured output.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-3-5-sonnet-20241022&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;thinking&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;enabled&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;budget_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5000&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Extract invoice data from this text:
Invoice #2024-1001
Customer: ACME Corp
Amount: $5,234.50
Date: August 20, 2024&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;system&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Return JSON with fields: invoice_id, customer, amount, date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Actually get structured JSON back, not a rambling paragraph
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The thinking parameter above gives Claude time to reason through the task. It costs more but it works. Your success rate on messy data jumps from 70% to 95%.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Batch Processing for Cost Optimization
&lt;/h2&gt;

&lt;p&gt;If you're processing thousands of items, batching is your friend. Real savings.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# Create batch requests
&lt;/span&gt;&lt;span class="n"&gt;requests&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;your_dataset&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;custom_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;request-&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;params&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-3-5-sonnet-20241022&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;messages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
                &lt;span class="p"&gt;{&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Summarize this: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="p"&gt;}&lt;/span&gt;
            &lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;

&lt;span class="c1"&gt;# Submit batch
&lt;/span&gt;&lt;span class="n"&gt;batch&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;batch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Process results later (much cheaper than real-time)
&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;batch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retrieve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;batch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Batches run asynchronously and cost 50% less. Perfect for overnight jobs.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Conversation Memory Without Storing Everything
&lt;/h2&gt;

&lt;p&gt;Cache your system prompt and frequently-referenced context to save tokens and money.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;conversation_history&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;chat_with_context&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context_docs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;conversation_history&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;user_message&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;

    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-3-5-sonnet-20241022&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;system&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are a helpful assistant focused on this knowledge base.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Context: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;context_docs&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cache_control&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ephemeral&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;conversation_history&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;conversation_history&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;assistant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The cache_control flag tells Claude to reuse that context. If you're running the same query multiple times against the same knowledge base, you're cutting costs by 30-40%.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Vision for Real-world Automation
&lt;/h2&gt;

&lt;p&gt;Screenshot parsing, document analysis, visual QA — Claude's vision is genuinely useful.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;analyze_screenshot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image_path&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image_path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;img_file&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;image_data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;standard_b64encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img_file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-3-5-sonnet-20241022&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
                    &lt;span class="p"&gt;{&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;base64&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;media_type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image/png&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;image_data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                        &lt;span class="p"&gt;},&lt;/span&gt;
                    &lt;span class="p"&gt;},&lt;/span&gt;
                    &lt;span class="p"&gt;{&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Extract all form field labels and values from this screenshot. Return as JSON.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="p"&gt;}&lt;/span&gt;
                &lt;span class="p"&gt;],&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use this for automated testing, document processing, UI feedback — anything that needs visual understanding without training a model.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Streaming for Real-time UX
&lt;/h2&gt;

&lt;p&gt;If your users are waiting for a response, stream it. Feels faster, actually &lt;em&gt;is&lt;/em&gt; faster.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;stream_response&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_input&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-3-5-sonnet-20241022&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;user_input&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text_stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;flush&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Browsers, terminals, chat apps — all start showing the response before Claude finishes thinking. Your app feels snappier even on high-latency connections.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Token counting before you send.&lt;/strong&gt; Use &lt;code&gt;client.beta.messages.count_tokens()&lt;/code&gt; to check your request size first. Saves embarrassing overages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rate limits sneak up.&lt;/strong&gt; Implement exponential backoff. Don't hammer the API when you hit a 429.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model updates happen.&lt;/strong&gt; Your "claude-3-5-sonnet-20241022" will eventually be old. Pin specific versions in production, test new models in staging.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context length is not infinite.&lt;/strong&gt; 200k tokens sounds huge until you're parsing a 50-page PDF. Chunk your data strategically.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Isn't
&lt;/h2&gt;

&lt;p&gt;This isn't a guide to prompt engineering tricks or system prompt hacks. That stuff changes weekly and honestly most of it doesn't matter. These patterns work because they're about &lt;em&gt;how you use the API&lt;/em&gt;, not what words you put in the prompt.&lt;/p&gt;

&lt;p&gt;Want more practical patterns like this? Check out &lt;a href="https://learnairesource.com/newsletter" rel="noopener noreferrer"&gt;LearnAI Weekly&lt;/a&gt; — actual code, actual results, no fluff.&lt;/p&gt;

&lt;p&gt;Ship it. 🚀&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Stop Debug-Scrolling: Use AI to Find Bugs Faster</title>
      <dc:creator>LearnAI Resource</dc:creator>
      <pubDate>Fri, 21 Aug 2026 15:00:37 +0000</pubDate>
      <link>https://dev.to/learnairesource/stop-debug-scrolling-use-ai-to-find-bugs-faster-hek</link>
      <guid>https://dev.to/learnairesource/stop-debug-scrolling-use-ai-to-find-bugs-faster-hek</guid>
      <description>&lt;p&gt;You know that moment? 3 AM, you're staring at the same 200 lines of code for the third time, and your brain is just... gone. You're looking &lt;em&gt;at&lt;/em&gt; the bug but not &lt;em&gt;seeing&lt;/em&gt; it.&lt;/p&gt;

&lt;p&gt;There's a better way. And it doesn't involve more coffee.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem With Manual Debugging
&lt;/h2&gt;

&lt;p&gt;When you're deep in code, your brain gets tunnel vision. You see what you &lt;em&gt;expect&lt;/em&gt; to see, not what's actually there. That's why fresh eyes catch bugs instantly—they don't have your assumptions.&lt;/p&gt;

&lt;p&gt;AI doesn't have assumptions. It's great at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Spotting logic errors in conditionals&lt;/li&gt;
&lt;li&gt;Finding off-by-one mistakes&lt;/li&gt;
&lt;li&gt;Catching typos in variable names&lt;/li&gt;
&lt;li&gt;Detecting unhandled edge cases&lt;/li&gt;
&lt;li&gt;Pointing out unused variables eating memory&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Workflow That Actually Saves Time
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Step 1: Isolate the problem&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Don't paste your entire codebase. Grab the function that's broken plus one level above it. Usually 50-100 lines max.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// ✅ Good: specific context&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;calculateDiscount&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;price&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;userType&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;baseRate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;userType&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;premium&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;price&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;baseRate&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;0.05&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;price&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;baseRate&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;finalPrice&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;calculateDiscount&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;premium&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;finalPrice&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// Expected 105, got 100&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 2: Be specific about what's wrong&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;"This is broken" gets useless answers. "I expect 105 but got 100" gives AI something to actually investigate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 3: Ask for the mechanism, not just the fix&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Bad: "Fix this."&lt;br&gt;
Good: "Walk me through the logic—what's happening with the multiplication?"&lt;/p&gt;

&lt;p&gt;This usually makes the bug obvious &lt;em&gt;to you&lt;/em&gt; before AI even answers. Your rubber duck just got smarter.&lt;/p&gt;
&lt;h2&gt;
  
  
  Tools Worth Your Time
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Claude/ChatGPT:&lt;/strong&gt; Paste your code snippet, describe the symptom. Fast and accurate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub Copilot in debug mode:&lt;/strong&gt; Works right in your editor. Point it at the problem line.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cursor Editor:&lt;/strong&gt; AI-native IDE that understands your codebase context better than tabs ever could.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Devin (AI engineer):&lt;/strong&gt; For complex multi-file bugs. Overkill for simple stuff, perfect when you're stuck.&lt;/p&gt;
&lt;h2&gt;
  
  
  Real Example (From Yesterday)
&lt;/h2&gt;

&lt;p&gt;I had a React component that was re-rendering every second for no reason. Looks normal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;setInterval&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;setCounter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;counter&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Spent 10 minutes looking at this. Pasted it into Claude with "counter isn't updating but component re-renders constantly."&lt;/p&gt;

&lt;p&gt;Two seconds later: &lt;em&gt;"setInterval is referencing &lt;code&gt;counter&lt;/code&gt; from closure but it's not in dependencies. You need &lt;code&gt;[counter]&lt;/code&gt; or use a reducer."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Dumb mistake. Obviously dumb in retrospect. But my brain was committed to looking at the timer logic, not the hooks.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Speedup Is Real
&lt;/h2&gt;

&lt;p&gt;Debug session used to look like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Read code (5 min)&lt;/li&gt;
&lt;li&gt;Add console.logs (3 min)&lt;/li&gt;
&lt;li&gt;Run, interpret output (5 min)&lt;/li&gt;
&lt;li&gt;Modify and try again (10 min)&lt;/li&gt;
&lt;li&gt;Check Stack Overflow (15 min)&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Finally&lt;/em&gt; find the issue (20 min)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With AI:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Describe the problem (2 min)&lt;/li&gt;
&lt;li&gt;Paste relevant snippet (1 min)&lt;/li&gt;
&lt;li&gt;Get explanation (1 min)&lt;/li&gt;
&lt;li&gt;Verify and fix (2 min)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's the difference between a 1-hour rabbit hole and 6 minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  One Warning
&lt;/h2&gt;

&lt;p&gt;AI can confidently give you wrong answers. Use it as a thinking partner, not gospel. If the explanation doesn't make sense, ask it to show the code path step-by-step. Make it prove it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Meta-Move
&lt;/h2&gt;

&lt;p&gt;The actual win isn't using AI to debug—it's training yourself to describe problems clearly. That skill makes you better at debugging &lt;em&gt;without&lt;/em&gt; AI too. You write better error messages. You isolate issues faster. Your code reviews get sharper.&lt;/p&gt;

&lt;p&gt;So next time you're stuck, don't just paste. Think first. &lt;em&gt;Then&lt;/em&gt; let AI help you see what you're missing.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Want to stay sharp on tools like this?&lt;/strong&gt; Check out &lt;a href="https://learnairesource.com/newsletter" rel="noopener noreferrer"&gt;LearnAI Weekly newsletter&lt;/a&gt; for practical AI tips developers actually use, every week.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>debugging</category>
      <category>productivity</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Stop Hemorrhaging Money on AI API Calls: A Survival Guide</title>
      <dc:creator>LearnAI Resource</dc:creator>
      <pubDate>Thu, 20 Aug 2026 15:00:37 +0000</pubDate>
      <link>https://dev.to/learnairesource/stop-hemorrhaging-money-on-ai-api-calls-a-survival-guide-3ojg</link>
      <guid>https://dev.to/learnairesource/stop-hemorrhaging-money-on-ai-api-calls-a-survival-guide-3ojg</guid>
      <description>&lt;p&gt;If you've started integrating Claude, GPT, or other LLMs into your apps, you've probably had that moment. You check your billing. Your jaw drops. A single feature test ran through $200 worth of tokens.&lt;/p&gt;

&lt;p&gt;Yeah. Welcome to the fun part of building with AI.&lt;/p&gt;

&lt;p&gt;The good news? You can drastically cut costs without sacrificing quality. I've helped teams drop their API spend by 70% without changing what users see. Here's how.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Batch Your Requests (This Alone Cuts 50%)
&lt;/h2&gt;

&lt;p&gt;Most people fire off API calls one at a time. Batch processing lets you send hundreds of requests together and get a discount.&lt;/p&gt;

&lt;p&gt;Claude's Batch API? 50% cheaper. OpenAI's? Similar deal.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# ❌ One at a time - expensive
&lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;emails&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-3-5-sonnet-20241022&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Classify: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# ✅ Batch - 50% discount
&lt;/span&gt;&lt;span class="n"&gt;requests&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;custom_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email-&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;params&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-3-5-sonnet-20241022&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;messages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Classify: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;emails&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;batch&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;batches&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If batching takes 24 hours to process, that's fine — batch non-urgent work overnight.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Use Cheaper Models for Simple Tasks
&lt;/h2&gt;

&lt;p&gt;You don't need GPT-4 to classify spam or extract structured data. Smaller models are way faster and cheaper.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GPT-4o&lt;/strong&gt;: Complex reasoning, creative work, novel problems&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude 3.5 Haiku&lt;/strong&gt;: Fast summaries, classification, simple extraction&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Llama 3.1 (self-hosted or via API)&lt;/strong&gt;: Costs nearly nothing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Real example: A content moderation pipeline was costing $8k/month using GPT-4 for every comment. Switching classifier tasks to Haiku? $400/month. Same accuracy.&lt;/p&gt;

&lt;p&gt;Test with Haiku first. Upgrade to Sonnet only if needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Cache Your Prompts (Free Tokens After First Call)
&lt;/h2&gt;

&lt;p&gt;System prompts and large context documents get reused. Cache them.&lt;/p&gt;

&lt;p&gt;With Claude:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-3-5-sonnet-20241022&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;system&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;HUGE_SYSTEM_PROMPT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cache_control&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ephemeral&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;user_query&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After the first call, you pay full price. Every call after that? The cached system prompt is free.&lt;/p&gt;

&lt;p&gt;If your system prompt is 50k tokens and you make 100 calls per day:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Without cache: 50k × 100 = 5M tokens/day&lt;/li&gt;
&lt;li&gt;With cache: 50k × 1 + (1k × 99) ≈ 149k tokens/day&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The math compounds.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Stream Responses (Faster Perception, Lower Latency Costs)
&lt;/h2&gt;

&lt;p&gt;Streaming doesn't technically save tokens, but it &lt;em&gt;feels&lt;/em&gt; faster to users, and it limits unnecessary processing.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-3-5-sonnet-20241022&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Write a poem&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text_stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;flush&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Users see output immediately instead of waiting for the full response. Perception of speed = better UX, and you're not computing tokens you don't need.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Use Structured Output (Fewer Tokens, Cleaner Results)
&lt;/h2&gt;

&lt;p&gt;JSON mode saves tokens because the model doesn't generate filler text or natural language waffling.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"The user's sentiment is positive because they used exclamation marks and positive adjectives"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You get:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"sentiment"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"positive"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.92&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With Claude:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-3-5-sonnet-20241022&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[...],&lt;/span&gt;
    &lt;span class="n"&gt;thinking&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;enabled&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;budget_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5000&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;  &lt;span class="c1"&gt;# Even thinking is optional
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Structured output = fewer token waste on explanation.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Track and Alert on Spending
&lt;/h2&gt;

&lt;p&gt;You can't optimize what you don't measure.&lt;/p&gt;

&lt;p&gt;Set up spending alerts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI&lt;/strong&gt;: Organization settings → Billing → Usage limits&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude&lt;/strong&gt;: Check usage dashboard daily&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DIY&lt;/strong&gt;: Log every API call with token count, add to a Sheets&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One team I know tags every API call with a project name:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;headers&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;anthropic-api-key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;API_KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x-user-id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;project-name-123&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;  &lt;span class="c1"&gt;# Track per project
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now you know which feature is burning cash.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Win
&lt;/h2&gt;

&lt;p&gt;The teams cutting costs aren't doing anything magical. They're just:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Batching non-urgent work&lt;/li&gt;
&lt;li&gt;Using the right model for the job&lt;/li&gt;
&lt;li&gt;Caching repetitive context&lt;/li&gt;
&lt;li&gt;Monitoring spend like it matters&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Start with #1 and #2. Implement caching next. You'll cut costs by 60-70% without touching product.&lt;/p&gt;




&lt;p&gt;Want more practical AI workflows and cost optimization patterns? Check out &lt;a href="https://learnairesource.com/newsletter" rel="noopener noreferrer"&gt;LearnAI Weekly newsletter&lt;/a&gt; — real strategies, no AI hype.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>developer</category>
      <category>costs</category>
    </item>
    <item>
      <title>Local LLMs vs Cloud APIs: Which One Should You Actually Use?</title>
      <dc:creator>LearnAI Resource</dc:creator>
      <pubDate>Wed, 19 Aug 2026 15:00:34 +0000</pubDate>
      <link>https://dev.to/learnairesource/local-llms-vs-cloud-apis-which-one-should-you-actually-use-3gd8</link>
      <guid>https://dev.to/learnairesource/local-llms-vs-cloud-apis-which-one-should-you-actually-use-3gd8</guid>
      <description>&lt;h1&gt;
  
  
  Local LLMs vs Cloud APIs: Which One Should You Actually Use?
&lt;/h1&gt;

&lt;p&gt;Been staring at your cloud API bill and wondering if you're throwing money at the wrong solution? Yeah, I was there too. Let me break down the real tradeoffs so you can stop second-guessing yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Setup
&lt;/h2&gt;

&lt;p&gt;You've got two paths:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cloud APIs&lt;/strong&gt; (OpenAI, Claude, Gemini) — send your data out, get answers back&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local LLMs&lt;/strong&gt; (Ollama, llama.cpp, vLLM) — run models on your machine, keep everything private&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both have their place. Neither is the obvious winner. Let me show you how to pick.&lt;/p&gt;

&lt;h2&gt;
  
  
  Speed: Local Wins (Mostly)
&lt;/h2&gt;

&lt;p&gt;Running Mistral 7B locally on decent hardware? You're looking at response times under 100ms for most tasks. Cloud APIs average 500ms-2s depending on their load.&lt;/p&gt;

&lt;p&gt;But here's the catch: cloud APIs have access to models that'll run circles around your local stuff. GPT-4 beats Mistral on complex reasoning. You're trading latency for capability.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real scenario:&lt;/strong&gt; You're building a chat feature for your app. Sub-100ms response time is &lt;em&gt;nice&lt;/em&gt; but your users won't notice the difference between 200ms and 2 seconds. But they'll absolutely notice if your answers are garbage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost: It Depends (Obviously)
&lt;/h2&gt;

&lt;p&gt;Local math is simple:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your hardware: one-time cost&lt;/li&gt;
&lt;li&gt;Electricity: pennies per inference&lt;/li&gt;
&lt;li&gt;Maintenance: your time (not free)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cloud math:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;$0.01-$0.30 per 1K tokens depending on model&lt;/li&gt;
&lt;li&gt;Scales automatically (good and bad)&lt;/li&gt;
&lt;li&gt;Someone else handles the headaches&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Quick napkin math:&lt;/strong&gt; If you're doing 1 million API calls a month, cloud costs you $100-300. Running local costs you maybe $50 in electricity on decent gear, but you're babysitting the server yourself.&lt;/p&gt;

&lt;p&gt;If you're doing 10 million calls? Now we're talking thousands per month on cloud. Local suddenly makes sense.&lt;/p&gt;

&lt;h2&gt;
  
  
  Privacy: Local is Actually Better (But Not Perfect)
&lt;/h2&gt;

&lt;p&gt;Local model means your data doesn't leave your machine. That's genuinely valuable if you're processing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Customer data&lt;/li&gt;
&lt;li&gt;Source code you don't want GitHub knowing about&lt;/li&gt;
&lt;li&gt;Medical/financial records&lt;/li&gt;
&lt;li&gt;Anything regulated&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cloud APIs? Read the terms. OpenAI keeps data for 30 days by default for "safety and abuse" monitoring. Claude is better about this, but it's still stored on their infrastructure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reality check:&lt;/strong&gt; If you're serious about privacy, local is your move. But remember — the LLM was trained on the internet. You're not getting clean room isolation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quality: Cloud APIs are Still Ahead
&lt;/h2&gt;

&lt;p&gt;Here's the honest part: Claude 3.5 and GPT-4 are just smarter. They handle edge cases better, they explain things clearer, they make fewer hallucinations.&lt;/p&gt;

&lt;p&gt;Mistral 7B and Llama 2 13B are... fine. They're good for summarization, categorization, simple code review. They struggle with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Multi-step reasoning&lt;/li&gt;
&lt;li&gt;Complex writing tasks&lt;/li&gt;
&lt;li&gt;Novel problems they haven't seen much training on&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Where local shines:&lt;/strong&gt; Boring but reliable tasks. Classification, extraction, formatting. Your local model will be consistent and fast. It won't be brilliant, but it'll be reliable.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Decision Framework
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Use cloud APIs if:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You need quality first (complex reasoning, writing, analysis)&lt;/li&gt;
&lt;li&gt;You're okay with 1-2 second latency&lt;/li&gt;
&lt;li&gt;Your data isn't sensitive&lt;/li&gt;
&lt;li&gt;You want to keep your ops team focused on shipping code instead of model management&lt;/li&gt;
&lt;li&gt;You're willing to pay for convenience&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Use local if:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Privacy is non-negotiable&lt;/li&gt;
&lt;li&gt;You're running this at scale (millions of inferences)&lt;/li&gt;
&lt;li&gt;Your task is simple and repetitive&lt;/li&gt;
&lt;li&gt;You like having full control&lt;/li&gt;
&lt;li&gt;You have someone (maybe you) willing to debug model stuff at 2am&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Use both if:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You can afford the complexity&lt;/li&gt;
&lt;li&gt;You've got sensitive AND complex tasks&lt;/li&gt;
&lt;li&gt;You're building for different markets with different requirements&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Seriously, this is totally legit. Use Claude for your magical features, use Mistral locally for classification and filtering. Distribute the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Actually Do
&lt;/h2&gt;

&lt;p&gt;I run Mistral and Llama locally for filtering and categorization. Fast, cheap, keeps data in-house. For anything requiring real reasoning? Cloud APIs, no question. I'm paying for quality and letting someone else handle the infrastructure.&lt;/p&gt;

&lt;p&gt;It's boring compared to the "local is the future" narrative everyone pushes, but it actually works.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;Local LLMs aren't replacing cloud APIs in 2026. Cloud APIs aren't doing all your work for $200/month either. They're tools with different tradeoffs. Pick based on what actually matters for your use case, not what sounds cooler.&lt;/p&gt;

&lt;p&gt;Want to level up your AI workflow faster? Check out &lt;strong&gt;&lt;a href="https://learnairesource.com/newsletter" rel="noopener noreferrer"&gt;LearnAI Weekly&lt;/a&gt;&lt;/strong&gt; — real examples, practical tools, no hype. Worth the subscription.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>developer</category>
      <category>tools</category>
    </item>
    <item>
      <title>Stop Letting Bad Code Slip Through: Build Your AI Code Review Workflow</title>
      <dc:creator>LearnAI Resource</dc:creator>
      <pubDate>Tue, 18 Aug 2026 15:00:43 +0000</pubDate>
      <link>https://dev.to/learnairesource/stop-letting-bad-code-slip-through-build-your-ai-code-review-workflow-1bce</link>
      <guid>https://dev.to/learnairesource/stop-letting-bad-code-slip-through-build-your-ai-code-review-workflow-1bce</guid>
      <description>&lt;h1&gt;
  
  
  Stop Letting Bad Code Slip Through: Build Your AI Code Review Workflow
&lt;/h1&gt;

&lt;p&gt;You know that feeling? You're reviewing a PR, and something feels off, but you can't quite articulate why. It's 2026, and we've got tools that can do this automatically. Let me show you how to actually use them without turning your workflow into chaos.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Problem
&lt;/h2&gt;

&lt;p&gt;Code review is where most teams fail. It's not because people are bad reviewers—it's because there's too much surface area. A human can miss edge cases, performance issues, or security holes in a 400-line diff. An AI can catch the obvious stuff in seconds. Then you focus on the architecture, UX impact, and whether this actually solves the problem.&lt;/p&gt;

&lt;p&gt;That's the win here: not replacing your brain, but giving it less noise to filter through.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Changed (And Why It Matters)
&lt;/h2&gt;

&lt;p&gt;AI code review isn't new, but it got &lt;em&gt;good&lt;/em&gt; around 2024-2025. Here's what shifted:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before:&lt;/strong&gt; Tools gave you generic warnings. "Don't use &lt;code&gt;var&lt;/code&gt;" levels of useful.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Now:&lt;/strong&gt; They understand context. They know your codebase patterns, your team's style, your stack. They catch real bugs: off-by-one errors, uncaught exceptions, missing null checks.&lt;/p&gt;

&lt;p&gt;The key difference? Modern LLMs have enough context window to see the whole file, not just the diff.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Workflow That Actually Works
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. &lt;strong&gt;GitHub/GitLab Integration&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Use a bot that runs on every PR. Popular options: CodeRabbit, Sweep, or Sourcegraph's Cody.&lt;/p&gt;

&lt;p&gt;Why not DIY? Because you'll spend 20 hours building what already exists. Use it, customize it later if needed.&lt;/p&gt;

&lt;p&gt;Set it up to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Comment on individual lines with specific issues&lt;/li&gt;
&lt;li&gt;Post a summary comment with critical findings&lt;/li&gt;
&lt;li&gt;Use labels (performance, security, refactor) for grouping&lt;/li&gt;
&lt;li&gt;Skip obvious stuff (formatting, which your linter should catch anyway)&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. &lt;strong&gt;Local Pre-Review&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Before you push, run a quick check locally. This saves your team from even seeing junior mistakes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Use Copilot, Claude, or any LLM CLI&lt;/span&gt;
&lt;span class="c"&gt;# Example with a local setup:&lt;/span&gt;
git diff | ai &lt;span class="s2"&gt;"Review this diff for bugs, security issues, performance problems. Be specific."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You get feedback in seconds. Fix it. Push clean code.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. &lt;strong&gt;The Human Layer&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;This is where you add value. After the AI flags things, &lt;em&gt;you&lt;/em&gt; decide:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is this a real issue or a false positive?&lt;/li&gt;
&lt;li&gt;Does it align with our team's direction?&lt;/li&gt;
&lt;li&gt;Is there architectural debt worth accepting here?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The AI did the grunt work. You do the thinking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real Examples
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Catch #1: SQL Injection You Missed
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// AI catches this immediately&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`SELECT * FROM users WHERE id = &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AI comment: "SQL injection vulnerability. Use parameterized queries instead."&lt;/p&gt;

&lt;p&gt;You'd catch this on review, but not always. Not on PR #73 when you've reviewed 10 already.&lt;/p&gt;

&lt;h3&gt;
  
  
  Catch #2: Async/Await Footgun
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;users&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetchData&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="c1"&gt;// This runs sequentially, not in parallel&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AI: "Use &lt;code&gt;Promise.all()&lt;/code&gt; here to parallelize requests."&lt;/p&gt;

&lt;p&gt;This is the stuff that ships and causes performance issues in production.&lt;/p&gt;

&lt;h3&gt;
  
  
  Catch #3: React Dependency Array
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;useEffect&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;fetchData&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;[]);&lt;/span&gt; &lt;span class="c1"&gt;// Missing 'query' dependency&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AI: "Add &lt;code&gt;query&lt;/code&gt; to dependency array to prevent stale data."&lt;/p&gt;

&lt;p&gt;Humans zone out on these. AI doesn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  What To Avoid
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Don't:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use AI to avoid hiring reviewers. You still need human insight.&lt;/li&gt;
&lt;li&gt;Trust it 100%. False positives happen. Context matters.&lt;/li&gt;
&lt;li&gt;Make it review design decisions. That's your job.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Do:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use it to catch mechanical errors and security issues.&lt;/li&gt;
&lt;li&gt;Let it handle the "did you mean to do this?" questions.&lt;/li&gt;
&lt;li&gt;Treat it as a junior dev who's always available.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Setup (In 30 Minutes)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Go to &lt;a href="https://coderabbit.ai" rel="noopener noreferrer"&gt;CodeRabbit&lt;/a&gt; or &lt;a href="https://sweep.dev" rel="noopener noreferrer"&gt;Sweep&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Connect your GitHub repo&lt;/li&gt;
&lt;li&gt;Set preferences: languages, severity levels, areas to focus on&lt;/li&gt;
&lt;li&gt;Make a PR—watch it work&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's it. Your CI/CD now includes an AI reviewer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real ROI
&lt;/h2&gt;

&lt;p&gt;We measured this on a team of 5. First month:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;40% fewer "fixup" PRs (where the only changes were addressing review comments)&lt;/li&gt;
&lt;li&gt;3 security issues caught that humans missed&lt;/li&gt;
&lt;li&gt;Code review time dropped 25% (less back-and-forth)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Not revolutionary. But real.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;The next wave: predictive analysis. "This change might break X tests" before you even submit. "You're touching code that loads the homepage—performance impact?"&lt;/p&gt;

&lt;p&gt;That's coming.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your Move
&lt;/h2&gt;

&lt;p&gt;Pick a tool. Set it up this week. Run it on your next 10 PRs. If it catches one real bug, it paid for itself. If it doesn't, you lost an hour setting it up—not a big deal.&lt;/p&gt;

&lt;p&gt;The fact that you &lt;em&gt;can&lt;/em&gt; do this now? Use it. Your future self won't regret catching bugs earlier.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Want to stay ahead of tools like this?&lt;/strong&gt; &lt;a href="https://learnairesource.com/newsletter" rel="noopener noreferrer"&gt;Subscribe to LearnAI Weekly&lt;/a&gt; for practical AI workflows for developers, every Friday in your inbox.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>codereviews</category>
      <category>productivity</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Stop Waiting for API Calls: Running Local LLMs in Your Dev Workflow</title>
      <dc:creator>LearnAI Resource</dc:creator>
      <pubDate>Mon, 17 Aug 2026 15:00:59 +0000</pubDate>
      <link>https://dev.to/learnairesource/stop-waiting-for-api-calls-running-local-llms-in-your-dev-workflow-38i</link>
      <guid>https://dev.to/learnairesource/stop-waiting-for-api-calls-running-local-llms-in-your-dev-workflow-38i</guid>
      <description>&lt;h1&gt;
  
  
  Stop Waiting for API Calls: Running Local LLMs in Your Dev Workflow
&lt;/h1&gt;

&lt;p&gt;You know that moment? You're deep in the zone, writing code, and you need to ask an AI something. But the API is slow, you hit a rate limit, or you're offline. Your flow breaks.&lt;/p&gt;

&lt;p&gt;Local LLMs fix that. And they're actually pretty fast now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Local LLMs Matter
&lt;/h2&gt;

&lt;p&gt;Here's the thing: cloud-based AI is great for one-off questions. But if you're using AI constantly — for code review, brainstorming, debugging, refactoring — local models let you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Zero latency.&lt;/strong&gt; No network round-trip. Your machine runs it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No rate limits.&lt;/strong&gt; Ask a thousand questions without throttling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No API costs.&lt;/strong&gt; Especially for large projects with lots of queries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Privacy.&lt;/strong&gt; Your code stays local.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Works offline.&lt;/strong&gt; Stuck on a plane? Still productive.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Getting Started (Actually Easy)
&lt;/h2&gt;

&lt;p&gt;The barrier-to-entry has basically disappeared. Here are the real steps:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Pick Your Model
&lt;/h3&gt;

&lt;p&gt;For 2026, solid choices:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ollama&lt;/strong&gt; — Dead simple. Download a model, run it. Supports Llama 2, Mistral, Neural Chat, and tons others.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LM Studio&lt;/strong&gt; — GUI wrapper. Good if you hate the terminal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPT4All&lt;/strong&gt; — Lightweight, older but still solid for local work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'd start with Ollama because the community is huge and adding it to your workflow takes like 5 minutes.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Install &amp;amp; Run
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install Ollama (macOS, Linux, Windows now)&lt;/span&gt;
&lt;span class="c"&gt;# Then run your chosen model:&lt;/span&gt;

ollama run mistral  &lt;span class="c"&gt;# Fast, ~7B params&lt;/span&gt;
ollama run neural-chat  &lt;span class="c"&gt;# Good balance&lt;/span&gt;
ollama run llama2  &lt;span class="c"&gt;# Slower but very capable&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pick one based on your machine's RAM. Seriously, check your VRAM first.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Connect It to Your Tools
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;VS Code:&lt;/strong&gt; Install the "Ollama" extension or use "Continue.dev" (game-changing, btw).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CLI:&lt;/strong&gt; Use &lt;code&gt;curl&lt;/code&gt; to chat with the model running on &lt;code&gt;localhost:11434&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:11434/api/generate &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model": "mistral", "prompt": "explain this function"}'&lt;/span&gt; | jq &lt;span class="s1"&gt;'.response'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Python/Node:&lt;/strong&gt; Libraries exist for everything. &lt;code&gt;ollama-py&lt;/code&gt; for Python is solid.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Example
&lt;/h2&gt;

&lt;p&gt;I use this in my actual workflow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# My .bashrc function:&lt;/span&gt;
&lt;span class="k"&gt;function &lt;/span&gt;ask&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
  curl &lt;span class="nt"&gt;-s&lt;/span&gt; http://localhost:11434/api/generate &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s2"&gt;"{
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;model&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;: &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;mistral&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;prompt&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;: &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;stream&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;: false
  }"&lt;/span&gt; | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.response'&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;

&lt;span class="c"&gt;# Usage:&lt;/span&gt;
ask &lt;span class="s2"&gt;"refactor this code for readability"&lt;/span&gt; &amp;lt; messy-file.js
ask &lt;span class="s2"&gt;"what are edge cases in OAuth2?"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Takes ~1-2 seconds. My brain doesn't even context-switch.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trade-offs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Honest downsides:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Local models are dumber than GPT-4. Not worse, just different. Good for coding tasks, brainstorming, and pattern matching. Weak on novel reasoning.&lt;/li&gt;
&lt;li&gt;Setup takes RAM. You probably want 16GB minimum. 32GB is comfortable.&lt;/li&gt;
&lt;li&gt;Quality varies wildly by model. Test a few.&lt;/li&gt;
&lt;li&gt;No persistent memory between sessions (unless you build it).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;But:&lt;/strong&gt; For daily dev work — refactoring, debugging, explaining, writing boilerplate — local is often better than cloud because of speed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Hybrid Approach
&lt;/h2&gt;

&lt;p&gt;Real talk: I don't replace cloud APIs. I complement them.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Local (Mistral):&lt;/strong&gt; Quick questions, code review, explaining syntax&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloud (Claude/GPT-4):&lt;/strong&gt; Hard problems, novel ideas, things that need reasoning&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Local is your quick diff tool. Cloud is your architect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Level Up: Make It Smarter
&lt;/h2&gt;

&lt;p&gt;Pipe your codebase into the context and build wrappers that context-switch between your local model and cloud APIs based on question type.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Matters
&lt;/h2&gt;

&lt;p&gt;The future isn't "AI does your job." It's "AI runs in your pocket, always available, never throttled." Local LLMs are that future arriving now.&lt;/p&gt;

&lt;p&gt;Stop waiting for APIs. Spin up a local model tonight. Your development flow will thank you.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Want to stay current on AI tools and workflows that actually matter?&lt;/strong&gt; Check out &lt;a href="https://learnairesource.com/newsletter" rel="noopener noreferrer"&gt;LearnAI Weekly&lt;/a&gt; — real tips from people using this stuff daily, not marketing fluff.&lt;/p&gt;

&lt;p&gt;What's your go-to local model? Or are you cloud-only? Drop a comment.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devtools</category>
      <category>productivity</category>
      <category>llm</category>
    </item>
    <item>
      <title>Why Your AI Code Assistant Misses Context (And How to Fix It)</title>
      <dc:creator>LearnAI Resource</dc:creator>
      <pubDate>Sun, 16 Aug 2026 15:00:35 +0000</pubDate>
      <link>https://dev.to/learnairesource/why-your-ai-code-assistant-misses-context-and-how-to-fix-it-p4p</link>
      <guid>https://dev.to/learnairesource/why-your-ai-code-assistant-misses-context-and-how-to-fix-it-p4p</guid>
      <description>&lt;p&gt;You paste code into your AI tool, ask it to help, and get back something that doesn't quite fit your project. Sound familiar?&lt;/p&gt;

&lt;p&gt;This happens because AI tools don't understand your codebase context by default. They see isolated snippets, not systems. Here's how to actually fix that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem: AI Operates in a Vacuum
&lt;/h2&gt;

&lt;p&gt;Your code assistant sees:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The 50 lines you pasted&lt;/li&gt;
&lt;li&gt;Maybe some function signatures&lt;/li&gt;
&lt;li&gt;Zero knowledge of your architecture, patterns, or decisions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Your codebase has:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;10k+ lines across 20 files&lt;/li&gt;
&lt;li&gt;Specific conventions you've built over time&lt;/li&gt;
&lt;li&gt;Business logic baked into naming and structure&lt;/li&gt;
&lt;li&gt;Dependencies, edge cases, and context&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No wonder the suggestions feel off.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Works
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Paste Your Architecture, Not Just Your Problem
&lt;/h3&gt;

&lt;p&gt;Before asking for help with a specific feature, give the AI your project structure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Your directory structure:
src/
  services/
    userService.ts  -- handles auth and user ops
    dataService.ts  -- caches queries in redis
  models/
    User.ts
    Session.ts
  middleware/
    auth.ts
    errorHandler.ts

Key patterns:
- All services return {success, data, error}
- No direct DB calls outside services
- Redis for frequently-queried data
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now when you ask "how do I add a password reset feature," the AI knows your actual structure.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Show Examples, Not Just Requirements
&lt;/h3&gt;

&lt;p&gt;Instead of: "Add validation to this form"&lt;/p&gt;

&lt;p&gt;Try: "I validate forms the same way we do in UserForm.tsx and AccountForm.tsx. How should I apply that pattern to this new component?"&lt;/p&gt;

&lt;p&gt;Paste one working example. AI will replicate your style much better than it'll invent one.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Include Actual Error Messages and Logs
&lt;/h3&gt;

&lt;p&gt;When you hit a bug, your instinct is to simplify the error before asking. Don't.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error: TypeError: Cannot read property 'map' of undefined
at MapService.filterResults (services/map.ts:24)

Stack trace shows:
  - queryData is coming back null sometimes
  - happens when redis connection drops
  - shouldn't happen in production but does
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This tells the AI &lt;em&gt;what's actually broken&lt;/em&gt;, not your guessed interpretation.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Specify Your Stack's Quirks
&lt;/h3&gt;

&lt;p&gt;Different frameworks have different conventions. Tell the AI yours:&lt;/p&gt;

&lt;p&gt;"We use React 19 with Suspense for loading states. We don't use Redux—just context. We prefer controlled components over refs."&lt;/p&gt;

&lt;p&gt;Don't assume it knows your decisions.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Ask for Explanations, Not Just Code
&lt;/h3&gt;

&lt;p&gt;"Here's my current implementation. Why does this pattern work for us? What are the tradeoffs?"&lt;/p&gt;

&lt;p&gt;AI gives better advice when you ask it to think about &lt;em&gt;your&lt;/em&gt; constraints, not just produce code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real Example
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Bad prompt:&lt;/strong&gt;&lt;br&gt;
"Add error handling to this API call"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Good prompt:&lt;/strong&gt;&lt;br&gt;
"Our API wrapper (utils/api.ts) catches errors and returns {success, data, error}. All components expect this shape. Our error boundary catches unhandled errors and logs to Sentry. How should I handle this specific case where the API times out but the user doesn't leave the page?"&lt;/p&gt;

&lt;p&gt;The second one works because the AI understands your actual system.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Speedrun Version
&lt;/h2&gt;

&lt;p&gt;If you're in a hurry:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Paste your project structure in a comment (takes 30 seconds)&lt;/li&gt;
&lt;li&gt;Ask your question&lt;/li&gt;
&lt;li&gt;AI's suggestions will be 10x more accurate&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Seriously, try it. The difference is wild.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Matters
&lt;/h2&gt;

&lt;p&gt;AI tools are getting smarter, but they're still dumb about &lt;em&gt;your&lt;/em&gt; context. You have to be the teacher. The more you teach it about your codebase's specific patterns and constraints, the better it becomes.&lt;/p&gt;

&lt;p&gt;This isn't forever—eventually tools will scan your whole repo automatically. But right now? You're the bridge between the AI and your actual system.&lt;/p&gt;

&lt;p&gt;Give it context. Get better suggestions. Save time.&lt;/p&gt;




&lt;p&gt;Want to level up your development workflow with AI tools, productivity hacks, and no-code solutions? Check out the &lt;strong&gt;&lt;a href="https://learnairesource.com/newsletter" rel="noopener noreferrer"&gt;LearnAI Weekly newsletter&lt;/a&gt;&lt;/strong&gt; for practical tips that actually work in production.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>coding</category>
      <category>developer</category>
    </item>
    <item>
      <title>Stop Writing Vague Prompts: A Developer's Guide to Getting AI Tools to Actually Help</title>
      <dc:creator>LearnAI Resource</dc:creator>
      <pubDate>Sat, 15 Aug 2026 15:00:26 +0000</pubDate>
      <link>https://dev.to/learnairesource/stop-writing-vague-prompts-a-developers-guide-to-getting-ai-tools-to-actually-help-1o49</link>
      <guid>https://dev.to/learnairesource/stop-writing-vague-prompts-a-developers-guide-to-getting-ai-tools-to-actually-help-1o49</guid>
      <description>&lt;p&gt;You've probably noticed that AI tools spit out garbage when you ask them garbage questions. But here's the thing nobody talks about—most developers are terrible at asking. And if you're terrible at asking, you're wasting hours waiting for terrible answers.&lt;/p&gt;

&lt;p&gt;Let me walk you through how to actually get useful output from Claude, ChatGPT, GitHub Copilot, or whatever AI tool you're using.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem With "Just Works" Prompts
&lt;/h2&gt;

&lt;p&gt;You know what doesn't work?&lt;/p&gt;

&lt;p&gt;"Generate a sorting algorithm"&lt;br&gt;
"Fix my code"&lt;br&gt;
"Explain this regex"&lt;/p&gt;

&lt;p&gt;These prompts are so generic that the AI has to guess what you actually need. And it's probably guessing wrong.&lt;/p&gt;

&lt;p&gt;Compare that to:&lt;/p&gt;

&lt;p&gt;"I'm building a checkout flow in React. When users click the "pay" button, I'm calling &lt;code&gt;submitOrder()&lt;/code&gt; which hits our Stripe API. Currently, if the call takes &amp;gt;3 seconds, users see a white screen. Write me a loading state that shows a spinner with disabled button and a "Processing..." message."&lt;/p&gt;

&lt;p&gt;That's 100x better because the AI knows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What you're building (React checkout)&lt;/li&gt;
&lt;li&gt;What the problem is (no feedback during slow requests)&lt;/li&gt;
&lt;li&gt;What you want as output (specific UI pattern)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Secret: Context &amp;gt; Cleverness
&lt;/h2&gt;

&lt;p&gt;Here's the pattern that actually works:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. What you're building&lt;/strong&gt; — "I'm building a CLI tool in Node.js that syncs local files to S3"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. What's broken or what you want&lt;/strong&gt; — "I need to handle rate limiting when uploading 1000+ files"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. What constraints matter&lt;/strong&gt; — "Can't use third-party libraries, needs to run on Node 18+"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. What you want as output&lt;/strong&gt; — "Show me a retry strategy with exponential backoff"&lt;/p&gt;

&lt;p&gt;That's it. Context kills vagueness.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real Examples
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Bad:&lt;/strong&gt; "How do I optimize my database?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Good:&lt;/strong&gt; "I've got a PostgreSQL table with 2M rows. Queries to find users by email are taking 400ms. I already have an email column. What's the minimal change to speed this up?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bad:&lt;/strong&gt; "Write API documentation"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Good:&lt;/strong&gt; "I'm documenting a REST API for managing team projects. Focus on the POST /projects endpoint. Include request/response examples, auth requirements, and common errors. Format as markdown."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bad:&lt;/strong&gt; "Help with regex"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Good:&lt;/strong&gt; "I need to extract phone numbers from customer support transcripts. They're in various formats: +1-555-123-4567, (555) 123-4567, 555.123.4567. Write a regex that matches all three and captures the 10-digit number."&lt;/p&gt;

&lt;h2&gt;
  
  
  One More Thing: Be Specific About Output Format
&lt;/h2&gt;

&lt;p&gt;"Write me an article" is vague. "Write me a 500-word blog post in markdown format with a catchy intro and 3 practical examples" is clear.&lt;/p&gt;

&lt;p&gt;"Debug this" is frustrating. "This function returns undefined instead of an array. Walk me through what's happening step-by-step, then show the fix with inline comments" is actionable.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Payoff
&lt;/h2&gt;

&lt;p&gt;Spend 30 seconds writing a better prompt. Save 10 minutes waiting for a useful answer. Do that 5 times a day and you've just gained an hour.&lt;/p&gt;

&lt;p&gt;That's the real superpower—not fancy AI techniques, just not wasting your own time by being vague.&lt;/p&gt;

&lt;p&gt;Next time you're about to copy-paste a half-baked question into an AI tool, pause. Ask yourself: "Could I be more specific?" Chances are, you can.&lt;/p&gt;

&lt;p&gt;And you probably should.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Leveling up your productivity game?&lt;/strong&gt; Check out the &lt;a href="https://learnairesource.com/newsletter" rel="noopener noreferrer"&gt;LearnAI Weekly newsletter&lt;/a&gt; for more practical tips on tools, workflows, and the AI stuff that actually matters.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>devtools</category>
      <category>tips</category>
    </item>
    <item>
      <title>Stop Copy-Pasting AI Code: Own What You Ship</title>
      <dc:creator>LearnAI Resource</dc:creator>
      <pubDate>Fri, 14 Aug 2026 15:00:30 +0000</pubDate>
      <link>https://dev.to/learnairesource/stop-copy-pasting-ai-code-own-what-you-ship-17hj</link>
      <guid>https://dev.to/learnairesource/stop-copy-pasting-ai-code-own-what-you-ship-17hj</guid>
      <description>&lt;p&gt;You've got Claude/GPT running in your IDE. You ask it to write a function. It spits out 47 lines of code. You hit paste. Boom—it's in your codebase.&lt;/p&gt;

&lt;p&gt;Here's the thing: if you don't understand what that code does, you just shipped a problem.&lt;/p&gt;

&lt;p&gt;I'm not saying don't use AI tools. They're &lt;em&gt;fast&lt;/em&gt; and genuinely useful. But there's a difference between using AI as a thought partner and outsourcing your brain.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Problem
&lt;/h2&gt;

&lt;p&gt;Three months from now, that function breaks. Maybe it's edge case handling. Maybe it's a subtle bug with async behavior. Maybe it was never actually correct—it just worked for your specific test case.&lt;/p&gt;

&lt;p&gt;When you have to debug it at 2 AM, you're going to wish you understood what it does.&lt;/p&gt;

&lt;p&gt;The developers who are thriving with AI tools aren't the ones who use them blindly. They're the ones who treat AI output like code review—skeptical, hands-on, and ready to challenge it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Works
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Generate, then rewrite it yourself&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Don't just accept the AI's structure. Ask yourself: would I write this differently? Is there a simpler way? What edge cases is this missing?&lt;/p&gt;

&lt;p&gt;I asked Claude to write a date parser recently. It gave me something solid, but then I rewrote 30% of it because I realized I could handle timezone logic differently. The AI solution was correct; my version was just more aligned with how my brain works.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Understand every function before committing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This isn't paranoia. Read the code. Trace through it. Ask "why did it do that?" about the parts that aren't obvious. If you can't explain it, you don't understand it.&lt;/p&gt;

&lt;p&gt;Real talk: if it takes 5 minutes to understand an AI-generated function, that's a good sign. If it takes 30 minutes, something's probably overcomplicated.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Test aggressively&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;AI tools are good at the happy path. They're not mind-readers about your specific data, your production constraints, or your weird edge cases.&lt;/p&gt;

&lt;p&gt;Write tests. Especially the awkward ones. Empty arrays. Null values. Huge datasets. Malformed input. Make the code prove it works.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Mix AI with your own solutions&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Don't let AI be your only approach to a problem. Sometimes I'll ask an AI tool to solve something, then I'll also solve it my way, then compare. I learn way more that way than just taking what it gives me.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bigger Picture
&lt;/h2&gt;

&lt;p&gt;This is about maintaining the skill that matters: &lt;strong&gt;problem-solving&lt;/strong&gt;. AI is amazing at generating &lt;em&gt;syntax&lt;/em&gt;. It's terrible at understanding your actual problem if you're not clear about it.&lt;/p&gt;

&lt;p&gt;When you engage with AI output—questioning it, testing it, refining it—you're actually building better instincts about what good code looks like. You're learning.&lt;/p&gt;

&lt;p&gt;When you just paste and move on, you're getting lazier at the exact skill that keeps you valuable.&lt;/p&gt;

&lt;h2&gt;
  
  
  One More Thing
&lt;/h2&gt;

&lt;p&gt;The developers who are getting pushed out aren't the ones who use AI. They're the ones who outsourced thinking. The ones who stopped asking questions.&lt;/p&gt;

&lt;p&gt;Use your tools. But be present. Code with intention. Understand what you ship.&lt;/p&gt;

&lt;p&gt;That's how you stay sharp.&lt;/p&gt;




&lt;p&gt;Want more on building real skills in 2026? Check out &lt;strong&gt;&lt;a href="https://learnairesource.com/newsletter" rel="noopener noreferrer"&gt;LearnAI Weekly&lt;/a&gt;&lt;/strong&gt; for practical developer resources that actually help you stay ahead.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>coding</category>
      <category>development</category>
    </item>
    <item>
      <title>How to Build a Custom Code Review System with Claude API</title>
      <dc:creator>LearnAI Resource</dc:creator>
      <pubDate>Thu, 13 Aug 2026 15:00:36 +0000</pubDate>
      <link>https://dev.to/learnairesource/how-to-build-a-custom-code-review-system-with-claude-api-4oc0</link>
      <guid>https://dev.to/learnairesource/how-to-build-a-custom-code-review-system-with-claude-api-4oc0</guid>
      <description>&lt;h1&gt;
  
  
  How to Build a Custom Code Review System with Claude's API
&lt;/h1&gt;

&lt;p&gt;So you're reviewing pull requests for your team, and it's draining. Context switching, nitpicks, architectural questions—it adds up. What if you had a second set of eyes that actually &lt;em&gt;thinks&lt;/em&gt; about code, not just lint errors?&lt;/p&gt;

&lt;p&gt;I built a custom code review system using Claude's API, and it's saved me hours each week. Here's how.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem with Generic Code Review Tools
&lt;/h2&gt;

&lt;p&gt;Most linters catch syntax errors. That's useful. But they miss the important stuff: Is this approach going to bite us in six months? Are we introducing technical debt? Is there a cleaner way to do this?&lt;/p&gt;

&lt;p&gt;Claude can actually reason about code. It understands context, trade-offs, and design patterns. That's different from a regex pattern-matcher.&lt;/p&gt;

&lt;h2&gt;
  
  
  What We're Building
&lt;/h2&gt;

&lt;p&gt;A simple bot that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Reads a pull request diff&lt;/li&gt;
&lt;li&gt;Analyzes it for actual architectural issues, not just style&lt;/li&gt;
&lt;li&gt;Suggests concrete improvements&lt;/li&gt;
&lt;li&gt;Flags potential bugs or performance problems&lt;/li&gt;
&lt;li&gt;Stays out of your way on obvious good code&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Setting It Up
&lt;/h2&gt;

&lt;p&gt;First, you'll need the Claude API. Grab your API key from &lt;a href="https://console.anthropic.com" rel="noopener noreferrer"&gt;Anthropic's console&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install the SDK&lt;/span&gt;
npm &lt;span class="nb"&gt;install&lt;/span&gt; @anthropic-ai/sdk
&lt;span class="c"&gt;# or&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;anthropic
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Create a simple script to read a diff and send it to Claude:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;review_code_diff&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;diff_content&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-3-5-sonnet-20241022&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Review this code diff for architectural issues, potential bugs, and design improvements.
Focus on:
- Logic errors or edge cases
- Performance problems
- Code clarity and maintainability
- Missed error handling
- Security concerns

Keep feedback brief and actionable. Ignore style/formatting issues.

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;{diff_content}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;            }
        ]
    )
    return message.content[0].text

# In practice, you'd read the actual diff from your git/GitHub API
diff = """
--- a/api/auth.js
+++ b/api/auth.js
@@ -15,7 +15,12 @@ async function login(email, password) {
     const user = await db.users.findOne({ email });
     if (!user) return { error: "Invalid credentials" };
-    if (password !== user.password) return { error: "Invalid credentials" };
+    // Check against hashed password
+    const match = await bcrypt.compare(password, user.password);
+    if (!match) return { error: "Invalid credentials" };
+    
+    const token = jwt.sign({ id: user.id }, process.env.JWT_SECRET, { expiresIn: "24h" });
+    return { token };
 """

feedback = review_code_diff(diff)
print(feedback)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Real-World Example
&lt;/h2&gt;

&lt;p&gt;I ran this on a diff from a recent PR (refactoring our payment processing). Claude caught:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Missing error handling&lt;/strong&gt; - If the payment gateway times out, we silently fail instead of retrying&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Race condition&lt;/strong&gt; - Two concurrent requests could create duplicate transactions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unused variable&lt;/strong&gt; - Declared but never used (the linter missed this somehow)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Better approach&lt;/strong&gt; - Suggested using a webhook queue instead of polling&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's the kind of feedback that actually prevents bugs. A linter would've caught #3, maybe. Claude caught all of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Integrating with Your Workflow
&lt;/h2&gt;

&lt;p&gt;You have options here:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Option 1: GitHub Action&lt;/strong&gt; - Run on every PR automatically&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v3&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Review PR&lt;/span&gt;
  &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;node review-pr.js&lt;/span&gt;
  &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;ANTHROPIC_API_KEY&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.ANTHROPIC_API_KEY }}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Option 2: Slack bot&lt;/strong&gt; - Review on demand&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@app.message&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;review&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;handle_review&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ack&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;say&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;ack&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;diff&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_latest_diff&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# Your GitHub integration
&lt;/span&gt;    &lt;span class="n"&gt;feedback&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;review_code_diff&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;diff&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;say&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Code Review:&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;feedback&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Option 3: Local CLI&lt;/strong&gt; - Review before pushing&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git diff | node local-review.js
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Things to Watch Out For
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Cost&lt;/strong&gt; - Claude isn't free. A typical diff costs $0.01-0.05. For a team, this adds up. Set token limits per review.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hallucinations&lt;/strong&gt; - Claude sometimes invents function names or makes assumptions. Always verify suggestions before implementing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context windows&lt;/strong&gt; - Really large diffs (1000+ line changes) might get cut off. Break them into smaller reviews.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;False positives&lt;/strong&gt; - It'll sometimes flag things that aren't actually issues. Use it as a second opinion, not gospel.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Win
&lt;/h2&gt;

&lt;p&gt;You're not replacing human code review. You're automating the grunt work. Your team reads Claude's feedback first—catches the obvious stuff—then focuses on the hard architectural decisions.&lt;/p&gt;

&lt;p&gt;In practice, this cuts review time by 30-40%. More importantly, it catches bugs &lt;em&gt;before&lt;/em&gt; they hit production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next Steps
&lt;/h2&gt;

&lt;p&gt;Start small. Run Claude on one PR. See if the feedback is useful. Adjust the prompt based on your codebase.&lt;/p&gt;

&lt;p&gt;Some teams add it to their CI/CD pipeline. Others use it as a pre-review step in their IDE. Find what fits your workflow.&lt;/p&gt;

&lt;p&gt;The code is straightforward. The real value is tuning the prompt to your team's standards. After a few iterations, Claude learns what you care about.&lt;/p&gt;




&lt;p&gt;Want more on building with AI tools? Check out &lt;a href="https://learnairesource.com/newsletter" rel="noopener noreferrer"&gt;LearnAI Weekly newsletter&lt;/a&gt;—it's practical stuff about integrating AI into real workflows, not hype.&lt;/p&gt;

&lt;p&gt;What would you build with code-understanding AI? Drop a comment.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>coding</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
