<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: zunairah</title>
    <description>The latest articles on DEV Community by zunairah (@zunairah_bfe3d030a9be261c).</description>
    <link>https://dev.to/zunairah_bfe3d030a9be261c</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4117361%2Fa5a0a0b7-0122-4582-9e76-b0132812533d.png</url>
      <title>DEV Community: zunairah</title>
      <link>https://dev.to/zunairah_bfe3d030a9be261c</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/zunairah_bfe3d030a9be261c"/>
    <language>en</language>
    <item>
      <title>6 Sneaky LLM Bugs That Only Show Up After You Ship</title>
      <dc:creator>zunairah</dc:creator>
      <pubDate>Sat, 12 Sep 2026 11:38:27 +0000</pubDate>
      <link>https://dev.to/zunairah_bfe3d030a9be261c/6-sneaky-llm-bugs-that-only-show-up-after-you-ship-3bh0</link>
      <guid>https://dev.to/zunairah_bfe3d030a9be261c/6-sneaky-llm-bugs-that-only-show-up-after-you-ship-3bh0</guid>
      <description>&lt;p&gt;Okay, real talk. You know that feeling when your AI feature works perfectly in testing, you ship it on a Friday feeling like a genius, and then Monday morning something's on fire and you have no idea why?&lt;/p&gt;

&lt;p&gt;Yeah. That feeling has a name, and it's usually one of these six things.&lt;/p&gt;

&lt;p&gt;None of these are exotic. There's no galaxy-brain fix here. They're just small, easy-to-miss decisions that look totally fine in a demo and quietly turn into 2am pages once real users show up. Let's fix them now so you don't have to learn them the hard way.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Stop hiding "streaming or not" behind a flag&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Raise your hand if you've written something like this:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
def ask(prompt, stream=False):&lt;br&gt;
    if stream:&lt;br&gt;
        # return a generator&lt;br&gt;
    else:&lt;br&gt;
        # return a string&lt;/p&gt;

&lt;p&gt;Feels efficient, right? One function, does everything. Except now every single place that calls ask() has to remember what stream=True does to the return type, and the day someone forgets, they're trying to call .upper() on a generator and wondering what they did wrong.&lt;/p&gt;

&lt;p&gt;Just split it into two functions:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
def ask(prompt: str) -&amp;gt; str:&lt;br&gt;
    """Always returns a full string."""&lt;br&gt;
    ...&lt;/p&gt;

&lt;p&gt;def ask_streaming(prompt: str):&lt;br&gt;
    """Always yields chunks. Always a generator."""&lt;br&gt;
    ...&lt;/p&gt;

&lt;p&gt;Boring? Sure. But now the function name tells you exactly what you're getting, and nobody has to hold extra state in their head to use it correctly. Future-you will say thanks.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Streaming calls need a context manager, not just a loop&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Here's a fun one. You write a loop to print tokens as they stream in. It works great — until the loop throws an error halfway through (a rendering bug, a network blip, whatever). Does the connection actually close?&lt;/p&gt;

&lt;p&gt;With a naive setup: often, no. It just... sits there. Leaking. Quietly. Until you're staring at your server's connection count going up and up with no idea why.&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
with client.messages.stream(...) as stream:&lt;br&gt;
    for text in stream.text_stream:&lt;br&gt;
        yield text&lt;/p&gt;

&lt;p&gt;Wrapping it in a context manager means cleanup happens no matter what — even if things blow up mid-stream. It's a one-line change that you'll never notice... until the one day it saves you.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;When trimming chat history, count in pairs, not messages&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you're building any kind of chatbot, you eventually need to cap how much history you send back to the model (context windows aren't infinite, and neither is your API bill). The natural instinct is "just keep the last N messages."&lt;/p&gt;

&lt;p&gt;Here's the trap: if N lands you in the middle of a user/assistant back-and-forth, you get an orphaned assistant message with no user message before it — and a lot of APIs will just reject that outright.&lt;/p&gt;

&lt;p&gt;The fix is almost embarrassingly simple once you see it:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
recent = self.history[-(self.max_history_turns * 2):]&lt;/p&gt;

&lt;p&gt;Multiply by 2, slice from the end — now you're always keeping whole conversational turns, never a half-finished one.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Forgetting to normalize vectors = wrong search results, zero errors&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is my favorite one because it's sneaky. No crash, no error message, nothing that tells you something's wrong. Just... search results that feel a little off.&lt;/p&gt;

&lt;p&gt;If you're using FAISS with IndexFlatIP to do cosine similarity search, that trick only works if your vectors are normalized first. Skip that step, and you're silently doing plain inner-product search instead — which ranks things differently, and there's no red flag telling you why your "most similar" results feel not-quite-right.&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
norms = np.linalg.norm(vectors, axis=1, keepdims=True)&lt;br&gt;
normalized = vectors / np.clip(norms, 1e-10, None)&lt;/p&gt;

&lt;p&gt;One line. Easy to forget. Impossible to notice until you go digging.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Not every error deserves a retry&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Retrying failed API calls feels responsible — rate limits happen, timeouts happen, stuff breaks sometimes and trying again is the grown-up thing to do. But if you retry everything indiscriminately, you'll also retry your own bugs. Sent a malformed request? Cool, now you're going to fail the exact same way four times in a row, burning time and rate-limit budget for absolutely nothing.&lt;/p&gt;

&lt;p&gt;Be picky about what you retry:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
@retry(&lt;br&gt;
    retry=retry_if_exception_type((RateLimitError, APITimeoutError, APIError)),&lt;br&gt;
    stop=stop_after_attempt(4),&lt;br&gt;
    wait=wait_exponential(multiplier=1, min=1, max=20),&lt;br&gt;
)&lt;br&gt;
def call_model(...):&lt;br&gt;
    ...&lt;/p&gt;

&lt;p&gt;Transient stuff (rate limits, timeouts)? Worth a retry with backoff. Your own bug? Let it fail fast so you actually see it and fix it, instead of hiding it behind four identical failed attempts.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Have a plan B model, not just a retry loop&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Retries are great for "oops, blip" moments. They don't help much if a model provider is having a genuinely bad day for an extended stretch. That's where a fallback model earns its keep — try the fast/cheap one first, and if it keeps failing, fall back to a stronger (or just different) model instead of just erroring out on your users.&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
try:&lt;br&gt;
    return call_model(primary_model, messages)&lt;br&gt;
except RetryableErrors:&lt;br&gt;
    return call_model(fallback_model, messages)&lt;/p&gt;

&lt;p&gt;It's the same trick most production AI gateways use behind the scenes. Costs you a few extra lines, saves you a very bad afternoon.&lt;/p&gt;

&lt;p&gt;Honestly, none of these are hard once you know them — that's kind of the whole point. They're just the small stuff that's easy to skip when you're moving fast, and painful to debug when they finally bite. (If you want the full runnable versions of these patterns plus a few more — RAG, tool calling, memory — they're all written up in the AI &amp;amp; LLM Integration Cookbook, but the six above will already save you a rough night regardless.)&lt;/p&gt;

&lt;p&gt;Now go add that context manager before you forget. 😄&lt;br&gt;
Get the AI and LLM cookbook at 20% off&lt;br&gt;
product link-&lt;a href="https://payhip.com/b/0dUzx" rel="noopener noreferrer"&gt;https://payhip.com/b/0dUzx&lt;/a&gt;&lt;/p&gt;

</description>
      <category>code</category>
      <category>coding</category>
      <category>llm</category>
      <category>ai</category>
    </item>
    <item>
      <title>6 LLM Integration Mistakes That Look Fine in a Demo and Break in Production</title>
      <dc:creator>zunairah</dc:creator>
      <pubDate>Sat, 12 Sep 2026 11:21:34 +0000</pubDate>
      <link>https://dev.to/zunairah_bfe3d030a9be261c/6-llm-integration-mistakes-that-look-fine-in-a-demo-and-break-in-production-429b</link>
      <guid>https://dev.to/zunairah_bfe3d030a9be261c/6-llm-integration-mistakes-that-look-fine-in-a-demo-and-break-in-production-429b</guid>
      <description>&lt;p&gt;Most LLM code you find online works great in a Jupyter notebook and falls apart the moment real traffic hits it. The bugs aren't exotic — they're small, structural decisions that don't show up until something goes wrong at the worst possible time. Here are six of them, and the fix for each.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Don't hide streaming behind a boolean flag&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;It's tempting to write one function with an if stream: branch. Don't. It means every caller has to know, out of band, which type they're going to get back — a plain string or a generator — and a wrong guess fails silently or crashes deep in your UI code.&lt;/p&gt;

&lt;p&gt;Split it into two functions instead:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
def ask(prompt: str) -&amp;gt; str:&lt;br&gt;
    """Returns the full text response."""&lt;br&gt;
    ...&lt;/p&gt;

&lt;p&gt;def ask_streaming(prompt: str):&lt;br&gt;
    """Yields text chunks as they arrive."""&lt;br&gt;
    ...&lt;/p&gt;

&lt;p&gt;Now the function signature is the documentation. ask() gives you something easy to log, cache, and unit test. ask_streaming() is unambiguously a generator, built for UIs. No flag, no ambiguity.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Wrap streaming calls in a context manager&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If your loop over a streaming response throws an exception halfway through — network hiccup, a bug in your own rendering code, whatever — does the underlying HTTP connection actually close? With a naive implementation, often not. Connections leak quietly until you're wondering why your process is hoarding sockets.&lt;/p&gt;

&lt;p&gt;The fix is boring but effective:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
with client.messages.stream(...) as stream:&lt;br&gt;
    for text in stream.text_stream:&lt;br&gt;
        yield text&lt;/p&gt;

&lt;p&gt;The context manager guarantees cleanup runs even on an error mid-stream. This is a one-line difference that only matters the day something actually goes wrong — which is exactly when you don't want to be debugging a connection leak too.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Trim chat history on turn pairs, not message count&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A common way to cap conversation memory is to just keep "the last N messages." The bug: if N happens to land in the middle of a user/assistant pair, you end up with an orphaned assistant message with no matching user turn — and some APIs will outright reject that payload.&lt;/p&gt;

&lt;p&gt;Slice on pairs instead:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
def _trimmed_messages(self) -&amp;gt; list[dict]:&lt;br&gt;
    recent = self.history[-(self.max_history_turns * 2):]&lt;br&gt;
    return [{"role": "system", "content": self.system_prompt}] + recent&lt;/p&gt;

&lt;p&gt;Multiplying by 2 and slicing keeps whole user→assistant exchanges intact, so you never truncate mid-conversation in a way the API can't parse.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Normalize your vectors before cosine similarity search&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This one is sneaky because it fails silently. If you're using FAISS's IndexFlatIP (inner product) to approximate cosine similarity, that approximation is only valid if your vectors are unit-normalized first. Skip it, and you don't get an error — you get search results ranked in the wrong order, with no signal that anything's broken.&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
@staticmethod&lt;br&gt;
def _normalize(vectors: np.ndarray) -&amp;gt; np.ndarray:&lt;br&gt;
    norms = np.linalg.norm(vectors, axis=1, keepdims=True)&lt;br&gt;
    return vectors / np.clip(norms, 1e-10, None)&lt;/p&gt;

&lt;p&gt;If your semantic search results look "close but weirdly off," this is one of the first things worth checking.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Only retry the errors that are actually transient&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Blanket "retry on any exception" logic is a trap. If a request fails because you sent a malformed parameter, retrying it four times just guarantees the same failure four times — burning latency and rate-limit budget for nothing.&lt;/p&gt;

&lt;p&gt;Scope retries to the errors that can plausibly resolve themselves:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
@retry(&lt;br&gt;
    stop=stop_after_attempt(4),&lt;br&gt;
    wait=wait_exponential(multiplier=1, min=1, max=20),&lt;br&gt;
    retry=retry_if_exception_type((RateLimitError, APITimeoutError, APIError)),&lt;br&gt;
    reraise=True,&lt;br&gt;
)&lt;br&gt;
def _call_model(...):&lt;br&gt;
    ...&lt;/p&gt;

&lt;p&gt;Rate limits and timeouts are worth retrying with backoff. Your own bugs are not — those should fail fast so you actually see them.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Give yourself a fallback model, not just a retry loop&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Retries handle transient failures. They don't help if a model is degraded or down for an extended window. Pairing retries with a fallback to a secondary model is the same pattern most production LLM gateways use under the hood:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
try:&lt;br&gt;
    return _call_model(primary_model, messages)&lt;br&gt;
except RETRYABLE_ERRORS:&lt;br&gt;
    return _call_model(fallback_model, messages)&lt;/p&gt;

&lt;p&gt;Cheap/fast model first, stronger model as a safety net. Your app degrades gracefully instead of just erroring out.&lt;/p&gt;

&lt;p&gt;Where these came from&lt;/p&gt;

&lt;p&gt;I pulled these six lessons out of The AI &amp;amp; LLM Integration Cookbook — a set of 10 complete, copy-pasteable Python templates covering both the OpenAI and Anthropic APIs (quick-starts, tool calling, RAG over PDFs with LangChain and LlamaIndex, embeddings/vector search, and the production wrapper above). Each recipe includes a "why this pattern" note like the ones above, so you're not just getting code — you're getting the reasoning that usually only shows up after something's already broken in production once.&lt;/p&gt;

&lt;p&gt;Worth a look if you're wiring up your first LLM feature or auditing an existing integration for exactly these kinds of gaps.&lt;/p&gt;

&lt;p&gt;Get the AI and LLM cookbook at 20% off&lt;br&gt;
product link- &lt;a href="https://payhip.com/besttechbooks" rel="noopener noreferrer"&gt;https://payhip.com/besttechbooks&lt;/a&gt;&lt;/p&gt;

</description>
      <category>programming</category>
      <category>api</category>
      <category>learning</category>
      <category>development</category>
    </item>
    <item>
      <title>Most "production-ready" LLM code I see in tutorials isn't production-ready. Here's what's usually missing.</title>
      <dc:creator>zunairah</dc:creator>
      <pubDate>Sat, 12 Sep 2026 10:43:40 +0000</pubDate>
      <link>https://dev.to/zunairah_bfe3d030a9be261c/most-production-ready-llm-code-i-see-in-tutorials-isnt-production-ready-heres-whats-usually-2m1</link>
      <guid>https://dev.to/zunairah_bfe3d030a9be261c/most-production-ready-llm-code-i-see-in-tutorials-isnt-production-ready-heres-whats-usually-2m1</guid>
      <description>&lt;p&gt;I spent the last few weeks pulling together 10 Python patterns I keep rebuilding on every LLM project, and three mistakes showed up over and over — including in my own early code.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Mixing sync and streaming behind a flag.&lt;br&gt;
It's tempting to write one ask() function with an if stream: branch. Don't. Callers of the blocking version want a plain string they can log, cache, and unit test. Callers of the streaming version want a generator built for a UI. Collapsing both into one function with a boolean flag means every caller has to know which mode they're in — and testing gets messy fast. Two small functions beat one clever one.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Retrying every failure the same way.&lt;br&gt;
Wrapping an API call in a retry loop feels like "production hardening" — until you realize you're retrying a malformed request four times with exponential backoff instead of failing fast. The fix is boring but important: only retry transient errors (rate limits, timeouts), and let everything else surface immediately. Pair that with a fallback model (cheap model first, stronger model if it keeps failing), and you've got the pattern most real LLM gateways actually use in production.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Re-embedding your entire document library on every restart.&lt;br&gt;
This one's a silent cost killer. In RAG demos, it's common to load PDFs, chunk them, embed them, and query — all in one script, every single run. In production, you build the index once, persist it to disk, and load it on startup. Skipping this step is the single most common reason RAG demos rack up huge embedding bills and never make it past week one.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of these are exotic. They're just the difference between "code that works when I run it" and "code that doesn't wake me up at 2am."&lt;/p&gt;

&lt;p&gt;I ended up writing these patterns down properly — full runnable files, not fragments, covering both the OpenAI and Anthropic SDKs: streaming wrappers, tool/function calling, dynamic system prompts, RAG with LangChain and LlamaIndex, conversational memory with context trimming, vector search from scratch with FAISS, and the retry/fallback wrapper above.&lt;/p&gt;

&lt;p&gt;Put them together into The AI &amp;amp; LLM Integration Cookbook — 10 copy-paste-adapt templates, each with a short "why this pattern" note so you're not just copying code, you understand the tradeoff behind it.&lt;/p&gt;

&lt;p&gt;If you're tired of rebuilding the same LLM plumbing from scratch on every project, it might save you a weekend. Link in the comments — happy to answer questions about any of the patterns above in the meantime.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>python</category>
    </item>
  </channel>
</rss>
