<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Salmon Joy</title>
    <description>The latest articles on DEV Community by Salmon Joy (@salmonjoy).</description>
    <link>https://dev.to/salmonjoy</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4166308%2Fa607206d-8f31-4a30-a0de-630ec1ef677c.png</url>
      <title>DEV Community: Salmon Joy</title>
      <link>https://dev.to/salmonjoy</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/salmonjoy"/>
    <language>en</language>
    <item>
      <title>Structured outputs as application contracts</title>
      <dc:creator>Salmon Joy</dc:creator>
      <pubDate>Wed, 07 Oct 2026 05:10:32 +0000</pubDate>
      <link>https://dev.to/salmonjoy/structured-outputs-as-application-contracts-34gd</link>
      <guid>https://dev.to/salmonjoy/structured-outputs-as-application-contracts-34gd</guid>
      <description>&lt;h1&gt;
  
  
  Your LLM Output Should Have a Contract, Not a Vibe
&lt;/h1&gt;

&lt;p&gt;So there I was, an AI engineer at a fintech startup, trying to turn a GPT‑4 chat into a credit‑score suggestion engine. The user typed “I just got a promotion, how much should I loan?” and the model spat back a paragraph that looked like poetry: “Congrats! Maybe $5‑10k could work, but keep your debt‑to‑income ratio low…” I tried to yank the numbers out with a regex like &lt;code&gt;/\$(\d+)-(\d+)k/&lt;/code&gt;. Spoiler: it broke the moment someone said “$5‑10 k” with a thin space or used “5‑10k” without the dollar sign. One mis‑typed hyphen and my downstream service crashed, mis‑classifying a safe applicant as risky. It felt like trying to catch a greased pig with a colander.&lt;/p&gt;

&lt;p&gt;That’s when I discovered the power of structured outputs as application contracts. Instead of hoping the model “gives me what I need”, I told it &lt;em&gt;exactly&lt;/em&gt; what shape the response must have – a JSON object that matches a JSON Schema and a typed TypeScript interface. The prompt became:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"loan_range"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"min"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"max"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;int&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"high"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"medium"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"low"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the model can’t accidentally sprinkle a poem in there; it either obeys or says “I’m sorry, I can’t comply”. That refusal is a feature: my code checks &lt;code&gt;if (response.refusal) …&lt;/code&gt; and falls back to a safe default.&lt;/p&gt;

&lt;p&gt;Why is this better than regex? Regex is a brittle detective that looks for patterns in free text. Anything outside the pattern – extra whitespace, emojis, a different phrasing – sends it into a dead end. A schema validator, on the other hand, parses a proper data structure and instantly tells you which fields are missing or of the wrong type. It’s like checking a passport instead of eyeballing a face.&lt;/p&gt;

&lt;p&gt;In my edutech side‑project, we needed a list of quiz questions with “question”, “options”, and “answer”. By versioning the schema (&lt;code&gt;"$schema": "http://json-schema.org/draft-07/schema#", "title": "QuizV2"&lt;/code&gt;), we could roll out a new field “explanation” without breaking older clients. Downstream safety checks – e.g., “no option string longer than 200 chars” – are baked into the schema, so the LLM can’t hand us a monster answer that blows up the UI.&lt;/p&gt;

&lt;p&gt;A quick case study: after switching to structured outputs and enabling refusals, our fintech loan‑approval pipeline’s error rate dropped from 12% to under 1%. The only time we got a refusal was when the user asked for advice on illegal activity, and we gracefully returned a polite error message.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; treat the LLM like a junior dev who follows a contract, not a poet who vibes with your mood. Define a JSON Schema, generate typed interfaces, validate, version, and let refusals be your safety net.&lt;/p&gt;

&lt;p&gt;Got a funny regex story or a schema win? Drop a comment – I’d love to hear how you’re taming your LLMs!&lt;/p&gt;

&lt;p&gt;If you are someone who loves to know the technical work and architecture design I have shared more details based on my experience on this here: &lt;a href="https://github.com/SalmonJoy/My_guide_for_building_AI_systems/blob/main/Structured_outputs_as_application_contracts.md" rel="noopener noreferrer"&gt;https://github.com/SalmonJoy/My_guide_for_building_AI_systems/blob/main/Structured_outputs_as_application_contracts.md&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>machinelearning</category>
      <category>enterprise</category>
    </item>
    <item>
      <title>Context engineering vs prompt engineering</title>
      <dc:creator>Salmon Joy</dc:creator>
      <pubDate>Tue, 06 Oct 2026 13:17:36 +0000</pubDate>
      <link>https://dev.to/salmonjoy/context-engineering-vs-prompt-engineering-1id</link>
      <guid>https://dev.to/salmonjoy/context-engineering-vs-prompt-engineering-1id</guid>
      <description>&lt;h1&gt;
  
  
  Prompt Engineering Is Too Small a Mental Model,&amp;nbsp;Think Context Architecture
&lt;/h1&gt;

&lt;p&gt;I was knee‑deep in a fintech app that tried to answer user questions about credit‑score changes. My first version was a classic “prompt engineering” hack:  &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“You are a friendly finance bot, explain the score drop in plain English.”  &lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It worked for the first few queries, then the answers got weird – sometimes the bot repeated the same disclaimer, other times it started spilling the policy docs verbatim. I was losing token budget fast and the model kept “remembering” the old disclaimer even after I removed it. I realized I was treating the whole prompt as a monolith, when in reality I was juggling five different pieces of context that should have their own lanes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enter context architecture
&lt;/h2&gt;

&lt;p&gt;Think of an LLM conversation like a pizza kitchen.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;System instructions&lt;/strong&gt; – the chef’s recipe book; they stay on the board all night.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User state&lt;/strong&gt; – the current order ticket (what the customer just asked, plus any previous choices).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieved knowledge&lt;/strong&gt; – the pantry shelf you pull ingredients from (a database lookup of the user’s credit history).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool results&lt;/strong&gt; – the side‑dish the sous‑chef prepares (e.g., a risk‑score API call).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Policy&lt;/strong&gt; – the health‑inspection sticker you slap on the box to make sure you never serve raw chicken.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The magic happens in precedence. The model first looks at &lt;strong&gt;system instructions&lt;/strong&gt;, then layers &lt;strong&gt;user state&lt;/strong&gt;, then &lt;strong&gt;fetched knowledge&lt;/strong&gt;, then &lt;strong&gt;tool results&lt;/strong&gt;, and finally &lt;strong&gt;policy&lt;/strong&gt; overrides anything that would break compliance. If you drop a policy clause at the bottom of a huge prompt, it might never be seen because the model runs out of tokens before it gets there. That’s why you always put the “stable‑prefix” (&lt;strong&gt;system + policy&lt;/strong&gt;) at the front – it’s cached, cheap, and guarantees the model never forgets the rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  Token budget = pizza dough budget
&lt;/h2&gt;

&lt;p&gt;You only have so many tokens per request; if you waste them on stale context (like a disclaimer you already sent), you run out of room for the juicy answer.  &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contamination risk&lt;/strong&gt; is when old user state leaks into a new conversation – imagine the kitchen forgetting to clear the ticket and adding yesterday’s pepperoni to today’s vegan pizza. To avoid that, we clear the user state each session or use a fresh context window.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stable‑prefix caching
&lt;/h2&gt;

&lt;p&gt;(OpenAI Prompt Caching) lets you store the &lt;strong&gt;system + policy&lt;/strong&gt; block once and reuse it across millions of calls. In my edutech project, we cached a 120‑token prefix that described the “tutor bot” persona and compliance rules. The cache cut latency by &lt;strong&gt;30 %&lt;/strong&gt; and saved us &lt;strong&gt;$2 K a month&lt;/strong&gt; on token spend.&lt;/p&gt;

&lt;h2&gt;
  
  
  Case study: Fintech fraud‑alert bot
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Piece&lt;/th&gt;
&lt;th&gt;Content&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;System&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;“You are a compliance‑aware fraud analyst.”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Policy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;“Never reveal raw transaction IDs.”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;User&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;“My card was declined twice.”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Retrieved&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;“Last 5 transactions: $23, $0, $0.”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tool&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;“Risk engine returns confidence 0.92.”&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The final prompt stayed under &lt;strong&gt;350 tokens&lt;/strong&gt;, the answer was spot‑on, and no policy leakage occurred.&lt;/p&gt;




&lt;p&gt;So the next time you’re tempted to dump a giant prompt into the model, pause and ask yourself:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What belongs in &lt;strong&gt;system instructions&lt;/strong&gt;?
&lt;/li&gt;
&lt;li&gt;What’s the &lt;strong&gt;user state&lt;/strong&gt;?
&lt;/li&gt;
&lt;li&gt;What do I need to &lt;strong&gt;fetch&lt;/strong&gt;?
&lt;/li&gt;
&lt;li&gt;What &lt;strong&gt;tool results&lt;/strong&gt; am I adding?
&lt;/li&gt;
&lt;li&gt;What &lt;strong&gt;policy&lt;/strong&gt; must sit at the very front?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Build that context architecture, and you’ll spend less time chasing bugs and more time building cool features.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What weird context mishap have you run into?&lt;/strong&gt; Drop a comment and let’s swap stories!&lt;/p&gt;

&lt;p&gt;If you are someone who loves to know the technical work and architecture design I have shared more details based on my experience on this here:&lt;br&gt;&lt;br&gt;
&lt;a href="https://github.com/SalmonJoy/My_guide_for_building_AI_systems/blob/main/Context_engineering_vs_prompt_engineering.md" rel="noopener noreferrer"&gt;https://github.com/SalmonJoy/My_guide_for_building_AI_systems/blob/main/Context_engineering_vs_prompt_engineering.md&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>machinelearning</category>
      <category>enterprise</category>
    </item>
    <item>
      <title>Model selection is an architecture decision</title>
      <dc:creator>Salmon Joy</dc:creator>
      <pubDate>Tue, 06 Oct 2026 12:51:02 +0000</pubDate>
      <link>https://dev.to/salmonjoy/model-selection-is-an-architecture-decision-2k05</link>
      <guid>https://dev.to/salmonjoy/model-selection-is-an-architecture-decision-2k05</guid>
      <description>&lt;h1&gt;
  
  
  Stop Asking “Which LLM Is Best?” – Ask Which Model Fits This Workload
&lt;/h1&gt;

&lt;p&gt;Last quarter I was elbow‑deep in a fintech product that had to turn noisy bank statements into clean expense reports. My first instinct was to grab the “biggest” LLM, fire‑off GPT‑4, Claude‑2, Gemini‑1.5 and hope it would magically understand every column, currency and weird abbreviation. After a weekend of latency checks, cost tallies and exploding error logs I realized the “best” model on every public leaderboard still missed our real KPI: &lt;strong&gt;cost per successful parsing&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The fix was to stop treating LLMs like shiny toys and start treating them as replaceable services. I listed the dimensions that actually mattered:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;capability&lt;/strong&gt; (financial jargon, tables)
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;reasoning&lt;/strong&gt; (multi‑step extraction without hallucination)
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;latency&lt;/strong&gt; (&amp;lt;300 ms for a smooth UI)
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;context length&lt;/strong&gt; (statements can be 10 k tokens)
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;modality&lt;/strong&gt; (sometimes we get a PDF image, so OCR + generation matters)
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;cost&lt;/strong&gt; (every API call eats our margin)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then I built a tiny benchmark that mirrors the exact workflow: feed a real statement, ask for JSON line items, measure success (JSON parses, totals match). I ran it against three services: &lt;strong&gt;OpenAI gpt‑3.5‑turbo‑16k&lt;/strong&gt;, &lt;strong&gt;AWS Bedrock Claude‑3 Opus&lt;/strong&gt;, and &lt;strong&gt;Google Gemini‑1.5‑Flash&lt;/strong&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Accuracy&lt;/th&gt;
&lt;th&gt;Cost / 1k‑tok&lt;/th&gt;
&lt;th&gt;Latency&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;gpt‑3.5‑turbo&lt;/td&gt;
&lt;td&gt;92 %&lt;/td&gt;
&lt;td&gt;$0.012&lt;/td&gt;
&lt;td&gt;420 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude‑3 Opus&lt;/td&gt;
&lt;td&gt;95 %&lt;/td&gt;
&lt;td&gt;$0.018&lt;/td&gt;
&lt;td&gt;280 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini‑Flash&lt;/td&gt;
&lt;td&gt;88 %&lt;/td&gt;
&lt;td&gt;$0.006&lt;/td&gt;
&lt;td&gt;150 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Cost‑per‑successful‑task&lt;/strong&gt; = (cost per call) ÷ (success rate).  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPT‑3.5: $0.013
&lt;/li&gt;
&lt;li&gt;Claude‑Opus: $0.019
&lt;/li&gt;
&lt;li&gt;Gemini‑Flash: $0.007
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Suddenly Gemini‑Flash became the clear winner for our budget‑tight pipeline, even though its raw accuracy looked lower on paper.&lt;/p&gt;

&lt;p&gt;The trick is to &lt;strong&gt;version‑track the exact model you use&lt;/strong&gt; (e.g., &lt;code&gt;gpt-3.5-turbo-0613&lt;/code&gt;) and re‑run the benchmark whenever a new upgrade lands. In a later ed‑tech project I swapped Claude‑2 for Claude‑3‑Sonnet after a single test showed a &lt;strong&gt;30 % latency drop&lt;/strong&gt; and a &lt;strong&gt;5 % reasoning bump&lt;/strong&gt;, shaving &lt;strong&gt;$0.003 per student quiz&lt;/strong&gt; from the cost‑per‑task metric.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;So the next time someone asks “Which LLM is best?” answer with a checklist and a cost‑per‑successful‑task chart. It turns brag‑rights into real engineering decisions.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;What’s the weirdest metric you’ve used to pick a model?&lt;/strong&gt; Drop a comment, I’m curious about your war stories!&lt;/p&gt;

&lt;p&gt;If you are someone who loves to know the technical work and architecture design I have shared details based on my experience on this here:&lt;br&gt;&lt;br&gt;
&lt;a href="https://github.com/SalmonJoy/My_guide_for_building_AI_systems/blob/main/Model_selection_is_an_architecture_decision.md" rel="noopener noreferrer"&gt;https://github.com/SalmonJoy/My_guide_for_building_AI_systems/blob/main/Model_selection_is_an_architecture_decision.md&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>machinelearning</category>
      <category>enterprise</category>
    </item>
  </channel>
</rss>
