<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: galian</title>
    <description>The latest articles on DEV Community by galian (@galian).</description>
    <link>https://dev.to/galian</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3827330%2F5a53ab61-2fc1-4072-a44e-873913dd8cd7.png</url>
      <title>DEV Community: galian</title>
      <link>https://dev.to/galian</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/galian"/>
    <language>en</language>
    <item>
      <title>The EU AI Act's Transparency Rules Went Live — Here's What You Actually Have to Ship</title>
      <dc:creator>galian</dc:creator>
      <pubDate>Sun, 09 Aug 2026 21:42:39 +0000</pubDate>
      <link>https://dev.to/cursuri-ai/the-eu-ai-acts-transparency-rules-went-live-heres-what-you-actually-have-to-ship-3i2d</link>
      <guid>https://dev.to/cursuri-ai/the-eu-ai-acts-transparency-rules-went-live-heres-what-you-actually-have-to-ship-3i2d</guid>
      <description>&lt;p&gt;Most AI regulation, for most developers, has been something that happens to other people. Risk classifications, conformity assessments, notified bodies — the kind of thing where you nod, assume legal will handle it, and go back to your streaming handler.&lt;/p&gt;

&lt;p&gt;Article 50 is different, and it went live on &lt;strong&gt;2 August 2026&lt;/strong&gt;. It's the part of the EU AI Act that doesn't care whether your system is "high risk." It cares about one thing: &lt;strong&gt;can the person on the other end tell that this is AI?&lt;/strong&gt; If the answer is no, you owe them a disclosure — and the disclosure is a product decision, a UI decision, and in one case a file-format decision. All three land on engineering.&lt;/p&gt;

&lt;p&gt;The Commission adopted its guidelines on Article 50 on &lt;strong&gt;20 July 2026&lt;/strong&gt;, and the AI Office published a &lt;strong&gt;Code of Practice on Transparency of AI-Generated Content&lt;/strong&gt; on &lt;strong&gt;10 June 2026&lt;/strong&gt;. Between those two documents and the article text itself, the shape of what you have to build is now reasonably clear. This is that shape, translated into work items.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcursuri-ai.ro%2Fimages%2Fblog%2Feu-ai-act-article-50-en.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcursuri-ai.ro%2Fimages%2Fblog%2Feu-ai-act-article-50-en.svg" alt="The four Article 50 transparency obligations mapped to what each one requires in a product" width="1200" height="660"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  First: does this apply to you?
&lt;/h2&gt;

&lt;p&gt;Two questions, and you're probably in scope on both.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are you a provider or a deployer?&lt;/strong&gt; The Act splits duties. A &lt;strong&gt;provider&lt;/strong&gt; develops an AI system (or has one developed) and places it on the market under its own name — if you built the chatbot or the image generator, that's you. A &lt;strong&gt;deployer&lt;/strong&gt; uses an AI system under its own authority — if you dropped someone else's model into your support widget, that's you. Article 50 assigns two obligations to providers and two to deployers, and plenty of teams are both.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does the EU reach you?&lt;/strong&gt; The Act's territorial hooks (Article 2) are not "are you an EU company." They're closer to "is the system placed on the EU market, or is its output used in the EU." A US startup with EU users is in scope. This is the GDPR pattern, and it caught a lot of people by surprise the first time.&lt;/p&gt;

&lt;p&gt;Not in scope: things that don't interact with people or generate content for them. The Commission's guidelines explicitly put spam filters and automated translation tools &lt;em&gt;outside&lt;/em&gt; Article 50(1), and voice assistants and chatbots &lt;em&gt;inside&lt;/em&gt; it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Obligation 1 — tell people they're talking to an AI
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;Providers shall ensure that AI systems intended to interact directly with natural persons are designed and developed in such a way that the natural persons concerned are informed that they are interacting with an AI system, unless this is obvious from the point of view of a natural person who is reasonably well-informed, observant and circumspect. — Article 50(1)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the one that hits the most products, and it's the easiest to get wrong in a way that looks compliant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The "obvious" carve-out is narrower than you want it to be.&lt;/strong&gt; The test isn't "our users are technical." It's a reasonably well-informed, observant and circumspect person &lt;em&gt;in the circumstances and context of use&lt;/em&gt;. A widget labelled "AI Assistant" in a developer tool is plausibly obvious. The same engine answering an inbound phone call in a warm, human-sounding voice is not — and voice is exactly where the gap is widest right now, because &lt;a href="https://cursuri-ai.ro/en/courses/voice-ai-and-realtime-multimodal-agents" rel="noopener noreferrer"&gt;realtime voice agents&lt;/a&gt; are good enough that the "obvious" defence has quietly stopped being true.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where the disclosure has to live.&lt;/strong&gt; Article 50(5) settles the argument you're about to have with someone in a planning meeting: the information must be provided &lt;strong&gt;at the latest at the time of the first interaction or exposure&lt;/strong&gt;, and it must be &lt;strong&gt;clear and distinguishable&lt;/strong&gt;. That rules out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;burying it in the Terms of Service&lt;/li&gt;
&lt;li&gt;a footnote in the privacy policy&lt;/li&gt;
&lt;li&gt;a tooltip behind a hover on a mobile UI&lt;/li&gt;
&lt;li&gt;disclosing on turn three, after the user has already asked something personal&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Article 50(5) also requires conformity with applicable accessibility requirements — which in practice means your disclosure has to survive a screen reader, not just a design review.&lt;/p&gt;

&lt;p&gt;What that looks like in a chat surface is boring, and boring is the point:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="c"&gt;&amp;lt;!-- Rendered before the first assistant message, not after it. --&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;div&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"ai-disclosure"&lt;/span&gt; &lt;span class="na"&gt;role=&lt;/span&gt;&lt;span class="s"&gt;"note"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;strong&amp;gt;&lt;/span&gt;You're chatting with an AI assistant.&lt;span class="nt"&gt;&amp;lt;/strong&amp;gt;&lt;/span&gt;
  Answers are generated automatically.
  &lt;span class="nt"&gt;&amp;lt;a&lt;/span&gt; &lt;span class="na"&gt;href=&lt;/span&gt;&lt;span class="s"&gt;"/support/human"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;Talk to a person&lt;span class="nt"&gt;&amp;lt;/a&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/div&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For voice, the equivalent is a spoken line in the &lt;strong&gt;first&lt;/strong&gt; turn — before you collect anything, and in the language of the call. For an API you sell to other developers, the honest move is to pass the obligation downstream explicitly: document it, and give integrators a disclosure string they can render, because when they ship your model to end users under their own brand, the deployer duties become theirs and the design duty stays yours.&lt;/p&gt;

&lt;p&gt;One more thing worth building while you're in there: a &lt;strong&gt;handoff path to a human&lt;/strong&gt;. Article 50 doesn't mandate it. But the disclosure lands very differently when it's followed by an escape hatch, and support teams that ship &lt;a href="https://cursuri-ai.ro/en/courses/ai-for-customer-support-and-service" rel="noopener noreferrer"&gt;AI assistants without one&lt;/a&gt; tend to discover the reason the hard way.&lt;/p&gt;

&lt;h2&gt;
  
  
  Obligation 2 — mark synthetic output so machines can detect it
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;Providers of AI systems […] generating synthetic audio, image, video or text content, shall ensure the outputs […] are marked in a machine-readable format and detectable as artificially generated or manipulated. — Article 50(2)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the genuinely hard engineering item, and it's the one with a different deadline (more on that below).&lt;/p&gt;

&lt;p&gt;Two properties are required, and they're not the same thing:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Machine-readable marking&lt;/strong&gt; — metadata that a downstream system can parse.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Detectability&lt;/strong&gt; — the output can be recognised as artificially generated or manipulated.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The Act asks for solutions that are "effective, interoperable, robust and reliable as far as this is technically feasible" — a standard that explicitly bends to the state of the art. It also carves out &lt;strong&gt;assistive editing functions&lt;/strong&gt; and systems that &lt;strong&gt;do not substantially alter the input data&lt;/strong&gt;. Your auto-crop and your denoise filter are not in scope. Your "generate a product photo from this prompt" endpoint is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the Code of Practice points at.&lt;/strong&gt; The AI Office's Code, published 10 June 2026, describes a layered approach rather than a single mechanism: signed, timestamped provenance metadata — where &lt;strong&gt;C2PA&lt;/strong&gt; is the standard identified as meeting those criteria — &lt;em&gt;plus&lt;/em&gt; an imperceptible watermark embedded in the content itself, robust enough to survive ordinary transformations like compression, cropping, scaling and format conversion. The Code is voluntary and, at the time of writing, going through an adequacy assessment by the Commission and the AI Board. Adhering to it is a route to demonstrating compliance; not adhering means you have to show equivalently adequate means of your own.&lt;/p&gt;

&lt;p&gt;In practice, for images and video, that means attaching &lt;strong&gt;C2PA Content Credentials&lt;/strong&gt; at generation time. The Content Authenticity Initiative ships open-source tooling for this — &lt;a href="https://github.com/contentauth/c2pa-rs" rel="noopener noreferrer"&gt;&lt;code&gt;c2pa-rs&lt;/code&gt;&lt;/a&gt; with Python, JS, C++, Swift and Android bindings, plus a &lt;code&gt;c2patool&lt;/code&gt; CLI — so this is a library integration, not a research project.&lt;/p&gt;

&lt;p&gt;The assertion that carries "this was AI-generated" is the IPTC digital source type, referenced inside a &lt;code&gt;c2pa.actions&lt;/code&gt; assertion:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"assertions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"label"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"c2pa.actions"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"data"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"actions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"c2pa.created"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"digitalSourceType"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="s2"&gt;"http://cv.iptc.org/newscodes/digitalsourcetype/trainedAlgorithmicMedia"&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;trainedAlgorithmicMedia&lt;/code&gt; is the IPTC code for content created by a generative model. There are neighbouring codes for composites and for algorithmically edited media — pick the one that actually describes what your pipeline did, because "created" on a system that only retouched is its own kind of wrong. Verify the current assertion shape against the C2PA spec before you ship; the standard is still moving.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three things nobody tells you in the compliance deck:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Metadata gets stripped.&lt;/strong&gt; Plenty of platforms re-encode uploads and discard provenance metadata on the way in. Signing at generation is necessary; assuming it survives the internet is not. This is precisely why the Code pairs metadata with a watermark instead of trusting either alone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Text is the weak link.&lt;/strong&gt; Machine-readable marking of &lt;em&gt;text&lt;/em&gt; has no equivalent of C2PA that works after a copy-paste. Statistical watermarking of token distributions exists, degrades under paraphrase, and doesn't survive a user retyping the paragraph. The Act's "as far as technically feasible" language is doing real work here — but "hard" is not "exempt," and documenting your reasoning is part of the deliverable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sign server-side.&lt;/strong&gt; Any marking applied in the browser is marking a determined user can skip. The signature belongs on the generation path, before the bytes reach a client.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're generating &lt;a href="https://cursuri-ai.ro/en/courses/ai-image-generation" rel="noopener noreferrer"&gt;images&lt;/a&gt; or video in a product today, this obligation is now a line item in your media pipeline, not a policy question.&lt;/p&gt;

&lt;h2&gt;
  
  
  Obligation 3 — emotion recognition and biometric categorisation
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;Deployers of an emotion recognition system or a biometric categorisation system shall inform the natural persons exposed thereto of the operation of the system, and shall process the personal data in accordance with [the GDPR and related instruments]. — Article 50(3)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Shorter, and mostly a matter of knowing that it applies to you. If your product infers emotional state from voice, face or text, or sorts people into categories from biometric data, you inform the people exposed to it — and you're squarely in GDPR territory on top, usually with special-category data.&lt;/p&gt;

&lt;p&gt;Before you scope the disclosure: check Article 5 first. Some emotion recognition — in the &lt;strong&gt;workplace&lt;/strong&gt; and in &lt;strong&gt;education&lt;/strong&gt; — is &lt;em&gt;prohibited&lt;/em&gt; outright, not merely subject to transparency, and has been since February 2025. Article 50 is the wrong chapter to be reading if that's your use case.&lt;/p&gt;

&lt;h2&gt;
  
  
  Obligation 4 — deepfakes and public-interest text
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;Deployers of an AI system that generates or manipulates image, audio or video content constituting a deep fake, shall disclose that the content has been artificially generated or manipulated. — Article 50(4)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Note the split from Obligation 2: marking the file is the &lt;strong&gt;provider's&lt;/strong&gt; duty; disclosing to the audience is the &lt;strong&gt;deployer's&lt;/strong&gt;. If you use a third-party model to produce a synthetic spokesperson for a campaign, the vendor's C2PA manifest doesn't discharge your obligation. You still have to tell the audience.&lt;/p&gt;

&lt;p&gt;Two carve-outs matter:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Artistic and satirical works.&lt;/strong&gt; Where the content is part of an evidently artistic, creative, satirical or fictional work, the disclosure shrinks to revealing the existence of generated content &lt;strong&gt;in an appropriate manner that does not hamper the display or enjoyment of the work&lt;/strong&gt;. A film doesn't need a permanent banner across the frame.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Text on matters of public interest.&lt;/strong&gt; AI-generated or manipulated &lt;em&gt;text&lt;/em&gt; published to inform the public on matters of public interest must be disclosed — &lt;strong&gt;unless&lt;/strong&gt; it has undergone human review or editorial control and a natural or legal person holds editorial responsibility. This is the clause every content-heavy site should read twice. An unreviewed AI-written news summary needs a label. The same article with a named editor who checked it and owns it does not. If your publishing workflow can't currently prove which of those two happened, that's the actual gap — and it's a workflow problem before it's a legal one, which is why &lt;a href="https://cursuri-ai.ro/en/courses/ai-for-content-creation-and-copywriting" rel="noopener noreferrer"&gt;content operations built on AI&lt;/a&gt; now need an audit trail as much as a style guide.&lt;/p&gt;

&lt;h2&gt;
  
  
  The deadlines, which are not all the same
&lt;/h2&gt;

&lt;p&gt;The Digital Omnibus on AI — published in the Official Journal on 24 July 2026 and in force since 27 July — shifted several AI Act dates. Article 50 came out of it mostly intact:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What&lt;/th&gt;
&lt;th&gt;Applies from&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Article 50(1) AI-interaction disclosure&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2 August 2026&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Article 50(3) emotion recognition / biometric categorisation&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2 August 2026&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Article 50(4) deepfake and public-interest text disclosure&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2 August 2026&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Article 50(2) marking, systems placed on the market &lt;strong&gt;on or after&lt;/strong&gt; 2 Aug 2026&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2 August 2026&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Article 50(2) marking, systems placed on the market &lt;strong&gt;before&lt;/strong&gt; 2 Aug 2026&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2 December 2026&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Annex III high-risk obligations (Chapter III)&lt;/td&gt;
&lt;td&gt;deferred to &lt;strong&gt;2 December 2027&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;High-risk AI embedded in regulated products (Annex I)&lt;/td&gt;
&lt;td&gt;deferred to &lt;strong&gt;2 August 2028&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last pair is the source of most of the confusion in the room right now: the &lt;em&gt;high-risk&lt;/em&gt; regime got a long deferral, and a lot of teams heard "the AI Act got pushed back" and stopped reading. Article 50 did not get pushed back. The only grace period is the marking obligation for generative systems that were already on the market, and it expires on 2 December 2026.&lt;/p&gt;

&lt;p&gt;Enforcement sits with national market surveillance authorities, the AI Office and — for EU institutions — the European Data Protection Supervisor. Breaching Article 50 carries fines of up to &lt;strong&gt;€15 million or 3% of worldwide annual turnover&lt;/strong&gt;, whichever is higher.&lt;/p&gt;

&lt;h2&gt;
  
  
  A ship checklist
&lt;/h2&gt;

&lt;p&gt;Pin this to the epic:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;[ ] Inventory every surface where an AI system talks to a person — chat, voice, email autoresponders, in-app agents. Each one gets a first-interaction disclosure or a written argument for why it's obvious.&lt;/li&gt;
&lt;li&gt;[ ] Disclosure rendered &lt;strong&gt;before&lt;/strong&gt; the first AI output, accessible, in the user's language, not in the ToS.&lt;/li&gt;
&lt;li&gt;[ ] Every generative endpoint identified as provider-side or deployer-side. Write it down; vendor contracts should say the same thing.&lt;/li&gt;
&lt;li&gt;[ ] C2PA Content Credentials signed &lt;strong&gt;server-side&lt;/strong&gt; on image/audio/video generation, with the correct IPTC &lt;code&gt;digitalSourceType&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;[ ] Watermarking assessed for each modality; where it isn't feasible, the reasoning is documented rather than assumed.&lt;/li&gt;
&lt;li&gt;[ ] Deepfake disclosure at the &lt;strong&gt;publication&lt;/strong&gt; surface, not just in the file metadata.&lt;/li&gt;
&lt;li&gt;[ ] Editorial-review provenance recorded for AI-assisted public-interest text — who reviewed, when, who owns it.&lt;/li&gt;
&lt;li&gt;[ ] Emotion recognition / biometric categorisation checked against &lt;strong&gt;Article 5 prohibitions&lt;/strong&gt; before anything else.&lt;/li&gt;
&lt;li&gt;[ ] Disclosure copy and placement covered by a test, so the next redesign doesn't silently delete it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The thing that makes Article 50 unusual, as regulation goes, is how little of it is paperwork. There's no conformity assessment here, no technical documentation dossier, no notified body. There's a banner that has to render before the first message, a signature that has to happen on the generation path, a label that has to reach the audience, and a record of who reviewed what. Four engineering tickets, roughly, and none of them are hard.&lt;/p&gt;

&lt;p&gt;They're just easy to defer — and the deferral is what gets expensive. Article 50 is now enforceable, the guidelines are published, the Code of Practice exists, and the tooling for the hard part is open source. Compliance here is mostly a question of whether someone put it in the sprint.&lt;/p&gt;

&lt;p&gt;The teams that will have the least trouble with this aren't the ones with the biggest legal department. They're the ones that were already willing to tell users, plainly, what the machine was doing. Turns out that was always the good product decision — it just became the required one.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written by the team behind &lt;a href="https://cursuri-ai.ro/en/" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt;, an AI education platform with hands-on English-language courses on &lt;a href="https://cursuri-ai.ro/en/courses/ai-data-privacy-and-eu-ai-act-compliance" rel="noopener noreferrer"&gt;AI, data privacy and EU AI Act compliance&lt;/a&gt;, production LLM integration, and shipping AI products.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources &amp;amp; further reading:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;EU AI Act — &lt;a href="https://artificialintelligenceact.eu/article/50/" rel="noopener noreferrer"&gt;Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems&lt;/a&gt; (full text, exemptions, codes of practice)&lt;/li&gt;
&lt;li&gt;European Commission — &lt;a href="https://digital-strategy.ec.europa.eu/en/policies/guidelines-transparency-ai-generated-content" rel="noopener noreferrer"&gt;Guidelines on transparency obligations for providers and deployers of certain AI systems&lt;/a&gt; (adopted 20 July 2026)&lt;/li&gt;
&lt;li&gt;European Commission — &lt;a href="https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content" rel="noopener noreferrer"&gt;Code of Practice on Transparency of AI-generated Content&lt;/a&gt; (published 10 June 2026)&lt;/li&gt;
&lt;li&gt;Content Authenticity Initiative — &lt;a href="https://opensource.contentauthenticity.org/docs/introduction/" rel="noopener noreferrer"&gt;open-source C2PA SDKs and &lt;code&gt;c2patool&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;IPTC — &lt;a href="https://cv.iptc.org/newscodes/digitalsourcetype/" rel="noopener noreferrer"&gt;Digital Source Type NewsCodes vocabulary&lt;/a&gt; (&lt;code&gt;trainedAlgorithmicMedia&lt;/code&gt; and related values)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This article is educational content written by engineers, not legal advice. Article 50 interacts with the GDPR, the DSA, national implementing rules and sector regulation, and the Digital Omnibus changed several dates in 2026 — verify against current official sources and your own counsel before shipping.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Claude Opus 5 Is Out — Here Are the API Changes That Will Actually Break Your Code</title>
      <dc:creator>galian</dc:creator>
      <pubDate>Tue, 28 Jul 2026 22:38:55 +0000</pubDate>
      <link>https://dev.to/cursuri-ai/claude-opus-5-is-out-here-are-the-api-changes-that-will-actually-break-your-code-312k</link>
      <guid>https://dev.to/cursuri-ai/claude-opus-5-is-out-here-are-the-api-changes-that-will-actually-break-your-code-312k</guid>
      <description>&lt;p&gt;Anthropic shipped &lt;strong&gt;Claude Opus 5&lt;/strong&gt; on July 24, 2026, and if you build on the Claude API there's a version of this launch story you can safely skip: the benchmark charts. The version you can't skip is the API contract, because for the first time in a while an Opus release changes how existing requests behave — and one parameter combination that was perfectly valid on Opus 4.8 now returns a &lt;strong&gt;400 error&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I write and teach about AI engineering at &lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt;, Eastern Europe's AI education platform, and this post is the writeup I wish every model launch came with: not "how smart is it," but "what do I need to change in my code, in what order, and what silently behaves differently if I change nothing." Everything below comes from Anthropic's official "What's new in Claude Opus 5" documentation — no leaked numbers, no vibes.&lt;/p&gt;

&lt;p&gt;Quick disclaimer: this space moves monthly and this is a launch-window snapshot. Verify against the &lt;a href="https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5" rel="noopener noreferrer"&gt;official docs&lt;/a&gt; before you wire anything to production.&lt;/p&gt;

&lt;h2&gt;
  
  
  The spec sheet, in one table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Spec&lt;/th&gt;
&lt;th&gt;Claude Opus 5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model ID&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;claude-opus-5&lt;/code&gt; (&lt;code&gt;anthropic.claude-opus-5&lt;/code&gt; on Bedrock)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context window&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1M tokens — both default and maximum&lt;/strong&gt; (no smaller variant)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max output&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;128k tokens&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thinking&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;On by default&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pricing&lt;/td&gt;
&lt;td&gt;$5 / $25 per million input/output tokens — unchanged from Opus 4.8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fast mode&lt;/td&gt;
&lt;td&gt;$10 / $50, research preview, &lt;strong&gt;Claude API only&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two of these are bigger deals than they look. The 1M context window isn't an opt-in variant — it's just what the model has, and the docs specifically claim consistent instruction following, tool calling, and reasoning throughout the window. And 128k max output makes single-request long deliverables (big refactors, full reports) realistic — with a caveat about &lt;code&gt;max_tokens&lt;/code&gt; we'll get to, because it's now doing more work than it used to.&lt;/p&gt;

&lt;p&gt;Opus 4.8 stays available on every platform, so nothing forces a same-day migration. But the changes below are the kind you want to understand &lt;em&gt;before&lt;/em&gt; your first "why is this request failing" incident, not after.&lt;/p&gt;

&lt;h2&gt;
  
  
  Change #1: thinking is on by default
&lt;/h2&gt;

&lt;p&gt;On Opus 4.8, a request without a &lt;code&gt;thinking&lt;/code&gt; field ran &lt;strong&gt;without&lt;/strong&gt; extended thinking. You opted in with &lt;code&gt;thinking: {"type": "adaptive"}&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;On Opus 5, that same field-less request now runs &lt;strong&gt;with thinking on&lt;/strong&gt;. The model decides when and how much to think on each turn, and the &lt;code&gt;effort&lt;/code&gt; parameter is the knob that controls thinking depth. If your code already sends &lt;code&gt;thinking: {"type": "adaptive"}&lt;/code&gt;, you're fine — that value remains valid and is equivalent to the new default.&lt;/p&gt;

&lt;p&gt;Why this can bite you even though it sounds like a free upgrade:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;max_tokens&lt;/code&gt; is a hard limit on &lt;em&gt;total&lt;/em&gt; output — thinking plus response text.&lt;/strong&gt; A workload that ran happily with &lt;code&gt;max_tokens: 2000&lt;/code&gt; on Opus 4.8 (no thinking, short answers) can now spend a chunk of that budget on reasoning before it writes a single visible token. The official guidance is explicit: revisit &lt;code&gt;max_tokens&lt;/code&gt; for every workload that previously ran without thinking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Change #2: the actual breaking change — disabled thinking + high effort = 400
&lt;/h2&gt;

&lt;p&gt;Here's the one to grep your codebase for today:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;thinking: {"type": "disabled"}&lt;/code&gt; is accepted &lt;strong&gt;only when effort is &lt;code&gt;high&lt;/code&gt; or below&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;thinking: {"type": "disabled"}&lt;/code&gt; combined with effort &lt;code&gt;xhigh&lt;/code&gt; or &lt;code&gt;max&lt;/code&gt; returns a &lt;strong&gt;400 error&lt;/strong&gt;, on every request.&lt;/li&gt;
&lt;li&gt;This is generally available behavior (not beta) from Opus 5 onward — and it's a breaking change from Opus 4.8, where disabling thinking was independent of the effort level.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you currently run thinking-disabled at high effort levels, you have exactly two exits: keep thinking disabled and drop effort to &lt;code&gt;high&lt;/code&gt; or below, or keep your effort level and &lt;strong&gt;delete the &lt;code&gt;thinking&lt;/code&gt; field entirely&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;There's a second, subtler reason to just let thinking stay on. The docs note that with thinking disabled, Opus 5 can occasionally write a tool call &lt;strong&gt;into its text output&lt;/strong&gt; instead of emitting a proper &lt;code&gt;tool_use&lt;/code&gt; block, or leak internal XML tags into the visible response. If you parse tool calls in production, that's not a cosmetic footnote — it's a failure mode your handler needs to survive. Anthropic's recommendation is to keep thinking enabled and control cost with lower effort levels instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  Change #3: &lt;code&gt;effort&lt;/code&gt; is now the control lever that matters
&lt;/h2&gt;

&lt;p&gt;With thinking on by default, &lt;code&gt;effort&lt;/code&gt; becomes the central parameter of your integration. The full ladder on Opus 5: &lt;code&gt;low&lt;/code&gt;, &lt;code&gt;medium&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt; (the default), &lt;code&gt;xhigh&lt;/code&gt;, and &lt;code&gt;max&lt;/code&gt;. The docs make a claim worth taking seriously: Opus 5 converts additional effort into better results &lt;em&gt;more reliably than any earlier Opus model&lt;/em&gt; — which means the level you pick carries more weight than it did on 4.8. At the other end, &lt;code&gt;low&lt;/code&gt; and &lt;code&gt;medium&lt;/code&gt; are explicitly called out for producing strong quality at a fraction of the tokens and latency.&lt;/p&gt;

&lt;p&gt;A request with everything turned up looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;64000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;output_config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;effort&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_final_message&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three deliberate details in that snippet: there's &lt;strong&gt;no &lt;code&gt;thinking&lt;/code&gt; field&lt;/strong&gt; (the default is what you want), &lt;code&gt;max_tokens&lt;/code&gt; is large (at &lt;code&gt;xhigh&lt;/code&gt;/&lt;code&gt;max&lt;/code&gt; the model needs room to think and act across tool calls), and it's &lt;strong&gt;streamed&lt;/strong&gt; — at 64k-token budgets, non-streaming requests can hit the time limit.&lt;/p&gt;

&lt;p&gt;The sane strategy is the documented one: start at the default &lt;code&gt;high&lt;/code&gt;, then adjust based on &lt;strong&gt;your own evals&lt;/strong&gt;, not vibes. Step down where quality holds — you pocket the tokens and latency. Step up to &lt;code&gt;xhigh&lt;/code&gt;/&lt;code&gt;max&lt;/code&gt; for the genuinely hard work. If you don't have an eval harness that can answer "did quality hold when I dropped effort?", that's the single highest-leverage thing to build this quarter — it's the discipline we treat as foundational in our &lt;a href="https://cursuri-ai.ro/courses/ai-evals-llm-productie" rel="noopener noreferrer"&gt;course on LLM evals in production&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Change #4: the model &lt;em&gt;behaves&lt;/em&gt; differently, even if you change nothing
&lt;/h2&gt;

&lt;p&gt;The docs have a refreshingly honest section on differences you'll notice without touching your code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Default responses run longer&lt;/strong&gt; — both user-facing answers and written deliverables. If your product has strict length constraints, enforce them in the prompt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;In agentic sessions, the model narrates its progress more often.&lt;/strong&gt; Good for UX transparency; tune it down where you want silence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;In multi-agent frameworks, it delegates to subagents more readily&lt;/strong&gt; — budget accordingly, subagents are tokens too.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It verifies its own work without being told to.&lt;/strong&gt; This is the one that demands action: verification instructions inherited from older prompts — "include a final verification step," "use a subagent to verify" — should be &lt;strong&gt;removed&lt;/strong&gt;, because on Opus 5 they cause over-verification. You pay twice, in tokens and latency, for work the model already does.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is a pattern every migration teaches: prompts get calibrated against the previous model's weaknesses, and those calibrations become friction on the next model. A model migration without a prompt audit is half a migration. That workflow — prompts, tools, verification, on real repos — is exactly what we drill in our &lt;a href="https://cursuri-ai.ro/courses/claude-code-mastery-coding-agentic" rel="noopener noreferrer"&gt;Claude Code and agentic coding course&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The smaller changes you'll actually be happy about
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prompt caching minimum drops to 512 tokens&lt;/strong&gt;, from 1,024 on Opus 4.8. System prompts that were too short to cache start caching with zero code changes. If you run compact prompts at volume, this shows up on your invoice by itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mid-conversation tool changes (beta)&lt;/strong&gt;: you can add or remove tools between turns &lt;strong&gt;while preserving the prompt cache&lt;/strong&gt;, with the &lt;code&gt;mid-conversation-tool-changes-2026-07-01&lt;/code&gt; beta header. For phase-based agents (explore → edit → verify, different tools each), this removes a whole architectural compromise — no more front-loading every tool "just in case."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A &lt;code&gt;"default"&lt;/code&gt; mode for fallbacks (beta)&lt;/strong&gt;: the &lt;code&gt;fallbacks&lt;/code&gt; parameter can now apply Anthropic's recommended fallback models by refusal category, instead of a hand-maintained model list. Use the &lt;code&gt;server-side-fallback-2026-07-01&lt;/code&gt; header (the older &lt;code&gt;2026-06-01&lt;/code&gt; header only accepts explicit lists).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fast mode exists but is API-only&lt;/strong&gt;: $10/$50 per million tokens, research preview, not currently on Bedrock, Google Cloud, or Microsoft Foundry. If you route through a cloud provider and were counting on it, adjust your plan.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The migration checklist, in the order things break
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Grep for the fatal combination&lt;/strong&gt;: &lt;code&gt;thinking: {"type": "disabled"}&lt;/code&gt; anywhere near effort &lt;code&gt;xhigh&lt;/code&gt;/&lt;code&gt;max&lt;/code&gt; → that's a 400 now. Decide per integration: drop the &lt;code&gt;thinking&lt;/code&gt; field, or lower effort.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Revisit &lt;code&gt;max_tokens&lt;/code&gt;&lt;/strong&gt; on every request that ran without thinking on 4.8 — the budget now covers reasoning + response.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Swap the model ID&lt;/strong&gt; to &lt;code&gt;claude-opus-5&lt;/code&gt; — from a config variable, not a hardcoded string. (Model retirements turn hardcoded IDs into production 404s; ask me how I know.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit your prompts for verification instructions&lt;/strong&gt; and remove them — over-verification is pure waste on Opus 5.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Turn on streaming&lt;/strong&gt; for anything with a large &lt;code&gt;max_tokens&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you must run thinking-disabled&lt;/strong&gt;, make your parser survive tool calls in text and stray XML tags.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-run your evals before shifting traffic.&lt;/strong&gt; "Better on average" is not "better on your tasks."&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If your architecture makes step 3 scary — model IDs scattered across services, no config layer, no eval gate — that's a structural problem worth fixing once, properly. Going from "API calls sprinkled through the codebase" to "an application where a model swap is a config change" is the arc of our &lt;a href="https://cursuri-ai.ro/courses/construire-aplicatii-ai-python-sdk" rel="noopener noreferrer"&gt;building AI applications with Python course&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bottom line
&lt;/h2&gt;

&lt;p&gt;Opus 5 doesn't ask you to learn a new API — it asks you to &lt;strong&gt;re-check the assumptions&lt;/strong&gt; your existing code was built on. Thinking is on unless you say otherwise, &lt;code&gt;effort&lt;/code&gt; is the lever that matters, &lt;code&gt;max_tokens&lt;/code&gt; now pays for cognition too, and one previously-valid parameter combo is a hard error. In exchange you get a 1M-token window as the default, 128k output, cheaper caching, and betas that clean up real architectural pain.&lt;/p&gt;

&lt;p&gt;Teams that treat this as a checklist migration — rather than a find-and-replace on the model name — will come out with integrations that are cheaper and more reliable than what they had. And that discipline, built once, pays out again on every launch that follows.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If you want structured, hands-on training on any of this — evals, agentic coding, production LLM apps — that's what we build at &lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>javascript</category>
      <category>opus</category>
    </item>
    <item>
      <title>Multi-Agent AI Systems: When One Agent Isn't Enough — and When Two Is Already Too Many</title>
      <dc:creator>galian</dc:creator>
      <pubDate>Wed, 22 Jul 2026 20:54:34 +0000</pubDate>
      <link>https://dev.to/cursuri-ai/multi-agent-ai-systems-when-one-agent-isnt-enough-and-when-two-is-already-too-many-38kj</link>
      <guid>https://dev.to/cursuri-ai/multi-agent-ai-systems-when-one-agent-isnt-enough-and-when-two-is-already-too-many-38kj</guid>
      <description>&lt;p&gt;Somewhere around the third time your agent's context window filled up with search results it no longer needed, you had the thought every agent builder eventually has: &lt;em&gt;what if I split this into multiple agents?&lt;/em&gt; One to research, one to write, one to review. Maybe a "manager" agent on top. It sounds like an org chart, and org charts feel like architecture.&lt;/p&gt;

&lt;p&gt;Sometimes that instinct is exactly right — Anthropic's research system, built as an orchestrator with parallel subagents, outperformed its best single-agent setup by a wide margin on internal evals. And sometimes it's exactly wrong — Cognition (the team behind Devin) published a piece bluntly titled &lt;a href="https://cognition.ai/blog/dont-build-multi-agents" rel="noopener noreferrer"&gt;"Don't Build Multi-Agents"&lt;/a&gt;, arguing that for their domain, splitting the work is how you manufacture inconsistency. Both teams are right, because they're solving different problems.&lt;/p&gt;

&lt;p&gt;This article is the decision framework between those two positions: what multi-agent architectures actually buy you, what they cost, the orchestration pattern that works in production, and the failure modes that don't make it into the launch posts.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a multi-agent system actually is
&lt;/h2&gt;

&lt;p&gt;Strip the buzz and it's simple: a &lt;strong&gt;multi-agent system&lt;/strong&gt; is one where an LLM-driven agent can delegate work to &lt;em&gt;other&lt;/em&gt; LLM-driven agents, each running in its &lt;strong&gt;own context window&lt;/strong&gt;, and use their results. The common production shape is &lt;strong&gt;orchestrator–worker&lt;/strong&gt;: a lead agent owns the goal, decomposes it into subtasks, spawns subagents to execute them (often in parallel), and synthesizes what comes back.&lt;/p&gt;

&lt;p&gt;That "own context window" clause is the entire point. It's not about the anthropomorphic org chart — it's about &lt;strong&gt;context isolation&lt;/strong&gt;. Everything else follows from it.&lt;/p&gt;

&lt;p&gt;Notice what this is &lt;em&gt;not&lt;/em&gt;: a pipeline of prompts you call in sequence from your own code is not a multi-agent system — it's just a program (and often, that's exactly what you should build instead). The multi-agent label earns its keep when the &lt;em&gt;decomposition itself is dynamic&lt;/em&gt; — when an agent decides at runtime what to delegate, to whom, and with what instructions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three things fan-out actually buys you
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Context isolation.&lt;/strong&gt; A subagent that greps a codebase or reads twenty web pages burns thousands of tokens on intermediate noise — raw file contents, search results, dead ends. In a single-agent design, all of that lands in the one context window you have, where it dilutes attention for every subsequent step. In an orchestrator–worker design, the noise lives and &lt;em&gt;dies&lt;/em&gt; in the subagent's window; the orchestrator receives a compressed conclusion. You've effectively multiplied your usable context by the number of workers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Parallelism.&lt;/strong&gt; Independent subtasks — "evaluate these five libraries," "check each of these twelve services for the deprecated call" — can run concurrently instead of serially. For research-shaped and audit-shaped work, this is a wall-clock improvement measured in multiples, not percent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Separation of concerns.&lt;/strong&gt; Each subagent gets a narrow prompt, a curated toolset, and one job. A focused agent with six relevant tools reliably beats a generalist juggling twenty — the same reason you curate tool sets aggressively in &lt;a href="https://cursuri-ai.ro/en/courses/context-engineering-and-memory-for-ai-agents" rel="noopener noreferrer"&gt;context engineering for agents&lt;/a&gt;, applied at the architecture level.&lt;/p&gt;

&lt;p&gt;Anthropic's engineering write-up on &lt;a href="https://www.anthropic.com/engineering/built-multi-agent-research-system" rel="noopener noreferrer"&gt;their multi-agent research system&lt;/a&gt; put numbers on this: the orchestrator-with-parallel-subagents architecture outperformed their single-agent baseline by &lt;strong&gt;90.2%&lt;/strong&gt; on an internal research eval, and they found that token usage — how much &lt;em&gt;relevant&lt;/em&gt; thinking-and-reading the system could do — explained most of the performance variance. Fan-out is, at bottom, a way to spend more tokens productively on one problem than a single context window physically allows.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs you
&lt;/h2&gt;

&lt;p&gt;The same write-up is refreshingly honest about the bill: in their data, a single agent used roughly &lt;strong&gt;4× more tokens than a chat interaction, and multi-agent systems roughly 15×&lt;/strong&gt;. That's the entry fee. Multi-agent architectures only make economic sense when the task's value clears it.&lt;/p&gt;

&lt;p&gt;The subtler cost is &lt;strong&gt;coordination&lt;/strong&gt;. The orchestrator communicates with subagents through &lt;em&gt;task descriptions&lt;/em&gt; — and every ambiguity in those descriptions becomes a subagent doing confidently the wrong thing in a context you can't see. Early versions of Anthropic's system had subagents duplicating each other's work, wandering off-scope, and searching for things other subagents had already found — not because the model was weak, but because the delegation instructions were vague. The fix was unglamorous prompt engineering: each delegation needs an explicit objective, output format, tool guidance, and &lt;em&gt;boundaries&lt;/em&gt; where the subtask ends.&lt;/p&gt;

&lt;h2&gt;
  
  
  The case against — and it's a strong one
&lt;/h2&gt;

&lt;p&gt;Cognition's counter-argument deserves to be taken seriously, because it identifies the precise condition under which multi-agent designs fail: &lt;strong&gt;when subtasks are not actually independent.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Their core principles: share full context and trace history, and remember that &lt;em&gt;actions carry implicit decisions&lt;/em&gt;. When two subagents work in parallel from partial views of the task, each makes small implicit decisions — naming, structure, interpretation of an ambiguous requirement. When their outputs meet, those decisions conflict. For coding work, where part A and part B must compile, link, and agree on conventions, parallel agents with partial context produce exactly the inconsistency you'd expect from two contractors who never spoke to each other.&lt;/p&gt;

&lt;p&gt;So the honest synthesis of the two positions isn't "Anthropic vs. Cognition" — it's a property of the &lt;em&gt;task&lt;/em&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Read-heavy, decomposable, breadth-first&lt;/strong&gt; work (research, audits, evaluation sweeps, codebase exploration) parallelizes beautifully. Subagent outputs are &lt;em&gt;findings&lt;/em&gt; that an orchestrator can merge, and conflicts between them are cheap to resolve.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write-heavy, interdependent, convention-laden&lt;/strong&gt; work (feature implementation, refactoring, anything where outputs must cohere) punishes fan-out. Here a single agent with a well-managed context — compaction, external memory, just-in-time retrieval — is the stronger architecture.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A useful mnemonic: &lt;strong&gt;fan out to read, stay single to write.&lt;/strong&gt; It's not absolute (isolated write tasks in separate worktrees parallelize fine), but it's the right default, and knowing when to break it is precisely the judgment that &lt;a href="https://cursuri-ai.ro/en/courses/ai-agents-architecture-and-automation" rel="noopener noreferrer"&gt;a serious course on AI agent architecture&lt;/a&gt; spends real time building.&lt;/p&gt;

&lt;h2&gt;
  
  
  The orchestration patterns that survive production
&lt;/h2&gt;

&lt;p&gt;If your task passes the test above, these are the shapes that hold up:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Orchestrator–worker (the default).&lt;/strong&gt; One lead agent decomposes, delegates, synthesizes. Workers don't talk to each other — results flow through the orchestrator. This keeps the communication topology a star, not a mesh, which is the difference between debuggable and not. Resist the peer-to-peer "agents chatting with agents" fantasy; in production it mostly produces token-expensive misunderstandings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pipeline with isolated stages.&lt;/strong&gt; Each item flows through fixed stages (find → verify → summarize), each stage its own agent call with its own clean context. No barrier between items — item A can be in stage 3 while item B is in stage 1. Deterministic control flow in &lt;em&gt;your&lt;/em&gt; code, LLM judgment inside each stage. This hybrid — code decides the structure, agents fill the slots — is the most underrated pattern in the space, and it's where most "do we even need a framework?" questions dissolve.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Adversarial verification.&lt;/strong&gt; Spawn N independent agents to &lt;em&gt;refute&lt;/em&gt; a finding rather than confirm it, and keep only what survives. This is fan-out used not for speed but for &lt;strong&gt;confidence&lt;/strong&gt; — independent contexts mean independent mistakes, which is exactly what majority voting needs to work. It's the cheapest defense against the plausible-but-wrong output that a single agent will happily double down on.&lt;/p&gt;

&lt;p&gt;Three implementation rules that pay for themselves regardless of pattern:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Delegate like you're writing a ticket for a contractor.&lt;/strong&gt; Objective, output format, tools to use, scope boundaries. If a competent human would need to ask a clarifying question, your subagent will instead guess — silently.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Return data, not prose.&lt;/strong&gt; Subagents should report structured results (schema-validated if your stack allows it), because the orchestrator is a program consuming output, not a reader enjoying it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Match effort to value.&lt;/strong&gt; Simple lookups get one worker with a small budget; open-ended research earns a fleet. An orchestrator that spawns ten subagents for a question one search would answer is the multi-agent version of the microservices monorepo with three users.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The failure modes nobody demos
&lt;/h2&gt;

&lt;p&gt;Every one of these is a thing that happens &lt;em&gt;after&lt;/em&gt; the demo works:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Duplicated work.&lt;/strong&gt; Two subagents independently discover the same fact at full token price, because the orchestrator's task descriptions didn't partition the space. Partition explicitly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compounding errors.&lt;/strong&gt; In a pipeline, a stage-one hallucination becomes stage-two's trusted input. Multi-agent systems don't just propagate errors — they &lt;em&gt;launder&lt;/em&gt; them: by the final synthesis, the wrong fact arrives with the confidence of having been "processed" three times. This is why verification stages belong &lt;em&gt;early&lt;/em&gt;, not only at the end.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The telephone game.&lt;/strong&gt; Every orchestrator-mediated hop compresses information. Three hops in, "the test fails intermittently on CI under load" has become "the tests are broken." Keep hierarchies shallow — one level of delegation covers almost every real use case.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lost partial work.&lt;/strong&gt; Ten parallel subagents, one crashes at minute eight. If your harness can't resume without re-running the other nine, you'll pay the 15× token bill more than once. Checkpointing and resumability are boring, and they are the difference between a system and a demo.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And the meta-failure that enables all of the above: &lt;strong&gt;not measuring.&lt;/strong&gt; A multi-agent system has more knobs than anything else you'll build with LLMs — decomposition granularity, worker count, delegation prompts, synthesis strategy — and every knob is a chance to regress silently. The teams that ship these systems treat &lt;a href="https://cursuri-ai.ro/en/courses/llm-evaluation-and-testing" rel="noopener noreferrer"&gt;evaluation as first-class infrastructure&lt;/a&gt;: a representative task set, end-to-end outcome scoring (with an LLM judge where rubrics allow), and a re-run on every architectural change. "It seemed better on the three queries I tried" is how 15× token bills get approved for 1× results.&lt;/p&gt;

&lt;h2&gt;
  
  
  A decision checklist
&lt;/h2&gt;

&lt;p&gt;Before you split one agent into several, put the task through this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;If yes →&lt;/th&gt;
&lt;th&gt;If no →&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Can subtasks run with &lt;strong&gt;genuinely independent context&lt;/strong&gt;?&lt;/td&gt;
&lt;td&gt;fan-out is viable&lt;/td&gt;
&lt;td&gt;stay single-agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Is the work &lt;strong&gt;read/analyze-heavy&lt;/strong&gt; rather than write/produce-heavy?&lt;/td&gt;
&lt;td&gt;fan-out helps&lt;/td&gt;
&lt;td&gt;single agent + context engineering&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Does the task &lt;strong&gt;value clear a ~15× token cost&lt;/strong&gt;?&lt;/td&gt;
&lt;td&gt;proceed&lt;/td&gt;
&lt;td&gt;simplify the architecture&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Would a fixed pipeline in code + single agent calls do it?&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;build that instead&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;orchestrator–worker&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Do you have evals to catch coordination regressions?&lt;/td&gt;
&lt;td&gt;ship&lt;/td&gt;
&lt;td&gt;build evals first&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Note how often the honest answer is the boring one. A single agent with disciplined context management — compaction, external memory, sub-tasking only when context demands it — is the right architecture for most of what gets built today, and wiring either variant into a real product with retries, budgets, and observability is &lt;a href="https://cursuri-ai.ro/en/courses/advanced-llm-integration-in-production" rel="noopener noreferrer"&gt;its own production discipline&lt;/a&gt; on top of the architecture choice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Multi-agent systems aren't the next tier above single agents — they're a specialized tool that trades tokens and coordination complexity for context capacity and parallelism. That trade is spectacular for breadth-first, read-heavy, decomposable work, and actively harmful for interdependent, convention-heavy production of artifacts. Fan out to read; stay single to write; let deterministic code own the control flow either way; and put an eval behind every knob.&lt;/p&gt;

&lt;p&gt;The teams that get this right aren't the ones with the most agents. They're the ones who can articulate, for their specific task, what a second context window &lt;em&gt;buys&lt;/em&gt; — and who kept the architecture exactly one notch more complex than that answer requires.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build and teach this material at &lt;a href="https://cursuri-ai.ro/en/" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt;, an AI education platform with hands-on English-language tracks on AI agent architecture, context engineering, LLM evaluation, and production integration.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources &amp;amp; further reading:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Anthropic — &lt;a href="https://www.anthropic.com/engineering/built-multi-agent-research-system" rel="noopener noreferrer"&gt;How we built our multi-agent research system&lt;/a&gt; (orchestrator–worker pattern, 90.2% eval result, token-cost data, delegation lessons)&lt;/li&gt;
&lt;li&gt;Anthropic — &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;Building effective agents&lt;/a&gt; (workflows vs. agents, orchestrator–workers, "simplest thing that works")&lt;/li&gt;
&lt;li&gt;Cognition — &lt;a href="https://cognition.ai/blog/dont-build-multi-agents" rel="noopener noreferrer"&gt;Don't Build Multi-Agents&lt;/a&gt; (context-sharing principles, the case for single-threaded agents)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This article is educational content. Architectures, model capabilities, and costs evolve quickly — validate against your own workloads and current documentation.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>We just launched our AI course catalog in English — 20 hands-on courses with an AI professor in every lesson</title>
      <dc:creator>galian</dc:creator>
      <pubDate>Fri, 17 Jul 2026 07:21:16 +0000</pubDate>
      <link>https://dev.to/galian/we-just-launched-our-ai-course-catalog-in-english-20-hands-on-courses-with-an-ai-professor-in-46ip</link>
      <guid>https://dev.to/galian/we-just-launched-our-ai-course-catalog-in-english-20-hands-on-courses-with-an-ai-professor-in-46ip</guid>
      <description>&lt;p&gt;I'm the founder of &lt;a href="https://cursuri-ai.ro/en" rel="noopener noreferrer"&gt;Cursuri AI&lt;/a&gt;, an AI e-learning platform for professionals. "Cursuri" is simply Romanian for "courses" — the platform started in Romania, where it now runs a catalog of 50 AI courses in Romanian, from prompt engineering fundamentals to building production AI agents.&lt;/p&gt;

&lt;p&gt;This week we shipped the thing people kept asking for: &lt;strong&gt;the platform is now available in English&lt;/strong&gt;, with an international catalog of 20 courses built for people who want to actually &lt;em&gt;use&lt;/em&gt; AI at work — not just read about it.&lt;/p&gt;

&lt;p&gt;This post is the announcement, but I also want to be concrete about what the platform does and what's in the catalog, so you can decide in two minutes whether it's for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  The idea: learning AI should feel like working with AI
&lt;/h2&gt;

&lt;p&gt;Most online courses are passive: you watch or read, you nod along, and a week later you remember 10% of it. That model is especially broken for AI, where the whole point is &lt;em&gt;interaction&lt;/em&gt; — prompting, iterating, getting feedback, correcting course.&lt;/p&gt;

&lt;p&gt;So we built the platform around active learning instead:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;An AI professor in every lesson.&lt;/strong&gt; Not a generic chatbot bolted onto the side — an assistant that knows exactly which lesson you're in and answers questions in that context. You can type or just talk to it: it works in &lt;strong&gt;text and voice&lt;/strong&gt;. Stuck on why your RAG retrieval returns garbage, or what a "system prompt" actually is? Ask mid-lesson and get a precise answer. There's a full overview of how it works on the &lt;a href="https://cursuri-ai.ro/en/ai-professor" rel="noopener noreferrer"&gt;AI professor page&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interactive AI quizzes.&lt;/strong&gt; Questions adapt to your level, and every answer comes with a detailed explanation — including &lt;em&gt;why the wrong options are wrong&lt;/em&gt;, which is where most of the learning actually happens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hands-on exercises on the platform.&lt;/strong&gt; Coding challenges for the engineering track, realistic work scenarios for the business track. You apply things as you learn them, not "someday later."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automatic AI summaries.&lt;/strong&gt; Key points extracted from every lesson, so reviewing before an interview or a meeting takes minutes, not hours.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A personal dashboard.&lt;/strong&gt; Progress, scores, and insights, so you can see whether you're actually getting better or just clicking "next."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And because AI moves absurdly fast, &lt;strong&gt;content is kept up to date&lt;/strong&gt; — courses get refreshed as the tools and models change, and new courses are added to the bundles at no extra cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's in the English catalog
&lt;/h2&gt;

&lt;p&gt;The 20 launch courses are split into two tracks, because "learning AI" means very different things for a backend engineer and a marketing manager.&lt;/p&gt;

&lt;h3&gt;
  
  
  IT &amp;amp; Engineering track (10 courses)
&lt;/h3&gt;

&lt;p&gt;This is the track I'm personally most excited about, because it covers the stack of skills that 2026-era AI engineering actually demands:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Introduction to AI Engineering&lt;/strong&gt; — the on-ramp if you're a developer new to the field&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The Complete Prompt Engineering Masterclass&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;RAG: Retrieval-Augmented Generation in Practice&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Advanced LLM Integration in Production Applications&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI Agents: Architecting and Automating Autonomous Systems&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context Engineering and Memory for AI Agents&lt;/strong&gt; — beyond prompting: managing what your agent knows and remembers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://cursuri-ai.ro/en/courses/claude-code-mastery-agentic-coding" rel="noopener noreferrer"&gt;Claude Code Mastery: Agentic Coding from the Terminal&lt;/a&gt;&lt;/strong&gt; — multi-file work, git workflows, CI, MCP&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;MCP (Model Context Protocol): Building Servers and Integrations&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cursor as a Pro: AI-Native IDE, Composer and Multi-Agent workflows&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vibe Coding: From Prompt to Application&lt;/strong&gt; with Lovable, v0, Bolt and Replit Agent&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you've been meaning to move from "I use ChatGPT sometimes" to "I ship AI features and agentic workflows," this track is a structured path through exactly that.&lt;/p&gt;

&lt;h3&gt;
  
  
  Business &amp;amp; Professionals track (10 courses)
&lt;/h3&gt;

&lt;p&gt;No code required — built for managers, marketers, founders, analysts, and office professionals:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;AI for Business Leaders&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Manager in the AI Era: Leading Your Team Through the AI Transformation&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI for Entrepreneurs and Startups: The Complete Guide&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI for Digital Marketing&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI for Content Creation and Copywriting&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI for Sales and CRM&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SEO and AEO/GEO in the AI Era&lt;/strong&gt; — optimizing for Google, AI Overviews and generative engines&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Microsoft 365 Copilot for Office Work: Role-Based Productivity&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No-Code Data Analysis with AI&lt;/strong&gt; — ChatGPT, Excel and SQL for non-programmers&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI Image Generation: The Complete Guide from Prompt to Publishing&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The full list with detailed curricula is on the &lt;a href="https://cursuri-ai.ro/en/courses" rel="noopener noreferrer"&gt;course catalog&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pricing — simple and honest
&lt;/h2&gt;

&lt;p&gt;We deliberately kept it simple, and every plan is &lt;strong&gt;billed monthly with cancel-anytime&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;th&gt;What you get&lt;/th&gt;
&lt;th&gt;Price&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Single course&lt;/td&gt;
&lt;td&gt;Any one course, full access + AI professor&lt;/td&gt;
&lt;td&gt;€49/month + VAT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Business&lt;/td&gt;
&lt;td&gt;All 10 Business &amp;amp; Professionals courses&lt;/td&gt;
&lt;td&gt;€199/month + VAT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IT Pro&lt;/td&gt;
&lt;td&gt;All 10 IT &amp;amp; Engineering courses&lt;/td&gt;
&lt;td&gt;€399/month + VAT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;All Access&lt;/td&gt;
&lt;td&gt;The entire international catalog, both tracks&lt;/td&gt;
&lt;td&gt;€499/month + VAT&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things worth calling out:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;You're not forced into a bundle.&lt;/strong&gt; If you only need the RAG course or the Copilot course, subscribe to just that one for €49/month and cancel when you're done.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;New courses are included.&lt;/strong&gt; Bundle subscribers get every new course we release, as it's released.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Full details and a plan comparison are on the &lt;a href="https://cursuri-ai.ro/en/pricing" rel="noopener noreferrer"&gt;pricing page&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this is for (and who it isn't for)
&lt;/h2&gt;

&lt;p&gt;It's for you if you're a professional who wants a structured, hands-on path to using AI in your actual job — whether that job is writing code, running a team, or growing a business. It works well for companies too: there's a dedicated track structure and team offering for organizations that want to upskill whole departments.&lt;/p&gt;

&lt;p&gt;It's probably &lt;em&gt;not&lt;/em&gt; for you if you want academic ML theory — we don't teach you to derive backpropagation. The focus is applied: tools, workflows, and skills you use the same week you learn them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ask
&lt;/h2&gt;

&lt;p&gt;If this sounds useful, take a look at &lt;a href="https://cursuri-ai.ro/en" rel="noopener noreferrer"&gt;cursuri-ai.ro/en&lt;/a&gt; — and if you check out a course, I'd genuinely love feedback. We're a team from Romania, competing on quality and depth of content, so every piece of honest criticism from this community makes the platform better.&lt;/p&gt;

&lt;p&gt;And if you have questions about the catalog, the AI professor, or where we're taking the platform next — ask away in the comments. I'll answer everything.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Fine-Tuning vs RAG vs Prompting: How to Actually Decide in 2026</title>
      <dc:creator>galian</dc:creator>
      <pubDate>Wed, 15 Jul 2026 11:49:01 +0000</pubDate>
      <link>https://dev.to/cursuri-ai/fine-tuning-vs-rag-vs-prompting-how-to-actually-decide-in-2026-11gj</link>
      <guid>https://dev.to/cursuri-ai/fine-tuning-vs-rag-vs-prompting-how-to-actually-decide-in-2026-11gj</guid>
      <description>&lt;p&gt;There's a predictable arc to most LLM projects. Something doesn't work, someone says "we should fine-tune it," a month disappears into dataset wrangling and GPU bills, and the model comes back... about as wrong as before — because the actual problem was that it never had the right facts in front of it. Fine-tuning was never going to fix that.&lt;/p&gt;

&lt;p&gt;The three techniques — &lt;strong&gt;prompting&lt;/strong&gt;, &lt;strong&gt;retrieval-augmented generation (RAG)&lt;/strong&gt;, and &lt;strong&gt;fine-tuning&lt;/strong&gt; — are not a ladder you climb from cheap to fancy. They solve &lt;em&gt;different problems&lt;/em&gt;, and choosing the wrong one is expensive in exactly the way that's hard to notice: it looks like progress while it burns weeks.&lt;/p&gt;

&lt;p&gt;This is a decision framework. Not "here's what each one is" — you can get that anywhere — but the concrete questions that tell you which one your problem actually needs, and the failure signatures that mean you picked wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one distinction that resolves most arguments
&lt;/h2&gt;

&lt;p&gt;Before any framework, internalize this split, because it settles 80% of the "should we fine-tune?" debates on its own:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;RAG changes what the model &lt;em&gt;knows&lt;/em&gt;.&lt;/strong&gt; It injects facts into the context at inference time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fine-tuning changes how the model &lt;em&gt;behaves&lt;/em&gt;.&lt;/strong&gt; It adjusts the weights to shift style, format, structure, and task-specific skill.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompting changes what the model is &lt;em&gt;told to do&lt;/em&gt; right now&lt;/strong&gt;, using the knowledge and behavior it already has.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the first question is never "which technique is best?" It's "&lt;strong&gt;is my problem a knowledge gap or a behavior gap?&lt;/strong&gt;"&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The model gives outdated, made-up, or "I don't have access to that" answers about &lt;em&gt;your&lt;/em&gt; data → &lt;strong&gt;knowledge gap → RAG.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;The model knows the facts but won't reliably produce the &lt;em&gt;format, tone, or task structure&lt;/em&gt; you need → &lt;strong&gt;behavior gap → fine-tuning (maybe).&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;You haven't seriously tried telling it clearly what to do yet → &lt;strong&gt;prompting, first, always.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Get this wrong and no amount of engineering saves you. Fine-tuning a model to "know your docs" is the classic error: you can bake a &lt;em&gt;few&lt;/em&gt; facts into weights, but they go stale the moment your docs change, you can't cite sources, and you've spent training compute to build a worse version of a lookup. Knowledge that changes belongs in retrieval, not in weights.&lt;/p&gt;

&lt;h2&gt;
  
  
  Always start with prompting (yes, even now)
&lt;/h2&gt;

&lt;p&gt;Prompting is not the beginner tier you graduate from. In 2026, with frontier models, a well-constructed prompt plus a few good examples solves a startling share of problems that teams &lt;em&gt;assume&lt;/em&gt; need training. It's the fastest, cheapest, most inspectable option, and it should be your baseline before you're allowed to say the word "fine-tune."&lt;/p&gt;

&lt;p&gt;Reach for prompting when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You're still discovering what "good output" even looks like. Prompts are editable in seconds; datasets are not.&lt;/li&gt;
&lt;li&gt;The task is reasoning, transformation, or generation the model already broadly knows how to do.&lt;/li&gt;
&lt;li&gt;You need to ship this week.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The techniques that make prompting punch above its weight are unglamorous but real: precise role and task framing, &lt;strong&gt;few-shot examples&lt;/strong&gt; that demonstrate the exact output shape, chain-of-thought for multi-step reasoning, and rigid output contracts (structured/JSON) so downstream code can trust the result. Most "the model can't do this" conclusions are actually "we asked badly" conclusions. Squeezing the ceiling out of prompting before spending on anything heavier is a discipline in itself — it's the whole point of a &lt;a href="https://cursuri-ai.ro/courses/prompt-engineering-masterclass" rel="noopener noreferrer"&gt;prompt engineering masterclass&lt;/a&gt;, and the ROI of getting it right first is enormous because everything downstream inherits a better baseline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The prompting ceiling — how you know you've hit it:&lt;/strong&gt; you've iterated seriously, added good examples, and the model &lt;em&gt;still&lt;/em&gt; fails — and the failure is either (a) it doesn't know facts it couldn't possibly know, or (b) it can't hold a consistent behavior across inputs no matter how you phrase the instruction. (a) points to RAG. (b) &lt;em&gt;might&lt;/em&gt; point to fine-tuning. Not before.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reach for RAG when the problem is knowledge
&lt;/h2&gt;

&lt;p&gt;RAG is the answer whenever the model needs to work with information it wasn't trained on: your internal documentation, a product catalog, last week's tickets, a knowledge base that updates daily, anything private or fresh.&lt;/p&gt;

&lt;p&gt;Choose RAG when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Answers must be &lt;strong&gt;grounded in a specific corpus&lt;/strong&gt; and you need to &lt;strong&gt;cite sources&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;The knowledge &lt;strong&gt;changes&lt;/strong&gt; — pricing, policies, docs, inventory. You update an index, not a model.&lt;/li&gt;
&lt;li&gt;Hallucination on facts is unacceptable and you need an audit trail of &lt;em&gt;where&lt;/em&gt; an answer came from.&lt;/li&gt;
&lt;li&gt;The knowledge base is large, or partly access-controlled per user.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The reason RAG beats fine-tuning for knowledge isn't subtle: updating a document store is trivial and instant; updating weights is a training run. RAG gives you freshness, provenance, and per-user access control for free — none of which fine-tuning can offer. When your facts have a shelf life, retrieval is the &lt;em&gt;only&lt;/em&gt; correct architecture, and building it well (chunking, hybrid search, re-ranking) is where the real engineering lives — the substance of a dedicated &lt;a href="https://cursuri-ai.ro/courses/rag-retrieval-augmented-generation" rel="noopener noreferrer"&gt;course on RAG and retrieval-augmented generation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RAG's own ceiling:&lt;/strong&gt; retrieval fixes &lt;em&gt;what the model knows&lt;/em&gt;, not &lt;em&gt;how it behaves&lt;/em&gt;. If your RAG answers are factually correct but come out in the wrong format, wrong tone, or don't follow your house style no matter how you prompt — that residual behavior gap is exactly where fine-tuning finally earns its place, &lt;em&gt;on top of&lt;/em&gt; RAG, not instead of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fine-tune when the problem is behavior — and only then
&lt;/h2&gt;

&lt;p&gt;Fine-tuning is the right tool, but for a narrower set of problems than its reputation suggests. It shines at teaching &lt;em&gt;consistent behavior&lt;/em&gt; that's hard to specify in a prompt: a very specific output structure, a domain's tone and terminology, a classification or extraction task where you have lots of labeled examples, or a skill the base model does clumsily.&lt;/p&gt;

&lt;p&gt;Legitimately reach for fine-tuning when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You need &lt;strong&gt;consistent style, format, or structure&lt;/strong&gt; at a level prompting can't hold across the full input distribution.&lt;/li&gt;
&lt;li&gt;You have a &lt;strong&gt;narrow, high-volume, well-defined task&lt;/strong&gt; (classification, extraction, a specific transformation) and enough quality labeled data.&lt;/li&gt;
&lt;li&gt;You want to &lt;strong&gt;bake in a behavior&lt;/strong&gt; so you can drop it from the prompt — shorter prompts, lower per-call cost, faster responses at scale.&lt;/li&gt;
&lt;li&gt;Latency or cost at scale matters and a smaller fine-tuned model can match a bigger prompted one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two things make modern fine-tuning far less scary than its reputation. First, &lt;strong&gt;you almost never do full fine-tuning&lt;/strong&gt; — parameter-efficient methods like &lt;strong&gt;LoRA/QLoRA&lt;/strong&gt; train a tiny set of adapter weights, cutting the compute and memory cost by orders of magnitude while getting most of the benefit. Second, the bottleneck is &lt;strong&gt;data quality, not model choice&lt;/strong&gt;: a few hundred to a few thousand &lt;em&gt;clean, consistent, representative&lt;/em&gt; examples beat a huge noisy pile every time. The hard part of fine-tuning was never running the training job — it's building the dataset, choosing PEFT trade-offs, and evaluating the result without fooling yourself, which is precisely the ground a &lt;a href="https://cursuri-ai.ro/courses/fine-tuning-modele-ai" rel="noopener noreferrer"&gt;fine-tuning course&lt;/a&gt; has to cover to be worth anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When fine-tuning is the &lt;em&gt;wrong&lt;/em&gt; answer — the red flags:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"We'll fine-tune it on our docs so it knows them." → No. That's RAG. Fine-tuned facts go stale and can't be cited.&lt;/li&gt;
&lt;li&gt;"We haven't really tried prompting." → Do that first; you may not need to train at all.&lt;/li&gt;
&lt;li&gt;"The requirements change weekly." → Fine-tuning bakes behavior in; if the target moves, you're re-training constantly. Keep it in the prompt until it stabilizes.&lt;/li&gt;
&lt;li&gt;"We have 40 examples." → Usually not enough for reliable behavior change; strong prompting with those 40 as few-shot examples will likely beat it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The combinations are the real answer
&lt;/h2&gt;

&lt;p&gt;Framing these as rivals is the beginner mistake. In production, the strongest systems &lt;strong&gt;combine&lt;/strong&gt; them, because they operate on different axes — knowledge, behavior, and instruction — and stack cleanly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;RAG + prompting&lt;/strong&gt; is the workhorse for most knowledge-grounded assistants: retrieve the right context, then a well-engineered prompt instructs the model to answer &lt;em&gt;only&lt;/em&gt; from it and cite sources. No training required.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fine-tuning + RAG&lt;/strong&gt; is the high end: fine-tune for the domain's &lt;em&gt;behavior&lt;/em&gt; (tone, format, task skill), and use RAG for the &lt;em&gt;facts&lt;/em&gt;. The model behaves exactly right &lt;em&gt;and&lt;/em&gt; stays current — behavior in the weights, knowledge in the index.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fine-tuning + prompting&lt;/strong&gt; collapses a long, brittle instruction into learned behavior, so your prompts get short and your inference gets cheaper.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Orchestrating these — deciding which layer owns which responsibility, and routing a request through retrieval, tools, and the model in the right order — is its own engineering discipline, and it's the core of a &lt;a href="https://cursuri-ai.ro/courses/advanced-llm-integration" rel="noopener noreferrer"&gt;course on advanced LLM integration&lt;/a&gt;. The mental model to keep: &lt;strong&gt;knowledge → retrieval, behavior → weights, instruction → prompt.&lt;/strong&gt; Put each requirement on the axis it actually lives on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The decision, in one pass
&lt;/h2&gt;

&lt;p&gt;Run your problem through this, in order. Stop at the first that fits:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Have you genuinely exhausted prompting&lt;/strong&gt; — clear instructions, good few-shot examples, structured output? If not → &lt;strong&gt;prompt.&lt;/strong&gt; (This is where most projects should still be.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is the failure a knowledge gap&lt;/strong&gt; — missing, stale, or private facts; needs citations? → &lt;strong&gt;RAG.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is the failure a behavior gap&lt;/strong&gt; — format/tone/task consistency the prompt can't hold, &lt;em&gt;and&lt;/em&gt; you have quality labeled data &lt;em&gt;and&lt;/em&gt; the target is stable? → &lt;strong&gt;fine-tune&lt;/strong&gt; (LoRA first).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is it both?&lt;/strong&gt; → &lt;strong&gt;RAG for the facts, fine-tuning for the behavior.&lt;/strong&gt; In that order.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And underneath all of it: &lt;strong&gt;you cannot make this decision without evaluation.&lt;/strong&gt; "It seems better" is not data. Before you choose, build a small eval set — representative inputs with known-good outputs — so you can measure whether prompting already clears the bar, whether RAG actually retrieves the right context, and whether a fine-tune moved the metric or just moved the failures around. Teams that skip this pick techniques by vibes and discover the mistake in production; teams that treat &lt;a href="https://cursuri-ai.ro/courses/ai-evals-llm-productie" rel="noopener noreferrer"&gt;evals as first-class&lt;/a&gt; make the cheap correct choice on purpose. The eval set is what turns this framework from an opinion into a decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The reason so many LLM projects stall isn't a shortage of technique — it's reaching for the wrong one and mistaking motion for progress. Fine-tuning a model to "learn facts," RAG-ing a problem that was really a bad prompt, or grinding on prompts when the model fundamentally lacks the data: each fails in a way that looks like effort.&lt;/p&gt;

&lt;p&gt;Anchor on the split and you'll rarely go wrong. &lt;strong&gt;Knowledge that changes → RAG. Behavior you can't prompt into place → fine-tuning. Everything else → prompt, and prompt well.&lt;/strong&gt; Start cheap, measure honestly, and add complexity only when an eval — not a hunch — tells you the current layer has topped out. The best architecture isn't the most sophisticated one; it's the one that puts each requirement on the axis where it actually belongs.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Sources &amp;amp; further reading:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Lewis et al. — &lt;a href="https://arxiv.org/abs/2005.11401" rel="noopener noreferrer"&gt;Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Hu et al. — &lt;a href="https://arxiv.org/abs/2106.09685" rel="noopener noreferrer"&gt;LoRA: Low-Rank Adaptation of Large Language Models&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Dettmers et al. — &lt;a href="https://arxiv.org/abs/2305.14314" rel="noopener noreferrer"&gt;QLoRA: Efficient Finetuning of Quantized LLMs&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This article is educational content. Models, tooling, and cost trade-offs evolve quickly; validate any approach against your own data and current provider documentation before committing to it in production.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>monitoring</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Vector Embeddings Explained: Build Semantic Search in Python</title>
      <dc:creator>galian</dc:creator>
      <pubDate>Mon, 13 Jul 2026 11:25:24 +0000</pubDate>
      <link>https://dev.to/cursuri-ai/vector-embeddings-explained-build-semantic-search-in-python-2og1</link>
      <guid>https://dev.to/cursuri-ai/vector-embeddings-explained-build-semantic-search-in-python-2og1</guid>
      <description>&lt;p&gt;Search for "reset my password" in a keyword-based system and a help article titled "How to recover your account credentials" won't match — not one word overlaps. Yet any human knows they mean the same thing. Closing that gap between &lt;em&gt;characters&lt;/em&gt; and &lt;em&gt;meaning&lt;/em&gt; is what &lt;strong&gt;vector embeddings&lt;/strong&gt; do, and they're the quiet engine behind semantic search, RAG, recommendation systems, and most of the "AI that understands you" experiences shipped since 2023.&lt;/p&gt;

&lt;p&gt;This is a practical guide. We'll cover what an embedding actually is, why cosine similarity is the operation you'll use constantly, and then build a real, working semantic search engine in Python — first with pure NumPy so you see the mechanics, then with the tools you'd actually reach for in production. By the end you'll have code that runs and a mental model that transfers to every embedding-powered feature you build next.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an embedding actually is
&lt;/h2&gt;

&lt;p&gt;An &lt;strong&gt;embedding&lt;/strong&gt; is a list of numbers — a vector — that represents a piece of content as a point in high-dimensional space. An embedding model (a neural network trained on enormous text corpora) reads your text and outputs, say, 384 or 1,536 floating-point numbers. The magic isn't the numbers themselves; it's the property the training instills: &lt;strong&gt;texts with similar meaning land close together in that space, and unrelated texts land far apart.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's the whole idea. "How do I reset my password?" and "I forgot my login credentials" produce vectors that sit near each other. "What's the weather in Cluj?" produces a vector off in a completely different region. The model has effectively turned &lt;em&gt;meaning&lt;/em&gt; into &lt;em&gt;geometry&lt;/em&gt; — and geometry is something a computer can measure with plain arithmetic.&lt;/p&gt;

&lt;p&gt;A few properties worth internalizing before we write code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Dimensionality is fixed per model.&lt;/strong&gt; A given model always outputs the same length (e.g. 384 for &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt;, 1,536 for OpenAI's &lt;code&gt;text-embedding-3-small&lt;/code&gt;). You can't mix vectors from different models — they live in different spaces.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The individual numbers are not interpretable.&lt;/strong&gt; Dimension 200 doesn't mean "formality" or "topic." Meaning is distributed across all dimensions at once. Don't try to read them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Distance is the entire point.&lt;/strong&gt; You almost never care about a vector's absolute position — only how close it is to &lt;em&gt;other&lt;/em&gt; vectors.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Cosine similarity: the one operation you'll use everywhere
&lt;/h2&gt;

&lt;p&gt;To ask "how similar are these two texts?", you compare their vectors. The near-universal choice for text embeddings is &lt;strong&gt;cosine similarity&lt;/strong&gt;: it measures the angle between two vectors, ignoring their magnitude.&lt;/p&gt;

&lt;p&gt;Why the angle and not, say, straight-line (Euclidean) distance? Because for text embeddings, &lt;em&gt;direction&lt;/em&gt; encodes meaning while &lt;em&gt;length&lt;/em&gt; often encodes uninteresting things like text length or token count. Two documents about the same topic point the same way even if one is a sentence and the other a paragraph. Cosine similarity captures exactly that.&lt;/p&gt;

&lt;p&gt;The formula is just the dot product of the two vectors divided by the product of their magnitudes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cos(θ) = (A · B) / (‖A‖ · ‖B‖)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It returns a value from &lt;strong&gt;-1 to 1&lt;/strong&gt;, though for most modern text embeddings you'll see results land in roughly the 0-to-1 range:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;~1.0&lt;/strong&gt; — nearly identical meaning&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;~0.5&lt;/strong&gt; — loosely related&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;~0.0&lt;/strong&gt; — unrelated&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That single number, computed against a corpus of stored vectors, &lt;em&gt;is&lt;/em&gt; semantic search. Everything else is optimization.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build a semantic search engine from scratch
&lt;/h2&gt;

&lt;p&gt;Let's make it concrete. We'll use &lt;a href="https://www.sbert.net/" rel="noopener noreferrer"&gt;&lt;code&gt;sentence-transformers&lt;/code&gt;&lt;/a&gt;, which runs a capable embedding model locally — no API key, no network calls, so you can run this offline right now.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;sentence-transformers numpy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Step 1 — Embed a corpus
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sentence_transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SentenceTransformer&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;

&lt;span class="c1"&gt;# A small, fast, widely used model. 384-dimensional output.
&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SentenceTransformer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;all-MiniLM-L6-v2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;documents&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;How to reset your account password&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Refund policy for annual subscriptions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Setting up two-factor authentication&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Our office hours and contact details&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;How to recover a locked account&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="c1"&gt;# Encode the whole corpus once. Shape: (5, 384)
&lt;/span&gt;&lt;span class="n"&gt;doc_embeddings&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc_embeddings&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# (5, 384)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;doc_embeddings&lt;/code&gt; array is your search index. In a real app you compute it &lt;strong&gt;once&lt;/strong&gt;, when a document is created or updated, and store it — never on every query.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2 — Cosine similarity in NumPy
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;cosine_similarity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ndarray&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ndarray&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ndarray&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Cosine similarity between vector `a` and each row of matrix `b`.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;a_norm&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;linalg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;b_norm&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;linalg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;axis&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;keepdims&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;b_norm&lt;/span&gt; &lt;span class="o"&gt;@&lt;/span&gt; &lt;span class="n"&gt;a_norm&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ten lines, no dependencies beyond NumPy. This is the actual core of semantic search — the rest is plumbing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3 — Search
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;query_vec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;scores&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;cosine_similarity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query_vec&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;doc_embeddings&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;ranked&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;argsort&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;)[::&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;][:&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;  &lt;span class="c1"&gt;# top-k, highest first
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[(&lt;/span&gt;&lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;ranked&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;I forgot my login credentials&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;  &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run it, and you'll get something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0.62  How to reset your account password
0.55  How to recover a locked account
0.31  Setting up two-factor authentication
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice what happened: the query &lt;strong&gt;"I forgot my login credentials"&lt;/strong&gt; shares &lt;em&gt;zero words&lt;/em&gt; with "How to reset your account password," yet it ranked first. A keyword search would have returned nothing. That's the payoff — you matched on meaning, not on string overlap. This shift from lexical to semantic matching is the foundation every retrieval-augmented system builds on, and it's the starting point of a structured &lt;a href="https://cursuri-ai.ro/courses/rag-retrieval-augmented-generation" rel="noopener noreferrer"&gt;course on RAG and retrieval-augmented generation&lt;/a&gt; that goes from this toy index to production retrieval.&lt;/p&gt;

&lt;h2&gt;
  
  
  From toy to production: what changes
&lt;/h2&gt;

&lt;p&gt;The NumPy version is perfect for learning and fine for a few thousand documents. Past that, three things force an upgrade.&lt;/p&gt;

&lt;h3&gt;
  
  
  You need a vector database
&lt;/h3&gt;

&lt;p&gt;Computing cosine similarity against &lt;em&gt;every&lt;/em&gt; stored vector on every query is &lt;code&gt;O(n)&lt;/code&gt; — fine at 5 documents, painful at 5 million. Vector databases solve this with &lt;strong&gt;Approximate Nearest Neighbor (ANN)&lt;/strong&gt; indexes (HNSW is the common one) that trade a sliver of accuracy for enormous speed, returning near-neighbors in milliseconds over huge corpora.&lt;/p&gt;

&lt;p&gt;You have good open-source options:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;pgvector&lt;/strong&gt; — a Postgres extension. If your data already lives in Postgres, this is often the pragmatic choice: vectors and relational data in one place, one backup story.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chroma / Qdrant / Weaviate / Milvus&lt;/strong&gt; — purpose-built vector stores with richer filtering and scaling stories.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;FAISS&lt;/strong&gt; — a library (not a server) from Meta for fast similarity search when you want to manage the index yourself.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A minimal Chroma example shows how little the mental model changes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;chromadb&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chromadb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Client&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;collection&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create_collection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;docs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ids&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;d&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="p"&gt;))])&lt;/span&gt;

&lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query_texts&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;I forgot my login credentials&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;n_results&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;documents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Chroma embeds the text for you and handles the index. Same concept, production ergonomics.&lt;/p&gt;

&lt;h3&gt;
  
  
  Chunking matters more than the model
&lt;/h3&gt;

&lt;p&gt;You rarely embed whole documents. A 30-page PDF becomes one vector that's an average of everything and a good match for nothing. In practice you &lt;strong&gt;chunk&lt;/strong&gt; documents into passages (a few hundred tokens, often with slight overlap) and embed each chunk. Get chunking wrong and even a great embedding model returns mush — which is one of the most common reasons retrieval systems quietly underperform. Chunking strategy, overlap, and metadata are exactly the unglamorous details that separate a demo from a dependable system, and they're covered in depth in a &lt;a href="https://cursuri-ai.ro/courses/advanced-llm-integration" rel="noopener noreferrer"&gt;course on advanced LLM integration for production apps&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Choosing and swapping embedding models
&lt;/h3&gt;

&lt;p&gt;Embedding models differ in dimensionality, speed, cost, and language coverage — and critically, &lt;strong&gt;you must embed your corpus and your queries with the same model.&lt;/strong&gt; Change the model and you re-embed everything. For multilingual apps (Romanian included), pick a model with strong multilingual training rather than an English-first one, or your non-English recall will suffer silently. Public benchmarks like the &lt;strong&gt;MTEB&lt;/strong&gt; leaderboard on Hugging Face are the sane starting point for comparing models on retrieval quality rather than vibes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hybrid search: when semantic alone isn't enough
&lt;/h2&gt;

&lt;p&gt;Here's a lesson that surprises people: pure semantic search is &lt;em&gt;worse&lt;/em&gt; than keyword search for certain queries. Ask for an exact product code, a specific error like &lt;code&gt;ERR_CONN_REFUSED&lt;/code&gt;, a person's name, or an acronym, and embeddings can betray you — they match on &lt;em&gt;meaning&lt;/em&gt;, and a precise identifier has little semantic meaning to spread around. The embedding for &lt;code&gt;ERR_CONN_REFUSED&lt;/code&gt; sits near "connection problems" generally, so a document about a &lt;em&gt;different&lt;/em&gt; connection error can outrank the exact match.&lt;/p&gt;

&lt;p&gt;The production answer is &lt;strong&gt;hybrid search&lt;/strong&gt;: run both a keyword search (classic lexical scoring like BM25) and a semantic search, then combine the rankings. Keyword search nails exact terms, identifiers, and rare words; semantic search nails paraphrase and intent. Together they cover each other's blind spots.&lt;/p&gt;

&lt;p&gt;The standard way to merge the two result lists is &lt;strong&gt;Reciprocal Rank Fusion (RRF)&lt;/strong&gt; — a simple, robust formula that scores each document by its &lt;em&gt;rank&lt;/em&gt; in each list rather than by raw scores that live on incompatible scales:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;reciprocal_rank_fusion&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rankings&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Fuse multiple ranked lists of doc IDs into one score per doc.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;ranking&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rankings&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;            &lt;span class="c1"&gt;# e.g. [keyword_results, semantic_results]
&lt;/span&gt;        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;rank&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ranking&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;rank&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;reverse&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because RRF only looks at positions, you don't have to normalize BM25 scores against cosine similarities — a notorious apples-to-oranges trap. Most serious vector databases (Weaviate, Qdrant, and others) now offer hybrid search with RRF built in, precisely because "just embeddings" quietly underperforms on real, messy query logs. If you take one production lesson from this article, make it this: &lt;strong&gt;measure your recall on real queries, and reach for hybrid the moment exact-match queries show up.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where embeddings show up beyond search
&lt;/h2&gt;

&lt;p&gt;Semantic search is the gateway, but the same vectors power a lot more:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;RAG (retrieval-augmented generation)&lt;/strong&gt; — retrieve relevant chunks by embedding similarity, then feed them to an LLM as grounding context. Embeddings are the &lt;em&gt;retrieval&lt;/em&gt; half; without good retrieval, generation hallucinates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deduplication &amp;amp; clustering&lt;/strong&gt; — near-duplicate detection and topic clustering fall out of distances almost for free.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recommendations&lt;/strong&gt; — "items similar to this one" is a nearest-neighbor query in embedding space.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Classification&lt;/strong&gt; — embed labeled examples, then classify new items by nearest neighbors, often without training a dedicated model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The through-line: any time you need "similar in meaning" rather than "matches exactly," embeddings are the tool. Building these features end to end — API to product, with the retrieval and orchestration wired up properly — is the spine of a hands-on &lt;a href="https://cursuri-ai.ro/courses/construire-aplicatii-ai-python-sdk" rel="noopener noreferrer"&gt;course on building AI applications in Python with the OpenAI and Anthropic SDKs&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common mistakes that cost you hours
&lt;/h2&gt;

&lt;p&gt;A few traps that catch almost everyone the first time:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Re-embedding the corpus on every query.&lt;/strong&gt; Embed documents once, store the vectors, embed only the incoming query at search time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mixing models.&lt;/strong&gt; Query vectors and document vectors must come from the &lt;em&gt;same&lt;/em&gt; embedding model. A silent mismatch produces garbage rankings with no error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Forgetting to normalize.&lt;/strong&gt; If you compute raw dot products instead of cosine similarity (and your vectors aren't already unit-normalized), longer texts get an unfair boost. Normalize, or use a library that does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embedding documents that are too large.&lt;/strong&gt; One vector per giant document averages meaning into uselessness. Chunk first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trusting a single similarity threshold forever.&lt;/strong&gt; The "good enough" cutoff depends on your model and data. Measure it on real queries; don't hardcode 0.8 because a blog post said so.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Vector embeddings are one of those ideas that feels like magic until you see the mechanics — and then it's just geometry. Text becomes a point in space, meaning becomes distance, and search becomes "find the nearest points." You built exactly that in a few lines of Python, from raw NumPy cosine similarity to a Chroma-backed index, and the same core idea scales from a toy corpus to millions of documents behind an ANN index.&lt;/p&gt;

&lt;p&gt;Start where we started: embed a handful of your own documents, run a query that shares no keywords with the right answer, and watch it surface anyway. Once that clicks, RAG, recommendations, and semantic features stop looking like separate topics and start looking like one tool applied five ways. If you want the structured path from here — retrieval, chunking, evaluation, and production wiring — &lt;a href="https://cursuri-ai.ro/courses/rag-retrieval-augmented-generation" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt; builds it step by step.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Sources &amp;amp; further reading:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Sentence-Transformers — &lt;a href="https://www.sbert.net/" rel="noopener noreferrer"&gt;Official documentation&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Hugging Face — &lt;a href="https://huggingface.co/spaces/mteb/leaderboard" rel="noopener noreferrer"&gt;MTEB: Massive Text Embedding Benchmark leaderboard&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;pgvector — &lt;a href="https://github.com/pgvector/pgvector" rel="noopener noreferrer"&gt;Open-source vector similarity search for Postgres&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Chroma — &lt;a href="https://www.trychroma.com/" rel="noopener noreferrer"&gt;Open-source embedding database&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This article is educational content. Model names, dimensions, and library APIs evolve; verify current details in the official documentation before building production systems.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Context Engineering for AI Agents: Beyond Prompt Engineering</title>
      <dc:creator>galian</dc:creator>
      <pubDate>Thu, 09 Jul 2026 13:08:24 +0000</pubDate>
      <link>https://dev.to/cursuri-ai/context-engineering-for-ai-agents-beyond-prompt-engineering-14ep</link>
      <guid>https://dev.to/cursuri-ai/context-engineering-for-ai-agents-beyond-prompt-engineering-14ep</guid>
      <description>&lt;p&gt;You wrote a great prompt. It worked beautifully in the playground — one question, one clean answer. Then you wired the same model into an agent that runs twenty steps, calls six tools, and reads back their output, and somewhere around step twelve it started forgetting the goal, calling the wrong tool, or confidently acting on something it misread three steps ago. The prompt didn't get worse. The &lt;em&gt;context&lt;/em&gt; did.&lt;/p&gt;

&lt;p&gt;This is the gap that context engineering fills. Prompt engineering is about writing one good instruction. Context engineering is about managing the entire set of tokens a model sees at inference time — across a long, multi-step run — so the signal stays high and the model keeps making good decisions. Anthropic frames it as &lt;a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents" rel="noopener noreferrer"&gt;the natural progression of prompt engineering&lt;/a&gt;, and if you're building anything agentic in 2026, it's the discipline that separates a demo from a system.&lt;/p&gt;

&lt;h2&gt;
  
  
  What context engineering actually is
&lt;/h2&gt;

&lt;p&gt;Start with a precise definition. &lt;strong&gt;Context&lt;/strong&gt; is the full set of tokens you include when you sample from a large language model. Not just your prompt — the system instructions, the tool definitions, the examples, the running message history, the retrieved documents, the tool results fed back in. Everything in the window.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context engineering&lt;/strong&gt; is the set of strategies for curating and maintaining the &lt;em&gt;optimal&lt;/em&gt; set of those tokens during inference. The goal, in one line: find the smallest set of high-signal tokens that reliably produces the outcome you want.&lt;/p&gt;

&lt;p&gt;The reason this is a distinct discipline from prompt engineering is the shape of the problem. A prompt is something you write once and it stays put. Context in an agent is &lt;em&gt;dynamic&lt;/em&gt; — it grows on every turn as the model reads files, calls tools, and accumulates history. You're not authoring a static string anymore; you're managing a budget that fills up on its own, and deciding continuously what earns a place in it and what gets thrown out. That's an engineering problem, and it's &lt;a href="https://cursuri-ai.ro/courses/prompt-engineering-masterclass" rel="noopener noreferrer"&gt;the foundation the whole prompt-to-production journey builds on&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "just add more context" is the wrong instinct
&lt;/h2&gt;

&lt;p&gt;The intuitive move, when an agent makes a mistake, is to give it more: more instructions, more examples, more history, more retrieved documents. Sometimes that helps. Very often it makes things worse, and here's why.&lt;/p&gt;

&lt;p&gt;A model's effective attention is a finite resource. Every token you add competes with every other token for the model's limited ability to attend to what matters. Past a certain point, adding context doesn't add capability — it dilutes it. The relevant fact is now buried among ten irrelevant ones, and the model attends to the wrong thing.&lt;/p&gt;

&lt;p&gt;This shows up empirically. The "lost in the middle" effect — &lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;documented by Liu et al.&lt;/a&gt; — found that models attend most reliably to information at the &lt;em&gt;start&lt;/em&gt; and &lt;em&gt;end&lt;/em&gt; of a long context, and least reliably to what's stuck in the middle. As context grows, retrieval of any single fact inside it gets less reliable, a degradation sometimes called &lt;em&gt;context rot&lt;/em&gt;. A 200K-token window does not mean you should put 200K tokens in it. Capacity is not the same as attention.&lt;/p&gt;

&lt;p&gt;So the mental model to internalize: &lt;strong&gt;context is a budget, not a bucket.&lt;/strong&gt; You're not trying to fill it. You're trying to spend it on the highest-signal tokens available and refuse everything else. Every technique below is a way to enforce that discipline.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four things competing for your window
&lt;/h2&gt;

&lt;p&gt;Before the techniques, know your spenders. In a running agent, four categories of tokens fight for the same finite budget:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The system prompt&lt;/strong&gt; — your instructions, role, constraints. Usually small, high-value, and stable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool definitions&lt;/strong&gt; — the schemas describing every tool the agent can call. These are sneakily expensive: each tool's description sits in context on &lt;em&gt;every&lt;/em&gt; turn, whether or not it's used.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Message history&lt;/strong&gt; — the accumulating transcript of the conversation and the agent's own steps. This is the one that grows without bound and quietly eats the window.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieved / external data&lt;/strong&gt; — documents, search results, file contents, database rows pulled in to ground the model.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The overall guidance from Anthropic is worth memorizing: keep each of these &lt;strong&gt;informative yet tight&lt;/strong&gt;. Not empty — an under-specified system prompt or a missing tool leaves the model guessing. But not bloated either. The art is the calibration, and it's different for each category. Let's work through the techniques that manage them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Technique 1: Compaction — summarize the history before it drowns you
&lt;/h2&gt;

&lt;p&gt;Message history is the runaway spender. A long agent run accumulates hundreds of turns of "called tool, got 4KB of JSON back, reasoned about it, called the next tool." Most of those raw tool outputs are dead weight three steps later — you needed the &lt;em&gt;conclusion&lt;/em&gt;, not the 4KB.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compaction&lt;/strong&gt; is the fix: periodically replace a chunk of verbose history with a tight summary that preserves the decisions, the key facts, and the current state, while dropping the raw noise. When the agent has finished investigating something, you don't need the full transcript of the investigation in context — you need "here's what I found and what it means for the task."&lt;/p&gt;

&lt;p&gt;Practical rules:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Compact at natural boundaries&lt;/strong&gt;, not mid-reasoning — after a sub-task completes, when a phase ends.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Preserve the load-bearing details&lt;/strong&gt;: open questions, decisions made, constraints discovered, current state. Drop the intermediate chatter and the raw dumps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the goal pinned.&lt;/strong&gt; The single most common long-run failure is the agent losing the plot on the original objective. The goal should survive every compaction.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Done well, compaction is what lets an agent run for hundreds of steps without either overflowing its window or forgetting why it started.&lt;/p&gt;

&lt;h2&gt;
  
  
  Technique 2: External memory — let the agent write things down
&lt;/h2&gt;

&lt;p&gt;The window is not the only place to keep information. The most effective long-running agents treat the context window as &lt;em&gt;working memory&lt;/em&gt; and offload durable state to &lt;strong&gt;external memory&lt;/strong&gt; — a file, a scratchpad, a structured store the agent reads from and writes to deliberately.&lt;/p&gt;

&lt;p&gt;Instead of carrying every fact in-context forever, the agent writes a note ("the auth module uses JWT with 15-minute expiry; the bug is in the refresh path") to a persistent store, and pulls it back only when relevant. The context window stays lean; the knowledge doesn't get lost. This is exactly how a human engineer works — you don't hold the entire codebase in your head, you keep notes and open the file when you need it.&lt;/p&gt;

&lt;p&gt;This pattern — persistent, deliberately-managed memory outside the window — is &lt;a href="https://cursuri-ai.ro/courses/context-engineering-memorie-agenti" rel="noopener noreferrer"&gt;the core of building agents that hold state over long horizons&lt;/a&gt;, and it's what turns a stateless model into something that can work a problem across a session without drowning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Technique 3: Sub-agents — isolate context so the main thread stays clean
&lt;/h2&gt;

&lt;p&gt;Here's a structural technique that most people never turn on. When a sub-task is going to burn a lot of tokens — "find every call site of this function across the repo," "research these five libraries" — don't do it in the main agent's context. Delegate it to a &lt;strong&gt;sub-agent&lt;/strong&gt;: a separate agent instance with its own isolated window that does the noisy work and reports back a clean result.&lt;/p&gt;

&lt;p&gt;The win is context hygiene. The thousands of tokens of file contents and search output that the investigation churns through stay in the &lt;em&gt;sub-agent's&lt;/em&gt; context and die with it. The main agent gets back a two-paragraph summary, and its own window stays focused on the actual task instead of silting up with intermediate noise. As a bonus, independent sub-tasks run in parallel instead of serially.&lt;/p&gt;

&lt;p&gt;The rule of thumb: delegate when work is &lt;strong&gt;independent, parallelizable, or context-heavy&lt;/strong&gt;; keep it inline when it's sequential and cheap. Knowing when to fan out to sub-agents and how to orchestrate them without stepping on each other is &lt;a href="https://cursuri-ai.ro/courses/ai-agents-automatizare" rel="noopener noreferrer"&gt;a central skill in building AI agents and automation&lt;/a&gt;, and it's one of the highest-leverage context moves you have.&lt;/p&gt;

&lt;h2&gt;
  
  
  Technique 4: Just-in-time retrieval — pull data when needed, not upfront
&lt;/h2&gt;

&lt;p&gt;There are two ways to get external data into an agent. The naive way is to &lt;strong&gt;preload&lt;/strong&gt;: at the start, dump everything the agent might conceivably need into context — the whole document, all the schemas, every config file. The problem is obvious once you see it: you're spending your budget on maybes, and most of it goes unused while crowding out what matters.&lt;/p&gt;

&lt;p&gt;The better pattern is &lt;strong&gt;just-in-time retrieval&lt;/strong&gt;: give the agent the &lt;em&gt;ability&lt;/em&gt; to fetch data (a search tool, a file reader, a database query) and let it pull exactly what it needs, exactly when it needs it. Instead of "here are all 40 files," it's "here's a tool to read a file — go get the one you need." The agent loads the relevant chunk into context at the moment of use, acts on it, and (with compaction) lets it fall away afterward.&lt;/p&gt;

&lt;p&gt;This mirrors how retrieval-augmented systems already work, and getting the retrieval layer right — what to fetch, how to rank it, how much to bring back — is &lt;a href="https://cursuri-ai.ro/courses/advanced-llm-integration" rel="noopener noreferrer"&gt;where advanced LLM integration earns its keep&lt;/a&gt;. Preloading feels safer; just-in-time is what actually scales.&lt;/p&gt;

&lt;h2&gt;
  
  
  Technique 5: Tool curation — the failure mode hiding in plain sight
&lt;/h2&gt;

&lt;p&gt;Remember that tool definitions sit in context on every turn. That makes a bloated tool set a double tax: it burns tokens continuously, &lt;em&gt;and&lt;/em&gt; it degrades decisions. Anthropic calls out one of the most common failure modes directly — tool sets that cover too much functionality or create ambiguous, overlapping choices about which tool to use.&lt;/p&gt;

&lt;p&gt;The tell is a sharp one: &lt;strong&gt;if a human engineer can't say for certain which tool should be used in a given situation, an AI agent can't either.&lt;/strong&gt; Fifteen tools with fuzzy, overlapping responsibilities will produce worse behavior than six sharp, non-overlapping ones — and cost more tokens doing it.&lt;/p&gt;

&lt;p&gt;So curate the toolbox like you'd curate an API:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fewer, sharper tools.&lt;/strong&gt; Each with a clear, distinct job and an unambiguous "use this when…"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No overlap.&lt;/strong&gt; Two tools that could both plausibly handle the same request is a decision point where the agent will sometimes pick wrong.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prune ruthlessly.&lt;/strong&gt; A tool that's rarely the right choice is paying rent in your context on every single turn. Cut it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Tool curation is the least glamorous technique here and often the highest-ROI. It's pure subtraction, and subtraction is exactly what context engineering rewards.&lt;/p&gt;

&lt;h2&gt;
  
  
  How you know it's working: measure it
&lt;/h2&gt;

&lt;p&gt;Every technique above is a change to your context, and changes to context are exactly the kind of thing that &lt;em&gt;feels&lt;/em&gt; better while being worse — or vice versa. If your only signal is "I ran a few queries and it seemed fine," you're tuning blind, and you'll ship a regression the day you compact one turn too aggressively and the agent starts forgetting a constraint.&lt;/p&gt;

&lt;p&gt;The fix is the same as anywhere in production ML: an evaluation set. Assemble representative tasks with known good outcomes, and re-run them every time you change how context is managed. Then you can say "compaction at phase boundaries held task success at 0.9 while cutting average tokens 40%" instead of "I think it's better now." Treating &lt;a href="https://cursuri-ai.ro/courses/ai-evals-llm-productie" rel="noopener noreferrer"&gt;evaluation as first-class rather than an afterthought&lt;/a&gt; is what turns context engineering from a craft into engineering — a number that moves when you change something, not a vibe.&lt;/p&gt;

&lt;h2&gt;
  
  
  Putting it together
&lt;/h2&gt;

&lt;p&gt;Context engineering isn't a framework you install; it's a posture you adopt toward the model's window. The whole discipline collapses to one principle applied relentlessly: &lt;strong&gt;spend the finite budget on the smallest set of high-signal tokens that does the job.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In practice, for a real agent, that means:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Compact&lt;/strong&gt; the history at natural boundaries so a long run doesn't drown in its own transcript.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Offload&lt;/strong&gt; durable state to external memory instead of carrying it forever.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delegate&lt;/strong&gt; noisy, independent work to sub-agents so the main window stays clean.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieve just-in-time&lt;/strong&gt; instead of preloading everything you might need.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Curate the tools&lt;/strong&gt; hard — fewer, sharper, no overlap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure&lt;/strong&gt; with evals so every change is verified, not hoped.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of these are exotic, and that's the point. The model is already capable. What makes the capability &lt;em&gt;hold up&lt;/em&gt; over a twenty-step run isn't a cleverer prompt — it's disciplined management of everything the model reads along the way.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Prompt engineering taught us to write one good instruction. Context engineering is what you need the moment that instruction has to survive a long, tool-using, self-accumulating agent run — which is to say, the moment you build anything real. The failure you saw at step twelve was never the model getting dumber. It was the context getting noisier, and no one curating it.&lt;/p&gt;

&lt;p&gt;Adopt the budget mindset, apply the five techniques, and put an eval set behind every change. Do that, and the same model that fell apart at step twelve will run to step fifty and still know exactly what it's doing — because you engineered what it was looking at the whole way.&lt;/p&gt;

&lt;p&gt;The courses linked throughout are part of &lt;a href="https://cursuri-ai.ro/courses/context-engineering-memorie-agenti" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt;, an AI-learning platform with hands-on, current tracks on context engineering, AI agents, LLM integration, and evaluating AI systems in production.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Sources &amp;amp; further reading:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Anthropic — &lt;a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents" rel="noopener noreferrer"&gt;Effective context engineering for AI agents&lt;/a&gt; (definition, the four context components, tool-set failure modes, compaction and memory)&lt;/li&gt;
&lt;li&gt;Anthropic — &lt;a href="https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents" rel="noopener noreferrer"&gt;Effective harnesses for long-running agents&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Liu et al. — &lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;Lost in the Middle: How Language Models Use Long Contexts&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This article is educational content. Model behavior, context limits, and tooling evolve quickly; validate approaches against your own workloads and current official documentation.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Run LLMs Locally with Ollama in 2026: The Practical Developer Guide</title>
      <dc:creator>galian</dc:creator>
      <pubDate>Sun, 05 Jul 2026 21:11:24 +0000</pubDate>
      <link>https://dev.to/cursuri-ai/run-llms-locally-with-ollama-in-2026-the-practical-developer-guide-48n</link>
      <guid>https://dev.to/cursuri-ai/run-llms-locally-with-ollama-in-2026-the-practical-developer-guide-48n</guid>
      <description>&lt;p&gt;For years, "run the model locally" was the option you mentioned and then didn't take: the models were too weak, the tooling too fiddly, and the cloud APIs too convenient. In 2026 that calculus has genuinely shifted. Open-weight models in the 12–35B range now handle real coding and agent workloads, Apple Silicon got a dedicated inference engine, and Ollama quietly became a drop-in backend for the same tools you already use against cloud APIs — including Claude Code.&lt;/p&gt;

&lt;p&gt;I teach practical AI engineering at &lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt;, Eastern Europe's AI education platform, and local inference has gone from a curiosity module to one of the questions companies ask us most — usually spelled "how do we use LLMs without sending our data anywhere?" This guide is the answer I give developers: what changed, what hardware you actually need, which models are worth pulling, and how to plug it all into a real workflow.&lt;/p&gt;

&lt;p&gt;As always with this space: versions and model rankings move monthly. Everything below is verified against Ollama's official blog and docs as of early July 2026 — re-check before you build a budget or an architecture on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why local, and why now
&lt;/h2&gt;

&lt;p&gt;Three arguments have survived contact with production; the rest is mostly vibes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Privacy and data residency.&lt;/strong&gt; With a local model, prompts and outputs never leave your machine (or your VPC, if you self-host on a server). For anyone dealing with client data, medical text, legal documents, or GDPR-sensitive workloads, this eliminates the entire "what does the provider do with my data" conversation instead of managing it through contracts. This is the single biggest adoption driver we see in Europe, and it's the backbone of our &lt;a href="https://cursuri-ai.ro/courses/llm-locale-ollama-privacy-self-hosting" rel="noopener noreferrer"&gt;course on local LLMs, self-hosting, and privacy&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost shape.&lt;/strong&gt; Cloud APIs bill per token; local inference bills you once, in hardware you may already own. For high-volume, latency-tolerant workloads — batch classification, summarization pipelines, internal tooling — a mid-range GPU that's already on someone's desk can absorb work that would otherwise be a real monthly line item. (For low-volume or frontier-quality work, cloud still wins. More on that below.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No external dependency.&lt;/strong&gt; A local model doesn't get deprecated, rate-limited, price-changed, or suspended out from under you. After the model-availability surprises of the last year, "at least one workload runs on weights we control" has become a reasonable line item in a resilience plan, not paranoia.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually changed in Ollama in 2026
&lt;/h2&gt;

&lt;p&gt;If you last touched Ollama when it was "a nice wrapper around llama.cpp," the 2026 releases are the reason to look again. All of this is from &lt;a href="https://ollama.com/blog" rel="noopener noreferrer"&gt;Ollama's official blog&lt;/a&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Anthropic API compatibility (January, v0.14.0).&lt;/strong&gt; Ollama now exposes a native Anthropic-style &lt;code&gt;/v1/messages&lt;/code&gt; endpoint. This is the sleeper feature of the year: Anthropic-native tools — most notably Claude Code — can talk to a local model directly, with no proxy or translation layer. There's a matching OpenAI-compatible endpoint too, so Codex and OpenAI-SDK apps work the same way.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ollama launch&lt;/code&gt; (January).&lt;/strong&gt; A single command that configures and starts a coding agent against a local model — &lt;code&gt;ollama launch claude&lt;/code&gt; sets up Claude Code, prompts you to pick a model, and you're in.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Experimental image generation (January).&lt;/strong&gt; Early days, but the scope of "local model" is no longer text-only.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MLX engine on Apple Silicon (March preview → June release).&lt;/strong&gt; Ollama moved its Mac inference path to Apple's MLX framework, which exploits unified memory. Ollama's own framing for the June release: its highest performance on Apple Silicon yet — faster output with reduced memory usage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ollama 0.30 and 0.31 (June).&lt;/strong&gt; Version 0.30 brought improved performance and broader GGUF model compatibility through llama.cpp; 0.31 made Gemma 4 significantly faster on Apple Silicon via multi-token prediction (MTP), enabled by default.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The theme is clear: Ollama is positioning itself less as a hobbyist toy and more as the standard local backend for agentic tooling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting started in five minutes
&lt;/h2&gt;

&lt;p&gt;Install (macOS and Windows have installers at &lt;a href="https://ollama.com/download" rel="noopener noreferrer"&gt;ollama.com/download&lt;/a&gt;; on Linux):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://ollama.com/install.sh | sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pull and run a model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull gemma4
ollama run gemma4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's a working local chat. Ollama also starts a local server on port &lt;code&gt;11434&lt;/code&gt;, which is where the interesting part begins — every API-based tool you have can point at it.&lt;/p&gt;

&lt;p&gt;Useful daily commands: &lt;code&gt;ollama ls&lt;/code&gt; (installed models), &lt;code&gt;ollama ps&lt;/code&gt; (what's loaded and where — CPU vs GPU), &lt;code&gt;ollama rm &amp;lt;model&amp;gt;&lt;/code&gt; (free disk space; models are multi-gigabyte).&lt;/p&gt;

&lt;h2&gt;
  
  
  Hardware: the honest sizing guide
&lt;/h2&gt;

&lt;p&gt;The rule of thumb that matters: a model quantized to 4 bits needs very roughly &lt;strong&gt;0.5–0.7 GB of memory per billion parameters&lt;/strong&gt;, plus overhead for context. Everything else follows from that.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Your hardware&lt;/th&gt;
&lt;th&gt;What runs comfortably&lt;/th&gt;
&lt;th&gt;Experience&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;8 GB RAM, no GPU&lt;/td&gt;
&lt;td&gt;3–8B models, quantized&lt;/td&gt;
&lt;td&gt;Fine for chat, drafting, classification. Slow but usable on CPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16 GB RAM (Apple Silicon)&lt;/td&gt;
&lt;td&gt;8–14B models&lt;/td&gt;
&lt;td&gt;Good daily-driver territory; MLX made this tier notably faster in 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;24 GB+ (M-series Pro/Max or a 24 GB GPU)&lt;/td&gt;
&lt;td&gt;27–35B models&lt;/td&gt;
&lt;td&gt;Where local coding models get genuinely useful&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;48 GB+ unified memory / multi-GPU&lt;/td&gt;
&lt;td&gt;Large MoE models&lt;/td&gt;
&lt;td&gt;Server-class local inference&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two nuances that save people disappointment:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Quantization is why any of this works.&lt;/strong&gt; Models ship in compressed 4–8 bit variants (the GGUF ecosystem) that trade a small quality loss for a 2–4× memory reduction. Ollama's default tags are already quantized — you rarely need to think about it, but it explains why a "27B model" fits in 24 GB.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mixture-of-experts (MoE) models need memory for their &lt;em&gt;total&lt;/em&gt; parameters but compute like their &lt;em&gt;active&lt;/em&gt; subset.&lt;/strong&gt; NVIDIA's Nemotron-3-Super, for example, is a 120B model with 12B active parameters: it runs faster than its size suggests, but you still need the RAM to hold it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Context length eats memory too — an agent session with 32K+ tokens of context adds real overhead on top of the weights. If you're sizing for coding agents, budget for that, not just the model file.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mid-2026 open-weight lineup worth knowing
&lt;/h2&gt;

&lt;p&gt;Rankings churn monthly, so treat this as a map, not a leaderboard. From &lt;a href="https://ollama.com/search" rel="noopener noreferrer"&gt;Ollama's model library&lt;/a&gt;, the families that matter right now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemma 4&lt;/strong&gt; (12B–31B) — Google's open family, currently the most-pulled model on Ollama. Multimodal, tuned for reasoning and agentic work, and the main beneficiary of the MLX/MTP speedups on Macs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qwen3.5 / Qwen3.6&lt;/strong&gt; (0.8B–122B) — the ecosystem's Swiss army knife. Qwen3.5 spans everything from edge-tiny to server-large; Qwen3.6 (27B–35B) focuses on agentic coding. &lt;strong&gt;qwen3-coder&lt;/strong&gt; is Ollama's own recommendation for coding-agent use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GLM-5 family&lt;/strong&gt; — flagship-class open weights (GLM-5 is 744B total / 40B active); strong at coding and long-horizon tasks. Too big for most desktops locally, but available as &lt;code&gt;:cloud&lt;/code&gt; variants (see below) and self-hostable on serious hardware.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nemotron-3-Super&lt;/strong&gt; (120B MoE, 12B active) — NVIDIA's entry, aimed at multi-agent applications.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MiniMax-M3&lt;/strong&gt; — notable for a 1M-token context window, if your workload is long-document analysis.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Specialists&lt;/strong&gt;: GLM-OCR for document understanding, TranslateGemma (4B–27B, 55 languages) for translation, LFM2 (24B) for on-device deployment, Ornith (9B–35B) for agentic coding.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Sensible defaults: on a 16 GB Mac, start with &lt;code&gt;gemma4:12b&lt;/code&gt;. On 24 GB+, try &lt;code&gt;qwen3-coder&lt;/code&gt; for code and &lt;code&gt;gemma4:27b&lt;/code&gt; for general work. Then run &lt;em&gt;your&lt;/em&gt; tasks on them — a model's rank on someone's benchmark tells you little about your use case.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that changes your workflow: Ollama as a drop-in API
&lt;/h2&gt;

&lt;p&gt;Ollama's server speaks both major API dialects, which means "switch to a local model" is now a base-URL change, not a rewrite.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OpenAI-compatible&lt;/strong&gt; (&lt;code&gt;/v1&lt;/code&gt;) — any OpenAI-SDK app works:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:11434/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ollama&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemma4&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Explain GGUF quantization in one paragraph.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Anthropic-compatible&lt;/strong&gt; (&lt;code&gt;/v1/messages&lt;/code&gt;) — and this is the one with teeth, because it means &lt;strong&gt;Claude Code runs against local models&lt;/strong&gt;. Per &lt;a href="https://docs.ollama.com/api/anthropic-compatibility" rel="noopener noreferrer"&gt;Ollama's official docs&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_AUTH_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;ollama       &lt;span class="c"&gt;# accepted but not validated&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;http://localhost:11434

claude &lt;span class="nt"&gt;--model&lt;/span&gt; qwen3-coder
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or let Ollama do the wiring for you:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama launch claude
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Honest caveats before you get excited:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ollama recommends &lt;strong&gt;at least 32K tokens of context&lt;/strong&gt; for Claude Code, and its model suggestions for coding are &lt;code&gt;qwen3-coder&lt;/code&gt; locally (30B — you want 24 GB+ of VRAM/unified memory) or &lt;code&gt;glm-4.7:cloud&lt;/code&gt; / &lt;code&gt;minimax-m2.1:cloud&lt;/code&gt; via Ollama's cloud, which keeps the same API surface but runs the weights remotely.&lt;/li&gt;
&lt;li&gt;The compatibility layer doesn't cover everything: &lt;strong&gt;no prompt caching, no token-counting endpoint, no forced tool selection, no batches API, no PDF inputs&lt;/strong&gt; (images must be base64). If your workflow leans on those, you'll feel it.&lt;/li&gt;
&lt;li&gt;A 30B local model is not Opus, and it isn't trying to be. It's "capable pair of hands on an airplane / on confidential code," not "frontier reasoning."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern that actually works in practice is routing: local models for the private, high-volume, or offline work; frontier cloud models for the hard reasoning. Deciding which tier a task belongs to — and building the escalation path — is an architecture skill, and it's exactly the kind of decision we drill in our &lt;a href="https://cursuri-ai.ro/courses/ai-system-architecture" rel="noopener noreferrer"&gt;AI system architecture course&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  When local is the wrong choice
&lt;/h2&gt;

&lt;p&gt;Being a fan of local inference means knowing where it loses:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Frontier-quality reasoning.&lt;/strong&gt; For the hardest tasks, top cloud models remain clearly ahead of anything you can run on a workstation. If wrong answers are expensive, don't fight this.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Low-volume workloads.&lt;/strong&gt; If you make a few thousand API calls a month, per-token billing is cheaper than any GPU. Local pays off at volume, at privacy constraints, or at both.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ops you don't want.&lt;/strong&gt; A self-hosted model is a service you now run: updates, monitoring, capacity. &lt;code&gt;ollama run&lt;/code&gt; on a laptop is trivial; a team-wide inference server is real infrastructure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multimodal breadth and long-tail capabilities.&lt;/strong&gt; Cloud APIs still bundle more (native PDF understanding, larger tool ecosystems, batch APIs) than the local stack replicates.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One more thing people conflate: running a model locally is different from &lt;em&gt;customizing&lt;/em&gt; one. If your actual goal is a model that speaks your domain language or follows your house style, that's a fine-tuning question — LoRA adapters on an open-weight base, then serving the result through Ollama. That pipeline (when to fine-tune vs when to just engineer the prompt) is its own discipline, covered in our &lt;a href="https://cursuri-ai.ro/courses/fine-tuning-modele-ai" rel="noopener noreferrer"&gt;fine-tuning course&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is Ollama free?
&lt;/h3&gt;

&lt;p&gt;The tool itself is open source and free. The models each carry their own licenses — Gemma, Qwen, GLM and friends have different terms, some with restrictions on commercial use. Check the license tab on the model's Ollama page before you ship a product on it. Ollama's optional cloud models are a paid service.&lt;/p&gt;

&lt;h3&gt;
  
  
  What hardware do I need to run LLMs locally in 2026?
&lt;/h3&gt;

&lt;p&gt;As a rule of thumb at 4-bit quantization: 8 GB of RAM runs 3–8B models, 16 GB runs 8–14B comfortably (especially on Apple Silicon with the MLX engine), and 24 GB+ opens up the 27–35B class where local coding models get genuinely useful. More context = more memory on top of the weights.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can a local model replace GPT or Claude?
&lt;/h3&gt;

&lt;p&gt;For a growing set of tasks — summarization, classification, drafting, routine coding on mid-size codebases — yes, credibly. For frontier reasoning and the highest-stakes accuracy, no. Production teams typically route: local for private/high-volume work, cloud for the hard 10%.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I really use Claude Code with Ollama?
&lt;/h3&gt;

&lt;p&gt;Yes. Since Ollama v0.14.0 (January 2026) there's native Anthropic Messages API compatibility: set &lt;code&gt;ANTHROPIC_BASE_URL=http://localhost:11434&lt;/code&gt;, run &lt;code&gt;claude --model qwen3-coder&lt;/code&gt; — or just &lt;code&gt;ollama launch claude&lt;/code&gt;. Expect a capable assistant, not Opus-level reasoning, and note that prompt caching and a few other API features aren't supported through the compatibility layer.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ollama vs llama.cpp vs vLLM — which should I use?
&lt;/h3&gt;

&lt;p&gt;Ollama for developer experience: one command, model management, dual API compatibility. llama.cpp (which powers Ollama's GGUF path) for maximum control and minimal footprint. vLLM for high-throughput multi-user serving on server GPUs. Most developers should start with Ollama and only move down the stack when they hit a concrete limit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The skill underneath the tool
&lt;/h2&gt;

&lt;p&gt;Here's the uncomfortable part: pulling a model is the easy 5%. The value shows up when you can answer the questions around it — which model for which task, how to measure whether the local model is &lt;em&gt;good enough&lt;/em&gt; for your workload instead of guessing, how to build the routing and fallback so privacy-sensitive work stays local while hard problems escalate to a frontier model. That's engineering judgment, not tooling trivia.&lt;/p&gt;

&lt;p&gt;That judgment is what we teach at &lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;our AI education platform&lt;/a&gt; — hands-on courses built around real repositories and an interactive AI instructor, covering the full local-to-cloud spectrum: self-hosting and privacy, fine-tuning, architecture, and the agentic workflow on top.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;In 2026, local LLMs crossed the line from hobby to infrastructure option. Ollama's dual API compatibility means your existing tools — including Claude Code — can run against open weights with a base-URL change; the MLX engine made a 16 GB MacBook a legitimate inference machine; and the open-weight lineup in the 12–35B range is good enough for a real slice of production work.&lt;/p&gt;

&lt;p&gt;The play isn't "cancel your API keys." It's knowing which slice of your workload belongs on weights you control — then running it there deliberately, measured, with an escalation path for everything else. Start with &lt;code&gt;ollama run gemma4&lt;/code&gt; tonight; you're one evening away from having an informed opinion.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written by the team at &lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt; — practical, hands-on AI engineering courses for developers and professionals across Eastern Europe, from local LLMs and self-hosting to agentic coding, evals, and AI system architecture.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources:&lt;/strong&gt; &lt;a href="https://ollama.com/blog" rel="noopener noreferrer"&gt;Ollama Blog&lt;/a&gt; · &lt;a href="https://docs.ollama.com/api/anthropic-compatibility" rel="noopener noreferrer"&gt;Anthropic API compatibility — Ollama Docs&lt;/a&gt; · &lt;a href="https://ollama.com/blog/claude" rel="noopener noreferrer"&gt;Claude Code with Anthropic API compatibility — Ollama&lt;/a&gt; · &lt;a href="https://ollama.com/search" rel="noopener noreferrer"&gt;Ollama Model Library&lt;/a&gt; · &lt;a href="https://ollama.com/download" rel="noopener noreferrer"&gt;Download Ollama&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Claude Sonnet 5 Just Made Running Agents Cheap — What Builders Actually Need to Know</title>
      <dc:creator>galian</dc:creator>
      <pubDate>Tue, 30 Jun 2026 22:06:37 +0000</pubDate>
      <link>https://dev.to/galian/claude-sonnet-5-just-made-running-agents-cheap-what-builders-actually-need-to-know-11j7</link>
      <guid>https://dev.to/galian/claude-sonnet-5-just-made-running-agents-cheap-what-builders-actually-need-to-know-11j7</guid>
      <description>&lt;p&gt;Anthropic shipped &lt;strong&gt;Claude Sonnet 5&lt;/strong&gt; on June 30, 2026, and the framing in the announcement is unusually blunt for a model launch: it's pitched as the most &lt;em&gt;agentic&lt;/em&gt; Sonnet yet — a model built to make plans, drive tools like browsers and terminals, and run autonomously at a level that, a few months ago, took something bigger and more expensive.&lt;/p&gt;

&lt;p&gt;For anyone building on top of these models — agents, pipelines, coding tools — that's the headline that matters. Not "it's smarter," but "near-frontier capability just got cheaper to run in a loop." I write and teach about agentic engineering at &lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt;, Eastern Europe's AI education platform, so I'll keep this grounded in what changes for people who actually ship on these APIs — not the launch-day benchmark theater.&lt;/p&gt;

&lt;p&gt;One disclaimer up front: model pricing and availability in this space change almost monthly, and this is a day-one snapshot. Verify the current numbers on Anthropic's official pages before you wire anything to a budget. I'm deliberately &lt;em&gt;not&lt;/em&gt; quoting benchmark scores here — the launch materials presented them in a way that's easy to misread, so for hard numbers go straight to the &lt;a href="https://www.anthropic.com/claude-sonnet-5-system-card" rel="noopener noreferrer"&gt;Sonnet 5 System Card&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-sentence version
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Sonnet 5 moves "good enough to run agents autonomously" down a price tier — and ships a new tokenizer that can quietly inflate your token counts by up to 35%.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Both halves of that sentence matter, and the second one is the part nobody puts on a launch slide. Let's take them in order.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's actually new
&lt;/h2&gt;

&lt;p&gt;Stripping the marketing down to verifiable claims from Anthropic's own announcement, here's what Sonnet 5 is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The most agentic Sonnet so far.&lt;/strong&gt; It's described as able to "make plans, use tools like browsers and terminals, and run autonomously," with improvements specifically in multi-step tool use — the exact workload that defines an agent rather than a chatbot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Close to Opus 4.8 — at a lower price.&lt;/strong&gt; Anthropic's own phrasing is that its "performance is close to that of Opus 4.8, but at lower prices." That's the whole pitch: most of the capability, a fraction of the cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A real step up from Sonnet 4.6.&lt;/strong&gt; Called a "substantial improvement over its predecessor, Sonnet 4.6, on important aspects of agentic performance like reasoning, tool use, coding, and knowledge work."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Safer in agentic contexts.&lt;/strong&gt; Anthropic reports an "overall lower rate of undesirable behaviors than Sonnet 4.6," plus lower rates of hallucination and sycophancy — which matters more than it sounds when a model is acting in a loop without a human reading every step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deliberately weaker at offensive cyber.&lt;/strong&gt; It shows "substantially poorer performance than models such as Opus 4.8" on dangerous cyber tasks and was "never able to develop a full working exploit." That's a safety design choice, not an oversight — worth knowing if security tooling is your domain.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two things Anthropic did &lt;strong&gt;not&lt;/strong&gt; publish that I'm not going to invent for you: an official &lt;strong&gt;context window&lt;/strong&gt; and &lt;strong&gt;max output token&lt;/strong&gt; figure for Sonnet 5 weren't stated in the launch materials at the time of writing. If you need those for capacity planning, pull them from the official API docs rather than trusting a blog (including this one). Guessing is how teams ship broken truncation logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  The economics shift is the real story
&lt;/h2&gt;

&lt;p&gt;Here's why builders should care more than end users.&lt;/p&gt;

&lt;p&gt;When you chat with a model, price-per-token is almost noise — you send a few thousand tokens and read the answer. When you run an &lt;strong&gt;agent&lt;/strong&gt;, the model is in a loop: read context, call a tool, read the result, reason, call another tool, repeat. A single "task" can burn hundreds of thousands of tokens across dozens of turns. At that volume, the price-per-million-tokens line &lt;em&gt;is&lt;/em&gt; your unit economics.&lt;/p&gt;

&lt;p&gt;So a model that lands near Opus-4.8 quality at Sonnet pricing doesn't just make chat cheaper — it changes which agent designs are economically viable at all. Workflows you'd previously gate behind Opus (multi-step research, autonomous refactors, long tool-using runs) become defensible on a Sonnet budget. That's the unlock.&lt;/p&gt;

&lt;p&gt;Here's the day-one pricing picture, with the rest of the current Anthropic lineup for context:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input / 1M tokens&lt;/th&gt;
&lt;th&gt;Output / 1M tokens&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Sonnet 5&lt;/strong&gt; (intro, through Aug 31 2026)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$2&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$10&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Promotional launch pricing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Sonnet 5&lt;/strong&gt; (standard, from Sep 1 2026)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$15&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Same as Sonnet 4.6's tier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Opus 4.8&lt;/td&gt;
&lt;td&gt;$5&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;Top accuracy; default in Claude Code&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Haiku 4.5&lt;/td&gt;
&lt;td&gt;$1&lt;/td&gt;
&lt;td&gt;$5&lt;/td&gt;
&lt;td&gt;Cheapest / fastest tier&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A few honest notes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The &lt;strong&gt;introductory $2 / $10&lt;/strong&gt; runs through &lt;strong&gt;August 31, 2026&lt;/strong&gt;, then settles to &lt;strong&gt;$3 / $15&lt;/strong&gt; — the same standard tier Sonnet has occupied. So the long-run story isn't "Sonnet got cheaper"; it's "the Sonnet tier got dramatically more capable for the same price."&lt;/li&gt;
&lt;li&gt;Sonnet 5 is the &lt;strong&gt;default model on Free and Pro plans&lt;/strong&gt;, and is available to Max, Team, and Enterprise users — in Claude Code, the Claude platform, and the API. So if you're on Claude Code, you may already be one model-switch away from it.&lt;/li&gt;
&lt;li&gt;Against Opus 4.8 the price ratio is roughly &lt;strong&gt;1.7×&lt;/strong&gt; (output $25 vs $15). When you're running agents at scale, that multiple compounds fast — which is exactly why the "close to Opus" claim is worth pressure-testing on &lt;em&gt;your&lt;/em&gt; workload, not taking on faith.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The tokenizer gotcha that will mess up your cost math
&lt;/h2&gt;

&lt;p&gt;This is the part I most want builders to internalize, because it's the easiest way to get a nasty surprise on your next invoice.&lt;/p&gt;

&lt;p&gt;Sonnet 5 ships with an &lt;strong&gt;updated tokenizer&lt;/strong&gt;. Anthropic states that the same input text now maps to &lt;strong&gt;roughly 1.0–1.35× as many tokens&lt;/strong&gt; as before, depending on content type. Read that again: identical prompts can cost up to &lt;strong&gt;35% more tokens&lt;/strong&gt; on Sonnet 5 than the token count you measured on an older model — &lt;em&gt;before&lt;/em&gt; any change in per-token price.&lt;/p&gt;

&lt;p&gt;Why it bites:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your cost dashboards, budget alerts, and per-request estimates were calibrated on the old tokenizer. Swap the model without re-measuring and your "same" workload silently costs more.&lt;/li&gt;
&lt;li&gt;Code, structured data (JSON/XML), and non-English text tend to sit at the higher end of that multiplier — and those are precisely the inputs agentic and coding workloads are made of.&lt;/li&gt;
&lt;li&gt;It interacts with context windows and truncation: more tokens for the same text means you hit limits sooner than your old math predicts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The fix is boring and non-negotiable: re-baseline.&lt;/strong&gt; Before you flip production traffic to Sonnet 5, measure real token counts on a representative sample of &lt;em&gt;your&lt;/em&gt; prompts with the new tokenizer, recompute cost per task, and update your budgets and alerts. The headline price drop is real — but the effective saving is &lt;code&gt;(price delta) × (token inflation)&lt;/code&gt;, and you can't know the second factor without measuring. Anyone who tells you "it's 33% cheaper" did half the arithmetic.&lt;/p&gt;

&lt;p&gt;This is also where good &lt;strong&gt;evals&lt;/strong&gt; earn their keep. A model swap isn't just a cost change; it's a behavior change. Run your task suite on Sonnet 5 against the model you're replacing before you commit — quality, tool-call success rate, and cost together. If you don't have an eval harness yet, this is the launch that should convince you to build one; it's a discipline we treat as core, not optional, in our &lt;a href="https://cursuri-ai.ro/courses/ai-evals-llm-productie" rel="noopener noreferrer"&gt;course on building LLM evals for production&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to still reach for Opus 4.8
&lt;/h2&gt;

&lt;p&gt;"Close to Opus" is not "Opus." The honest read on where Sonnet 5 fits:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reach for Sonnet 5&lt;/strong&gt; as your default agent workhorse: high-volume tool-using loops, coding assistance, research and summarization, anything where you're paying per turn and the marginal quality of Opus isn't worth ~1.7× the output cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stay on Opus 4.8&lt;/strong&gt; for the hardest reasoning, the highest-stakes accuracy, and security-sensitive work where Sonnet 5 is &lt;em&gt;intentionally&lt;/em&gt; weaker (offensive-cyber tasks). When a wrong answer is expensive, the price gap is cheap insurance.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern most production teams land on isn't "pick one." It's a &lt;strong&gt;router&lt;/strong&gt;: Sonnet 5 handles the bulk of turns, and you escalate to Opus 4.8 for the steps that genuinely need it — with a human in the loop on the consequential ones. Getting that routing logic right (and knowing which task belongs in which tier) is a real engineering skill, and it's the through-line of our &lt;a href="https://cursuri-ai.ro/courses/comparatie-modele-ai" rel="noopener noreferrer"&gt;model-comparison course&lt;/a&gt;, which treats "which model for which job" as a decision you make with data rather than vibes.&lt;/p&gt;

&lt;h2&gt;
  
  
  A pragmatic migration checklist
&lt;/h2&gt;

&lt;p&gt;If you're considering moving an agent or pipeline to Sonnet 5, here's the order I'd do it in:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Re-baseline tokens.&lt;/strong&gt; Run a representative sample through the new tokenizer. Recompute cost per task. Update budget alerts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run your evals.&lt;/strong&gt; Quality, tool-call success, latency, and cost, head-to-head against the model you're replacing. No eval suite? Build a small one first — even 30 representative tasks beats a gut call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shadow, then canary.&lt;/strong&gt; Route a slice of real traffic to Sonnet 5, compare outputs, then scale gradually. Don't flip 100% on day one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep an escalation path.&lt;/strong&gt; Wire Opus 4.8 as the fallback for tasks that fail Sonnet 5's quality bar. Routing beats an all-or-nothing bet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-read your safety posture.&lt;/strong&gt; Lower hallucination and sycophancy is good news for autonomous runs, but "safer" isn't "supervise nothing." Keep guardrails and human checkpoints where consequences are real.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of this is exotic. It's the same discipline that separates teams who run agents in production from teams who demo them — and it's exactly the muscle we build in our hands-on track on &lt;a href="https://cursuri-ai.ro/courses/ai-agents-automatizare" rel="noopener noreferrer"&gt;AI agents and automation&lt;/a&gt;, taught around real repositories rather than toy notebooks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is Claude Sonnet 5 better than Opus 4.8?
&lt;/h3&gt;

&lt;p&gt;Not across the board. Anthropic positions Sonnet 5's performance as &lt;em&gt;close to&lt;/em&gt; Opus 4.8 at a lower price — so for high-volume agentic and coding work it's often the better &lt;em&gt;value&lt;/em&gt;, but Opus 4.8 still leads on the hardest reasoning, top-end accuracy, and (deliberately) on offensive-cyber capability. Match the tier to the task instead of picking a favorite.&lt;/p&gt;

&lt;h3&gt;
  
  
  How much does Claude Sonnet 5 cost?
&lt;/h3&gt;

&lt;p&gt;It launched with introductory pricing of &lt;strong&gt;$2 per million input tokens and $10 per million output tokens through August 31, 2026&lt;/strong&gt;, then moves to a standard &lt;strong&gt;$3 / $15&lt;/strong&gt; — the same tier Sonnet 4.6 occupied. Your &lt;em&gt;effective&lt;/em&gt; cost also depends on the new tokenizer (see below), so measure before you budget.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does the new tokenizer really change my costs?
&lt;/h3&gt;

&lt;p&gt;Yes. Anthropic states the same input can map to roughly &lt;strong&gt;1.0–1.35× as many tokens&lt;/strong&gt; under Sonnet 5's updated tokenizer, depending on content type — code and structured data sit at the higher end. Re-measure your real prompts before assuming the headline price drop equals your actual saving.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I use Sonnet 5 in Claude Code?
&lt;/h3&gt;

&lt;p&gt;Yes. It's available in Claude Code, the Claude platform, and the API, and it's the default model on Free and Pro plans (and available to Max, Team, and Enterprise). If you're already in Claude Code, switching is a model selection, not a migration.&lt;/p&gt;

&lt;h3&gt;
  
  
  Should I migrate my agents to Sonnet 5 immediately?
&lt;/h3&gt;

&lt;p&gt;Don't flip production on day one. Re-baseline token counts, run your eval suite head-to-head against your current model, then canary a slice of traffic before scaling — and keep an escalation path to Opus 4.8 for tasks that need it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The skill underneath the model
&lt;/h2&gt;

&lt;p&gt;Here's the part the launch posts skip: a cheaper, more agentic model doesn't make anyone a better builder. It just makes the &lt;em&gt;consequences&lt;/em&gt; of your design bigger — cheaper to be right at scale, and cheaper to be confidently wrong at scale. Point Sonnet 5's autonomy at a vague spec and you get a fast, plausible wall of actions you didn't design and can't fully audit.&lt;/p&gt;

&lt;p&gt;The developers getting real leverage from this launch aren't the ones who memorized the new price-per-token. They're the ones who understand agent architecture, context engineering, evals, and cost modeling well enough to know &lt;em&gt;when&lt;/em&gt; the cheap-and-autonomous option is the right call and when it's a trap. That foundation — taught around real repositories with an interactive AI instructor, not slide decks — is what we build at our &lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;Eastern European AI education platform&lt;/a&gt;, including a dedicated, hands-on track on &lt;a href="https://cursuri-ai.ro/courses/claude-code-mastery-coding-agentic" rel="noopener noreferrer"&gt;agentic coding with Claude Code&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Claude Sonnet 5 is a genuinely significant release for builders, but not for the reason most coverage leads with. The story isn't a benchmark number — it's that near-frontier agentic capability just moved down a price tier, which changes which agent designs are economically worth shipping. The catch is the new tokenizer: the real saving is the price drop &lt;em&gt;minus&lt;/em&gt; token inflation, and you only learn the second number by measuring.&lt;/p&gt;

&lt;p&gt;So don't migrate on the headline. Re-baseline your tokens, run your evals, canary your traffic, and keep Opus 4.8 one route away for the work that needs it. Do that, and Sonnet 5 is one of the better deals in the 2026 model lineup. Skip it, and you'll find out the hard way — on your invoice.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written by the team at &lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt; — practical, hands-on AI engineering courses for developers and professionals across Eastern Europe, from agentic coding and AI agents to evals, context engineering, and the modern AI-native workflow.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources:&lt;/strong&gt; &lt;a href="https://www.anthropic.com/news/claude-sonnet-5" rel="noopener noreferrer"&gt;Introducing Claude Sonnet 5 — Anthropic&lt;/a&gt; · &lt;a href="https://www.anthropic.com/claude-sonnet-5-system-card" rel="noopener noreferrer"&gt;Claude Sonnet 5 System Card&lt;/a&gt; · &lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;Claude Platform — Pricing&lt;/a&gt; · &lt;a href="https://claude.com/pricing" rel="noopener noreferrer"&gt;Claude Pricing&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>cursor</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Cursor vs GitHub Copilot vs Claude Code: Which AI Coding Tool in 2026?</title>
      <dc:creator>galian</dc:creator>
      <pubDate>Mon, 29 Jun 2026 14:59:29 +0000</pubDate>
      <link>https://dev.to/cursuri-ai/cursor-vs-github-copilot-vs-claude-code-which-ai-coding-tool-in-2026-6c8</link>
      <guid>https://dev.to/cursuri-ai/cursor-vs-github-copilot-vs-claude-code-which-ai-coding-tool-in-2026-6c8</guid>
      <description>&lt;p&gt;If you write code for a living in 2026, you're not asking &lt;em&gt;whether&lt;/em&gt; to use an AI coding tool — you're asking &lt;em&gt;which one&lt;/em&gt;. And the three names that dominate every team's Slack debate are &lt;strong&gt;Cursor&lt;/strong&gt;, &lt;strong&gt;GitHub Copilot&lt;/strong&gt;, and &lt;strong&gt;Claude Code&lt;/strong&gt;. They look similar from a distance (type intent, get code) but they're built on three genuinely different bets about how software gets written.&lt;/p&gt;

&lt;p&gt;I've spent serious time in all three on real, multi-file, multi-repo work — not toy demos — and this is the comparison I wish someone had handed me before I burned a month figuring it out. I write and teach about agentic engineering at &lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt;, Eastern Europe's AI education platform, so I'll keep this grounded in how these tools actually behave in production, not in launch-day marketing.&lt;/p&gt;

&lt;p&gt;A note before we start: pricing and features in this category change almost monthly. Everything below is a mid-2026 snapshot — verify the current numbers on each tool's official page before you budget for a team.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR — three different philosophies
&lt;/h2&gt;

&lt;p&gt;Here's the one-sentence version of each, before we go deep:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; is an &lt;strong&gt;AI-native editor&lt;/strong&gt; — it rebuilt the IDE around the agent. Best for developers who want fast, fluid, in-the-flow generation with deep editor integration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GitHub Copilot&lt;/strong&gt; is the &lt;strong&gt;ecosystem play&lt;/strong&gt; — it lives where your code, issues, and PRs already are. Best for teams standardized on GitHub who want AI woven through the whole SDLC.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; is the &lt;strong&gt;terminal-first agent&lt;/strong&gt; — it treats the command line as the primary surface and excels at autonomous, multi-step, multi-file work. Best for engineers comfortable orchestrating agents rather than babysitting autocomplete.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of them is "the best." They optimize for different moments, and the real skill is knowing which to reach for. Let's break down why.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Cursor?
&lt;/h2&gt;

&lt;p&gt;Cursor is an AI-native IDE built as a fork of VS Code, so the editor feels instantly familiar — your extensions, keybindings, and themes mostly carry over. What's different is that the AI isn't bolted on as a plugin; the whole editing experience is designed around it.&lt;/p&gt;

&lt;p&gt;Its signature features:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tab completion&lt;/strong&gt; — a multi-line, context-aware autocomplete that predicts your &lt;em&gt;next edit&lt;/em&gt;, not just the next token. It's the feature people miss most when they switch away.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Composer&lt;/strong&gt; — Cursor's agentic, multi-file editing mode. You describe a change in natural language and it edits across files, runs commands, and iterates. Cursor now ships &lt;strong&gt;Composer 2.5&lt;/strong&gt;, its own model trained specifically for agentic coding, alongside routing to frontier models from Anthropic, OpenAI, and Google.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloud Agents&lt;/strong&gt; — introduced in the Cursor 3.5 release (May 20, 2026), these run in isolated cloud VMs with terminal and browser access, can work across multiple repos in parallel, and report results back to your IDE asynchronously. It's Cursor's answer to "I want the agent working while I do something else."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cursor's center of gravity is &lt;strong&gt;in-the-flow coding&lt;/strong&gt;: you stay in the editor, you see every diff, and the AI keeps pace with your thinking. It rewards developers who want speed without giving up granular control over the code.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is GitHub Copilot?
&lt;/h2&gt;

&lt;p&gt;Copilot is the most widely deployed of the three, and its biggest advantage is gravitational: it lives inside the tools and platform most teams already use. It runs in VS Code, JetBrains IDEs, Visual Studio, and on GitHub itself.&lt;/p&gt;

&lt;p&gt;By 2026 Copilot has grown well past autocomplete:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agent mode&lt;/strong&gt; became generally available across both VS Code and JetBrains in March 2026 (previously VS Code only) — a multi-step agent that plans, edits across files, and runs commands inside your editor.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The autonomous coding agent&lt;/strong&gt; is the standout. You assign a GitHub issue to Copilot, and it works asynchronously in the background — analyzing the repo, making changes, and opening a ready-to-review pull request. Assign, walk away, come back to a PR. It's the closest any mainstream tool comes to "fire-and-forget" feature work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agentic code review&lt;/strong&gt; gathers full project context before suggesting changes and can hand fixes straight to the coding agent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GitHub Spark&lt;/strong&gt; lets you describe an app in plain English and get generated code with a live preview.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The strategic point: Copilot's value isn't any single feature — it's that AI is now threaded through the entire GitHub-centric SDLC, from issue to PR to review. If your team lives on GitHub, that integration is hard to beat.&lt;/p&gt;

&lt;p&gt;One billing change worth flagging: as of June 1, 2026, GitHub moved to &lt;strong&gt;GitHub AI Credits&lt;/strong&gt; (token-based billing) in place of the older Premium Request Units. You're now billed by tokens processed at published model rates, which makes heavy agent usage more transparent — and easier to accidentally overspend if you're not watching.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Claude Code?
&lt;/h2&gt;

&lt;p&gt;Claude Code, from Anthropic, takes the opposite stance from Cursor: instead of building an editor, it makes the &lt;strong&gt;terminal&lt;/strong&gt; the primary surface (with IDE extensions available on top). That sounds minimalist until you see what it does with full shell access.&lt;/p&gt;

&lt;p&gt;Its defining strengths:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agentic, multi-file, repo-aware work&lt;/strong&gt; from the command line — it reads your codebase, makes coordinated changes across many files, runs your tests, and handles git operations and CI-aware workflows natively.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Subagents&lt;/strong&gt; — reusable agent configurations with their own custom prompts and tool access, so you can define a "reviewer," a "test-writer," or a "migration" agent and invoke it on demand.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent teams and multi-agent orchestration&lt;/strong&gt; — coordinate multiple agent sessions working in parallel, with an agent view dashboard to manage them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Claude Code runs on Anthropic's models — currently Claude Opus 4.8 as the default, with the newer Claude Fable 5 as the most capable tier — and it's deliberately model-opinionated rather than a router. The tradeoff is real: it's the most powerful for autonomous, complex tasks, and the least hand-holdy. It assumes you're comfortable thinking like an &lt;em&gt;orchestrator of agents&lt;/em&gt; rather than a writer of lines.&lt;/p&gt;

&lt;p&gt;A word of caution that applies to every agent platform but bites hardest here: &lt;strong&gt;parallel agents multiply your token spend.&lt;/strong&gt; Running ten agents at once consumes your quota roughly ten times faster. The autonomy is exhilarating; the bill is real. Set limits before you scale up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Head-to-head: the dimensions that actually matter
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The editing model
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; wins on &lt;em&gt;in-editor flow&lt;/em&gt;. Tab completion and inline diffs keep you in control of every change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Copilot&lt;/strong&gt; wins on &lt;em&gt;breadth of surface&lt;/em&gt; — it's good everywhere your code already is.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; wins on &lt;em&gt;autonomous depth&lt;/em&gt; — it goes furthest without supervision, but you give up the inline, line-by-line feel.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Agents and autonomy
&lt;/h3&gt;

&lt;p&gt;All three now have agents, but the philosophy differs. Cursor's Cloud Agents and Copilot's coding agent are both "assign work, get a result later." Claude Code goes further with explicit multi-agent orchestration and reusable subagents. If your work is increasingly &lt;em&gt;delegating&lt;/em&gt; rather than &lt;em&gt;typing&lt;/em&gt;, this is the dimension to weigh most — and it's exactly the shift that makes understanding &lt;a href="https://cursuri-ai.ro/courses/ai-agents-automatizare" rel="noopener noreferrer"&gt;AI agent architecture and automation&lt;/a&gt; a genuine career edge rather than a nice-to-have.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ecosystem and integration
&lt;/h3&gt;

&lt;p&gt;This is Copilot's home turf. The issue-to-PR loop, native code review, and presence across every major IDE make it the path of least resistance for GitHub-standardized teams. Cursor integrates deeply but inside &lt;em&gt;its&lt;/em&gt; editor; Claude Code integrates deeply with your &lt;em&gt;shell and git&lt;/em&gt;, which is either liberating or intimidating depending on your comfort with the command line.&lt;/p&gt;

&lt;h3&gt;
  
  
  Models
&lt;/h3&gt;

&lt;p&gt;Cursor routes across many frontier models and adds its own Composer model. Copilot offers a model picker. Claude Code is Anthropic-only by design. If model choice matters to you (and for some workloads it genuinely does), Cursor and Copilot give you more knobs; Claude Code bets that a tightly-integrated, top-tier model beats a buffet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pricing, side by side (mid-2026 snapshot)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Entry&lt;/th&gt;
&lt;th&gt;Mid tier&lt;/th&gt;
&lt;th&gt;Power / team&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cursor&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Hobby (free)&lt;/td&gt;
&lt;td&gt;Pro — $20/user/mo&lt;/td&gt;
&lt;td&gt;Teams — $40/user/mo (Standard), $120/user/mo (Premium)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GitHub Copilot&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;td&gt;Pro — $10/mo · Pro+ — $39/mo&lt;/td&gt;
&lt;td&gt;Max — $100/mo · Business / Enterprise seats&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Claude Code&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Pro — $20/mo&lt;/td&gt;
&lt;td&gt;Max 5× — $100/mo&lt;/td&gt;
&lt;td&gt;Max 20× — $200/mo · API pay-per-token&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A few honest caveats on cost:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Copilot&lt;/strong&gt; has the cheapest entry paid tier ($10), but token-based AI Credits mean heavy agent use can climb fast beyond the included allotment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cursor's&lt;/strong&gt; $20 Pro includes a fixed amount of frontier-model usage; power users hit the ceiling and either upgrade or switch to its cheaper Auto/Composer routing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code's&lt;/strong&gt; Max tiers are priced for sustained, agent-heavy sessions — and again, parallel agents are a multiplier, not an add.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Prices and tiers shift constantly in this category. Treat the table as a snapshot, not a quote, and confirm before committing a team budget.&lt;/p&gt;

&lt;h2&gt;
  
  
  So which one should you choose?
&lt;/h2&gt;

&lt;p&gt;Here's the honest, persona-based answer:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Choose Cursor if&lt;/strong&gt; you want the best in-editor experience, you value fast inline generation and tight control over every diff, and you're happy living inside a (very good) VS Code fork. It's the most natural upgrade for a developer who loves their editor and wants AI to keep pace with their flow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Choose GitHub Copilot if&lt;/strong&gt; your team is standardized on GitHub and you want AI woven through the entire lifecycle — issues, PRs, reviews — across whatever IDEs your team already uses. The issue-to-PR autonomous agent alone can change how a team ships. It's the safest institutional bet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Choose Claude Code if&lt;/strong&gt; you're comfortable in the terminal, your work skews toward complex multi-file refactors and autonomous tasks, and you want to orchestrate agents rather than supervise autocomplete. It has the highest ceiling for autonomy — and asks the most of you in return.&lt;/p&gt;

&lt;p&gt;And the answer most senior engineers actually land on? &lt;strong&gt;More than one.&lt;/strong&gt; Plenty of us keep Cursor open for flow-state editing, lean on Copilot inside the GitHub workflow, and fire up Claude Code for the gnarly autonomous jobs. The tools overlap, but they're not redundant — they're a toolkit. The real meta-skill isn't loyalty to one editor; it's &lt;strong&gt;fluency across the category&lt;/strong&gt; so you instinctively reach for the right one per task.&lt;/p&gt;

&lt;h2&gt;
  
  
  The skill underneath the tools
&lt;/h2&gt;

&lt;p&gt;Here's the uncomfortable truth that the demos hide: these tools amplify the engineer you already are. Point a powerful agent at a vague intent and you get a fast, confident wall of code you didn't design and can't fully maintain. The developers getting outsized leverage from Cursor, Copilot, and Claude Code aren't the ones who learned the keyboard shortcuts — they're the ones who understand agent architecture, context engineering, and how to specify intent precisely enough that autonomy becomes an asset instead of a liability.&lt;/p&gt;

&lt;p&gt;That foundation is exactly what we build at &lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;our AI education platform&lt;/a&gt; for Eastern Europe — practical, project-based courses taught around real repositories with an interactive AI instructor, not slide decks. If you want to go from "I use these tools" to "I get serious leverage from them," we maintain dedicated, hands-on tracks for &lt;a href="https://cursuri-ai.ro/courses/cursor-pro" rel="noopener noreferrer"&gt;using Cursor as a pro&lt;/a&gt; and for &lt;a href="https://cursuri-ai.ro/courses/claude-code-mastery-coding-agentic" rel="noopener noreferrer"&gt;agentic coding with Claude Code&lt;/a&gt; — both built around real multi-file, real-repo workflows rather than toy examples.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;In 2026, "AI coding tool" isn't one product category — it's three philosophies wearing similar clothes. Cursor bet on the editor, Copilot bet on the ecosystem, and Claude Code bet on the terminal-native agent. Each is genuinely excellent at the thing it optimized for, and genuinely compromised at the things it didn't.&lt;/p&gt;

&lt;p&gt;So don't ask "which is best." Ask "best at what, for whom, doing which task" — and then build the judgment to switch fluently between them. That judgment, not the tool, is what compounds over a career. Try each one on a real feature, not a demo, and you'll feel the differences fast.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written by the team at &lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt; — practical, hands-on AI engineering courses for developers and professionals across Eastern Europe, from agentic coding and AI agents to context engineering and the modern AI-native IDE workflow.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources:&lt;/strong&gt; &lt;a href="https://cursor.com/docs/models-and-pricing" rel="noopener noreferrer"&gt;Cursor Models &amp;amp; Pricing&lt;/a&gt; · &lt;a href="https://github.com/features/copilot/plans" rel="noopener noreferrer"&gt;GitHub Copilot Plans &amp;amp; Pricing&lt;/a&gt; · &lt;a href="https://docs.github.com/en/copilot/get-started/plans" rel="noopener noreferrer"&gt;GitHub Copilot Plans (Docs)&lt;/a&gt; · &lt;a href="https://claude.com/pricing" rel="noopener noreferrer"&gt;Claude Pricing&lt;/a&gt; · &lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;Claude Platform Docs — Pricing&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Kiro: A Practical Guide to AWS's Spec-Driven Agentic IDE"</title>
      <dc:creator>galian</dc:creator>
      <pubDate>Fri, 26 Jun 2026 21:45:36 +0000</pubDate>
      <link>https://dev.to/galian/kiro-a-practical-guide-to-awss-spec-driven-agentic-ide-26o9</link>
      <guid>https://dev.to/galian/kiro-a-practical-guide-to-awss-spec-driven-agentic-ide-26o9</guid>
      <description>&lt;p&gt;If you've spent any time with AI coding assistants, you know the failure mode: you write a vague prompt, the agent generates a wall of plausible-looking code, and twenty minutes later you're debugging something you didn't design and don't fully understand. &lt;strong&gt;Kiro&lt;/strong&gt;, the agentic IDE from AWS, is a bet that the fix isn't a smarter autocomplete — it's making a &lt;em&gt;specification&lt;/em&gt; the unit of work instead of a prompt.&lt;/p&gt;

&lt;p&gt;I've been digging into how Kiro actually works, and this is the practical guide I wish I'd had on day one: what spec-driven development really means, how agent hooks and steering files change your workflow, where Kiro fits next to tools like Cursor and Claude Code, and when it's worth it. I write and teach about agentic engineering at &lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt;, Eastern Europe's AI education platform, so I'll keep this grounded in how these tools behave in real projects rather than in launch-day hype.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Kiro?
&lt;/h2&gt;

&lt;p&gt;Kiro is an agentic IDE built on the Code OSS platform — the same open-source foundation behind VS Code — which means the editor itself feels immediately familiar. What's different is the engine. Instead of treating each request as a one-off chat turn, Kiro is designed to turn a high-level prompt into a structured &lt;strong&gt;spec&lt;/strong&gt;, then drive implementation, tests, and documentation from that spec.&lt;/p&gt;

&lt;p&gt;The headline idea, in Kiro's own framing, is "moving beyond AI coding to agentic engineering." That sounds like marketing until you see the artifacts it produces. A feature request doesn't become a blob of code — it becomes three reviewable files: requirements, design, and tasks. You stay in the loop at each stage. The agent does the typing; you keep the judgment.&lt;/p&gt;

&lt;p&gt;It's worth being precise about what Kiro is &lt;em&gt;not&lt;/em&gt;: it isn't an AWS cloud service you provision in a console, and it doesn't lock you into AWS infrastructure to write code. It's a desktop IDE. You can point it at any project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Spec-driven development: the core idea
&lt;/h2&gt;

&lt;p&gt;Most AI coding tools optimize for speed-to-first-keystroke. Spec-driven development optimizes for &lt;em&gt;correctness-to-intent&lt;/em&gt; — does the code match what you actually meant? Kiro does this by formalizing the part of engineering we usually skip when we're moving fast: writing down what we're building before we build it.&lt;/p&gt;

&lt;p&gt;When you describe a feature, Kiro generates a spec in three phases:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Requirements
&lt;/h3&gt;

&lt;p&gt;Kiro turns your prompt into user stories with explicit acceptance criteria, written in &lt;strong&gt;EARS notation&lt;/strong&gt; (Easy Approach to Requirements Syntax). EARS is a lightweight, real technique for writing testable requirements — patterns like &lt;em&gt;"When [trigger], the system shall [response]"&lt;/em&gt;. The value is that ambiguity gets surfaced &lt;em&gt;before&lt;/em&gt; code exists. If your one-line prompt was underspecified, you'll see it in the requirements draft and can correct it in seconds, not after a debugging session.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Design
&lt;/h3&gt;

&lt;p&gt;Next, Kiro produces a technical design: the architecture, the components, the data flow, and the implementation approach. This is the document a senior engineer would normally write (or wish a junior had written) before touching the codebase. You review it, push back, and refine.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Tasks
&lt;/h3&gt;

&lt;p&gt;Finally, the design becomes a sequenced task list — discrete, trackable units of work the agent implements in order. Because tasks are explicit, you get accountability: you can see what's done, what's in progress, and what's left, instead of trusting a black box.&lt;/p&gt;

&lt;p&gt;The payoff is maintainability. A spec that lives in your repo is documentation that doesn't rot, because it &lt;em&gt;is&lt;/em&gt; the thing the agent built from. Six months later, the requirements and design files explain &lt;em&gt;why&lt;/em&gt; the code looks the way it does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agent hooks: automation that runs itself
&lt;/h2&gt;

&lt;p&gt;The second pillar is &lt;strong&gt;agent hooks&lt;/strong&gt; — automated triggers that fire agent prompts or shell commands when something happens in your IDE. Instead of remembering to run the linter, regenerate tests, or scan for secrets, you wire those actions to events once and forget about them.&lt;/p&gt;

&lt;p&gt;Hooks can be triggered by:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;File events&lt;/strong&gt; — a file is created, saved, or deleted&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompt and agent lifecycle events&lt;/strong&gt; — prompt submit, agent stop, pre/post tool use&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spec task events&lt;/strong&gt; — before or after a task executes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Manual triggers&lt;/strong&gt; — a button you press on demand&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Under the hood, hooks are just JSON files. Workspace-level hooks live in &lt;code&gt;.kiro/hooks/&lt;/code&gt;, and user-level hooks in &lt;code&gt;~/.kiro/hooks/&lt;/code&gt;. You can create them three ways: describe what you want in plain English and let Kiro generate the JSON, fill out a form, or write the JSON by hand. The practical version of this: every time you save a file, a hook can run your tests and a security scan automatically, so problems surface the moment they're introduced — not in CI an hour later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Steering files: stop repeating yourself
&lt;/h2&gt;

&lt;p&gt;If you've ever pasted "remember, we use tabs not spaces, we use Vitest not Jest, and never import from the legacy module" into chat for the hundredth time, &lt;strong&gt;steering files&lt;/strong&gt; are the fix. Steering gives Kiro persistent knowledge about your project through markdown files, so your conventions, libraries, and standards are applied consistently without re-explaining them every session.&lt;/p&gt;

&lt;p&gt;Steering files can be scoped two ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Workspace steering&lt;/strong&gt; lives in &lt;code&gt;.kiro/steering/&lt;/code&gt; and applies only to that project&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Global steering&lt;/strong&gt; applies across everything you build&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is essentially context engineering applied to a coding agent — encoding the durable knowledge an agent needs so it behaves like a teammate who's read your style guide, not a contractor seeing the repo for the first time. If you want to go deep on the discipline behind this, persistent memory and context strategy for agents is exactly what our &lt;a href="https://cursuri-ai.ro/courses/context-engineering-memorie-agenti" rel="noopener noreferrer"&gt;context engineering and agent memory course&lt;/a&gt; covers end to end.&lt;/p&gt;

&lt;h2&gt;
  
  
  MCP and agentic chat
&lt;/h2&gt;

&lt;p&gt;Beyond specs, hooks, and steering, Kiro ships the features you'd expect from a modern AI editor. It supports the &lt;strong&gt;Model Context Protocol (MCP)&lt;/strong&gt; for connecting external tools and data sources to the agent, and it includes an agentic chat with context providers for files, URLs, and docs for the ad-hoc work that doesn't justify a full spec.&lt;/p&gt;

&lt;p&gt;MCP support matters more than it sounds. It's the open standard that lets an agent reach your database, your ticketing system, your internal docs — without bespoke glue for each one. If MCP is new to you, building and integrating MCP servers is its own skill set; our &lt;a href="https://cursuri-ai.ro/courses/mcp-model-context-protocol" rel="noopener noreferrer"&gt;MCP course&lt;/a&gt; walks through standing up real servers and wiring them into agentic workflows.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your first hour with Kiro
&lt;/h2&gt;

&lt;p&gt;The fastest way to understand the workflow is to feel the loop once on a real, small feature. In practice it looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Describe the feature in plain language.&lt;/strong&gt; Not "build an app" — something concrete like "add an endpoint that returns a user's last five orders, paginated."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review the requirements.&lt;/strong&gt; Kiro drafts user stories and acceptance criteria. This is where you catch the ambiguity: did you mean five orders total, or five per page? Fix it in the spec, where it costs nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review the design.&lt;/strong&gt; Check that the proposed architecture matches your codebase's conventions — and if it doesn't, that's a sign your steering files need to capture those conventions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Let it work the task list.&lt;/strong&gt; The agent implements tasks in sequence; you watch and intervene where judgment is needed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wire one hook.&lt;/strong&gt; Even a single "run tests on save" hook changes how the session feels — feedback becomes immediate instead of deferred.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Do that once and the abstract pitch — "specs as the unit of work" — turns concrete. The discipline isn't heavy; it's front-loaded, and the front-loading is where the bugs you didn't ship were quietly avoided.&lt;/p&gt;

&lt;h2&gt;
  
  
  Kiro vs. vibe coding
&lt;/h2&gt;

&lt;p&gt;"Vibe coding" — prompting your way to an app on feel, accepting whatever the model produces — is genuinely useful for prototypes, throwaway scripts, and learning. It's also where a lot of teams get burned when that "prototype" quietly becomes production.&lt;/p&gt;

&lt;p&gt;Kiro is, in a sense, the structured opposite. The spec phase forces the requirements-and-design thinking that vibe coding skips. That doesn't make vibe coding wrong — it makes them tools for different moments. Reaching for a spec to build a one-off script is overkill; vibe-coding a payment flow is asking for trouble. Knowing &lt;em&gt;which mode fits which task&lt;/em&gt; is the actual skill, and it's the through-line of our &lt;a href="https://cursuri-ai.ro/courses/vibe-coding-prompt-la-aplicatie" rel="noopener noreferrer"&gt;vibe coding course&lt;/a&gt;, which treats prompt-to-app speed and structured engineering as complementary, not rival, philosophies.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Kiro compares to other agentic tools
&lt;/h2&gt;

&lt;p&gt;Kiro isn't alone — the agentic IDE space is crowded, and the tools overlap. A few honest distinctions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; is an AI-native editor built around fast in-editor generation, multi-file edits, and an agent mode. Its center of gravity is fluid, in-the-flow coding.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; is a terminal-first agentic tool that excels at multi-file changes, git operations, and CI-aware work from the command line.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kiro&lt;/strong&gt; distinguishes itself by making the &lt;em&gt;spec&lt;/em&gt; the artifact — front-loading requirements and design before implementation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These aren't mutually exclusive; plenty of engineers use more than one and switch by task. The meta-skill is fluency across the category rather than loyalty to one editor. If you want a structured path through these tools, we maintain dedicated, hands-on courses on &lt;a href="https://cursuri-ai.ro/courses/claude-code-mastery-coding-agentic" rel="noopener noreferrer"&gt;agentic coding with Claude Code&lt;/a&gt; and on &lt;a href="https://cursuri-ai.ro/courses/cursor-pro" rel="noopener noreferrer"&gt;Cursor as a pro&lt;/a&gt; — both built around real multi-file, real-repo workflows rather than toy demos.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pricing
&lt;/h2&gt;

&lt;p&gt;At the time of writing, Kiro uses a credit-based model measured in &lt;strong&gt;agent interactions&lt;/strong&gt;, with no daily or weekly rate limits and pre-paid overages so you don't hit a hard wall mid-task:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Free&lt;/strong&gt; — 50 agent interactions per user per month (fine for experimentation, not serious daily work)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pro&lt;/strong&gt; — $19 per user per month for 1,000 agent interactions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pro+&lt;/strong&gt; — $39 per user per month for 3,000 interactions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Pricing and tiers for tools in this category change often, so verify the current numbers on Kiro's official pricing page before you budget for a team. Treat the figures above as a snapshot, not a contract.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Kiro is worth it — and when it isn't
&lt;/h2&gt;

&lt;p&gt;Spec-driven development has a cost: the spec phase is overhead. That overhead pays off when the work is durable and shared, and it's pure friction when the work is disposable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Kiro shines when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You're building features meant to live and be maintained, not prototypes you'll throw away&lt;/li&gt;
&lt;li&gt;More than one person (or one agent) touches the codebase and conventions matter&lt;/li&gt;
&lt;li&gt;You want an auditable trail of &lt;em&gt;why&lt;/em&gt; the code is the way it is&lt;/li&gt;
&lt;li&gt;You're tired of re-explaining your standards every session&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Reach for something lighter when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You're exploring, prototyping, or scripting something you'll delete tomorrow&lt;/li&gt;
&lt;li&gt;The task is small enough that writing the spec costs more than writing the code&lt;/li&gt;
&lt;li&gt;You just need a quick answer or a one-file change&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The honest take: spec-driven development is a discipline, and Kiro is tooling that makes the discipline cheaper to follow. The tool won't supply the engineering judgment — it removes the excuse not to apply it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Want to go deeper?
&lt;/h2&gt;

&lt;p&gt;Tools like Kiro lower the cost of doing engineering properly, but they reward people who already understand specs, agent architecture, context management, and the MCP ecosystem underneath. That foundation is what turns an agentic IDE from a faster autocomplete into genuine leverage.&lt;/p&gt;

&lt;p&gt;At &lt;strong&gt;&lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt;&lt;/strong&gt;, Eastern Europe's AI education platform, we build practical, project-based courses on exactly this stack — agentic coding, MCP, context engineering, and the modern AI-native IDE workflow — taught around real repositories with an interactive AI instructor, not slide decks. If Kiro made you curious about &lt;em&gt;agentic engineering&lt;/em&gt; as a craft rather than a buzzword, that's the rabbit hole our catalog is built for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Kiro's core bet is simple and, I think, correct: the bottleneck in AI-assisted development was never typing speed — it was the gap between what you meant and what the model built. By making specs the unit of work, adding agent hooks for automation and steering files for persistent context, Kiro turns "AI coding" into something closer to engineering with an agent.&lt;/p&gt;

&lt;p&gt;It won't replace judgment, and it isn't the right tool for every task. But for durable, maintainable software built with an AI in the loop, spec-driven development is a genuinely different — and more accountable — way to work. Try it on a real feature, not a toy, and you'll feel the difference fast.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written by the team at &lt;a href="https://cursuri-ai.ro" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt; — practical, hands-on AI engineering courses for developers and professionals across Eastern Europe, from agentic coding and MCP to context engineering and AI-native IDE workflows.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources:&lt;/strong&gt; &lt;a href="https://kiro.dev/" rel="noopener noreferrer"&gt;kiro.dev&lt;/a&gt; · &lt;a href="https://kiro.dev/docs/specs/" rel="noopener noreferrer"&gt;Kiro Specs docs&lt;/a&gt; · &lt;a href="https://kiro.dev/docs/hooks/" rel="noopener noreferrer"&gt;Kiro Hooks docs&lt;/a&gt; · &lt;a href="https://kiro.dev/docs/steering/" rel="noopener noreferrer"&gt;Kiro Steering docs&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Stop Vibe-Checking Your LLM: A Developer's Guide to Evals</title>
      <dc:creator>galian</dc:creator>
      <pubDate>Mon, 22 Jun 2026 08:24:22 +0000</pubDate>
      <link>https://dev.to/cursuri-ai/stop-vibe-checking-your-llm-a-developers-guide-to-evals-3oed</link>
      <guid>https://dev.to/cursuri-ai/stop-vibe-checking-your-llm-a-developers-guide-to-evals-3oed</guid>
      <description>&lt;p&gt;You tweaked the system prompt, ran the same two test questions you always run, the answers looked good, and you shipped. A week later support is forwarding you screenshots of the model confidently doing the exact thing your prompt was supposed to stop. You never saw it, because "did it get better?" was answered by vibes.&lt;/p&gt;

&lt;p&gt;This is the single most common failure mode in shipping LLM features, and it has nothing to do with which model you picked. &lt;strong&gt;If your only quality gate is reading a handful of outputs and nodding, every change you make is a coin flip.&lt;/strong&gt; You can't tell whether a prompt edit helped, hurt, or just moved the failures somewhere you didn't look. Evals are how you replace the nod with a number.&lt;/p&gt;

&lt;p&gt;This is a practical guide to building that number — from a 30-row eval set you can write this afternoon, through code-based checks and LLM-as-judge scoring, to wiring the whole thing into CI so regressions get blocked instead of discovered by users. No new framework to adopt; just the discipline that separates a demo from a system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why you can't just &lt;code&gt;assert output == expected&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Traditional tests work because the output space is small and exact. &lt;code&gt;add(2, 2)&lt;/code&gt; is &lt;code&gt;4&lt;/code&gt; or it's a bug. LLM output breaks all three assumptions that make &lt;code&gt;assertEqual&lt;/code&gt; work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It's non-deterministic.&lt;/strong&gt; The same prompt can produce different text on two calls. Even at temperature 0 you are not guaranteed byte-identical output across runs or model versions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It's open-ended.&lt;/strong&gt; "Summarize this ticket" has thousands of correct answers. None of them are string-equal to your reference, and that's fine — a good summary isn't &lt;em&gt;the&lt;/em&gt; summary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It fails softly.&lt;/strong&gt; A wrong answer isn't a stack trace. It's a fluent, plausible, well-formatted paragraph that happens to be incorrect. Nothing crashes. Nothing logs an error.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the goal of an eval isn't "is the output identical to the expected string." It's "does the output satisfy the &lt;em&gt;properties&lt;/em&gt; I care about" — is it grounded in the provided context, does it stay on policy, does it actually answer the question, is it valid JSON. You're testing behavior against criteria, not bytes against bytes. Once that clicks, the rest is mechanics.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the eval set, not the metric
&lt;/h2&gt;

&lt;p&gt;The instinct is to reach for a fancy metric first. Wrong order. The asset that makes everything else work is a small, representative &lt;strong&gt;eval set&lt;/strong&gt;: a fixed collection of inputs paired with what a good output looks like (or the criteria a good output must meet). This is your golden dataset, your regression suite, your source of truth.&lt;/p&gt;

&lt;p&gt;You do not need thousands of examples to start. &lt;strong&gt;Thirty to fifty well-chosen pairs&lt;/strong&gt; turn LLM tuning from vibes into engineering, because now every change is measured against the same fixed bar. Build the set like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Mine real failures.&lt;/strong&gt; Every time the system gets something wrong in dev or prod, that exact input goes into the eval set with a note on what the right behavior is. Your bug reports &lt;em&gt;are&lt;/em&gt; your test cases. This is the highest-signal source you have.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cover the categories, not just the happy path.&lt;/strong&gt; Easy questions, ambiguous ones, adversarial ones, out-of-scope ones ("I don't know" is the correct answer and you should test that it says so), and the edge cases specific to your domain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Freeze it and version it.&lt;/strong&gt; The eval set lives in your repo next to the code. When you add a case, that's a commit. A moving target can't measure progress.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep a holdout.&lt;/strong&gt; If you start tuning prompts &lt;em&gt;against&lt;/em&gt; the eval set, you'll overfit to it. Keep a slice you don't look at until you think you're done.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A minimal eval set is just data — JSON, a CSV, a Python list. Here's the shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# evals/dataset.py
&lt;/span&gt;&lt;span class="n"&gt;EVAL_SET&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refund-window-basic&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;question&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What is our refund window?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;context&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Refunds are accepted within 14 days of purchase.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;expected&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;14 days&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;must_not_say&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;30 days&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;no refunds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;out-of-scope&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;question&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s the weather in Cluj tomorrow?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;context&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Refunds are accepted within 14 days of purchase.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;expected&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;REFUSE&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# correct behavior: decline, don't invent
&lt;/span&gt;    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="c1"&gt;# ... 30-50 of these, grown from real failures
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the foundation. Everything below scores outputs against this set.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two halves of every LLM eval
&lt;/h2&gt;

&lt;p&gt;Separate two questions that get mushed together when you eval by eyeball, because they have different fixes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Did the system retrieve / set up the right context?&lt;/strong&gt; (a retrieval or pipeline question)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Given that context, did the model produce a good answer?&lt;/strong&gt; (a generation question)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you're building RAG, the first half is its own discipline — measuring recall@k and precision@k on questions with known relevant documents tells you whether the right chunk even reached the prompt. That's a deep enough topic that it deserves its own treatment; a dedicated &lt;a href="https://cursuri-ai.ro/courses/rag-retrieval-augmented-generation" rel="noopener noreferrer"&gt;course on RAG and retrieval-augmented generation&lt;/a&gt; spends real time there, and the failure modes are different from the ones below. This guide focuses on the second half: scoring the generated answer. The techniques split into two families — &lt;strong&gt;code-based checks&lt;/strong&gt; and &lt;strong&gt;model-based judges&lt;/strong&gt; — and you want both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code-based checks: cheaper and more reliable than you think
&lt;/h2&gt;

&lt;p&gt;Before you reach for an LLM to grade an LLM, a surprising amount of quality is checkable with plain code. These checks are deterministic, free, instant, and never hallucinate. Use them for everything they can cover:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Structural validity.&lt;/strong&gt; If the output should be JSON matching a schema, validate it. A response that doesn't parse is a hard failure, no judgment call needed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Must-contain / must-not-contain.&lt;/strong&gt; The answer about a 14-day refund window must contain "14" and must not contain "30." Keyword and regex assertions catch a whole class of factual regressions for free.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Format and bounds.&lt;/strong&gt; Length limits, required citations present, no leaked system-prompt text, no forbidden phrases (the "as an AI language model" tax), valid enum values.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Semantic similarity.&lt;/strong&gt; For open-ended answers, embed the output and your reference answer and check cosine similarity passes a threshold. It's fuzzy, but it catches "the answer wandered off topic" without needing a judge model.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# evals/checks.py
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check_structural&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;schema_keys&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;JSONDecodeError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;schema_keys&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check_must_not_say&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;banned&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;low&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;any&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;low&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;banned&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rule of thumb: &lt;strong&gt;anything a regex or a schema can catch, don't pay a model to catch.&lt;/strong&gt; Reserve the expensive, fuzzy judge for the genuinely subjective stuff.&lt;/p&gt;

&lt;h2&gt;
  
  
  LLM-as-judge: powerful, biased, and fixable
&lt;/h2&gt;

&lt;p&gt;For the subjective half — "is this answer faithful to the source?", "is this helpful?", "is the tone right?" — you use a strong model to grade outputs. This is &lt;strong&gt;LLM-as-judge&lt;/strong&gt;, and it scales human-quality judgment to thousands of examples for the price of an API call. Two metrics carry most of the weight for RAG-style apps:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Faithfulness / groundedness&lt;/strong&gt; — does every claim in the answer trace back to the provided context, or did the model invent things? This is your hallucination detector.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Answer relevance&lt;/strong&gt; — does the response actually address the question that was asked, or is it a fluent dodge?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The catch: &lt;strong&gt;LLM judges have well-documented biases&lt;/strong&gt;, and if you ignore them your eval numbers are noise dressed up as signal. The big ones, all reported in the research on using models as evaluators:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Position bias&lt;/strong&gt; — when comparing two answers, judges favor the one shown first (or in a fixed slot) regardless of quality.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verbosity bias&lt;/strong&gt; — judges tend to rate longer, more elaborate answers higher even when a short answer is more correct.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Self-preference&lt;/strong&gt; — a judge model can favor text written in its own style or by its own family.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You don't abandon the technique; you engineer around the bias:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Score against a rubric, not a vibe.&lt;/strong&gt; Ask for a 1–5 score with explicit criteria for each level, and require the judge to output its reasoning &lt;em&gt;before&lt;/em&gt; the score. A judge forced to justify itself is more consistent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For pairwise comparisons, randomize and swap.&lt;/strong&gt; Run each comparison twice with the order flipped; only count it as a win if the judge picks the same answer both times. This cancels position bias directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Calibrate against humans.&lt;/strong&gt; Hand-label 20–30 examples yourself, run the judge on them, and check it agrees with you. If it doesn't, fix the rubric before trusting it on 2,000. An uncalibrated judge is a random number generator with good grammar.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use a strong model as the judge.&lt;/strong&gt; Grading is harder than answering. Use a current frontier model for the judge even if your app runs on a smaller, cheaper one.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# evals/judge.py — sketch of a rubric-based faithfulness judge
&lt;/span&gt;&lt;span class="n"&gt;JUDGE_PROMPT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;You are grading whether an ANSWER is fully supported by the CONTEXT.

CONTEXT:
{context}

ANSWER:
{answer}

Rules:
- A claim is &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;supported&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; only if the CONTEXT states or directly implies it.
- Outside knowledge does NOT count as support.

First write one sentence of reasoning. Then output a JSON object:
{{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reasoning&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;faithful&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;: true|false}}&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;judge_faithfulness&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;JUDGE_PROMPT&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;format&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;faithful&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Designing judges that hold up — picking the rubric, calibrating, knowing when a model is the wrong tool for the grade — is exactly the muscle a &lt;a href="https://cursuri-ai.ro/courses/ai-evals-llm-productie" rel="noopener noreferrer"&gt;course on AI evals in production&lt;/a&gt; builds, because it's the difference between "the new prompt feels better" and "faithfulness went from 0.78 to 0.91 on the holdout."&lt;/p&gt;

&lt;h2&gt;
  
  
  Wire it into CI, or it won't survive contact with deadlines
&lt;/h2&gt;

&lt;p&gt;An eval you run by hand when you remember to is an eval you'll stop running the week things get busy. The whole point is to make regressions &lt;em&gt;impossible to ship silently&lt;/em&gt;, and that means the eval runs automatically on every change to a prompt, a retrieval setting, or a model version.&lt;/p&gt;

&lt;p&gt;The pattern is a regression gate: run the eval set, compute the aggregate score, and &lt;strong&gt;fail the build if the score drops below a threshold&lt;/strong&gt; (or below the last known-good baseline). It looks like an ordinary test suite, because that's what it is.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# tests/test_evals.py
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pytest&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;evals.dataset&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;EVAL_SET&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;evals.checks&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;check_must_not_say&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;myapp&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;answer_question&lt;/span&gt;

&lt;span class="n"&gt;PASS_THRESHOLD&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.90&lt;/span&gt;  &lt;span class="c1"&gt;# 90% of eval cases must pass to ship
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;run_case&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;answer_question&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;question&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;context&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;expected&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;REFUSE&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;i don&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;t know&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;can&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;t&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;check_must_not_say&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;must_not_say&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[])):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;expected&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_eval_suite_meets_threshold&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;run_case&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;EVAL_SET&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;failed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ok&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;EVAL_SET&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;PASS_THRESHOLD&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Eval score &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; below &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;PASS_THRESHOLD&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;. Failed: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;failed&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few practical notes that keep this sane in CI:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pin the model version.&lt;/strong&gt; Provider model IDs update, and an unpinned model means your eval baseline shifts under you for reasons unrelated to your code. Pin it, and treat a model upgrade as its own deliberate eval run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Budget for cost and flakiness.&lt;/strong&gt; LLM calls cost money and occasionally time out. Cache where you can, run the judge-heavy suite on a schedule rather than every commit if needed, and set a slightly forgiving threshold so one stochastic blip doesn't red-X a good PR.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log the failures, not just the score.&lt;/strong&gt; When the gate trips, the output should name &lt;em&gt;which&lt;/em&gt; cases regressed so the fix is obvious. A bare "0.86 &amp;lt; 0.90" sends you debugging blind.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now a prompt change is a PR with a number attached. The reviewer sees faithfulness went up and refusal rate held steady, or they see it tanked and the build is red. That's the entire difference between hoping and knowing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five mistakes that quietly poison your evals
&lt;/h2&gt;

&lt;p&gt;Even teams that build evals often undermine them. Watch for these:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Testing only the happy path.&lt;/strong&gt; If every case in your set is a question the system already answers well, your score is a flattering lie. Adversarial and out-of-scope cases are where the signal is.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tuning on your test set.&lt;/strong&gt; Optimize prompts against the same examples you grade on and you'll overfit to them. Keep a holdout you don't peek at.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An uncalibrated judge.&lt;/strong&gt; Trusting an LLM judge you never checked against your own labels is trusting a number you made up. Calibrate first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One giant blended score.&lt;/strong&gt; A single average hides that faithfulness improved while refusals broke. Track metrics &lt;em&gt;separately&lt;/em&gt; so a regression in one can't be masked by a gain in another.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Letting the set rot.&lt;/strong&gt; Your product changes; cases that no longer reflect real usage drag the signal down. Prune and grow the set as part of normal work, the same way you maintain any test suite.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of these are exotic. They're the eval equivalent of not testing error paths — obvious in hindsight, easy to skip under deadline.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this connects to the rest of your LLM stack
&lt;/h2&gt;

&lt;p&gt;Evals aren't a standalone chore; they're the measurement layer that makes every other improvement legible. When you tighten a prompt, evals tell you if it worked — which is why &lt;a href="https://cursuri-ai.ro/courses/prompt-engineering-masterclass" rel="noopener noreferrer"&gt;structured prompt engineering&lt;/a&gt; and a real eval loop are two halves of the same skill. When you redesign what goes into the context window — what to include, what to cut, how to order it — evals are how you know the redesign helped rather than just &lt;em&gt;felt&lt;/em&gt; cleaner; that discipline of deciding what earns a place in the prompt is increasingly called context engineering and has &lt;a href="https://cursuri-ai.ro/courses/context-engineering-memorie-agenti" rel="noopener noreferrer"&gt;its own dedicated course&lt;/a&gt;. And when you wire up function calling, multi-tool orchestration, and the production concerns of a real integration, evals are what keep the whole pipeline honest as it grows — the kind of end-to-end build covered in a deeper &lt;a href="https://cursuri-ai.ro/courses/advanced-llm-integration" rel="noopener noreferrer"&gt;course on advanced LLM integration&lt;/a&gt;. The pattern is always the same: build the measurement first, then every change becomes verifiable instead of hopeful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The teams whose LLM features actually hold up in production aren't using a secret model or a magic prompt. They're disciplined about measurement. They have a versioned eval set grown from real failures, code-based checks for everything a regex can catch, calibrated LLM judges for the subjective rest, and a CI gate that blocks regressions before users find them.&lt;/p&gt;

&lt;p&gt;Start smaller than you think you can. Write thirty cases this afternoon — half of them things your system currently gets &lt;em&gt;wrong&lt;/em&gt; — add three code checks and one rubric-based judge, and put a threshold in your test suite. The first time a red build stops you from shipping a prompt change that would have quietly broken refusals, you'll never go back to vibe-checking. That's the moment an LLM demo becomes an LLM system people can trust.&lt;/p&gt;

&lt;p&gt;The courses linked throughout are part of &lt;a href="https://cursuri-ai.ro/courses/ai-evals-llm-productie" rel="noopener noreferrer"&gt;Cursuri-AI.ro&lt;/a&gt;, an AI-learning platform with hands-on, current tracks on evaluating AI systems in production, prompt engineering, RAG, and advanced LLM integration.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Sources &amp;amp; further reading:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Zheng et al. — &lt;a href="https://arxiv.org/abs/2306.05685" rel="noopener noreferrer"&gt;Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena&lt;/a&gt; (documents position, verbosity, and self-enhancement bias in LLM judges)&lt;/li&gt;
&lt;li&gt;Liu et al. — &lt;a href="https://arxiv.org/abs/2303.16634" rel="noopener noreferrer"&gt;G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Liang et al. — &lt;a href="https://arxiv.org/abs/2211.09110" rel="noopener noreferrer"&gt;Holistic Evaluation of Language Models (HELM)&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This article is educational content. Techniques and tooling evolve quickly; validate approaches against your own data and current library documentation.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
