<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Bry</title>
    <description>The latest articles on DEV Community by Bry (@brywritescode).</description>
    <link>https://dev.to/brywritescode</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4009547%2Fe6e5c498-b824-493d-b06e-fef6445e4ef7.jpeg</url>
      <title>DEV Community: Bry</title>
      <link>https://dev.to/brywritescode</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/brywritescode"/>
    <language>en</language>
    <item>
      <title>Serverless vs Containers: A Manager's Guide</title>
      <dc:creator>Bry</dc:creator>
      <pubDate>Tue, 06 Oct 2026 15:00:00 +0000</pubDate>
      <link>https://dev.to/brywritescode/serverless-vs-containers-a-managers-guide-3220</link>
      <guid>https://dev.to/brywritescode/serverless-vs-containers-a-managers-guide-3220</guid>
      <description>&lt;h2&gt;
  
  
  Key Points
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Serverless reduces operational burden — your team writes code, not infrastructure. Containers give you more control, but someone has to own that control.&lt;/li&gt;
&lt;li&gt;The cost winner depends on your traffic pattern, not your team's preference. Low or spiky traffic favors serverless. High, steady traffic favors containers.&lt;/li&gt;
&lt;li&gt;If your application needs persistent connections or long-running processes, serverless cannot do the job. Containers win by default.&lt;/li&gt;
&lt;li&gt;Most mature engineering organizations use both. Serverless handles events and background jobs; containers run the core API.&lt;/li&gt;
&lt;li&gt;Start with serverless. Migrate to containers when you hit its limits — not before.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;Every engineering team reaches a moment where the architecture question becomes real: do we go serverless or do we run containers? The answer your engineers give will sound technical. The decision you make should be business-driven. I have sat in that room — first as the engineer making the case, later as the person holding the budget — and the framing that matters is almost never the one the engineers lead with. This is not a question of which technology is more sophisticated. It is a question of which model reduces your operational burden, matches your cost structure, and fits the expertise you have on the team today. The wrong choice does not break your product — but it does burn money and time that you cannot get back.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Core Question
&lt;/h2&gt;

&lt;p&gt;Engineering teams often frame this debate around technical elegance. That is the wrong frame for a leadership conversation.&lt;/p&gt;

&lt;p&gt;The right questions are: How much of your team's time do you want spent managing infrastructure versus building product? What does your traffic look like — steady and predictable, or bursty and unpredictable? And what does your cost structure look like at the scale you expect to reach in 12 to 18 months?&lt;/p&gt;

&lt;p&gt;Those three questions will do more to guide your decision than any benchmark comparison.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Analogy That Makes This Clear
&lt;/h2&gt;

&lt;p&gt;Think of serverless as hiring contract workers. When work arrives, they show up — instantly, in whatever number the job requires. When work stops, they go home and you stop paying. The tradeoff: each new contractor takes a moment to get oriented before they start producing (this is the "cold start" delay engineers worry about). For most workloads, that delay is imperceptible. For some, it matters.&lt;/p&gt;

&lt;p&gt;Containers are more like full-time staff. They are always at their desks, warmed up, ready to respond the moment a request arrives. Performance is consistent and predictable. The tradeoff: you pay their salary whether they are busy or not. If your servers run at 10% capacity most of the day, you are paying full-time wages for part-time work.&lt;/p&gt;

&lt;p&gt;Neither model is universally better. The right one depends on what your workload actually looks like.&lt;/p&gt;




&lt;h2&gt;
  
  
  What This Means for Your Team
&lt;/h2&gt;

&lt;p&gt;The operational difference between these two approaches is significant — and it falls on your engineering team, not on you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;With serverless&lt;/strong&gt;, your team does not manage servers, operating systems, or capacity planning. The cloud provider handles all of that. The result is a smaller operational burden, faster deployment cycles, and less specialized infrastructure expertise required. A team of three engineers can run a meaningful serverless application without a dedicated DevOps hire. The tradeoff is that debugging and monitoring are harder. There is no persistent process to inspect, no server to SSH into. Your engineers need to think differently about troubleshooting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;With containers&lt;/strong&gt;, your team owns the runtime environment — which means they control everything, including the things that go wrong. This is powerful for complex applications with specific runtime requirements. It also means someone on your team needs to understand container orchestration: the tooling that runs, scales, and connects your containers across a fleet of servers. That is a real skill that takes time to develop and retain. In practice, I have seen teams underestimate this consistently — they provision containers, get them running, and then discover six months later that nobody owns the upgrade cycle or the on-call rotation for the orchestration layer. If your team already has the expertise, containers unlock significant capability. If they do not, you are taking on an infrastructure learning curve while trying to ship product.&lt;/p&gt;

&lt;p&gt;A 2024 platform engineering survey found that organizations running container orchestration required an average of 2.5 dedicated platform engineers per 100 application developers. Serverless architectures required 0.5. That difference has a salary cost.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Cost Model Without the Math
&lt;/h2&gt;

&lt;p&gt;Both pricing models are simple in principle. Serverless charges per request and per millisecond of execution time. If no one is using your application, you pay nothing. Containers charge per running hour — the server is on whether requests are flowing or not.&lt;/p&gt;

&lt;p&gt;The cost winner is almost entirely determined by your traffic pattern:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Traffic Level&lt;/th&gt;
&lt;th&gt;Serverless&lt;/th&gt;
&lt;th&gt;Containers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Low / spiky (under 1M requests/month)&lt;/td&gt;
&lt;td&gt;Cheaper — you pay only for actual use&lt;/td&gt;
&lt;td&gt;Wasteful — servers sit idle most of the time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Medium (1M–100M requests/month)&lt;/td&gt;
&lt;td&gt;Competitive&lt;/td&gt;
&lt;td&gt;Competitive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;High / steady (over 100M requests/month)&lt;/td&gt;
&lt;td&gt;Can get expensive&lt;/td&gt;
&lt;td&gt;Cheaper — cost amortized across constant load&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A useful real-world reference point: a nightly batch job that runs two hours per day costs roughly a third less on serverless than on a dedicated container. A high-volume production API serving 500 requests per second costs roughly half as much on containers. The inflection point is utilization — how much of your provisioned capacity you are actually using at any given moment.&lt;/p&gt;

&lt;p&gt;One important nuance: serverless has hidden costs that do not appear in the headline pricing. API gateways, data transfer, and third-party integrations all add to the bill. Factor these in before drawing conclusions from a back-of-napkin estimate. I recommend building a 90-day cost model that includes those line items before you take either option to a budget conversation — the headline number is almost always optimistic.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Serverless Cannot Do
&lt;/h2&gt;

&lt;p&gt;Serverless is not the right tool for every job, and knowing its hard limits will save you from an expensive architectural mistake.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Long-running processes.&lt;/strong&gt; Most serverless platforms impose a maximum execution time of 15 minutes. If your application does anything that runs longer — large data processing jobs, video transcoding, complex ML inference — serverless will cut it off. Containers have no such limit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Persistent connections.&lt;/strong&gt; Serverless functions do not hold open connections between requests. If your application relies on WebSockets (live chat, real-time dashboards, multiplayer features) or long-lived TCP connections (certain financial data feeds, streaming protocols), serverless cannot support them. This is an architectural constraint, not a configuration option. I have seen this surface mid-project when a team built their real-time notification feature on serverless, got it working in testing, and then watched it fail immediately under a live user load that expected sustained connections — a costly rebuild two months before launch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In-memory session state.&lt;/strong&gt; Serverless functions are stateless by design. They do not remember anything between requests. Applications that rely on keeping data in memory across a user's session need a different approach — either an external cache or a stateful server, which points toward containers.&lt;/p&gt;

&lt;p&gt;If your application requires any of these capabilities, containers are not just a preference. They are a requirement.&lt;/p&gt;




&lt;h2&gt;
  
  
  How Most Organizations Actually Do This
&lt;/h2&gt;

&lt;p&gt;The serverless-versus-containers debate often presents as a binary choice. In practice, most mature engineering organizations use both — deliberately.&lt;/p&gt;

&lt;p&gt;The common pattern: serverless handles everything event-driven and asynchronous. File uploads trigger a serverless function that resizes images and updates a database. Webhooks from payment processors invoke serverless logic that updates order status. Scheduled jobs run as serverless functions on a timer. None of this requires a persistent server.&lt;/p&gt;

&lt;p&gt;Meanwhile, the core API — the service that handles user-facing requests in real time — runs in containers. It has predictable traffic, consistent performance requirements, and stateful session logic that makes serverless a poor fit.&lt;/p&gt;

&lt;p&gt;This hybrid approach is not a compromise. It is a rational use of each tool where it performs best. If you are starting fresh, there is no rule that says you must choose one and abandon the other. I default to this split on greenfield projects: serverless for anything async and event-driven, containers for the core API — and revisit the boundaries only when cost data or a hard technical limit gives me a reason to.&lt;/p&gt;




&lt;h2&gt;
  
  
  Decision Flowchart
&lt;/h2&gt;

&lt;p&gt;Use this to orient your conversation with your engineering team. It does not replace a proper architectural review, but it will tell you which direction to lean.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flzjjgfgzemyzwak4lxr3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flzjjgfgzemyzwak4lxr3.png" alt="Decision Flowchart" width="800" height="1255"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Adoption Roadmap
&lt;/h2&gt;

&lt;p&gt;If you are moving toward either architecture deliberately — rather than inheriting whatever a previous team built — a phased approach reduces risk.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 1 — Evaluate (Weeks 1–4).&lt;/strong&gt; Map your existing workloads. Identify which are event-driven and short-lived versus long-running and stateful. Check your current monthly request volume and traffic distribution. Audit your team's infrastructure skills. This phase produces a recommendation, not a commitment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 2 — Pilot (Weeks 5–10).&lt;/strong&gt; Choose one low-risk, non-critical workload and run it on the target architecture. A background job or an internal webhook handler is ideal. Measure actual cost against your estimate. Have your team document what was harder than expected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 3 — Expand (Months 3–6).&lt;/strong&gt; Move additional workloads based on what you learned in the pilot. Establish internal patterns and templates so that each new workload does not require a new architectural conversation. Define your monitoring and alerting standards for the chosen model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 4 — Standardize (Month 6+).&lt;/strong&gt; Codify the decision criteria your team uses. Document which workload types go serverless and which go containers. New engineers joining the team should be able to read one internal document and understand why things are built the way they are. That is the sign that the architecture is settled.&lt;/p&gt;




&lt;h2&gt;
  
  
  Questions to Ask Your Engineering Team
&lt;/h2&gt;

&lt;p&gt;Before your team commits to an architecture, make sure you can get clear answers to these seven questions. If the answers are vague or contradictory, you need more time in Phase 1 before moving forward.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;What does our traffic actually look like?&lt;/strong&gt; Ask for a graph of requests per hour over the last 30 days. The shape of that graph — spiky or flat — is the single most important input to this decision.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Does our application need to hold open connections, and for how long?&lt;/strong&gt; This determines whether serverless is technically viable, full stop. If the answer is yes, the discussion is over — you need containers.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Who on the team owns infrastructure, and what is their current skill set?&lt;/strong&gt; If the answer is "nobody, really," that tells you something important about the operational risk of adopting containers. In my experience, this question surfaces the most organizational debt — teams that say they "use Kubernetes" but have one person who set it up two years ago and nobody who can troubleshoot it at 2 a.m.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;What is our expected request volume in 12 months?&lt;/strong&gt; Cost models diverge significantly at scale. An architecture that looks cheaper today may reverse at 50 million requests per month. Make sure the projection includes growth.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;What is our tolerance for cold start latency?&lt;/strong&gt; Serverless functions can take a fraction of a second to initialize when they have been idle. For most web APIs, this is invisible. For latency-sensitive applications — real-time bidding, financial transactions, gaming — it may be unacceptable.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Do any of our processes run longer than 15 minutes?&lt;/strong&gt; ETL jobs, report generation, large file processing. If the answer is yes for any workload, that workload needs containers. It does not necessarily mean everything needs containers.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;What does our current infrastructure bill look like, and what is it buying us?&lt;/strong&gt; Understanding your baseline makes it possible to evaluate whether a new architecture is actually cheaper or just feels cheaper because it is new.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Serverless and containers are not rivals — they are tools with different strengths, suited to different conditions. The teams I have seen make this decision well are the ones who start from their actual workload and traffic data, not from what is trending at the time. Start with serverless if your team is small, your traffic is variable, and you want to minimize operational overhead. Move to containers — or add them alongside — when the data tells you to, not when someone in a conference room makes an argument that sounds good. The seven questions above are not a formality: if your engineers cannot answer them cleanly, you are not ready to commit to either architecture, and that is the most important thing to know before you approve a roadmap.&lt;/p&gt;




&lt;h2&gt;
  
  
  Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://aws.amazon.com/blogs/aws-cloud-financial-management/a-finops-guide-to-comparing-containers-and-serverless-functions-for-compute/" rel="noopener noreferrer"&gt;AWS FinOps Guide: Comparing Containers and Serverless for Compute&lt;/a&gt; — AWS Cloud Financial Management Blog&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.cloudflare.com/learning/serverless/serverless-vs-containers/" rel="noopener noreferrer"&gt;Serverless vs. Containers — Cloudflare Learning Center&lt;/a&gt; — Cloudflare&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.readysetcloud.io/blog/allen.helton/when-is-serverless-more-expensive/" rel="noopener noreferrer"&gt;When Is Serverless More Expensive Than Containers?&lt;/a&gt; — Ready, Set, Cloud&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.datadoghq.com/knowledge-center/serverless-architecture/serverless-vs-containers/" rel="noopener noreferrer"&gt;Serverless vs. Containers: Which is Right for You?&lt;/a&gt; — Datadog Knowledge Center&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.digitalocean.com/resources/articles/serverless-vs-containers" rel="noopener noreferrer"&gt;Serverless vs Containers: Which is best for your needs?&lt;/a&gt; — DigitalOcean&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;If this helped, a like and a follow are appreciated — and if you've solved this differently, drop a comment, I'd like to hear it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Bry Writes Code — cloud and AI infrastructure specialist. Evaluating your cloud architecture? &lt;a href="mailto:brywritescode@gmail.com"&gt;Let's talk&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>aws</category>
      <category>devops</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Junior Engineer Pipeline Collapse: Who Trains the Next Generation When AI Does the Grunt Work</title>
      <dc:creator>Bry</dc:creator>
      <pubDate>Wed, 30 Sep 2026 15:00:00 +0000</pubDate>
      <link>https://dev.to/brywritescode/junior-engineer-pipeline-collapse-who-trains-the-next-generation-when-ai-does-the-grunt-work-1fik</link>
      <guid>https://dev.to/brywritescode/junior-engineer-pipeline-collapse-who-trains-the-next-generation-when-ai-does-the-grunt-work-1fik</guid>
      <description>&lt;h2&gt;
  
  
  Key Points
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Entry-level hiring has fallen sharply and fast: roughly 65% down at major tech companies and 76% at early-stage startups, with new software engineering postings dropping 15% in just the first two months of this year.&lt;/li&gt;
&lt;li&gt;The logic driving it is straightforwardly rational at the individual-firm level. One senior engineer with an AI coding assistant now ships what used to take a senior plus a junior, and when one person does the work of one and a half, nobody needs the extra half.&lt;/li&gt;
&lt;li&gt;Junior work has always been how engineers learn the specific, nameable judgment that became the valuable thing. Cutting junior hiring isn't just saving money short-term. It's quietly cutting the only pipeline that reliably produces the senior judgment firms will need five to ten years from now.&lt;/li&gt;
&lt;li&gt;We now open with the pipeline problem specifically because every other Delivery Practice including RFPs, fixed-price risk, client trust, insourcing, maintenance economics, assumes a supply of experienced engineers that this collapse is putting in real doubt for firms that don't deliberately intervene.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;A senior engineer said something to a hiring committee I sat on this year that was blunt enough to stick with me: why hire a junior for $90,000 a year when the AI coding assistant that makes any one of us more productive costs $10 a month? Nobody on the committee had a good rebuttal in the moment, because at the level of that single hiring decision, he wasn't wrong. The math really did favor not hiring.&lt;/p&gt;

&lt;p&gt;The data backs up how widespread that individual logic has become. Entry-level hiring has fallen roughly 65% at major tech companies and 76% at early-stage startups this year, and new software engineering postings dropped 15% in just the first two months, hitting entry-level roles hardest of all. Unemployment for recent computer science and computer engineering graduates has climbed well above the general population rate, 6-7% against roughly 4.3% overall, a specific, measurable cost landing on a specific cohort. The mechanism is simple and, again, locally rational: companies report 40-55% more code output per sprint after adopting AI coding tools, and one senior engineer with those tools now ships what used to require a senior plus a junior. When one person does the work of one and a half, the extra half doesn't get hired.&lt;/p&gt;

&lt;p&gt;What that hiring committee, and a lot of firms like it, aren't fully reckoning with yet is that junior roles were never just cheap labor. They were the mechanism by which the industry manufactured senior judgment, the specific, nameable pattern-recognition, the kind that only develops by spending years making, and fixing, the mistakes junior engineers make on real systems. Cutting that pipeline doesn't just save payroll in the short term. It's quietly cutting the supply of the exact thing that is becoming the valuable, billable thing: engineers who've built up enough hard-won pattern recognition to catch a subtly wrong AI-proposed solution before it reaches production. Some experts are flagging this directly already, warning that fewer new graduates entering the field could create real senior-developer shortages five to ten years out. That's the kind of warning that's easy to discount right now and expensive to have ignored later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Firms Cutting Junior Hiring vs. Firms Restructuring It
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;What It Looks Like Right Now&lt;/th&gt;
&lt;th&gt;Where This Is Likely to Lead&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pure hiring freeze&lt;/td&gt;
&lt;td&gt;Stopped junior hiring entirely, redirected the savings to senior headcount and AI tooling licenses&lt;/td&gt;
&lt;td&gt;A real, measurable senior-pipeline gap five to seven years out, competing hard for a shrinking pool of experienced engineers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"AI does the grunt work, juniors do nothing new"&lt;/td&gt;
&lt;td&gt;Kept some junior hiring, but gives juniors only the residual work AI tooling doesn't already handle&lt;/td&gt;
&lt;td&gt;Juniors underdeveloped, exposed to too little real judgment-building work to mature into the seniors the firm will need&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Restructured junior roles around AI-output review and constrained real ownership&lt;/td&gt;
&lt;td&gt;Redesigned entry-level work explicitly around reviewing AI output, owning small but real production components, and structured mentorship time&lt;/td&gt;
&lt;td&gt;A genuine, if smaller, pipeline of engineers developing judgment faster than the old apprenticeship model, because they're reviewing more decisions per year, not writing more boilerplate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New entry-level categories: AI-output QA, data labeling and curation, model-behavior validation&lt;/td&gt;
&lt;td&gt;Creating explicitly new junior roles rather than shrinking old ones&lt;/td&gt;
&lt;td&gt;Some firms positioned to develop judgment through a different but real apprenticeship path, though how well this generalizes to full engineering judgment remains genuinely contested&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Recommendation:&lt;/strong&gt; if your firm's response to AI-driven productivity gains was simply "hire fewer juniors," check what your senior bench is likely to look like five years from now, not just this year's payroll. Firms restructuring junior work around real ownership and AI-output review, instead of eliminating it, are the ones building an actual pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rebuilding a Junior Pipeline That Actually Produces Judgment
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Stop giving juniors only residual work.&lt;/strong&gt; If AI tooling handles the boilerplate and juniors get whatever's left over, they're not building judgment. They're doing chores. Deliberately assign juniors ownership of small, real, production-consequential components, with AI as a tool they direct, not a replacement for their decisions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build structured review of AI output into junior roles explicitly, not as an afterthought.&lt;/strong&gt; Reviewing and validating AI-generated code against real production constraints is turning out to accelerate judgment development faster than writing routine implementation ever did. Treat it as core training, not busywork.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Create a real mentorship cadence, and protect senior engineers' time to do it.&lt;/strong&gt; A senior doing the work of one and a half people has no slack left for mentorship unless that time is explicitly protected and counted as part of their role, not an unpaid extra.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure junior progression by judgment indicators, not implementation speed.&lt;/strong&gt; Catching a subtly wrong AI suggestion, correctly pushing back on a requirement, recognizing a system's known failure mode: track these, not lines shipped, as the signal that the pipeline is actually working.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accept that this pipeline will be smaller and more deliberate than the old one, and budget for it as a real cost center.&lt;/strong&gt; The old apprenticeship model scaled with headcount almost automatically. The new one requires an explicit, funded decision to keep training people, because the market no longer forces it on you by default.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Questions to Ask Your Team
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;If we've cut junior hiring this year, do we know what our senior engineering bench is likely to look like five years from now, or are we assuming the market will supply seniors when we need them?&lt;/li&gt;
&lt;li&gt;Are our current juniors getting real ownership and judgment-building work, or just the residual tasks AI tooling doesn't already handle?&lt;/li&gt;
&lt;li&gt;Is mentorship time for senior engineers protected and counted as part of their role, or treated as something they're supposed to absorb on top of an already AI-accelerated workload?&lt;/li&gt;
&lt;li&gt;Would we recognize a senior-engineer shortage forming in our own pipeline before it became an active hiring crisis, or only after?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The individual hiring-committee math this year isn't wrong. One senior engineer with AI tooling genuinely does the work of one and a half, and paying for the extra half stops making sense at the level of a single decision. What that math doesn't price in is that junior roles were never just labor capacity. They're the industry's mechanism for manufacturing the exact judgment this whole thread has been arguing is becoming the valuable, billable thing. Firms that simply stop hiring juniors get the short-term savings and, on a five-to-ten-year timeline, are likely to face a real gap in the judgment supply they'll need to compete on the deliverable. Firms restructuring junior work, around AI-output review, real ownership, and protected mentorship, are building something smaller than the old pyramid, but real.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.cio.com/article/4062024/demand-for-junior-developers-softens-as-ai-takes-over.html" rel="noopener noreferrer"&gt;CIO: Demand for Junior Developers Softens as AI Takes Over&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ardura.consulting/blog/junior-developer-crisis-2026-why-companies-stopped-hiring-entry-level/" rel="noopener noreferrer"&gt;ARDURA Consulting: Junior Developer Crisis 2026, Why Hiring Dropped 50%&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://codeconductor.ai/blog/future-of-junior-developers-ai/" rel="noopener noreferrer"&gt;CodeConductor: The Future of Junior Developers in the Age of AI&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;If this helped, a like and a follow are appreciated — and if you've solved this differently, drop a comment, I'd like to hear it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Bry Writes Code; cloud and AI infrastructure specialist. Worried about what your senior bench looks like in five years? &lt;a href="mailto:brywritescode@gmail.com"&gt;Let's talk&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>business</category>
      <category>ai</category>
      <category>hiring</category>
      <category>beginners</category>
    </item>
    <item>
      <title>An Engineer's Reflection: Watching the SI Business Model Change From the Inside</title>
      <dc:creator>Bry</dc:creator>
      <pubDate>Wed, 23 Sep 2026 15:00:00 +0000</pubDate>
      <link>https://dev.to/brywritescode/an-engineers-reflection-watching-the-si-business-model-change-from-the-inside-4adn</link>
      <guid>https://dev.to/brywritescode/an-engineers-reflection-watching-the-si-business-model-change-from-the-inside-4adn</guid>
      <description>&lt;h2&gt;
  
  
  Key Points
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The identity discomfort is real and widely reported this year. Commentators describe it as bordering on depression for engineers who feel their role slipping from skilled creator to what one report calls "cleanup crew." I don't think that framing is wrong. I think it's incomplete.&lt;/li&gt;
&lt;li&gt;What I'm watching happen to the engineers coming through this well isn't that judgment is replacing coding as some abstract virtue. It's that the specific, nameable things they know, a legacy system's quirks, a client's real constraints, when an AI-proposed solution is subtly wrong, are becoming the paid part of the job, while the part that used to define professional identity, writing the implementation, is turning into infrastructure.&lt;/li&gt;
&lt;li&gt;Fujitsu's own internal reforms are a useful, concrete marker of how far this is going beyond rhetoric: the company has moved from uniform starting salaries for new graduates to job-competency pay from day one, restructuring roughly 650 roles around it, an explicit break from seniority-based structure that would have been unthinkable at a firm like Fujitsu a few years ago.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;I want to talk about a specific conversation, not a general trend, because the trend only ever shows up to me one uncomfortable conversation at a time. This one was with a colleague, a genuinely excellent engineer who's spent fifteen years getting better at exactly the kind of implementation work AI tooling is rapidly making cheap. He asked me, not rhetorically, whether he'd wasted those fifteen years. I don't have a fully clean answer for him yet.&lt;/p&gt;

&lt;p&gt;The discomfort he feels is widely shared and worth taking seriously on its own terms, not just as a business-model footnote. Commentary this year describes something close to an identity crisis for engineers whose professional self-image has been built on skilled implementation. One widely circulated framing put it as "bordering on depression," with veteran engineers feeling their role slip from skilled creator to what amounts to a cleanup crew reviewing AI output. That's not a business-press exaggeration. I'm watching it happen to people I respect, including, some days, myself.&lt;/p&gt;

&lt;p&gt;Here's where I've landed after watching this from inside a role that straddles engineering and business: the resolution isn't that "judgment" swoops in as an abstract replacement virtue for "coding." It's narrower and more specific than that. The things turning out to be worth paying for are nameable: knowing that a particular legacy system behaves unpredictably under a specific load pattern, knowing which of a client's stated requirements are actually politically immovable versus negotiable, catching the specific moment an AI-proposed integration is subtly, dangerously wrong in a way that would only surface months into production. None of that is "judgment" as a vague virtue. It's specific, earned pattern-recognition that's becoming the billable part of the job, because the part that used to be billable, writing the implementation, got cheap. Fujitsu's own workforce reforms are the clearest evidence I've seen that this isn't just rhetoric: the company has broken from uniform, seniority-linked starting salaries toward job-competency pay from day one, restructuring around 650 roles in the process. That's an institution built on decades of seniority-based structure admitting, structurally and publicly, that what it pays for has changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  An Engineer's Position: What Made You Valuable, Then and Now
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;The Old Basis of Value&lt;/th&gt;
&lt;th&gt;Where This Is Heading&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What made an engineer valuable&lt;/td&gt;
&lt;td&gt;Consistent, correct implementation output&lt;/td&gt;
&lt;td&gt;Specific domain and system knowledge, judgment about AI-proposed solutions, willingness to own an outcome&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Career ladder shape&lt;/td&gt;
&lt;td&gt;Seniority and tenure-linked, fairly uniform within a cohort&lt;/td&gt;
&lt;td&gt;Competency-linked from early career, per Fujitsu's own 2026 reform pattern&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Most exposed cohort right now&lt;/td&gt;
&lt;td&gt;Mid-career implementation specialists with deep tooling expertise, less domain-specific judgment&lt;/td&gt;
&lt;td&gt;Likely to settle, but only for engineers who deliberately build domain-specific judgment during the transition&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What "keeping up" means&lt;/td&gt;
&lt;td&gt;Learning the newest AI coding tool&lt;/td&gt;
&lt;td&gt;Deepening specific, nameable expertise a tool can't substitute for&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Honest emotional reality right now&lt;/td&gt;
&lt;td&gt;Real anxiety, reasonably described as an identity crisis for a meaningful share of the workforce&lt;/td&gt;
&lt;td&gt;Likely to settle over time, but the transition is costing real careers and real confidence along the way. Not a costless story&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Recommendation:&lt;/strong&gt; if you're an engineer reading this mid-transition, the actionable version of this table isn't "learn to use AI tools better." It's "get specific about what you know that a tool doesn't": a client's real constraints, a system's real failure modes, a domain's real edge cases. That's the column most likely to survive.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm Telling That Colleague
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The fifteen years aren't wasted, but the specific skill that felt most valuable, fast, correct implementation, really is stopping being the paid part of the job.&lt;/strong&gt; That's a real loss worth naming, not minimizing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What those fifteen years actually built, underneath the implementation skill, is pattern-recognition about when something is subtly wrong.&lt;/strong&gt; That's turning out to be the transferable part, and it wasn't obvious that it would be.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The engineers adapting fastest aren't out-learning AI tooling. They're out-specifying their own domain knowledge.&lt;/strong&gt; They're turning "I've seen this before" into something they can name, price, and sell, instead of leaving it as unpriced background competence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The institutional signal to watch for isn't a memo about "AI transformation." It's compensation structure actually changing.&lt;/strong&gt; Fujitsu breaking seniority-linked pay isn't a communications exercise. It's the company admitting what it now pays for, and that kind of structural signal tells you more than any strategy announcement.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Questions to Ask Yourself, Not Just Your Team
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;What do I know, specifically and by name, that isn't just "I'm experienced": a system's real failure pattern, a client's real constraints, a domain's real edge cases?&lt;/li&gt;
&lt;li&gt;Am I still measuring my own value by implementation speed, in a market that's increasingly not paying for that specifically?&lt;/li&gt;
&lt;li&gt;If my compensation structure hasn't changed to reflect a competency-based model, is that because my organization hasn't caught up yet, or because it doesn't need to for the kind of work I actually do?&lt;/li&gt;
&lt;li&gt;Would I recognize the moment an AI-proposed solution was subtly wrong in my own domain, and could I explain why, specifically, to someone who trusted the AI's output at face value?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;This thread started with a pricing unit dying and ends, for now, with what that dying pricing unit is doing to the people underneath it. I don't think the honest reflection is either the doom-laden "identity collapse" framing circulating this year or the tidy "engineers move up the value chain" framing that shows up in a lot of retrospective business writing, including some of my own earlier pieces in this thread. It's both, at once: a real, disorienting loss for people whose professional identity was built on a skill that's stopping being scarce, and a real, achievable adaptation for the people who find the specific, nameable thing they know that a tool doesn't. My colleague hasn't found his yet, as far as I know. I think he will. I don't think it'll be comfortable getting there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://beyondplm.com/2026/05/06/engineering-identity-in-the-age-of-ai-from-knowledge-holder-to-judgment-owner/" rel="noopener noreferrer"&gt;Beyond PLM: Engineering Identity in the Age of AI, From Knowledge Holder to Judgment Owner&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.blog/news-insights/octoverse/the-new-identity-of-a-developer-what-changes-and-what-doesnt-in-the-ai-era/" rel="noopener noreferrer"&gt;The GitHub Blog: The New Identity of a Developer, What Changes and What Doesn't in the AI Era&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://completeaitraining.com/news/fujitsu-goes-ai-first-with-nvidia-breaks-seniority-and/" rel="noopener noreferrer"&gt;CompleteAITraining: Fujitsu Goes AI-First With NVIDIA, Breaks Seniority, and Tests Its Reach Beyond Japan&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Bry Writes Code; cloud and AI infrastructure specialist. Figuring out what you know that a tool doesn't, professionally or for your team? &lt;a href="mailto:brywritescode@gmail.com"&gt;Let's talk&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>business</category>
      <category>ai</category>
      <category>career</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Azure Service Bus: Queues and Topics for When Event Grid Isn't Enough</title>
      <dc:creator>Bry</dc:creator>
      <pubDate>Tue, 22 Sep 2026 15:00:00 +0000</pubDate>
      <link>https://dev.to/brywritescode/azure-service-bus-queues-and-topics-for-when-event-grid-isnt-enough-2jh8</link>
      <guid>https://dev.to/brywritescode/azure-service-bus-queues-and-topics-for-when-event-grid-isnt-enough-2jh8</guid>
      <description>&lt;h2&gt;
  
  
  Key Points
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Basic tier supports queues only — no topics, no subscriptions, no sessions, at any price. It's a feature gate, not a cheaper version of Standard, and Standard is the practical floor the moment any of those three enter the picture.&lt;/li&gt;
&lt;li&gt;Sessions are Service Bus's actual differentiator over Event Grid: they group related messages and guarantee in-order, single-consumer processing within a session — Event Grid has no equivalent.&lt;/li&gt;
&lt;li&gt;Premium tier drops per-message transaction billing entirely in favor of a flat daily rate per messaging unit — the economics flip once you're at consistent high volume.&lt;/li&gt;
&lt;li&gt;A subscription with no filter rule receives every message on the topic via the implicit &lt;code&gt;1=1&lt;/code&gt; default rule — easy to forget you're relying on that default.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;CLI/SDK version tested against: &lt;code&gt;az-cli 2.6x.x&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;An Azure resource group (&lt;code&gt;az group create&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Standard tier or higher for the topic/subscription/session examples below — Basic tier will reject topic creation outright&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;Article 352 covered Event Grid and drew a clear line: it's a push-based router for discrete events, not a queue and not a stream. This article covers the service that actually is a queue — with topics, ordering guarantees, and transactional semantics Event Grid was never built to provide.&lt;/p&gt;

&lt;p&gt;The single most common surprise I've watched people hit with Service Bus isn't conceptual, it's the tier structure. Someone provisions a Basic-tier namespace to save money, tries to create a topic, and gets a flat rejection — topics simply don't exist at that tier. Not throttled, not limited. Absent. That's worth knowing before you provision anything, not after.&lt;/p&gt;

&lt;p&gt;This article builds a session-enabled queue (the feature that actually justifies choosing Service Bus over Event Grid for ordered, related work) and a filtered topic subscription, entirely from the CLI.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Tier Gate
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fronvshumu11m4sr4ubz8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fronvshumu11m4sr4ubz8.png" alt="The Tier Gate" width="784" height="247"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Diagram: Basic tier is a genuinely different feature set, not a scaled-down Standard — this is the detail that catches people off guard.&lt;/em&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;Queues&lt;/th&gt;
&lt;th&gt;Topics/Subscriptions&lt;/th&gt;
&lt;th&gt;Sessions&lt;/th&gt;
&lt;th&gt;Billing model&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Basic&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Per-operation, lowest rate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Per-operation, moderate rate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Premium&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Flat daily rate per messaging unit (1, 2, or 4 units); no per-message charge&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If your design needs a topic or a session — and sessions are the reason to pick Service Bus at all in most cases — Basic tier is off the table before you write a single line of application code.&lt;/p&gt;




&lt;h2&gt;
  
  
  Building a Session-Enabled Queue
&lt;/h2&gt;

&lt;p&gt;Sessions group related messages so they're delivered in order to exactly one consumer at a time, identified by a session ID you assign when sending. This is the feature Event Grid genuinely cannot do — it has no concept of message grouping or delivery ordering across a related set.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az servicebus namespace create &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--resource-group&lt;/span&gt; rg-servicebus-demo &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; sb-demo-namespace &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--location&lt;/span&gt; eastus &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--sku&lt;/span&gt; Standard

az servicebus queue create &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--resource-group&lt;/span&gt; rg-servicebus-demo &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace-name&lt;/span&gt; sb-demo-namespace &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; orders-sessions-queue &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--enable-session&lt;/span&gt; &lt;span class="nb"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4upgzij8xzhnmngx94l3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4upgzij8xzhnmngx94l3.png" alt="Building a Session-Enabled Queue&lt;br&gt;
" width="784" height="291"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Diagram: session locking guarantees in-order delivery within a session without blocking unrelated sessions from being processed concurrently by other consumers.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Sending a message into a session queue requires setting the session ID explicitly — omit it and the send fails, because &lt;code&gt;--enable-session true&lt;/code&gt; makes session ID mandatory for every message on that queue, not optional.&lt;/p&gt;




&lt;h2&gt;
  
  
  Building a Filtered Topic Subscription
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az servicebus topic create &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--resource-group&lt;/span&gt; rg-servicebus-demo &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace-name&lt;/span&gt; sb-demo-namespace &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; orders-topic

az servicebus topic subscription create &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--resource-group&lt;/span&gt; rg-servicebus-demo &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace-name&lt;/span&gt; sb-demo-namespace &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--topic-name&lt;/span&gt; orders-topic &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; high-value-orders

&lt;span class="c"&gt;# Without this rule, the subscription's default "1=1" filter&lt;/span&gt;
&lt;span class="c"&gt;# receives every message on the topic — not just high-value orders.&lt;/span&gt;
az servicebus topic subscription rule create &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--resource-group&lt;/span&gt; rg-servicebus-demo &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace-name&lt;/span&gt; sb-demo-namespace &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--topic-name&lt;/span&gt; orders-topic &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--subscription-name&lt;/span&gt; high-value-orders &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; HighValueFilter &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--filter-sql-expression&lt;/span&gt; &lt;span class="s2"&gt;"Total &amp;gt; 500"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Any subscription created without an explicit rule keeps the implicit default rule, which matches everything. That's fine when you genuinely want every message — the risk is forgetting the default is there at all and being surprised when a subscription you intended to be narrow is receiving the full topic firehose.&lt;/p&gt;




&lt;h2&gt;
  
  
  When to Reach for Service Bus Instead of Event Grid
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Service Bus&lt;/th&gt;
&lt;th&gt;Event Grid&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Delivery model&lt;/td&gt;
&lt;td&gt;Pull (consumer receives/completes)&lt;/td&gt;
&lt;td&gt;Push (webhook delivery)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ordering&lt;/td&gt;
&lt;td&gt;Yes, via sessions&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Message grouping&lt;/td&gt;
&lt;td&gt;Yes, via sessions&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Duplicate detection&lt;/td&gt;
&lt;td&gt;Yes (Standard+)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transactions&lt;/td&gt;
&lt;td&gt;Yes (Standard+)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;Ordered work queues, transactional messaging&lt;/td&gt;
&lt;td&gt;React-to-discrete-event fan-out&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Recommendation:&lt;/strong&gt; If a stakeholder describes a requirement involving "process these in order" or "these messages belong together and must not interleave," that's sessions, which means Service Bus. If the requirement is "notify several things when X happens" with no ordering constraint, Event Grid remains the lighter, cheaper choice from article 352.&lt;/p&gt;




&lt;h2&gt;
  
  
  Common Mistakes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Mistake 1: Provisioning Basic tier, then needing a topic&lt;/strong&gt;&lt;br&gt;
There's no in-place upgrade path that preserves existing resources across a tier change in every case — plan tier selection before provisioning, not as a fix-it-later decision.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mistake 2: Enabling sessions without assigning session IDs on every send&lt;/strong&gt;&lt;br&gt;
&lt;code&gt;--enable-session true&lt;/code&gt; makes the session ID mandatory. A send call missing it fails outright — this isn't a soft warning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mistake 3: Assuming a subscription with no rule is empty until configured&lt;/strong&gt;&lt;br&gt;
It's not empty — it has the default &lt;code&gt;1=1&lt;/code&gt; rule and receives everything. If you wanted a narrow subscription, add the filter before anything starts publishing, not after noticing unexpected volume.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mistake 4: Reaching for Service Bus when Event Grid would do&lt;/strong&gt;&lt;br&gt;
If there's no ordering or grouping requirement, Service Bus's added complexity and Standard-tier cost floor isn't buying you anything Event Grid didn't already provide cheaper.&lt;/p&gt;




&lt;h2&gt;
  
  
  Production Considerations
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Performance:&lt;/strong&gt; Standard tier is shared-capacity and can experience throttling under sustained high load. Premium's dedicated messaging units remove that variability — worth the flat rate once volume is consistent and predictable rather than spiky.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Security:&lt;/strong&gt; Use Azure AD (Entra ID) authentication with managed identities rather than shared access signature connection strings where possible — connection strings embed credentials that are easy to leak into logs or source control.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost:&lt;/strong&gt; Premium's flat daily rate per messaging unit becomes cheaper than Standard's per-operation billing once your message volume is consistently high — model both before committing, since Premium at low volume is a worse deal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Monitoring:&lt;/strong&gt; Track &lt;code&gt;ActiveMessageCount&lt;/code&gt; and &lt;code&gt;DeadLetterMessageCount&lt;/code&gt; per queue and subscription. A rising dead-letter count with no alert is the Service Bus equivalent of the EventBridge and Event Grid dead-letter blind spots covered in this series' earlier articles.&lt;/p&gt;




&lt;h2&gt;
  
  
  Full Example: Send and Receive with Sessions (TypeScript)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;ServiceBusClient&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@azure/service-bus&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;connectionString&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;SERVICEBUS_CONNECTION_STRING&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ServiceBusClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;connectionString&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="cm"&gt;/** Sends two related messages into the same session, in order. */&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;sendSessionMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sessionId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="k"&gt;void&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sender&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;createSender&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;orders-sessions-queue&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;sender&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendMessages&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;orderId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;sessionId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;step&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;created&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="nx"&gt;sessionId&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;orderId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;sessionId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;step&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;payment-confirmed&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="nx"&gt;sessionId&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;]);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;sender&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="cm"&gt;/** Receives and processes messages for a specific session in order. */&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;receiveSessionMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sessionId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="k"&gt;void&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;receiver&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;acceptSession&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;orders-sessions-queue&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;sessionId&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;receiver&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;receiveMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;maxWaitTimeInMs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5000&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Session &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;sessionId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;receiver&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;completeMessage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;receiver&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sessionId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;order-abc123&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sendSessionMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sessionId&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;receiveSessionMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sessionId&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;Full source with the topic/subscription example: &lt;a href="https://github.com/brywritescode/bry-writes-code-examples.git" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; → &lt;code&gt;cloud-apis/azure-service-bus-cli/&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Reach for Service Bus over Event Grid specifically when ordering or message grouping matters — sessions are the concrete feature that justifies it, not a vague "more enterprise" gesture. Standard tier is the practical floor the moment topics or sessions enter the picture; Basic tier's queue-only feature gate isn't a limitation you can budget around, it's absent entirely. Know which one you actually need before you provision the namespace — moving between tiers after the fact costs more time than picking correctly up front.&lt;/p&gt;




&lt;h2&gt;
  
  
  Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://azure.microsoft.com/en-us/pricing/details/service-bus/" rel="noopener noreferrer"&gt;Pricing - Service Bus&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://learn.microsoft.com/en-us/azure/service-bus-messaging/service-bus-premium-messaging" rel="noopener noreferrer"&gt;Azure Service Bus Premium Messaging Features&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://learn.microsoft.com/en-us/azure/service-bus-messaging/service-bus-tutorial-topics-subscriptions-cli" rel="noopener noreferrer"&gt;Use the Azure CLI to create Service Bus topics and subscriptions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://learn.microsoft.com/en-us/azure/service-bus-messaging/service-bus-quickstart-cli" rel="noopener noreferrer"&gt;Quickstart - Use the Azure CLI to create a Service Bus queue&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;If this helped, a like and a follow are appreciated — and if you've solved this differently, drop a comment, I'd like to hear it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Bry Writes Code — cloud and API infrastructure specialist. Choosing between Event Grid and Service Bus for your next Azure project? &lt;a href="mailto:brywritescode@gmail.com"&gt;Get in touch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>azure</category>
      <category>typescript</category>
      <category>backend</category>
    </item>
    <item>
      <title>Data Residency and APPI Compliance: What Cloud Vendors Entering Japan Need to Know</title>
      <dc:creator>Bry</dc:creator>
      <pubDate>Thu, 17 Sep 2026 15:00:00 +0000</pubDate>
      <link>https://dev.to/brywritescode/data-residency-and-appi-compliance-what-cloud-vendors-entering-japan-need-to-know-3jmj</link>
      <guid>https://dev.to/brywritescode/data-residency-and-appi-compliance-what-cloud-vendors-entering-japan-need-to-know-3jmj</guid>
      <description>&lt;h2&gt;
  
  
  Key Points
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;APPI applies extraterritorially — if you handle personal data of individuals in Japan, you're in scope regardless of where your company is incorporated, and the PPC has authority to order compliance from overseas.&lt;/li&gt;
&lt;li&gt;Cross-border transfer requires either the data subject's opt-in consent (with specific disclosures) or a documented, annually-monitored equivalent-protection system at the receiving party — "we have a privacy policy" is not either of those.&lt;/li&gt;
&lt;li&gt;Breach notification runs on a tight two-report timeline: an initial report within 3–5 days of recognizing a qualifying breach, a final report within 30 days (60 for cyberattack-related incidents).&lt;/li&gt;
&lt;li&gt;Japan has no blanket data-localization statute, but specific sectors (finance, critical infrastructure, government-adjacent) function as if it did — treat "no hard law" as "no automatic pass," not as "no requirement."&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;I've reviewed vendor security questionnaires where the compliance answer for Japan reads "N/A — Japan does not require data localization." That's technically true and functionally the wrong answer, because it treats an absence of a blanket statute as an absence of a real obligation. Article 431 in this series flagged that data residency has become a procurement gate; this article is the regulatory substance behind that gate — what APPI actually requires, and where "no hard localization law" quietly turns into a hard requirement anyway once you're selling into specific sectors.&lt;/p&gt;

&lt;p&gt;Nothing here substitutes for a licensed opinion from Japan-qualified counsel — treat this as the shape of the requirements so your legal and engineering teams know what questions to bring to that counsel, not as the final word on your specific situation.&lt;/p&gt;




&lt;h2&gt;
  
  
  APPI's Extraterritorial Reach
&lt;/h2&gt;

&lt;p&gt;The Act on the Protection of Personal Information, enforced by the Personal Information Protection Commission, applies to any business handling the personal information of individuals in Japan — regardless of where that business is incorporated. Article 171 of the amended APPI explicitly grants the PPC authority over foreign operators, including the power to compel reports and issue orders to a company with no physical presence in Japan at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The problem this solves for the regulator, and the problem it creates for you:&lt;/strong&gt; a foreign SaaS vendor with zero Japan entity, serving Japanese users entirely from a US data center, is still squarely in scope. "We're not a Japanese company" is not a compliance strategy — it's the specific gap Article 171 was written to close.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnx9wjdxeeoxpu4oo28bo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnx9wjdxeeoxpu4oo28bo.png" alt="APPI's Extraterritorial Reach" width="713" height="270"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Cross-Border Transfer: The Two Legitimate Paths
&lt;/h2&gt;

&lt;p&gt;This is the mechanic most foreign vendors get wrong, because the intuitive assumption — "we have a privacy policy, we're covered" — isn't one of the two paths APPI actually recognizes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Path 1: Opt-in consent.&lt;/strong&gt; You obtain the data subject's affirmative consent before transferring their personal information outside Japan, and at the point of obtaining that consent you disclose three specific things: the name of the destination country, that country's data protection legal framework, and the specific data protection measures the receiving party has in place. A generic "by using this service you consent to our terms" does not meet this bar — the disclosure requirements are specific and checkable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Path 2: Equivalent-protection system.&lt;/strong&gt; You establish a documented personal-information-protection system at the receiving party that meets APPI's substantive standard, and — this is the part vendors most often miss — you're required to &lt;strong&gt;monitor that system's continued adequacy at least once a year&lt;/strong&gt;, not just certify it once at setup and move on.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff21pc0zekt0asje5025v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff21pc0zekt0asje5025v.png" alt="Cross-Border Transfer: The Two Legitimate Paths" width="586" height="846"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Most SaaS vendors operating multi-region infrastructure default to Path 2 because Path 1's per-transfer consent flow doesn't fit a normal signup funnel. If you choose Path 2, budget for the annual monitoring obligation as an ongoing compliance line item, not a one-time setup task — this is where "we did this once during onboarding" quietly becomes non-compliant eighteen months later.&lt;/p&gt;




&lt;h2&gt;
  
  
  Breach Notification: A Tighter Timeline Than It Looks
&lt;/h2&gt;

&lt;p&gt;APPI's breach notification runs on two reports with two different clocks, and both are shorter than most Western equivalents foreign compliance teams are used to planning around.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Report&lt;/th&gt;
&lt;th&gt;Deadline&lt;/th&gt;
&lt;th&gt;Trigger&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Initial report to the PPC&lt;/td&gt;
&lt;td&gt;3–5 calendar days from recognizing the breach&lt;/td&gt;
&lt;td&gt;Breach involves sensitive personal information, risk of property damage, improper/unauthorized use, or more than 1,000 affected data subjects&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Final report to the PPC&lt;/td&gt;
&lt;td&gt;30 calendar days (60 days if the breach involved an "unjust purpose," e.g. a cyberattack)&lt;/td&gt;
&lt;td&gt;Same qualifying triggers as the initial report&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The reporting &lt;em&gt;obligation&lt;/em&gt; itself isn't universal — it's triggered by the specific conditions in the table above, not every incident. But the practical guidance is to build your incident response runbook assuming you'll need to move on the 3–5 day clock, because determining whether an incident qualifies is itself something you often can't finish in 3–5 days if you're starting that assessment from zero at the moment of the breach. The teams that hit this deadline comfortably are the ones who decided their qualifying criteria and report template in advance, not during the incident.&lt;/p&gt;




&lt;h2&gt;
  
  
  Comparison: No Data Residency Commitment vs Multi-Region Flexibility vs Japan-Only Hosting
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criteria&lt;/th&gt;
&lt;th&gt;No Residency Commitment&lt;/th&gt;
&lt;th&gt;Multi-Region Flexibility&lt;/th&gt;
&lt;th&gt;Japan-Only Hosting&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;APPI cross-border transfer obligation&lt;/td&gt;
&lt;td&gt;Triggered on every transfer, by default&lt;/td&gt;
&lt;td&gt;Triggered, but manageable via Path 2 + annual monitoring&lt;/td&gt;
&lt;td&gt;Largely avoided — data doesn't leave Japan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise procurement pass rate&lt;/td&gt;
&lt;td&gt;Low — increasingly a hard "no" at the security-questionnaire stage&lt;/td&gt;
&lt;td&gt;Moderate — depends on how well Path 2 is documented and monitored&lt;/td&gt;
&lt;td&gt;High — clears the tier where residency is treated as a gate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Regulated-sector eligibility (finance, gov-adjacent)&lt;/td&gt;
&lt;td&gt;Effectively disqualifying&lt;/td&gt;
&lt;td&gt;Sector-dependent, often insufficient&lt;/td&gt;
&lt;td&gt;Required baseline for most regulated-sector deals&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Engineering cost&lt;/td&gt;
&lt;td&gt;None (but highest compliance/deal-risk exposure)&lt;/td&gt;
&lt;td&gt;Moderate — regional data-partitioning work&lt;/td&gt;
&lt;td&gt;Highest — dedicated Japan-region infrastructure&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; most horizontal SaaS products should target multi-region flexibility with a properly documented and annually-monitored Path 2 system. Reserve full Japan-only hosting for when a specific regulated-sector deal or government-adjacent procurement actually requires it — building it speculatively is expensive infrastructure ahead of a deal that may not materialize.&lt;/p&gt;




&lt;h2&gt;
  
  
  Pros and Cons
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Advantages of Getting This Right Early
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Clears the security questionnaire faster:&lt;/strong&gt; a vendor who can name their transfer mechanism (Path 1 or Path 2) and point to a monitoring record answers in one email what an unprepared vendor spends weeks scrambling to produce.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Avoids the annual-monitoring trap:&lt;/strong&gt; teams that document the Path 2 monitoring obligation as a recurring calendar item, not a one-time setup task, don't discover the gap during a customer's annual vendor security review.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Extraterritorial exposure is manageable, not existential:&lt;/strong&gt; once you accept APPI applies regardless of your entity's location, the compliance program itself is a known, bounded scope of work — not an open-ended risk.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Disadvantages and Risks
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The 3–5 day initial-report clock is unforgiving if unprepared:&lt;/strong&gt; an incident response plan that doesn't already define your qualifying criteria will burn most of that window just deciding whether you need to report at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Path 2's annual monitoring is easy to let lapse:&lt;/strong&gt; it's not a dramatic one-time certification, which is exactly why it's the compliance item most likely to quietly go stale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sector-specific expectations aren't written into APPI itself:&lt;/strong&gt; finance and critical-infrastructure buyers apply a stricter bar than the statute's floor, and a vendor reading only the statute (not the sector's actual procurement practice) will underestimate what's actually required to close that specific deal.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Is This Right for You?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxegtyxawvmtwy86v72y2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxegtyxawvmtwy86v72y2.png" alt="Is This Right for You?" width="784" height="1162"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prioritize Japan-only hosting if:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You're targeting finance, healthcare, or critical-infrastructure-adjacent customers specifically.&lt;/li&gt;
&lt;li&gt;A specific large deal has already surfaced data residency as a hard requirement, not a nice-to-have.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Multi-region with Path 2 is likely sufficient if:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your product is horizontal SaaS without sector-specific regulatory exposure.&lt;/li&gt;
&lt;li&gt;You can commit real process (not just a policy document) to annual monitoring.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Revisit your posture if:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your current answer to "how do we handle cross-border transfer" is a generic privacy policy clause — that's not one of APPI's two legitimate paths.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Implementation Approach
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Phase 1: Scope and Gap Assessment (Weeks 1–3)
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Confirm whether you process personal data of individuals in Japan today, not just whether you have a Japan entity.&lt;/li&gt;
&lt;li&gt;Identify your current cross-border transfer mechanism, if any — most foreign vendors discover at this step that they have none.&lt;/li&gt;
&lt;li&gt;Engage Japan-qualified counsel for a gap assessment against Path 1 and Path 2 requirements specifically.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Phase 2: Choose and Implement a Transfer Mechanism (Weeks 4–10)
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Decide Path 1 (consent-based) or Path 2 (equivalent-protection system) based on your product's signup and data-flow architecture.&lt;/li&gt;
&lt;li&gt;If Path 2: document the receiving-party protection measures formally, not informally, and calendar the first annual monitoring review now.&lt;/li&gt;
&lt;li&gt;Update privacy disclosures if pursuing Path 1, including the three required disclosure elements.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Phase 3: Breach Response Readiness (Weeks 8–12, overlapping Phase 2)
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Define your qualifying-breach criteria in advance (sensitive data, property-damage risk, unauthorized use, 1,000+ subjects) so an incident doesn't start with a definitional debate.&lt;/li&gt;
&lt;li&gt;Draft the initial and final report templates now, so the 3–5 day clock is a fill-in-the-template exercise, not a from-scratch drafting exercise.&lt;/li&gt;
&lt;li&gt;Assign a named incident owner responsible for the PPC reporting timeline specifically.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Phase 4: Ongoing Compliance (Month 4+)
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Execute the annual Path 2 monitoring review on a fixed calendar date, tracked the same way you'd track a compliance certification renewal.&lt;/li&gt;
&lt;li&gt;Revisit sector-specific expectations whenever you pursue a new vertical (finance, healthcare, government-adjacent) — APPI's floor and a specific sector's procurement bar are not the same thing.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Cost Considerations
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cost Type&lt;/th&gt;
&lt;th&gt;What to Budget For&lt;/th&gt;
&lt;th&gt;Typical Range&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Legal gap assessment&lt;/td&gt;
&lt;td&gt;Japan-qualified counsel review of current transfer mechanism&lt;/td&gt;
&lt;td&gt;$10,000–$25,000 one-time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Path 2 documentation and setup&lt;/td&gt;
&lt;td&gt;Formal protection-system documentation, initial assessment&lt;/td&gt;
&lt;td&gt;$15,000–$40,000 one-time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Annual monitoring (Path 2)&lt;/td&gt;
&lt;td&gt;Recurring review of receiving-party protection measures&lt;/td&gt;
&lt;td&gt;$5,000–$15,000/year&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Incident response readiness&lt;/td&gt;
&lt;td&gt;Runbook, report templates, named ownership&lt;/td&gt;
&lt;td&gt;2–3 weeks engineering/legal time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Japan-only hosting (if pursued)&lt;/td&gt;
&lt;td&gt;Dedicated Japan-region infrastructure&lt;/td&gt;
&lt;td&gt;Highly variable — model against your existing multi-region cost baseline&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;ROI signal:&lt;/strong&gt; the legal gap assessment and Path 2 setup cost is small relative to a single enterprise contract lost at the security-questionnaire stage — treat it as a cost of being enterprise-sales-ready in Japan, not a discretionary compliance nice-to-have.&lt;/p&gt;




&lt;h2&gt;
  
  
  Questions to Ask Your Team
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;"Do we actually know whether we're on Path 1 or Path 2 for cross-border transfer, or is our current answer a generic privacy policy clause?"&lt;/li&gt;
&lt;li&gt;"When was our Path 2 protection-system monitoring last actually reviewed, and is it on a calendar for next year?"&lt;/li&gt;
&lt;li&gt;"Could we produce an initial PPC breach report within 3–5 days today, or would we spend that window just deciding if we need to report?"&lt;/li&gt;
&lt;li&gt;"Are we treating 'Japan has no blanket data-localization law' as 'we have no requirement,' when our target sector's actual procurement bar says otherwise?"&lt;/li&gt;
&lt;li&gt;"Does our incident response plan name a specific owner for the PPC reporting timeline, or is that an assumption nobody's confirmed?"&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;APPI's extraterritorial reach means "we're not a Japanese company" was never a valid compliance position, and its lack of a blanket data-localization statute was never a reason to skip a real transfer mechanism. Pick Path 1 or Path 2 deliberately, calendar the Path 2 annual monitoring obligation before it lapses quietly, and build your incident response runbook around the 3–5 day initial-report clock before you need it under pressure. The vendors who get burned here aren't the ones facing a hostile regulator — they're the ones who read "no hard localization law" as "no real requirement" and never picked a transfer mechanism at all.&lt;/p&gt;




&lt;h2&gt;
  
  
  Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.ppc.go.jp/en/" rel="noopener noreferrer"&gt;Personal Information Protection Commission, Japan&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://iapp.org/news/a/practical-notes-for-japans-important-updates-of-the-appi-guidelines-and-qas" rel="noopener noreferrer"&gt;IAPP — Practical Notes for Japan's Amended APPI Guidelines and Q&amp;amp;As&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://resourcehub.bakermckenzie.com/en/resources/global-data-and-cyber-handbook/asia-pacific/japan/topics/security-requirements-and-breach-notification" rel="noopener noreferrer"&gt;Baker McKenzie — Security Requirements and Breach Notification: Japan&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;If this helped, a like and a follow are appreciated — and if you've solved this differently, drop a comment, I'd like to hear it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Bry Writes Code — cloud and AI infrastructure specialist, 15 years in IT, based in Tokyo. Building out your Japan data-compliance posture? &lt;a href="mailto:brywritescode@gmail.com"&gt;Let's talk&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>devops</category>
      <category>webdev</category>
      <category>beginners</category>
    </item>
    <item>
      <title>The New SI Deliverable: Selling Judgment and Integration, Not Headcount</title>
      <dc:creator>Bry</dc:creator>
      <pubDate>Wed, 16 Sep 2026 15:00:00 +0000</pubDate>
      <link>https://dev.to/brywritescode/the-new-si-deliverable-selling-judgment-and-integration-not-headcount-5a8o</link>
      <guid>https://dev.to/brywritescode/the-new-si-deliverable-selling-judgment-and-integration-not-headcount-5a8o</guid>
      <description>&lt;h2&gt;
  
  
  Key Points
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;For most of the industry's history, an SI's deliverable was, functionally, staffed capacity: a defined number of qualified people, for a defined duration, executing a defined scope. AI-assisted delivery is making that unit cheap to produce and hard to sell at the old price.&lt;/li&gt;
&lt;li&gt;What's replacing it won't be a single new product. Roughly 85% of major tech firms are already turning to specialist providers for data curation and model fine-tuning, a services category that doesn't route through headcount at all, but through platform access, curated data, and judgment applied to a client's specific constraints.&lt;/li&gt;
&lt;li&gt;Japan's largest SIers are placing genuinely different bets on what to sell instead: NTT Data toward "proposal-based SI that redesigns entire business processes," Fujitsu and Hitachi toward vertical, industry-specific judgment paired with AI (what Hitachi calls "Physical AI" for operational technology), NEC toward recurring revenue as its stated biggest management challenge.&lt;/li&gt;
&lt;li&gt;The through-line across all of these bets is that none of them are headcount-denominated. A client won't be buying a number of engineer-months for much longer. They'll be buying a specific, demonstrable judgment applied to their specific problem, sold as a defined package rather than a staffing commitment.&lt;/li&gt;
&lt;li&gt;SIs still trying to sell headcount with an "AI-powered" label attached to the invoice are already losing ground to competitors who've rebuilt their deliverable around something AI can't produce directly: legacy-system interpretation, cross-system governance, and the willingness to own the outcome, not just the hours.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;I sat in on a proposal review recently where a prospective client asked, pointedly, what exactly they were buying if not "a team of engineers for six months." That's been the honest answer to that question for as long as anyone in the room has worked in the industry. The account lead's answer, which took longer to land than it should have, was that the client was buying a specific integration outcome across three legacy systems that didn't talk to each other cleanly, delivered by people who'd solved that exact class of problem before, priced against the result rather than the hours it took to get there. That answer, awkward the first few times an SI has to give it out loud, is close to the actual new deliverable the entire industry is having to build toward.&lt;/p&gt;

&lt;p&gt;The old deliverable is staffed capacity: a defined headcount, at a defined skill level, for a defined duration, executing a scope the client (or the SI, on the client's behalf) has already worked out. AI-assisted development is making the execution portion of that deliverable cheap enough that selling it at the old price is stopping being credible, which article 500 in this thread covers from the pricing side. What's worth examining directly is what firms are actually building to replace it, because "sell judgment, not headcount" is easy to say and hard to operationalize into an actual, purchasable thing.&lt;/p&gt;

&lt;p&gt;The data-and-model-services category is the cleanest example of a genuinely new deliverable, not a relabeled old one. Roughly 85% of major tech firms are already turning to specialist providers for data curation and model fine-tuning, work that doesn't route through a headcount line item at all, but through platform access, curated proprietary data, and applied judgment about what a specific client's models actually need. Japan's largest SIers are making this concrete and public. NTT Data's stated shift is toward "proposal-based SI that redesigns entire business processes," backed by a specific 300-billion-yen AI-revenue target rather than a headcount target. Fujitsu and Hitachi are going the other direction, not toward infrastructure scale, but toward vertical, industry-specific judgment, with Hitachi explicitly betting on "Physical AI" applied to operational technology systems it already understands better than any AI-native competitor could. NEC has framed its own transition bluntly: recurring revenue, not staffing utilization, as its central management challenge going forward.&lt;/p&gt;

&lt;h2&gt;
  
  
  Old SI Deliverable vs. New SI Deliverable
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criteria&lt;/th&gt;
&lt;th&gt;Staffed-Capacity Deliverable (legacy)&lt;/th&gt;
&lt;th&gt;Judgment/Integration Deliverable (where this is heading)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What the client is actually buying&lt;/td&gt;
&lt;td&gt;A defined headcount for a defined duration&lt;/td&gt;
&lt;td&gt;A specific outcome, applied judgment, or a platform capability configured to the client's constraints&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;How it's priced&lt;/td&gt;
&lt;td&gt;Rate card times hours times headcount&lt;/td&gt;
&lt;td&gt;Outcome fee, usage-based platform access, or a defined asset license&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What makes the vendor hard to replace&lt;/td&gt;
&lt;td&gt;Availability of enough qualified staff&lt;/td&gt;
&lt;td&gt;Depth of judgment in a specific vertical or system class the client can't easily replicate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vendor's real product&lt;/td&gt;
&lt;td&gt;Labor supply&lt;/td&gt;
&lt;td&gt;Curated data, applied domain judgment, integration ownership, or a licensed platform capability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Effect of a client bringing AI tooling in-house&lt;/td&gt;
&lt;td&gt;Direct threat: replaces the exact thing being sold&lt;/td&gt;
&lt;td&gt;Limited threat: the client still lacks the specific judgment or curated asset being sold&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Recommendation:&lt;/strong&gt; if a client could replicate your entire deliverable just by licensing the same AI tools you use, you're still selling headcount with different branding. The deliverable that survives client-side AI adoption is the one built on something the client genuinely can't get by buying the tool directly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building a Judgment-Based Deliverable: What Firms Are Changing Internally
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Firms are inventorying what they know that AI tooling doesn't.&lt;/strong&gt; Legacy-system quirks, a specific regulator's actual enforcement patterns, a client's internal politics around a prior failed migration: these are turning out to be real, sellable assets once firms stop treating them as background context and start treating them as the product.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Firms are building defined, packaged offerings instead of open-ended staffing.&lt;/strong&gt; NTT Data's "redesigns entire business processes" framing and Hitachi's vertical Physical AI bet are both packaged, named offerings, not "send us your requirements and we'll staff a team," but a defined scope built around a specific judgment claim.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Firms are changing who gets hired and promoted.&lt;/strong&gt; Selling judgment instead of headcount means fewer people needed for routine execution and a premium on people who can own client relationships, interpret ambiguous legacy systems, and make architectural calls AI tooling can propose but not confidently commit to.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Firms are accepting that the new deliverable serves fewer clients per unit of revenue at first.&lt;/strong&gt; A judgment-based engagement doesn't scale the way staffed-capacity did. Replacing headcount-scaling with judgment-scaling is a real, temporarily uncomfortable revenue-model change, not a free upgrade.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Questions to Ask Your Team
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;If a client asked us directly what they're buying if not a team of engineers, do we have a specific answer, a named judgment, a curated asset, an integration commitment, or would we describe our own deliverable in headcount terms without meaning to?&lt;/li&gt;
&lt;li&gt;What do we know, about our clients' systems or industry, that AI tooling genuinely doesn't have access to, and are we pricing that specifically, or treating it as unpriced background expertise?&lt;/li&gt;
&lt;li&gt;Have we built a packaged, named offering around that judgment, the way NTT Data and Hitachi are, or are we still proposing open-ended staffing with an AI label attached?&lt;/li&gt;
&lt;li&gt;Are we prepared for a judgment-based deliverable to serve fewer clients per unit of revenue than staffed capacity did, at least initially?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The systems integrators most likely to make it through this transition won't do it by getting better at the old deliverable faster than competitors. They're building a different one, priced against judgment, curated assets, and integration ownership instead of engineer-hours, and Japan's largest firms are already making genuinely different bets on what that judgment should be: process redesign for NTT Data, vertical Physical AI for Hitachi, recurring revenue as the explicit target for NEC. None of those bets are headcount-denominated, and that's the actual signal. The firms still describing their deliverable as "a team of qualified engineers" are selling something a client can increasingly get more cheaply somewhere else. The ones that aren't are figuring out, specifically and early, what they know that AI doesn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://envistacorp.com/blog/the-value-of-a-systems-integrator-in-the-ai-era/" rel="noopener noreferrer"&gt;enVista: The Value of a Systems Integrator in the AI Era&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.altmansolon.com/thought-leadership/enterprise-ai-systems-integrators" rel="noopener noreferrer"&gt;Altman Solon: Reinventing Systems Integration in the Gen AI Era&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://note.com/zephel01/n/n4993598ac473?hl=en" rel="noopener noreferrer"&gt;note.com: Generative AI, Strategies of 6 Major Japanese Companies (2026 Edition)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Bry Writes Code; cloud and AI infrastructure specialist. Still describing your deliverable as a headcount commitment? &lt;a href="mailto:brywritescode@gmail.com"&gt;Let's talk&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>business</category>
      <category>ai</category>
      <category>consulting</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Building a RAG Pipeline from Scratch: Embeddings, Retrieval, and Claude</title>
      <dc:creator>Bry</dc:creator>
      <pubDate>Tue, 15 Sep 2026 15:00:00 +0000</pubDate>
      <link>https://dev.to/brywritescode/building-a-rag-pipeline-from-scratch-embeddings-retrieval-and-claude-1h96</link>
      <guid>https://dev.to/brywritescode/building-a-rag-pipeline-from-scratch-embeddings-retrieval-and-claude-1h96</guid>
      <description>&lt;h2&gt;
  
  
  Key Points
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;RAG (Retrieval-Augmented Generation)&lt;/strong&gt; connects an LLM to your private data at query time — no fine-tuning, no retraining, no data leakage into model weights.&lt;/li&gt;
&lt;li&gt;The pipeline has five stages: &lt;strong&gt;Ingest → Chunk → Embed → Store → Query&lt;/strong&gt; (retrieve, augment, generate). Getting chunking and retrieval right matters more than which LLM you pick.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ChromaDB&lt;/strong&gt; runs in-process for local dev with zero infrastructure — &lt;code&gt;collection.add()&lt;/code&gt; to insert, &lt;code&gt;collection.query()&lt;/code&gt; to retrieve. Production options include Pinecone and pgvector.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Common failure modes&lt;/strong&gt; — chunks too large, no deduplication, no query expansion, skipping evaluation — are all avoidable. This article shows how.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;Large language models hallucinate when they don't know the answer. The standard fix — fine-tuning on your private data — costs tens of thousands of dollars, takes weeks, and produces a static model that goes stale as soon as your data changes. RAG solves a different problem: instead of baking knowledge into weights, it retrieves the relevant facts at query time and gives them to the model as context. The model answers from evidence, not memory. I've built RAG pipelines for internal knowledge bases and customer-facing chat tools — the architecture is the same whether you're indexing 500 internal docs or 500,000 support tickets; the parameters are what change.&lt;/p&gt;

&lt;p&gt;This article builds a complete RAG pipeline in Python using ChromaDB for vector storage, &lt;code&gt;sentence-transformers&lt;/code&gt; for local embeddings, and Claude for generation. You will understand each stage, why the design choices matter, and what breaks in production if you skip them. The full working implementation is in &lt;code&gt;src/pipeline.py&lt;/code&gt; and &lt;code&gt;src/main.py&lt;/code&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  What RAG Is — and What It Replaces
&lt;/h2&gt;

&lt;p&gt;Before writing code, align on why RAG exists.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fine-tuning&lt;/strong&gt; trains the model on your data. It is expensive (typically $5,000–$50,000+ for a production run), takes days to weeks, and produces a checkpoint that bakes in your data as of the training date. If your documentation changes next week, the model doesn't know. Fine-tuning is correct when you need the model to learn a style, a domain vocabulary, or a task format — not when you need it to answer questions from a document set.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt-only (stuffing)&lt;/strong&gt; puts your documents directly into the context window. It works for small document sets (tens of pages) but breaks on large corpora: context windows have limits, filling them with irrelevant text degrades answer quality, and at scale the cost per query becomes prohibitive.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RAG&lt;/strong&gt; indexes your documents, retrieves only the chunks relevant to each query, and gives the model a focused context. It handles large corpora, stays current as documents are updated, and costs a fraction of fine-tuning. The tradeoff is a retrieval layer you have to build and maintain.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Data scale&lt;/th&gt;
&lt;th&gt;Latency&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;th&gt;Freshness&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Fine-tuning&lt;/td&gt;
&lt;td&gt;Any&lt;/td&gt;
&lt;td&gt;Low (no retrieval)&lt;/td&gt;
&lt;td&gt;High (training + inference)&lt;/td&gt;
&lt;td&gt;Static — retrain to update&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt-only&lt;/td&gt;
&lt;td&gt;Small (&amp;lt; 50 pages)&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Fresh&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;RAG&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Large (any)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Medium (retrieval + LLM)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Low–Medium&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Fresh (re-embed on update)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  RAG Pipeline Architecture
&lt;/h2&gt;

&lt;p&gt;A RAG pipeline has two sides: the &lt;strong&gt;ingest side&lt;/strong&gt; (run once per document update) and the &lt;strong&gt;query side&lt;/strong&gt; (run on every user request). Together they form five stages.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv5t12u3s1nqh0mr6vmpp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv5t12u3s1nqh0mr6vmpp.png" alt="RAG Pipeline Architecture" width="800" height="1290"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Diagram: Ingest side (left) runs document processing once. Query side (right) runs on every user request.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The five stages:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Ingest&lt;/strong&gt; — load raw documents. Source can be text files, PDFs, database records, or API responses.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chunk&lt;/strong&gt; — split documents into pieces small enough to embed meaningfully. Chunk size is the most consequential parameter in the pipeline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embed&lt;/strong&gt; — convert each chunk to a vector. The embedding model maps semantic meaning to a point in high-dimensional space.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Store&lt;/strong&gt; — persist vectors (with their source text and metadata) in a vector database.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Query&lt;/strong&gt; — embed the user's question, find the nearest chunks, inject them into a prompt, and generate an answer.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Chunking Strategies
&lt;/h2&gt;

&lt;p&gt;Chunking is where most RAG pipelines fail. A bad chunking strategy produces irrelevant retrievals; irrelevant retrievals produce hallucinated answers. The model can't conjure information that wasn't in the retrieved chunks. In practice, I spend more time tuning chunk size and overlap than on any other pipeline parameter — getting it wrong is invisible until a user catches the model confidently citing the wrong passage.&lt;/p&gt;

&lt;h3&gt;
  
  
  Fixed-Size Chunking
&lt;/h3&gt;

&lt;p&gt;Split every N characters or tokens. Simple to implement, ignores document structure.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Fixed-size chunking — simple but blunt
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;chunk_fixed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;overlap&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;end&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt;
        &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;overlap&lt;/span&gt;  &lt;span class="c1"&gt;# overlap preserves context at boundaries
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The overlap matters: without it, a sentence split across two chunks retrieves each half separately and both halves are incomplete.&lt;/p&gt;

&lt;h3&gt;
  
  
  Recursive Chunking
&lt;/h3&gt;

&lt;p&gt;Try splitting on paragraph breaks (&lt;code&gt;\n\n&lt;/code&gt;), then line breaks (&lt;code&gt;\n&lt;/code&gt;), then spaces. Each level is a fallback when the previous splitter still produces chunks over the target size. This respects document structure — paragraphs before sentences before words.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_chunk&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Recursively split text into chunks, preferring natural boundaries.

    Args:
        text: The input text to split.
        max_tokens: Approximate maximum chunk size in tokens (1 token ≈ 4 chars).

    Returns:
        List of text chunks, each within the max_tokens limit.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;max_chars&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;  &lt;span class="c1"&gt;# rough token → char estimate
&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;max_chars&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()]&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

    &lt;span class="c1"&gt;# Try splitting on paragraph breaks first, then newlines, then spaces
&lt;/span&gt;    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;separator&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="n"&gt;parts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;separator&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
            &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;candidate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;separator&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;candidate&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;max_chars&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;candidate&lt;/span&gt;
                &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                        &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                    &lt;span class="c1"&gt;# Recurse on oversized parts
&lt;/span&gt;                    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;max_chars&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                        &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;extend&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_chunk&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
                        &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;
                    &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                        &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="c1"&gt;# No separator found — hard split
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="n"&gt;max_chars&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;max_chars&lt;/span&gt;&lt;span class="p"&gt;:].&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Semantic Chunking
&lt;/h3&gt;

&lt;p&gt;Group sentences by semantic similarity — a sentence joins the current chunk if its embedding is similar enough; otherwise it starts a new chunk. Produces the most coherent chunks but is slower (requires embedding every sentence during ingestion).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkt3abefi2ps8o68zskdr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkt3abefi2ps8o68zskdr.png" alt="Semantic Chunking" width="800" height="1371"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Diagram: Chunking strategy selection — recursive approach, falling back from paragraphs to lines to words.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Embeddings
&lt;/h2&gt;

&lt;p&gt;An embedding model converts text into a dense vector — a list of floats (e.g., 384 dimensions for &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt;) where similar texts produce similar vectors. Semantic similarity becomes geometric proximity.&lt;/p&gt;

&lt;h3&gt;
  
  
  Local Embeddings: &lt;code&gt;sentence-transformers&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;sentence-transformers&lt;/code&gt; runs entirely locally — no API key, no per-token cost, no latency from network calls.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sentence_transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SentenceTransformer&lt;/span&gt;

&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SentenceTransformer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;all-MiniLM-L6-v2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;texts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Python is a high-level programming language.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Guido van Rossum created Python in 1991.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The weather is sunny today.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;embeddings&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;texts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# shape: (3, 384)
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;embeddings&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;           &lt;span class="c1"&gt;# (3, 384)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt; is 80MB on disk, encodes ~14,000 sentences per second on CPU, and produces 384-dimensional vectors. It is the right default for local development and prototyping.&lt;/p&gt;

&lt;h3&gt;
  
  
  API Embeddings
&lt;/h3&gt;

&lt;p&gt;For production, Anthropic's and OpenAI's embedding APIs trade local cost for higher-quality vectors at scale. Use them when your retrieval accuracy on domain-specific content drops below acceptable thresholds.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Dims&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sentence-transformers&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;384&lt;/td&gt;
&lt;td&gt;Free (local)&lt;/td&gt;
&lt;td&gt;Best for dev/prototyping&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sentence-transformers&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;all-mpnet-base-v2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;768&lt;/td&gt;
&lt;td&gt;Free (local)&lt;/td&gt;
&lt;td&gt;Higher quality, 3× slower&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI API&lt;/td&gt;
&lt;td&gt;&lt;code&gt;text-embedding-3-small&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1536&lt;/td&gt;
&lt;td&gt;$0.02 / 1M tokens&lt;/td&gt;
&lt;td&gt;Good price/quality tradeoff&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic API&lt;/td&gt;
&lt;td&gt;Voyage-3 (via Voyage AI)&lt;/td&gt;
&lt;td&gt;1024&lt;/td&gt;
&lt;td&gt;$0.06 / 1M tokens&lt;/td&gt;
&lt;td&gt;SOTA quality for RAG&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Keep your embedding model consistent between ingestion and query time. If you embed documents with &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt; and query with &lt;code&gt;text-embedding-3-small&lt;/code&gt;, your similarity scores will be meaningless.&lt;/p&gt;




&lt;h2&gt;
  
  
  Vector Storage with ChromaDB
&lt;/h2&gt;

&lt;p&gt;ChromaDB is an in-process vector database for Python. It requires no server, no Docker container, no cloud account — import it and use it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;chromadb&lt;/span&gt;

&lt;span class="c1"&gt;# In-memory client — resets between runs (good for unit tests)
&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chromadb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Client&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# Persistent client — stores to disk (good for development)
&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chromadb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;PersistentClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;./chroma_db&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;collection&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_or_create_collection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;my_docs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Add documents with pre-computed embeddings
&lt;/span&gt;&lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;ids&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chunk_0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chunk_1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chunk_2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;embeddings&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="mf"&gt;0.1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...],&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...],&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mf"&gt;0.9&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...]],&lt;/span&gt;
    &lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Python is...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Guido van Rossum...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Weather is...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;metadatas&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;python_intro.txt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;python_intro.txt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;weather.txt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Query — returns top-3 nearest chunks
&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;query_embeddings&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="mf"&gt;0.15&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.18&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...]],&lt;/span&gt;
    &lt;span class="n"&gt;n_results&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;documents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;metadatas&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;distances&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dist&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;documents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;metadatas&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;distances&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;dist&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;] &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;source&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;ChromaDB uses cosine similarity by default. Lower distance = more similar.&lt;/p&gt;

&lt;h3&gt;
  
  
  Choosing a Vector Database
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;ChromaDB&lt;/th&gt;
&lt;th&gt;Pinecone&lt;/th&gt;
&lt;th&gt;pgvector&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Setup&lt;/td&gt;
&lt;td&gt;Zero (in-process)&lt;/td&gt;
&lt;td&gt;Managed cloud&lt;/td&gt;
&lt;td&gt;Add extension to Postgres&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scale&lt;/td&gt;
&lt;td&gt;Millions of vectors&lt;/td&gt;
&lt;td&gt;Billions&lt;/td&gt;
&lt;td&gt;Millions (with tuning)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost&lt;/td&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;td&gt;$70/month+ (managed)&lt;/td&gt;
&lt;td&gt;Postgres hosting cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Filtering&lt;/td&gt;
&lt;td&gt;Metadata filters&lt;/td&gt;
&lt;td&gt;Metadata filters&lt;/td&gt;
&lt;td&gt;Full SQL WHERE clauses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Persistence&lt;/td&gt;
&lt;td&gt;Local disk&lt;/td&gt;
&lt;td&gt;Cloud-managed&lt;/td&gt;
&lt;td&gt;Postgres storage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;Local dev, prototypes&lt;/td&gt;
&lt;td&gt;Production at scale&lt;/td&gt;
&lt;td&gt;Teams already on Postgres&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Start with ChromaDB locally — it removes all infrastructure friction during development, which is where you need to iterate fastest. I reach for pgvector over Pinecone when the team is already on Postgres: the operational overhead is near-zero and full SQL filtering eliminates an entire class of retrieval bugs that metadata-only filters can't handle.&lt;/p&gt;




&lt;h2&gt;
  
  
  Retrieval Strategies
&lt;/h2&gt;

&lt;p&gt;Retrieval is not just "find the nearest vectors." Three strategies matter in practice.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cosine Similarity Search
&lt;/h3&gt;

&lt;p&gt;The default. Compute the cosine similarity between the query vector and every stored vector; return the top-k most similar chunks. Fast, well-understood, and works well when the query language matches the document language.&lt;/p&gt;

&lt;h3&gt;
  
  
  Maximal Marginal Relevance (MMR)
&lt;/h3&gt;

&lt;p&gt;Standard top-k retrieval can return five chunks that all say the same thing — high similarity, low diversity. MMR trades some similarity for diversity: each additional chunk is selected to be both similar to the query &lt;em&gt;and&lt;/em&gt; different from already-selected chunks.&lt;/p&gt;

&lt;p&gt;ChromaDB does not implement MMR natively. Implement it post-retrieval:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_mmr_rerank&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;query_embedding&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;candidate_embeddings&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt;
    &lt;span class="n"&gt;candidate_docs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;lambda_&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Select top-k chunks using Maximal Marginal Relevance.

    Args:
        query_embedding: Embedded user query.
        candidate_embeddings: Embeddings of candidate chunks (over-fetch, e.g., top-20).
        candidate_docs: Text of each candidate chunk.
        k: Number of chunks to return.
        lambda_: Trade-off between relevance (1.0) and diversity (0.0).

    Returns:
        k chunks selected for both relevance and diversity.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;

    &lt;span class="n"&gt;q&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query_embedding&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;cands&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;candidate_embeddings&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Cosine similarity: query vs candidates
&lt;/span&gt;    &lt;span class="n"&gt;sim_to_query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cands&lt;/span&gt; &lt;span class="o"&gt;@&lt;/span&gt; &lt;span class="n"&gt;q&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;linalg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cands&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;axis&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;linalg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;q&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;1e-9&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;selected_indices&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="n"&gt;remaining&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;candidate_docs&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;remaining&lt;/span&gt;&lt;span class="p"&gt;))):&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;selected_indices&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="c1"&gt;# First pick: highest similarity to query
&lt;/span&gt;            &lt;span class="n"&gt;best&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;remaining&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;sim_to_query&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="c1"&gt;# Subsequent picks: balance relevance vs redundancy
&lt;/span&gt;            &lt;span class="n"&gt;selected_vecs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cands&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;selected_indices&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
            &lt;span class="n"&gt;scores&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;remaining&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;sim_q&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sim_to_query&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
                &lt;span class="n"&gt;sim_selected&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cands&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;@&lt;/span&gt; &lt;span class="n"&gt;selected_vecs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt;
                    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;linalg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cands&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;linalg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;selected_vecs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;1e-9&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;selected_vecs&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
                &lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lambda_&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;sim_q&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;lambda_&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;sim_selected&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;best&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;remaining&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;index&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;))]&lt;/span&gt;

        &lt;span class="n"&gt;selected_indices&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;best&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;remaining&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;remove&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;best&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;candidate_docs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;selected_indices&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Hybrid Search (BM25 + Semantic)
&lt;/h3&gt;

&lt;p&gt;Semantic search handles paraphrase and synonymy well. BM25 (keyword search) handles exact terms — product codes, names, technical identifiers — better. Hybrid search runs both and combines scores (typically via Reciprocal Rank Fusion). Use hybrid when your documents contain precise identifiers that semantic search might miss. I default to MMR over top-k similarity the moment a domain has redundant content — policy documentation, API reference pages, and anything generated from a template will poison straight similarity retrieval with near-duplicate chunks every time.&lt;/p&gt;




&lt;h2&gt;
  
  
  Context Augmentation
&lt;/h2&gt;

&lt;p&gt;Retrieved chunks are useful only if the prompt tells Claude how to use them. A weak prompt ("here are some documents, answer the question") produces weak answers. A well-structured prompt produces grounded, citable answers.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_build_prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Format retrieved chunks and question into a grounded generation prompt.

    Args:
        question: The user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s original question.
        chunks: Retrieved document chunks, ordered by relevance (most relevant first).

    Returns:
        A formatted prompt that instructs Claude to cite sources and acknowledge gaps.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;context_block&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s"&gt;---&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[Source &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;]&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;You are a helpful assistant. Answer the question below using only the provided sources.

Rules:
- Cite which source(s) support each claim (e.g., &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;According to Source 2...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;).
- If the sources do not contain enough information to answer the question, say so explicitly.
- Do not use knowledge outside the provided sources.

Sources:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;context_block&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Question: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Answer:&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The three rules matter. Citing sources forces the model to ground claims in retrieved text rather than training data. Acknowledging gaps prevents confident-sounding hallucinations. Forbidding outside knowledge keeps the model from blending retrieved context with parametric knowledge unpredictably.&lt;/p&gt;




&lt;h2&gt;
  
  
  Generation with Claude
&lt;/h2&gt;

&lt;p&gt;The query method ties everything together: embed the question, retrieve chunks, build the prompt, call Claude, and return the answer.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzt2xkkavrboj8ao6kvgd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzt2xkkavrboj8ao6kvgd.png" alt="Generation with Claude" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Diagram: Full query path from user question to Claude-generated answer.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# reads ANTHROPIC_API_KEY from environment
&lt;/span&gt;
&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-sonnet-4-6&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Check why generation stopped
&lt;/span&gt;&lt;span class="n"&gt;match&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stop_reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;case&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;end_turn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;
    &lt;span class="n"&gt;case&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# Response was cut off — increase max_tokens or reduce prompt length
&lt;/span&gt;        &lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s"&gt;[Response truncated — max_tokens reached]&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;case&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[Generation stopped: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stop_reason&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;]&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Always check &lt;code&gt;stop_reason&lt;/code&gt;. &lt;code&gt;end_turn&lt;/code&gt; means the model finished naturally. &lt;code&gt;max_tokens&lt;/code&gt; means the answer was cut off — a common silent failure when retrieved chunks are large and the combined prompt + answer exceeds your token budget.&lt;/p&gt;

&lt;p&gt;For cheap classification tasks within a RAG system (e.g., query routing, intent detection, relevance filtering), use &lt;code&gt;claude-haiku-4-5-20251001&lt;/code&gt; instead of &lt;code&gt;claude-sonnet-4-6&lt;/code&gt;. Haiku is significantly faster and cheaper for short-context classification where generation quality is less critical.&lt;/p&gt;




&lt;h2&gt;
  
  
  The &lt;code&gt;RAGPipeline&lt;/code&gt; Class
&lt;/h2&gt;

&lt;p&gt;The full implementation encapsulates all stages into a single class with three public methods: &lt;code&gt;ingest&lt;/code&gt;, &lt;code&gt;query&lt;/code&gt;, and the private helpers.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frp6rmf8y460i0uw0gyu1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frp6rmf8y460i0uw0gyu1.png" alt="RAGPipeline Class" width="726" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Diagram: &lt;code&gt;RAGPipeline&lt;/code&gt; class structure — three public methods, three private helpers.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The constructor initializes all three dependencies. Chunk and embed are private because callers should not need to call them directly — they are implementation details of &lt;code&gt;ingest&lt;/code&gt; and &lt;code&gt;query&lt;/code&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Evaluation
&lt;/h2&gt;

&lt;p&gt;A RAG pipeline with no evaluation is a guess. Add measurement before shipping.&lt;/p&gt;

&lt;p&gt;Three metrics from the &lt;a href="https://docs.ragas.io/" rel="noopener noreferrer"&gt;RAGAS framework&lt;/a&gt; cover the core failure modes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;What it measures&lt;/th&gt;
&lt;th&gt;How it catches failures&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Faithfulness&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Is every claim in the answer supported by a retrieved chunk?&lt;/td&gt;
&lt;td&gt;Catches hallucinations — the model added facts not in context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Answer Relevancy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Does the answer actually address the question?&lt;/td&gt;
&lt;td&gt;Catches topic drift — the model answered a different question&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Context Recall&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Did retrieval surface the chunks needed to answer?&lt;/td&gt;
&lt;td&gt;Catches retrieval failures — the right chunks weren't found&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Run evaluation on a golden dataset — 20–50 question/answer pairs you've manually verified. Automate it in CI so a configuration change that breaks retrieval doesn't silently ship.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# RAGAS evaluation — requires: pip install ragas
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;ragas&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;evaluate&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;ragas.metrics&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;faithfulness&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;answer_relevancy&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context_recall&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datasets&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Dataset&lt;/span&gt;

&lt;span class="n"&gt;eval_data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;question&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What is Python?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Who created Python?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;answer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Python is a programming language.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Guido van Rossum.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;contexts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Python is a high-level language...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Guido van Rossum created Python in 1991...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ground_truth&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Python is a high-level programming language.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Guido van Rossum.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;evaluate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;Dataset&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;eval_data&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;metrics&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;faithfulness&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;answer_relevancy&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context_recall&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Common Mistakes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Mistake 1: Chunks larger than 600 tokens drown the retrieval signal.&lt;/strong&gt;&lt;br&gt;
A 600-token chunk covers multiple topics. When you embed it, the vector averages over those topics and becomes less specific to any one of them. Similarity search returns chunks that are vaguely related to the query, not precisely relevant. Keep chunks under 400 tokens. If a passage requires more context, overlap consecutive chunks by 50–100 tokens rather than enlarging the chunk size.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mistake 2: Not deduplicating similar chunks before storing.&lt;/strong&gt;&lt;br&gt;
If your corpus contains repeated passages (headers, boilerplate, legal disclaimers), those chunks will dominate retrieval — every query returns five variations of the same paragraph. Deduplicate before ingestion: compute a hash of each chunk's text and discard exact duplicates, then use a similarity threshold (cosine similarity &amp;gt; 0.95) to collapse near-duplicates. I've seen this sink a support-bot demo: the corpus was a product manual where the safety warning appeared verbatim on 40 of 200 pages — top-5 retrieval returned the warning for nearly every query, and the model dutifully answered "please refer to a qualified technician" regardless of what was asked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mistake 3: Embedding the raw query without expansion.&lt;/strong&gt;&lt;br&gt;
A user asking "how does Python handle memory?" may not use the same words as your documentation ("garbage collection", "reference counting", "memory management"). Query expansion generates alternative phrasings before embedding: &lt;code&gt;original_query + hypothetical_answer_keywords&lt;/code&gt;. Hypothetical Document Embeddings (HyDE) generates a fake answer to the question and embeds that instead — the fake answer often matches document language better than the raw question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mistake 4: Skipping evaluation entirely.&lt;/strong&gt;&lt;br&gt;
Most RAG pipelines are shipped based on developer intuition ("the answers look right in testing"). Without a golden dataset and automated metrics, you won't know when a configuration change — a new chunk size, a different top-k, an updated embedding model — makes retrieval worse. Define your evaluation dataset before you tune parameters, not after. I've watched a team spend two weeks tuning chunk size upward because answers "felt" more complete — then discover via RAGAS that faithfulness had dropped from 0.89 to 0.71 because larger chunks were diluting the retrieved context with off-topic sentences.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mistake 5: Using the same model for embedding and generation.&lt;/strong&gt;&lt;br&gt;
Anthropic's Claude models are generation models, not embedding models. Using a generation model's hidden states as embeddings produces inferior vectors compared to models trained specifically for semantic similarity (sentence-transformers, OpenAI's &lt;code&gt;text-embedding-3-*&lt;/code&gt;, Voyage AI). Keep embedding and generation as separate concerns: sentence-transformers (or a dedicated embedding API) for vectors, Claude for generation.&lt;/p&gt;


&lt;h2&gt;
  
  
  Full Example
&lt;/h2&gt;

&lt;p&gt;The complete implementation lives in two files:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;src/pipeline.py&lt;/code&gt;&lt;/strong&gt; — &lt;code&gt;RAGPipeline&lt;/code&gt; class with all methods documented and implemented.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;src/main.py&lt;/code&gt;&lt;/strong&gt; — demo that ingests five documents about Python and runs three queries.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Setup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv
&lt;span class="nb"&gt;source&lt;/span&gt; .venv/bin/activate        &lt;span class="c"&gt;# Windows: .venv\Scripts\activate&lt;/span&gt;
pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt

&lt;span class="nb"&gt;cp&lt;/span&gt; .env.example .env
&lt;span class="c"&gt;# Add your ANTHROPIC_API_KEY to .env&lt;/span&gt;

python src/main.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Expected output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Ingested 5 documents → 12 chunks stored.

Q: What is Python's GIL?
A: According to Source 1, Python's Global Interpreter Lock (GIL) is a mutex that protects access to Python objects, preventing multiple threads from executing Python bytecodes simultaneously...

Q: Who created Python and when?
A: According to Source 3, Python was created by Guido van Rossum and first released in 1991...

Q: What is list comprehension?
A: According to Source 2, list comprehension is a concise syntax for creating lists based on existing iterables...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;Full source: &lt;a href="https://github.com/brywritescode/bry-writes-code-examples.git" rel="noopener noreferrer"&gt;GitHub link&lt;/a&gt; → &lt;code&gt;ai-integration/rag-pipeline/&lt;/code&gt; — see README for setup steps.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;RAG is the right architecture when your LLM needs access to private, large, or frequently updated data. The implementation is not what makes or breaks it — the parameters are: chunk size, overlap, retrieval strategy, and whether you bother to measure any of it. Most pipelines that fail in production fail because they were tuned by intuition and never measured. Add the RAGAS evaluation step before you ship anything to users, keep your chunks under 400 tokens, and you will avoid the failure modes that make RAG look unreliable. The architecture isn't fragile — the shortcuts are.&lt;/p&gt;




&lt;h2&gt;
  
  
  Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://docs.anthropic.com/en/api/messages" rel="noopener noreferrer"&gt;Anthropic API Documentation — Messages&lt;/a&gt; — official reference for the &lt;code&gt;messages.create&lt;/code&gt; endpoint, &lt;code&gt;stop_reason&lt;/code&gt; values, and model IDs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.trychroma.com/" rel="noopener noreferrer"&gt;ChromaDB Documentation&lt;/a&gt; — official docs covering collections, embedding functions, metadata filtering, and persistence&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.sbert.net/" rel="noopener noreferrer"&gt;sentence-transformers Documentation&lt;/a&gt; — model selection guide, semantic similarity, and batch encoding&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.ragas.io/" rel="noopener noreferrer"&gt;RAGAS Documentation&lt;/a&gt; — RAG evaluation framework; faithfulness, answer relevancy, and context recall metrics&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.voyageai.com/docs/embeddings" rel="noopener noreferrer"&gt;Voyage AI Embeddings (Anthropic ecosystem)&lt;/a&gt; — production-quality embedding API recommended by Anthropic for RAG applications&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;If this helped, a like and a follow are appreciated — and if you've solved this differently, drop a comment, I'd like to hear it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Bry Writes Code — cloud and AI infrastructure specialist. Building a RAG system? &lt;a href="mailto:brywritescode@gmail.com"&gt;Let's talk&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>llm</category>
      <category>programming</category>
    </item>
    <item>
      <title>Who Absorbs the Margin — Client, SIer, or the AI Vendor: Renegotiating SI Economics</title>
      <dc:creator>Bry</dc:creator>
      <pubDate>Wed, 09 Sep 2026 15:00:00 +0000</pubDate>
      <link>https://dev.to/brywritescode/who-absorbs-the-margin-client-sier-or-the-ai-vendor-renegotiating-si-economics-2ce2</link>
      <guid>https://dev.to/brywritescode/who-absorbs-the-margin-client-sier-or-the-ai-vendor-renegotiating-si-economics-2ce2</guid>
      <description>&lt;h2&gt;
  
  
  Key Points
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Every dollar an AI tool saves an SI on delivery cost is a dollar three different parties have a plausible claim to: the client (who's aware operating costs went down), the AI platform vendor (whose licensing terms increasingly shift toward capturing usage-based value), and the SI (whose judgment turned a raw capability into a delivered outcome).&lt;/li&gt;
&lt;li&gt;Procurement teams are learning to sort renewals into three buckets: renew and absorb, renegotiate to consumption, or replace and build. Expect more SIs to get pushed into the second and third buckets as clients understand AI has genuinely lowered the SI's own cost base.&lt;/li&gt;
&lt;li&gt;93% of sellers reported struggling to quantify and defend the value they'd actually added, which puts most SIs into these renegotiations without a real number to counter a client's discount demand.&lt;/li&gt;
&lt;li&gt;SIs that come through this pressure well are the ones that stop treating margin as something to defend by obscuring cost structure, and start treating it as something to defend by pricing a specific, demonstrable judgment contribution the client couldn't get from the AI platform directly.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;A client's CFO said something to me directly not long ago that a lot of SI account teams are starting to hear some version of: "I know what your AI licensing costs. I know what it used to cost you to deliver this. Explain to me why I'm still paying the old number." It's not an unreasonable question, and at this point in the AI-services transition, most SIs don't have a rehearsed answer for it. For a while, plenty of firms have been quietly pocketing the AI-driven cost reduction as pure margin, hoping nobody with enough visibility into the underlying economics would ask.&lt;/p&gt;

&lt;p&gt;The honest framing is that AI-driven delivery savings don't belong to any one party by default. They're contested, three ways. Clients have a real claim: operational costs genuinely came down, and clients increasingly know it, because AI tool pricing itself became public and metered rather than opaque. AI platform vendors have a claim too, since a growing share of enterprise software licensing is shifting toward capturing usage-based value directly, rather than just charging a flat seat fee and letting downstream margin sit wherever it lands. And the SI has a claim, the one most vulnerable to getting argued away in a renegotiation: the actual judgment, integration work, and delivery risk absorbed in turning a raw AI capability into something a client can actually rely on in production.&lt;/p&gt;

&lt;p&gt;Procurement organizations are getting sophisticated about this fast. A fairly standard renewal framework is already sorting vendor relationships into three buckets: renew and absorb (accept the current pricing, usually for low-volatility, well-understood work), renegotiate to consumption (move rigid seat- or hour-based pricing to a usage-based model that tracks real utilization), or replace and build (walk away from the vendor relationship entirely and bring the capability in-house, now that AI has made that a realistic option for more capabilities than it used to be). SIs that show up to these conversations without a defensible, quantified account of their own value contribution are getting sorted into renegotiate-to-consumption or replace-and-build far more often than SIs that can show precisely what they still add. A Holden Advisors study found 93% of sellers admitted they struggled to accurately quantify and defend their value, which means most SIs are walking into these renegotiations structurally unprepared for a client who's done their homework.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the Margin Actually Goes: Three Claims on the Same Savings
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Claimant&lt;/th&gt;
&lt;th&gt;Basis for the Claim&lt;/th&gt;
&lt;th&gt;What Wins the Argument in Practice&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Client&lt;/td&gt;
&lt;td&gt;Aware that AI genuinely lowered the vendor's delivery cost; expects some pass-through&lt;/td&gt;
&lt;td&gt;A vendor with a metered, transparent cost basis wins credibility; one hiding behind an unchanged invoice loses it fast&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI platform vendor&lt;/td&gt;
&lt;td&gt;Owns the underlying capability; increasingly prices on usage rather than flat seat fees&lt;/td&gt;
&lt;td&gt;Whatever margin the platform vendor's own licensing terms capture directly. This claim is largely settled by contract, not negotiation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Systems integrator&lt;/td&gt;
&lt;td&gt;Provided judgment, integration, delivery risk absorption, and a working result the client couldn't have gotten by licensing the AI tool directly&lt;/td&gt;
&lt;td&gt;A quantified account of that judgment, specific defect rates avoided, integration complexity resolved, risk absorbed, wins. A vague "we add value" claim loses&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Recommendation:&lt;/strong&gt; don't walk into a renewal conversation assuming your margin is safe because it always has been. Assume the client already has a rough estimate of your AI-driven cost reduction, and come with a specific, demonstrable account of what you add beyond the tool itself, or expect to get renegotiated into the consumption bucket.&lt;/p&gt;

&lt;h2&gt;
  
  
  Defending Margin With a Real Number Instead of a Vague Claim
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Quantify your own AI-driven cost reduction honestly, before the client does it for you.&lt;/strong&gt; If you know the number, you control how it's framed. If the client calculates it independently and you're caught unprepared, you're negotiating from a defensive position for the rest of the relationship.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Separate "what the AI platform did" from "what we did with it," in terms specific enough to price.&lt;/strong&gt; Defect rates avoided, integration issues resolved that the AI output alone would have missed, delivery risk actually absorbed: name the specific contribution, not a general appeal to expertise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Offer the renegotiation before the client demands it.&lt;/strong&gt; SIs that proactively bring a fair, transparent pricing update to a renewal conversation are keeping far more of their margin than SIs that wait to be confronted, because proactive framing reads as partnership and reactive framing reads as having been caught.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Know which of your engagements are genuinely replace-and-build candidates, and don't fight that battle.&lt;/strong&gt; Some capabilities really have become cheap enough for a client to bring in-house. Contesting that reality burns credibility you'll need for the engagements where your judgment genuinely isn't replaceable.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Questions to Ask Your Team
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Do we know our own AI-driven delivery cost reduction, on our top client relationships, well enough to defend a number if the client brings one first?&lt;/li&gt;
&lt;li&gt;Can we name, specifically, what we add beyond the AI platform's raw output on our current engagements, or would we be making a vague appeal to "expertise" in a real renegotiation?&lt;/li&gt;
&lt;li&gt;Are we proactively bringing pricing conversations to clients as AI changes our cost base, or waiting for a CFO to ask the question first?&lt;/li&gt;
&lt;li&gt;Which of our current engagements are honest replace-and-build candidates for the client, and are we prepared for that conversation instead of resisting it?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;There's no default answer to who keeps the margin AI creates. It depends on who shows up to the renegotiation with a real number. Clients who understand the underlying cost shift will ask for a share of it, and they're not wrong to. AI platform vendors are already capturing their share through licensing terms most SIs don't control. What's actually contestable is the SI's own piece, and 93% of sellers walking into that conversation without a quantified value story is exactly why so much margin is likely to get renegotiated away from firms that could have kept more of it. The ones that hold their ground will be the ones treating their judgment as a line item with a number attached, not an assumption everyone keeps taking on faith.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.baytechconsulting.com/blog/saas-pricing-shift-negotiate-ai-driven-renewals" rel="noopener noreferrer"&gt;BayTech Consulting: SaaS Pricing Shift, How to Negotiate AI-Driven Renewals&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.hirefraction.com/blog/ai-is-killing-saas-margins-outcome-based-pricing-is-how-you-get-them-back/" rel="noopener noreferrer"&gt;Hire Fraction: AI Is Killing SaaS Margins. Outcome-Based Pricing Is How You Get Them Back&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.manufacturingtomorrow.com/article/2026/06/how-manufacturers-and-system-integrators-can-build-pricing-power-in-a-commodity-market/27634" rel="noopener noreferrer"&gt;Manufacturing Tomorrow: How Manufacturers and System Integrators Can Build Pricing Power in a Commodity Market&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;If this helped, a like and a follow are appreciated — and if you've solved this differently, drop a comment, I'd like to hear it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Bry Writes Code; cloud and AI infrastructure specialist. Heading into a renewal conversation without a defensible value number? &lt;a href="mailto:brywritescode@gmail.com"&gt;Let's talk&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>business</category>
      <category>ai</category>
      <category>consulting</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Amazon SQS vs SNS: Queues, Fan-Out, and Picking the Right One from the CLI</title>
      <dc:creator>Bry</dc:creator>
      <pubDate>Tue, 08 Sep 2026 15:00:00 +0000</pubDate>
      <link>https://dev.to/brywritescode/amazon-sqs-vs-sns-queues-fan-out-and-picking-the-right-one-from-the-cli-3f6m</link>
      <guid>https://dev.to/brywritescode/amazon-sqs-vs-sns-queues-fan-out-and-picking-the-right-one-from-the-cli-3f6m</guid>
      <description>&lt;h2&gt;
  
  
  Key Points
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;SQS is a queue — one message, delivered to one consumer, held until processed. SNS is pub/sub — one message, fanned out to every subscriber. The fan-out pattern combining both is the standard, not an either/or choice.&lt;/li&gt;
&lt;li&gt;SNS-to-SQS delivery requires the queue's access policy to explicitly allow the topic to call &lt;code&gt;SendMessage&lt;/code&gt; — skip this and messages vanish with no error on the publish side.&lt;/li&gt;
&lt;li&gt;FIFO queues cap at 3,000 messages/second batched (300/s unbatched) unless you enable high-throughput mode, which raises that to 70,000/s. Both SQS and SNS also bill in 64 KB payload chunks, same mechanic as EventBridge.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;CLI/SDK version tested against: &lt;code&gt;aws-cli/2.35.x&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;An IAM role with &lt;code&gt;sqs:*&lt;/code&gt; and &lt;code&gt;sns:*&lt;/code&gt; permissions&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;jq&lt;/code&gt; installed for parsing CLI JSON output in the examples below&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;I still get asked "should I use SQS or SNS?" as though it's a fork in the road. Most of the time the honest answer is both, wired together — SNS decides who gets notified, SQS makes sure each of them actually gets to process the message at their own pace without losing it if they're briefly unavailable. Treating them as competing choices is how teams end up building a fan-out mechanism inside application code that SNS already does for free.&lt;/p&gt;

&lt;p&gt;The part that trips people up isn't the concept. It's the access policy. Subscribe an SQS queue to an SNS topic without granting the topic permission to send to that queue, and the subscription succeeds, the topic publish succeeds, and the message simply never arrives — no error surfaces anywhere in that chain. I've debugged this exact silent failure in a client's fan-out pipeline, and it's the single most common gap in CLI tutorials for this pattern.&lt;/p&gt;

&lt;p&gt;This article builds a standalone queue, a standalone topic, and the full fan-out pattern with the access policy step included, then covers where FIFO changes the throughput math.&lt;/p&gt;




&lt;h2&gt;
  
  
  Queue vs Topic, Conceptually
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqpehoogg5nma3czt5vrc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqpehoogg5nma3czt5vrc.png" alt="Queue vs Topic, Conceptually" width="631" height="864"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Diagram: SQS delivers each message to exactly one consumer; SNS delivers a copy to every subscriber.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A message sitting in an SQS queue waits until something polls for it. A message published to an SNS topic is pushed immediately to every current subscriber — there's no "waiting" concept on the topic itself, which is exactly why SNS alone is a poor fit for a subscriber that might be temporarily down. That's what the fan-out pattern fixes.&lt;/p&gt;




&lt;h2&gt;
  
  
  Standalone SQS
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Standard queue — no ordering guarantee, near-unlimited throughput.&lt;/span&gt;
aws sqs create-queue &lt;span class="nt"&gt;--queue-name&lt;/span&gt; orders-standard-queue

&lt;span class="nv"&gt;QUEUE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;aws sqs get-queue-url &lt;span class="nt"&gt;--queue-name&lt;/span&gt; orders-standard-queue &lt;span class="nt"&gt;--query&lt;/span&gt; QueueUrl &lt;span class="nt"&gt;--output&lt;/span&gt; text&lt;span class="si"&gt;)&lt;/span&gt;

aws sqs send-message &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--queue-url&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$QUEUE_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--message-body&lt;/span&gt; &lt;span class="s1"&gt;'{"orderId":"order-abc123","status":"pending"}'&lt;/span&gt;

aws sqs receive-message &lt;span class="nt"&gt;--queue-url&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$QUEUE_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--max-number-of-messages&lt;/span&gt; 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# FIFO queue — strict ordering within a message group, capped throughput.&lt;/span&gt;
aws sqs create-queue &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--queue-name&lt;/span&gt; orders-fifo-queue.fifo &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--attributes&lt;/span&gt; &lt;span class="s1"&gt;'{"FifoQueue":"true","ContentBasedDeduplication":"true"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Queue type is permanent. There's no &lt;code&gt;update-queue-type&lt;/code&gt; command — you create a new queue and migrate, you don't convert an existing one. Decide standard vs FIFO before you have production traffic depending on the answer.&lt;/p&gt;




&lt;h2&gt;
  
  
  Standalone SNS
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sns create-topic &lt;span class="nt"&gt;--name&lt;/span&gt; orders-topic

&lt;span class="nv"&gt;TOPIC_ARN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;aws sns list-topics &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s2"&gt;"Topics[?contains(TopicArn,'orders-topic')].TopicArn"&lt;/span&gt; &lt;span class="nt"&gt;--output&lt;/span&gt; text&lt;span class="si"&gt;)&lt;/span&gt;

aws sns publish &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--topic-arn&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TOPIC_ARN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--message&lt;/span&gt; &lt;span class="s1"&gt;'{"orderId":"order-abc123","status":"pending"}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--subject&lt;/span&gt; &lt;span class="s2"&gt;"Order Created"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That publish immediately pushes to every current subscriber — email, SMS, HTTPS endpoint, Lambda, or SQS. Nothing is retained after delivery attempts complete. If you need durability — a subscriber that's slow or briefly offline shouldn't lose the message — that's the case for combining SNS with SQS, not using SNS alone.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Fan-Out Pattern — Including the Step Everyone Skips
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F76bu37p489hayovbos3g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F76bu37p489hayovbos3g.png" alt="The Fan-Out Pattern — Including the Step Everyone Skips" width="784" height="289"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Diagram: fan-out delivery depends entirely on each queue's access policy explicitly trusting the topic — there's no other permission gate in this chain.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sqs create-queue &lt;span class="nt"&gt;--queue-name&lt;/span&gt; orders-notifications-queue
&lt;span class="nv"&gt;QUEUE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;aws sqs get-queue-url &lt;span class="nt"&gt;--queue-name&lt;/span&gt; orders-notifications-queue &lt;span class="nt"&gt;--query&lt;/span&gt; QueueUrl &lt;span class="nt"&gt;--output&lt;/span&gt; text&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;QUEUE_ARN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;aws sqs get-queue-attributes &lt;span class="nt"&gt;--queue-url&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$QUEUE_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--attribute-names&lt;/span&gt; QueueArn &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s1"&gt;'Attributes.QueueArn'&lt;/span&gt; &lt;span class="nt"&gt;--output&lt;/span&gt; text&lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="c"&gt;# This is the step that gets skipped. Without it, the subscription&lt;/span&gt;
&lt;span class="c"&gt;# below will succeed and messages will vanish with zero errors.&lt;/span&gt;
aws sqs set-queue-attributes &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--queue-url&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$QUEUE_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--attributes&lt;/span&gt; &lt;span class="s2"&gt;"{&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;Policy&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;{&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;Version&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;2012-10-17&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;Statement&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;:[{&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;Effect&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;Allow&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;Principal&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;:{&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;Service&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;sns.amazonaws.com&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;},&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;Action&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;sqs:SendMessage&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;Resource&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;QUEUE_ARN&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;Condition&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;:{&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;ArnEquals&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;:{&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;aws:SourceArn&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;TOPIC_ARN&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\\\"&lt;/span&gt;&lt;span class="s2"&gt;}}}]}&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;}"&lt;/span&gt;

aws sns subscribe &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--topic-arn&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TOPIC_ARN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--protocol&lt;/span&gt; sqs &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--notification-endpoint&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$QUEUE_ARN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--attributes&lt;/span&gt; &lt;span class="nv"&gt;RawMessageDelivery&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;RawMessageDelivery=true&lt;/code&gt; matters as much as the policy step. Without it, the queue receives SNS's full JSON envelope — &lt;code&gt;Type&lt;/code&gt;, &lt;code&gt;MessageId&lt;/code&gt;, &lt;code&gt;TopicArn&lt;/code&gt;, and the actual message nested inside a &lt;code&gt;Message&lt;/code&gt; string field — instead of your original payload directly. Consumers that expect the raw body will parse the wrong structure and either error out or silently read &lt;code&gt;undefined&lt;/code&gt; fields.&lt;/p&gt;




&lt;h2&gt;
  
  
  FIFO Throughput Limits
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Configuration&lt;/th&gt;
&lt;th&gt;Throughput ceiling&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;FIFO queue, batched&lt;/td&gt;
&lt;td&gt;3,000 messages/second&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FIFO queue, unbatched&lt;/td&gt;
&lt;td&gt;300 messages/second&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FIFO queue, high-throughput mode&lt;/td&gt;
&lt;td&gt;70,000 messages/second&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SNS FIFO topic&lt;/td&gt;
&lt;td&gt;3,000 messages/second or 20 MB/second, whichever is hit first&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard queue/topic&lt;/td&gt;
&lt;td&gt;No published ceiling&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;High-throughput mode isn't automatic — it's an explicit setting (&lt;code&gt;DeduplicationScope&lt;/code&gt; and &lt;code&gt;FifoThroughputLimit&lt;/code&gt; attributes set to &lt;code&gt;messageGroup&lt;/code&gt; and &lt;code&gt;perMessageGroupId&lt;/code&gt; respectively) that trades a small amount of strict cross-group ordering guarantee for the throughput increase. If you need strict global ordering across all message groups, don't enable it. If your ordering requirement is per-customer or per-order (a common case), high-throughput mode is almost always the right call.&lt;/p&gt;




&lt;h2&gt;
  
  
  Common Mistakes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Mistake 1: Subscribing SQS to SNS without a queue access policy&lt;/strong&gt;&lt;br&gt;
The subscription API call succeeds regardless. The failure is invisible until someone notices messages aren't arriving — which could be minutes or weeks later depending on traffic patterns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mistake 2: Forgetting &lt;code&gt;RawMessageDelivery&lt;/code&gt;&lt;/strong&gt;&lt;br&gt;
Consumers built against the raw payload shape break silently or throw confusing parsing errors when they receive SNS's wrapped envelope instead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mistake 3: Choosing FIFO by default "to be safe"&lt;/strong&gt;&lt;br&gt;
FIFO's throughput ceiling and stricter deduplication requirements are a real cost. Most systems don't actually need strict ordering — verify the requirement before paying for it in complexity and throughput headroom.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mistake 4: Trying to convert a standard queue to FIFO&lt;/strong&gt;&lt;br&gt;
There's no such command. You create a new FIFO queue and migrate producers and consumers to it — plan for that as a deploy, not a config change.&lt;/p&gt;




&lt;h2&gt;
  
  
  Production Considerations
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Performance:&lt;/strong&gt; Long polling (&lt;code&gt;--wait-time-seconds 20&lt;/code&gt; on &lt;code&gt;receive-message&lt;/code&gt;) eliminates the empty-receive charges that pile up from aggressive short polling — this alone is one of the two biggest SQS cost levers, the other being batching sends and receives.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Security:&lt;/strong&gt; Scope queue access policies to the specific topic ARN, not a wildcard principal. The AWS Tip source in this article's research log describes a real incident where a queue policy pointed at a stale topic ARN after a topic recreation — silent failure, same root cause as skipping the policy entirely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost:&lt;/strong&gt; Both services bill in 64 KB payload chunks. A consistently large message body (nested JSON, embedded metadata) multiplies your bill the same way it does on EventBridge — trim payloads, pass references where you can.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Monitoring:&lt;/strong&gt; Alarm on &lt;code&gt;ApproximateAgeOfOldestMessage&lt;/code&gt; for SQS queues — a rising value means consumers aren't keeping up, well before the queue depth itself looks alarming.&lt;/p&gt;




&lt;h2&gt;
  
  
  Full Example: Fan-Out Setup Script
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-euo&lt;/span&gt; pipefail

&lt;span class="nv"&gt;TOPIC_NAME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;TOPIC_NAME&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;-topic&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nv"&gt;QUEUE_NAME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;QUEUE_NAME&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;-notifications-queue&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

aws sns create-topic &lt;span class="nt"&gt;--name&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TOPIC_NAME&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;/dev/null
&lt;span class="nv"&gt;TOPIC_ARN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;aws sns list-topics &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s2"&gt;"Topics[?contains(TopicArn,'&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;TOPIC_NAME&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;')].TopicArn"&lt;/span&gt; &lt;span class="nt"&gt;--output&lt;/span&gt; text&lt;span class="si"&gt;)&lt;/span&gt;

aws sqs create-queue &lt;span class="nt"&gt;--queue-name&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$QUEUE_NAME&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;/dev/null
&lt;span class="nv"&gt;QUEUE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;aws sqs get-queue-url &lt;span class="nt"&gt;--queue-name&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$QUEUE_NAME&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--query&lt;/span&gt; QueueUrl &lt;span class="nt"&gt;--output&lt;/span&gt; text&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;QUEUE_ARN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;aws sqs get-queue-attributes &lt;span class="nt"&gt;--queue-url&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$QUEUE_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--attribute-names&lt;/span&gt; QueueArn &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s1"&gt;'Attributes.QueueArn'&lt;/span&gt; &lt;span class="nt"&gt;--output&lt;/span&gt; text&lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="nv"&gt;POLICY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt; &lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt;
{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Principal":{"Service":"sns.amazonaws.com"},"Action":"sqs:SendMessage","Resource":"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;QUEUE_ARN&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;","Condition":{"ArnEquals":{"aws:SourceArn":"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;TOPIC_ARN&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"}}}]}
&lt;/span&gt;&lt;span class="no"&gt;EOF
&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;

aws sqs set-queue-attributes &lt;span class="nt"&gt;--queue-url&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$QUEUE_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--attributes&lt;/span&gt; &lt;span class="s2"&gt;"{&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;Policy&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$POLICY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | jq &lt;span class="nt"&gt;-Rs&lt;/span&gt; .&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;}"&lt;/span&gt;
aws sns subscribe &lt;span class="nt"&gt;--topic-arn&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TOPIC_ARN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--protocol&lt;/span&gt; sqs &lt;span class="nt"&gt;--notification-endpoint&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$QUEUE_ARN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--attributes&lt;/span&gt; &lt;span class="nv"&gt;RawMessageDelivery&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true

echo&lt;/span&gt; &lt;span class="s2"&gt;"Fan-out ready: &lt;/span&gt;&lt;span class="nv"&gt;$TOPIC_NAME&lt;/span&gt;&lt;span class="s2"&gt; -&amp;gt; &lt;/span&gt;&lt;span class="nv"&gt;$QUEUE_NAME&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;Full source including a batched consumer with long polling: &lt;a href="https://github.com/brywritescode/bry-writes-code-examples.git" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; → &lt;code&gt;cloud-apis/amazon-sqs-vs-sns-cli/&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Stop framing SQS and SNS as competing choices — SNS decides who hears about something, SQS makes sure each listener actually gets to act on it without losing the message if they're briefly unavailable. The fan-out pattern combining both is the default for a reason. The one step that will actually cost you debugging time if skipped is the queue access policy: get it wrong and everything upstream reports success while the message goes nowhere.&lt;/p&gt;




&lt;h2&gt;
  
  
  Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://aws.amazon.com/sqs/pricing/" rel="noopener noreferrer"&gt;Amazon SQS Pricing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/sns/latest/dg/subscribe-sqs-queue-to-sns-topic.html" rel="noopener noreferrer"&gt;Subscribing an Amazon SQS queue to an Amazon SNS topic&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/cli/latest/reference/sqs/create-queue.html" rel="noopener noreferrer"&gt;create-queue — AWS CLI Command Reference&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/sns/latest/dg/fifo-topic-code-examples.html" rel="noopener noreferrer"&gt;Amazon SNS code examples for FIFO topics&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;If this helped, a like and a follow are appreciated — and if you've solved this differently, drop a comment, I'd like to hear it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Bry Writes Code — cloud and API infrastructure specialist. Designing a messaging or fan-out architecture on AWS? &lt;a href="mailto:brywritescode@gmail.com"&gt;Get in touch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>The Multi-Tier Subcontracting Pyramid Under Pressure: What Happens to Body-Shop SIers When AI Writes the Code</title>
      <dc:creator>Bry</dc:creator>
      <pubDate>Fri, 04 Sep 2026 13:30:00 +0000</pubDate>
      <link>https://dev.to/brywritescode/the-multi-tier-subcontracting-pyramid-under-pressure-what-happens-to-body-shop-siers-when-ai-1ocp</link>
      <guid>https://dev.to/brywritescode/the-multi-tier-subcontracting-pyramid-under-pressure-what-happens-to-body-shop-siers-when-ai-1ocp</guid>
      <description>&lt;h2&gt;
  
  
  Key Points
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Japan's &lt;em&gt;tajuu-shitauke&lt;/em&gt; (multi-layer subcontracting) structure routes enterprise IT work from a prime contractor (&lt;em&gt;motoke&lt;/em&gt;) down through one or more subcontractor tiers to the engineers who actually write and test the code, a pattern that's persisted, largely unchanged, since the mainframe era.&lt;/li&gt;
&lt;li&gt;Each tier historically takes its margin off the top before passing reduced-price work downward. A client-paid ¥1,000,000 monthly rate can leave the engineer actually doing the work with roughly ¥350,000 to ¥400,000 after cascading through two or three intermediary layers.&lt;/li&gt;
&lt;li&gt;AI-assisted coding tools compress exactly the layer this structure depends on most: routine implementation, testing, and documentation work performed by second- and third-tier subcontractors and individual SES engineers.&lt;/li&gt;
&lt;li&gt;The pyramid's top (client relationships, architecture, governance) looks likely to hold up. Its middle and bottom, firms and engineers whose entire value proposition is executing defined tasks at a lower price than the tier above, are facing the sharpest margin compression in the industry.&lt;/li&gt;
&lt;li&gt;Japan's 2026 Subcontract Act reform, which expanded transparency requirements to creative and consulting services, gives lower-tier firms a real legal lever to push back on unfair pass-through pricing right as AI is already reshaping what "fair" pricing even means.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;A subcontractor relationship I've watched come under real strain recently is a fairly typical version of what's starting to happen across the industry. This is a second-tier firm that built a stable, unglamorous business supplying detailed-design and coding-and-testing engineers to a prime contractor's enterprise projects, at a rate that undercuts the tier above it. That's the entire business model: execute defined tasks, reliably, for less than the next tier up would charge. As the prime contractor's own AI-assisted tooling starts producing that same detailed-design and coding output directly, at a fraction of the cost and turnaround time, there's no lower price the second-tier firm can offer to stay competitive. It isn't underpriced. It's becoming structurally obsolete.&lt;/p&gt;

&lt;p&gt;The pyramid this firm sits inside is one of the most distinctive features of Japan's IT industry. The prime contractor, or &lt;em&gt;motoke&lt;/em&gt;, wins the client relationship and the overall project, then delegates substantial portions of detailed design, implementation, integration testing, and commissioning to one or more subcontractor tiers below it. Prime contractors capture the largest margins, commonly cited around 30-40%. First-tier subcontractors run 20-30%. Second-tier and below drop to 10-15%. Individual engineers, often working under System Engineering Service (SES) staffing arrangements, sit at the bottom, with compensation shrinking at each layer above them. A documented rate cascade shows the mechanism concretely: a client paying roughly ¥1,000,000 a month for an engineer's time can leave that engineer with something in the range of ¥350,000 to ¥400,000 after the intermediary tiers each take their cut.&lt;/p&gt;

&lt;p&gt;This structure has survived for decades because it solved a real coordination problem: large enterprise projects needed a way to scale headcount up and down without every firm in the chain carrying the client relationship or the delivery risk directly. AI-assisted coding isn't attacking that coordination function. It's attacking the thing the lower tiers are actually selling, person-hours of implementation and testing labor, priced progressively lower the further down the pyramid you look. As a prime contractor, or increasingly the client itself, can generate a larger share of that same output directly, the lower tiers' entire value proposition, execute cheaply, stops being a viable business on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pyramid Layer Economics Under AI Compression
&lt;/h2&gt;



&lt;pre data-lang="mermaid"&gt;&lt;code&gt;graph TD
    Client[Client] --&amp;gt;|pays full rate| Prime[Prime Contractor Motoke]
    Prime --&amp;gt;|passes down reduced rate, keeps 30 to 40 percent margin| Tier1[First-Tier Subcontractor]
    Tier1 --&amp;gt;|passes down reduced rate, keeps 20 to 30 percent margin| Tier2[Second-Tier Subcontractor]
    Tier2 --&amp;gt;|passes down reduced rate, keeps 10 to 15 percent margin| SES[Individual Engineer SES]
    AI[AI-Assisted Coding Tooling] -.-&amp;gt;|replaces routine output of| Tier2
    AI -.-&amp;gt;|replaces routine output of| SES&lt;/code&gt;&lt;/pre&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Historical Role&lt;/th&gt;
&lt;th&gt;Historical Margin&lt;/th&gt;
&lt;th&gt;Effect of AI-Assisted Coding&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Client&lt;/td&gt;
&lt;td&gt;Pays full contracted rate&lt;/td&gt;
&lt;td&gt;N/A&lt;/td&gt;
&lt;td&gt;Increasingly aware of the gap between what AI can produce and what's still being billed for&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prime contractor (motoke)&lt;/td&gt;
&lt;td&gt;Client relationship, architecture, governance, overall delivery risk&lt;/td&gt;
&lt;td&gt;~30-40%&lt;/td&gt;
&lt;td&gt;Largely intact so far. Judgment and relationship work AI doesn't replace&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;First-tier subcontractor&lt;/td&gt;
&lt;td&gt;Mid-scale delivery management, some architecture&lt;/td&gt;
&lt;td&gt;~20-30%&lt;/td&gt;
&lt;td&gt;Under pressure, but retains value where it manages integration complexity AI can't own alone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Second-tier subcontractor and below&lt;/td&gt;
&lt;td&gt;Detailed design, coding, testing execution&lt;/td&gt;
&lt;td&gt;~10-15%&lt;/td&gt;
&lt;td&gt;Hardest hit. This is exactly the work AI-assisted tooling compresses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Individual engineer (SES)&lt;/td&gt;
&lt;td&gt;Task-level implementation&lt;/td&gt;
&lt;td&gt;Remainder after cascading cuts&lt;/td&gt;
&lt;td&gt;Displaced where the task is routine; in higher demand where verification and AI-output review are the actual job&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Recommendation:&lt;/strong&gt; if your firm's position in this chain has always been "we execute the same task the tier above us does, for less," that position won't survive AI-assisted coding regardless of how thin you cut your margin further. The layers most likely to survive are the ones pricing judgment, integration complexity, or client trust, not marginal labor cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Likely to Survive, and What a Lower-Tier Firm Should Change Now
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Firms moving up the value chain, not just down on price.&lt;/strong&gt; The subcontractors most likely to survive won't compete on being cheaper than the tier above. They're taking on integration and governance responsibility that used to sit with the prime contractor, becoming harder to replace with AI output alone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Firms specializing in verifying AI-generated output, not just producing more of it.&lt;/strong&gt; Reviewing and validating AI-generated code against a client's actual production constraints is turning out to require exactly the kind of experienced-engineer judgment that routine implementation work didn't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Firms using Japan's 2026 Subcontract Act reform actively.&lt;/strong&gt; The reform's expanded transparency requirements give lower-tier firms real standing to contest unfair pass-through pricing. Firms that understand and use this leverage should renegotiate from a stronger position than firms that don't know the reform applies to them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Firms consolidating instead of competing on the same shrinking margin.&lt;/strong&gt; Some second- and third-tier firms are merging capabilities, combining a compressed coding-and-testing practice with a smaller firm's domain expertise, to offer something closer to the first-tier's integration value than to commodity execution.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Firms that change nothing.&lt;/strong&gt; These are the most exposed to winding down, getting acquired for their client relationships and remaining engineers, or shrinking into a much smaller commodity-execution niche with correspondingly thinner margins.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Questions to Ask Your Team
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Where does our firm actually sit in our clients' delivery chains, and is our value proposition still "we execute this cheaper than the tier above," or has it genuinely moved to something AI-assisted tooling can't produce directly?&lt;/li&gt;
&lt;li&gt;If a prime contractor above us started generating our layer's output directly with AI tooling, what would we still have left to sell them?&lt;/li&gt;
&lt;li&gt;Are we aware of what Japan's 2026 Subcontract Act reform actually changed for transparency and fair pricing in multi-tier arrangements, and have we used it in a renegotiation?&lt;/li&gt;
&lt;li&gt;Have we tested whether our engineers are more valuable reviewing and validating AI-generated output than producing new output themselves, and are we pricing that shift yet?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The multi-tier subcontracting pyramid isn't collapsing, and it's not going to. What it built its lower tiers on, the assumption that there's always a cheaper price for the same routine task one layer further down, is what's breaking. Firms most likely to weather the pressure are moving toward judgment, integration, and verification work. Firms trying to survive by cutting their price further into an already-thin margin are the most exposed. The prime-contractor layer looks likely to hold up fine, because it was never really selling person-hours in the first place. Everyone underneath it needs to figure out, quickly, whether they've been selling person-hours the whole time, and if so, what else they actually have.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.youngju.dev/transcribe/culture/2026-03-19-japan-it-industry-structure-subcontracting.en" rel="noopener noreferrer"&gt;youngju.dev: Japan IT Industry Subcontracting Structure, SIer, SES, and the Multi-Layer Pyramid&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://abe-legal.jp/en/news/subcontract-act-reform-2026" rel="noopener noreferrer"&gt;Abe Legal: Japan Subcontract Act Reform 2026, Expanded Scope to Creative &amp;amp; Consulting Services&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.dualbootpartners.com/insights/the-talent-pyramid/" rel="noopener noreferrer"&gt;Dual Boot Partners: The Talent Pyramid Is Crumbling, Why Traditional IT Services Can't Survive AI&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;If this helped, a like and a follow are appreciated — and if you've solved this differently, drop a comment, I'd like to hear it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Bry Writes Code; cloud and AI infrastructure specialist. Trying to figure out where your firm's real value sits in a compressed delivery chain? &lt;a href="mailto:brywritescode@gmail.com"&gt;Let's talk&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>business</category>
      <category>ai</category>
      <category>consulting</category>
      <category>beginners</category>
    </item>
    <item>
      <title>AI Cost Optimization: A Business Guide to LLM API Budgeting</title>
      <dc:creator>Bry</dc:creator>
      <pubDate>Tue, 01 Sep 2026 15:00:00 +0000</pubDate>
      <link>https://dev.to/brywritescode/ai-cost-optimization-a-business-guide-to-llm-api-budgeting-gdl</link>
      <guid>https://dev.to/brywritescode/ai-cost-optimization-a-business-guide-to-llm-api-budgeting-gdl</guid>
      <description>&lt;p&gt;&lt;strong&gt;Key Points&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;LLM API costs double every six months for most growing organizations — model selection alone can reduce your bill by 50–80% without changing what you build.&lt;/li&gt;
&lt;li&gt;Prompt caching and response caching together eliminate the majority of repeated compute costs, saving 60–90% on input tokens for the right workloads.&lt;/li&gt;
&lt;li&gt;Batch processing cuts per-token prices by 50% for any task that doesn't require a real-time answer.&lt;/li&gt;
&lt;li&gt;Visibility comes first: you cannot optimize spend you cannot see — budget alerts, cost attribution by feature, and a 4-phase adoption roadmap make savings stick.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The AI Bill Nobody Expected
&lt;/h2&gt;

&lt;p&gt;The CFO opens the cloud invoice. Line 37: LLM API usage — $48,000. Last month it was $22,000. Three months ago it was $4,000.&lt;/p&gt;

&lt;p&gt;This is not a hypothetical. Enterprise LLM API spending doubled in under six months as teams moved from experiments to production workloads, and most organizations had no framework in place to manage the acceleration. Unlike compute or storage — where costs scale predictably with users — AI API costs are shaped by choices made at the code level: which model you pick, how you structure requests, whether you cache responses, and whether you batch non-urgent jobs. Those choices are invisible to finance and often undiscussed with engineering leadership.&lt;/p&gt;

&lt;p&gt;I've watched this play out across organizations at different scales — from startups where one engineer's prototype quietly became a $30K/month production workload, to larger teams where cost attribution was a year-long engineering project after the fact. This guide gives business leaders — CTOs, product managers, and anyone signing cloud invoices — a clear, non-technical framework for getting AI spend under control without sacrificing the capabilities your teams depend on.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why AI Costs Spiral
&lt;/h2&gt;

&lt;p&gt;The pattern is consistent across organizations of every size: costs spiral because of three compounding defaults.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Teams default to the most powerful model.&lt;/strong&gt; When engineers prototype an AI feature, they reach for the frontier model — it produces the best output, reduces debugging time, and avoids internal debates about quality. That default rarely gets revisited after launch. The same flagship model handling nuanced legal analysis ends up powering an FAQ bot that answers "What are your business hours?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Nobody tracks token usage by feature.&lt;/strong&gt; Unlike server costs that map neatly to infrastructure, LLM costs are buried in a single API line item. Without attribution — which feature consumed how many tokens — there is no feedback loop. Expensive patterns persist indefinitely because nobody can see them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Users trigger expensive chains without limits.&lt;/strong&gt; In production, one user action can silently trigger a cascade of API calls: an initial query, a summarization step, a classification call, a response synthesis step. Each hop multiplies the token count. Without circuit breakers or cost caps, a single power user can consume more than the entire intended monthly budget in a week. I've seen a five-step document analysis pipeline — perfectly reasonable in isolation — go uncapped in production and rack up more in a weekend than the team had budgeted for the entire month, because one user stress-tested it with 400-page PDFs.&lt;/p&gt;




&lt;h2&gt;
  
  
  How LLM Pricing Works
&lt;/h2&gt;

&lt;p&gt;AI APIs charge by the token — a chunk of text roughly equivalent to 0.75 words in English. A 100-word paragraph is approximately 130 tokens. Two separate charges apply to every request: input tokens (what you send to the model) and output tokens (what the model generates in response).&lt;/p&gt;

&lt;p&gt;Output tokens are significantly more expensive. Across all major providers, output pricing runs 4–5 times higher per token than input pricing. This matters because it means that asking a model to write a long report costs far more than asking it to classify a sentence — even if both requests contain the same amount of input text.&lt;/p&gt;

&lt;p&gt;Here is the current pricing landscape for the major models your engineering team is most likely using:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input $/1M tokens&lt;/th&gt;
&lt;th&gt;Output $/1M tokens&lt;/th&gt;
&lt;th&gt;Context&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 4.8&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;td&gt;$25.00&lt;/td&gt;
&lt;td&gt;200K&lt;/td&gt;
&lt;td&gt;Complex reasoning, multi-step analysis, legal/financial review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;$3.00&lt;/td&gt;
&lt;td&gt;$15.00&lt;/td&gt;
&lt;td&gt;200K&lt;/td&gt;
&lt;td&gt;Balanced quality and cost; most production workloads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;$1.00&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;td&gt;200K&lt;/td&gt;
&lt;td&gt;High-volume, simple tasks: classification, extraction, FAQ&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-4o&lt;/td&gt;
&lt;td&gt;$2.50&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;td&gt;General-purpose; strong tool-use and structured output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Pro&lt;/td&gt;
&lt;td&gt;$1.25&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;Very long documents; high-context summarization&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Pricing as of June 2026. Verify current rates at official provider pricing pages before committing to budget projections.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Context window refers to the maximum amount of text a model can consider in a single request. Larger context windows are critical for processing lengthy contracts, codebases, or conversation histories — but they also increase token costs proportionally.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Biggest Lever: Model Selection
&lt;/h2&gt;

&lt;p&gt;If you take one action from this guide, make it this: match task complexity to model tier.&lt;/p&gt;

&lt;p&gt;The cost gap between tiers is not incremental — it is multiplicative. Using Claude Opus for a simple classification task costs roughly 20 times more per request than using Claude Haiku for the same task. At scale, that ratio becomes a six-figure annual line item for a mid-sized product.&lt;/p&gt;

&lt;p&gt;The business principle is straightforward: the model's capability should match the task's demand. Sending a routine FAQ to your most powerful model is the equivalent of deploying a senior partner to answer reception phone calls.&lt;/p&gt;

&lt;p&gt;Use this decision framework when evaluating which model tier a task belongs in:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Furrje63wqp1xuxhqkrm6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Furrje63wqp1xuxhqkrm6.png" alt="Model Selection" width="800" height="1200"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;My recommendation: start the audit with your highest-volume features, not your most complex ones. FAQ responses, form classification, data extraction, and templated summaries almost always belong in the Haiku tier — and those tend to be your volume leaders. Complex document analysis, strategic synthesis, and nuanced content generation belong in the Sonnet or Opus tier, but they're rarely the source of the runaway bill.&lt;/p&gt;




&lt;h2&gt;
  
  
  Prompt Caching: Pay Once, Reuse Many Times
&lt;/h2&gt;

&lt;p&gt;Every AI application sends instructions to the model with every request. A customer service bot might include a 2,000-word system prompt describing the company's products, policies, and tone guidelines — and that prompt gets sent, and charged, with every single user message.&lt;/p&gt;

&lt;p&gt;Prompt caching solves this. When you send the same instructions repeatedly, the provider stores that content in a temporary cache. Subsequent requests that use the same cached prefix are charged at 10% of the normal input price — a 90% discount on those tokens.&lt;/p&gt;

&lt;p&gt;For applications with long, stable system prompts — support bots, document processors, product assistants — prompt caching typically reduces input token costs by 60–90%. The cached content stays valid for a configurable duration (5 minutes or 1 hour on Anthropic's API) and is automatically refreshed when accessed.&lt;/p&gt;

&lt;p&gt;From a business perspective: if your engineering team is not using prompt caching on any application that includes a system prompt longer than a few hundred words, you are paying full price for tokens you have already paid for. In practice, this is the first thing I check when a team tells me their AI costs feel out of control — it is almost always uncached, and enabling it is usually a one-day engineering task that pays for itself within the first billing cycle.&lt;/p&gt;




&lt;h2&gt;
  
  
  Response Caching: Skip the API Call Entirely
&lt;/h2&gt;

&lt;p&gt;Prompt caching operates at the provider level and reduces the cost of repeated inputs. Response caching operates at your application level and eliminates the API call entirely for repeated questions.&lt;/p&gt;

&lt;p&gt;The concept is simple: when a user asks a question, your application checks a local database before calling the AI API. If the same question has been asked before and the answer is still valid, return the stored answer. No tokens consumed, no API cost, no latency.&lt;/p&gt;

&lt;p&gt;This approach works well for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;FAQ bots:&lt;/strong&gt; The same 200 questions account for 80% of support volume in most businesses&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Product descriptions:&lt;/strong&gt; Thousands of users view the same product content&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Templated reports:&lt;/strong&gt; Weekly summaries generated from the same data schema&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It does not work well for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Real-time data requests:&lt;/strong&gt; Questions about live inventory, current prices, or today's metrics require fresh API calls&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User-specific personalization:&lt;/strong&gt; Responses that incorporate individual user history or preferences cannot be safely reused across users&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The infrastructure investment is modest — a simple key-value store or database table — and the return on high-repetition workloads is significant. I use response caching as the first conversation to have with any team running a support bot or product FAQ: the ROI calculation is fast, the engineering effort is low, and approval is easy when you can show finance that 60% of calls will cost nothing after day three. On FAQ-heavy applications, organizations typically see 40–70% of requests served from cache after the first few days of production traffic.&lt;/p&gt;




&lt;h2&gt;
  
  
  Batch Processing: Trade Speed for 50% Off
&lt;/h2&gt;

&lt;p&gt;Real-time API calls — where your application waits for an immediate response — carry a premium price. For tasks where a response is not needed within seconds, every major provider offers a batch processing API at 50% off standard pricing.&lt;/p&gt;

&lt;p&gt;Anthropic's Batch API and OpenAI's Batch API both accept large volumes of requests submitted at once and return results within hours, typically overnight. The same models, the same quality — at half the cost.&lt;/p&gt;

&lt;p&gt;Batch processing is well-suited for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Overnight report generation&lt;/li&gt;
&lt;li&gt;Bulk document classification or extraction&lt;/li&gt;
&lt;li&gt;Weekly content summarization jobs&lt;/li&gt;
&lt;li&gt;Training data generation or quality review&lt;/li&gt;
&lt;li&gt;Large-scale sentiment analysis&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The key trade-off is latency. Batch jobs are asynchronous — you submit the work and retrieve results later. For any workflow where a user is waiting for a response, batch processing is not appropriate. For workflows that run on a schedule or process accumulated data, it is one of the simplest cost reductions available.&lt;/p&gt;

&lt;p&gt;A team spending $10,000 per month on overnight AI processing jobs that currently use the real-time API can reduce that line item to $5,000 with a single architectural change. If your organization runs any scheduled AI jobs — weekly digests, monthly classification sweeps, overnight data enrichment — batch processing should be a standing agenda item in your next budget review: the savings are predictable, the risk is low, and the approval case is straightforward.&lt;/p&gt;




&lt;h2&gt;
  
  
  Budget Controls and Monitoring
&lt;/h2&gt;

&lt;p&gt;Cost optimization techniques only work if you can see whether they are working. Visibility is the prerequisite for everything else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Set spend alerts in provider dashboards.&lt;/strong&gt; Both the Anthropic Console and the OpenAI usage dashboard allow you to configure email alerts when monthly spend crosses a threshold. Set alerts at 50%, 75%, and 100% of your planned monthly budget — not just at the limit. Early warning gives engineering time to investigate before costs become a crisis.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Implement soft limits in application code.&lt;/strong&gt; Provider-level alerts fire after costs have already accumulated. Application-level limits stop the accumulation. Work with your engineering team to implement per-user, per-feature, and per-workflow token budgets. When a workflow hits its limit, it either degrades gracefully (using a cheaper model) or queues the request for batch processing rather than failing loudly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Track cost per feature, per user, per workflow.&lt;/strong&gt; A single API line item tells you nothing actionable. Attribution — which feature consumed which tokens — is what makes optimization possible. Organizations that implement cost attribution consistently report identifying two or three "cost sink" features that account for the majority of spend, often features that had never been considered high-cost during development.&lt;/p&gt;




&lt;h2&gt;
  
  
  Adoption Roadmap: 4 Phases to Controlled AI Spend
&lt;/h2&gt;

&lt;p&gt;Cost optimization is not a one-time project. It is an ongoing practice. Organizations that achieve and sustain 50–80% cost reductions follow a consistent phased approach.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F85w15bir670g4utyh0zc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F85w15bir670g4utyh0zc.png" alt="Adoption Roadmap" width="799" height="165"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 1 — Measure your baseline.&lt;/strong&gt; Before changing anything, establish what you are spending, by feature and model. This takes two to four weeks and requires engineering effort to add cost attribution to your existing AI calls. Without this data, you are optimizing blind.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 2 — Identify expensive patterns.&lt;/strong&gt; With attribution in place, surface the top cost drivers. Typical findings: a high-volume feature using the wrong model tier, a long system prompt without caching, a batch-eligible workflow running in real-time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 3 — Apply targeted optimizations.&lt;/strong&gt; Address findings in order of impact. Model right-sizing is usually first — it requires minimal engineering effort and delivers immediate, compounding returns. Prompt caching is second. Response caching and batch processing follow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 4 — Monitor and iterate.&lt;/strong&gt; New features introduce new cost patterns. Set a recurring monthly review cadence where engineering and finance align on spend vs. budget and flag new anomalies. The loop between Phase 4 and Phase 2 is what prevents costs from spiraling again after the initial optimization effort.&lt;/p&gt;




&lt;h2&gt;
  
  
  Questions to Ask Your Engineering Team
&lt;/h2&gt;

&lt;p&gt;Before your next budget review or AI project kickoff, bring these questions to your engineering leadership:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Which AI features are using which models?&lt;/strong&gt; Can you show me a list of every production AI feature and the model tier it currently runs on?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Do we have cost attribution?&lt;/strong&gt; Can we see our monthly AI spend broken down by feature, workflow, or user segment — not just as a single total?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Are we using prompt caching on any feature with a long system prompt?&lt;/strong&gt; If a feature sends the same instructions with every request and is not using prompt caching, what is the estimated monthly savings from enabling it?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Which AI workflows run in real-time that could run overnight?&lt;/strong&gt; Is there a list of report generation, bulk processing, or classification jobs that currently use the real-time API but don't need to?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Do we have spend alerts configured?&lt;/strong&gt; At what thresholds do we receive notifications, and who receives them?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Are there application-level rate limits or cost caps per user?&lt;/strong&gt; What prevents a single user or workflow from consuming an outsized share of our monthly token budget?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;When did we last review whether each feature is on the right model tier?&lt;/strong&gt; Has any feature been moved from a flagship model to a lower-cost tier after initial development — or do we still run the same models we used during prototyping?&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;AI API costs are not a fixed overhead — they are a function of architectural decisions your engineering team makes every day. The organizations I've seen get this right share one thing: they treated visibility as non-negotiable from the start, not as a cleanup project after the bill became alarming. Once you can see where the money goes, the optimizations follow naturally — and they tend to be faster and cheaper to implement than anyone expected. The goal is not to spend less on AI. It is to stop funding waste so the budget can go toward the AI capabilities that actually move your business forward.&lt;/p&gt;




&lt;h2&gt;
  
  
  Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;Anthropic API Pricing — Official rates for Claude models&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching" rel="noopener noreferrer"&gt;Anthropic Prompt Caching Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.anthropic.com/en/docs/build-with-claude/batch-processing" rel="noopener noreferrer"&gt;Anthropic Batch API Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openai.com/api/pricing/" rel="noopener noreferrer"&gt;OpenAI API Pricing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;Google Gemini API Pricing&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;If this helped, a like and a follow are appreciated — and if you've solved this differently, drop a comment, I'd like to hear it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Bry Writes Code — cloud and AI infrastructure specialist. Managing AI infrastructure costs? &lt;a href="mailto:brywritescode@gmail.com"&gt;Let's talk&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cloud</category>
      <category>productivity</category>
      <category>programming</category>
    </item>
    <item>
      <title>From Billing Hours to Billing Outcomes: New Contract Structures for AI-Assisted SI Projects</title>
      <dc:creator>Bry</dc:creator>
      <pubDate>Wed, 26 Aug 2026 15:00:00 +0000</pubDate>
      <link>https://dev.to/brywritescode/from-billing-hours-to-billing-outcomes-new-contract-structures-for-ai-assisted-si-projects-2m53</link>
      <guid>https://dev.to/brywritescode/from-billing-hours-to-billing-outcomes-new-contract-structures-for-ai-assisted-si-projects-2m53</guid>
      <description>&lt;h2&gt;
  
  
  Key Points
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Time-and-materials billing gives the vendor no financial incentive to finish faster. Every extra hour is another invoice line. AI-accelerated delivery is making that misalignment obvious to clients who can see the work speeding up while the invoice logic stays the same.&lt;/li&gt;
&lt;li&gt;73% of consulting clients already tell researchers they prefer value-based or outcome-driven pricing over hourly rates. That preference is likely to harden into the default contract structure across most of the systems-integration market within the next several years.&lt;/li&gt;
&lt;li&gt;Outcome-based contracts don't eliminate risk. They move it. Vendors absorb schedule risk on defined deliverables, which is forcing real changes to scoping discipline, delivery governance, and which projects an SI will even accept.&lt;/li&gt;
&lt;li&gt;Time-and-materials isn't disappearing. It survives, correctly, for genuinely exploratory work where the scope can't be defined upfront. The mistake to watch for is applying it to well-defined work out of habit, not because T&amp;amp;M is inherently wrong.&lt;/li&gt;
&lt;li&gt;IDC's forecast, 30% of IT services contracts outcome-based by 2029, is starting to look conservative. If AI-driven delivery speed keeps compounding, outcome-based and hybrid subscription-plus-usage structures could cover a majority of new systems-integration engagements within the decade, in markets where clients have any real pricing leverage.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://medium.com/@brywritescode/why-the-man-month-is-dying-how-ai-broke-it-services-oldest-pricing-unit-d8121156c937" rel="noopener noreferrer"&gt;Why the Man-Month Is Dying: How AI Broke IT Services' Oldest Pricing Unit&lt;/a&gt; in this thread covered why the man-month is dying as a pricing unit. This one covers what's actually starting to get written into contracts once firms stop defaulting to it, because "bill outcomes, not hours" is a slogan, and slogans don't survive contact with a real statement of work.&lt;/p&gt;

&lt;p&gt;The economic case against pure time-and-materials was already documented before AI made it urgent. PMI's Pulse of the Profession data shows T&amp;amp;M engagements running 23% over budget on average, which on a $50,000 project is $11,500 of unplanned client spend, with no penalty to the vendor for the overrun. AI-assisted delivery hasn't fixed that structural misalignment. If anything, it's making it worse in the near term, because a vendor billing by the hour has an active disincentive to let AI tools cut delivery time, and clients increasingly can tell. Futurum Research already found 73% of consulting clients favor value-based or outcome-driven pricing over hourly rates, largely because AI's delivery-speed gains have made time-based billing look indefensible rather than merely inefficient.&lt;/p&gt;

&lt;p&gt;What's replacing it won't be a single template. Outcome-based contracts, payment tied to specific, measurable deliverables with acceptance criteria, are taking over the well-defined end of the market: data migrations, defined integrations, modernization projects with a clear "done" state. Hybrid structures, pairing a base subscription for fixed costs with a usage or outcome-based component above a baseline, are showing up in ongoing managed-service relationships. Time-and-materials should survive where it always made sense: genuinely exploratory work, early-stage product discovery, engagements where the client themselves doesn't yet know the final scope. The mistake to avoid over the next few years isn't choosing outcome-based pricing. It's applying whichever model a firm already knows how to bill, regardless of which one actually fits the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Contract Model Fit by Project Type
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Project Type&lt;/th&gt;
&lt;th&gt;Best-Fit Model&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Defined data migration or system integration&lt;/td&gt;
&lt;td&gt;Outcome-based&lt;/td&gt;
&lt;td&gt;Scope and success criteria are specifiable in advance; AI-accelerated delivery becomes vendor margin, not client discount&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Early-stage product discovery, undefined scope&lt;/td&gt;
&lt;td&gt;Time-and-materials&lt;/td&gt;
&lt;td&gt;Scope genuinely can't be fixed upfront; forcing outcome pricing here just relabels the guesswork&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ongoing managed services / maintenance&lt;/td&gt;
&lt;td&gt;Hybrid subscription + usage&lt;/td&gt;
&lt;td&gt;Base cost stays predictable for the client; usage component tracks real variability instead of headcount&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fixed-scope modernization with a hard deadline&lt;/td&gt;
&lt;td&gt;Fixed-price with milestone acceptance&lt;/td&gt;
&lt;td&gt;Client needs cost certainty; vendor absorbs schedule risk in exchange for a defined, unchanging scope&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Regulatory-driven compliance rebuild&lt;/td&gt;
&lt;td&gt;Outcome-based, tied to audit-passable state&lt;/td&gt;
&lt;td&gt;Client cares about a certifiable end state, not hours logged getting there&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Recommendation:&lt;/strong&gt; match the contract model to how well-defined the scope actually is, not to which model your firm is most comfortable billing. An SI that only offers T&amp;amp;M today is telling well-defined-scope clients, correctly, that it hasn't done the scoping work outcome-based pricing requires.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making the Switch: A Phased Approach
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Build the measurement infrastructure before changing the invoice template.&lt;/strong&gt; Outcome-based billing requires a checkable definition of "done": acceptance criteria, a metering system for usage, or an audit standard. Firms that flip their contract language before building this lose money on undefined "outcomes."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Start with your most well-understood, most frequently repeated project type.&lt;/strong&gt; A data migration you've delivered fifty times is far easier to price on outcome than a novel integration you've never scoped before. Don't lead the transition with your hardest, most novel work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Renegotiate existing T&amp;amp;M relationships in the open, not by stealth.&lt;/strong&gt; Clients who discover a vendor quietly pocketing AI-driven speed gains under an unchanged T&amp;amp;M invoice react far worse than clients told directly that pricing is moving to outcomes and shown why.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep a genuine T&amp;amp;M option for genuinely exploratory engagements.&lt;/strong&gt; Retiring T&amp;amp;M entirely just pushes clients with undefined scope toward vendors willing to be honest that some work can't be priced by outcome yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Revisit pricing on every engagement renewal, not just new business.&lt;/strong&gt; The real lesson behind IDC's forecast isn't the 30% number. It's that firms which wait for contract renewal cycles to force the pricing conversation, rather than initiating it, cede the framing to whichever competitor gets there first.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Questions to Ask Your Team
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;For our current book of business, how many engagements are still billed by the hour purely out of habit, on work that's actually well-defined enough to price on outcome?&lt;/li&gt;
&lt;li&gt;If a client asked us directly why our pricing model doesn't reward them for AI-driven delivery speed, do we have an honest answer, or an evasive one?&lt;/li&gt;
&lt;li&gt;Do we have real acceptance criteria and measurement infrastructure in place before we've committed to an outcome-based number, or are we guessing at "outcomes" the same way we used to guess at hours?&lt;/li&gt;
&lt;li&gt;Are we keeping time-and-materials available for the engagements that genuinely need it, or treating it as a legacy model to be phased out everywhere regardless of fit?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The contract structures most likely to survive the man-month's decline won't be chosen because they sound better in a sales deck. Outcome-based pricing is winning the well-defined end of the market because it aligns vendor incentive with AI-driven speed instead of punishing it. Time-and-materials should survive at the genuinely exploratory end because forcing a false "outcome" definition onto undefined scope just relabels the same uncertainty. The firms most likely to struggle through this transition aren't the ones that pick the wrong model. They're the ones that pick one model and apply it everywhere, regardless of whether the work in front of them actually fits it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.wednesday.is/writing-articles/time-and-materials-vs-fixed-price-vs-outcome-based-contracts" rel="noopener noreferrer"&gt;Wednesday Solutions: Time and Materials vs Fixed Price vs Outcome-Based Contracts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://birdviewpsa.com/blog/outcome-based-contracts/" rel="noopener noreferrer"&gt;Birdview PSA: Outcome-Based Contracts, KPIs, Milestones, and Margin Control&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://hig.com/news/it-services-in-the-age-of-agentic-ai-underwriting-through-a-structural-shift/" rel="noopener noreferrer"&gt;H.I.G. Capital: IT Services in the Age of Agentic AI&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;If this helped, a like and a follow are appreciated — and if you've solved this differently, drop a comment, I'd like to hear it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Bry Writes Code; cloud and AI infrastructure specialist. Deciding which of your engagements are actually ready for outcome-based pricing? &lt;a href="mailto:brywritescode@gmail.com"&gt;Let's talk&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>business</category>
      <category>ai</category>
      <category>consulting</category>
      <category>beginners</category>
    </item>
  </channel>
</rss>
