<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: karmendra pandey</title>
    <description>The latest articles on DEV Community by karmendra pandey (@karmendra_pandey_43ac6983).</description>
    <link>https://dev.to/karmendra_pandey_43ac6983</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3278591%2F16b6feb4-84a4-4586-9576-f2fb003c3a68.jpg</url>
      <title>DEV Community: karmendra pandey</title>
      <link>https://dev.to/karmendra_pandey_43ac6983</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/karmendra_pandey_43ac6983"/>
    <language>en</language>
    <item>
      <title>The Receipt: I Priced Every Token of One AI Agent Task</title>
      <dc:creator>karmendra pandey</dc:creator>
      <pubDate>Tue, 06 Oct 2026 14:16:57 +0000</pubDate>
      <link>https://dev.to/karmendra_pandey_43ac6983/the-receipt-i-priced-every-token-of-one-ai-agent-task-4ibp</link>
      <guid>https://dev.to/karmendra_pandey_43ac6983/the-receipt-i-priced-every-token-of-one-ai-agent-task-4ibp</guid>
      <description>&lt;p&gt;Everyone talks about AI costs in aggregate. Dashboards show monthly spend, per-model totals, team budgets. Nobody shows you a receipt.&lt;/p&gt;

&lt;p&gt;So here's one. A single AI agent task, priced line by line — every token, every tool call, every retry. This is a worked example from production-style workloads I describe in my reference architecture work (full paper: &lt;a href="https://doi.org/10.5281/zenodo.23178788" rel="noopener noreferrer"&gt;https://doi.org/10.5281/zenodo.23178788&lt;/a&gt;). The numbers are representative of a mid-complexity support-triage agent running on AWS Bedrock. Your numbers will differ. The shape of the receipt won't.&lt;/p&gt;

&lt;h2&gt;
  
  
  The task
&lt;/h2&gt;

&lt;p&gt;A customer support triage agent: read an incoming ticket, pull the customer's history, check the knowledge base, decide whether to auto-resolve or escalate, and draft the response. One task, one receipt.&lt;/p&gt;

&lt;h2&gt;
  
  
  The receipt
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Line item&lt;/th&gt;
&lt;th&gt;Tokens in&lt;/th&gt;
&lt;th&gt;Tokens out&lt;/th&gt;
&lt;th&gt;Unit cost basis&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;System prompt (loaded once, cached)&lt;/td&gt;
&lt;td&gt;1,850&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;cached input&lt;/td&gt;
&lt;td&gt;$0.0011&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Ticket text + customer history&lt;/td&gt;
&lt;td&gt;2,300&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;standard input&lt;/td&gt;
&lt;td&gt;$0.0069&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Knowledge base retrieval (3 chunks)&lt;/td&gt;
&lt;td&gt;4,100&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;standard input&lt;/td&gt;
&lt;td&gt;$0.0123&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Reasoning: triage decision&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;420&lt;/td&gt;
&lt;td&gt;standard output&lt;/td&gt;
&lt;td&gt;$0.0050&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Tool call: CRM lookup&lt;/td&gt;
&lt;td&gt;180&lt;/td&gt;
&lt;td&gt;90&lt;/td&gt;
&lt;td&gt;in/out&lt;/td&gt;
&lt;td&gt;$0.0016&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Reasoning: draft response&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;680&lt;/td&gt;
&lt;td&gt;standard output&lt;/td&gt;
&lt;td&gt;$0.0082&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Retry: first draft failed tone eval, regenerated&lt;/td&gt;
&lt;td&gt;6,200&lt;/td&gt;
&lt;td&gt;710&lt;/td&gt;
&lt;td&gt;in (cached) + out&lt;/td&gt;
&lt;td&gt;$0.0121&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Embedding: ticket for memory store&lt;/td&gt;
&lt;td&gt;340&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;embedding&lt;/td&gt;
&lt;td&gt;$0.0000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;Final response + audit log write&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;350&lt;/td&gt;
&lt;td&gt;standard output&lt;/td&gt;
&lt;td&gt;$0.0042&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~15,000&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~2,250&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$0.051&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Wait — $0.051, not $0.23? The $0.23 figure from my paper is the fully-loaded cost: it includes the amortized cost of the eval harness, the retrieval infrastructure, and the human review queue time for escalations. The raw inference receipt is five cents. The &lt;em&gt;governed&lt;/em&gt; cost is twenty-three cents. Both numbers matter, and most teams track neither.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading the receipt
&lt;/h2&gt;

&lt;p&gt;Three things jump out:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The retry cost more than the original attempt.&lt;/strong&gt; Line 7 — a failed tone evaluation forced a regeneration — cost $0.0121, more than lines 4+6 combined. Retries are the silent budget killer in agent systems. Every eval gate you add has a cost; every gate you skip has a bigger one. The receipt lets you price the tradeoff instead of guessing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Context is the bulk of input spend.&lt;/strong&gt; Lines 2+3 are 6,400 input tokens — 43% of the total input. That knowledge-base retrieval is doing real work, but it's also the first place to optimize: better chunking, semantic caching of frequent queries, and prompt compression all attack this line directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The system prompt is nearly free — because it's cached.&lt;/strong&gt; Line 1 costs a tenth of a cent thanks to prompt caching. Without caching, that 1,850-token system prompt gets re-priced on every call in the chain. If your provider supports caching and you're not using it, you're donating money.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the receipt is the unit of cost governance
&lt;/h2&gt;

&lt;p&gt;You can't govern what you can't itemize. Monthly dashboards tell you &lt;em&gt;that&lt;/em&gt; you spent $40K. The receipt tells you &lt;em&gt;why&lt;/em&gt; — which line items to attack, which eval gates earn their keep, which retries are worth preventing.&lt;/p&gt;

&lt;p&gt;Three practices that fall out of this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Emit a receipt per task.&lt;/strong&gt; Log input/output tokens per span — per agent, per tool call, per retry. If your framework doesn't do this natively, add a callback handler.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Budget at the task level, not the month level.&lt;/strong&gt; A per-task budget with a circuit breaker ("halt if projected cost exceeds $X") catches runaway loops in seconds instead of at invoice time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Price your eval gates.&lt;/strong&gt; Every quality check costs tokens. The receipt tells you whether the gate is cheaper than the failure it prevents.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The next time someone asks what your AI agents cost, don't open the dashboard. Show them a receipt.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Karmendra Pandey is a Practice Architect in AI &amp;amp; ML at TEK Systems, building production agentic AI on AWS. His reference architecture for cost-governed AI agents is at &lt;a href="https://doi.org/10.5281/zenodo.23178788" rel="noopener noreferrer"&gt;https://doi.org/10.5281/zenodo.23178788&lt;/a&gt;, and the companion implementation at &lt;a href="https://github.com/karmendra8386/agentic-ai-aws-reference" rel="noopener noreferrer"&gt;https://github.com/karmendra8386/agentic-ai-aws-reference&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>aws</category>
      <category>costoptimization</category>
    </item>
    <item>
      <title>I Audited 5 Agent Frameworks for Cost Visibility. Here's the Scorecard.</title>
      <dc:creator>karmendra pandey</dc:creator>
      <pubDate>Tue, 06 Oct 2026 14:16:15 +0000</pubDate>
      <link>https://dev.to/karmendra_pandey_43ac6983/i-audited-5-agent-frameworks-for-cost-visibility-heres-the-scorecard-4pe4</link>
      <guid>https://dev.to/karmendra_pandey_43ac6983/i-audited-5-agent-frameworks-for-cost-visibility-heres-the-scorecard-4pe4</guid>
      <description>&lt;p&gt;Every agent framework promises observability. Traces, spans, dashboards. But ask a harder question — &lt;em&gt;how much did that agent run cost, and who do I bill?&lt;/em&gt; — and the answers get thin fast.&lt;/p&gt;

&lt;p&gt;So I audited five of the most-used agent frameworks against one rubric: cost visibility. Not general observability — the specific ability to see, attribute, and control spend. I read the official docs for each (sources linked throughout). No guessing, no marketing pages. Here's the scorecard.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rubric
&lt;/h2&gt;

&lt;p&gt;Four questions, scored 1–5:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Does it track token usage natively, per run / per agent / per tool call?&lt;/li&gt;
&lt;li&gt;Does it convert tokens to dollars?&lt;/li&gt;
&lt;li&gt;Can you attribute spend to an agent, session, or user out of the box?&lt;/li&gt;
&lt;li&gt;Does it support budgets or cost guardrails natively?&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The scorecard
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Framework&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;LangGraph + LangSmith&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;The gold standard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Amazon Bedrock AgentCore&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;Zero-wiring cost visibility&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Strands Agents SDK&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;td&gt;Best guardrails, priced in tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CrewAI&lt;/td&gt;
&lt;td&gt;3/5&lt;/td&gt;
&lt;td&gt;Tokens yes, dollars DIY&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AutoGen (v0.4)&lt;/td&gt;
&lt;td&gt;2/5&lt;/td&gt;
&lt;td&gt;Went backwards&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  LangGraph + LangSmith: 5/5
&lt;/h2&gt;

&lt;p&gt;LangSmith does the thing everyone else stops short of: it converts tokens to dollars automatically, using a customizable per-model pricing table (docs: &lt;a href="https://docs.langchain.com/langsmith/cost-tracking" rel="noopener noreferrer"&gt;https://docs.langchain.com/langsmith/cost-tracking&lt;/a&gt;). Custom or non-linear pricing goes through &lt;code&gt;usage_metadata&lt;/code&gt;. The trace tree breaks cost down per node, so you can separate what the supervisor spent from what the specialist agents spent. Metadata-based attribution covers users and sessions, and there are dashboards on top.&lt;/p&gt;

&lt;p&gt;The blind spot: no native budget enforcement. You can &lt;em&gt;see&lt;/em&gt; every dollar with precision — you just can't make the framework &lt;em&gt;stop&lt;/em&gt; spending them. Visibility without teeth.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bedrock AgentCore: 5/5
&lt;/h2&gt;

&lt;p&gt;The managed CloudWatch GenAI dashboard shows per-invocation token usage &lt;strong&gt;and cost&lt;/strong&gt; with zero wiring — no callbacks, no exporters, no pricing tables to maintain. Per-user traceability comes through span attributes, and tenant attribution via OpenTelemetry baggage.&lt;/p&gt;

&lt;p&gt;The blind spot: spend control is external. It's CloudWatch alarms plus AWS Budgets — both notify-only. The platform will tell you, beautifully, exactly how much you overspent. It won't pull the plug.&lt;/p&gt;

&lt;h2&gt;
  
  
  Strands Agents SDK: 4/5
&lt;/h2&gt;

&lt;p&gt;Every &lt;code&gt;AgentResult&lt;/code&gt; carries automatic per-invocation token metrics — input, output, total, cache tokens, plus per-tool stats (docs: &lt;a href="https://strandsagents.com/docs/user-guide/sdk/observability-evaluation/metrics/" rel="noopener noreferrer"&gt;https://strandsagents.com/docs/user-guide/sdk/observability-evaluation/metrics/&lt;/a&gt;). And Strands is the &lt;strong&gt;only&lt;/strong&gt; framework in this group with native budget caps: &lt;code&gt;limits={"turns", "output_tokens", "total_tokens"}&lt;/code&gt; halts the agent loop with explicit stop reasons.&lt;/p&gt;

&lt;p&gt;The blind spot: the caps are token-denominated, and there's no dollar conversion. A 10,000-token cap means wildly different dollars on Haiku vs. Opus. It's the only framework with a native kill-switch — it just thinks in the wrong currency.&lt;/p&gt;

&lt;h2&gt;
  
  
  CrewAI: 3/5
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;token_usage&lt;/code&gt; on every &lt;code&gt;kickoff()&lt;/code&gt; result is genuinely zero-setup, and per-task usage rides along on &lt;code&gt;TaskOutput&lt;/code&gt;. The event bus plus OpenTelemetry export gives you the plumbing to build whatever you want (docs: &lt;a href="https://docs.crewai.com/v1.15.12/en/observability/overview" rel="noopener noreferrer"&gt;https://docs.crewai.com/v1.15.12/en/observability/overview&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The blind spot: the framework stops at tokens. No pricing, no dashboard, no budgets. Dollar cost is a DIY project, and per-agent attribution needs custom event listeners. You get the raw material of cost visibility and none of the product.&lt;/p&gt;

&lt;h2&gt;
  
  
  AutoGen (v0.4): 2/5
&lt;/h2&gt;

&lt;p&gt;Here's the irony: AutoGen &lt;strong&gt;had&lt;/strong&gt; built-in cost logging. Version 0.2 shipped &lt;code&gt;start_logging&lt;/code&gt; / &lt;code&gt;print_usage_summary&lt;/code&gt; with cost computation. Then came the v0.4 rewrite — and cost visibility went backwards. The current telemetry story is native OpenTelemetry with GenAI semantic-convention spans (docs: &lt;a href="https://microsoft.github.io/autogen/stable/user-guide/core-user-guide/framework/telemetry.html):" rel="noopener noreferrer"&gt;https://microsoft.github.io/autogen/stable/user-guide/core-user-guide/framework/telemetry.html):&lt;/a&gt; token data exists as span attributes for &lt;em&gt;your&lt;/em&gt; backend to aggregate. No first-party cost computation, no dashboard, no budgets.&lt;/p&gt;

&lt;p&gt;The blind spot: everything. The most popular multi-agent framework in the ecosystem currently treats cost as somebody else's problem — and it used to be better at this.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three patterns
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. The industry stops at tokens.&lt;/strong&gt; Only LangSmith and AgentCore convert to dollars natively. Everyone else hands you token counts and wishes you luck with the pricing page. Tokens are an implementation detail; dollars are the budget. The gap between them is where finance teams lose trust in AI projects.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Nobody ships a dollar kill-switch.&lt;/strong&gt; Strands has token caps. AutoGen has a token-limited context. Everything else is alarms and custom logic. As of today, no major framework lets you say "halt this agent if it spends more than $2" — in dollars, natively, enforced. That missing primitive is the single biggest gap in agent cost governance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Visibility went backwards on at least one framework.&lt;/strong&gt; AutoGen's regression from v0.2 to v0.4 is a cautionary tale: cost features are the first thing cut in a rewrite because they're nobody's launch-blocking requirement — until the invoice arrives.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do about it
&lt;/h2&gt;

&lt;p&gt;If you're picking a framework this quarter: LangSmith or AgentCore if you want cost visibility today; Strands if you want guardrails and can live with token-denominated caps. If you're already on CrewAI or AutoGen, budget a sprint for the DIY layer — token export, a pricing table, per-agent attribution, and an external budget alarm — because the framework isn't going to do it for you.&lt;/p&gt;

&lt;p&gt;And framework maintainers: the tokens-to-dollars gap is the most valuable ten lines of code in your roadmap. Ship the pricing table. Ship the dollar cap. Your users' CFOs will thank you.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Methodology: scored against official documentation as of October 2026. Karmendra Pandey is a Practice Architect in AI &amp;amp; ML at TEK Systems, working on cost-governed AI agents on AWS. Reference architecture: &lt;a href="https://doi.org/10.5281/zenodo.23178788" rel="noopener noreferrer"&gt;https://doi.org/10.5281/zenodo.23178788&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>costoptimization</category>
    </item>
    <item>
      <title>How to Audit an ML Design Before It Ships</title>
      <dc:creator>karmendra pandey</dc:creator>
      <pubDate>Tue, 06 Oct 2026 05:45:52 +0000</pubDate>
      <link>https://dev.to/karmendra_pandey_43ac6983/how-to-audit-an-ml-design-before-it-ships-1f95</link>
      <guid>https://dev.to/karmendra_pandey_43ac6983/how-to-audit-an-ml-design-before-it-ships-1f95</guid>
      <description>&lt;h1&gt;
  
  
  How to Audit an ML Design Before It Ships
&lt;/h1&gt;

&lt;p&gt;Most ML failures are not model failures. They are &lt;em&gt;design&lt;/em&gt; failures — and they are almost always visible before launch to anyone who knows where to look.&lt;/p&gt;

&lt;p&gt;Training-serving skew nobody checked. An evaluation set that doesn't match production. A feedback loop that will quietly poison the next retraining run. Cost and latency budgets nobody wrote down. I've watched these kill systems that had perfectly good models. After 19+ years of shipping software and reviewing AI research, I now run every ML design through the same five-gate audit before it ships. Here it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate 1: Data — provenance, leakage, drift
&lt;/h2&gt;

&lt;p&gt;Ask three questions. Where did every training example come from, and can you prove it? Could any information from the future — or from the test set — have leaked into training? And what happens when the real world drifts away from your training distribution?&lt;/p&gt;

&lt;p&gt;The classic killer here is eval leakage: the "99% accurate" classifier that fails on its first real day because the eval set was accidentally drawn from training data. It happens more than anyone admits. Demand a written data provenance statement and a leakage check as merge requirements, not nice-to-haves.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate 2: Evaluation — does the metric match the mission?
&lt;/h2&gt;

&lt;p&gt;A model can ace its metric and still fail its job. Offline accuracy means nothing if the production decision has different costs for different errors — a fraud model optimized for accuracy will happily approve everything when fraud is rare.&lt;/p&gt;

&lt;p&gt;For every model, write down: what decision does this model actually drive, what does a mistake cost in each direction, and does our eval set look like production traffic? If the answers are vague, the design isn't ready.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate 3: Operations — latency, cost, fallback
&lt;/h2&gt;

&lt;p&gt;Nobody writes down the latency budget until the first user complaint. Nobody prices the prediction until the first cloud bill. For each model, record: p95 latency budget, cost per 1,000 predictions at expected volume, and what happens when the model is down or too slow. "The page breaks" is not a fallback strategy — cache the last good prediction, degrade to a simpler model, or fail open explicitly and loudly.&lt;/p&gt;

&lt;p&gt;For agentic systems this gate matters 10x more: every unnecessary tool call burns tokens &lt;em&gt;and&lt;/em&gt; inflates every downstream context window. Cost discipline starts with call discipline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate 4: Security — adversarial inputs and access control
&lt;/h2&gt;

&lt;p&gt;Who can feed inputs to this model, and what can they make it do? Cover the basics: input validation, rate limiting, access controls on the model endpoint and its training data. For LLM-backed systems, add prompt-injection review — treat every tool the agent can call as a privilege, and scope it like one. I model agents as IAM principals: unique identity, least-privilege roles, session-scoped credentials.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate 5: Accountability — who owns the decision?
&lt;/h2&gt;

&lt;p&gt;When the model is wrong — and it will be — who is accountable, and what is the remediation path? Every production ML system needs a named owner, a rollback plan, and a written policy for the failure modes you &lt;em&gt;know&lt;/em&gt; exist. "The model decided" is not an accountability structure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-page audit
&lt;/h2&gt;

&lt;p&gt;Turn these gates into a single page your team fills out at design review:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Data:&lt;/strong&gt; provenance documented? leakage check passed? drift monitors planned?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evaluation:&lt;/strong&gt; metric matches the mission? eval set production-representative?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Operations:&lt;/strong&gt; latency budget, cost per prediction, fallback behavior defined?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security:&lt;/strong&gt; inputs validated? access scoped? agent tools least-privilege?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accountability:&lt;/strong&gt; named owner? rollback plan? known failure modes written down?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Any blank answer blocks the ship. It takes an hour, and it catches the failures that postmortems always describe as "obvious in hindsight."&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Karmendra Pandey is a Practice Architect in AI &amp;amp; ML at TEKsystems. He builds production agentic AI systems on AWS and peer-reviews AI research on PREreview.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>mlops</category>
      <category>datascience</category>
    </item>
    <item>
      <title>I Found a Cost Blind Spot in an Open-Source Agent Framework</title>
      <dc:creator>karmendra pandey</dc:creator>
      <pubDate>Tue, 06 Oct 2026 05:20:49 +0000</pubDate>
      <link>https://dev.to/karmendra_pandey_43ac6983/i-found-a-cost-blind-spot-in-an-open-source-agent-framework-34b9</link>
      <guid>https://dev.to/karmendra_pandey_43ac6983/i-found-a-cost-blind-spot-in-an-open-source-agent-framework-34b9</guid>
      <description>&lt;h1&gt;
  
  
  I Found a Cost Blind Spot in an Open-Source Agent Framework
&lt;/h1&gt;

&lt;p&gt;I was building cost attribution for AI agents on AWS when I ran into a gap in the open-source Strands Agents SDK: the framework knew how to talk to Amazon Bedrock models, but it didn't know which &lt;em&gt;family&lt;/em&gt; some of them belonged to. That sounds like trivia. It turned out to be a cost-governance blind spot — and fixing it taught me something about how open-source contributions actually happen.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;The Strands harness SDK keeps a registry of model families for Bedrock — it uses the family to make decisions about capabilities and pricing behavior. But several Bedrock model families were missing from the registry. When your agent ran on one of those models, the framework couldn't classify it. Depending on the code path, that meant falling back to defaults that didn't match the model's real cost or capability profile.&lt;/p&gt;

&lt;p&gt;In a cost-governed setup, "unknown model family" is not a minor metadata gap. If you can't identify the model, you can't attribute its cost correctly, you can't route to or away from it intelligently, and your burn-rate dashboard has a hole in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finding it
&lt;/h2&gt;

&lt;p&gt;I found it the unglamorous way: reading the source. I was tracing how the SDK resolved model metadata for a Bedrock deployment and noticed the family lookup silently returning nothing for models I knew existed. No error, no warning — just a quiet gap. That's the worst kind of bug in cost infrastructure: it doesn't fail, it just misleads.&lt;/p&gt;

&lt;p&gt;I filed it as issue #4852 in the strands-agents/harness-sdk repository, documenting exactly which families were missing and what the downstream effects were.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;The fix itself was small — extending the family registry to cover the missing Bedrock families, with tests proving each new mapping resolved correctly. I added four new tests, ran the full suite (152 passed), and confirmed the 13 pre-existing failures also failed on clean main, so they weren't mine.&lt;/p&gt;

&lt;p&gt;Small fix, but the verification mattered more than the code. In cost infrastructure, a wrong mapping is worse than a missing one — it makes your dashboards lie with confidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;p&gt;Three things worth sharing:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Read the source before the docs.&lt;/strong&gt; The docs described the registry as complete. The source said otherwise. For cost-critical paths, the source is the only documentation you can trust.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;File the issue before the PR.&lt;/strong&gt; Writing up #4852 first forced me to articulate the &lt;em&gt;impact&lt;/em&gt;, not just the bug. That made the fix obviously correct to reviewers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test the mapping, not just the code.&lt;/strong&gt; The interesting assertions weren't "does it run" but "does this model resolve to the right family" — the semantic correctness that cost attribution depends on.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Why this matters beyond one PR
&lt;/h2&gt;

&lt;p&gt;Every team running agents in production is building cost governance on top of frameworks like this one. The frameworks are young, the model landscape changes monthly, and gaps like this one are everywhere — quiet, unremarkable, and expensive in aggregate.&lt;/p&gt;

&lt;p&gt;If you're running agents on Bedrock, check your model metadata paths. The bug you find might be as boring as a missing registry entry. Fix it anyway. Your burn-rate dashboard will thank you.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The fix is PR #4853 in strands-agents/harness-sdk. Karmendra Pandey is a Practice Architect in AI &amp;amp; ML at TEKsystems.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>ai</category>
      <category>aws</category>
      <category>programming</category>
    </item>
    <item>
      <title>Your AI Agent Has a Burn Rate: Governing the Cost of Production AI Agents on AWS</title>
      <dc:creator>karmendra pandey</dc:creator>
      <pubDate>Tue, 06 Oct 2026 05:19:55 +0000</pubDate>
      <link>https://dev.to/karmendra_pandey_43ac6983/your-ai-agent-has-a-burn-rate-governing-the-cost-of-production-ai-agents-on-aws-1jil</link>
      <guid>https://dev.to/karmendra_pandey_43ac6983/your-ai-agent-has-a-burn-rate-governing-the-cost-of-production-ai-agents-on-aws-1jil</guid>
      <description>&lt;h1&gt;
  
  
  Your AI Agent Has a Burn Rate: Governing the Cost of Production AI Agents on AWS
&lt;/h1&gt;

&lt;p&gt;Every AI pilot I have reviewed in the last two years had the same surprise waiting in production: the invoice. A demo that cost $40 to run suddenly costs $4,000 a month at real traffic. Nobody planned for it because nobody was measuring it. Your AI agent has a burn rate — cost per task, per user, per day — and if you are not governing it, you are flying blind.&lt;/p&gt;

&lt;p&gt;This is the practitioner's guide to cost-governed agents on AWS: what to measure, where the money actually goes, and the architecture patterns that keep the burn rate under control.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the money goes
&lt;/h2&gt;

&lt;p&gt;For a typical agent built on Amazon Bedrock, the cost stack looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Model inference&lt;/strong&gt; — the big one. Every token in and out is metered. Long system prompts, verbose reasoning traces, and multi-step agent loops multiply this fast.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval&lt;/strong&gt; — OpenSearch or Kendra queries per agent step. Cheap per call, expensive at scale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool calls&lt;/strong&gt; — Lambda invocations, API calls, data transfers. Each agent step can trigger several.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retries and loops&lt;/strong&gt; — the silent killer. An agent that retries a failing step five times just quintupled that task's cost.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Most teams only see line 1 on the bill. The real burn rate is the sum of all four, attributed per task.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pattern 1: Attribute cost per task
&lt;/h2&gt;

&lt;p&gt;You cannot govern what you cannot attribute. Every agent execution should log: input tokens, output tokens, model used, tool calls made, and wall-clock time — tagged with a task ID and a user or tenant ID. On AWS, this means structured logging from your Bedrock invocations into CloudWatch, with a pipeline that rolls it up per task.&lt;/p&gt;

&lt;p&gt;The moment you can say "this customer support task cost $0.42 and this one cost $3.80," you can start asking why — and that question is where the savings live.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pattern 2: Route, don't default
&lt;/h2&gt;

&lt;p&gt;Not every task needs the flagship model. A classification step, a format check, a simple extraction — these run fine on smaller, cheaper models at a tenth of the cost. The expensive model should be reserved for the steps that actually need judgment.&lt;/p&gt;

&lt;p&gt;This is the routing pattern: a lightweight decision at each step picks the cheapest model capable of handling it. In practice, teams that add routing cut inference spend 40–70% with no measurable quality drop — because most agent steps were never hard enough to justify the big model in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pattern 3: Cap the loop
&lt;/h2&gt;

&lt;p&gt;Agents loop. Sometimes they loop forever. Every production agent needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A max-iteration cap&lt;/strong&gt; per task (if it hasn't converged in N steps, it won't).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A cost ceiling per task&lt;/strong&gt; — abort or escalate to a human when the spend crosses the line.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Idempotent tool calls&lt;/strong&gt; — so a retry doesn't double-charge a payment API or re-send a customer email.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are not optimizations. They are the guardrails that keep one bad task from burning a day's budget.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pattern 4: Cache the predictable
&lt;/h2&gt;

&lt;p&gt;A surprising share of agent work is repetitive: the same policy lookup, the same product description, the same classification of the same ticket type. Cache retrieval results and deterministic step outputs aggressively. Bedrock's prompt caching and a simple Redis layer in front of your retrieval calls compound into real savings.&lt;/p&gt;

&lt;h2&gt;
  
  
  The burn-rate dashboard
&lt;/h2&gt;

&lt;p&gt;Put it together and you get one view that matters: cost per task, trended over time, broken down by model, step type, and tenant. When a new model version ships or a prompt changes, the dashboard tells you within hours whether the burn rate moved — before the invoice does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start this week
&lt;/h2&gt;

&lt;p&gt;You don't need a platform rebuild. Pick one agent, add per-task cost logging, and look at the numbers for a week. The waste will introduce itself. Then route one step to a smaller model, cap one loop, and cache one retrieval call. Measure again.&lt;/p&gt;

&lt;p&gt;Cost governance isn't about spending less on AI. It's about knowing what each dollar bought — and making sure the expensive dollars went to the steps that deserved them.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Karmendra Pandey is a Practice Architect in AI &amp;amp; ML at TEKsystems, where he designs production AI agent systems on AWS.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>aws</category>
      <category>devops</category>
    </item>
    <item>
      <title>How AI is changing data processing and insight</title>
      <dc:creator>karmendra pandey</dc:creator>
      <pubDate>Fri, 20 Jun 2025 05:22:33 +0000</pubDate>
      <link>https://dev.to/karmendra_pandey_43ac6983/how-ai-is-changing-data-processing-and-insight-413h</link>
      <guid>https://dev.to/karmendra_pandey_43ac6983/how-ai-is-changing-data-processing-and-insight-413h</guid>
      <description>&lt;p&gt;🔄 &lt;strong&gt;From Static Pipelines to Adaptive Intelligence&lt;/strong&gt;&lt;br&gt;
Traditional ETL pipelines were rigid: predefined schemas, manual transformations, static schedules. Now, AI introduces:&lt;br&gt;
• Auto-adjusting workflows: Pipelines adapt based on schema drift, error rates, or downstream demand.&lt;br&gt;
• Semantic understanding: LLMs can interpret column meanings, detect sensitive data, or auto-document datasets.&lt;br&gt;
It’s not just about moving data anymore—it’s about understanding it as it moves.&lt;/p&gt;

&lt;p&gt;⚙️ &lt;strong&gt;Processing Power Meets Prediction&lt;/strong&gt;&lt;br&gt;
AI accelerates and enriches processing by:&lt;br&gt;
• Pattern recognition: Detecting anomalies in real-time log streams or IoT sensor feeds.&lt;br&gt;
• Entity extraction: Parsing legal contracts, medical records, or video transcripts at scale (think: media analytics + Rekognition).&lt;br&gt;
• Data summarization: Quickly distilling terabytes into consumable formats—crucial for analysts, execs, or even fine-tuned dashboards.&lt;/p&gt;

&lt;p&gt;🔍 &lt;strong&gt;Insight Generation Becomes Conversational&lt;/strong&gt;&lt;br&gt;
No more staring at dashboards hoping the insights jump out. AI enables:&lt;br&gt;
• Natural language Q&amp;amp;A over your data (via Bedrock agents or BI copilots).&lt;br&gt;
• Automated storytelling: Reports that highlight why KPIs changed, not just what changed.&lt;br&gt;
• Forecasting &amp;amp; simulation: What-if models baked into insight tools, not buried in notebooks.&lt;br&gt;
It’s the difference between “What are the numbers?” and “What do I do next?”&lt;/p&gt;

&lt;p&gt;🔁 &lt;strong&gt;Continuous Learning &amp;amp; Feedback&lt;/strong&gt;&lt;br&gt;
AI-driven systems evolve:&lt;br&gt;
• User behavior fine-tunes what’s shown or flagged.&lt;br&gt;
• Drift detection retriggers retraining or pipeline updates.&lt;br&gt;
• Closed-loop systems (especially in DevOps or media pipelines) correct issues without manual intervention.&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
