DEV Community

Bims-creator
Bims-creator

Posted on

When the linter and AWS disagree, an agent that shows you both sides

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

I maintain iam-lint, a small static analyzer for AWS IAM policies. It's good at catching obvious over-permissioning — wildcard actions, wildcard resources, unrestricted PassRole — but like any linter, it's blunt. It reads a JSON policy document and nothing else, so it can't tell a genuine misconfiguration from a pattern AWS itself documents as intentional. ec2:DescribeInstances gets flagged for Resource: "*" even though that action has no other option. An MFA-condition rule fires on a Lambda execution role even though a service identity has no way to present MFA in the first place.

So I built an agent that sits on top of the linter's output and actually reasons about whether a flagged pattern is a real risk. It pairs each of five iam-lint rules against what AWS's own documentation and the CIS AWS Foundations Benchmark say about the same pattern, and four of the five pairs genuinely disagree with the linter's blanket severity. The agent doesn't pick a side — it walks through the disagreement and tells you which context factors resolve it.

Demo

This is a CLI tool, not a hosted app, so here it is running against four questions, one for each place the linter and AWS's guidance part ways:

Query: Is a policy with Resource: '*' on ec2:DescribeInstances a real finding or a false positive?

Short answer: false positive — and this is the KB's canonical example of one. iam-lint's WILDCARD_RESOURCE rule is explicitly action-blind — it doesn't check whether the action even supports resource-level permissions. AWS's Service Authorization Reference names ec2:DescribeInstances directly as an action where Resource: "*" is the only value IAM will accept.

Query: Does an IAM policy attached to a Lambda execution role need an MFA condition on sts:AssumeRole?

Short answer: No — this is a documented false positive. AWS's MFA recommendation targets human users signing in interactively; a Lambda execution role authenticates via short-lived STS credentials with no login step. The agent backs this with a full decision table by identity type, and flags that the same policy's PassRole scoping would need a separate check.

Query: A policy grants iam:PassRole scoped to one specific role ARN, with no iam:PassedToService condition. iam-lint doesn't flag it. Is it fully safe?

Short answer: no — it's hardened against one threat, not both. This is the one case where the linter under-checks instead of over-firing: it clears the statement once the resource is a specific ARN, but AWS also recommends an iam:PassedToService condition to close the confused-deputy gap. The agent calls it a gap in coverage, not a false positive, and points to a manual fix instead of an exception.

Query: An S3 bucket policy has Principal: '*' but with an aws:SourceArn condition restricting it to one CloudFront distribution. iam-lint flags it as critical. Real risk or false positive?

Verdict: very likely a false positive, but document it instead of suppressing it. The linter stops reading at Principal and never looks at the Condition block, while AWS documents a wildcard principal with a tight condition as a normal pattern. The agent lists what to verify before closing it, such as whether the ARN names one concrete distribution, and says plainly which parts it can't judge without seeing the policy.

Full unedited transcripts, including one where the agent caught and flagged a citation bug in its own source data rather than silently mis-citing, are in transcripts.md.

Code

github.com/Bims-creator/sanity-iam-agent

How I Used Sanity

The Knowledge Base is built from five markdown files (knowledge-base/*.md in the repo), one per IAM finding, added as a Files source in Sanity Context. Each file states the linter's claim, what AWS/CIS says about the same pattern, and where the two land — agreeing, partially conflicting, or (in one case, full admin access) not in conflict at all, included on purpose as a control so the agent doesn't manufacture disagreement where none exists.

The build did more than I expected. Beyond one entry per source file, it synthesized a cross-cutting entry — "True-Risk vs False-Positive Evaluation Method" — that pulled a shared decision framework out of the five files: four context factors (action type, identity type, condition blocks, service-specific semantics), a false-positive/real-risk signal pair for each rule, and an exception-documentation template. That entry is doing most of the actual work when the agent answers a question.

The build also caught something I didn't plan for: two of my source files got their citation numbering crossed in a couple of the generated entries (a wildcard_principal source cited as if it were the passrole source, and vice versa). Context flagged all four instances as critical conflicts in its Issues review, I resolved them by picking the correct source-to-claim mapping, and a full rebuild propagated the fix. Worth knowing if you're building on this: the automated conflict detection covers structural mismatches like this reliably, but I still had the agent itself catch one straggler in a newly-synthesized entry before I rebuilt — which, if anything, is a decent argument for the "surface the disagreement instead of hiding it" design.

The agent connects to the Knowledge Base through the Sanity Context MCP endpoint, using Claude's native MCP connector (mcp_servers + an mcp_toolset tool) so the tool calls happen server-side — no manual MCP client loop in the agent code.

Sanity Project Details

Project ID: r3gwx3ch

Agent Session

Sanity IAM Policy Review Agent — build session claude-sonnet-5
You

Can we give this a go?

Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent

The Sanity Challenge is real and live (Sept 18 – Oct 4, 2026, $2,500 across 5 winners). Two paths:

  • Path One — Ship an Agent: build any agent, point it at a Sanity Context MCP endpoint (Knowledge Base over structured content), judged on how well structured content beats plain keyword search.
  • Path Two — Vibe-Code: prompt your way to a Next.js/Astro app with Sanity as the backend, judged more on the build process/writeup than polish.

Worth flagging: Sanity's stack is TypeScript/JS-native (schemas in TS, GROQ queries, Next.js/Astro frontends) — a departure from your usual AWS/Python/cloud-infra portfolio track, and submissions are due Oct 4 (about 2.5 weeks out), which is tight alongside your interview prep pipeline. Path One (an agent over an MCP endpoint) is the closer fit to what you've been building lately (agentic/AI-adjacent work) and doesn't require deep Sanity fluency — you'd mostly need a Knowledge Base config and any agent framework talking to it via MCP.

Agent
Agent

Path One (Agent + MCP) is the stronger Global Talent play. Here's why:

  • Judging criteria matches evidence criteria. Path One is scored on technical implementation, meaningful use of structured content, and whether the result required more than keyword search — that's "exceptional technical contribution," one of the actual Tech Nation/digital-tech assessment categories. Path Two is scored on writeup honesty and creativity of a CMS-driven app — more portfolio filler than evidence of technical depth.
  • It's public, judged, and prize-bearing. A DEV.to post + a sponsor (Sanity) judging it is third-party recognition, however modest — placing or even a strong published submission is citable external evidence, unlike a solo repo nobody reviewed.
  • It can extend your existing narrative instead of starting a new one. Your portfolio is cloud/security tooling (iam-lint, aws-audit-platform, k8s-multicloud). A Knowledge Base built from, say, AWS IAM policy docs / CIS benchmarks / your own audit rule set, with an agent that flags where two sources contradict each other (a real problem in security compliance), is a natural extension — not a detour into an unrelated CMS demo.

Tradeoff: Path One still requires learning Sanity's Content Lake/MCP setup from scratch under a ~2.5 week deadline, and it competes for time with your Accenture/Deloitte/GSK interview prep.

Want me to sketch the concrete build (e.g. a Knowledge Base over AWS/IAM security docs + an agent that surfaces contradicting guidance) before we start, so you can sanity-check the scope first?

You

Let's build the Path One

Agent

Two things before we start:

  1. You'll need to create the Sanity account yourself — I can't create accounts or handle sign-ups on your behalf. I'll guide you through it step by step (it's free), but you'll click through the signup.
  2. Content domain matters a lot for the "Global Talent" angle — the strongest submission reuses content you already authored, so the Knowledge Base has real depth and provenance instead of generic docs.
Agent
Agent
Agent
Agent
Agent
Agent

This is a multi-day build across a new stack (TypeScript/Sanity/MCP) plus content authoring, so let me lay out the full plan before we start writing code.

Agent
Agent

Here's how to begin, in order:

1. Create your accounts (you do this — I can't create accounts for you)
- sanity.io → sign up free, create a new project (I'll walk you through the CLI setup once you're in)
- dev.to → sign up if you don't have an account (submissions are published there with the #sanitychallenge tag)

2. While you do that, I'll start on the content — the Knowledge Base source material (the IAM security guidance docs with the deliberate contradictions we discussed: wildcard resources, PassRole, MFA-on-service-roles, wildcard principals). This is plain Python/Markdown work in your existing comfort zone, no Sanity needed yet.

3. Once you have a Sanity project, we run sanity init/configure Sanity Context together, pointing it at the content from step 2 to build the Knowledge Base and get an MCP endpoint.

4. Then we build the agent that queries that MCP endpoint and answers questions by surfacing contradicting sources.

5. Write up + submit on DEV.to with the submission template, tagging #sanitychallenge, optionally embedding this Claude Code session as the agent-session transcript.

Want me to start drafting the IAM guidance content now (step 2) while you go create the Sanity account?

You

Can we beat the deadline?

Agent

Good — looks like you're already signed into DEV.to, so that step's done. Sixteen days (today to Oct 4) is workable if we keep scope tight and start immediately. Rough breakdown:

Days Phase
1–2 Content: write the IAM security guidance docs with the built-in contradictions (I draft, you review)
3–5 Sanity account/project setup, learn just enough schema + Sanity Context to ingest the content into a Knowledge Base
6–8 Get the MCP endpoint live, verify it actually surfaces the contradictions correctly
9–12 Build the agent (Python + Claude API + MCP client) that queries it and answers questions
13–14 Polish, test edge cases, capture the agent session transcript
15–16 Write the DEV.to submission, tag #sanitychallenge, submit

That leaves ~2–3 days of slack. The real risk isn't the calendar — it's that Sanity/TypeScript/MCP is new stack for you, so days 3–8 could run long if the learning curve bites, and you've got interview prep (GSK, Deloitte) competing for the same hours. Keeping the agent itself simple (a CLI script, not a polished app — Path One is judged on the Knowledge Base and technical use of structured content, not UI) is what protects the buffer.

Want me to start on the content today so we're not losing any of those 16 days?

You

Please write the review to sound human and natural rather than AI generated.

Agent

Just to make sure I write the right thing — what "review" are you referring to? We haven't drafted a submission write-up yet, so I want to confirm before I start:

Agent
Agent
Agent
Agent
Agent
Agent

Got it — noted for good. From here on, any write-up, README, or submission text I draft for this project will be in a natural voice, no AI tells (formulaic openers, bullet-stuffed prose, hedging, summary-paragraph wrap-ups). Code stays as-is since that instruction was about prose, not comments.

Since there's no draft yet to rewrite, want me to go ahead and start on the IAM content docs now (the step we agreed on), writing the rationale/caveat text for each entry in that natural voice from the start?

You

Join the Sanity Challenge: $2,500 in prizes for FIVE winners!
#sanitychallenge#devchallenge#agents#webdev
We're excited to announce our newest challenge with Sanity!
The Sanity Challenge runs September 18 to October 4. Build an AI agent on structured content, or vibe-code an app with Sanity behind it. $2,500 in prizes.
New to Sanity? Sanity is the AI Content Operating System. Your content lives in the Content Lake as JSON documents, with schemas you define in TypeScript and query with GROQ. You can use Sanity to build anything that you can imagine with structured content as your starting point.
Read on to learn more.
Our Prompts
Path one: Ship an agent that queries real content
Build an agent, then point it at a Sanity Context MCP endpoint backed by a Knowledge Base. Any agent framework, hosted anywhere.
Build anything that needs an answer it can't afford to get wrong. A board game companion that knows the errata contradicts the rulebook. A better interface to your favorite open-source docs. A eurorack planner that knows what actually fits in your case. An award-travel agent that untangles which transfer partner story is current. A car repair agent? A camera gear-head compendium? The sky is the limit.
Point Sanity Context at a website, a set of files, or your own Sanity content, and it distills a navigable Knowledge Base your agent reads through MCP. Every entry stays linked to the source it came from. When two sources contradict each other, both claims surface side by side with their sources, and the decision you make carries across future builds. It all lives in your Sanity Dashboard.

Path One Submission Template

The strongest submissions will show an agent that only works because the content was structured. If a keyword search would have gotten you the same answer, aim higher.
Path two: Vibe-code something strange
Prompt your way to a working app. Any AI-native IDE, Next.js or Astro on the front, Sanity behind it.
This one is judged on the build as much as the result. How deep did you get into Sanity's features? Did you customize the interface? Build a new component to turn videos into gifs? Create a workflow that kicks off an external API call? A rough app with an honest writeup beats a polished one with three sentences.
Bonus points for reaching past the Studio. Two things we'd especially like to see prompted into existence:

  • App SDK: build a custom app on top of your content, with real-time data and your own interface, instead of another read-only frontend.
  • Workflows: model a process (content reviews, translations, and so on) as data next to the content, so an agent can move a draft forward and a person can approve it through the same transitions.

Neither is required. A submission that uses one well will stand out from a pile of blog templates.

Path Two Submission Template
One Extra Submission Requirement 📌
Every submission, both paths, needs to include your Sanity project ID or a link to a public dataset URL.
This lets the Sanity team look at how you actually modeled and used your structured content, which is a real part of judging. There's a section for it in both templates above. Submissions without it may be considered incomplete, so don't skip it.
Judging Criteria
Path one: Ship an agent that queries real content submissions will be evaluated for:

  • Meaningful use of Sanity Context and structured content
  • Technical implementation and code quality
  • Use of Knowledge Bases
  • Usability

Path two: Vibe-code something strange submissions will be evaluated for:

  • Quality and honesty of the build process writeup
  • Functionality of the finished app
  • Thoughtfulness of the schema behind it
  • Creativity and originality

Prizes 🏆
Three winners from Path one, "Ship an agent that queries real content," will receive:

  • $500 USD cash prize
  • DEV++ Membership
  • Exclusive DEV Badge

Two winners from Path two, "Vibe-code something strange," will receive:

  • $500 USD cash prize
  • DEV++ Membership
  • Exclusive DEV Badge

All participants with a valid submission will receive a completion badge on their DEV profile.
How To Participate
Publish a post on DEV using the prompt submission templates above and be sure to include the required challenge tag: #sanitychallenge. You may submit to both paths, but you must create a separate post for each.
If your app requires logging in, please provide testing credentials in your submission and/or instructions on how to best test your application for judges.
Optional but encouraged: embed your agent session. Upload a transcript from Claude Code, Gemini CLI, Codex, GitHub Copilot CLI, or Pi, curate or slice the parts worth showing, and drop it straight into your post. It beats a screenshot, and for Path two it's the most direct way to show your work.
Share, Embed, and Curate Agent Sessions on DEV [Beta]
Jonathan Gottfried
for The DEV Team
4 Mar
Share, Embed, and Curate Agent Sessions on DEV [Beta]
#ai#forem#devto#gemini
76 reactions
20 comments
3 min read
Head to the Agent Sessions uploader to get started. Two things to remember: uploads are unlisted by default, so use the Make Public button or judges won't be able to open your session. And check your transcript for keys and sensitive data before publishing. There's a redaction layer, but it won't catch every edge case.
Please review our judging criteria, rules, guidelines, and FAQ page before submitting so you understand our participation guidelines and official contest rules such as eligibility requirements.
Getting Started With Sanity 🚀

Important Dates

  • September 18: Sanity Challenge begins!
  • October 4: Submissions due at 11:59 PM PDT
  • October 22: Winners Announced

We can't wait to see what you build. Questions about the challenge? Drop them in the comments below.
Good luck and happy coding!

Agent

This confirms a few important details I'll fold into the plan: separate DEV.to post per path (if submitting both), the required tags (devchallenge, sanitychallenge, sanity, ai), and that the Sanity project ID / public dataset URL is a hard requirement, not optional. Also useful — Sanity Context can ingest "a set of files" directly, so we may not need heavy custom TypeScript schema work up front.

Agent
Agent

Now let's start on the content. I'll set up the project folder and draft the first Knowledge Base entries — each one pairs a claim from your iam-lint rule set against real AWS/CIS guidance, deliberately including spots where they disagree.

Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent

Yes — separate from DEV.to. The "Sign up" button in that screenshot is DEV.to's own challenge-notification signup (and that page looks like a stale/cached pre-launch version, since it still says "Upcoming" and "Submissions: TBA" — the live challenge we found earlier already has real dates). You'll additionally need a free Sanity.io account to actually build the Knowledge Base and get a project ID, which the submission requires either way.

To sign up:
1. Go to sanity.io and click Get started free
2. Sign up with GitHub, Google, or email — your call
3. Once you're in, you'll land on a dashboard prompting you to create a project — just tell me when you're there and I'll walk you through naming it and picking a dataset

Meanwhile, I've drafted the first 5 Knowledge Base entries at knowledge-base/ — each one pairs an iam-lint rule against real AWS/CIS guidance, with 4 of the 5 containing a genuine point of disagreement (wildcard resources, PassRole, MFA on service roles, wildcard principals) and one control case (full admin access) where everyone agrees. Want to review those before you go sign up, or sign up first and we'll review in parallel?

Agent

For "What will you use Sanity for?" — pick Personal. This is a portfolio/challenge project, not something tied to an employer or client work, and Personal typically skips team/billing prompts that Work or Agency would trigger, which keeps onboarding faster.

Click Personal, then let me know what step 2 asks for — I'll guide you through each one as it comes up (this wizard usually moves into project name, dataset visibility, and template/schema choice next).

Agent

Pick Beginner — that's the honest answer given you're coming from Python/AWS rather than JS/TypeScript, and it should steer Sanity's onboarding toward more guided defaults instead of assuming you already know GROQ/TS schema syntax (which fits our plan anyway, since we're avoiding heavy custom schema work).

Agent

This one's literally on the tin — pick Knowledge base ("Create documentation and support content"). That's exactly what we're building, and it should point Sanity's onboarding at Knowledge Base features and starter templates instead of e-commerce or blog scaffolding we don't need.

Agent

Name it something that reads clearly on a submission — I'd go with iam-security-agent (matches the local project folder name style you already use, and tells a judge exactly what it is at a glance). Type that in and hit Create project.

Agent

Pick "I don't know" — none of these fit. We're not building a frontend at all for Path One; the agent is a standalone script that queries the MCP endpoint directly. Selecting a framework here would just make Sanity scaffold a Next.js/Astro starter app we don't need. "I don't know" should keep it to a bare Studio + Content Lake setup, which is exactly what we want to build the Knowledge Base on top of.

Agent

Pick "No, building from scratch" — this is a brand-new Sanity project for the challenge, not an integration into something existing.

Agent

Pick "No, starting from scratch" — same logic, no migration involved here.

Agent

Let me check Sanity's actual docs for Knowledge Bases / Sanity Context before telling you what to click here — this dashboard defaults to a Next.js scaffold, but I want to confirm we actually need that for an agent-only submission before you commit to it.

Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent

That "Getting started" widget with Next.js is for Path Two (a full app) — we don't need it. Good news: for Path One, we don't need a Studio or any frontend framework at all. Here's what the actual docs say the real path is:

  1. Enable Context (Knowledge Bases) for your org first — it's an opt-in beta feature, turned on from the Labs page under your organization's Manage settings (org: oyryduzu). This is a prerequisite before anything else works.
  2. Then in the Sanity Dashboard, open the Context app → New knowledge base → give it a title and a one-two sentence purpose (something like "Security-review Knowledge Base for AWS IAM policies — helps decide whether a static-analysis finding is a real risk or a false positive.")
  3. Add source → Files, and point it at our knowledge-base/*.md files
  4. Click Build entries — Sanity reads the files, builds a topic tree, and should flag the contradictions we wrote in as Issues to resolve
  5. Finally, create an org-level API token with Context Viewer permission (under Manage → API → Tokens) to serve the Knowledge Base to our agent over MCP

So: skip "Agent setup"/"Guided Installation" on the screen you're on. Instead, click on Bims-creator at the top-left (or look for a "Manage organization" link) to find the Labs page. Screenshot what you see there and I'll point you to the exact toggle.

Agent

There's the Labs tab right in that nav bar — click it.

Agent

Exactly what we need. Click "Enable for Organization" next to Context.

Agent

Context is enabled (green checkmark). Now click "Open in dashboard" to get into the Context app where we'll create the Knowledge Base.

Agent

We're right where we need to be. Click "New knowledge base".

Agent

Fill it in like this:

Title:

IAM policy review

Purpose:

Answer whether an IAM policy finding from static analysis is a real risk or a false positive. Lead with: the linter's rule and severity, what AWS/CIS guidance says about the same pattern, and whether they agree or conflict. Leave out: IAM topics unrelated to the five reviewed patterns (wildcard resource, PassRole, MFA, wildcard principal, full admin access).

That mirrors the placeholder's structure (who asks, what leads, what's excluded) and should steer the build toward surfacing exactly the agree/disagree comparisons we wrote into the files. Type those in and let's see what the "add sources" step looks like next.

Agent

That's filled in correctly. Click "Create knowledge base".

Agent

Knowledge base created. Click the "Files" tab (next to Dataset/Website) — that's the source type we want, since our content is the five markdown files in knowledge-base/.

Agent

Click "Browse files", and in the file picker navigate to:

[REDACTED]\sanity-iam-agent\knowledge-base

Select all 5 files there (wildcard-resource.md, passrole-unrestricted.md, mfa-sensitive-actions.md, wildcard-principal.md, full-admin-access.md) and upload them together.

Agent

All 5 files uploaded and marked Ready. Now click "Build entries" — this is where Sanity reads them, builds the topic tree, and should flag the deliberate contradictions as Issues.

Agent

This is a great sign — exactly 4 issues flagged, matching our 4 deliberately-contradictory topics (wildcard resource, PassRole, MFA, wildcard principal), while full-admin-access.md — our control case where every source agrees — correctly generated no issue. That's the build behaving exactly as designed.

Click "Review" (top right) or the Issues tab in the sidebar to see how it framed each conflict.

Agent

That's a good sign the build is treating "IAM policy review" as one interconnected topic rather than 5 separate silos — but look closely: all 4 flagged issues are about citation numbering being crossed between WILDCARD_PRINCIPAL and PASSROLE_UNRESTRICTED (e.g., a rule reference table citing source [5] for wildcard_principal when [5] is actually the passrole source). That's a structural/citation mix-up, not the semantic disagreements we deliberately wrote in (wildcard-resource-as-false-positive, PassRole needing iam:PassedToService, MFA-on-service-roles, wildcard-principal-with-condition).

That suggests the build consolidated all 5 rules into one combined "rule reference table" entry and mislabeled which numbered source belongs to which rule — likely a quirk of this being a beta feature, or of our files sharing similar structure (each has a "Claim: iam-lint" / "Claim: AWS..." pattern) confusing the source-matching step.

Click "View entries" in that email (or open the Entries tab in the dashboard) so we can see what got generated — I want to check whether the real content-level contradictions still made it into the entry text correctly despite the citation numbering being off, before deciding whether to just fix the citations or rebuild.

Agent

Good news — this confirms the build actually did pick up our substantive framing (see that "True-Risk vs False-Positive Evaluation Method" entry, with a "Where iam-lint is correct but incomplete" section — that's exactly the "Where this actually lands" reasoning we wrote into the files). The 4 issues are all the same narrow bug: two source UUIDs got their numeric labels swapped in a couple of cross-referencing tables — specifically the wildcard_principal source (UUID starting a076...) and the passrole source (UUID starting cb28...) — never anything involving the real semantic contradictions we designed.

For all 4, the fix is the same: pick the lower box in each (the one that correctly matches the UUID to its actual topic — e.g. "Source [5] is cb28... (passrole)... correct source for WILDCARD_PRINCIPAL is [3]"), then click "Resolve issue". Do that on all 4:

  1. "iam-lint Rule Reference" (WILDCARD_PRINCIPAL→[5] mixup) → select the box citing a076... as [3]
  2. "True-Risk vs False-Positive Evaluation Method" (Context factor 3) → select the box saying WILDCARD_PRINCIPAL should cite [2]
  3. "True-Risk vs False-Positive Evaluation Method" (PassRole/[2] mixup) → select the box saying PassRole should cite [4]
  4. "iam-lint Rule Reference" (PassRole→[3] mixup) → select the box saying PassRole should cite [5]

Once all 4 are resolved, go back to Entries and let's actually read the full "True-Risk vs False-Positive Evaluation Method" entry — that's the one I want to check word-for-word to confirm it's surfacing our real contradictions (wildcard-resource false positives, MFA-on-service-roles, etc.), not just the citation table.

Agent

All 4 resolved exactly as planned — nice work. Now let's go check the actual content. Head to the Entries tab and open "True-Risk vs False-Positive Evaluation Method" — screenshot that full entry so we can confirm it's actually reasoning through the wildcard-resource, MFA-on-service-roles, and PassRole disagreements the way we designed, not just listing the rule table.

You

True-Risk vs False-Positive Evaluation Method
Framework for deciding whether an iam-lint finding is a real risk or false positive: context factors (action type, identity type, condition blocks), systematic false-positive patterns, exception documentation criteria
The core evaluation question for any iam-lint finding is whether the flagged pattern is a genuine misconfiguration or a false positive produced by the tool's inability to read context that a JSON policy document alone doesn't supply. Four context factors determine this for every finding:Rewrite…

  1. Action type — does the action even support resource-level permissions? Some actions have no alternative but "*".
  2. Identity type — is the policy attached to a human user or to a machine identity (EC2 instance profile, Lambda execution role, service principal)?
  3. Condition block — is a wildcard narrowed by aws:SourceArn, aws:SourceAccount, aws:MultiFactorAuthPresent, iam:PassedToService, or similar conditions?
  4. Service-specific semantics — does the AWS service in question document a particular pattern (e.g., Principal: "*" for static website hosting) as intentional and supported?

Rewrite…
Systematic false-positive patterns by ruleRewrite…
WILDCARD_RESOURCE — high severityRewrite…
iam-lint fires on any Allow statement where Resource contains "*", without checking which action is involved . This produces a background false-positive rate on read-only, account-wide actions — ec2:DescribeInstances, iam:ListRoles, s3:ListAllMyBuckets — where the AWS Service Authorization Reference explicitly marks the wildcard as the only accepted value; the API has no concept of a resource-scoped ARN for these calls . The resolution is to cross-check the flagged action against AWS's resource-type table before treating the wildcard as a finding.Rewrite…
False positive signal: action is a List/Describe/Get that AWS marks as not supporting resource-level restrictions.
Real risk signal: action is a mutating call (s3:PutObject, ec2:TerminateInstances) with no ARN scope.Rewrite…
PASSROLE_UNRESTRICTED — high severityRewrite…
iam-lint stops flagging once Resource is scoped to a specific role ARN . That removes the "any role in the account" risk and is a genuine improvement. However, AWS's own PassRole guidance recommends also adding an iam:PassedToService condition to restrict which service can assume the passed role . Without that condition, a narrowly-scoped role can still be handed to a service it was never intended to run under (the confused-deputy angle). A policy that passes iam-lint's check can still be missing half of what AWS considers the complete mitigation — the rule's naming implies "fully safe" when the check only verifies "not wildcard" .Rewrite…
False positive signal: resource is already scoped to a specific ARN and the intended service is tightly controlled by other means.
Real risk signal: wildcard resource, or specific ARN but no iam:PassedToService condition (per AWS's fuller recommendation).Rewrite…
MISSING_MFA_CONDITION — medium severityRewrite…
iam-lint flags any sensitive action (iam:CreateAccessKey, iam:AttachUserPolicy, iam:PutUserPolicy, sts:AssumeRole, iam:PassRole) that lacks aws:MultiFactorAuthPresent, applied identically regardless of who the policy is attached to . AWS's MFA guidance targets interactive human logins — a person supplying a hardware key or OTP. Machine identities (EC2 instance roles, Lambda execution roles, service principals) authenticate via short-lived STS credentials with no login step and no mechanism to present MFA . Applying the condition to an automated role would break the automation without adding security .Rewrite…
False positive signal: policy is attached to a service role, instance profile, or Lambda execution role.
Real risk signal: policy is attached to a human IAM user or a role assumed interactively via the console or CLI.
Diagnosis gap: a JSON policy document alone does not carry the identity type — determining true/false positive requires checking what the policy is actually attached to.Rewrite…
WILDCARD_PRINCIPAL — critical severityRewrite…
iam-lint fires at the highest severity on any Principal: "*" or equivalent in a trust or resource-based policy, with no examination of accompanying conditions . AWS documents Principal: "*" paired with a restrictive Condition (e.g., aws:SourceArn, aws:SourceAccount) as a normal, explicitly supported pattern — SNS, SQS, S3 static website hosting, and CloudFront origin access all use this shape in AWS's own example policies . A bare wildcard with no condition is a genuine critical finding. The same wildcard beside a tight condition block is a deliberate, documented pattern .Rewrite…
False positive signal: Principal: "*" is accompanied by a Condition block that enforces a specific source ARN, account, or service.
Real risk signal: Principal: "*" with no condition — genuinely open to any unauthenticated principal.
Operator note (source assignment): WILDCARD_PRINCIPAL findings should cite source (a07619d4); PassRole findings cite source (cb28a96d). See the operator instructions for context.Rewrite…
Exception documentation criteriaRewrite…
When a finding is confirmed as a false positive, document the exception with:Rewrite…

  • The specific rule ID and the exact statement (SID or action + resource) being excepted.
  • Which context factor resolves it (action lacks resource-level support; identity is a machine role; condition block narrows the wildcard; pattern is AWS-documented for the service).
  • A pointer to the AWS reference that confirms the pattern is intentional (service authorization table entry, Well-Architected guidance, or resource policy documentation).
  • A review date — AWS occasionally adds resource-level support for previously wildcard-only actions, or updates service policy patterns.

Rewrite…
Remediation steps and policy rewriting are out of scope here; see findings/* entries for rule-specific remediation.

Agent

This is exactly the differentiator the challenge is judging on — it's not just listing rules, it built an actual decision framework (4 context factors), a false-positive/real-risk signal pair for each rule, and an exception-documentation template. That "Diagnosis gap" line under MFA — "a JSON policy document alone does not carry the identity type" — is precisely the "why keyword search wouldn't get you this" argument the judges are looking for. Full admin access correctly doesn't appear here since it has no false-positive angle. This entry alone could anchor the whole submission write-up.

Two things next:

1. Create an organization API token (needed before the MCP endpoint will work):
Go to Manage → API → Tokens at the organization level (not the project level) → create a new token → give it Context Viewer permission. Copy the token somewhere safe — Sanity only shows it once.

2. Create the MCP endpoint:
Back in the Context app sidebar, click "New endpoint", add the "IAM policy review" Knowledge Base as its source. That gives us a URL we'll point the agent at.

Go ahead and create the token first, then create the endpoint — screenshot each step and I'll confirm as we go.

You

Point me to where to find manage

Agent

From the org overview page you were on earlier (the one showing "Bims-creator" with the Overview/Projects/Labs/Members/Activity/Usage/Billing/API/Security/Settings tabs) — that page is the "Manage" console. Click "API" in that same tab row, then look for a "Tokens" sub-tab.

If you've navigated away from it: click the Bims-creator org switcher at the top-left of any Sanity screen → it drops into that same organization overview → then API → Tokens.

Agent

Right where we need to be. Click "Add API token".

Agent

Fill it in like this:

  1. Name: iam-security-agent-mcp (so it's identifiable later)
  2. Expiration: leave as "No expiration" — fine for a short-lived challenge project
  3. Under Organization permissions, check the Context checkbox, then select the Viewer radio button underneath it (not Editor — the agent only needs to read)
  4. Leave everything else unchecked, and skip Project access — Context tokens work at the org level
  5. Click Save token

Important: the token value will only be shown once. Copy it immediately and save it somewhere — don't paste it into this chat, just confirm you've saved it once you see it.

You

Token saved, what's next

Agent

Good — keep that token safe, don't paste it into chat. Now let's create the MCP endpoint:

  1. Go back into the Context app (org sidebar icon we used earlier, the sunburst/asterisk icon, or via Open Dashboard → Context)
  2. In the left sidebar under MCP ENDPOINTS, click "New endpoint"
  3. Give it a name (e.g. iam-policy-review-endpoint)
  4. When it asks for a source, add the "IAM policy review" Knowledge Base (not a dataset or website) — this is what makes the endpoint serve Knowledge Base tools instead of raw GROQ queries
  5. Save it, and it should give you an MCP endpoint URL

Screenshot each step as you go — once we have that URL plus your Context Viewer token, we can start wiring up the actual agent.

You

Cool

Agent

Go ahead and walk through those steps — send a screenshot once you're creating the endpoint or if anything looks different from what I described.

Agent

That's set up correctly — Knowledge base source selected, "IAM policy review" is checked and Ready. Fill in the Instructions box with something that steers the agent toward the actual value of this Knowledge Base:

You are a security review assistant for AWS IAM policies. When asked about a static-analysis finding (e.g. from iam-lint), check the Knowledge Base for the relevant rule and apply its context factors (action type, identity type, condition block, service-specific semantics) to judge whether it's a real risk or a false positive. When sources disagree or a finding is more nuanced than the rule's severity implies, say so explicitly and cite where each claim comes from — do not silently pick one.

Scroll down for a Save or Create endpoint button and click it — then share the full MCP URL it generates (the one starting https://api.sanity.io/v1/context/organizations/oyryduzu/...) so we can wire the agent to it.

Agent

We've got everything needed now: the endpoint URL and your Context Viewer token. Before I write the agent code, let me pull up the current Claude API reference for the MCP connector feature — Anthropic's Messages API can talk to a remote MCP server like this one directly, which would mean less custom client code than a manual MCP loop.

Agent
You

Base directory for this skill: [REDACTED]\AppData\Local\Temp\claude\bundled-skills\2.1.274\707515a2f06def1c7b0d42b02e1c13d2\claude-api

Building LLM-Powered Applications with Claude

This skill helps you build LLM-powered applications with Claude. Choose the right surface based on your needs, detect the project language, then read the relevant language-specific documentation.

Before You Start

Scan the target file (or, if no target file, the prompt and project) for non-Anthropic provider markers - import openai, from openai, langchain_openai, OpenAI(, gpt-4, gpt-5, file names like agent-openai.py or *-generic.py, or any explicit instruction to keep the code provider-neutral. If you find any, stop and tell the user that this skill produces Claude/Anthropic SDK code; ask whether they want to switch the file to Claude or want a non-Claude implementation. Do not edit a non-Anthropic file with Anthropic SDK calls. (Exception: the prompt-audit subcommand is non-interactive and does not stop here - it records non-Anthropic provider markers in its report's stated assumptions and never proposes switching a non-Anthropic file to the Anthropic SDK.)

Output Requirement

When the user asks you to add, modify, or implement a Claude feature, your code must call Claude through one of:

  1. The official Anthropic SDK for the project's language (anthropic, @anthropic-ai/sdk, com.anthropic.*, etc.). This is the default whenever a supported SDK exists for the project.
  2. Raw HTTP (curl, requests, fetch, httpx, etc.) - only when the user explicitly asks for cURL/REST/raw HTTP, the project is a shell/cURL project, or the language has no official SDK.

Never mix the two - don't reach for requests/fetch in a Python or TypeScript project just because it feels lighter. Never fall back to OpenAI-compatible shims.

Never guess SDK usage. Function names, class names, namespaces, method signatures, and import paths must come from explicit documentation - either the {lang}/ files in this skill or the official SDK repositories or documentation links listed in shared/live-sources.md. If the binding you need is not explicitly documented in the skill files, WebFetch the relevant SDK repo from shared/live-sources.md before writing code. Do not infer Ruby/Java/Go/PHP/C# APIs from cURL shapes or from another language's SDK.

If WebFetch or repository access fails (network restricted, timeouts, clone blocked): do not keep retrying - write code from the patterns and namespace/package tables in the {lang}/ file, run the compiler or interpreter on it, and iterate on the error output. For statically-typed SDKs (C#, Java, Go) a compile-fix loop against local errors reaches working code faster than blocked network research.

Defaults

Unless the user requests otherwise:

For the Claude model version, please use Claude Opus 5, which you can access via the exact model string claude-opus-5. Please default to using adaptive thinking (thinking: {type: "adaptive"}) for anything remotely complicated. And finally, please default to streaming for any request that may involve long input, long output, or high max_tokens - it prevents hitting request timeouts. Use the SDK's .get_final_message() / .finalMessage() helper to get the complete response if you don't need to handle individual stream events. When a streaming request defines user-defined (client) tools, set eager_input_streaming: true on each of those tools so large tool inputs (file contents, code, documents) stream as they are generated instead of arriving in one burst after the server finishes buffering them; the client then owns validation: the SDKs' tolerant parsers can return a silently truncated input instead of raising, so validate each parsed tool input against its schema before running it (the typed runner helpers such as betaZodTool / typed @beta_tool do this; betaTool() JSON-Schema tools and manual loops must validate themselves), treat a failure like invalid JSON (INVALID_JSON error tool_result when you hold the block, re-issue otherwise), check max_tokens / refusal stop reasons before running tools, and catch only the SDK's JSON error, never its typed API errors - pattern in shared/tool-use-concepts.md -> Eager input streaming. Leave it off for non-streaming requests, for server tools, and when the request goes through a proxy or an older Bedrock model deployment that rejects the field.

Warning: API Drift - Your Training Prior May Be Stale

Several common Claude API shapes changed in 2025-2026. If you recall a pattern from training, verify it against the {lang}/ files in this skill before writing - the rows below are the most frequent drift points:

Area Stale prior Current API
Extended thinking thinking: {type: "enabled", budget_tokens: N} On Claude 4.6+ models: thinking: {type: "adaptive"}. budget_tokens is deprecated on Opus 4.6 / Sonnet 4.6 and rejected with a 400 on Fable 5/5.1 / Sonnet 5 / Opus 5 / 4.8 / 4.7. Pre-4.6 models still use budget_tokens.
Web search / web fetch tool type web_search_20250305, web_fetch_20250910 web_search_20260209, web_fetch_20260209 (dynamic filtering) on Opus 5/4.8/4.7/4.6, Sonnet 5, and Sonnet 4.6. Older models keep the basic variants; on Vertex AI only basic web_search_20250305 is available (web fetch is not on Vertex) - see the Server Tools QR below.
PHP parameter names snake_case wire names as named args (max_tokens) Top-level named args are camelCase (maxTokens). Nested array keys vary by feature (e.g. 'taskBudget', 'skillID', 'mcp_server_name') - copy the exact key from the documented example; do not bulk-convert.
Managed Agents credentials Keep secrets host-side via custom tools (the only option before vaults shipped) Vault environment_variable credentials - stored by Anthropic, substituted at egress, never visible in the sandbox (shared/managed-agents-tools.md -> Vaults). Host-side custom tools remain the fallback for self-hosted sandboxes.
Files API / Skills client.beta.files.* / client.beta.skills.* with beta files-api-2025-04-14 / skills-2025-10-02 Out of beta: client.files.* / client.skills.*, no beta header. In current SDKs client.beta.files / client.beta.skills have breaking shape changes from previous versions, matching the stable namespaces - migrate per shared/live-sources.md -> Files API / Skills Guide.

The {lang}/ files in this skill are authoritative over recalled patterns.


Subcommands

If the User Request at the bottom of this prompt is a bare subcommand string (no prose), search every Subcommands table in this document - including any in sections appended below - and follow the matching Action column directly. This lets users invoke specific flows via /claude-api <subcommand>. If no table in the document matches, treat the request as normal prose.

Subcommand Action
migrate Migrate existing Claude API code to a newer model. Read shared/model-migration.md immediately and follow it in order: Step 0 (confirm scope - ask which files/directories before any edit), Step 1 (classify each file), then the per-target breaking-changes section. Do not summarize the guide - execute it. If the user did not name a target model, ask which model to migrate to in the same turn as the scope question. After the per-target changes are applied, audit the in-scope prompt text, tool descriptions, and request code against shared/prompt-audit.md - prompting written for the source model is part of every migration, and it does not announce itself.
prompt-audit Audit existing prompts, skills, and tool descriptions for dated patterns ("cruft") written for older models. Read shared/prompt-audit.md immediately and follow it in order: Step 0 (establish scope and target model from the request and the repository - state the assumptions in the report, do not stop to ask), inventory, provenance, then the pattern scan. Produce both deliverables in full - the audit report (findings with file:line, pattern, why it's obsolete for the target model, confidence) and a proposed diff - without pausing for confirmation; apply edits only if the request explicitly asked for them. Do not summarize the guide - execute it.
upgrade Upgrade the project's Anthropic SDK dependency across a major version - currently the Python SDK, anthropic 0.x -> 1.x. Trailing words may name the language and/or a scope (upgrade python, upgrade python sdk src/). Read python/claude-api/sdk-upgrade.md immediately and follow it in order: Step 0 (confirm scope, then establish the current and target versions - a published 1.x must exist before you write a pin), the Step 1 inventory, each numbered section, then verification and the report. Do not summarize the guide - execute it. If the detected or named language has no sdk-upgrade.md in this skill, say that no major-version upgrade guide is bundled for that SDK yet and point the user at that SDK's CHANGELOG (repositories in shared/live-sources.md); do not improvise one from the Python guide. This is not model migration - to move code to a newer Claude model, use migrate.
cost-optimize Reduce what existing Claude API code costs to run, without sacrificing output quality. Read shared/cost-optimization.md immediately and follow it in order: Step 0 (establish scope, quality bar, and baseline), the token profile - measured through the Usage and Cost Admin API when the user has an Admin API key, from the app's own response.usage logs when it has those (ask), or estimated from the code otherwise - then a savings-ranked shortlist of levers (quoted in dollars, % of bill, or relative buckets depending on which of those data sources you have), free wins (caching, input-token hygiene, loop hygiene, output-token hygiene, batch) before tradeoffs (budgets, effort, model choice, multi-model); any lever that earns a place becomes its own diff - proposed by default, applied and measured against the eval covering the traffic it touches when the user asks and approves - and "no changes recommended" is a valid outcome. Two standing rules: every run that exercises the model spends real money, so get the user's approval first; and when context for a lever is missing, work through it interactively with the user - this workflow is not expected to one-shot the audit. Do not summarize the guide - execute it; presenting the profile and the ranked plan to the user is part of executing it.
build-eval Help the user build an eval set for their Claude-powered app. Read shared/evals/build-eval.md immediately and run its interview: Step 0 (what's being evaluated), Step 1 (source the prompts - existing eval / transcripts / synthesized), Step 2 (grading method), Step 3 (runnable script + measured cost). Get the user's explicit sign-off on the inputs, the grading method, and the cost before producing the eval.
hillclimb Iteratively improve the user's app against an existing eval. Read shared/evals/eval-hillclimb.md immediately and follow it: Step 0 (confirm a runnable eval exists - if not, route to build-eval), Step 1 (what to change / what's off-limits), Step 2 (budget + stopping condition from measured per-run cost), get the plan approved, then the read->propose->apply->run->record loop with on-disk state and a train/validation/test split.

Language Detection

Before reading code examples, determine which language the user is working in (exception: for the prompt-audit subcommand, skip this section's ask steps - the audit is non-interactive and its inventory is language-agnostic; when no language is inferable, proceed without asking and state the assumption in the report):

  1. Look at project files to infer the language:
  • *.py, requirements.txt, pyproject.toml, setup.py, Pipfile -> Python - read from python/
  • *.ts, *.tsx, package.json, tsconfig.json -> TypeScript - read from typescript/
  • *.js, *.jsx (no .ts files present) -> TypeScript - JS uses the same SDK, read from typescript/
  • *.java, pom.xml, build.gradle -> Java - read from java/
  • *.kt, *.kts, build.gradle.kts -> Java - Kotlin uses the Java SDK, read from java/
  • *.scala, build.sbt -> Java - Scala uses the Java SDK, read from java/
  • *.go, go.mod -> Go - read from go/
  • *.rb, Gemfile -> Ruby - read from ruby/
  • *.cs, *.csproj -> C# - read from csharp/
  • *.php, composer.json -> PHP - read from php/
  1. If multiple languages detected (e.g., both Python and TypeScript files):
  • Check which language the user's current file or question relates to
  • If still ambiguous, ask: "I detected both Python and TypeScript files. Which language are you using for the Claude API integration?"
  1. If language can't be inferred (empty project, no source files, or unsupported language):
  • Use AskUserQuestion with options: Python, TypeScript, Java, Go, Ruby, cURL/raw HTTP, C#, PHP
  • If AskUserQuestion is unavailable, default to Python examples and note: "Showing Python examples. Let me know if you need a different language."
  1. If unsupported language detected (Rust, Swift, C++, Elixir, etc.):
  • Suggest cURL/raw HTTP examples from curl/ and note that community SDKs may exist
  • Offer to show Python or TypeScript examples as reference implementations
  1. If user needs cURL/raw HTTP examples, read from curl/.

Language-Specific Feature Support

Every SDK language above supports both the beta Tool Runner and Managed Agents (beta) - Python (@beta_tool decorator), TypeScript (betaZodTool + Zod), Java (annotated classes), Go (BetaToolRunner in the toolrunner pkg), Ruby (BaseTool + tool_runner), C# (BetaToolRunner + raw JSON schema), PHP (BetaRunnableTool + toolRunner()); code entry points are in the Tool Use Patterns quick reference below. cURL is raw HTTP (no SDK features) and supports Managed Agents.

Managed Agents code examples: see the reading guide in the ## Managed Agents (Beta) section below.


Which Surface Should I Use?

Start simple. Default to the simplest tier that meets your needs. Single API calls and workflows handle most use cases - only reach for agents when the task genuinely requires open-ended, model-driven exploration. "Simplest" means the least code you own: for a hosted, scheduled, or memory-backed agent, Managed Agents is usually the simplest option (no loop code, no state files, no scheduler), even though it's a bigger platform.

Use Case Tier Recommended Surface Why
Classification, summarization, extraction, Q&A Single LLM call Claude API One request, one response
Batch processing or embeddings Single LLM call Claude API Specialized endpoints
Multi-step pipelines with code-controlled logic Workflow Claude API + tool use You orchestrate the loop
Custom agent with your own tools Agent Claude API + tool use Maximum flexibility
Server-managed stateful agent with workspace Agent Managed Agents Anthropic runs the loop and hosts the tool-execution sandbox
Persisted, versioned agent configs Agent Managed Agents Agents are stored objects; sessions pin to a version
Long-running multi-turn agent with file mounts Agent Managed Agents Per-session containers, SSE event stream, Skills + MCP
Agent that runs on a schedule (cron, "every night") Agent Managed Agents - scheduled deployments Deployments fire sessions autonomously; no client-side scheduler
Agent work that must meet a quality bar ("until it's right") Agent Managed Agents - outcomes A separate grader iterates the agent against your rubric until it passes

Note: Managed Agents is the right choice when you want Anthropic to run the agent loop and host the container where tools execute - file ops, bash, code execution all run in the per-session workspace. If you want to host the compute yourself or run your own custom tool runtime, Claude API + tool use is the right choice - use the tool runner for the agentic loop - its per-turn hooks still give you approval gates, logging, error interception, and conditional execution (see shared/tool-use-concepts.md) - or the manual loop when you want to own the entire loop yourself.

Cloud-provider access. Claude Platform on AWS is Anthropic-operated with same-day API parity - see shared/claude-platform-on-aws.md for client setup. For per-feature availability on Claude Platform on AWS, Amazon Bedrock, Google Vertex AI, and Microsoft Foundry, see shared/platform-availability.md - that table is the single source of truth in this skill; do not infer availability from anywhere else.

Building an Agent: Four Approaches

Once you've decided you actually need an agent (open-ended, model-driven tool use), there are four distinct ways to build one. Two independent questions separate them: who supplies the harness (the agent loop + context management) and who supplies the deployment (the infra the agent runs on). The Tool Runner and the Claude Agent SDK both supply a harness only - you still host and deploy them yourself - which is why they're easy to conflate. Managed Agents (CMA) is the only option that supplies both the harness and managed deployment; the manual loop supplies neither.

# Approach You write Harness & deployment Tools available Use when
1 Claude API - manual loop The while stop_reason == "tool_use" loop yourself You build the harness; you host Only tools you define You want to own the entire loop - no beta dependency, or a control flow the Tool Runner's per-turn hooks don't fit
2 Claude API - Tool Runner (client.beta.messages.tool_runner + @beta_tool / betaZodTool) Just the tool functions SDK supplies the loop (harness only); you host Only tools you define A custom-tool agent without hand-writing the loop (most cases). Per-turn hooks still give you approval gates, error interception, result modification (e.g. cache_control), retries, streaming, and compaction
3 Managed Agents (REST, beta) Agent config + your tool results Anthropic supplies the harness and hosts a per-session sandbox (harness + deployment) Anthropic-hosted sandbox (bash, files, code exec) + Skills/MCP + your tools You want Anthropic to run the loop and host the per-session workspace; persisted/versioned configs; long-running sessions
4 Claude Agent SDK - separate product (claude-agent-sdk / @anthropic-ai/claude-agent-sdk) A prompt + options SDK supplies the Claude Code harness + built-in tools (harness only); you host Built-in Read/Write/Edit/Bash/Glob/Grep/WebSearch/WebFetch + MCP + subagents You want a batteries-included coding/filesystem agent running on your own infra

The harness/deployment split is the key mental model: options 1, 2, and 4 all leave deployment to you; only option 3 (CMA) adds managed deployment. Options 1-3 are what this skill generates; option 4 is a different library with its own docs - see the disambiguation below.

Tool Runner != Claude Agent SDK. These sound alike but are different packages:
- Tool Runner is part of the regular Anthropic API SDK (anthropic / @anthropic-ai/sdk), reached via client.beta.messages.tool_runner. It automates the request -> execute -> loop cycle for tools you define. No built-in tools, no filesystem access, no sandbox - you supply every tool and host the compute. It is option 2 above, a thin helper over POST /v1/messages.
- Claude Agent SDK (claude-agent-sdk / @anthropic-ai/claude-agent-sdk) is Claude Code packaged as a library. It ships built-in tools (file read/write/edit, bash, grep, web search), the full agent loop, context management, hooks, subagents, permissions, and sessions. You call query(prompt, options) and it drives everything.

Both are harness-only - you host and deploy them. The difference is scope of harness: the Tool Runner loops over tools you define (with per-turn hooks for approval, interception, result modification, and retries - but no built-in tools); the Agent SDK is the full Claude Code harness with built-in tools. Neither provides managed deployment - that's what Managed Agents (CMA) adds (Anthropic hosts the loop and a per-session sandbox).

This skill covers the Claude API and Managed Agents (options 1-3); it does not generate Claude Agent SDK code. If the user actually wants the Claude Agent SDK, point them to its docs (code.claude.com/docs/en/agent-sdk) - don't substitute the API Tool Runner for it, or vice-versa.

Should I Build an Agent?

Before choosing the agent tier, check all four criteria:

  • Complexity - Is the task multi-step and hard to fully specify in advance? (e.g., "turn this design doc into a PR" vs. "extract the title from this PDF")
  • Value - Does the outcome justify higher cost and latency?
  • Viability - Is Claude capable at this task type?
  • Cost of error - Can errors be caught and recovered from? (tests, review, rollback)

If the answer is "no" to any of these, stay at a simpler tier (single call or workflow).


Architecture

Everything goes through POST /v1/messages. Tools and output constraints are features of this single endpoint - not separate APIs.

User-defined tools - You define tools (via decorators, Zod schemas, or raw JSON), and the SDK's tool runner handles calling the API, executing your functions, and looping until Claude is done. For full control, you can write the loop manually.

Server-side tools - Anthropic-hosted tools that run on Anthropic's infrastructure. Code execution is fully server-side (declare it in tools, Claude runs code automatically). Computer use can be server-hosted or self-hosted.

Structured outputs - Constrains the Messages API response format (output_config.format) and/or tool parameter validation (strict: true). The recommended approach is client.messages.parse() which validates responses against your schema automatically. Note: the old output_format parameter is deprecated; use output_config: {format: {...}} on messages.create().

Supporting endpoints - Batches (POST /v1/messages/batches), Files (POST /v1/files), Token Counting (POST /v1/messages/count_tokens - see shared/token-counting.md), and Models (GET /v1/models, GET /v1/models/{id} - live capability/context-window discovery) feed into or support Messages API requests.


Current Models (cached: 2026-06-24)

Model Model ID Context Input $/1M Output $/1M
Claude Fable 5.1 claude-fable-5-1 1M $10.00 $50.00
Claude Mythos 5.1 (Project Glasswing only) claude-mythos-5-1 1M $10.00 $50.00
Claude Fable 5 claude-fable-5 1M $10.00 $50.00
Claude Opus 5 claude-opus-5 1M $5.00 $25.00
Claude Opus 4.8 claude-opus-4-8 1M $5.00 $25.00
Claude Opus 4.7 claude-opus-4-7 1M $5.00 $25.00
Claude Opus 4.6 claude-opus-4-6 1M $5.00 $25.00
Claude Sonnet 5 claude-sonnet-5 1M $2.00 $10.00
Claude Sonnet 4.6 claude-sonnet-4-6 1M $3.00 $15.00
Claude Haiku 4.5 claude-haiku-4-5 200K $1.00 $5.00

Partner pricing: The prices above are Anthropic first-party API rates - they also apply to Claude on Microsoft Foundry, which is billed through the Microsoft Marketplace at standard API rates. Claude on Amazon Bedrock and Vertex AI is partner-operated with separate pricing - see Bedrock or Vertex AI. For WebFetch, use the Pricing row in shared/live-sources.md.

ALWAYS use claude-opus-5 unless the user explicitly names a different model. This is non-negotiable. Do not use claude-sonnet-5, claude-sonnet-4-6, or any other model unless the user literally says "use sonnet" or "use haiku". Never downgrade for cost - that's the user's decision, not yours. Where a second, cheaper model is in play alongside the main one (worker or sub-agent threads, bulk extractors, LLM judges, the executor under an advisor) - because the user asked for one or a guide in this skill calls for it - or the user says "sonnet" or "haiku" without a version, that means the current generation from the table above (claude-sonnet-5, claude-haiku-4-5); previous-generation IDs such as claude-sonnet-4-6 are only for users who name that version. Use claude-fable-5-1 only when the user explicitly asks for Claude Fable 5.1, "fable", or Anthropic's most capable model - it has different API behavior than the Opus family (see below) and pricing that exceeds Opus-tier. Use only the exact model ID strings from the table - they are complete as-is; never append date suffixes (claude-opus-5, never claude-opus-5-20260401 or any other date-suffixed variant you might recall from training data). If the user requests an older model not in the table (e.g., "opus 4.5", "sonnet 3.7"), read shared/models.md for the exact ID - do not construct one yourself.

Claude Fable 5.1 (claude-fable-5-1) - most capable widely released model

Claude Fable 5.1 is Anthropic's most capable widely released model, for the most demanding reasoning and long-horizon agentic work; everything below also applies to Claude Mythos 5.1 (claude-mythos-5-1, Project Glasswing - same capabilities, pricing, and API surface; it runs safeguards that depend on the access program, so the refusal handling below applies there too; successor to Claude Mythos 5, which ran no safety classifiers). 1M context window (the maximum is also the default), 128K max output. Key API differences from Opus-tier - see shared/model-migration.md -> Migrating to Claude Fable 5.1 for details:

  • Thinking is always on - omit the thinking parameter entirely (or send {type: "adaptive"}). Any other explicit configuration is rejected: {type: "disabled"} and {type: "enabled", budget_tokens: N} both return a 400. Control depth with output_config.effort (supports low through xhigh and max).
  • The raw chain of thought is never returned - responses carry regular thinking blocks (not redacted_thinking): display: "summarized" returns a readable summary, "omitted" (the default) leaves the thinking field as an empty string. Replay rules: pass thinking blocks back unchanged on the same model; other models drop them silently (unbilled - nothing to strip; Claude Mythos 5.1 instead reads them); details in shared/model-migration.md.
  • Tokenizer - same tokenizer as Opus 4.8 (introduced with Opus 4.7). Token counts are roughly unchanged when migrating from Opus 4.7/4.8; per-token pricing differs. Coming from Opus 4.6, Sonnet, Haiku, or older, re-baseline with count_tokens (the Opus 4.7 tokenizer uses ~1×-1.35× as many tokens).
  • refusal stop reason - handle it, and opt into fallbacks by default - safety classifiers may decline a request (HTTP 200, stop_reason: "refusal", with a stop_details category); always check stop_reason before reading content. When you write claude-fable-5-1 or claude-opus-5 code, include the server-side fallbacks parameter by default. Simplest form: betas: ["server-side-fallback-2026-07-01"] + fallbacks: "default", which routes by refusal category so you never maintain a model list. (The older array form - betas: ["server-side-fallback-2026-06-01"] + fallbacks: [{"model": "claude-opus-4-8"}] - still works; Claude API and Claude Platform on AWS - on Bedrock, Vertex and Foundry, use the SDKs' client-side BetaRefusalFallbackMiddleware + BetaFallbackState). Tell the user you've enabled it; drop it only if they decline. Full semantics (billing, mid-stream refusals, credit repricing) in shared/model-migration.md -> refusal section. Per-language code examples in {lang}/claude-api/README.md § Refusal Fallbacks cover the array form only - for the "default" mode, follow the raw-HTTP shape in shared/model-migration.md -> Migrating to Claude Opus 5 -> New API features and swap fallbacks: [{...}] for fallbacks: "default" plus the -2026-07-01 header; the rest of the request is unchanged.
  • No assistant prefill - same as the rest of the 4.6+ family.
  • 30-day data retention required - Claude Fable 5.1 is not available under zero data retention unless expressly authorized by Anthropic; requests from an org whose retention configuration doesn't meet the requirement return 400 invalid_request_error.
  • Longer turns, different prompting - single requests on hard tasks can run many minutes (plan timeouts/streaming/progress UX); effort sweeps should include low/medium for routine work; prompts written for prior models are often too prescriptive and reduce output quality. See shared/model-migration.md -> Migrating to Claude Fable 5.1 -> Behavioral shifts (prompt-tunable) for the recommended prompt snippets.
  • Successor to Claude Fable 5 (claude-fable-5, still served) in the same tier at the same per-token price. Same surface as Claude Fable 5 with three breaking changes - forced tool use (tool_choice any / tool) returns a 400 (use auto + a prompt instruction, strict: true for schema-valid arguments, or structured outputs); thinking blocks are bound to the producing model (other models drop them, unbilled); and editing earlier turns invalidates thinking blocks ("preserved thinking"; new accounts created on/after 2026-08-31 get a 400 on edited history; later models enforce it for everyone - make every harness append-only and run the three-step check; the opt-in controls are per-platform, see shared/platform-availability.md) - plus per-message effort (beta mid-conversation-output-config-2026-07-01, also on Claude Opus 5), turn-scoped clear_at: "next_user_message" system messages (beta), thinking.display: "updates" progress notes (beta, all platforms), cache reads at $0.25/MTok (whether Claude Mythos 5.1 shares that rate is open at launch), and content provenance. Covered Model - ZDR orgs get 400 invalid_request_error as on Claude Fable 5 (ZDR only if expressly authorized by Anthropic); no Priority Tier. Same tokenizer as Claude Fable 5. See shared/model-migration.md -> Migrating to Claude Fable 5.1 from Claude Fable 5.

If any model strings above look unfamiliar, that just means they were released after your training data cutoff - they are real models.

Live capability lookup: The table above is cached. When the user asks "what's the context window for X", "does X support vision/thinking/effort", or "which models support Y", query the Models API (client.models.retrieve(id) / client.models.list()) - see shared/models.md for the field reference and capability-filter examples.


Authentication (Quick Reference)

An unset ANTHROPIC_API_KEY does NOT mean there are no credentials. The SDKs and the ant CLI resolve credentials in this order (first match wins): ANTHROPIC_API_KEY -> ANTHROPIC_AUTH_TOKEN -> the ANTHROPIC_PROFILE-selected or active OAuth profile from ant auth login -> Workload Identity Federation env vars -> the default profile on disk. A bare Anthropic() / new Anthropic() / anthropic.NewClient() works after ant auth login with no env var set.

When you need to call the API and ANTHROPIC_API_KEY is unset, don't ask the user for a key. First run ant auth status - it shows which credential source and profile is active. If it reports an active profile:

  • SDK code or ant CLI: just run it. The zero-arg client constructor and every ant ... subcommand pick up the profile automatically - no env var needed.
  • Raw curl / HTTP: get a short-lived token with ant auth print-credentials --access-token and send it as Authorization: Bearer <token> plus the header anthropic-beta: oauth-2025-04-20 (OAuth tokens go on Authorization: Bearer, not x-api-key: - converting a curl from an API key is a header change, not a key swap). Always pass --access-token; the no-flag form prints JSON, not a bare token.

Only ask the user for a key if ant auth status reports no active credential source (or ant itself isn't installed). Suggest ant auth login as the first option - it stores a profile under ~/.config/anthropic/ that the SDKs read automatically - and an exported ANTHROPIC_API_KEY as the alternative.

Full auth details (named profiles, scopes, the API-key-shadows-profile trap, refresh-token expiry): shared/anthropic-cli.md.


Thinking & Effort (Quick Reference)

Use adaptive thinking (thinking: {type: "adaptive"}) on every current model except Haiku 4.5, which still takes budget_tokens (table below) - Claude dynamically decides when and how much to think. Per-model rules:

Model Thinking config Omitting thinking budget_tokens Sampling (temperature/top_p/top_k) Effort levels
Fable 5 / Claude Fable 5.1 (and the Mythos counterparts) {type: "adaptive"} or omit; explicit {type: "disabled"} returns 400 - omit the param instead (Claude Fable 5.1 / Claude Mythos 5.1 also 400 on forced tool_choice any/tool, and run preserved thinking's history-editing check on replayed thinking blocks) Runs adaptive (thinking is always on) Removed - {type: "enabled", budget_tokens: N} returns 400 Removed - 400 low/medium/high/xhigh/max
Claude Opus 5 {type: "adaptive"} or omit; {type: "disabled"} accepted only at effort high or below - 400 at xhigh/max, and see the disabled-thinking pitfall below Runs adaptive (thinking is on by default - unlike Opus 4.8/4.7) Removed - 400 Removed - 400 low-max (all five)
Opus 4.8 / 4.7 {type: "adaptive"} is the only on-mode; {type: "disabled"} accepted Runs without thinking - set {type: "adaptive"} explicitly Removed - 400 Removed - 400 low/medium/high/xhigh/max
Sonnet 5 {type: "adaptive"} is the only on-mode; {type: "disabled"} accepted Runs adaptive Removed - 400 Removed - 400 low/medium/high/xhigh/max
Opus 4.6 / Sonnet 4.6 {type: "adaptive"} (recommended; auto-enables interleaved thinking, no beta header) Set {type: "adaptive"} explicitly Deprecated - do not use in new code; transitional escape hatch only (see below) Allowed low/medium/high/max (xhigh arrived with Opus 4.7)
Haiku 4.5; older models (Sonnet 4.5, ...) only if explicitly requested {type: "enabled", budget_tokens: N} No thinking Required for thinking; must be less than max_tokens, minimum 1024 - errors otherwise Allowed effort works on Opus 4.5 (low/medium/high only - no xhigh/max); errors on Sonnet 4.5 / Haiku 4.5

Opus 4.8 keeps the same request surface as 4.7 (no new breaking changes) - see shared/model-migration.md -> Migrating to Opus 4.8 for the behavioral re-tuning, and -> Migrating to Opus 4.7 for the full breaking-change list when coming from 4.6 or earlier. With thinking disabled, Opus 4.8 may write longer reasoning into the visible response - leave adaptive thinking on, or add a final-answer-only instruction (see the migration guide).

  • Effort (GA, no beta header): output_config: {effort: "low"|"medium"|"high"|"xhigh"|"max"} - inside output_config, not top-level; default high (equivalent to omitting it). Controls thinking depth and overall token spend; combine with adaptive thinking for the best cost-quality tradeoffs. xhigh (added on Opus 4.7, between high and max) is the best setting for most coding and agentic use cases on Fable 5 / Opus 4.7/4.8 / Sonnet 5, and the default in Claude Code; effort matters more on those models than on any prior model in their tier - re-tune it when migrating, and run long-horizon/agentic tasks at high/xhigh with the full task spec given up front. Use a minimum of high for intelligence-sensitive work, max when correctness matters more than cost, and low for subagents or simple tasks - lower effort means fewer and more-consolidated tool calls, less preamble, and terser confirmations (high is often the sweet spot balancing quality and token efficiency).
  • Choosing an effort level (cost tuning): Effort is the first quality-trading lever, after the free wins (caching first) - it trades thoroughness against token spend within one model, and the top of the range earns its cost only on hard problems (raise to max only when measurement shows headroom at the level below). Which workloads repay higher effort is a property of the workload: coding and long-horizon agentic work respond strongly; chat, classification, and high-volume or latency-sensitive routes often don't and do well at low, with medium as the cost-saving step-down where quality holds (the per-level defaults above cover the rest). Measure on a sample of real requests before raising a default, and tune per route rather than globally. Before building a multi-model cost cascade, measure the simpler alternative first - the most capable model at lower effort on the same tasks: lower effort on the newest models often matches or exceeds prior-generation performance at high effort (on Fable 5, lower effort often exceeds xhigh on prior models), and one model means one cache namespace (caches are model-scoped, so a cascade forfeits cache reuse across its models; a mid-conversation top-level effort change still invalidates the messages cache, though the per-message effort system message avoids that on Claude Fable 5.1 / Claude Mythos 5.1 / Claude Opus 5 - shared/prompt-caching.md § Invalidation hierarchy). Judge cost per completed task, not per request - a cheaper request that needs more turns or retries to finish the job isn't cheaper. For the measured effort/cost tradeoffs by workload and the full lever order, shared/cost-optimization.md § 2.6.
  • Thinking display - "omitted" by default on Fable 5 / Claude Fable 5.1 / Mythos 5 / Claude Mythos 5.1 / Opus 5 / 4.8 / 4.7 / Sonnet 5: display: "summarized" returns a readable summary of the reasoning; "omitted" (the default on all eight - a silent change from Opus 4.6 and Sonnet 4.6, where it was "summarized") streams thinking blocks with empty text. display controls visibility only - thinking happens and is billed the same under every setting; the raw chain of thought is never exposed on any model. If you stream reasoning to users, the default looks like a long pause before output - set thinking: {type: "adaptive", display: "summarized"} explicitly. (Independent of display, echo thinking blocks back unchanged when continuing on the same model; other models silently ignore them (Claude Fable 5.1 / Claude Mythos 5.1 read them) - see the migration guide.) On Claude Fable 5.1 / Claude Mythos 5.1 / Claude Fable 5, display: "updates" (beta thinking-display-updates-2026-08-18, every platform) hides reasoning like "omitted" but returns the model's between-tool-call progress notes as short thinking block summaries - see shared/model-migration.md -> Migrating to Claude Fable 5.1 from Claude Fable 5 -> New API features.
  • When the user asks for "extended thinking", a "thinking budget", or budget_tokens: always use Fable 5/5.1, Opus 5, 4.8, 4.7, or 4.6 with thinking: {type: "adaptive"} - the fixed thinking-token-budget concept is deprecated and adaptive thinking replaces it. Do NOT use budget_tokens for new 4.6/4.7/4.8 code and do NOT switch to an older model just because the user mentions it. Gradual-migration carve-out: budget_tokens is still functional on Opus 4.6 and Sonnet 4.6 only, as a transitional escape hatch for existing code that needs a hard token ceiling before you've tuned effort - see shared/model-migration.md -> Transitional escape hatch. It is fully removed on Fable 5/5.1, Opus 5/4.7/4.8, and Sonnet 5.

Compaction (Quick Reference)

Beta, Fable 5/5.1, Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5, and Sonnet 4.6. For long-running conversations that may exceed the 1M context window, enable server-side compaction. The API automatically summarizes earlier context when it approaches the trigger threshold (default: 150K tokens). Requires beta header compact-2026-01-12.

Critical: Append response.content (not just the text) back to your messages on every turn. Compaction blocks in the response must be preserved - the API uses them to replace the compacted history on the next request. Extracting only the text string and appending that will silently lose the compaction state.

See {lang}/claude-api/README.md (Compaction section) for code examples. Full docs via WebFetch in shared/live-sources.md.


Prompt Caching (Quick Reference)

Prefix match. Any byte change anywhere in the prefix invalidates everything after it. Render order is tools -> system -> messages. Keep stable content first (frozen system prompt, deterministic tool list), put volatile content (timestamps, per-request IDs, varying questions) after the last cache_control breakpoint.

Mid-conversation operator instructions (Claude Opus 5, Claude Opus 4.8, Claude Fable 5, Claude Fable 5.1, Claude Mythos 5, Claude Mythos 5.1; not Claude Sonnet 5; no beta header): append {"role": "system", ...} to messages[] instead of editing top-level system. Preserves the cached history prefix and is the prompt-injection-safe operator channel. See shared/prompt-caching.md § Mid-conversation system messages.

Top-level auto-caching (cache_control: {type: "ephemeral"} on messages.create()) is the simplest option when you don't need fine-grained placement. Max 4 breakpoints per request. Minimum cacheable prefix is model-dependent (512-4096 tokens - see shared/prompt-caching.md § API reference) - shorter prefixes silently won't cache.

Verify with usage.cache_read_input_tokens - if it's zero across repeated requests, a silent invalidator is at work (datetime.now() in system prompt, unsorted JSON, varying tool set).

For placement patterns, architectural guidance, and the silent-invalidator audit checklist: read shared/prompt-caching.md. Language-specific syntax: {lang}/claude-api/README.md (Prompt Caching section).


Fast Mode (Quick Reference)

Research preview, Claude Opus 5 / Opus 4.8 only - Claude API and Managed Agents, not Bedrock / Google Cloud / Foundry. Opus 4.7 fast mode has been removed: speed: "fast" on 4.7 returns an error. Fast mode on Claude Opus 5 is priced at $10 / $50 per MTok. Fast mode runs the same model at up to 2.5x higher output tokens per second, at premium pricing. Three things are required on every request: use the beta messages endpoint (client.beta.messages....), pass the beta flag fast-mode-2026-02-01, and set speed: "fast" as a top-level request parameter (not a header, not in extra_body).

client.beta.messages.create(
    model="claude-opus-5", max_tokens=4096,
    speed="fast", betas=["fast-mode-2026-02-01"],
    messages=[...],
)
Language Beta flag Speed parameter
Python betas=["fast-mode-2026-02-01"] speed="fast"
TypeScript / Ruby betas: ["fast-mode-2026-02-01"] speed: "fast"
Go []anthropic.AnthropicBeta{anthropic.AnthropicBetaFastMode2026_02_01} Speed: anthropic.BetaMessageNewParamsSpeedFast
Java .addBeta(AnthropicBeta.FAST_MODE_2026_02_01) .speed(MessageCreateParams.Speed.FAST)
C# Betas = ["fast-mode-2026-02-01"] Speed = Speed.Fast (Anthropic.Models.Beta.Messages)
PHP betas: ['fast-mode-2026-02-01'] speed: 'fast'
cURL anthropic-beta: fast-mode-2026-02-01 header "speed": "fast" in body

response.usage.speed reports which speed was used. Fast mode has its own rate limit separate from standard Opus; on 429, either retry after the retry-after delay or drop speed and fall back to standard (note: switching speed invalidates prompt cache). Not available with Batch API, Priority Tier, Claude Platform on AWS, or third-party platforms.

Priority Tier is not supported on every current model. It is supported on Claude Fable 5, Opus 4.8, and the older current models, but Claude Opus 5, Claude Sonnet 5, Claude Fable 5.1, Claude Mythos 5.1, Claude Mythos 5, and Mythos Preview are excluded - a Priority Tier request naming one of them fails validation.


Task Budgets (Quick Reference)

Beta, Claude Opus 5 / Fable 5 / Claude Fable 5.1 (confirm at launch) / Sonnet 5 / Opus 4.8 / 4.7. A task budget gives Claude a token ceiling for an agentic loop so it paces itself and finishes gracefully instead of being cut off - distinct from max_tokens, which is an enforced per-response ceiling the model is not aware of. Minimum total: 20,000. Set task_budget inside output_config on client.beta.messages.stream(...) with beta flag task-budgets-2026-03-13 - use streaming so the large max_tokens doesn't hit HTTP timeouts (full details: shared/model-migration.md -> Task Budgets):

with client.beta.messages.stream(
    model="claude-opus-5", max_tokens=128000,
    output_config={"effort": "high", "task_budget": {"type": "tokens", "total": 64000}},
    betas=["task-budgets-2026-03-13"],
    messages=[...], tools=[...],
) as stream:
    response = stream.get_final_message()

task_budget fields: type (always "tokens"), total, and optional remaining (defaults to total). The server injects a countdown marker Claude sees during generation; the budget counts what Claude generates and the tool results it reads this turn - not the full history you resend each request. Not the same thing as Managed Agents session budgets - those are hard, dollar-denominated, platform-enforced caps on one CMA session (shared/managed-agents-core.md § Session budgets); a task budget is advisory and token-denominated.

Observing spend: accumulate response.usage.output_tokens (plus the token count of the tool-result blocks you append) across loop iterations if you want to display progress. Leave remaining unset in the normal loop - the server tracks the countdown itself, and passing a client-computed remaining while also resending full history under-reports the budget. Only pass remaining when you compact or rewrite history between requests and the server can no longer derive prior spend.


Provider Clients (Quick Reference)

When targeting Claude on a third-party platform, use that platform's dedicated client class - not the first-party Anthropic() client with a base_url override. After construction the client exposes the same messages.create / .stream surface as the first-party SDK.

Amazon Bedrock

Use the Mantle client (Messages-API Bedrock endpoint). Bedrock model IDs take an anthropic. prefix (e.g. "anthropic.claude-opus-5"). Region is required.

Language Client
Python from anthropic import AnthropicBedrockMantle -> AnthropicBedrockMantle(aws_region="...")
TypeScript import { AnthropicBedrockMantle } from "@anthropic-ai/bedrock-sdk" -> new AnthropicBedrockMantle({ awsRegion: "..." })
Go bedrock.NewMantleClient(ctx, bedrock.MantleClientConfig{ AWSRegion: "..." })
Java AnthropicOkHttpClient.builder().backend(BedrockMantleBackend.fromEnv()).build() (from com.anthropic.bedrock.backends)
C# new AnthropicBedrockMantleClient(new() { AwsRegion = "..." }) (package Anthropic.Bedrock)
PHP use Anthropic\Bedrock\MantleClient; -> new MantleClient(awsRegion: '...')
Ruby Anthropic::BedrockMantleClient.new(aws_region: "...")

AnthropicBedrock / BedrockClient / BedrockBackend (without Mantle) are the legacy bedrock-runtime InvokeModel path - prefer the Mantle client for new code.

Microsoft Foundry

Language Client
Python from anthropic import AnthropicFoundry -> AnthropicFoundry(api_key=..., resource="...")
TypeScript import AnthropicFoundry from "@anthropic-ai/foundry-sdk" -> new AnthropicFoundry({ ... })
Java AnthropicOkHttpClient.builder().backend(FoundryBackend.fromEnv()).build() (from com.anthropic.foundry.backends)
C# new AnthropicFoundryClient(new AnthropicFoundryApiKeyCredentials(...)) (package Anthropic.Foundry)
PHP Foundry\Client::withCredentials(...)

The Go and Ruby SDKs do not currently support Foundry. For Ruby, use the standard Anthropic::Client.new(base_url: "<foundry endpoint>") as a fallback (Entra ID auth is not built in). For Claude Platform on AWS, see shared/claude-platform-on-aws.md.

Google Cloud Vertex AI

Two required constructor args: GCP project_id and region. Vertex model IDs take no prefix - current-generation models (Opus 4.8/4.7/4.6, Sonnet 5, Sonnet 4.6) use the bare first-party ID (e.g. "claude-opus-5"); dated-snapshot models use an @ version separator (e.g. claude-opus-4-5@20251101, not claude-opus-4-5-20251101). Auth is GCP ADC (gcloud auth application-default login); no Anthropic API key. region can be "global" (recommended), a multi-region ("us"/"eu"), or a specific region. After construction, use the same messages.create / .stream surface.

Language Client
Python from anthropic import AnthropicVertex -> AnthropicVertex(project_id="...", region="...") (install "anthropic[vertex]")
TypeScript import { AnthropicVertex } from "@anthropic-ai/vertex-sdk" -> new AnthropicVertex({ projectId, region })
Go import "github.com/anthropics/anthropic-sdk-go/vertex" -> anthropic.NewClient(vertex.WithGoogleAuth(ctx, region, projectID))
Java AnthropicOkHttpClient.builder().backend(VertexBackend.builder().region("...").project("...").build()).build() (from com.anthropic.vertex.backends)
C# new AnthropicClient { Backend = new VertexBackend(projectId, region) } (package Anthropic.Vertex)
PHP use Anthropic\Vertex; -> Vertex\Client::fromEnvironment(location: '...', projectId: '...') - note location, not region
Ruby Anthropic::VertexClient.new(region: "...", project_id: "...")

Context Editing (Quick Reference)

Beta. Context editing clears old tool results or thinking blocks from the conversation before the model sees it; it is not compaction (which summarizes). On client.beta.messages.* with beta context-management-2025-06-27, pass context_management.edits with a strategy type:

client.beta.messages.create(
    model="claude-opus-5", max_tokens=4096,
    betas=["context-management-2025-06-27"],
    context_management={"edits": [{"type": "clear_tool_uses_20250919"}]},
    tools=[...], messages=[...],
)

Strategy types: clear_tool_uses_20250919 (clears old tool results; optional clear_tool_inputs: true also clears the tool_use params) and clear_thinking_20251015 (clears thinking blocks). Do not use compact_20260112 or beta compact-2026-01-12 - those are the separate compaction feature.


Mid-Conversation System Messages (Quick Reference)

Claude Opus 5, Claude Opus 4.8, Claude Fable 5, Claude Fable 5.1, Claude Mythos 5, and Claude Mythos 5.1; not Claude Sonnet 5; no beta header. Append {"role": "system", "content": "..."} to the messages array (not the top-level system field) to add an operator instruction mid-conversation without invalidating the cached prefix. Use the regular client.messages.create - there is no beta. A mid-conversation system message must follow a user message (or an assistant message ending in server-tool use), and must be either the last entry in messages or be followed by an assistant turn - it cannot be messages[0]. Availability: shared/platform-availability.md. See shared/prompt-caching.md § Mid-conversation system messages. A beta extension shipped with Claude Fable 5.1: output_config: {effort: ...} with content: [] changes effort from that point on without a cache reset (beta mid-conversation-output-config-2026-07-01; Claude Fable 5.1, Claude Mythos 5.1, Claude Opus 5; Claude API). An effort-only message (empty content) is exempt from the placement rules above - it can sit anywhere in messages, including first or between an assistant turn and the next user turn; the rules apply to text and clear_at messages. For a per-turn reminder, give the message clear_at: "next_user_message" (beta mid-conversation-system-clear-at-2026-08-21): it renders for one turn, then stays in the transcript cleared - never delete earlier copies (on Claude Fable 5.1 deleting one invalidates later thinking blocks); without the beta, a text block after the tool results, earlier copies kept. See shared/model-migration.md -> Migrating to Claude Fable 5.1 from Claude Fable 5 -> New API features.


Managed Agents (Beta)

Managed Agents is a third surface: server-managed stateful agents with Anthropic-hosted tool execution. You create a persisted, versioned Agent config (POST /v1/agents), then start Sessions that reference it. Each session provisions a container as the agent's workspace - bash, file ops, and code execution run there; the agent loop itself runs on Anthropic's orchestration layer and acts on the container via tools. The session streams events; you send messages and tool results back.

Availability: shared/platform-availability.md. For agents on Bedrock / Vertex / Foundry (where Managed Agents is unsupported), use Claude API + tool use.

Mandatory flow: Agent (once) -> Session (every run). model/system/tools live on the agent, never the session. See shared/managed-agents-overview.md for the full reading guide, beta headers, and pitfalls.

Beta headers: managed-agents-2026-04-01 - the SDK sets this automatically for all client.beta.{agents,environments,sessions,vaults,memory_stores,deployments,deployment_runs}.* calls. Files API and Skills API are out of beta - no beta header needed (see the API Drift table above for the migration guides).

Subcommands - invoke directly with /claude-api <subcommand>:

Subcommand Action
managed-agents-onboard Walk the user through setting up a Managed Agent from scratch. Read shared/managed-agents-onboarding.md immediately and follow its interview script: describe -> configure the agent (propose, don't interrogate) -> environment -> session (same arc as the Console quickstart, auth deferred to the session step) - defaults and inline suggestions do the work, with a silent viability gate (job vs tools/credentials/data) before any code is emitted. Do not summarize - run the interview.

Reading guide: Start with shared/managed-agents-overview.md, then the topical shared/managed-agents-*.md files (core, environments, tools, events, outcomes, multiagent, webhooks, memory, scheduled-deployments, client-patterns, onboarding, api-reference). For Python, TypeScript, Go, Ruby, PHP, and Java, read {lang}/managed-agents/README.md for code examples. For cURL, read curl/managed-agents.md. Agents are persistent - create once, reference by ID. Define agents and environments as version-controlled YAML applied with the ant CLI - this is the recommended flow (see shared/anthropic-cli.md): the CLI owns the control plane (creating and updating agents), your code owns the data plane (sessions.create with the stored agent ID). Call agents.create() in code only when you must provision programmatically; either way, store the returned agent ID and pass it to every subsequent sessions.create; never call agents.create() in the request path. If a binding you need isn't shown in the language README, WebFetch the relevant entry from shared/live-sources.md rather than guess. C# has beta Managed Agents support via client.Beta.Agents and related namespaces - see csharp/claude-api/README.md for details, or curl/managed-agents.md for raw HTTP reference.

When the user wants to set up a Managed Agent from scratch (e.g. "how do I get started", "walk me through creating one", "set up a new agent"): read shared/managed-agents-onboarding.md and run its interview - same flow as the managed-agents-onboard subcommand.

When the user asks "how do I write the client code for X": reach for shared/managed-agents-client-patterns.md - covers lossless stream reconnect, processed_at queued/processed gate, interrupt, tool_confirmation round-trip, the correct idle/terminated break gate, post-idle status race, stream-first ordering, file-mount gotchas, etc. For credentials, lead with vault environment_variable credentials - the first-class mechanism; secrets are substituted at egress and never enter the sandbox (shared/managed-agents-tools.md -> Vaults). Keeping credentials host-side via custom tools is the fallback where vault credentials don't fit (e.g. self-hosted sandboxes).

When the task is a deliverable - default the kickoff to an outcome, not a plain message. If the session's job is to produce something checkable (an artifact, a report, a PR, a dataset, a fixed set of changes), read shared/managed-agents-outcomes.md and kick off with user.define_outcome plus a starter rubric you draft from the task (5-10 concrete, independently gradeable criteria; comment it as a starter to tune). Reserve plain user.message for genuinely conversational sessions. Trigger on intent, not just the word: "keep working until it's right", "make sure the output is actually good", "don't stop at a first draft" all mean outcomes.

When the user asks about tool approvals, permission policies, or "auto mode" (which tool calls need a human, letting the server evaluate calls, evaluated_permission / evaluation on tool-use events): read shared/managed-agents-tools.md § Permission Policies - always_allow / always_ask / auto and the three auto outcomes (runs, denied as high-risk, pauses when indeterminate). For attaching a terminal to a live session (ant beta:sessions connect): shared/anthropic-cli.md.

When the user wants the agent to run on a schedule (cron, "every night", "weekly report"): read shared/managed-agents-scheduled-deployments.md - deployments fire sessions autonomously on a cron cadence, with per-firing run records and lifecycle controls (pause/unpause/archive).

When the agent's work fans out (research across several sources, per-file or per-record work, "look into N things, then summarize") or one loop would fill its context with reading: read shared/managed-agents-multiagent.md and recommend a multiagent session - start with just {"type": "self"} in the roster so the agent can delegate to copies of itself, then move reading-heavy sub-tasks to a cheaper worker agent (e.g. Claude Haiku 4.5, or Claude Sonnet 5 when the worker needs more judgment) referenced by ID.


Server Tools (Quick Reference)

Server-side tools run on Anthropic's infrastructure - no client-side execution loop. Declare in tools; results arrive as content blocks in the same response. No beta header unless noted. Prefer the latest type variant your model supports. The _20260209 web search / web fetch variants below (dynamic filtering) require Opus 5/4.8/4.7/4.6, Sonnet 5, or Sonnet 4.6; the basic variants for older models are listed after the table.

Tool type name Key optional params Result block type
Web search web_search_20260209 web_search max_uses, allowed_domains/blocked_domains, user_location web_search_tool_result -> .content is a list of web_search_result
Web fetch web_fetch_20260209 web_fetch max_uses, allowed_domains/blocked_domains, citations, max_content_tokens web_fetch_tool_result -> .content is a web_fetch_result with a document block
Code execution code_execution_20260521 code_execution none bash_code_execution_tool_result -> .content.stdout / .stderr / .return_code
Tool search (regex) tool_search_tool_regex_20251119 tool_search_tool_regex mark other tools defer_loading: true tool_search_tool_result
Tool search (BM25) tool_search_tool_bm25_20251119 tool_search_tool_bm25 mark other tools defer_loading: true tool_search_tool_result

web_search_20260209 / web_fetch_20260209 have built-in dynamic filtering - code execution runs under the hood, so do not separately declare code_execution in tools (a second execution environment confuses the model). For models older than Opus 4.6 / Sonnet 4.6, use the basic variants web_search_20250305 / web_fetch_20250910 instead; on Vertex AI only basic web_search_20250305 is available. code_execution_20260120 (REPL persistence + programmatic tool calling) runs on Opus 4.5+ / Sonnet 4.5+. Go SDK only: code_execution_20260521 lives under client.Beta.Messages.New with Betas: []anthropic.AnthropicBeta{"code-execution-2025-08-25"} (other languages use plain client.messages.create); code_execution_20260120 uses the non-beta client.Messages.New in Go like everywhere else. Web fetch only fetches URLs already present in the conversation. Provider availability varies by tool - see shared/platform-availability.md. See shared/tool-use-concepts.md for pause_turn handling.

Document & File Input (Quick Reference)

PDF (base64, no beta): {"type": "document", "source": {"type": "base64", "media_type": "application/pdf", "data": <b64 string>}} in user content, placed before the text block. Base64 string must have no newlines. Limits: 32 MB request, 600 pages (100 for 200k-context models). Java: ContentBlockParam.ofDocument(DocumentBlockParam... Base64PdfSource.builder().data(...)).

Files API (no beta): upload via client.files.upload(...) -> response id is the file_id. Reference it as {"type": "document", "source": {"type": "file", "file_id": "..."}} for PDF/text, or {"type": "image", ...} for images - the content-block type must match the file's MIME type. To migrate code off files-api-2025-04-14, WebFetch the Files API row in shared/live-sources.md. Availability: shared/platform-availability.md.

Citations (no beta): set citations: {enabled: true} on each document content block (all or none). Response splits into multiple text blocks; cited blocks carry a citations array. Each citation has cited_text, document_index, document_title, and a location by type: char_location (start_char_index/end_char_index) for plain text, page_location (start_page_number/end_page_number, 1-indexed) for PDF, content_block_location for custom content. Incompatible with output_config.format (returns a 400).

Tool Use Patterns (Quick Reference)

Strict tool use (no beta): set strict: true as a top-level field on the tool definition (alongside name/description/input_schema), not on tool_choice. Schema must have additionalProperties: false + required. Guarantees tool_use.input validates exactly. Go: Strict: anthropic.Bool(true) + additionalProperties via InputSchema.ExtraFields; Java: .strict(true) + .putAdditionalProperty("additionalProperties", JsonValue.from(false)).

Parallel tool use (default on): one assistant message may contain multiple tool_use blocks. Execute them concurrently, then return all tool_result blocks in a single user message - splitting them across multiple messages silently trains Claude to stop making parallel calls. For a failed tool, return tool_result with is_error: true - don't drop it.

Tool Runner (SDK beta helper): drives the tool-call loop for you via client.beta.messages.*. Python: @beta_tool decorator + client.beta.messages.tool_runner(...) -> runner.until_done(). TypeScript: betaZodTool({...}) from @anthropic-ai/sdk/helpers/beta/zod + client.beta.messages.toolRunner(...) -> await runner. Go: toolrunner.NewBetaToolFromJSONSchema(...) + client.Beta.Messages.NewToolRunner(...) -> .RunToCompletion(ctx). Java requires .addBeta("structured-outputs-2025-11-13"). Ruby: Anthropic::BaseTool subclass + client.beta.messages.tool_runner(...). PHP: BetaRunnableTool + ->toolRunner(...). C#: raw JSON-schema tools + BetaToolRunner via client.Beta.Messages.ToolRunner(...).

Programmatic tool calling (no beta header): Claude calls your custom tool from inside code execution. Add {"type": "code_execution_20260120", "name": "code_execution"} and set "allowed_callers": ["code_execution_20260120"] on your custom tool. Opus 4.5+ / Sonnet 4.5+ (availability: shared/platform-availability.md). When responding to a pending programmatic call, the user message must contain only tool_result blocks (no text). Not compatible with strict: true, disable_parallel_tool_use, forced tool_choice, or MCP tools.

Other API Surfaces (Quick Reference)

Message Batches (no beta; availability: shared/platform-availability.md): client.messages.batches.create(requests=[{custom_id, params}, ...]) -> poll client.messages.batches.retrieve(id).processing_status until "ended" -> stream client.messages.batches.results(id). Each result has .custom_id + .result.type (succeeded/errored/canceled/expired); on success read .result.message.content. Python wraps requests as Request(custom_id=..., params=MessageCreateParamsNonStreaming(...)). Results arrive in any order - key by custom_id, never by position.

Models API (no beta; availability: shared/platform-availability.md): client.models.list() (auto-paginates) and client.models.retrieve("claude-opus-5"). Each model object has id, display_name, created_at, and - since Mar 2026 - max_input_tokens (the context window), max_tokens (the output cap), and capabilities. There is no context_window field.

Stop details (GA, Opus 4.7+): response.stop_details is populated only when stop_reason == "refusal" (fields: type: "refusal", category - an open set, e.g. "cyber", "bio", "reasoning_extraction", "frontier_llm", or null; see the docs for the full list - and explanation). It is null for every other stop_reason (end_turn, max_tokens, tool_use, pause_turn, ...) - always guard before reading.

Admin API (beta, since 2026-08-26): organization management - members, invites, workspaces and workspace members, API keys, rate limit reports, service accounts, federation issuers/rules, CMEK external keys - under client.beta.organization in all seven SDKs and ant beta:organization in the CLI. Requires an admin credential: an Admin API key (sk-ant-admin..., read from ANTHROPIC_API_KEY) or an org:admin OAuth token (ANTHROPIC_AUTH_TOKEN); regular API keys are rejected. Usage and cost reports and the Claude Enterprise user-management/analytics endpoints are not in the SDKs - raw HTTP only. See shared/admin-api.md.

Client config (no beta): timeout default 10 min; units differ by SDK - Python/Ruby: seconds; TypeScript: milliseconds; Go option.WithRequestTimeout(time.Duration); Java Duration; C# TimeSpan. TS scales the default up to 60 min for large max_tokens on non-streaming requests; Java does so for streaming requests (Java non-streaming scales 30s-10 min). max_retries/maxRetries default 2 (retries 408/409/429/5xx + connection errors). base_url (or ANTHROPIC_BASE_URL env). Per-request override: Python client.with_options(timeout=5.0).messages.create(...); TS client.messages.create({...}, {timeout: 5_000}); Ruby request_options: {timeout: 5}. Timeouts are retried - wall-clock can reach timeout × (max_retries+1).

Workload Identity Federation (Quick Reference)

GA, no beta header. Construct the normal zero-arg client (Anthropic() / new Anthropic() / anthropic.NewClient() / AnthropicOkHttpClient.fromEnv()); the SDK auto-detects WIF when all of ANTHROPIC_FEDERATION_RULE_ID, ANTHROPIC_ORGANIZATION_ID, ANTHROPIC_SERVICE_ACCOUNT_ID, and ANTHROPIC_IDENTITY_TOKEN_FILE (or ANTHROPIC_IDENTITY_TOKEN) are set, exchanges the JWT at /v1/oauth/token, and auto-refreshes. ANTHROPIC_WORKSPACE_ID does not gate activation - required only when the federation rule spans multiple workspaces (else 400 workspace_id_required), optional for single-workspace rules. ANTHROPIC_API_KEY or ANTHROPIC_AUTH_TOKEN (even empty) outrank WIF, and a set ANTHROPIC_PROFILE also wins over the federation env vars (a missing named profile is an error, not a fall-through) - unset all three.


Reading Guide

After detecting the language, read the relevant files based on what the user needs. Every {lang}/..., shared/..., and curl/... path cited in this document is relative to this skill's base directory, and none of those files' content is included above - Read each one on demand before relying on what it covers.

All SDK languages use the same multi-file layout - directory {lang}/claude-api/ containing README.md (install, client init, basic request, thinking, caching, stop details, misc), tool-use.md (tool definitions, agentic loop, Anthropic-defined tools, structured outputs), streaming.md, batches.md, files-api.md. Not every language has every file (e.g., Ruby has no batches.md); if a file is absent, that feature's example is not yet documented for that language - fall back to the cURL shape or WebFetch the SDK repo from shared/live-sources.md. cURL -> curl/examples.md.

The Quick Task Reference below uses the {lang}/claude-api/FILE.md path notation for all languages.

Quick Task Reference

Single text classification/summarization/extraction/Q&A:
-> Read only {lang}/claude-api/README.md - always read the README first for any task (installation, quick start, common patterns, error handling)

Chat UI or real-time response display:
-> Read {lang}/claude-api/README.md + {lang}/claude-api/streaming.md

Long-running conversations (may exceed context window):
-> Read {lang}/claude-api/README.md - see Compaction section
Migrating to a newer model (Fable 5.1 / Fable 5 / Opus 5 / Opus 4.8 / Opus 4.7 / Opus 4.6 / Sonnet 5 / Sonnet 4.6), replacing a retired model, or translating budget_tokens / prefill patterns to the current API:
-> Read shared/model-migration.md
Upgrading the Anthropic SDK package itself across a major version (anthropic 0.x -> 1.x: httpx2, awaited async .with_raw_response, removed deprecated parameters / aliases / Text Completions, Python >= 3.10) - or writing new code against a project already on 1.x:
-> Read {lang}/claude-api/sdk-upgrade.md (currently Python only; other SDKs have no bundled major-version guide yet - use that SDK's CHANGELOG via shared/live-sources.md)
Building an eval set for a Claude app (or "how do I know if my change helped"):
-> Read shared/evals/build-eval.md - it loads shared/evals/eval-audit.md (the health checklist every eval must satisfy) before Step 0.
Checking whether an existing eval is trustworthy ("is my eval any good?"):
-> Read shared/evals/eval-audit.md and run it against the eval; report per its section 6.
Iteratively improving an app against an eval (prompt tuning, hill-climbing):
-> Read shared/evals/eval-hillclimb.md - runs Step 0 -> Step 5 with a train/test split; test is scored every round and is the headline.
Rendering an eval-hillclimb HTML report:
-> Run shared/evals/report/build-report.mjs when it is on disk (EAP install), else shared/evals/report/build-report-lite.mjs (always extracted with this skill) - both consume the _state.json / vN/ layout produced by the hillclimb guide and write the same trajectory/scores.tsv. Don't write a parallel one.
Prompting or tuning Fable 5/5.1 (long turns, effort, verbosity, autonomous runs, sub-agents):
-> Read shared/model-migration.md -> Migrating to Claude Fable 5.1 -> Behavioral shifts (prompt-tunable) + Long-running agent recommendations
Prompting or tuning Claude Fable 5.1 (progress updates, parallel tool calls, writing density / formatting, autonomy, test sprawl, whole-file rewrites) or making a harness compatible with preserved thinking's history-editing check (history edits, compaction, per-turn reminders):
-> Read shared/model-migration.md -> Migrating to Claude Fable 5.1 from Claude Fable 5 -> New API features + Behavioral shifts (prompt-tunable); for the history-editing check itself (the three-step check, the append-only edit table, compaction shapes), Breaking change 3 in the same section
Prompt caching / optimize caching / "why is my cache hit rate low":
-> Read shared/prompt-caching.md (prefix-stability design, breakpoint placement, anti-patterns that silently invalidate cache) + {lang}/claude-api/README.md (Prompt Caching section)
Auditing or cleaning up prompts, skills, or tool descriptions ("is this prompt outdated", "remove the cruft", "this was written for an older model"):
-> Read shared/prompt-audit.md - dated-pattern tables with greppable signals, the keep list (what NOT to delete), and the report + proposed-diff output contract
Count tokens in a file / prompt / diff ("how many tokens is X"):
-> Read shared/token-counting.md - use messages.count_tokens, never tiktoken
Reducing or reviewing API spend ("the bill is too high", "make this cheaper", "am I overspending", cost per completed task, cheapest model or effort that holds quality):
-> Read shared/cost-optimization.md - baseline and token profile first, then the levers in order (free wins before tradeoffs) with measured expectations, and a workload-shape -> lever mapping table

Function calling / tool use / agents:
-> Read {lang}/claude-api/README.md + shared/tool-use-concepts.md (conceptual foundations: function calling, code execution, memory, structured outputs) + {lang}/claude-api/tool-use.md (language-specific code examples: tool runner, manual loop, code execution, memory, structured outputs)

Agent design (tool surface, context management, caching strategy):
-> Read shared/agent-design.md (bash vs. dedicated tools, programmatic tool calling, tool search/skills, context editing vs. compaction vs. memory, caching principles)

Batch processing (non-latency-sensitive; runs asynchronously at 50% cost):
-> Read {lang}/claude-api/README.md + {lang}/claude-api/batches.md

File uploads across multiple requests (same file without re-uploading):
-> Read {lang}/claude-api/README.md + {lang}/claude-api/files-api.md

Organization administration (members, invites, workspaces, API keys, rate limit reports, service accounts, WIF resources, CMEK):
-> Read shared/admin-api.md - client.beta.organization endpoint/method table, admin credentials, per-language naming and pagination, what stays curl-only

Debugging HTTP errors or implementing error handling:
-> Read shared/error-codes.md - per-SDK typed exception class table and the Go errors.As pattern

Latest official documentation:
-> WebFetch the URLs in shared/live-sources.md

Managed Agents (server-managed stateful agents with workspace):
-> See the reading guide in the ## Managed Agents (Beta) section above - it lists every shared/managed-agents-*.md file and the language-specific READMEs ({lang}/managed-agents/README.md, curl/managed-agents.md).


When to Use WebFetch

Use WebFetch to get the latest documentation when:

  • User asks for "latest" or "current" information
  • Cached data seems incorrect
  • User asks about features not covered here

Live documentation URLs are in shared/live-sources.md.

Common Pitfalls

  • Don't truncate inputs when passing files or content to the API. If the content is too long to fit in the context window, notify the user and discuss options (chunking, summarization, etc.) rather than silently truncating.
  • Prefill removed (Fable 5, Claude Fable 5.1, Opus 5, Sonnet 5, and the 4.6/4.7/4.8 family): Assistant message prefills (last-assistant-turn prefills) return a 400 error on Fable 5, Claude Fable 5.1, Opus 5, Sonnet 5, Opus 4.6, Opus 4.7, Opus 4.8, and Sonnet 4.6. Use structured outputs (output_config.format) or system prompt instructions to control response format instead. (One exception: the fallback-credit prefill claim - when redeeming a credit with fallback_has_prefill_claim: true, the server accepts the echoed assistant message; see the migration guide's refusal section.)
  • Confirm migration scope before editing: When a user asks to migrate code to a newer Claude model without naming a specific file, directory, or file list, ask which scope to apply first - the entire working directory, a specific subdirectory, or a specific set of files. Do not start editing until the user confirms. Imperative phrasings like "migrate my codebase", "move my project to X", "upgrade to Sonnet 4.6", or bare "migrate to Opus 4.8" are still ambiguous - they tell you what to do but not where, so ask. Proceed without asking only when the prompt names an exact file, a specific directory, or an explicit file list ("migrate app.py", "migrate everything under services/", "update a.py and b.py"). See shared/model-migration.md Step 0.
  • max_tokens defaults: Don't lowball max_tokens - hitting the cap truncates output mid-thought and requires a retry. For non-streaming requests, default to ~16000 (keeps responses under SDK HTTP timeouts). For streaming requests, default to ~64000 (timeouts aren't a concern, so give the model room). Only go lower when you have a hard reason: classification (~256), cost caps, deliberately short outputs, or max_tokens: 0 for cache pre-warming (see shared/prompt-caching.md -> Pre-warming).
  • Disabling thinking on Claude Opus 5 has two failure modes - prefer low/medium effort instead. Only affects code that explicitly opts out; thinking is on by default, so watch for a disabled-thinking setting carried forward from Opus 4.8. With thinking: {type: "disabled"}, the model occasionally writes a tool call into its visible text instead of a tool_use block: the turn succeeds, the call never runs, no error is raised, and in an agentic loop that text pollutes later turns. It can also leak <thinking> tags into the response. Turning thinking on and lowering effort fixes both and still cuts cost. If a route must stay thinking-off: delete any don't-think/don't-reason rule (it makes tag leakage worse), don't name thinking tags, and add the combined instruction "When you use a tool, you may say a brief sentence first. If no tool can express what the user asked for, say so instead of guessing. Do not include internal or system XML tags in your response." Details: shared/model-migration.md -> Two failure modes when thinking is disabled.
  • 128K output tokens: Fable 5, Claude Fable 5.1, Opus 5, Opus 4.6, Opus 4.7, Opus 4.8, Sonnet 5, and Sonnet 4.6 support up to 128K max_tokens, but the SDKs require streaming for values that large to avoid HTTP timeouts. Use .stream() with .get_final_message() / .finalMessage().
  • Forced tool use removed (Claude Fable 5.1 / Claude Mythos 5.1, as on Mythos Preview): tool_choice: {type: "any"} and {type: "tool", name: ...} return a 400 (tool_choice: type "tool" and "any" are not supported for this model.), on count_tokens and Batches too. Use {type: "auto"} plus an explicit instruction naming the tool, strict: true on the tool to keep schema-valid arguments, or structured outputs (output_config.format) when the forced call only existed to get JSON back. {type: "none"} is unaffected; disable_parallel_tool_use still works with auto (at most one call).
  • Tool call JSON parsing (Fable 5, Claude Fable 5.1, Opus 5, and the 4.6/4.7/4.8 family): Fable 5, Claude Fable 5.1, Opus 5, Opus 4.6, Opus 4.7, Opus 4.8, and Sonnet 4.6 may produce different JSON string escaping in tool call input fields (e.g., Unicode or forward-slash escaping). Always parse tool inputs with json.loads() / JSON.parse() - never do raw string matching on the serialized input.
  • Structured outputs (all models): Use output_config: {format: {...}} instead of the deprecated output_format parameter on messages.create(). This is a general API change, not 4.6-specific.
  • Don't reimplement SDK functionality: The SDK provides high-level helpers - use them instead of building from scratch. Specifically: use stream.finalMessage() instead of wrapping .on() events in new Promise(); use typed exception classes (Anthropic.RateLimitError, etc.) instead of string-matching error messages; use SDK types (Anthropic.MessageParam, Anthropic.Tool, Anthropic.Message, etc.) instead of redefining equivalent interfaces.
  • Error handling - catch a chain, not one broad class. A single except APIStatusError / catch (AnthropicServiceException) / rescue APIError loses the distinction between retryable (429, >=500, network) and non-retryable (400/404) failures. Write a most-specific-first chain - e.g. NotFoundError -> RateLimitError -> APIStatusError -> APIConnectionError (or the Go equivalent: errors.As into *anthropic.Error then switch apierr.StatusCode { case 404: ...; case 429: ...; default: ... }). Per-language class names and namespaces are in shared/error-codes.md.
  • Don't research SDK types - write first. If a type name isn't shown in the documentation included in this skill, write the code file from the namespace/package tables in the language-specific doc and let the compiler's error point you to the right name. Do not spend turns on WebFetch, SDK-repo clones, or compiling-and-running a separate reflection program to discover type names before writing - produce the source file first, then fix what the compiler reports. A quick strings / jar tf / javap against the installed SDK is acceptable for locating names (it returns in seconds), but don't escalate beyond that. A file with a wrong type name is recoverable; a session spent on discovery with no file written is not.
  • Bash and text editor tools are Anthropic-defined, schema-less. Declare {"type": "bash_20250124", "name": "bash"} / {"type": "text_editor_20250728", "name": "str_replace_based_edit_tool"} - no input_schema. A custom tool with your own schema named "bash" is a different tool. Handler paths and security checks are in shared/tool-use-concepts.md § Client-Side Tools.
  • Advisor tool model pairing. The advisor tool's model must be at least as capable as the request's top-level model - e.g. executor claude-sonnet-5 -> advisor claude-opus-5 or claude-opus-4-8. An invalid pair returns 400. Pairing table (and which advisors return plaintext vs encrypted advisor_redacted_result advice) in shared/tool-use-concepts.md § Advisor. Availability: shared/platform-availability.md.
  • Agent Skills != Managed Agents. To have Claude generate a .pptx/.xlsx/etc. via Agent Skills, call client.beta.messages.create with container={"skills": [...]}, the code_execution_20260521 tool, and the code-execution-2025-08-25 beta (Skills is out of beta - no skills-2025-10-02 header needed). Do not use client.beta.agents / sessions / environments here - those are the Managed Agents surface, not Agent Skills.
  • MCP connector needs both halves. mcp_servers=[{type:"url", url, name}] alone is rejected as a validation error - also add tools=[{type:"mcp_toolset", mcp_server_name:<same name>}] with beta mcp-client-2025-11-20. Availability: shared/platform-availability.md.
  • inference_geo is a direct top-level request parameter - client.messages.create(..., inference_geo="us") / .inferenceGeo("us"). Do not put it in extra_body / putAdditionalBodyProperty. (Messages API only - on Managed Agents, inference_geo instead nests inside the agent's model object, never top-level; see shared/managed-agents-core.md § Pinning inference geography.) Supported on Opus 4.6 / Sonnet 4.6 and later; availability: shared/platform-availability.md. response.usage.inference_geo reports where inference ran.
  • Fine-grained tool streaming is not a beta feature; this skill's default is to turn it on for streaming + client tools (the API itself still defaults to buffered). Set eager_input_streaming: true on the tool definition and call the regular client.messages.stream(...). There is no beta header and no client.beta.* path. Do not also send the legacy fine-grained-tool-streaming-2025-05-14 beta header. Python's @beta_tool(eager_input_streaming=True) accepts it directly; TypeScript's betaZodTool() does not, so spread it on: { ...betaZodTool({...}), eager_input_streaming: true }. With the field on, the API no longer coerces or validates the input, so the accumulated partial_json may be incomplete (max_tokens) or invalid - guard the parse (shared/tool-use-concepts.md -> Eager input streaming).
  • Cache diagnostics is beta. Use client.beta.messages.* with beta cache-diagnosis-2026-04-07. Pass diagnostics: {previous_message_id: null} on the first turn and diagnostics: {previous_message_id: <previous response id>} on subsequent turns; the result is on response.diagnostics. Availability: shared/platform-availability.md.
  • Memory tool type is memory_20250818. Declare {"type": "memory_20250818", "name": "memory"}. Go uses the beta-namespace type {OfMemoryTool20250818: &anthropic.BetaMemoryTool20250818Param{}} on client.Beta.Messages.New; Python/TypeScript/Ruby/PHP/C# use the non-beta client.messages.create; Java has both a non-beta MemoryTool20250818 and a beta tool-runner path. Python/TypeScript provide BetaAbstractMemoryTool / betaMemoryTool helpers for implementing the backend.
  • Use a model the feature actually supports. Some features are restricted to specific model tiers - fast mode is Claude Opus 5 / Opus 4.8 only (and Claude API only), task budgets (Messages API only - Managed Agents session budgets have no model-tier restriction) are Claude Opus 5 / Fable 5 / Claude Fable 5.1 (confirm at launch) / Sonnet 5 / Opus 4.8 / 4.7 only, and the advisor tool requires a valid executor<->advisor pair. If the user's prompt names a model that the feature doesn't support, use a supported model instead and note the substitution in the output.
  • Don't define custom types for SDK data structures: The SDK exports types for all API objects. Use Anthropic.MessageParam for messages, Anthropic.Tool for tool definitions, Anthropic.ToolUseBlock / Anthropic.ToolResultBlockParam for tool results, Anthropic.Message for responses. Defining your own interface ChatMessage { role: string; content: unknown } duplicates what the SDK already provides and loses type safety.
  • Report and document output: For tasks that produce reports, documents, or visualizations, the code execution sandbox has python-docx, python-pptx, matplotlib, pillow, and pypdf pre-installed. Claude can generate formatted files (DOCX, PDF, charts) and return them via the Files API - consider this for "report" or "document" type requests instead of plain stdout text.
  • Server-tool errors don't raise. Web search and web fetch errors return HTTP 200 with a web_search_tool_result / web_fetch_tool_result block whose content is a single error object (e.g. {error_code: "max_uses_exceeded"}) - not a raised exception. For web search, a success content is a list; an error content is an object - branch on that before indexing.
  • Managed Agents web tools ignore the environment's networking. web_search / web_fetch run on Anthropic's servers in cloud and self-hosted environments, and Console org-level web settings apply to the Messages API only. Restrict them per tool with allowed_domains or blocked_domains (never both; 1-64 plain hostnames per list, subdomains covered; IPs, bare TLDs, single-label and localhost-style names rejected on both tools; a path suffix is allowed only on web_search) on the toolset configs entry - shared/managed-agents-tools.md § Web search & web fetch settings.
  • Eval / hillclimb work has dedicated guides: If the user says "hillclimb", "improve my eval score", "iterate on my prompt against an eval", or "build me an eval" - load shared/evals/eval-hillclimb.md or shared/evals/build-eval.md rather than improvising. The bundled HTML report builder is shared/evals/report/build-report.mjs when it is on disk (EAP install), else shared/evals/report/build-report-lite.mjs (always extracted with this skill); don't write a parallel one.
  • Code execution output block type: code_execution_20260521 returns bash_code_execution_tool_result (with .content.stdout), not the legacy bare code_execution_tool_result. Iterate response.content and match on the correct type.
  • Tool search: never defer everything. The search tool itself must not have defer_loading: true, and at least one tool in tools must be non-deferred, or the API returns 400 All tools have defer_loading set.

No project language was auto-detected. Ask the user which language they are using (see Language Detection above), then Read the matching {lang}/claude-api/README.md (or curl/examples.md for cURL/raw HTTP or an unsupported language) from the base directory before anything else.

User Request

MCP connector: calling a remote MCP server (Sanity Context, HTTP/SSE, org API token in Authorization header) directly from the Messages API via mcp_servers, current beta header, field names (url, authorization_token, tool_configuration), and Python SDK usage example.

Agent
Agent
Agent
Agent
Agent

Got what I need — the MCP connector requires two paired fields (mcp_servers + a tools entry of type mcp_toolset) plus a beta header, and it runs server-side, so Claude calls the Knowledge Base tools directly without us writing a manual tool loop. Let me write the agent script now.

Agent
Agent
Agent
Agent
Agent

Now let's set up the environment and test it end-to-end.

Agent
Agent

The virtual environment is ready. For testing, you should run this yourself rather than pasting the Sanity token into our chat — even in a command I run, it'd land in my context, which isn't necessary here.

Open your own terminal and run:

cd /c[REDACTED]/sanity-iam-agent
source venv/Scripts/activate
export SANITY_API_TOKEN="your-context-viewer-token"
export SANITY_MCP_URL="https://api.sanity.io/v1/context/organizations/oyryuduzu/mcp/iam-policy-review-endpoint"
export ANTHROPIC_API_KEY="your-anthropic-api-key"
python agent.py "Is a policy with Resource: '*' on ec2:DescribeInstances a real finding or a false positive?"

(If you're in PowerShell instead of Git Bash, use $env:SANITY_API_TOKEN="..." style instead of export.)

That question is deliberately chosen to test the wildcard-resource false-positive case we built in — a good answer should say it's a false positive because DescribeInstances doesn't support resource-level permissions, citing both the iam-lint rule and the AWS reference. Paste me the output (safe to share — it won't contain your token) and we'll see how the agent actually performs.

You

PS [REDACTED]> cd /c[REDACTED]/sanity-iam-agent
cd : Cannot find path 'C:\c\Users\bimbo\sanity-iam-agent' because it does not exist.
At line:1 char:1
+ cd /c[REDACTED]/sanity-iam-agent
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+ CategoryInfo : ObjectNotFound: (C:\c\Users\bimbo\sanity-iam-agent:String) [Set-Location], ItemNotFoundE
xception
+ FullyQualifiedErrorId : PathNotFound,Microsoft.PowerShell.Commands.SetLocationCommand

PS [REDACTED]> source venv/Scripts/activate
source : The term 'source' is not recognized as the name of a cmdlet, function, script file, or operable program.
Check the spelling of the name, or if a path was included, verify that the path is correct and try again.
At line:1 char:1
+ source venv/Scripts/activate
+ ~~~~~~
+ CategoryInfo : ObjectNotFound: (source:String) [], CommandNotFoundException
+ FullyQualifiedErrorId : CommandNotFoundException

PS [REDACTED]> export SANITY_API_TOKEN="your-context-viewer-token"
export : The term 'export' is not recognized as the name of a cmdlet, function, script file, or operable program.
Check the spelling of the name, or if a path was included, verify that the path is correct and try again.
At line:1 char:1
+ export SANITY_API_TOKEN="your-context-viewer-token"
+ ~~~~~~
+ CategoryInfo : ObjectNotFound: (export:String) [], CommandNotFoundException
+ FullyQualifiedErrorId : CommandNotFoundException

PS [REDACTED]> export SANITY_MCP_URL="https://api.sanity.io/v1/context/organizations/oyryuduzu/mcp/iam-policy-review-endpoint"
export : The term 'export' is not recognized as the name of a cmdlet, function, script file, or operable program.
Check the spelling of the name, or if a path was included, verify that the path is correct and try again.
At line:1 char:1
+ export SANITY_MCP_URL="https://api.sanity.io/v1/context/organizations ...
+ ~~~~~~
+ CategoryInfo : ObjectNotFound: (export:String) [], CommandNotFoundException
+ FullyQualifiedErrorId : CommandNotFoundException

PS [REDACTED]> export ANTHROPIC_API_KEY="your-anthropic-api-key"
export : The term 'export' is not recognized as the name of a cmdlet, function, script file, or operable program.
Check the spelling of the name, or if a path was included, verify that the path is correct and try again.
At line:1 char:1
+ export ANTHROPIC_API_KEY="your-anthropic-api-key"
+ ~~~~~~
+ CategoryInfo : ObjectNotFound: (export:String) [], CommandNotFoundException
+ FullyQualifiedErrorId : CommandNotFoundException

PS [REDACTED]> python agent.py "Is a policy with Resource: '*' on ec2:DescribeInstances a real finding or a false positive?"

Agent

Right — you're in PowerShell, not Git Bash, so the syntax is different. Also, those "your-context-viewer-token" values were placeholders in my example — you need to swap in your actual token and API key. Try this in PowerShell:

cd [REDACTED]\sanity-iam-agent
venv\Scripts\Activate.ps1
$env:SANITY_API_TOKEN="paste-your-actual-context-viewer-token-here"
$env:SANITY_MCP_URL="https://api.sanity.io/v1/context/organizations/oyryuduzu/mcp/iam-policy-review-endpoint"
$env:ANTHROPIC_API_KEY="paste-your-actual-anthropic-api-key-here"
python agent.py "Is a policy with Resource: '*' on ec2:DescribeInstances a real finding or a false positive?"

If PowerShell blocks the Activate.ps1 script with an execution-policy error, run python.exe agent.py ... directly using the venv's own interpreter instead — no activation needed:

venv\Scripts\python.exe agent.py "Is a policy with Resource: '*' on ec2:DescribeInstances a real finding or a false positive?"

Paste the output once it runs (it won't contain your token, so that's safe to share).

You

I will change it to bash

Agent

Sounds good. Once you're in bash, here are the commands again:

cd /c[REDACTED]/sanity-iam-agent
source venv/Scripts/activate
export SANITY_API_TOKEN="your-actual-context-viewer-token"
export SANITY_MCP_URL="https://api.sanity.io/v1/context/organizations/oyryuduzu/mcp/iam-policy-review-endpoint"
export ANTHROPIC_API_KEY="your-actual-anthropic-api-key"
python agent.py "Is a policy with Resource: '*' on ec2:DescribeInstances a real finding or a false positive?"

Paste the output when it runs.

You

bimbo@Abimbola MINGW64 ~
$ cd /c[REDACTED]/sanity-iam-agent
source venv/Scripts/activate
export SANITY_API_TOKEN="your-context-viewer-token"
export SANITY_MCP_URL="https://api.sanity.io/v1/context/organizations/oyryuduzu/mcp/iam-policy-review-endpoint"
export ANTHROPIC_API_KEY="your-anthropic-api-key"
python agent.py "Is a policy with Resource: '*' on ec2:DescribeInstances a real finding or a false positive?"
Traceback (most recent call last):
File "[REDACTED]\sanity-iam-agent\agent.py", line 51, in <module>
print(ask(question))
^^^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\agent.py", line 33, in ask
response = client.beta.messages.create(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\venv\Lib\site-packages\anthropic_utils_utils.py", line 294, in wrapper
return func(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\venv\Lib\site-packages\anthropic\resources\beta\messages\messages.py", line 1167, in create
return self._post(
^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\venv\Lib\site-packages\anthropic_base_client.py", line 1511, in post
return cast(ResponseT, self.request(cast_to, opts, stream=stream, stream_cls=stream_cls))
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\venv\Lib\site-packages\anthropic_base_client.py", line 1195, in request
raise self._make_status_error_from_response(response) from None
anthropic.AuthenticationError: Error code: 401 - {'type': 'error', 'error': {'type': 'authentication_error', 'message': 'invalid x-api-key'}, 'request_id': 'req_011CfBf4CbvDfyLqfGTqRhMa'}
((venv) )
bimbo@Abimbola MINGW64 ~/sanity-iam-agent

Agent

That error is because the commands were run with the literal placeholder text "your-anthropic-api-key" instead of your real key — those were examples for you to swap in, not values to paste as-is.

You'll need two real values:

  1. Your Anthropic API key — if you don't already have one, get it from console.anthropic.com under API Keys (starts with sk-ant-...)
  2. Your Sanity Context Viewer token — the one you saved earlier when we created it under Manage → API → Tokens

Re-run with the actual values substituted:

export SANITY_API_TOKEN="<paste your real Sanity token here>"
export ANTHROPIC_API_KEY="<paste your real Anthropic key here>"
python agent.py "Is a policy with Resource: '*' on ec2:DescribeInstances a real finding or a false positive?"

(SANITY_MCP_URL is already set correctly from before, no need to re-export it unless you closed the terminal.) Go ahead and run it with your real credentials — I won't see them, just paste the output.

You

How do I know my anthropic key?

Agent

Let me check whether you already have Claude API credentials set up on this machine via the ant CLI — that's often already configured if you've used Claude Code before, and would save you from needing to get a separate key.

Agent
Agent

No ant CLI on this machine, so you'll need to get an API key directly. Note: this is a separate, pay-as-you-go API account — distinct from any Claude.ai/Claude Code subscription you have, and it bills per token used (this agent script will make real API calls that cost money, though for a handful of test questions it'll be cents).

Steps:
1. Go to console.anthropic.com and sign in (or create an account if you don't have one)
2. If prompted, add a payment method — the API needs billing set up before it'll issue keys
3. Go to API Keys in the left sidebar → Create Key
4. Copy the key (starts with sk-ant-...) — it's shown once

Once you have it, run the export and test command again with the real value.

Agent

This is set up fine as-is. Suggestions:

  • Name: sanity-iam-agent (so you can identify and revoke it later)
  • Expires: 30 days is fine — it covers you past the Oct 4 submission deadline
  • Linked account / Scope: leave as default (Default workspace)

Click Create key, copy the value immediately (shown once), then re-run the export command with it.

Agent

Auth is working now — this is a billing issue, not a code issue: "Your credit balance is too low to access the Anthropic API." The console account needs prepaid credits before it'll serve requests, separate from adding a payment method.

Go to console.anthropic.com → Plans & Billing → purchase credits (even $5 is plenty for testing this agent — each question costs a fraction of a cent to a few cents). Once that's done, re-run the same command — no code changes needed.

You

Traceback (most recent call last):
File "[REDACTED]\sanity-iam-agent\agent.py", line 51, in <module>
print(ask(question))
^^^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\agent.py", line 33, in ask
response = client.beta.messages.create(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\venv\Lib\site-packages\anthropic_utils_utils.py", line 294, in wrapper
return func(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\venv\Lib\site-packages\anthropic\resources\beta\messages\messages.py", line 1167, in create
return self._post(
^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\venv\Lib\site-packages\anthropic_base_client.py", line 1511, in post
return cast(ResponseT, self.request(cast_to, opts, stream=stream, stream_cls=stream_cls))
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\venv\Lib\site-packages\anthropic_base_client.py", line 1195, in request
raise self._make_status_error_from_response(response) from None
anthropic.BadRequestError: Error code: 400 - {'type': 'error', 'error': {'type': 'invalid_request_error', 'message': "mcp_servers[0] 'iam-policy-review': Authentication error while communicating with MCP server. Please check your authorization token."}, 'request_id': 'req_011CfBgcxhHyhF2a5WfM6ZFy'}
Traceback (most recent call last):
File "[REDACTED]\sanity-iam-agent\agent.py", line 51, in <module>
print(ask(question))
^^^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\agent.py", line 33, in ask
response = client.beta.messages.create(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\venv\Lib\site-packages\anthropic_utils_utils.py", line 294, in wrapper
return func(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\venv\Lib\site-packages\anthropic\resources\beta\messages\messages.py", line 1167, in create
return self._post(
^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\venv\Lib\site-packages\anthropic_base_client.py", line 1511, in post
return cast(ResponseT, self.request(cast_to, opts, stream=stream, stream_cls=stream_cls))
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "[REDACTED]\sanity-iam-agent\venv\Lib\site-packages\anthropic_base_client.py", line 1195, in request
raise self._make_status_error_from_response(response) from None
anthropic.BadRequestError: Error code: 400 - {'type': 'error', 'error': {'type': 'invalid_request_error', 'message': "mcp_servers[0] 'iam-policy-review': Authentication error while communicating with MCP server. Please check your authorization token."}, 'request_id': 'req_011CfBgdEvrecsgU5GMu2Y8N'}
((venv) )
bimbo@Abimbola MINGW64 ~/sanity-iam-agent
$
((venv) )
bimbo@Abimbola MINGW64 ~/sanity-iam-agent

Agent

That's Sanity rejecting the token itself, not a code problem. A few things to check:

  1. Confirm the env var actually holds the real token — run this (safe to share, it only reveals the first few characters):
    bash
    echo "${SANITY_API_TOKEN:0:6}... (length: ${#SANITY_API_TOKEN})"

    If that shows something short or garbled, the export didn't take (e.g. quotes got mangled, or it's still literally your-context-viewer-token).

  2. Double-check the token's permission in Sanity — go back to Manage → API → Tokens → find iam-security-agent-mcp, and confirm it actually shows Context: Viewer as its permission (not blank, not a different scope). It's possible the checkbox/radio selection didn't save the way we intended.

  3. Check the token hasn't been revoked or is scoped to the wrong org — the token must belong to org oyryuduzu (Bims-creator), same org the "iam-policy-review" Knowledge Base and endpoint live in.

Run step 1 first and tell me what it shows — that'll tell us whether this is an export problem or a Sanity-side token problem.

You

((venv) )
bimbo@Abimbola MINGW64 ~/sanity-iam-agent
$ echo "${SANITY_API_TOKEN:0:6}... (length: ${#SANITY_API_TOKEN})"
tokens... (length: 185)
((venv) )
bimb

Agent

That's the problem — a real token shouldn't start with the literal word tokens. It looks like extra text got copied along with the actual secret (maybe a label or page fragment got grabbed in the selection).

Since Sanity only shows the token value once, we can't recover the original — you'll need to create a new one:

  1. Go back to Manage → API → Tokens
  2. Click "Add API token" again, same settings as before (name it iam-security-agent-mcp-2, Context → Viewer)
  3. When the token appears, use the copy icon/button next to it if there is one, rather than manually selecting the text — that avoids grabbing surrounding label text
  4. Immediately export it fresh: bash export SANITY_API_TOKEN="<new token, copied via the copy button>"
  5. Re-run the same echo "${SANITY_API_TOKEN:0:6}..." check first — this time it should show random-looking characters, not a real word
  6. Then re-run python agent.py "..."
You

bimbo@Abimbola MINGW64 ~/sanity-iam-agent
$ echo "${SANITY_API_TOKEN:0:6}... (length: ${#SANITY_API_TOKEN})"
skUw74... (length: 180)
((venv) )

Agent

That looks like a real token now. Go ahead and re-run:

python agent.py "Is a policy with Resource: '*' on ec2:DescribeInstances a real finding or a false positive?"
You

bimbo@Abimbola MINGW64 ~/sanity-iam-agent
$ python agent.py "Is a policy with Resource: '*' on ec2:DescribeInstances a real finding or a false positive?"
I'll start by loading the session context and schema.
Now reading the entries relevant to wildcard resource findings.

Short answer: false positive — and this is the KB's canonical example of one.

What the linter says: WILDCARD_RESOURCE, high severity. It fires on any Allow statement where Resource contains "*", with no exceptions. The rule is explicitly action-blind — it doesn't inspect which action is granted, how broad it is, or whether resource-level permissions are even supported for it (reference/iam_lint_rules, findings/wildcard_resource).

What AWS says: The Service Authorization Reference documents that a large class of actions don't support resource-level permissions at all, and ec2:DescribeInstances is named directly as a canonical example (alongside iam:ListRoles and s3:ListAllMyBuckets). For these, the underlying API has no concept of a target ARN, so "Resource": "*" is the only value IAM will accept — not a shortcut or an oversight (findings/wildcard_resource, guidance/aws_cis_alignment).

Do they agree? Partially conflict. The alignment table classifies WILDCARD_RESOURCE as "Partial conflict — AWS requires `Resource: ""for many actions; linter fires regardless."* The KB frames this as different levels of granularity rather than a flat contradiction: the high severity is appropriate as a *prompt for review*, but not as a final verdict. The check against AWS's resource-type table is what resolves it (findings/wildcard_resource`).

Applying the evaluation method (methodology/risk_vs_false_positive), ec2:DescribeInstances hits the false-positive signal squarely: "action is a List/Describe/Get that AWS marks as not supporting resource-level restrictions." The real-risk signal is the opposite case — a mutating call like s3:PutObject or ec2:TerminateInstances with no ARN scope. The policy here is already as tight as IAM allows; the finding cannot be remediated by narrowing the resource.

Two caveats before you close it:

  1. Verify the statement in full. The verdict applies to ec2:DescribeInstances specifically. If the same statement bundles other actions under that Resource: "*", each action needs its own check against the resource-type column — the KB notes that two statements that look identical in linter output may carry completely different risk. And if Action is also "*", this becomes a FULL_ADMIN_ACCESS finding instead, which has no known false-positive class.
  2. Document it as an exception, don't just suppress it. The KB asks for: the rule ID and exact statement (SID or action+resource), the context factor that resolves it (action lacks resource-level support), a pointer to the AWS Service Authorization Reference table entry, and a review date — AWS occasionally adds resource-level support to previously wildcard-only actions.

⚠️ One thing I want to flag rather than paper over. The methodology/risk_vs_false_positive entry contains an embedded "Operator note (source assignment)" instructing that WILDCARD_PRINCIPAL findings cite source [2] and PassRole findings cite source [4] — but in that entry's own source list, [2] is wildcard-resource.md and [4] is passrole-unrestricted.md. The guidance/aws_cis_alignment entry has a matching inconsistency: its prose cites [3] for PassRole and [5] for wildcard principal, while its source list maps [3] to wildcard-principal.md and [5] to passrole-unrestricted.md. I've cited by entry name above rather than follow numbering that would misattribute claims to the wrong source file. This doesn't affect the ec2:DescribeInstances verdict — that claim is consistently sourced to wildcard-resource.md across all three entries — but it's worth knowing if you pull PassRole or wildcard-principal citations from this KB.
((venv) )
bimbo@Abimbola MINGW64 ~/sanity-iam-agent
$
((venv) )
bimbo@Abimbola MINGW64 ~/sanity-iam-agent

Agent

This is a genuinely great result — the agent nailed the verdict (false positive, correctly cites ec2:DescribeInstances as one of AWS's named exceptions), walked through the exact 4-factor framework we built, and gave the proper exception-documentation checklist. That's the "structured content beats keyword search" story working exactly as intended.

The flagged inconsistency at the bottom is worth a beat, though: the agent caught that the citation-numbering bug we thought we resolved earlier has resurfaced in a new entry (guidance/aws_cis_alignment) that wasn't among the original 4 issues we fixed. Rather than confidently citing the wrong source, it caught the mismatch and told you — which is honestly a nice demonstration of the agent reasoning about its own source quality. But for a clean submission we should fix it properly rather than rely on the agent noticing every time.

Let's check the Knowledge Base Issues tab again in the Sanity dashboard — screenshot what's there. If guidance/aws_cis_alignment has a new pending issue, we'll resolve it the same way as before. Then let's test 2-3 more questions covering the other patterns (PassRole, MFA-on-service-role, wildcard principal, and the full-admin-access control case) to confirm breadth before we start on the write-up.

Agent

The Sources list looks fine (still 5 files, all Ready), and there's no red badge on "Issues" right now — meaning Sanity's own conflict detector isn't currently flagging the citation mismatch the agent found. That's a gap between what Sanity checks for and what's actually in the generated entries.

Click into Issues to confirm the count, and then let's look at the Entries list — click on guidance/aws_cis_alignment and methodology/risk_vs_false_positive specifically so we can see the exact citation numbering in each. Since entries can't be hand-edited (only rebuilt from sources or corrected via Instructions), we may need to add an explicit Instruction pinning which source ID maps to which topic, then trigger a rebuild.

Agent

Good — confirms Sanity's own detector considers everything resolved (4/4, 0 pending), but the actual entry text still has the old mislabeling baked in. That likely means the "guidance/aws_cis_alignment" entry (part of the 8 total now, up from what we last saw) was written before or without picking up the fix we applied to the other entries.

Let's force a full rebuild to get the correction to propagate everywhere. Go to Sources, and check the "⋮" menu (or Settings) near "Entries up to date" for a "Rebuild" or "Force rebuild" option — click that. If there's no such option there, click into the Entries tab, open guidance/aws_cis_alignment, and look for a rebuild/regenerate action on that specific entry.

Agent

Exactly what we need — click "Rebuild knowledge base". It says instructions and resolved issues are kept, so this should reapply the citation fixes across every entry, including the ones the agent flagged. Given our KB is small (5 sources), this should finish quickly, not take hours.

Agent

Says "Entries up to date," so the rebuild should have completed. Let's re-test the same question from earlier to confirm the citation issue is gone, and also try one covering a different pattern. Run these in your terminal:

python agent.py "Does an IAM policy attached to a Lambda execution role need an MFA condition on sts:AssumeRole?"

Paste the output.

You

bimbo@Abimbola MINGW64 ~/sanity-iam-agent
$ python agent.py "Does an IAM policy attached to a Lambda execution role need an MFA condition on sts:AssumeRole?"
I'll start by loading the dataset context, then check the knowledge base for the relevant rule.

Short answer: No — this is a documented false positive

What the linter says

iam-lint's MISSING_MFA_CONDITION rule (severity: medium) fires on any statement granting one of five sensitive actions — iam:CreateAccessKey, iam:AttachUserPolicy, iam:PutUserPolicy, sts:AssumeRole, iam:PassRole — without an aws:MultiFactorAuthPresent condition. sts:AssumeRole is explicitly in that list, so your policy will be flagged.

Critically, the rule fires identically regardless of what identity the policy is attached to — human user, Lambda execution role, or EC2 instance profile all trigger it the same way, because a JSON policy document doesn't encode what it's attached to (findings/mfa_condition, reference/rule_severity_map).

What AWS guidance says

AWS's MFA recommendation targets human users signing in interactively. MFA is a second factor a person supplies at login; it has no analog for machine identities. A role assumed by Lambda authenticates via short-lived STS credentials — there's no login step and no human present to supply a hardware key or one-time code. For machine identities, AWS's guidance is short-lived credentials, least-privilege scoping, and rotation — not MFA (findings/mfa_condition).

Do the sources agree?

No — this is one of three patterns where the knowledge base says the linter over-fires. The severity matrix marks MISSING_MFA_CONDITION as "⚠️ Partial conflict — false positives on service/machine roles" against AWS Well-Architected machine-identity guidance (reference/rule_severity_map, guidance/linter_vs_aws).

The decision table is explicit:

Identity type MFA applicable? Finding
Human IAM user (interactive login) Yes True positive — add the condition
Service role (Lambda, EC2, etc.) No False positive — control cannot exist

Adding the condition to an automated role "would break the automation without adding any security" (findings/mfa_condition).

What to do

Per guidance/exception_criteria, this is classified as a documented false positive, not a risk acceptance — but it still needs documenting rather than silent suppression:

  • Identity type: service/machine (not human)
  • Compensating controls: short-lived credentials, least-privilege scoping, rotation policy
  • Confirmation the role is not assumable interactively by a human principal

Plus the standard workflow: written justification, named owner, review cadence (suggested 180 days for medium findings), and a tracker entry.

Two caveats worth checking

  1. Mixed-use roles. If humans can also assume this role interactively, the KB calls that "an architectural smell" — the correct fix is splitting into separate roles, not adding an MFA condition that would break service access. This is why the exception criteria require you to confirm no human principal can assume it.

  2. Scope limit. This answer covers only whether MFA is required. If the same policy grants iam:PassRole, that action appears in both rules' action lists — and PASSROLE_UNRESTRICTED is a separate finding where the KB says the linter under-checks (it clears on ARN-scoping, but AWS also wants an iam:PassedToService condition). A clean MFA verdict says nothing about PassRole scoping.
    ((venv) )
    bimbo@Abimbola MINGW64 ~/sanity-iam-agent

Agent

No citation warning this time — the rebuild fixed it, and this answer is even stronger than the first: a clean false-positive verdict, a decision table, proper exception criteria, and it correctly flagged that iam:PassRole in the same policy would need a separate check without conflating the two rules. The Knowledge Base is working exactly as designed.

We're in good shape to move toward the submission. Next up:

  1. Push the code to GitHub — I can do this, but since it makes the repo public, I want your go-ahead first. I'd push to Bims-creator/sanity-iam-agent (matching your existing repo naming pattern).
  2. Write the DEV.to submission post using the Path One template, in the natural voice we agreed on, including the required Sanity project ID (r3gwx3ch) and the #sanitychallenge tag.
  3. Optionally embed this Claude Code session as the agent-session transcript via dev.to/agent_sessions/new.

Want me to go ahead and create the GitHub repo and push now, or would you rather do that part yourself?

You

I hope it won't show you as a contributor in the repo?

Agent

By default, yes — my instructions have me add Co-Authored-By: Claude Sonnet 5 <[REDACTED]> to commit messages I create, which would show up as a co-author on those commits. Since you're asking, I'll take that as wanting it omitted — I can commit under just your authorship with no AI attribution line, which makes sense for a solo portfolio/competition submission where you want it read as your own work.

Want me to go ahead on that basis: create the repo and push with plain commits, no co-author trailer?

You

Yes please

Agent
Agent
Agent

Now let's set up the git repo and push it.

Agent
Agent

Good — git identity is already set to you (Bims-creator / [REDACTED]), and gh is authenticated. Let's commit and push.

Agent
Agent

Good — .env and venv/ correctly excluded via .gitignore, no secrets staged. Committing now.

Agent
Agent

No co-author trailer, plain commit under your identity. Now creating the GitHub repo and pushing.

Agent
Agent

Repo is live: github.com/Bims-creator/sanity-iam-agent — solo commit, no AI attribution.

Agent
Agent

Repo's pushed. Next natural step is drafting the DEV.to submission post (using the Path One template, in the human voice we agreed on) — I'd write it as a draft for you to review before anything gets published, since posting it is your call. Want me to draft that now?

You

Yes, draft it

Agent
Agent

Draft's saved at SUBMISSION.md. Here's the full text:

Agent
Agent

Two loose threads before this is submission-ready:

  1. The demo excerpts have literal [...] placeholder text where I truncated the real agent output — that reads as unfinished in a real post. I can either paste the full answers verbatim or trim them to clean, complete excerpts.
  2. I wrote "Full transcripts are in the repo" but haven't actually created that file yet.

Want me to add a transcripts.md with the two full Q&A exchanges from earlier and fix the demo section to reference it properly?

You

Yes, fix both

Agent
Agent

Now fixing the Demo section in SUBMISSION.md to reference this properly instead of the truncated placeholders.

Agent
Agent

Both fixed. Let's commit and push these two files.

Agent
Agent

Pushed to github.com/Bims-creator/sanity-iam-agent.

Where things stand: Knowledge Base is live and validated, agent works correctly and is committed, and the submission draft plus full transcripts are in the repo. What's left before you submit:

  1. Review SUBMISSION.md — read it over, tell me if anything should change (tone, title, detail level)
  2. Publish on DEV.to — since that's posting to your account, I'll leave that click to you: open the Path One submission template, paste in the content from SUBMISSION.md, and hit publish when ready
  3. Optional: upload this Claude Code session as an agent-session transcript at dev.to/agent_sessions/new, hit Make Public, then embed it in the post's Agent Session section

Take a look at the draft and let me know what you'd like adjusted.

You

What do you think about the title?

Agent

Honestly, it's fine but a little generic — "Is X real or just noise? I built an agent to tell you" is a common blog-hook template, and with judges skimming a pile of submissions, it doesn't signal the actual technical hook (structured content resolving real disagreements between sources) the way a more specific title could.

I'd lean toward something like "When the linter and AWS disagree, an agent that shows you both sides" or "Teaching an agent to catch its own linter's false positives" — either names the actual mechanism instead of just posing a question. Tradeoff: less catchy as a hook, more precise as a description — probably the right trade for a technical judging criteria over general reader engagement.

You

Pick the stronger option out of the above 2 suggested titles

Agent

The first one: "When the linter and AWS disagree, an agent that shows you both sides."

It describes what the agent actually does, and it lines up with the challenge's own framing of contradicting sources surfacing side by side. The second title overclaims. The agent judges findings and doesn't "catch" them, and "its own linter" reads as if the agent owns the linter, which it doesn't.

Want me to update the title in SUBMISSION.md and push it?

You

Yes please

Agent
Agent
Agent

Title updated and pushed, with a plain commit and no co-author trailer, as you asked.

What's left is on your side: paste the post into DEV.to and publish it, and optionally upload this session as an agent transcript first. Tell me if you want any other changes to the draft.

You

Is the task for the challenge completed?

1 message not shown
You

Yes, give me the two test queries

Agent

Run these one at a time from ~/sanity-iam-agent (env vars still need to be exported if you opened a new terminal).

PassRole:
bash
python agent.py "A policy grants iam:PassRole scoped to one specific role ARN, with no iam:PassedToService condition. iam-lint doesn't flag it. Is it fully safe?"

A good answer says the linter clears it, but AWS also recommends an iam:PassedToService condition, so it's only half-mitigated.

Wildcard principal:
bash
python agent.py "An S3 bucket policy has Principal: '*' but with an aws:SourceArn condition restricting it to one CloudFront distribution. iam-lint flags it as critical. Real risk or false positive?"

A good answer treats a wildcard principal with a tight condition as a documented pattern (false positive), and a bare wildcard with no condition as a genuine critical finding.

Paste the outputs here. I'll check them against the Knowledge Base and add them to transcripts.md if they're clean.

You

bimbo@Abimbola MINGW64 ~/sanity-iam-agent
$ python agent.py "Does an IAM policy attached to a Lambda execution role need an MFA condition on sts:AssumeRole?"
I'll start by loading the dataset context, then check the knowledge base for the relevant rule.

Short answer: No — this is a documented false positive

What the linter says

iam-lint's MISSING_MFA_CONDITION rule (severity: medium) fires on any statement granting one of five sensitive actions — iam:CreateAccessKey, iam:AttachUserPolicy, iam:PutUserPolicy, sts:AssumeRole, iam:PassRole — without an aws:MultiFactorAuthPresent condition. sts:AssumeRole is explicitly in that list, so your policy will be flagged.

Critically, the rule fires identically regardless of what identity the policy is attached to — human user, Lambda execution role, or EC2 instance profile all trigger it the same way, because a JSON policy document doesn't encode what it's attached to (findings/mfa_condition, reference/rule_severity_map).

What AWS guidance says

AWS's MFA recommendation targets human users signing in interactively. MFA is a second factor a person supplies at login; it has no analog for machine identities. A role assumed by Lambda authenticates via short-lived STS credentials — there's no login step and no human present to supply a hardware key or one-time code. For machine identities, AWS's guidance is short-lived credentials, least-privilege scoping, and rotation — not MFA (findings/mfa_condition).

Do the sources agree?

No — this is one of three patterns where the knowledge base says the linter over-fires. The severity matrix marks MISSING_MFA_CONDITION as "⚠️ Partial conflict — false positives on service/machine roles" against AWS Well-Architected machine-identity guidance (reference/rule_severity_map, guidance/linter_vs_aws).

The decision table is explicit:

Identity type MFA applicable? Finding
Human IAM user (interactive login) Yes True positive — add the condition
Service role (Lambda, EC2, etc.) No False positive — control cannot exist

Adding the condition to an automated role "would break the automation without adding any security" (findings/mfa_condition).

What to do

Per guidance/exception_criteria, this is classified as a documented false positive, not a risk acceptance — but it still needs documenting rather than silent suppression:

  • Identity type: service/machine (not human)
  • Compensating controls: short-lived credentials, least-privilege scoping, rotation policy
  • Confirmation the role is not assumable interactively by a human principal

Plus the standard workflow: written justification, named owner, review cadence (suggested 180 days for medium findings), and a tracker entry.

Two caveats worth checking

  1. Mixed-use roles. If humans can also assume this role interactively, the KB calls that "an architectural smell" — the correct fix is splitting into separate roles, not adding an MFA condition that would break service access. This is why the exception criteria require you to confirm no human principal can assume it.

  2. Scope limit. This answer covers only whether MFA is required. If the same policy grants iam:PassRole, that action appears in both rules' action lists — and PASSROLE_UNRESTRICTED is a separate finding where the KB says the linter under-checks (it clears on ARN-scoping, but AWS also wants an iam:PassedToService condition). A clean MFA verdict says nothing about PassRole scoping.
    ((venv) )
    bimbo@Abimbola MINGW64 ~/sanity-iam-agent
    $ python agent.py "A policy grants iam:PassRole scoped to one specific role ARN, with no iam:PassedToService condition. iam-lint doesn't flag it. Is it fully safe?"
    I'll start by pulling the dataset context and then look up the relevant rule.

    Short answer: no — it's hardened against one threat, not both.

What the linter did. PASSROLE_UNRESTRICTED (severity: high) fires only when iam:PassRole is granted with Resource: "*", and it clears as soon as Resource is scoped to a specific role ARN — at that point it stops checking (findings/passrole). So a clean result here means exactly one thing was verified.

What AWS says. AWS's PassRole documentation recommends two controls working together, not one (findings/passrole):

  1. Restrict Resource to a specific role ARN — prevents passing any role in the account.
  2. Add an iam:PassedToService condition — restricts which AWS service may receive that role.

AWS names the second gap the confused-deputy angle by name: "the ARN restriction limits which role can be passed; the condition restricts what can assume it. The two controls are not substitutes — both are needed for complete mitigation."

Do the sources agree? No — this is a partial conflict, and notably it's the one rule in the set where the linter under-checks rather than over-fires (guidance/linter_vs_aws, reference/rule_severity_map). The other four rules' tension is false positives; here it's a coverage gap in the opposite direction.

Your exact configuration is called out explicitly in the KB's risk table (findings/passrole):

Scenario Linter result Real risk level
iam:PassRole + Resource: "*" FLAGGED (high) True positive
specific role ARN, no iam:PassedToService Clean Residual risk — confused-deputy attack still possible
specific role ARN + iam:PassedToService Clean Fully hardened per AWS guidance

Important framing point: the KB is explicit that this is "a gap in coverage, not a false positive" — the linter isn't wrong, and ARN-scoping is a genuine improvement over a wildcard. The issue is that the rule's name ("unrestricted") implies broader verification than the check performs. So this isn't an exception-to-document case; guidance/exception_criteria lists no PassRole category. It's a manual remediation step: add the iam:PassedToService condition.

One secondary thing to check. iam:PassRole is also in the covered-action list for MISSING_MFA_CONDITION (severity: medium), which fires when any of its five actions appear without an aws:MultiFactorAuthPresent condition (findings/mfa_condition). You describe the statement as having no condition block, so if that rule didn't fire either, that's worth verifying. Note the verdict there depends on identity type, which the linter can't read from the JSON: true positive for a human IAM user, false positive for a Lambda/EC2/service role, where MFA is structurally inapplicable.

Bottom line: clean on PASSROLE_UNRESTRICTED ≠ compliant with AWS's recommended pattern. The KB's guidance is to manually verify iam:PassedToService on any PassRole grant that matters to your threat model. I can't tell you how exploitable the residual risk is in your environment — that depends on the target role's permissions and which services could receive it, which isn't in this dataset.
((venv) )
bimbo@Abimbola MINGW64 ~/sanity-iam-agent (master)
$
((venv) )
bimbo@Abimbola MINGW64 ~/sanity-iam-agent (master)

Agent

The PassRole answer is what the Knowledge Base should produce. It says the configuration is not fully safe, shows the exact "specific ARN, no iam:PassedToService" row in the risk table, and describes it as a gap in what the linter checks, not a false positive. It also noticed on its own that the MFA rule might apply to the same statement.

I'll add it to the transcripts now.

Agent
Agent

The PassRole run is in transcripts.md and pushed. The file's header still says "Two full runs," and I'll correct the count once the last run is added.

The last test is the wildcard-principal one:

python agent.py "An S3 bucket policy has Principal: '*' but with an aws:SourceArn condition restricting it to one CloudFront distribution. iam-lint flags it as critical. Real risk or false positive?"

Paste the output when it finishes.

You

bimbo@Abimbola MINGW64 ~/sanity-iam-agent (master)
$ python agent.py "An S3 bucket policy has Principal: '*' but with an aws:SourceArn condition restricting it to one CloudFront distribution. iam-lint flags it as critical. Real risk or false positive?"
I'll start by pulling the dataset context and then look up the relevant rule.

Verdict: Very likely a false positive — but document it, don't just suppress it

What the linter says

WILDCARD_PRINCIPAL, severity critical. The rule fires on any trust or resource-based policy where Principal is or includes "*". Critically, it stops reading at Principal and never inspects the Condition block — so a tightly-scoped policy produces the exact same critical finding as an accidentally public bucket (findings/wildcard_principal, reference/rule_severity_map).

What AWS guidance says about the same pattern

AWS explicitly documents Principal: "*" paired with a restrictive Condition (aws:SourceArn, aws:SourceAccount) as a normal, supported pattern for resource-based policies. In that shape the wildcard means "anyone who satisfies the condition," not "anyone on the internet." AWS's own example policies for SNS, SQS, and S3 use this form, and service-to-resource invocation restricted to a specific CloudFront distribution via aws:SourceArn is named directly as a legitimate case (findings/wildcard_principal).

The KB's assessment table maps your exact case: Principal: "*" + aws:SourceArn/aws:SourceAccount → "Likely false positive — deliberate, AWS-documented pattern." Note the KB's own hedge — it says "likely," not "definitely."

Do the sources agree?

No — this is a known, documented conflict. guidance/linter_vs_aws classifies WILDCARD_PRINCIPAL as "Partial — linter over-fires," with the key tension being "condition-mitigated wildcards are a documented AWS pattern." It adds a useful framing point: this is simultaneously the loudest finding in the rule set and the one most prone to false positives on correctly structured policies.

Critical severity remains the correct call for a bare Principal: "*" with no condition — the severity isn't wrong in general, it's just applied without the context needed to distinguish your case.

What to verify before closing it

The KB's exception criteria (guidance/exception_criteria) ask you to confirm and record:

  1. The specific condition key and value — e.g. aws:SourceArn: arn:aws:cloudfront::123456789012:distribution/ABCDEF. Confirm it names one concrete distribution ARN, not a wildcarded ARN pattern.
  2. The service/account the condition constrains access to.
  3. That the condition is not trivially bypassable — the KB's counter-example is a weak key like aws:RequestedRegion. (Its other "true positive" example is aws:SourceIp with an open CIDR — a condition that doesn't meaningfully narrow access.)

Process note

The KB is explicit that this should be a documented false positive, not a silent suppression. Its approval workflow calls for written justification, compensating controls, a named owner, a review cadence (suggested 90 days for critical findings), and a tracker entry.

Limits of what I can tell you

Everything above comes from the knowledge base's five-pattern rule coverage. Two things it does not address, so I can't speak to them from this data:

  • Whether your specific aws:SourceArn value is well-formed or scoped as you intend — I haven't seen the policy.
  • CloudFront-specific mechanics (e.g. OAC vs. legacy OAI behavior, or whether the granted S3 actions themselves are appropriately scoped). The KB treats CloudFront only as an example of a SourceArn-constrained service. A wildcard resource or overly broad action in the same statement would be a separate finding under WILDCARD_RESOURCE, which this answer doesn't cover. ((venv) ) bimbo@Abimbola MINGW64 ~/sanity-iam-agent (master)
Agent

This one is correct too. It gave "likely a false positive," listed the checks to run before closing the finding, and said which parts it can't judge from the Knowledge Base. I'll add it and fix the header count.

Agent
Agent

All four contradiction patterns are now tested and pushed, in transcripts.md. The two false-positive cases (wildcard resource, wildcard principal, MFA on service roles) came out as false positives, and PassRole came out as a coverage gap the linter doesn't check. That matches what we designed.

The Demo section of SUBMISSION.md still says "two questions," so I'd update it to cover all four, including the PassRole case since it's the one where the linter under-checks. Want me to do that?

After that, what's left is on your side. Publish the post on DEV.to before October 4 at 11:59 PM PDT, and revoke the unused first Sanity token.

You

Yes, update the demo section

Agent
Agent
Agent
Agent
Agent

The Demo section in SUBMISSION.md now covers all four queries, and I've pushed it with a plain commit.

The post is ready for you to publish on DEV.to before October 4 at 11:59 PM PDT. Before you do, revoke the unused first Sanity token under Manage → API → Tokens.

You

How do I upload the agent session?

Agent

Let me find where Claude Code saved this session's transcript file.

Agent
Agent

This session's transcript is the newest file, about 8 MB:

[REDACTED]\.claude\projects\C--Users-bimbo--github\660f6410-8ee9-473d-8d59-316a761c3444.jsonl

To upload it:
1. Go to dev.to/agent_sessions/new while signed in.
2. Choose Claude Code as the tool and upload that .jsonl file.
3. Use the curate/slice option to keep only the parts worth showing. Everything from picking Path One through the four test runs is the useful part.
4. Click Make Public. Uploads are unlisted by default, and judges can't open them until you do.
5. Copy the embed code or link it gives you, and paste it into the ## Agent Session section of your post, replacing the HTML comment.

Before you make it public:
- Slice it. At 8 MB, it holds the whole session, including some unrelated setup context at the start. I don't know exactly what the file contains beyond the visible conversation. It may hold your saved memory notes, which mention your interviews and visa plans, so check the start of the transcript for anything you don't want public.
- Search it for secrets. I never saw your Sanity token or Anthropic key. Still, DEV.to's redaction won't catch everything, so scan for sk- and skUw before publishing.

If you'd rather not deal with a large file, skip the upload. The transcripts and repo already cover the "show your work" side, and the session embed is optional.

You

Won't it be too early if I submit today given that the challenge just opened about a day ago?

Agent

No, submitting early isn't a problem. The challenge page lists only a deadline (October 4, 11:59 PM PDT) and a winners date (October 22), with no penalty or bonus for timing, so judging almost certainly happens after the close.

Reasons to publish now:
- It removes the risk of running out of time or hitting a problem on deadline day, which matters with interview prep also on your plate.
- You can still edit a published DEV.to post, so you can fix things or add the agent session embed later. I haven't seen the challenge rules say edits are frozen after submission, so check the Official Rules link on the challenge page if that matters to you.
- The post can pick up reactions and comments over the next two weeks. Those probably don't affect judging, but they're useful visibility for your portfolio.

Reasons to wait:
- Other entrants can see your approach. That's a small risk, since the value here is in your own iam-lint content and the write-up.
- If you want to record a short demo video or embed the session first, waiting a day or two lets the post go out complete. Neither is required.

I'd publish now. If you want the agent session in the post, do that first, since it takes about ten minutes.

You

I will wait a little

Agent

Fine. Everything on the build side is finished and pushed to github.com/Bims-creator/sanity-iam-agent, and the deadline is October 4 at 11:59 PM PDT.

Two things will help you when you come back:
- Revoke the unused first Sanity token. It's still active under Manage → API → Tokens.
- Watch your Anthropic API key. It expires in 30 days (October 18), which is fine for the deadline, and any credit you added is only spent when you run the agent.

I've saved the project status in memory. Message me when you're ready to publish or want changes to the draft.

You

Please rename this session

Agent
Agent
Agent

I've renamed this session to "Sanity Challenge: IAM policy review agent". It was previously "Session planning".

You

I am ready to submit my work

Agent
Agent
Agent

The submission file looks complete and consistent — title, tags, project ID, four demo queries, code and transcript links all check out, and the repo's fully pushed. One open item: the Agent Session section still has the placeholder comment since you held off on that upload last time.

Agent
Agent
Agent

Good, this is still the right file — same session, now ~8.8 MB with everything up through today included.

To upload:
1. Go to dev.to/agent_sessions/new while signed in.
2. Select Claude Code as the tool, upload this file:

[REDACTED]\.claude\projects\C--Users-bimbo--github\660f6410-8ee9-473d-8d59-316a761c3444.jsonl

3. Use the curator to slice it down — the useful stretch runs from choosing Path One through the four test queries at the end. Earlier unrelated context (interview prep, other projects) doesn't need to be in there.
4. Before clicking Make Public, scan whatever preview/export it shows you for sk- or skUw — those are the prefixes of the Anthropic key and the Sanity token. I never saw either value in this chat, but the raw session file might contain them if they appeared in your terminal output at any point.
5. Click Make Public, then copy the embed code it gives you.

Once you've got the embed code, paste it to me and I'll drop it into SUBMISSION.md in place of the placeholder comment, then we'll do the final publish.

You

[Image: source: [REDACTED]\AppData\Local\Temp\claude\C--Users-bimbo--github\660f6410-8ee9-473d-8d59-316a761c3444\images\1.png]

Agent

You're in the right place. Two things to fix here:

1. Set the title — click into "New post title here..." and type:

When the linter and AWS disagree, an agent that shows you both sides

2. Replace the whole body — click into the editor, select all (Ctrl+A), delete, and paste this in place of it:

Agent
Agent
*This is a submission for the [Sanity Challenge, Path One: Ship an Agent That Queries Real Content](https://dev.to/challenges/sanity-2026-09-16)*

## What I Built

I maintain [iam-lint](https://github.com/Bims-creator/iam-lint), a small static analyzer for AWS IAM policies. It's good at catching obvious over-permissioning — wildcard actions, wildcard resources, unrestricted `PassRole` — but like any linter, it's blunt. It reads a JSON policy document and nothing else, so it can't tell a genuine misconfiguration from a pattern AWS itself documents as intentional. `ec2:DescribeInstances` gets flagged for `Resource: "*"` even though that action has no other option. An MFA-condition rule fires on a Lambda execution role even though a service identity has no way to present MFA in the first place.

So I built an agent that sits on top of the linter's output and actually reasons about whether a flagged pattern is a real risk. It pairs each of five iam-lint rules against what AWS's own documentation and the CIS AWS Foundations Benchmark say about the same pattern, and four of the five pairs genuinely disagree with the linter's blanket severity. The agent doesn't pick a side — it walks through the disagreement and tells you which context factors resolve it.

## Demo

This is a CLI tool, not a hosted app, so here it is running against four questions, one for each place the linter and AWS's guidance part ways:

**Query:** `Is a policy with Resource: '*' on ec2:DescribeInstances a real finding or a false positive?`

> Short answer: false positive — and this is the KB's canonical example of one. iam-lint's `WILDCARD_RESOURCE` rule is explicitly action-blind — it doesn't check whether the action even supports resource-level permissions. AWS's Service Authorization Reference names `ec2:DescribeInstances` directly as an action where `Resource: "*"` is the only value IAM will accept.

**Query:** `Does an IAM policy attached to a Lambda execution role need an MFA condition on sts:AssumeRole?`

> Short answer: No — this is a documented false positive. AWS's MFA recommendation targets human users signing in interactively; a Lambda execution role authenticates via short-lived STS credentials with no login step. The agent backs this with a full decision table by identity type, and flags that the same policy's PassRole scoping would need a separate check.

**Query:** `A policy grants iam:PassRole scoped to one specific role ARN, with no iam:PassedToService condition. iam-lint doesn't flag it. Is it fully safe?`

> Short answer: no — it's hardened against one threat, not both. This is the one case where the linter under-checks instead of over-firing: it clears the statement once the resource is a specific ARN, but AWS also recommends an `iam:PassedToService` condition to close the confused-deputy gap. The agent calls it a gap in coverage, not a false positive, and points to a manual fix instead of an exception.

**Query:** `An S3 bucket policy has Principal: '*' but with an aws:SourceArn condition restricting it to one CloudFront distribution. iam-lint flags it as critical. Real risk or false positive?`

> Verdict: very likely a false positive, but document it instead of suppressing it. The linter stops reading at `Principal` and never looks at the `Condition` block, while AWS documents a wildcard principal with a tight condition as a normal pattern. The agent lists what to verify before closing it, such as whether the ARN names one concrete distribution, and says plainly which parts it can't judge without seeing the policy.

Full unedited transcripts, including one where the agent caught and flagged a citation bug in its own source data rather than silently mis-citing, are in [transcripts.md](https://github.com/Bims-creator/sanity-iam-agent/blob/master/transcripts.md).

## Code

[github.com/Bims-creator/sanity-iam-agent](https://github.com/Bims-creator/sanity-iam-agent)

## How I Used Sanity

The Knowledge Base is built from five markdown files (`knowledge-base/*.md` in the repo), one per IAM finding, added as a **Files** source in Sanity Context. Each file states the linter's claim, what AWS/CIS says about the same pattern, and where the two land — agreeing, partially conflicting, or (in one case, full admin access) not in conflict at all, included on purpose as a control so the agent doesn't manufacture disagreement where none exists.

The build did more than I expected. Beyond one entry per source file, it synthesized a cross-cutting entry — "True-Risk vs False-Positive Evaluation Method" — that pulled a shared decision framework out of the five files: four context factors (action type, identity type, condition blocks, service-specific semantics), a false-positive/real-risk signal pair for each rule, and an exception-documentation template. That entry is doing most of the actual work when the agent answers a question.

The build also caught something I didn't plan for: two of my source files got their citation numbering crossed in a couple of the generated entries (a `wildcard_principal` source cited as if it were the `passrole` source, and vice versa). Context flagged all four instances as critical conflicts in its Issues review, I resolved them by picking the correct source-to-claim mapping, and a full rebuild propagated the fix. Worth knowing if you're building on this: the automated conflict detection covers structural mismatches like this reliably, but I still had the agent itself catch one straggler in a newly-synthesized entry before I rebuilt — which, if anything, is a decent argument for the "surface the disagreement instead of hiding it" design.

The agent connects to the Knowledge Base through the Sanity Context MCP endpoint, using Claude's native MCP connector (`mcp_servers` + an `mcp_toolset` tool) so the tool calls happen server-side — no manual MCP client loop in the agent code.

## Sanity Project Details

Project ID: `r3gwx3ch`

## Agent Session

<!-- paste the embed code here after you finish the upload at dev.to/agent_sessions/new -->

---

Straightforward takeaway: a linter tells you what pattern matched. This Knowledge Base tells you whether the pattern is actually a problem, and shows its work either way.

Tags are already correct (devchallenge, sanitychallenge, sanity, ai) — leave those as-is. Once you've pasted this in and set the title, go do the agent session upload in a separate tab, then come back and swap that last HTML comment for the real embed code before hitting Publish.

You

What next?

Agent

Next steps, in order:

  1. Save this as a draft first — click Save Draft (not Publish yet) so you don't lose the title/body you just set.
  2. Open a new tab to dev.to/agent_sessions/new, upload the .jsonl file from earlier, curate it down to the useful stretch, scan it for sk-/skUw before making it public, then click Make Public and copy the embed code it gives you.
  3. Come back to this draft, find the <!-- paste the embed code here... --> line under Agent Session, and replace it with that embed code.
  4. Preview the post (the "Preview" tab at the top) — check headings render properly, the four demo queries are readable, and links work.
  5. Revoke the old, unused Sanity token if you haven't already (Manage → API → Tokens in Sanity).
  6. Publish.

Let me know once you've got the embed code, or if anything looks off in the preview.

You

[Image: source: [REDACTED]\AppData\Local\Temp\claude\C--Users-bimbo--github\660f6410-8ee9-473d-8d59-316a761c3444\images\2.png]

Agent

That's the wrong file — custom-title.json is only 59 bytes and detected as Gemini CLI, not our actual session. Click the X next to it to remove it, then browse to (or drag in) the real file:

[REDACTED]\.claude\projects\C--Users-bimbo--github\660f6410-8ee9-473d-8d59-316a761c3444.jsonl

It should be about 8.8 MB and get detected as Claude Code. Once that's loaded correctly, click Parse File.

You

[Image: source: [REDACTED]\AppData\Local\Temp\claude\C--Users-bimbo--github\660f6410-8ee9-473d-8d59-316a761c3444\images\3.png]

Agent

That folder named 660f6410-... is just this session's temp/scratch folder — not where the actual transcript lives. Don't open it. Instead, type the full path directly into the "File name:" box at the bottom and click Open:

[REDACTED]\.claude\projects\C--Users-bimbo--github\660f6410-8ee9-473d-8d59-316a761c3444.jsonl

That should jump straight to the right file regardless of which folder the dialog is currently browsing.

You

[Image: source: [REDACTED]\AppData\Local\Temp\claude\C--Users-bimbo--github\660f6410-8ee9-473d-8d59-316a761c3444\images\4.png]

Agent

Good, correct file loaded — 8.8 MB, detected as Claude Code. Two things before you click Parse:

  1. Update the Session Title — it still says "Gemini CLI Session" (leftover default). Change it to something like:

    Sanity IAM Policy Review Agent — build session

  2. Then click Parse File. That'll kick off the local parsing/scrubbing, and should bring you to the curation step next.

You

[Image: source: [REDACTED]\AppData\Local\Temp\claude\C--Users-bimbo--github\660f6410-8ee9-473d-8d59-316a761c3444\images\5.png]

Agent

That's likely a transient browser-side failure — 8.8 MB is a fair bit for client-side parsing. Try these in order:

  1. Click "Parse File" again — sometimes it just needs a retry.
  2. If that fails again, reload the dev.to page fresh and re-upload the same file.
  3. If it still errors, it may be hitting a size limit in the browser's local parser.

Since this step is optional and we already have thorough coverage in transcripts.md plus the full repo, I'd try it once or twice more but not burn a lot of time on it — if it keeps failing, I'd say skip the agent session embed and publish without it rather than risk missing other prep before the deadline.


Straightforward takeaway: a linter tells you what pattern matched. This Knowledge Base tells you whether the pattern is actually a problem, and shows its work either way.

Top comments (0)