This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
What I Built
Which AWS quota is actually current? You copied a default vCPU limit off the AWS docs, shipped it, and it broke in production. I have done this. The number was stale, and nothing warned me.
Here is what makes it nasty. For one AWS fact, three official pages can each give you a different number. An old User Guide says one thing. The Service Quotas console says another. A pricing page says a third. All look official. None of them tells you which is live today.
So I built an agent that answers "which AWS value is current?" and then proves it. It reads typed facts from a Sanity Context Knowledge Base, picks the winner with a deterministic rule the model is not allowed to override, and shows both the current value and the one it replaced, each with its source. Then it does the part most content agents skip: it calls the live AWS API (read-only) to check whether even the reconciled record is still true. On the EC2 vCPU quota it turns up three different numbers, docs 32, record 5, live 16, and puts all three on the screen.
Stack: Amazon Nova Pro on Bedrock, Strands Agents (Python), a Sanity Context Knowledge Base read over the hosted Context MCP, and a read-only live AWS cross-check.
Why a keyword search returns the wrong AWS value
The challenge sets a hard bar: "the strongest submissions show an agent that only works because the content was structured. If a keyword search would have gotten you the same answer, aim higher."
Fair. So I built a fact where keyword search gets it wrong on purpose. Real one too: the max IOPS per volume for an EBS general purpose SSD. The current number is 80,000 (gp3, from the EBS console). The old number is 16,000, the legacy gp2 ceiling, still sitting in an older SSD guide. That old guide repeats the exact phrase you would type into a search box: "maximum IOPS per volume for a general purpose SSD EBS volume." It is stale and wordy, so keyword scoring loves it.
Watch it pick the wrong answer. Same content, real TF-IDF, the question a person would actually ask:
Keyword/TF-IDF baseline for: "maximum IOPS per volume general purpose SSD"
0.8626 EBS/limit: Maximum IOPS per volume for a general purpose SSD ... = 16000 <-- WRONG (old)
0.2026 EBS/limit: Maximum provisioned IOPS per gp3 volume = 80000 <-- right, ranked lower
0.0218 S3/price: S3 Standard storage, first 50 TB / month = 0.023
0.0185 EC2/quota: Running On-Demand Standard ... = 5
...
The stale 16,000 wins by a mile. Keyword search hands you the wrong number and sounds sure of it.
Demo
Run the credential-free core yourself in under a minute (public dataset over anonymous GROQ, local reconcile, read-only live AWS check, no model or token needed):
# 1. clone + install
git clone https://github.com/simplynadaf/aws-source-of-truth-agent
cd aws-source-of-truth-agent
pip install -r requirements.txt
# 2. the credential-free path (public dataset + local reconcile + live AWS check)
python -m agent.reconcile_offline --service EBS --type limit --region us-east-1
# -> Verdict: 80000 IOPS (serviceQuotasConsole wins over officialDocs)
# 3. the keyword control that returns the WRONG answer:
python -m agent.baseline "maximum IOPS per volume general purpose SSD"
For the full LLM agent through the Context MCP, copy .env.example to .env and add a Bedrock region plus a Sanity Context Viewer token (the README walks through it). There is a JSON mode too, for piping into something else:
python -m agent.ask --json "What is the maximum IOPS per volume for a gp3 EBS volume?"
The real run, Nova Pro, us-east-1, read-only throughout:
| Fact | Reconciled (from the record) | Superseded | Live AWS | Result |
|---|---|---|---|---|
| EC2 On-Demand Standard vCPU quota | 5 (console) | 32 (old user guide) | 16 | DRIFT |
| EBS gp3 max IOPS per volume | 80,000 (console) | 16,000 (old SSD guide) | unavailable | trusted (no such live quota) |
| S3 Standard $/GB-mo | 0.023 (pricing page) | 0.021 (stale blog) | 0.023 | AGREE |
| RDS PostgreSQL oldest major | 13 (release notes) | 11 (old tutorial) | 11 | DRIFT |
| Lambda concurrent executions | 1000 (dev guide) | n/a | unavailable | trusted |
| Graviton4 (R8g) availability | Available (instance types) | n/a | Available | AGREE |
Look at the EC2 row. Docs say 32. The reconciled record says 5. The live account says 16. Three numbers for one quota, and the agent shows all three with their sources instead of picking one and hoping. The drift is not a failure. It is the honest answer: this is the current record, and here is where reality has already moved past it.
Code
π AWS Source of Truth
When your AWS docs, pricing page, and Service Quotas console disagree, an agent that knows which one is telling the truth, then checks the live API to see if even that record has drifted.
Sanity project id: 0q5ohtvv Β· β If a stale AWS number has ever bitten you in production, give this a star.
The Problem β’ Why Search Fails β’ How Structure Fixes It β’ The Twist β’ Getting Started β’ FAQ
π Table of Contents
- The Problem
- Why Keyword Search Can't Save You
- How the Structure Fixes It
- The Twist: Even the Reconciled Record Can Be Stale
- How It Works
- Tech Stack
- Prerequisites
- Getting Started
- Build the Knowledge Base
- Project Structure
- The Integrity Story (why the model can't fake it)
- Least-Privilege IAM Policy
- What Didn't Work
- Troubleshooting
- FAQ
- License
π€ The Problem
You copied a limit straight out of the AWS docsβ¦
Full source, plus the credential-free path judges can run with no token.
How I Used Sanity
What I pointed Sanity Context at: my own Sanity content. Every fact is a typed awsFact document in the production dataset, not a wall of prose:
{
"_type": "awsFact",
"service": "EBS", "factType": "limit", "region": "us-east-1",
"key": "Maximum provisioned IOPS per gp3 volume",
"currentValue": "80000", "unit": "IOPS",
"effectiveDate": "2026-01-15",
"source": { "name": "EBS gp3 volume limits (Service Quotas / EBS console)",
"kind": "serviceQuotasConsole", "url": "..." }
}
// ...plus a separate record for the old 16,000 (kind: officialDocs, 2020).
source.kind and effectiveDate are real fields, so the rule that picks the winner is dull and readable:
highest source precedence wins (console and pricing page beat changelog, which beats official docs, which beats a blog), and the newest effectiveDate breaks ties.
Prose cannot do this. The schema is the whole trick.
The Knowledge Base: a Sanity Context Knowledge Base indexes these facts. The build reads them ahead of time and writes short cited entries. Where two records fight, the entry keeps both numbers and both sources next to each other.
Which Context tools the agent used: the agent pulls facts over the hosted Context MCP with knowledge_base_search (ranked lookup) then knowledge_base_read (full entry). That is the proof the answer came from Sanity and not from the model's memory. (It can also fall back to anonymous GROQ over the public dataset for the no-token path.)
What the agent does with what it retrieves: Nova Pro runs the tools and writes the sentence, but it does not get to decide the number. A plain function reconciles the value, and a guard checks the model's answer against it. If they disagree, the guard throws out the model's prose and ships the deterministic answer instead. On one EBS run Nova muddled its own wording, the guard caught it, and swapped in the correct answer with no help from me. The model cannot invent the number even if it tries.
The tool trail proves the path through Sanity on every question:
- [Sanity Context MCP] knowledge_base_search('EC2 quota us-east-1 ...') -> top='ec2/quotas_and_limits'; knowledge_base_read(['ec2/quotas_and_limits'])
- fetch_candidate_facts(service='EC2', fact_type='quota', region='us-east-1') -> 2 rows
- reconcile_facts(n=2) -> current=5
- verify_live(EC2/quota) -> drift
Then: what if even the reconciled record is stale? (live AWS drift)
Reconciling the sources gives you the best answer the documents can offer. But documents rot. So the agent does one more thing a pure content agent will not: it calls the live AWS API, read-only, and asks whether the reconciled value is still true right now. Service Quotas, the Price List API, EC2, RDS. The results are in the Demo table above. The live layer is additive and clearly labelled; it is not part of Sanity and it is account-specific.
What didn't work (the honest part)
- I planned to screenshot a resolved conflict in the dashboard. The build never gave me one. The Knowledge Base read the typed records, reconciled the 32-vs-5 fight into a clean entry with both numbers cited, and left the Issues queue empty. That is the KB doing its job, but it means there is no "I resolved an Issue" screenshot to show, so I am not claiming one.
-
Nova dropped a tool argument on me.
fetch_candidate_factsreturned a row, then Nova passed an empty string toreconcile_facts, which saw zero facts. The guard fail-closed to "Not verified" instead of guessing. I fixed it by caching the last real fetch server-side, so the facts always come from Sanity, never from whatever the model relayed. -
A wrong
factTypeguess used to lose a real fact. Nova guessedinstanceTypewhen the field value wasregionalAvailability, the typed filter matched nothing, and the fact vanished. Now a zero-row typed fetch retries on service and region alone. - The live check is not part of Sanity, and it is account-specific. A judge on a different AWS account will see different live numbers. That layer is additive and labelled as such. Two facts (EBS max IOPS, Lambda) have no matching Service Quotas entry, so the agent says "unavailable" rather than fake a check. The seed is labelled demo data at the top.
Reusing the pattern beyond AWS
Drop the AWS parts and the shape is generic: type your sources, index them into a Knowledge Base, reconcile by precedence and date, guard the answer so the model cannot fake the winner, and check a live system of record when one exists. It fits API version support, pricing, compliance clauses, legal terms. Anywhere the real question is "which version is current?"
The thing I took away: the win here was not a smarter model. It was structured content, plus the nerve to say "the record says 5, but reality says 16."
Sanity Project Details
-
Sanity project ID:
0q5ohtvv -
Dataset:
production(public) - Public dataset query (no token): https://0q5ohtvv.api.sanity.io/v2023-05-03/data/query/production?query=*%5B_type%3D%3D%22awsFact%22%5D
If your docs and your console have ever disagreed with each other, tell me which number bit you.




Top comments (2)
This is a really solid approach to a problem that looks simple until you actually have to trust the answer.
I especially like that the model is not allowed to decide which number is correct. The flow is much more interesting: structured facts β source precedence β deterministic reconciliation β model response β guard β live AWS check.
The 32 vs 5 vs 16 example makes the point really well. Instead of hiding the conflict and returning one confident number, the system exposes the disagreement and then checks whether the reconciled value has drifted from the live environment.
That pattern is useful far beyond AWS. I can see the same architecture working for API versions, pricing, compliance rules, dependency versions, or even internal company policies where the question is basically: βWhich value is actually current?β
For me, the strongest part isn't the AI agent itself. It's the architecture around the agent that limits what the AI is allowed to claim. That's a much more practical way to build reliable AI systems.
Much better way to use Ai and limitations were sorted out in better way.