DEV Community

Tuấn Trần
Tuấn Trần

Posted on

WhatBin: Source-Backed Waste Guidance with Sanity Context MCP

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content.

What I Built

WhatBin helps people in Hanoi and Ho Chi Minh City figure out how to dispose of household items properly.

It sounds like a simple question:

How should I dispose of this?

But the answer can get complicated pretty quickly.

An old mattress, a household battery, and a broken fluorescent tube all need different handling. The answer can also change depending on the city and which local rules are currently in effect.

I built WhatBin to make that process easier:

  1. Describe an item or upload a photo.
  2. Choose Hanoi or Ho Chi Minh City and confirm the identified item.
  3. See the reviewed disposal guidance and the sources behind it.
  4. Ask follow-up questions about the evidence or compare the same item between the two cities.

One decision shaped most of the project: the AI should explain the guidance, not decide what the guidance is.

Published content in Sanity determines which disposal instruction applies. A server-side agent uses Sanity Context MCP to retrieve the supporting evidence and explain the result.

It cannot replace the reviewed instruction with its own recommendation.

WhatBin can return three results:

  • MATCHED — there is a reviewed rule for this item, city, and date.
  • UNKNOWN — WhatBin does not currently have enough reviewed content for this combination.
  • CONFLICT — there is an unresolved conflict or overlap between active records.

For UNKNOWN and CONFLICT, WhatBin does not show a disposal instruction.

I think this is important. Missing coverage does not mean there is no valid disposal route. It only means WhatBin does not have enough reviewed information to give one.

The goal was never to make the agent answer everything. I wanted the app to make it clear when the evidence supports an answer and when it does not.


Demo

Live demo:

https://bin.zura.tr

Select English from the language menu in the top-right corner.

A good example to start with is an old mattress.

  • Choose Hanoi and confirm the item.
  • Open the disposal guidance and source links.
  • Ask: "Why does this rule apply?"
  • Then ask: "How does this differ in Ho Chi Minh City?"

This example shows why location and source history matter.

The reviewed mattress guidance for Hanoi is different from the guidance for Ho Chi Minh City. In Ho Chi Minh City, a later decision also removed provisions related to local collection-point arrangements.

WhatBin does not try to fill that gap by making up an address, collection schedule, or fee.

You can also try:

  • a used household battery
  • a standalone rechargeable lithium-ion battery
  • a fluorescent lamp
  • a mercury thermometer

These are useful examples because item identification matters too. A standalone lithium-ion battery, for example, should not automatically be treated the same way as an entire power bank.

The video also shows what happens when WhatBin cannot safely return a disposal instruction.

Instead of asking the model to improvise, the system keeps the result as UNKNOWN or CONFLICT.


Code

GitHub repository:

https://github.com/Zura1555/whatbin

The project includes:

  • A browser UI for text input, photos, item confirmation, and follow-up questions.
  • A Node.js API for item recognition, deterministic rule resolution, and streamed explanations.
  • Sanity Studio for managing reviewed rules, supporting evidence, and conflicts.
  • An explanation agent built with the AI SDK and @ai-sdk/mcp.

The main MCP integration is here:

https://github.com/Zura1555/whatbin/blob/main/app/context-agent.mjs

The Sanity content models are here:

https://github.com/Zura1555/whatbin/tree/main/studio/schemaTypes


How I Used Sanity

Sanity holds the reviewed decision model

I am not using Sanity only as a place to store text for the LLM to search.

The records in Sanity decide whether WhatBin is allowed to show disposal guidance in the first place.

Each disposalRule contains information such as:

  • canonical item ID
  • human-readable item name
  • jurisdiction
  • disposal category
  • reviewed instruction
  • effective start date
  • optional end date
  • official source references
  • article or clause citations
  • supporting passages
  • source versions
  • related sources needed to understand the evidence

This means the application can work out whether a rule applies before the explanation agent runs.

The model does not have to infer applicability from a collection of loosely related paragraphs.

Source roles matter

I also keep different source roles separate.

For example, a source can represent:

  • a binding rule
  • an agency clarification
  • a record used to confirm whether another source is still current

This matters because an official logistics note may be useful evidence, but that does not automatically make it equivalent to a binding disposal requirement.

Currentness is handled before the agent runs

The resolver reads published content only and checks each rule against the selected city's local date.

Draft rules do not count.

Future rules do not count yet.

Expired rules no longer authorize current guidance.

This part is handled in normal application logic rather than by the LLM.


What I Pointed Sanity Context At

I pointed Sanity Context at the structured content in my own Sanity project.

  • Project ID: xqeddep2
  • Dataset: production
  • Dataset visibility: Public published content
  • Knowledge Base: WhatBin disposal guidance.
  • Knowledge Base ID: kbvvvJAbi13v
  • Context source: Sanity dataset import

The Knowledge Base has a dataset import using this GROQ projection:

*[_type in ["disposalRule", "disposalConflict"]]{
  _id,
  _type,
  canonicalItemId,
  jurisdiction,
  validFrom,
  validUntil,
  itemName,
  disposalCategory,
  instruction,
  sourceReferences,
  supportingPassages,
  summary,
  claims
}
Enter fullscreen mode Exit fullscreen mode

The configured import currently reports 10 successfully distilled documents.

The agent works with WhatBin's structured rule and conflict records, including their source references and curated supporting passages.

I did not separately import every linked government website or PDF into the Knowledge Base.

Instead, the Sanity records keep the source URLs, citations, source versions, and relevant passages needed to explain where the reviewed guidance came from.


How the Agent Uses Sanity Context MCP

The server connects to the Sanity Context MCP endpoint over HTTPS using a server-side Context Viewer token.

When someone asks for an explanation, the integration checks which MCP tools are available and follows a controlled retrieval sequence.

Tool How WhatBin uses it
initial_context Establishes the endpoint context first
knowledge_base_read Reads Knowledge Base content when available
schema_explorer Inspects the Sanity content structure
groq_query Queries relevant structured records
array_field_reader Reads nested fields when needed

The exact sequence depends on the tools exposed by the endpoint.

For required steps, I expose only the tool needed for that step. This prevents the model from skipping ahead and choosing another retrieval method simply because it is available.

Before starting an explanation, the server resolves the selected item and city again.

The agent receives:

  • the user's question
  • a limited amount of conversation history
  • the selected item
  • the relevant city
  • the date
  • the verified resolver result

It then retrieves evidence for that specific scope.

The agent does not decide which disposal rule should apply.


What the Agent Does With the Evidence

For a MATCHED result, the agent explains why the reviewed rule applies and which evidence supports it.

For a CONFLICT, it can explain the competing claims and their citations, but it cannot choose which one is correct.

For an UNKNOWN, it explains that WhatBin does not currently have enough reviewed coverage.

It does not invent a fallback disposal route.

The agent can cite source titles and article or clause names in its explanation.

Verified URLs are displayed separately by the app rather than generated by the model.


Comparing Hanoi and Ho Chi Minh City

WhatBin only allows a clean cross-city comparison when the same item independently resolves to MATCHED in both cities.

The Hanoi result and the Ho Chi Minh City result are resolved separately.

The agent cannot take a valid Hanoi rule and treat it as valid guidance for Ho Chi Minh City.

If one city returns UNKNOWN, that uncertainty stays visible.

If one city returns CONFLICT, the agent can explain the conflict but does not turn it into a confident answer.


The Main Result Does Not Depend on the LLM

I also wanted the reviewed disposal result to remain usable when the AI explanation layer is unavailable.

If the model provider or Sanity Context service fails, an existing MATCHED result still works.

Users can still see the reviewed instruction and direct source links. They only lose the additional explanation.

An LLM outage should not make already reviewed guidance disappear.


Why Keyword Search Wasn't Enough

Finding a paragraph containing the word "mattress" does not tell me whether that paragraph should be shown to someone in Hanoi today.

There are still several questions to answer:

  • Does it apply to Hanoi or Ho Chi Minh City?
  • When did it become effective?
  • Has it expired?
  • Was it changed by a later decision?
  • What type of source is it?
  • Is there an active conflict recorded by a reviewer?
  • Is it actually about the same item?

These relationships are represented directly in Sanity.

Canonical item IDs prevent similar products from automatically being treated as the same thing.

Jurisdictions keep local rules separate.

Effective dates distinguish current guidance from future or expired content.

Source roles and supporting passages make the basis of the guidance easier to inspect.

Conflict records make disagreement explicit instead of hiding it behind a confident AI answer.


Handling Conflicting Evidence

Conflicts have their own Sanity document type: disposalConflict.

A conflict can contain:

  • the affected item
  • jurisdiction
  • competing claims
  • source URLs
  • source versions
  • exact citations
  • review status

A published active conflict takes priority over an otherwise matching disposal rule.

The agent can explain the different claims and where they came from.

It cannot decide which claim wins.

That remains a human review decision.


Research and Publication Are Separate

Sanity Studio also includes a source-research workflow for coverage gaps.

The workflow can help an editor investigate possible sources and produce a preview or gap report.

It does not automatically publish the result as disposal guidance.

A human still has to review the evidence and publish the record.

I wanted to keep that boundary clear because finding a potentially relevant source and deciding that residents should act on it are two different things.


Sanity Project Details

  • Project ID: xqeddep2
  • Dataset: production
  • Dataset visibility: Public published content
  • Content models: disposalRule, disposalConflict
  • Knowledge Base: WhatBin disposal guidance.
  • Knowledge Base ID: kbvvvJAbi13v
  • Context source: Sanity dataset import

You can inspect published disposal rules through the public Sanity API:

https://xqeddep2.api.sanity.io/v2025-02-19/data/query/production?query=*%5B_type%20%3D%3D%20%22disposalRule%22%5D%7BcanonicalItemId%2Cjurisdiction%2CvalidFrom%2CvalidUntil%7D

The repository also contains the schemas, resolver, MCP integration, and the architecture behind the separation between reviewed guidance and agent-generated explanations.


Final Notes

The most important design question for me was not how much the agent could do.

It was what I should not let it do.

WhatBin does not ask the model to decide which disposal instruction is legally or operationally correct for a resident.

That decision comes from reviewed and published Sanity content plus deterministic application logic.

The agent has a smaller job: retrieve the relevant evidence and explain it in a way that is easier to understand.

And when WhatBin does not have enough reviewed evidence, it says so.

Sometimes UNKNOWN is the better answer.

Top comments (0)