DEV Community

Cover image for Precedent: an agent that knows which Solidity security advice has been overruled
Shikhar Verma
Shikhar Verma

Posted on

Precedent: an agent that knows which Solidity security advice has been overruled

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

Security advice for Solidity goes stale, and stale advice in this field is expensive. Docs, EIPs, audit reports and old blog posts contradict each other, and much of that disagreement is just time passing: the compiler and the EVM changed underneath the advice. A keyword search or a plain LLM returns the outdated answer with full confidence, because both only see text, not the fact that one piece of text replaced another for a given version.

Precedent treats security guidance like case law. Every claim carries a date and a compiler-version scope, and newer rulings overrule older ones. Ask "is transfer() still safe on 0.8.28?" and you get:

  • a ruling for your exact version,
  • a docket: what the older advice said (struck through) and what overruled it, each linked to its original source,
  • a contested status when sources genuinely disagree and nothing controls, instead of a guess.

It covers [[5]] patterns in depth: [[LIST, e.g. ETH sends with transfer/send, selfdestruct, SafeMath, ERC-4626 inflation attacks, and tx.origin as a no-change control]]. The control case matters: it checks that the agent doesn't invent drift where none exists.

Precedent summarizes published sources. It is not an audit, and every verdict says so.

Who it's for: Solidity developers who need a citable answer for their compiler version, reviewers doing quick triage, and hackathon builders who don't want to ship a pattern that was deprecated last year.

Demo

Live: [[DEPLOYED URL]]

[[VIDEO EMBED OR GIF: a verdict resolving, with overruled claims struck through and the controlling claim highlighted]]

Try these questions [[replace with the questions that work well in your build]]:

  • [[Question 1: a stale-trap pattern with a version]]
  • [[Question 2: a genuinely contested pattern]]
  • [[Question 3: the stable control, which should show no drift]]

Code

[[GITHUB REPO URL]]

How It Works

Question + compiler version
        │
        ▼
  Precedent agent (Vercel AI SDK, multi-step tool loop)
        │
        ├──► Context MCP endpoint A (Knowledge Base)
        │       what do the sources say? where do they conflict?
        │
        ├──► Context MCP endpoint B (dataset, GROQ)
        │       verified claims in scope for this version,
        │       with supersession and sources
        ▼
  Verdict: status, stance, claim IDs only
        │
        ▼
  UI fetches claim text, dates, URLs from Sanity by ID
        │
        ▼
  Ruling card + docket timeline
Enter fullscreen mode Exit fullscreen mode
  1. Identify the pattern and version from the question and the version selector. If no version is given, the agent assumes the latest in the dataset and says so.
  2. Read the sources through the Knowledge Base endpoint to see what they say and where they conflict.
  3. Query the claims through the GROQ endpoint: the verified claims whose version range covers the version, plus which claims supersede which.
  4. Resolve. Controlling claims are in-scope claims that no other in-scope claim supersedes. If their stances agree, the status is settled. If they differ, it's contested and no winner is picked. If none exist, it's out of scope.
  5. Return IDs, not prose. The agent emits a structured verdict containing claim IDs. The UI then loads the text, dates and URLs from Sanity.

Step 5 is deliberate. The model never writes a source, a URL or a date, so it can't fabricate a citation. The API route also re-checks that every ID exists and recomputes the status itself; on a mismatch the user sees an error, not a wrong answer.

How I Used Sanity

The agent reads Sanity through two Context MCP endpoints, because they do different jobs. An endpoint backed by a Knowledge Base serves Knowledge Base mode, and an endpoint backed by a dataset serves GROQ mode, so I use one of each.

Endpoint A: Knowledge Base

I pointed Sanity Context at [[N]] source documents: [[e.g. Solidity docs and changelogs, EIPs, OpenZeppelin docs and release notes, N public audit reports]]. Sanity distills them into a navigable Knowledge Base where each entry stays linked to the source it came from, and where conflicting claims surface side by side with their sources. [[Add what you saw: how many entries it produced, an example of a contradiction it surfaced on its own, and any decision you made about a conflict.]]

I stayed well within the beta's ~150-document budget on purpose: a small, curated set of primary sources beats a large noisy one.

Endpoint B: the dataset, in GROQ mode

Knowledge Base mode can't express "this claim applies only to 0.8.x and was replaced by that one", so I modeled that as structured content:

Type What it holds
pattern A named pattern, a short summary, and aliases
source Title, URL, publisher, kind (docs, EIP, audit, release notes, blog), publish date, license
claim A paraphrased statement (280 characters max), a stance (safe, unsafe, deprecated, mixed), fromVersion / toVersion, supersedes[], contradicts[], and a confidence flag

Compiler versions are stored as numbers (for example 0.8.28 becomes 8028) so GROQ can compare them. Only claims I had checked by hand against the primary source are marked verified, and only those are used. [[N]] claims across [[5]] patterns.

The ruling is computed, not stored. This query returns every in-scope claim and which other claims supersede it:

*[_type == "claim" && confidence == "verified" && pattern._ref == $patternId
  && (!defined(fromVersion) || fromVersion <= $v)
  && (!defined(toVersion)   || toVersion   >= $v)]{
    _id, statement, stance,
    "source": source->{title, url, publishedAt},
    "supersededBy": *[_type == "claim" && ^._id in supersedes[]._ref]._id
  }
Enter fullscreen mode Exit fullscreen mode

The controlling claims are the ones with no in-scope claim in supersededBy. Because the ruling is derived, fixing one claim or adding a newer one changes every affected answer immediately, with no re-curation.

Why structure matters here

Keyword search finds passages that mention a pattern. It has no way to know that one passage replaced another for a specific compiler version. The supersedes chain and the version scope are the content model, and the agent's correctness depends on them. The Knowledge Base shows what the sources say; the dataset says which statement controls.

Tools the agent used

[[List the actual tool names each endpoint exposed, discovered at runtime, and what the agent used each one for. Example format: tool_name (endpoint A): used to ... ]]

The Context endpoints are read-only, so the agent can't change the content. Claims are seeded by a script and edited in Studio.

Does It Actually Beat the Alternatives?

I wrote [[N]] "Stale-Trap" questions where the obvious answer is outdated: [[N]] stale traps, [[N]] genuinely contested cases, [[N]] stable controls, and [[N]] out-of-scope. I labeled the ground truth by hand from the verified claims, then ran three systems with the same model, same output schema, and [[3]] runs per question:

  • A: keyword search over the raw sources (top passages given to the model, no tools)
  • B: Knowledge Base only (Endpoint A without the claims layer)
  • C: Precedent (both endpoints)

Because every system returns the same structured verdict, grading is deterministic rather than a model judging a model: an answer is stale if it matches a superseded claim, and a citation is correct if it includes a controlling claim.

System Stale-answer rate (lower is better) Cites the controlling source Says "contested" when it should
A: keyword search [[X%]] [[X%]] [[X%]]
B: Knowledge Base only [[X%]] [[X%]] [[X%]]
C: Precedent [[X%]] [[X%]] [[X%]]

[[Two or three sentences in your own words: which system did best, where B failed and why, and one specific question that shows the difference.]]

Design Decisions

  • IDs only from the model. Prevents fabricated citations.
  • Contested is a first-class outcome. A security tool that picks a winner between genuinely conflicting sources is worse than one that says so.
  • Verified claims only. Unreviewed claims never reach the ruling path.
  • No fallback to model memory. If the endpoints fail, the user gets an error, not a guess from the model's training data.
  • Small and deep. [[5]] patterns done properly beat 50 done loosely.
  • A control case. tx.origin has no drift, so a correct agent should report a stable answer.

Guardrails

Precedent never writes exploit code and never audits user code. Every verdict shows "Summarizes published sources. Not an audit." Claims are my paraphrases with links to the originals, and I used permissively licensed sources where possible.

Challenges and What I Learned

[[Write 3 to 4 honest items from your actual build. Prompts to answer: What broke first? What did the Knowledge Base do better or worse than expected? Where did the model misread supersession, and how did you fix it? How long did verifying claims by hand take, and what did you find?]]

Limits, Honestly

  • [[N]] questions is a small sample, so treat the percentages as indicative, not conclusive.
  • Only [[5]] patterns are covered.
  • Claims are paraphrased and hand-verified, so errors are possible. Check the linked sources before relying on a ruling.
  • [[Knowledge Bases are in beta; mention anything that didn't work as expected.]]

What's Next

  • More patterns and more eras of the EVM
  • Paste a code snippet and detect which patterns it touches
  • A compare mode: the same pattern across two compiler versions
  • [[Anything else you genuinely plan]]

Built with: Next.js, TypeScript, Vercel AI SDK, Sanity Studio and Sanity Context, [[model name and provider]], deployed on [[Vercel]], with Google Antigravity as the coding environment. (keep only if true)

Sanity Project Details

Project ID: [[YOUR SANITY PROJECT ID]] (dataset: [[DATASET NAME]])
[[OR: public dataset URL, and confirm the dataset is public]]

Top comments (0)