DEV Community

MANIFESTA
MANIFESTA

Posted on

Bot Lawyer: an agent that knows which of Discord's own pages is lying

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

On June 10, 2026 Discord changed the rule every bot developer hits first. The privileged intent threshold went from "100 servers" to "10,000 unique users", verification got split from intent review, annual reapplication appeared, and about half of Discord's own pages got updated. In my 2026-09-19 snapshot the application flags table still says 100 servers, and the gateway page says 10,000 users at the top of a section and "after your app is verified, you can apply" 25 lines lower. Same site, same day, two rules ๐Ÿ‘€

I build and run communities for a living, developer communities included, and I build my own bots and tools for them because the ones I want never exist. So I have been on both sides of this: the mod who has to explain to a dev why their bot got denied, and the dev reading the old page while the rule changed underneath me. There is a specific feeling when a support article, the API docs and a Discord staff comment on GitHub all give you a different number and you have 100 servers waiting. Bot Lawyer is the lawyer I wanted at 2am that night, minus the invoice ๐Ÿ’€ You tell it your situation (users, servers, which intents, verified or not), it rules yes, no or depends, quotes the page verbatim with the URL and the effective date, and when two pages disagree it shows both sides and says which one it applied. It has a "Rule as of" date, so the same question as of 2026-05-01 and today gives two different rulings, both cited. It runs as a web app and as a real Discord app with a /ask command, and there is a public ledger of all 11 contradictions I found in Discord's docs.

It also does the thing people in my DMs actually ask for every week, which is never about intents: give it a server type and it proposes the whole channel structure with the emoji and separator naming people use (๐Ÿ‘‹โ”†welcome, ๐Ÿ“ขโ”†announcements, ๐Ÿ”ดโ”†live-streams), category headers, and every Discord constraint respected (1 to 100 characters, 50 channels per category, lowercase with no spaces for text channels). It says which parts are Discord rules and which parts are community taste, because it has both on file. Same for a welcome message or a rules post: you get a copyable block that renders correctly, with a character count against the 2000 limit.

Demo

Live: https://bot-lawyer.vercel.app

Same question, two dates, two rulings

If you want the five minute version:

  1. Ask "My bot is in 90 servers with about 4,000 users, do I need Message Content approved?" It says no, the threshold is 10,000 unique users since 2026-06-10, and under the answer you see the old 100 servers rule struck through with its dates.
  2. Set "Rule as of" to 2026-05-01 and ask about 120 servers. Now it says yes, and cites the rule that was in force from 2020-10-27 to 2026-06-09. Same bot, same data, different law.
  3. Ask "Is retry_after in seconds or milliseconds?" and watch it quote both pages with dates and pick one (C5 in the ledger).
  4. Ask for a welcome message for your gaming server. You get one copyable block and a character count against the 2000 limit, nothing else.
  5. Ask for channel names for a crypto community server: welcome, rules, announcements, live streams, general chat, trading, memes. You get a category layout with emoji, separator and lowercase names, and a line per constraint saying whether it is a Discord rule or just how people do it.
  6. Ask your own thing. Nothing is precomputed, every answer is a live GROQ query plus a knowledge base read.

Rule as of 2026-05-01 vs today
The unresolved C4 memo inside the Discord skin
Rules on file panel on a phone

Discord: the app is called BOT LAWYER, command /ask question:<...> as_of:<YYYY-MM-DD>. HTTP interactions endpoint, no gateway, no intents, so the bot obeys its own rulings. Add it to any server you own with this link and run /ask (global commands can take up to an hour to show up after install): https://discord.com/oauth2/authorize?client_id=1550945390789140661&scope=bot%20applications.commands&permissions=2048 . It only needs Send Messages. Counsel does not read replies, use /ask again.

Ledger: https://bot-lawyer.vercel.app/contradictions

Code

https://github.com/A1VARA5/bot-lawyer (MIT). Schema in studio/, seed generator and validator in seed/, the Next.js 16 app in web/, all 40 eval questions with grades in eval/, the contradiction file in research/. The README has the commands to run it yourself.

How I Used Sanity

I started by collecting Discord's public pages into a dated snapshot: the developer docs (they serve markdown on purpose and the source repo is MIT), the two help centers through Zendesk's public JSON API, the legal pages, and the discord-api-docs change log and issues. Snapshotted, not live, so the demo does not shift under you if Discord fixes a page next week; the snapshot stays out of the repo. Then I read all of it and wrote down every place two pages disagreed. That became research/CONTRADICTIONS.md, eleven of them, each with both quotes and both URLs. I have moderated arguments between humans for years and this was worse, because both sides are the same company and neither one shows up to the call. Then I modelled the content so an agent could reason over it instead of searching it.

The dataset: 136 rules, 58 sources, 11 policy areas, 40 eval questions, 11 documented contradictions (C1 to C11). One rule is one claim Discord makes, quoted verbatim from one page, with the conditions under which it applies: appliesWhen (users, servers, intents, verification state, region, owner age), effectiveFrom and effectiveTo, a source reference with an authority score (because Discord staff said on GitHub that support articles reflect user functionality, not bots), supersedes, conflictsWith, resolution and confidence. Of the first 100 rules, 55 are confirmed and 45 inferred, and the agent has to say "inferred" out loud when it uses one.

Two Sanity Context endpoints, because an endpoint with a dataset source ignores its knowledge base sources:

bot-lawyer-rules is GROQ mode over the dataset, filtered to _type in ["rule","source","policyArea"] so the agent cannot read the eval questions. Tools: initial_context, schema_explorer, groq_query, array_field_reader. It answers anything that depends on a number or a state.

bot-lawyer-kb is knowledge base mode over 90 snapshotted Discord pages (docs, both help centers, legal pages, GitHub issues). Tools: initial_context, knowledge_base_read. It answers "what does the page around the quote say" and holds facts I never turned into rules. Every ruling reads both: the rules give the verdict and the dates, the knowledge base entry gives the page around it, and the answer marks that part with "From the page:" so you can see which is which. Under each answer there is a "Pages read" line with the entries and their source pages, next to the "Rules on file" cards.

Rules on file and Pages read under the payout answer

The app merges both clients into one Vercel AI SDK tool set. Both initial_context payloads are inlined into the cached system prompt, and groq_query and knowledge_base_read are re described so the model routes by intent. The eval made 68 tool calls: 46 groq_query, 22 knowledge_base_read, zero schema_explorer or array_field_reader, because the inlined context already gave the schema.

The knowledge base took 3 builds and raised 8 issues. Build 2 raised 4: one real conflict Context found on its own (the Monetization Terms page says "Effective June 6, 2024" at the top and "As of June 4, 2024" for the Growth Tier lower down; two facts, not one, I resolved it as June 4 and it became an Instruction), and 3 merge suggestions I dismissed. Then I wrote a manual Instruction for the June 10 change, scoped to the application, gateway, verification, privileged intents and change log pages. Context replied that this rule contradicts facts in two other entries and offered to rebuild them. Yes please ๐Ÿค Build 3 raised 4 conflicts and Context found C1 by itself: "Thresholds: 10,000 unique users" against "Portal Issues: 100 servers", 8 sources. I resolved three. The fourth I left pending on purpose: App Directory says a support server is required before enabling discovery, Monetization SKUs never mentions one. Two processes, not a contradiction, and the only resolution offered is binary. Final count on the dashboard: 5 resolved, 3 dismissed, 1 pending, 5 active Instructions (4 from resolutions, 1 manual). Insights recorded 65 conversations.

Issues, one left pending on purpose
Five active Instructions

Quick sanity check I did on myself, because a bot that just searches the docs and reads them back is not worth anyone's time. Type "90 servers message content intent" into any search over Discord's pages and the top hit is the flags table: "100 or more servers". A real page, a real URL, a wrong answer, because that table was never updated in June. Bot Lawyer gets it right only because the old rule has effectiveTo set and the new one points at it with supersedes. Same with "retry_after seconds or milliseconds": search gives you both pages and a shrug. Here both rules carry conflictsWith, a written resolution, and each source has an authority score, so the answer is seconds and it can tell you why the support article lost. That is the whole point of putting this in Sanity instead of a vector database: the structure does the arguing.

What broke

  • The Studio deploy failed twice: studioHost missing, then @sanity/icons v5 has no named exports from the index, you import per icon. The best practices reference said exactly that. I did not read it closely enough.
  • The first knowledge base build was refused: "plan limit for indexed documents is reached (303 of 150 used)". My 177 page live crawl plus a ZIP of snapshots was double the cap. I deleted both and curated 80 files.
  • Build 1 ran 10 minutes and died with "The outline's structure stopped converging at thin_leaves (1 remaining) after 3 rounds". Build 2 resumed and passed.
  • Haiku skipped groq_query, answered the retry_after question from the knowledge base alone and said "both sources agree". They do not, that is C5. The prompt now mandates GROQ first and forbids "sources agree" unless the dataset shows no conflict. Haiku ignored the order anyway, so it runs on Sonnet.
  • My GROQ hint was wrong twice. area->slug == "rateLimits" returned nothing, because slug is an object and the seed used kebab case anyway. The ids are area.rateLimits.
  • The "case file" tool call lines never rendered. MCP tools arrive as dynamic-tool parts in the AI SDK stream, not tool-<name>.
  • The first design was black text on a black background. Very legal, very unreadable. I threw it out and rebuilt the whole UI as a Discord DM, because the audience lives in Discord and a memo embed with a green, red or yellow bar makes sense to them before they have read a word. The sidebar has a #billing-never channel. It is the only promise in this project I am fully confident about.
  • I burned a $7 test key in one afternoon of evals and took the public demo down with it ๐Ÿ’€ Every call was carrying both initial contexts uncached, about 15,000 input tokens, a hundred times over. Now the system prompt is cached, every request is metered, spend is capped at $3 a day, and if the key dies the route falls back to Vercel AI Gateway.
  • Anonymous reads on the public dataset come back empty. I thought I had broken something, so I created a fresh public dataset with one document: same result, "permission", while Sanity's own public demo project answers anonymous queries fine and the security page for my project says "2 public datasets". So it is the project, not the data, probably the trial. The ledger page reads with a server side viewer token for now and the question is in the Sanity Discord.
  • Bot Lawyer, the app that tells you to turn off privileged intents you do not use, had all three privileged intent toggles on in the Developer Portal until my own test pass caught it ๐Ÿ˜‚ Physician, heal thyself. It now runs with zero intents and mentions this in its own footer, which is the most passive aggressive thing I have shipped.

Eval

40 questions, written from the contradictions I had found, graded by a separate model against expected answers, spot checked by hand.

kind PASS PARTIAL FAIL items
trap 7 1 0 8
standard 22 6 1 29
unknown 2 1 0 3
all 31 8 1 40

Ruling right on 39 of 40. 0 of 8 traps fell for the stale 100 servers answer. 22 of 40 read the knowledge base, 12 surfaced a contradiction, 20.6 seconds average. The one FAIL was a real defect: a knowledge base summary sentence went into a blockquote attributed to a page that does not say it. I added a provenance rule (blockquotes only from a rule's quote field or the page's own words) and re ran the four answers that had that smell. All 12 blockquotes in the re run are verbatim in the page snapshots, checked by string match, not by eye. The table above is the original run; I am not going to re grade the whole thing to make the number prettier.

Honest limit: I wrote the questions and I built the dataset. This measures whether the agent reads the structure correctly, not whether the structure is complete or right. It is not a benchmark.

Cost: about 9 cents per question on Sonnet.

Sanity Project Details

Project ID v6745vem, dataset production (public). Studio: https://bot-lawyer.sanity.studio. Context endpoints bot-lawyer-rules (GROQ mode) and bot-lawyer-kb (knowledge base mode, 90 sources). Anonymous queries currently come back empty on this project even though the dataset is public (details under What broke); a read token or the Studio shows everything.

Agent Session

I built this in Claude Code and recorded the session. Being curated down to the bits worth watching: the icons fight, the thin_leaves build, the Haiku "sources agree" moment, and the afternoon I set fire to a $7 API key. Public before this post goes live.

Two things in here I did not invent: the two endpoint setup with merged tools and the initial context inlined into the prompt is from the Sanity Labs agent workshop, and every word of content is Discord's own docs, support articles, policies and change log. The opinions about which page is lying are mine.

The thing I still cannot solve: Context's pending issue wants me to pick a winner between two entries that are both right, which is also how most of my community disputes end. Is there a way to tell Context "not a conflict, keep both" that I missed, and if not, how are you handling it? ๐Ÿ”ฅ

And if you run a Discord server and your channel names are still general and general-2, you know where to find counsel.

sanitychallenge #devchallenge #buildinpublic #discord

Top comments (0)