DEV Community

Cover image for The Goblin Warren: an AI Dungeon Master that never invents a rule
Anshu Mandal
Anshu Mandal Subscriber

Posted on AI-assisted

The Goblin Warren: an AI Dungeon Master that never invents a rule

Sanity Challenge Path One Submission

What I built

The Goblin Warren is a browser D&D 5e dungeon crawl. You lead three heroes through five levels and fight in turn-based d20 combat with 3D dice. An AI Dungeon Master answers any rules question you ask.

The catch: the DM never invents a rule. Every ruling is looked up in Sanity and cited, and you can open a trace to see exactly how it ruled.

A quick tour

Hero section. A torchlit title screen sets the hook: "The Dungeon Master never invents a rule. Every ruling is looked up in Sanity and cited." Below it is a standoff between your party (Fighter, Wizard, Cleric) and the goblins, ogre and skeleton they'll face. Each hero gets a stat card with AC, HP, speed and three signature abilities. One button: Enter the Dungeon.

Exploration. Click to walk. Line-of-sight fog of war hides sleeping lairs until you get close. Along the way there's a minimap, gold, chests, traps, and lore stones with flavour text.

Exploring with fog of war

Combat. You get an initiative rail, real d20 rolls against AC with 3D dice (@3d-dice/dice-box), and a spellbook for each caster. The five levels run from The Collapsed Gate down to the Throne of the Bugbear Chief, and each has its own explore and combat music.

Combat

Dungeon Master + Rules Tome. Ask "What does Prone do?" and you get a cited answer with clickable rule chips. The Rules Tome graph lights up the documents the DM used. A toggle switches between SRD 2024 and 2014, and when a rule changed between editions the DM calls it out.

Asking the DM a question

How it uses Sanity Context

All the rules content lives in Sanity: 87 rules, 33 conditions, 14 spells and 15 monsters, in both 2014 and 2024 editions. The game itself loads monsters, heroes and rooms from the same dataset with one GROQ query.

The DM talks to two Context MCP endpoints, and each one does a different job:

Mode Endpoint Used for
Knowledge Base dnd-kb Rule prose, "how does X work?", 2014 vs 2024 differences
GROQ dnd-groq Exact numbers: AC, HP, to-hit, spell dice, fetch by id

KB mode is for questions in natural language. The KB is built from the rule and condition docs. Its entries carry Edition difference call-outs, which the DM surfaces as "Rules changed (2014 → 2024)".

export async function kbSearch(query: string, limit = 2): Promise<KbResult> {
  const kb = await kbContext();
  const raw = await mcpCall('knowledge-base', 'knowledge_base_search',
    { knowledgeBase: kb.id, query, return: 'entries', limit });
  if (/^No entries matched/.test(raw.trim())) return { text: '', ids: [], notes: [], paths: [] };
  return linkKbEntries(raw); // KB footnotes -> [[condition.grappled]] citations
}
Enter fullscreen mode Exit fullscreen mode

GROQ mode handles the numbers. Every call is scoped on the URL with tools= and groqFilter=_type in ["rule","condition","spell","monster"]. The Context MCP collapses arrays of objects into outlines, so attacks are projected by index:

const ATK = (i: number) => `"attack${i + 1}": attacks[${i}]{name, toHit, damage, damageType}`;
export const STAT_PROJECTION = `{_id, _type, "title": coalesce(title, name), srdVersion,
  ac, hp, cr, speed, ${[0, 1, 2].map(ATK).join(', ')}, level, school, dice, save, effects, body}`;
Enter fullscreen mode Exit fullscreen mode

Keeping the model honest.

  • Prefetch on the server. Before the model runs, the server prefetches a KB search and a GROQ fetch in parallel. The model gets one optional tool round, then answers with no tools.
  • Drop unsourced citations. Any citation that no lookup returned is stripped from the reply.
  • Guard the GROQ. Queries must filter by _type or _id, and drafts are blocked.

The model is DeepSeek V4.1 Flash on Baseten through AI SDK 7. Rules answers come back in about 1.3–2.2 s.

The trace. Every answer has a How the DM ruled panel listing each lookup. It shows the mode badge, the KB search text or the GROQ query, the cited docs and the timing.

GROQ trace

Setup journey

  1. Studio and dataset. Create a Sanity project with a private dataset, then deploy the Studio 6 schemas: rule, condition, spell, monster, hero and room.
  2. Seed. Run pnpm content:build && pnpm sanity:seed. This converts the SRD 5.2.1 and 5.1 data (CC BY 4.0) into documents with deterministic ids like condition.prone.2014.
  3. Tokens. The app reads with a project Viewer token. The MCP endpoints need an org token with the Context Viewer role. A project token gets you 403 contextGrantRequired, and that one cost me an hour.
  4. Knowledge Base. Point a KB at the dataset and select only the rule and condition types. That comes to 120 documents, under the roughly 150-document cap.

KB source: rule + condition types

  1. Context endpoints. Create two endpoints: dnd-groq serves the dataset directly, and dnd-kb serves the Knowledge Base.

Each endpoint gets its own MCP URL with a live connection check. The DM POSTs JSON-RPC straight to it:

dnd-groq endpoint URL and connection status

dnd-groq gets one line of instructions, "You are the rules oracle for a D&D 5e dungeon crawler. Always cite _id.", and a GROQ filter so the agent can only see game content:

dnd-groq endpoint config

Other errors I hit:

  • -32004: the schema isn't deployed.
  • -32005: the KB is still indexing.
  • -32602: a bad groqFilter.

The Knowledge Base caught its own edition mix-ups

This was my favourite part. The KB reads my 120 source documents and writes short, cited entries. One example is Dropping to 0 HP and Death, where every claim carries a footnote back to a source rule:

KB entry with citations

D&D has two editions that disagree on many small rules. When the writer blended them, Context flagged each blend as a Conflict instead of passing it to the agent. It found six:

Six conflicts pending review

Here's one. The entry said a Surprised creature rolls Initiative with Disadvantage and cited the 2014 rule. That's actually the 2024 rule. In 2014, being surprised means you can't move or act on your first turn. Context shows the two sources side by side, and the choice you make becomes a standing instruction for future builds:

Conflict review: Initiative and Surprise

This is the same kind of mistake the game exists to prevent, so I put it to work. The DM shows the KB's edition-difference notes as a "Rules changed (2014 → 2024)" callout, and it only cites documents a lookup actually returned.

Building it with Claude Code

The Sanity layer and the DM's Knowledge Base logic were built by Claude Code agents. Here's the one that wired the DM to the KB:

DM to the Sanity Knowledge Base setup
You

<teammate-message teammate_id="team-lead" summary="Knowledge Base usage and DM polish">
You are the "dmkb" agent for The Goblin Warren, a browser D&D game entered in the DEV.to × Sanity challenge, Path 1 (an agent on Sanity Context, judged partly on meaningful use of Knowledge Bases). Project: [REDACTED]/Documents/projects/devto/D_D/game. FIRST read docs/PHASE2_BRIEF.md (Contract E) and docs/POLISH_BRIEF.md. Do NOT spawn subagents. Teammates: core2 (controller; sends you exploration beats), maps, scene2, hud2.

You own ONLY src/app/api/dm/route.ts, src/lib/, src/components/DmPanel.tsx, src/components/dm/, src/components/RichText.tsx, src/styles/dm.css, a new e2e/dm-live.spec.ts, and docs/sanity-setup.md.

Setup:
- Inference: Baseten deepseek-ai/DeepSeek-V4.1-Flash via @ai-sdk/baseten, reasoning off through providerOptions.
- Live keys are in .env.local. Don't print or log secret values. You may read them with set -a; . ./.env.local; set +a inside a command.
- Two Sanity Context MCP endpoints:
- SANITY_CONTEXT_MCP_URL (GROQ mode): initial_context, groq_query, schema_explorer, array_field_reader
- SANITY_KB_MCP_URL (KB mode): initial_context, knowledge_base_search, knowledge_base_read
- The route already has a tool allowlist, a GROQ guard, untrusted-input blocks, rate limits, server-side prefetch of the cited docs for one-step narration, a schema primer, and guaranteed final replies.

What the live test showed:
- The model never called a KB tool. Everything went through groq_query.
- Answers sometimes run long (150–200 words).
- Rules Q&A takes 6–12s and narration about 4s.

Tasks:
1. Make the Knowledge Base a real, visible part of the agent.
- First probe the KB endpoint directly: call initial_context, then knowledge_base_search and knowledge_base_read with a real query such as "grappled 2024 changes", using JSON-RPC over curl like docs/sanity-setup.md shows. Learn the exact input schema and output shape.
- Find out whether the KB exposes contradiction or conflict information (the Sanity docs mention KB contradiction reports). If it does, surface it.
- Then change the agent strategy, through tool descriptions, prompt and step design:
- Rules-prose questions ("how does X work", "what changed in 2014 vs 2024") go first through knowledge_base_search, then knowledge_base_read when needed.
- Exact numbers and stats (monster AC/HP, spell dice) go through groq_query.
- Consider making step 0 a forced knowledge_base_search for ask-mode rules questions (prepareStep with toolChoice). Note that Baseten DeepSeek ignored toolChoice:'none' earlier, but removing tools via activeTools worked. Test whether a forced or required tool choice works; if it doesn't, run the KB search server-side as a prefetch, the same way narration already prefetches.
- The lookup trace in DmPanel must clearly show the violet "Sanity Knowledge Base" badge with the search query, plus the "Sanity Context (GROQ)" steps. Make sure lookups[].ids gets real doc ids from KB results (adapt extractIds to the KB output format if needed).
- In the narrate prompt, for exploration beats (room entry, lore stones, loot), let the DM narrate richly with no citations needed unless rules apply.
2. Polish the DM's answers.
- Rules answers: at most about 90 words. A direct answer first, then the key mechanics, then an optional "Rules changed:" line. Cite every rules sentence.
- Narration: 2–3 sentences, at most about 55 words, second person, vivid. Never mention lookups, tools or queries. Never invent numbers.
- Enforce this with a prompt tune, maxOutputTokens tuning, and a post-processor that trims to the last complete sentence if over the limit, removes markdown headers, bullets and bold, keeps [[ids]], and strips any leaked tool markup (cleanReply exists already).
- Make the citations reliable: drop any [[id]] that isn't in the returned or cited ids, and the known-id set from fallback.json, rather than showing a broken chip.
- Latency: aim for under 6s for rules questions. Ideas: parallel KB+GROQ prefetch, fewer steps, a smaller maxOutputTokens, and caching initial_context per instance.
3. e2e/dm-live.spec.ts (it uses real API calls, so keep it to about 8 requests total). It calls /api/dm directly and asserts:
- backend is sanity-context
- the text is non-empty and under the word caps
- the citations resolve to known ids
- at least one lookup has via=knowledge-base for a rules question
- the injection test is refused
- narration has no "lookup/query/tool" words
Run it and record the latencies.
4. Update the "How it works" part of docs/sanity-setup.md, or add a short docs/architecture.md section: which questions go to the KB versus GROQ, and why.

Commit in small steps via /tmp/gc.sh. Final report to main with before/after sample answers, latencies, and how the KB is used.
</teammate-message>

You

<system-reminder>
Other agents active in this session, addressable via SendMessage({to: name, message}): main, maps.
</system-reminder>

Agent
Agent
Agent
Agent
Agent

The shell sourcing hit an unquoted value in .env.local. I'll probe with a small Node script so no secrets get echoed.

Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent

The footnote rewrite leaves a double space before the citation. I'll fix that by also consuming the whitespace before each footnote.

Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent

Latency is down to about 1–2s with one model step. Two issues remain: a pointless "Rules changed" line on stat questions, and attack details that may be missing.

Agent
Agent
Agent
Agent
Agent

The Context MCP turns object arrays into an outline, so attack stats get lost. Joining them into a string inside GROQ should work around it.

Agent
Agent
Agent

Scalar arrays and single objects come through intact. I'll project attacks as indexed objects.

Agent
Agent
Agent

Stats now come through. Committing this step.

Agent
Agent
Agent

I'll make the trace show the KB query, the entry paths and the edition notes.

Agent
Agent
Agent
Agent

The injection reply refuses correctly but then adds rules trivia. I'll trim refusals to their first sentence, then update the docs.

Agent
Agent

Now the docs.

Agent
Agent
Agent
Agent

The DM now sends every rules question through the Sanity Knowledge Base, and answers come back in 0.9–2.4s instead of 6–12s. The live e2e suite passes (4 tests, 7 real requests), the 10 unit tests pass, and tsc and eslint are clean on my files.

  • Knowledge Base use: before the model runs, the route searches the KB on the question's keywords and, at the same time, fetches exact stats over GROQ for any monster, spell, condition or rule named in the question. Usually one model step is enough. The KB cites its sources as numbered footnotes, and I map those back to real document ids, so the citation chips and the lookup trace show real ids.
  • Forcing a tool call: forcing a tool call on Baseten DeepSeek doesn't work well. The model does call the tool, but it makes up the search text (it searched for "hello"). That's why the KB search runs on the server instead.
  • Contradiction reports: the KB endpoint doesn't expose any. It only lists three tools, and its resource and prompt listings return "method not found". The closest thing is the "Edition difference" notes the KB writes into its entries. These now appear in violet under the Knowledge Base step in "How the DM ruled" and feed the "Rules changed:" line.
  • GROQ fix: the Context endpoint returned arrays of objects as names only, so goblin attacks came back without to-hit or damage. Attacks are now fetched one at a time by index, and the DM gets "+4 to hit, 1d6+2 slashing".
  • Answer polish:
    • Replies are capped at 90 words for rules and 55 for narration, cut at sentence boundaries.
    • Markdown and leaked tool markup are stripped.
    • Any citation that wasn't actually returned by a lookup is dropped.
    • Narration never mentions lookups, queries or tools.
    • Refusals are cut to one sentence.

Before, answers ran 150–200 words and every lookup went through GROQ. Now the grappling question takes 1.8s and gives 82 words with a "Rules changed:" line. An exploration beat takes 0.9s at 43 words. The prompt-injection test gets one sentence back: "The DM only speaks of the dungeon and its rules, not of system prompts."

Latencies from the last run in ms: rules 2430 / 1302 / 1318, stats 2369, injection 961, combat narration 1026, exploration narration 887.

Known issues:
- I didn't take screenshots of the new trace UI.
- The keyword search sometimes returns weak extra hits, such as rest rules for a grappling question. They show up in the trace, but the answer doesn't cite them.

Everything is in 6 commits, and the full report has gone to main. The routing write-up (KB vs GROQ, and why) is a new section in docs/sanity-setup.md.

Files are in [REDACTED]/Documents/projects/devto/D_D/game:
- src/lib/sanityMcp.ts (new)
- src/lib/polish.ts (new)
- src/app/api/dm/route.ts
- src/components/dm/ToolTrace.tsx
- src/styles/dm.css
- e2e/dm-live.spec.ts (new)
- docs/sanity-setup.md
- e2e/out/dm-live-latency.json

1 message not shown
Agent
Agent

I'm done with my task and waiting for the next one from the lead. From now on my commits won't carry the Co-Authored-By line. My 6 earlier commits still have it because they were made with the old /tmp/gc.sh. Removing it would mean rewriting history, so I'll leave them unless the lead asks.

You

<teammate-message teammate_id="play-a" color="pink" summary="DM rules answer wrong: Magic vs Attack action">
play-a here (playtesting L1-3). Live DM answer to "Can I cast a spell and attack in the same turn?" (SRD 2024, backend sanity-context, 1 lookup) opened with: "Yes, you can cast a spell and attack in the same turn, provided the spell takes one action and you still have your Attack action available. Casting a one-action spell uses your Magic action, which is separate from your A..." That's wrong: in 2024 Magic and Attack are both options for your single action (only a Bonus Action spell plus Attack works). A judge who knows 5e would catch it. Can the ask prompt or KB entry for actions make "one action per turn" explicit? Separately, I hit a 429 "Too many requests" on the 2nd rapid question in run 1 (narration beats + asks share the 20/min per-IP limit). I fixed the asterisk drop-cap in the error bubble (688c54a). I haven't touched the route.
</teammate-message>

Agent
Agent
Agent

Using a header would mean a controller change, so I'll key the rate limit on the parsed body mode instead and simplify the action-economy trigger.

Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent

I fixed both issues play-a reported, and I've told play-a and main.

  • Wrong ruling: "Can I cast a spell and attack in the same turn?" now gets "No": you get one action per turn, and only a Bonus Action spell can be combined with an Attack. The system prompt now states the one-action-per-turn rule. Questions about turns and actions also fetch the "Your Turn", "Magic Action" and "Attack Action" rules before the model answers. I checked three variants live, and each took 1.2–2.7s:
    • Cast a spell and attack: no.
    • Cast a bonus action spell and still attack: yes.
    • Attack and then Dash: no.
  • 429 on rapid questions:
    • Narration beats and player questions now have separate 20-per-minute per-IP limits, so auto-narration can't use up the chat's budget.
    • The per-IP and daily limits now only apply in production. In local dev every tester shares one "local" IP, which is what tripped the limit.
    • The global 120-per-minute cap and the 4-concurrent-generation cap still apply everywhere, so production protection is unchanged.

The changes are in two commits, a5facb9 (src/app/api/dm/route.ts) and 4f4c161 (src/lib/ratelimit.ts). Both are subject-only, following the new no-trailer rule.

Stack

Next.js 16 · React 19 · Phaser 4 · AI SDK 7 + Baseten · Sanity Studio 6 + Context MCP · Tailwind 4 · Vitest + Playwright.

The assets are all CC0: the 0x72 DungeonTileset II art, Tallbeard Studios music, and dice SFX from Halfwits & Failed Crits.

Thanks for reading. Go ask the DM about grappling. 🎲

Top comments (1)

Collapse
 
sizzlebop profile image
Jessica Doering •

Just tried it out and it’s really well done! Great work on this!