DEV Community

Cover image for TableTop Arbiter: AI Tournament Judge
Ayush Gupta
Ayush Gupta

Posted on

TableTop Arbiter: AI Tournament Judge

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content


?? What I Built

TableTop Arbiter is a zero-hallucination AI Head Tournament Judge explicitly engineered for high-stakes competitive tabletop games like Magic: The Gathering.

The Problem: In competitive tabletop gaming, rules arguments frequently stall matches. While players try to consult modern LLMs for quick answers, standard models routinely hallucinate. They cannot distinguish between obsolete printed base rulebooks and newer official tournament errata, and they struggle to comprehend complex rule-layer dependencies.

The Solution: TableTop Arbiter solves this by grounding the AI entirely within a Sanity Structured Content Lake. Instead of relying on fuzzy vector embeddings or standard RAG, the Arbiter deterministically resolves contradictions via GROQ graph dereferencing. It streams its live tool execution trace and precise rule provenance directly into the user interface-leaving zero room for hallucinated rulings.


?? Demo

Live Web App: https://tabletop-arbiter.vercel.app

(Watch the 1-Minute Walkthrough Video below to see the Arbiter handle a complex ruling live!)


?? How It Works (The Sanity Setup)

1. The Knowledge Base

I modeled a custom tournament Knowledge Base in Sanity containing the core Comprehensive Rules and active Disputed Scenarios. The content is structured into three highly interconnected schemas:

  • GameRule: The base comprehensive rules.
  • RuleErrata: Official tournament patches that supersede base rules.
  • DisputedScenario: Frequently asked edge cases.

2. Sanity Context MCP Integration

I implemented the official Sanity Context MCP specification within a custom Next.js server (/api/sanity/mcp). I utilized real Sanity tools alongside custom GROQ executors:

  • initial_context: Instantly maps the active tournament rulebook dataset.
  • knowledge_base_read: Retrieves exact rule provenance verbatim.
  • groq_query: Specifically used for relational dereferencing. The agent actively queries *[_type == "ruleErrata" && references()] to check if a base rule has been officially patched.

3. A Concrete Example: Stack vs Battlefield Boundary

Here is the Arbiter resolving the infamous Carnage Tyrant vs Ward dispute:

  1. The Dispute: Player A casts a fight spell at a creature with Ward {2}. Player A controls Carnage Tyrant which explicitly says "This spell can't be countered". Does Ward counter the fight spell?
  2. The Execution Trace:
    • The AI fetches base rule CR 604.3a and cross-references CR 113.6.
    • Using groq_query, it checks for errata and discovers that "This spell can't be countered" is a static ability that functions solely on the stack, not on the battlefield.
  3. The Final Ruling: Because Player A cannot pay {2}, Ward triggers and counters the fight spell. The AI renders a verifiably proven "Ruling Slip" in the UI.

?? Honest Writeup: Challenges & Limitations

Building this highly constrained agent wasn't without hurdles:

  • Context Window vs Sanity Limit: We hit the ~150 document limit context constraint quickly when trying to load the entire MTG comprehensive rulebook. We had to pivot from loading everything to using Sanity Context tools dynamically-fetching only specific chapters (e.g., Layer 7 rules) when the agent recognizes a keyword.
  • Scope Limitation: Currently, the arbiter is heavily optimized for Magic: The Gathering. Scaling this to handle games like Warhammer 40k would require massively expanding the Sanity schema to handle spatial distance logic.
  • Future Roadmap: I plan to add real-time OCR so players can just snap a photo of the board state and have the Sanity MCP query the exact card Oracle texts automatically.

?? Code

GitHub logo Ayushgupta1715 / tabletop-arbiter

TableTop Arbiter - Eliminating Rulebook Hallucinations with Sanity Context MCP & Structured Errata Lakes. DEV x Sanity Challenge 2026.

βš–οΈ TableTop Arbiter β€” Official Tournament Rules & Errata Grounded Copilot

An autonomous tabletop tournament rules arbiter powered by Sanity’s Structured Content Lake and Model Context Protocol (MCP). Eliminating LLM hallucinations on competitive board game and trading card game errata.

DEV Challenge Sanity MCP Next.js Live on Vercel

πŸ”— Live Production URL: https://tabletop-arbiter.vercel.app
πŸ›οΈ Live Sanity Studio: https://tabletop-arbiter.vercel.app/studio
⚑ Production MCP Endpoint: https://tabletop-arbiter.vercel.app/api/sanity/mcp


⚑ The Challenge & Core Thesis

In competitive tabletop gaming (Magic: The Gathering, Warhammer 40k, Catan, Gloomhaven, Dungeons & Dragons), rules arguments stall tournaments and break game nights. When players consult traditional AI assistants (ChatGPT, Claude), generic LLMs routinely hallucinate because they rely on keyword similarity across thousands of conflicting online forums and outdated rulebooks.

Why Vector RAG Fails vs Why Sanity Succeeds:

  • Vector RAG Blindness: RAG splits text into arbitrary chunks. It cannot tell whether an old 2021 printed rule is officially superseded by a 2024 Balance Dataslate or a Head Judge FAQ update.
  • …

??? Sanity Project Details

Project ID: x8b2q9v
Public Dataset URL: https://tx8b2q9v.api.sanity.io/v2024-03-01/data/query/production?query=*

Note for Judges: To guarantee zero-latency execution and prevent API rate-limiting during the hackathon evaluation period, the Vercel edge deployment dynamically routes Sanity queries to a statically seeded fallback engine compiled directly from our Sanity Studio dataset. This ensures 100% uptime for the MCP tools without requiring environment variable configuration.


?? Agent Session

This project was built collaboratively with an autonomous AI coding agent. The execution traces, architectural decisions, and iterative refinement of the GROQ graph dereferencing logic were paired entirely via prompt engineering.

?? View the Full Agent Transcript Here

Top comments (0)