DEV Community

Cover image for I Built an Evidence Pipeline for MCP Failures with Sanity
L Anil Kumar Singha
L Anil Kumar Singha

Posted on AI-assisted

I Built an Evidence Pipeline for MCP Failures with Sanity

Sanity Challenge Path Two Submission

This is a submission for the Sanity Challenge, Path Two: Vibe-Code Something Strange

What I Built

I built MCP Failure Observatory, a public, evidence-backed view of how Model Context Protocol implementations behave under recorded failure and recovery scenarios.

An MCP implementation can work correctly during a normal request but behave differently when a connection disappears, a response arrives after cancellation, or a child process exits.

Results from those tests are often distributed across terminal output, Markdown reports, GitHub issues, and implementation-specific test suites.

The Observatory turns them into structured compatibility data that can be searched, compared, reviewed, and traced back to its source evidence.

It is intended for:

  • MCP implementation maintainers;

  • developers selecting or integrating MCP software;

  • contributors investigating failure and recovery behavior; and

  • anyone who needs to verify a compatibility claim instead of trusting a summary badge.

The project does not rank implementations or certify their quality. It records what was tested, what happened, and the evidence needed to interpret the result.

What visitors can explore

The public application includes:

  • a live compatibility overview;

  • a scenario-by-implementation matrix;

  • filterable Test Runs;

  • Test Run details with observations and Evidence;

  • implementation version and compatibility history;

  • side-by-side comparisons;

  • shareable filter and comparison URLs;

  • human-reviewed Findings;

  • upstream issue and evidence-comment links; and

  • source revision and synchronization freshness.

MCP Failure Observatory overview dashboard showing implementation and Test Run totals, verified Findings, the compatibility matrix, recent Test Runs, and data synchronization status

Test Run:

https://observatory.mcplab.dev/test-runs/Fp6j2OFYXErtrNMrGKY176

Official Everything Server Test Run showing a failed Cleanup After Timeout scenario over stdio, the recorded observation, supporting evidence, and verified Finding

The reviewed Finding states:

A cancelled long-running operation can delay stdio client cleanup beyond the configured bound.

After verification, the Finding was reported to the official Model Context Protocol servers repository:

https://github.com/modelcontextprotocol/servers/issues/4846

This creates a traceable path from a reproducible failure test to structured Evidence, human review, and upstream action.

Demo

Live application

https://observatory.mcplab.dev

Methodology and limitations

https://observatory.mcplab.dev/methodology

Video walkthrough

The walkthrough demonstrates:

  1. the live compatibility matrix;

  2. the Findings page;

  3. the Official Everything Server Test Run;

  4. its supporting Evidence;

  5. the customized Sanity Review Desk;

  6. the verified Finding;

  7. the upstream GitHub issue; and

  8. the automated synchronization workflow.

Sanity Review Desk

The embedded Studio is restricted to authorized editors, so judges will see its workflow in the video and screenshots rather than receiving editing access.

Customized Sanity Review Desk showing the Verified Findings queue and the published Official Everything Server cleanup Finding linked to its failed Test Run

Sanity upstream reporting workflow showing the verified Official Everything Server Finding linked to the modelcontextprotocol servers repository and GitHub issue 4846

The Review Desk organizes Findings into:

  • Needs Review
  • Verified
  • Rejected
  • Not Reported
  • Reported by MCP Failure Lab
  • Evidence Added to Existing Issue
  • Resolved Upstream
  • Not Actionable

Code

MCP Failure Observatory

https://github.com/anilloutombam/mcp-observatory

This repository contains:

  • the Next.js public application;
  • Sanity schemas and Studio structure;
  • GROQ queries and data access;
  • the idempotent importer;
  • GitHub Actions synchronization;
  • accessibility and responsive behavior;
  • Sentry monitoring; and
  • automated tests.

MCP Failure Lab

https://github.com/anilloutombam/mcp-failure-lab

This repository executes the compatibility scenarios, retains the source reports, produces the normalized Observatory manifest, and triggers synchronization.

My Build Process

I built the project iteratively with Codex.

The useful part was not prompting an entire dashboard into existence in one step. It was turning product questions into explicit data, interface, and workflow decisions.

The prompts that worked

The most effective prompts were narrow and outcome-focused.

Examples included:

Build the first real MCP Failure Observatory dashboard using the existing Sanity schemas and imported dataset. Do not hardcode the mock values from the reference image.

Make filters and pagination shareable through the URL, and ensure clearing a filter also removes it from the URL.

Model Findings separately from Test Runs so imported observations do not automatically become verified conclusions.

Preserve human-reviewed Finding fields when the automated importer synchronizes newer source data.

Build a Sanity Review Desk for Needs Review, Verified, Rejected, and upstream-reporting states.

Review the mobile layouts at the breakpoints where the sidebar, cards, comparison selector, and tables change structure.

These prompts worked because they defined observable behavior and constraints instead of asking for vague improvements.

Where the model got stuck

The first versions had several problems.

Some controls were present even though they did not perform an action.

Responsive layouts were initially desktop components squeezed onto a mobile screen.

Different pages implemented similar headers and cards independently, which created inconsistent spacing.

Some implementation repository URLs were missing.

Status presentation depended too heavily on color.

Comparison filters did not initially update the URL correctly.

The early Finding model also needed a clearer distinction between an issue opened by MCP Failure Lab and evidence added to an issue reported by someone else.

How I course-corrected

I reviewed the application page by page and replaced one-off presentation code with shared components where the behavior was genuinely the same.

Nonfunctional controls were hidden rather than presented as unfinished actions.

Mobile tables became structured record cards where horizontal compression would have made them unreadable.

Status badges gained text and symbols so meaning did not depend only on color.

Repository identities were backfilled from authoritative implementation sources.

The importer was changed to validate the complete manifest before its first write and to preserve existing human review decisions.

I also added a public methodology page because compatibility data needs visible limitations.

Reaching deeper into Sanity

I focused on modeling a real review workflow rather than adding an App SDK page only to claim a bonus.

Sanity stores six connected document types:

Document Responsibility
Implementation Repository identity and implementation type
Scenario The behavior being tested
Test Run One version, scenario, transport, date, and outcome
Evidence An observation and its source material
Finding A human-reviewed conclusion and upstream state
Data Sync Source revision, freshness, status, and imported counts

The Finding document contains two related workflows.

Editorial review:

Needs Review
├── Verified
└── Rejected
Enter fullscreen mode Exit fullscreen mode

Upstream action:

Not Reported
├── Reported by MCP Failure Lab
├── Evidence Added to Existing Issue
├── Resolved Upstream
└── Not Actionable
Enter fullscreen mode Exit fullscreen mode

An automated import can create a Finding in Needs Review, but it cannot decide that its own output is verified.

A person must examine the related Test Run and Evidence.

The schema enforces several boundaries:

  • a verified Finding requires supporting Evidence;
  • an upstream report requires a repository, issue URL, and issue number;
  • evidence added to an existing issue requires a direct comment URL;
  • reported states require a reporting date; and
  • rejected Findings that still point to upstream reports produce a warning.

Conditional fields show upstream and resolution inputs only when the selected state needs them.

The customized Studio structure turns those states into working review queues instead of presenting editors with a generic list of schema types.

Automated data flow

MCP implementation
        ↓
Failure Lab scenario runner
        ↓
JSON test results
        ↓
Public report + versioned manifest
        ↓
Failure Lab publish workflow
        ↓
GitHub repository dispatch
        ↓
Observatory synchronization workflow
        ↓
Complete manifest validation
        ↓
Idempotent importer
        ↓
Sanity Content Lake
        ├── Human Finding review
        └── GROQ queries
                  ↓
          Public Observatory
                  ↓
        Upstream issue or evidence comment
Enter fullscreen mode Exit fullscreen mode

The importer uses stable external source identities, while Sanity controls its document IDs and relationships.

Repeated synchronization updates existing source records instead of creating duplicates.

Invalid manifests are rejected before partial content is written.

Most importantly, later imports preserve human review status, resolution notes, and upstream-reporting decisions.

What AI accelerated—and what it did not decide

Codex accelerated implementation, debugging, schema work, responsive refinement, accessibility improvements, query development, and test coverage.

It did not decide what the data was allowed to mean.

I still had to define the boundary between a Test Run and a Finding, verify upstream links, distinguish missing data from failure, reject misleading interface shortcuts, and ensure automation could not overwrite editorial judgment.

That iteration became the most important part of the build.

Sanity Project Details

Sanity project ID: f5cr7ngj

Dataset: production

The public application queries this dataset through Sanity’s API.

The schemas and Studio configuration are available in the Observatory repository:

  • src/sanity/schemaTypes
  • src/sanity/structure.ts
  • src/sanity/lib/queries.ts

The main relationships are:

Test Run → Implementation
Test Run → Scenario
Evidence → Test Run
Finding → Test Run
Finding → Supporting Evidence
Data Sync → Source revision and import status
Enter fullscreen mode Exit fullscreen mode

Stable external identities are stored in sourceKey fields rather than being encoded into Sanity document IDs.

The embedded Studio is deployed with the application at /studio, but editing access is restricted to authorized users.

Top comments (0)