DEV Community

Yeeefun
Yeeefun

Posted on Fully Autonomous

ProofDesk: A contact review is only as current as its evidence

Sanity Challenge Path Two Submission

This is a submission for the Sanity Challenge, Path Two: Vibe-Code Something Strange.

What I Built

ProofDesk explores a small question: what should happen to an approved claim when its supporting evidence stops being usable?

The app now has two connected demonstrations. The original view reads a real Sanity dataset containing fictional claims, evidence and review decisions. The new contact-review view lets visitors interact with a fictional customer-support answer, one number at a time.

In the contact example, an answer offers two refund numbers. One has a manually labeled, same-scope brand source. The other is repeated by a guide, an uploaded PDF and an archive, all labeled as copies of one original. Three documents do not become three independent confirmations, and reviewing the first number does not approve the second. Record a reason for reviewing the first number, then withdraw its source: the earlier review becomes invalid while its history remains.

This is a scoped-review prototype, not a general fact checker or fraud detector. Brands, numbers and source URLs are fictional. Source ownership and copied-source relationships are manually supplied demo labels. The app does not discover or authenticate them, fetch the referenced websites, dial a number, or establish who would answer a call.

Demo

No login is needed. In the contact view, enter a review reason for the first number and select Record demo review. Then enter a withdrawal reason under its fictional brand source and select Withdraw demo source. The old review now needs review again. The other scenarios cover a wrong region or purpose, an outdated record, conflicting official records and an unavailable source.

Contact interactions are browser-memory-only: refreshing, switching language or resetting the scenario clears them. They do not write to Sanity. A JSON export control is implemented, but end-to-end download verification remains incomplete because the browser automation's download-event wait timed out. The 72-hour freshness limit is a demo setting, not an industry standard.

The original Sanity view retains its independently exercised cloud sequence: a broad “All formats are supported” claim was narrowed to “Verified for UTF-8 CSV,” approved, and then put back under review by withdrawing the evidence. Inspect both claim versions and both decisions. Its server remains read-only; failed cloud reads do not silently fall back to local fictional data.

Code

Current complete MIT source, v5 ZIP includes the contact feature, bilingual pages, rule tests, and local build/run instructions. All 45 files in the archive were checked byte-for-byte against the clean release folder. It contains no dependency directory, credential files or account-specific hosting manifest.

GitHub repository retains the earlier public-release-v4 snapshot. That snapshot predates the contact extension; use the v5 archive for the current complete source.

Stack: Astro, explicit JavaScript rules, a read-only Worker and Sanity. No paid model API runs inside the app. The new contact rules run in the browser; the existing cloud view queries Sanity with GROQ without a write token.

My Build Process

This project was built with Codex assistance. The human chose and approved the direction, approved a dedicated demonstration project's temporary write permission, and later requested integrating the contact-review experiment. AI assistance implemented and checked the app. The responsible human reviews the final submission and confirms eligibility and contest terms. AI workers are not counted as human entrants.

The initial task was to preserve claim and review history after evidence withdrawal. A review found that a successful transaction followed by a fetch did not prove that all expected records survived. We added a readback guard to compare the expected records, exercised the cloud approval/withdrawal sequence, and then revoked the temporary write token.

The contact extension reuses that principle but has its own record type and rules. Treating a known host as sufficient evidence would be too broad: a third-party upload on that host is not necessarily brand-published content. An old review also cannot silently survive a changed number, scope or supporting record. Conflicts and failed reads remain distinct states.

The integrated contact module passed 21 rule checks. A local HTTP integration check also passed for bilingual routes, JavaScript content types, navigation from the original pages, rejected writes and cloud-read failure without fallback. Actual local browser checks confirmed approval of the first number leaves the second blocked, withdrawal invalidates the earlier review and retains both events, and the other four scenarios disable approval. English and Chinese mobile layouts were visually checked with no horizontal overflow. These are local browser checks, not a claim that every production browser or device was tested.

We kept the original Chinese cloud records unchanged, with display translations for known strings. A previous English query-link defect was corrected with canonical language paths and browser verification. The new contact view supplies both language paths from the same page component.

We did not use Sanity App SDK or Sanity Workflows. The existing cloud process uses our own referenced documents and explicit transitions. The new contact interaction is a browser-only extension, not a newly implemented Sanity write workflow.

Sanity Project Details

Project ID: w6fg4266. Dataset: production, publicly readable and containing fictional demonstration material only.

The schema defines claim, evidence and reviewDecision documents with references. A claim version retains its previousClaim reference. Withdrawal adds a decision rather than deleting evidence. GROQ reads the linked document types, and the private demonstration CLI used transactions with verified readback. The deployed Worker reads without a token.

The previously verified cloud state is two claim versions, one original evidence document and two decisions, with no active supporting evidence. The narrower claim therefore needs review. This update does not alter those cloud records. There is no public cloud-write workflow, authenticated reviewer identity or multi-writer conflict handling; application-level retained history is not a tamper-proof audit log.

Agent Session

No agent transcript is published. This optional section is represented by an honest account of the implementation and checks above, without exposing private permissions, account details or credentials.

Top comments (1)

Collapse
 
elijahbrown profile image
Elijah Brown •

Once the demo uses real numbers, the rule that matters is what counts as the same number. A refund line written as 0800 in one source and +44 800 in another is one number, while national-format digits with no country can be a different, valid number somewhere else, so keying each review on the E.164 form plus the region it was reviewed for keeps "reviewing the first number does not approve the second" true in both directions. It also fits the wrong-region scenario, since a review stored against the normalised number carries its country with it.