DEV Community

Mike
Mike Subscriber

Posted on

✍️ MagenticCMS: Rehearse the Comments Before You Publish the Post

Sanity Challenge Path Two Submission

This is a submission for the Sanity Challenge, Path Two: Vibe-Code Something Strange.

What I Built

A post can sound perfectly reasonable to its author and land very differently with the people reading it.

MagenticCMS is a pre-publication rehearsal room for content editors. Write a draft, pick an audience, and watch simulated readers react to it before you decide what to change. The goal is not to predict virality. It is to surface the questions worth asking while the post is still editable.

What's inside:

  • A web composer: title, copy, an optional description of the image/video, platform, brand and test audience.
  • A live conversation feed: simulated readers scroll past, react, comment, and in later waves reply to each other. Reactions stream in with floating emoji and an optional beat, livestream-style.
  • A 3D audience table (Three.js): one figure per stored reaction, grouped by expressed sentiment, with a separate lane for people who scrolled past. Click a figure to read its comment and follow its real reply links.
  • An analyst: reads the reactions, the worst-reacting cohorts and the brand's own memory (red lines, past incidents stored in Sanity), quotes the exact phrase causing trouble, and drafts a rewrite.
  • An editorial workflow stored as data: agents may flag or clear a draft; only a human may approve or publish.

The audience is 379 persona records projected from a sample of the MatrAIx Persona-1M dataset (human-grounded and synthetic profiles). They are not real participants, not verified identities, and not a representative sample of any country.

The name nods to Microsoft's Magentic Marketplace, a multi-agent market simulation. MagenticCMS applies the idea to content, with the CMS behind it.

Demo

🔗 Live:

✍️ MagenticCMS ✍️

Sanity-powered content rehearsal: test drafts with AI personas, explore reactions and replies in 3D, and review suggested edits before publishing.
Live Demo 🔥 - https://magentic-cms.vercel.app

🎬 Video:

💻 Code: https://github.com/longphanquangminh/magentic-cms

There are only two controls: ▶ Run simulation (a new simulated audience; uses the LLM) and ↻ Replay (plays back a saved session; no LLM calls). Hype mode and sound are on by default and can be turned off with one click.

A quick tour that costs nothing:

  1. Open Bolt, Ramadan late-night promo and press Replay. You're watching stored results, not a new prediction.
  2. Drag the 3D table, click a figure, follow a reply link.
  3. Open Insights for the flagged phrases, the matching brand memory and the suggested rewrite.
  4. Open Workflow for the editorial history, with 🤖 and 🧑 entries side by side.

To test your own draft, use + Write a post, save it, then press Run simulation. Saving never calls the model. The demo is shared, so please don't enter private content. "Publish" here changes a CMS workflow state; nothing is sent to a social network.

What it caught

The seeded brands and their incident histories are fictional test fixtures. One deliberately bad Saveo post read:

"Still broke at 30? 😂 Maybe skip the daily latte, bestie… No excuses."

Saveo's brand record contains a fictional 2025 "avocado-toast ad" incident and a red line against shaming people for small pleasures. The analyst connected the new wording to that stored history instead of producing a generic sentiment summary. In two 40-persona development runs the backlash index was 89/100 for the original and 42/100 after applying the suggested rewrite. Those are heuristic simulation indices: not percentages, not validated predictions, and not a controlled A/B test (a new audience was sampled for the second run).

The Bolt Burgers fixture ("Midnight Bacon Feast… the perfect way to break your fast after iftar") scored 91–94 and was flagged against the red line "no pork promotions in halal-certified markets"; the worst cohorts were MENA, South Asia and Southeast Asia. The safe Saveo post ("No judgment, just receipts") scored 18 and was cleared to review. When I then asked the agent to approve it, the workflow document refused:

{"error": "\"Approve\" can only be taken by: human (not agent)"}
Enter fullscreen mode Exit fullscreen mode

Code

Repository: https://github.com/longphanquangminh/magentic-cms

magenticcms/
├─ web/                         Next.js 16 app: live room, composer, simulation engine
│  ├─ app/
│  │  ├─ page.tsx               Home: posts, latest scores, dataset counts
│  │  ├─ compose/page.tsx       Write / edit a draft
│  │  ├─ posts/[id]/page.tsx    The live room
│  │  └─ api/
│  │     ├─ composer/           Create/update posts (optimistic _rev check)
│  │     ├─ posts/[id]/         Post state · workflow transitions
│  │     └─ runs/               Create run · step · finalize · cancel · reactions · SSE stream
│  ├─ components/
│  │  ├─ LiveRoom.tsx           Orchestrates a run from the browser, reveals reactions
│  │  ├─ AudienceWorld.tsx      Three.js audience table + 2D fallback
│  │  ├─ PostComposer.tsx       The editor
│  │  └─ useLiveAudio.ts        Web Audio: synthesized cues + 128 BPM beat
│  ├─ lib/
│  │  ├─ engine.ts              Sampling, crowd batches, analyst, transitions
│  │  ├─ workflowEngine.ts      Enforces the workflow document
│  │  ├─ metrics.ts             Pure scoring, shared by browser and server
│  │  ├─ modelTransport.ts      Bounded retries / failover, no raw provider errors
│  │  ├─ crowdValidation.ts     Rejects incomplete or duplicated model output
│  │  ├─ heal.ts                Closes runs abandoned by a closed tab
│  │  └─ sanity.ts · queries.ts · gemini.ts
│  └─ tests/resilience.cjs      Offline tests for the failure paths
├─ studio/                      Sanity Studio v6
│  ├─ schemaTypes/              post · brand · audience · persona · simulationRun · reaction · workflow · workflowEvent
│  ├─ actions/workflowActions.tsx   Document actions generated from the workflow document
│  ├─ components/SimulationPane.tsx Live "Simulation" tab (client.listen)
│  ├─ components/StageBadge.tsx
│  └─ structure.ts              Posts grouped by stage, runs, personas, workflow log
├─ scripts/
│  ├─ build_personas.py         MatrAIx parquet → persona documents (NDJSON)
│  └─ seed.py                   Personas, workflow, brands, audiences, demo posts → Content Lake
└─ DATA_ATTRIBUTION.md          Dataset terms and responsible-use notes
Enter fullscreen mode Exit fullscreen mode

My Build Process

MagenticCMS is my project: the idea, the product decisions, the testing, and the long list of things I sent back because they weren't right. I built it by directing an AI coding agent, running every version locally, and deciding what shipped. The prompts quoted below are paraphrases of my real requests, not a verbatim transcript.

Where it started

Before this challenge I had prototyped a "simulate a population before you ship" idea for another hackathon. Its weakness was obvious once I used it: it only told you things. For this build my rule was that the system must move content through a real editorial process: a person approves, the agent can only flag or clear. That's why the workflow lives in Sanity next to the content instead of in code.

I chose MatrAIx persona records as the audience source because I wanted reactions conditioned on structured attributes (income, politics, religiosity, attitudes to brands and influencers) rather than a generic "the internet thinks…" paragraph.

Data first: Python → Content Lake

scripts/build_personas.py reads the MatrAIx sample shard (999 rows × 994 columns) with pandas/pyarrow, keeps adults only, drops the Wikipedia-derived profiles of public figures, ranks records by how many of the attitude/personality fields are actually populated, and projects ~40 fields (Big Five/BFI-2 traits, attitudes to social media, influencers, brands, consumerism; income, region, generation, religiosity, shopping style, tone…) into a persona document. Each document keeps matraix{source, recordId, rowIndex} so every simulated comment can be traced to its source record. Handles are deterministic aliases generated from a seeded hash, not names from the data.

scripts/seed.py writes everything through the Sanity HTTP mutation API: 379 personas in batches of 100 (createOrReplace), the workflow document, two fictional brands with red lines and past incidents, four audiences, and three demo posts (createIfNotExists, with a --reset-posts flag to put them back to draft).

The schema is the product

post ──► brand            (voice, redLines[], pastIncidents[{title, date, whatHappened, lesson}])
  │  └► audience          (title, groqFilter over personas, sampleSize)
  │
  ├── simulationRun       (revision, bodySnapshot, mediaSnapshot, queue[{persona→, wave}],
  │     │                  status, processed, metrics{…}, cohorts[], analysis{mode, verdict,
  │     │                  headline, summary, riskFlags[{phrase, why, severity, brandMemory,
  │     │                  evidence[→reaction]}], suggestedRevision, changes[]})
  │     └── reaction ×N   (persona→, wave, seq, action, emoji, text, sentiment, likes,
  │                        replyTo→reaction, flag{category, reason}, thought)
  │
  └── workflowEvent ×N    (transition, from, to, actor{kind: human|agent, name}, note, run→)

persona                   (379 docs, read-only; traits{}, attitudes{}, matraix{source, recordId})
workflow.socialPost       (stages[], transitions[{id, from[], to, actors[], requiresNote}])
Enter fullscreen mode Exit fullscreen mode

Decisions behind it:

  • An audience is a GROQ filter, not a list of IDs. "Budget-conscious" is literally householdIncome in ["<$25k","$25k-50k"] || shoppingStyle == "Bargain hunter". The engine samples personas by running that filter against the Content Lake, so an editor can define a new audience in Studio without touching code.
  • replyTo is a reference from one reaction to another. That reference is what makes the agents talk to each other rather than beside each other, and it is exactly what the 3D view draws as reply links.
  • Every run snapshots the copy and the media description. Revisions stay comparable after the post changes.
  • Each risk flag carries evidence[] references to the reactions that justify it. The phrase is highlighted in the post; the comments behind it are highlighted in the feed.
  • Every persona points back to its MatrAIx record. Click "why?" on a comment to see the persona card, its private "thought", and the source.

Workflow as data, enforced in one place

workflow.socialPost is a normal document you can open in Studio:

start_simulation   draft | needs_revision | ready_for_review → simulating   [human, agent]
flag_risk          simulating → needs_revision                              [agent]
clear_for_review   simulating → ready_for_review                            [agent]
apply_revision     needs_revision | ready_for_review → draft                [human, agent]
approve            ready_for_review → approved                              [human]
override_approve   needs_revision → approved                                [human, note required]
publish            approved → published                                     [human]
send_back          ready_for_review | approved | needs_revision → draft     [human]
Enter fullscreen mode Exit fullscreen mode

The analyst, the live-room buttons and the Studio document actions all call the same transition() function in web/lib/workflowEngine.ts. It reads the document, checks the current stage and the actor's kind, patches the post (and its draft, if one exists) and appends a workflowEvent, all in one transaction.

Customising the Studio

  • Document actions are generated from the workflow document: Run simulation, Apply suggested revision, Approve, Approve anyway (asks for a reason), Publish, Send back. Each one hides itself unless the current stage and a human actor allow it.
  • A "Simulation" view on every post: a custom pane that subscribes with client.listen() to the post's runs, so the risk/trend meters move while a run is in progress, even when it was started from the web app.
  • Stage badges and a structure tree that groups posts by stage with GROQ-filtered lists; separate lists for runs, reactions, personas and the workflow log; Vision for ad-hoc queries.
  • Personas, runs, reactions and the log are read-only and have no "create new" templates. Editors shouldn't hand-tune the crowd until it says what they want.

The engine

  1. Sample. Run the audience's GROQ filter, shuffle, cap at 120, split into three waves (50 / 30 / 20 %). The queue is stored on the simulationRun.
  2. Step. Each call to /api/runs/:id/step takes up to 16 personas from the same wave and plays them with Gemini (gemini-3.5-flash-lite, 8 personas per call, 2 calls in parallel, JSON mode with a response schema). Personas in waves 2 and 3 are shown the top and newest comments with IDs and may reply to or like them. Every model response is checked against the exact persona IDs requested; incomplete or duplicated batches are rejected. Accepted reactions, like increments and the queue update go into one transaction with an ifRevisionId lock on the run, so two tabs can't double-process.
  3. Stream. /api/runs/:id/stream subscribes to Sanity's listener on the server (the token never reaches the browser) and forwards "changed" pings; the browser pulls the delta by seq and reveals reactions one by one.
  4. Score. metrics.ts is a pure function used by both browser and server, so the live gauge equals the saved number. Backlash = 0.35·negative share + 0.25·flag rate + 0.15·angry share + 0.25·negative amplification (share of comment likes that went to negative comments). Trend = engagement, share rate, reply depth and intensity. Cohorts are grouped by generation, region, political lean, income and religion (min n = 3).
  5. Analyse. gemini-3.5-flash gets the metrics, worst/best cohorts, the 14 most negative and 6 most positive comments with IDs, and the brand document. It returns a verdict, flags (verbatim phrase + brand memory + evidence IDs) and a rewrite.
  6. Move. If the analyst is unavailable, risk ≥ 45, any flag is high severity or the verdict is kill → flag_risk; otherwise clear_for_review.

For 40 personas that is normally 6 crowd calls + 1 analyst call. This is batched persona role-play, not one autonomous process per person, and the reveal animation is not token streaming.

Failure handling: "never invent evidence"

Asking "what happens when the API fails?" exposed a real bug: an early version counted a missing model response as a scroll, so an outage looked like audience behaviour. Now:

  • modelTransport.ts has bounded timeouts and retries, an optional explicit fallback model, and never surfaces raw provider error bodies.
  • Failed or incomplete batches stay pending; the run pauses with saved progress and offers Retry pending agents, Stop & return to draft, or Replay last completed session.
  • If only the analyst fails, the UI shows METRICS ONLY: deterministic scores, no invented rewrite or approval, human review required.
  • A run abandoned by a closed tab is closed as interrupted after 3 minutes (heal.ts); its reactions are kept and the post returns to draft. A fresh run resumes automatically on reload.
  • tests/resilience.cjs covers missing key, auth and quota errors, failover, invalid JSON, omitted/duplicated personas, partial-batch pause/resume and the metrics-only analyst path, all against an in-memory store with no real API calls.

The 3D room

AudienceWorld.tsx renders a tabletop with an orthographic camera and OrbitControls. One miniature (head + torso + base) per stored reaction, coral/stone/lime by sentiment, a separate lane for scrolls. Rendering is on demand (no idle loop), hover and click use raycasting, the selected figure's replyTo parent and children are drawn as curved tubes, and everything is disposed on unmount. A keyboard-accessible participant menu and a 2D cohort fallback keep the same data when WebGL is unavailable.

What I pushed back on

  • "It looks busy but means nothing." The first room was a dense dashboard; the next put decorative dots behind a floating post that overlapped the UI on long copy. I asked for 3D only if it carried data.
  • "Why can't I write my own post?" Early builds could only test seeded examples. The composer came from that, including an editable media description, because removing "bacon" from the copy doesn't help if the image is still described as a bacon burger.
  • "Three buttons is confusing." I cut the controls to Run simulation and Replay, and moved the rewrite loop into Insights.
  • "It should feel like a livestream." Hype mode speeds up the reveal of stored reactions, emoji float up per real reaction, and a synthesized beat (Web Audio, no files) plays from the moment a run starts until it finishes. None of it adds model calls or fake participants.

Bugs I found while testing

  • Emoji vanished on my machine: Windows had animation effects off and the CSS hid them under prefers-reduced-motion.
  • Posts stuck on "simulating" after F5: fixed with resume-on-reload and the 3-minute interrupted-run cleanup.
  • The beat stopped the moment a simulation started: resetting the feed silenced it and nothing restarted it; now one rule decides when audio plays and I tested start, analysis, replay, pause, sound-off and error paths.
  • Sanity specifics: listener filters can't follow references (run->post._ref), so the Studio pane listens to simulationRun documents; @sanity/icons v5 needs per-icon imports.
  • Smaller things: an off-centre logo glyph, an unreadable dropdown placeholder, a missing favicon, demo content that wasn't in English.

What I didn't do

  • I modelled the workflow myself as documents rather than using the new Workflows beta, which shipped two weeks before the deadline. Moving workflow.socialPost onto it is the obvious next step.
  • No App SDK app; the live room uses @sanity/client. No Context MCP or Knowledge Base; this is a Path Two entry.
  • The agents can't see images; they read the media description.
  • The browser drives processing steps; durable background execution is next.

Limits I'm aware of

  • Scores are heuristic indices, not probabilities, and are not validated against real feedback.
  • Persona attributes and LLM outputs can be biased. "Human-grounded" is not "verified". Cohorts overlap, and n = 3 says nothing about real groups.
  • The demo is shared and unauthenticated. Human/agent labels demonstrate workflow logic, not a production authorization boundary. Real authentication and quotas are required before any production use.

Sanity Project Details

  • Project ID: bh5o6glp
  • Dataset: production (public, read-only)
  • Document types: post, brand, audience, persona, simulationRun, reaction, workflow, workflowEvent
  • Query it yourself: https://bh5o6glp.api.sanity.io/v2025-02-19/data/query/production?query=*[_id=="workflow.socialPost"][0].transitions

The project ID is shared for judging the schema; it is not a credential. No write tokens or LLM keys are published.

Data attribution: MatrAIx Persona-1M · MatrAIx paper. The dataset is released for non-commercial research use only, including subsets and derivatives; upstream source terms apply. This is a research prototype, not a claim of commercial data rights.

Agent Session

No transcript attached. The build process above is the honest account; AI assistance is disclosed, and nothing in this post is invented for the submission.

Top comments (0)