DEV Community

Cover image for My voice app's one job was to take back a wrong sentence. It was taking back sentences it never said.
Edy Cu
Edy Cu

Posted on

My voice app's one job was to take back a wrong sentence. It was taking back sentences it never said.

Ray is 68 and five days home after a hip replacement. He asks his assistant whether he can put weight on the new hip. Halfway through the answer, his physio changes the plan. Unsay exists for that moment: the assistant stops mid-sentence and says "Wait — don't do that. What I just told you is out of date."

Its one job is to take back a wrong sentence. When I had the code audited, it turned out it was also taking back sentences it had never said.

This post covers the four ways the retraction was wrong, and what fixing each one taught me about MCP resource subscriptions. The repo is public, with every fix below in its own commit. Try it live at api.unsay.edycu.dev/judge, and see the hash-chain proof at /verify.

The mechanism, in one paragraph

Unsay is an MCP server (spec 2025-11-25, Streamable HTTP). The care plan is not a PDF: each instruction is a versioned MCP resource such as care://ray/weight_bearing. A voice host subscribes with resources/subscribe. When a clinician publishes a signed revision, the server sends notifications/resources/updated. The host stops, re-reads the resource, and finds a server-rendered sentence in _meta["unsay/retraction"]:

const parts = [
  // 1 — the stop, before anything else.
  `Wait — don’t do that. What I just told you is out of date.`,
  // 2 — the sentence being withdrawn, so there is one instruction to drop.
  `I said ${quote(prev.value)}`,
  // …and the one that replaces it, with who changed it and when.
  `${next.authorLabel} changed it ${age}: ${quote(next.value)}`,
]
Enter fullscreen mode Exit fullscreen mode

The order is the safety property. A frightened person acts on the first clause, so the stop comes before any metadata.

For hosts that can't subscribe, there's a fallback tool, whats_changed. Everything below is about what prev should be.

Wrong #1: "the previous version" is not "what you said"

Here is the original line:

const previous = record.version > 1
  ? chain.find((r) => r.version === record.version - 1)
  : undefined
Enter fullscreen mode Exit fullscreen mode

Version n − 1 is what the chain said before. It is not what this listener heard. Those only coincide in a demo.

The seeded record starts at v2, so the very first read of care://ray/weight_bearing already carried:

Wait — don't do that. What I just told you is out of date. I said "No weight through the operated leg. Transfers with the frame only." …

Nothing had been said yet. And the proof file I had committed as evidence (docs/proof/live_run.jsonl) recorded exactly that sentence on a first read. The receipt documented the bug.

The resume case was worse. The host reads v2, the connection drops, v3 and v4 land, and the host re-reads. The retraction then quoted v3, a sentence the listener never heard.

The fix is a map per session from resource to the version this session was last served:

const key = uriFor(parsed!.subject, parsed!.topic, record.audience)
const lastHeard = heard.get(key)
const previous =
  !superseded && lastHeard !== undefined && lastHeard < record.version
    ? chain.find((r) => r.version === lastHeard)
    : undefined
if (!superseded) heard.set(key, record.version)
Enter fullscreen mode Exit fullscreen mode

There is no retraction on a first read, and none for an old version someone asked for by URI. That URI now says [SUPERSEDED — v1, replaced by v2; …] in its text, because annotations don't survive a read in every host.

Wrong #2: "since when you last spoke" is the wrong cursor

The fallback tool took a since timestamp, and I told hosts to pass "when you last spoke about the plan". That sounds right, and it's wrong twice over:

  1. A correction that lands while the host is still talking is written before the host finishes. So writtenAt > endOfAnswer filters out exactly the event the product exists for.
  2. since came from the host's clock and was compared against the server's writtenAt. A host clock running a minute fast returned changed: [], with no error.

The fix is the same one every sync API eventually makes: the server hands out the cursor.

// Server clock at the call, i.e. before the host speaks. Handed back as the next
// `since`, it catches a correction written while the host was still talking.
const payload = { changed, asOf: at.toISOString() }
Enter fullscreen mode Exit fullscreen mode

A since ahead of the server clock is now an error, not a silent "nothing changed". The reviewer then found one more hole: a write stamped in the same millisecond as asOf was never returned, because the filter was >. It's >= now. A record returned twice carries no second retraction, which brings us to #3.

Wrong #3: a re-confirmation is not a correction

The anticoagulant record ships stale on purpose. When the GP confirms it with a new review date, the words don't change, only staleAfter does. That is a real revision, so a notification fires, and the host dutifully stopped mid-word to say:

I said "Rivaroxaban 10mg once daily. Stop date: 25 September 2026." Dr Mensah, GP changed it just now: "Rivaroxaban 10mg once daily. Stop date: 25 September 2026."

The fallback path already had a same-words guard. The subscribed path didn't. Both now share one rule:

...(previous && previous.value !== record.value
Enter fullscreen mode Exit fullscreen mode

The Agent Skill tells the host what a missing retraction means: the words you were saying did not change: carry on.

Wrong #4: the gloss turned "do NOT" into "you can"

Clinical phrases get a plain-English gloss. "Full weight-bearing as tolerated" becomes "you can put as much weight through that leg as is comfortable". The glossary matched on text alone, so this instruction:

Do NOT progress to full weight-bearing as tolerated until the X-ray is reviewed.

was helpfully explained as permission.

My first fix was a list of negation words. The review came back with four sentences the list couldn't see: "Stop full weight-bearing…", "Wait for the X-ray before…", "…is cancelled", "Hold off…". No word list is ever complete, but position is checkable:

// A weight-bearing gloss is worded as permission, and no word list catches every
// negation ("Hold off…", "…is cancelled"). Position does: gloss only a value that
// OPENS with the term, with no action word over it. Definitions are unaffected.
return hits
  .filter((g) => !WEIGHT_BEARING.has(g.label) || (opens(g) && !ACTION.test(text)))
  .map(({ label, gloss }) => ({ label, gloss }))
Enter fullscreen mode Exit fullscreen mode

When in doubt, the sentence is quoted verbatim with no gloss. A missing explanation is an inconvenience. A wrong one is a fall.

How these were found

I gave another model read access to the repository and asked it for an architecture audit. The rule for every finding it raised: verify it against the code before touching anything, reject it with a file:line if it's wrong, and commit the fix with a regression test that fails first if it's right.

Three rounds produced 24 fixes. The four above are the ones that changed what the product says. Most of the rest were in the /write endpoint: it accepted "Assistant" as an audience and spoke it, a captured signed request could be replayed to roll back a newer correction, and a subscribe to …/weight_bearing/ returned {} and then never fired.

The part I'd repeat is the verification step, not the model. Every one of these bugs passed a green suite. The suite tested what I thought the product did.

The numbers, with their caveats

  • 344 tests, tsc --strict clean, plus 34 safety assertions run against a live server (npm run verify).
  • Local server path: signed write → notification → re-read, p95 3.0 ms, 200/200 inside a 3.4 s speech window. The window is an assumption (12 words at ~150 wpm), not a measurement of Alexa+ speech.
  • Over the public internet (client on another continent from the us-west2 server, both network legs paid twice): the committed receipt is 200/200 with p95 461 ms. Two re-runs today gave p95 710 ms and 966 ms, and one run landed 199/200. The server accounts for single-digit milliseconds, so that spread is the network.

What it doesn't do

  • It has never talked to Alexa+ itself. Whether Alexa+ declares resource subscriptions is untested. That's why the whats_changed fallback exists and is exercised end to end.
  • The patient is fictional, and this is not a medical device. The server renders the clinician's own words. It generates no clinical advice.
  • No one outside the project has used it yet. If you work with post-operative patients, I'd genuinely like to know where this breaks.

Repo: github.com/edycutjong/unsay (MIT) · Try it: api.unsay.edycu.dev/judge · Demo video: youtu.be/jJM5LWHONDU

Built for the Amazon Alexa+ hackathon track.

Top comments (0)