The problem
We build Mushare, a messaging network for AI assistants: you tell Meta's Muse "ask Alice whether Saturday or Sunday works", and Alice's Muse shows her a card she answers with a tap. On Muse, Mushare is a skill: a SKILL.md with references, plus a small Python CLI that Muse runs on its own cloud computer.
The CLI has 200 unit tests and they all pass. They say nothing about the part that fails most: whether Muse reads the skill the way we meant. Does it pick our skill for "tell Alice I'm late"? Does it run the right command? Does it show the card, or retell the card in its own words? Only a real assistant can answer that, and its answers vary from run to run.
So we built a regression suite that talks to Muse the way a user does, in the browser, and scores what comes back.
The harness
Each test account is a real Muse account, logged in once in its own Chrome profile. A Node script drives the browsers with Playwright over the Chrome DevTools Protocol. A test case is a sentence a real user would say, plus what should happen. For one case the script:
- Opens a new side chat, so no earlier conversation leaks in.
- Says the request, for example "Ask Alice: Costco on Saturday or Sunday? Give her two options", sometimes with a receipt image it drew for this run.
- Waits until Muse is done (more on that below).
- Collects the Mushare cards on the page (each card is an iframe our CLI renders), the reply text and a screenshot.
- After the run, pulls our server's tool-call log and matches each call to the case by account and time window.
- Scores each case: at least one card, the card contains the right words, the right server calls happened (send_message, recall_message), the content went out as a share card rather than plain text, and Muse did not say "I already sent this".
The output is a one-line-per-case summary (PASS, FAIL, FLAKY, with the reason and a ▲ or ▼ against the last run) and an HTML report with screenshots. We run only the affected cases while iterating on the skill text, and the full suite before each release.
Memory is state
The first surprise: the same SKILL.md scored 15/16 in the morning and 10/16 in the afternoon. Nothing had changed in our code. What had changed was Muse. After dozens of test runs a day, it remembered them. It answered "what do I need to deal with?" from memory instead of running our command. It told us "I already sent this receipt to Alice last night" and stopped. It had even stored a "preference" that the user wanted a hand-computed bill preview, learned from our own test sentence.
Three changes made runs comparable again:
- Every case starts with one line asking Muse to pause its memory for that chat. The only cases that keep memory are the ones that test "look it up now, don't answer from memory".
- Content is generated fresh for every run: random products to compare, a different restaurant and dishes on a receipt drawn as an image, times and reasons for each message. A repeated try within one run gets new content too.
- A reply that says "already sent" or "send it again?" is scored as stale state, not as a skill failure.
The lesson is that an assistant's memory is test state, like a database. Reset it, or your results drift through the day.
The 155-second wait
A full run of 21 cases took 62 minutes, and almost every case took exactly 154 or 155 seconds, even "what can Mushare do?". Assistants are slow, but not that regular.
We watched the page with a second, read-only script and found the cause in our own code. To read the chat, the harness asked Playwright for the text of main. Muse's web app no longer had a main element. So every look waited for Playwright's 30-second timeout, then fell back to reading the whole page. "Done" meant three unchanged looks in a row, so every wait was about 155 seconds. Muse itself had usually finished in 30 to 40.
The fix had two parts:
- Read the page text in place (
document.body.innerText) and cut off the side panel, instead of waiting for an element. - Decide "done" from what the app shows while it works: a Stop button next to the input, and the assistant's status line ("Drafting message", "Sending message") instead of "Connected". Done is no Stop button, a Connected status and six quiet seconds.
The Stop button alone was not enough: between tool steps it disappears for a while, and one case stopped early while Muse was still working. The status line closed that gap.
A full run now takes 9 to 10 minutes. We also log, per wait, when the chat changed and what changed, so the next slow step shows up in the report instead of in our patience.
Two account groups, and what fresh accounts revealed
To halve the run time we added a second pair of test accounts and ran the two pairs side by side. One account per browser profile: tabs in one browser share a login. Cases name roles (the sender, the friend), not accounts, and a case that depends on another ("confirm the bill Leo sent") runs after it in the same group.
The new pair did more than save time. On the old accounts, "tell Alice I'll be late" always went through Mushare. On the new ones, Muse looked for Messenger and WhatsApp, said it had no way to reach Alice, and offered to connect Messenger. The old accounts had months of memory that Alice is a Mushare friend; a new user has none. Our scores had been measuring returning users only.
The cause was one line in the skill's frontmatter: includeInPrompt: false. With it, Muse does not list the skill among the skills it has on hand; it only finds it if it decides to search. Setting it to true puts the description line in every conversation (not the whole SKILL.md, just the description), and we moved the routing rule to the front of that line: if the user names a person and no app, check Mushare first. By the next full run the fresh accounts went from 5/10 to 8/10.
We now treat the two groups as two kinds of user: the old pair is a returning user, the new pair is a new one, and we compare each group with itself.
Ask the assistant why
When a case kept failing, the most useful debugging step was to open a new chat on that account and ask Muse why. It cannot see its earlier chats with memory paused, but it can describe how it decides, and that was usually enough:
- "How are my Mushare reminders set?" got a description of Muse's own scheduled job instead of our card. Muse said it had never opened the skill: the question looked like something it could answer from its job list. We added that question to the routing rules in the description line.
- "What do I need to deal with?" was answered from Muse's own task tracker. Muse pointed out a conflict we had written ourselves: the card rule said "add at most one line", so it would not put its own tasks after our card, and wrote everything as text instead. We allowed one exception.
- "Send the school notice to Alice" went out as plain text, not a share card. Muse explained that the verb "send" mapped straight to our send command before it ever reached the sharing rules. The fix was to decide by the content first: content from a source, or with two or more facts, is a share.
One caution: the assistant's account of itself is a hypothesis, not a fact, so every fix still has to pass the cases.
Results
In two days the full suite went from 13/16 on one account pair to 22/22 on two, and a full run from about an hour to under ten minutes. Most of the score came from skill-text changes the suite made safe to try; one came from a test we had not written (a share rule that turned "confirm dinner at 7, add it to the calendar" into a share card instead of a calendar event, found while shooting screenshots and now its own case).
| Skill version | What changed | Cases passed | Full run |
|---|---|---|---|
| 0.61.2 | Skill rewritten in Simplified Technical English | 13/16 (one pair) | about 45 min |
| 0.61.6 | Second account pair added | 12/21 | 62 min |
| 0.61.9 | Skill listed in every chat; routing rule first | 17/21 | 9 min (wait fix) |
| 0.61.15 | General questions routed; send decides by content | 18/21 | 9 min |
| 0.61.16 | Settling a time is a message with a calendar event | 22/22 | about 10 min |
CI for the codebase went from 5.5 to 3 minutes the same week (three parallel shards), and releases that change only skill text no longer wait for it: CI cannot test how an assistant reads a sentence. The regression suite can.
What we would tell someone starting out
- Test the skill where it runs. Unit tests cover your code; only the real assistant tells you how it reads your words.
- Score from two sides. The page shows what the user saw; your server log shows what actually happened. A card without a send, or a send without a card, are both failures.
- Treat memory as state. Pause it, generate fresh content, and label "already done" replies as stale instead of failing the skill.
- Keep a fresh account. Your old test accounts are your most loyal users. A new one shows what a stranger gets.
- Time every wait. A constant duration is a bug in your harness, not in the model.
- Ask the agent, then verify. Its explanation is a good hypothesis; the cases decide.
- Write the skill plainly. We write ours in ASD-STE100 Simplified Technical English: short sentences, one instruction each, the condition first. Muse told us it misses conditions buried mid-paragraph, and the scores agreed.
Rong Zhou is the CTO and one of two founders of Mushare (https://mushare.ai), a messaging network for AI assistants.
Originally published at mushare.ai.
Top comments (1)
The fresh-account result makes memory isolation an important test assertion, not just a setup instruction. Asking the assistant to pause memory could itself be ignored, so I would include a canary fact learned in a previous chat and check whether it leaks into a supposedly isolated case.
The account-and-time-window match also seems worth strengthening with a per-case correlation marker in the generated content. With two groups and retries, a delayed send can fall into the next case's window. Joining page evidence and server calls on that marker would make a passing card harder to attribute to the wrong action.