Every release of Cá Viên Chiên ends with the same step: walk every changed flow on production, in Chromium and WebKit, at phone and desktop sizes, in both languages. Co-op rooms run on the real room server, share codes come from the real API, rewards are written to the real database.
That's a lot of fake players. Early on they all showed up in Google Analytics and in my own player count, so I wrote down one rule:
Test traffic never goes to Google Analytics. Our own API and database are used for real, with a test tag, and the tagged data is deleted when the tests are done.
Here is how that works on Cloudflare Pages + D1, and where it broke.
1. Nothing reaches Google Analytics
Every Playwright context goes through one helper before any page opens:
export async function guardContext(context, { presence = 'tag' } = {}) {
// GA4's own opt-out flag, so gtag sends nothing even where its script loads
await context.addInitScript({ content: `window['ga-disable-${GA_ID}'] = true;` });
// Silence: the game's Web Audio graph runs into a gain of 0
await context.addInitScript({ content: MUTE_INIT });
// Drop analytics and ad hosts before the request leaves the browser
for (const pattern of BLOCKED) await context.route(pattern, (r) => r.abort('blockedbyclient'));
// Our own player count goes through, tagged as a test browser
if (presence === 'tag') {
await context.route('**/api/presence**', (r) =>
r.continue({ headers: { ...r.request().headers(), 'x-cvc-test': '1' } }));
}
}
BLOCKED covers Tag Manager, google-analytics.com, analytics.google.com, DoubleClick and AdSense. A local astro build loads GA4 too, so the guard applies to localhost as well.
There is a test for the guard itself: it scans every script in scripts/ that opens the site and fails if one doesn't import guardContext.
Chrome DevTools tabs driven over MCP can't block requests, so they get the same thing as an init script: the GA opt-out flag, the mute, and a fetch wrapper that adds the header.
2. The real API, with a tag
I don't fake our own endpoints. A faked /api/presence means no play record, which means no share code, which means the referral reward flow can't be tested at all. I found that out the hard way, by skipping the reward test on production once.
So the server stores test browsers like anyone else, with a flag:
// functions/api/presence.ts
const test = request.headers.get('X-CVC-Test') === '1' ? 1 : null;
await db.prepare(
'INSERT OR IGNORE INTO installs (id, …, test) VALUES (?, …, ?)'
).bind(id, …, test).run();
The statistics page and pnpm stats skip installs.test = 1. Everything that hangs off an install (play days, first-session steps, replays, share codes, referral rewards, poster events) can be traced back to a tagged record.
3. Delete it when you're done
After every run, local or production:
node scripts/clean-test-data.mjs --remote # delete
node scripts/clean-test-data.mjs --remote --dry # count only: must be 0 everywhere
The script selects by the tag only, deletes dependents first, then the records, all in one call, then counts again and exits non-zero if anything is left:
const T = 'SELECT id FROM installs WHERE test = 1';
const TABLES = [
['referral_rewards', `newcomer_install_id IN (${T}) OR code IN (SELECT code FROM share_codes WHERE install_id IN (${T}))`],
['share_codes', `install_id IN (${T})`],
['replays', `install_id IN (${T})`],
['install_steps', `install_id IN (${T})`],
['poster_events', `install_id IN (${T})`],
['install_days', `install_id IN (${T})`],
['installs', 'test = 1'],
];
run(TABLES.map(([t, w]) => `DELETE FROM ${t} WHERE ${w};`).join(' '));
The release checklist has test-rules before the deploy and test-cleanup after it. The deploy script refuses to run until the before items are answered, and the after items must pass before the release record is committed.
After the 1.37.0 release walk it deleted 129 test installs and everything attached to them. A real player's row is never selected.
4. The table that had no tag
Then I tested the new rating stars.
The game page has a "rate this game" card. My check opened it, hovered, and voted 4 stars, at two widths, in two languages, in two engines. Eight votes. Locally that was fine. On production, the log told the story on its own:
4 stars saved | 4.42/5 · 57 player ratings → 4.41/5 · 58 player ratings
…
4 stars saved | 4.38/5 · 64 player ratings → 4.37/5 · 65 player ratings
The ratings table has a voter id, a score, a hashed IP and timestamps, but no install id and no test flag. My cleanup script had no way to find these rows, and every vote moved the public average shown on every page and in the structured data.
Cleaning it up carefully:
- List the newest rows, read-only. Eight rows from the same IP hash, alternating vi / en, all 4 stars, inside a three-minute window, matching the four Chromium and four WebKit runs exactly. The vote just before them came from a different address and stayed.
- Count before deleting, with exactly the conditions the delete would use:
SELECT COUNT(*) FROM ratings
WHERE ip_hash LIKE 'a2c680bee7%' AND score = 4
AND created_at BETWEEN strftime('%s','2026-10-03 13:50:00')*1000
AND strftime('%s','2026-10-03 13:53:30')*1000;
-- 8
- Run the delete with the same
WHERE:changes: 8. - Check the result: back to 57 votes and 4.42. The public summary is cached at the edge for 5 minutes, so the API caught up a few minutes later.
Then I made sure it can't happen the same way twice:
- The rating check now stubs the
POSTwhen it runs on production and reads the live summary for its assertions. The real vote is tested only on the local stack. - The testing rules say it in one line: a test never votes on production, because a rating can't be tagged or found afterwards.
- The script was run once more against production with the stub: 24 / 24 checks, still 57 votes.
What I'd tell myself a month ago
- Block analytics in the browser, not in a dashboard filter. GA4's internal-traffic filters only apply from the day you turn them on.
- Tag at the source. One request header, one column, and every derived row can be traced back to it.
- Make cleanup a script with a count, not a habit. Exit non-zero if anything is left.
- List every table a test can write to. Any table that can't be traced back to a tagged record needs its writes stubbed on production.
-
Count before you delete on production, with exactly the
WHEREyou're going to use.
The game is free in the browser at cavienchien.net.
Top comments (0)