Most software gets a second chance. A bad deploy on Tuesday is fixed by Wednesday morning, and by Friday nobody remembers it happened. That safety net is so ordinary that it disappears from the way we talk about quality: we argue about coverage percentages because, underneath, we assume anything we miss can be corrected later.
I spent the last months on an app where that assumption is false, and it changed how I decide what to test.
The deployment model is the requirement
Rueda de Actos is a Windows desktop app I built for a filà, one of the associations that take part in the Moros y Cristianos festivities in Alcoy, Spain. Once a year the board sits down and hands out the spots for five acts, eleven places each, following a rotation that decides who takes part and who waits another year.
That session is the entire product. It happens once. It happens live, with the screen projected in front of the room. And it happens on a machine that is not connected to anything.
Three properties fall out of that:
- No network. The app never opens an outbound connection. Everything lives in a local SQLite file inside the user's Windows profile.
- No auto-update. There is no update channel to check, because there is no network. A new version means someone runs a new installer over the old one.
- One live run per year. Between two real sessions there are twelve months, and I am not in the room for eleven and a half of them.
Together they remove the thing every web developer leans on without noticing: I cannot ship a fix before it matters. The build installed on that machine is the build that runs. By the time anyone notices a problem, the decision it affected has already been made, out loud, in front of the people it affects.
Coverage percentage is the wrong question
The question I actually needed to answer was not how much of this is covered. It was:
If this path is wrong, when do we find out, and what does it cost to fix it then?
That splits the codebase into three groups.
- Wrong and obvious immediately. The user sees it, retries, moves on. Cheap. A test here buys very little.
- Wrong and obvious later, with the data still intact. Annoying, recoverable, worth some tests.
- Wrong and plausible. No error, no crash, nothing red anywhere. The app produces a result that looks exactly like a correct result and isn't. Or data quietly stops existing.
Roughly 400 tests, and almost all of them live in the third group.
Where the tests actually are
1. The rotation
This is the critical one, and it is critical precisely because it never fails loudly.
If the wheel logic is wrong, nothing throws. The screen shows eleven names per act, correctly formatted, in a table that looks right. What has actually happened is that somebody was skipped, or got a second turn before someone else got a first. Nobody sees a stack trace. Someone notices months later that they have been left out two years running, and at that point the damage is social, not technical. In a volunteer association, fairness is the entire reason this software exists.
So the rotation is not tested as a function with a known return value. It is tested as a set of invariants over simulated years:
// simplified
it('never grants a second turn before everyone has had a first', () => {
let wheel = newWheel(makeMembers(40))
for (let year = 0; year < 5; year++) {
const assigned = runYear(wheel) // 5 acts, 11 places each
wheel = advance(wheel, assigned)
}
expect(maxTurns(wheel) - minTurns(wheel)).toBeLessThanOrEqual(1)
})
The valuable tests here run the wheel across several years and then assert a property of the outcome, never an exact list of names. Exact-list assertions break the moment a rule changes, and they teach you nothing when they do. Invariants survive rule changes, which is the whole point: they encode what fairness means, not what the current implementation happens to produce.
The same style covers the awkward cases that real use produced. Skipping an act must not advance anyone's turn. Closing an act half full must not pretend it was complete. Reopening a closed act to correct a mistake must leave the queue exactly as it would have been had the correction been made the first time.
2. The rules that block a spot
Gender rules for squads, payment status, sanctions, licence validity, and the rule against repeating an act you already did in a previous year. On top of that, three configuration switches that each turn one of those rules off, because different associations work differently.
They interact, so they are tested as a matrix rather than one by one. A member can be eligible under the gender rule and blocked by the licence, or blocked by both, or unblocked by a switch that only lifts one of the two.
One of these tests is not about logic at all. Payment and sanction status block a spot but must never be rendered in the projected table, because the room is watching. So there is a test asserting that the sensitive field does not appear in the table markup for a blocked member. It is a privacy requirement, and writing it as a test is what stops it being quietly undone by a future refactor of that component.
3. Schema migrations
The tests nobody applauds, and the ones I would keep if I could keep only one group.
Every year, the installed version is a year old. Upgrading means running the new installer over the previous one, and the database sitting in the user's profile is the real one, with the real history of who took part in what. There is no seed data to fall back on and no server-side copy anywhere.
So every schema version ships with a migration, and every migration ships with a test that starts from a fixture database at the previous version, migrates it, and asserts everything survived: members, years, acts, flags, corrections. Not "the migration completed without throwing", but the row counts and the row contents, plus whatever the new version adds.
Two more tests guard the edges. Migrations must be ordered and idempotent, and the app must refuse to open a database written by a newer version rather than trying its best. That last one sounds theoretical until you remember how these installers travel. Reinstalling last year's build from an old pendrive is a completely realistic accident.
4. Import and export, in both directions
The app starts empty. Real data gets in through an import from the spreadsheet the association already had, which is to say: a file a human maintained by hand, for years, with no schema in mind.
That means accents, trailing spaces, inconsistent capitalisation, blank rows in the middle, a header row that is not the first row, and a name column that might be called Nombre, NOMBRE or Nombre y apellidos. The alternative to importing that file is typing everyone in by hand, which kills a tool before it is ever used. The messy path is the main path.
Excel and JSON both work in both directions, and the tests treat that as a round trip: export the full wheel, re-import it into an empty database, assert the resulting state is identical. An export you cannot re-import is not a backup, it is a report.
The fixtures are deliberately dirty. The clean file passes on day one and never catches anything again.
5. The build itself
The suite runs on every build of the installer. The installer is not produced by hand from whatever happens to be in my working tree: an automated process builds it from source with the same pinned dependency versions the tests ran against. If the suite fails, no installer comes out.
That matters more here than in a service I can redeploy. The artefact is physically handed over. The file that reaches that machine is the last point at which anything can still be corrected, so it should not be possible for it to exist in an untested state.
It is not perfect. The installer is still unsigned, so Windows SmartScreen warns the first time it runs, and no amount of testing fixes that particular rough edge. It is a signing certificate, not a bug.
What I deliberately don't test
- Visual layout. Checked by using the app, not by asserting on pixels.
- Electron and SQLite themselves. They have their own test suites, written by people who know them better than I do.
- Components with no logic in them. A unit test there mostly proves I typed the props twice.
- Performance. Dozens of people and a handful of years of history. Nothing here will ever be slow, and a benchmark would only give me something to maintain.
None of those are unrecoverable if they are wrong. That was the only criterion.
The rule I ended up with
Test budget is finite, whatever anyone says about coverage. The way I spend it now is by asking, for every path, when a mistake would surface and what fixing it would cost at that point.
Loud failures are cheap, because the user becomes the test. Silent, plausible failures are expensive, because nothing tells you they happened, and by the time something does, the moment has passed. The rotation is the purest example I have worked on: no exception, no error state, a result that looks completely reasonable, and a consequence that arrives months later and lands on a person rather than on a log.
In a few weeks the app runs for real again, offline, on one machine, in front of a room. Feature work has already stopped. From here it is tests, polish and leaving things alone.
I will write about how it goes.
Top comments (0)