DEV Community

Cover image for Similarity Is Not Authority
Keenen Wilkins
Keenen Wilkins

Posted on

Similarity Is Not Authority

Week four of i.c.stars produced four documents. A defect log. A peer review of somebody else's defect log. Three runbooks. A run of show.

None of them are code. All four turned out to be the same document.

What the testing session actually taught

We were given a portfolio site with bugs planted in it and told to find them and write them up.

The format was given, when, then. Given each nav link opens the page it names, when I click Gallery, then the Gallery page opens. Then the defect log underneath it: steps to reproduce, expected, actual, severity.

My group ran three tests. All three failed, all three high severity. Gallery opened About Me. The third image opened the second. The contact email would not open a message.

Writing that up took longer than finding any of it.

Then the part that made the session worth it. We traded reports with another team and had to go reproduce their findings on the live site and mark each item verified or not verified.

Six of their eight held up. One I could not reproduce at all. One only reproduced on the machine of the person who reported it, and their write-up never said which machine, which browser, or which page they started from.

The verdict I wrote at the bottom of that review: strong at finding bugs, weak at recording them.

That is not an insult. It is the normal failure. Finding a bug feels like the work. Writing it so a stranger can hit it again in ninety seconds is the actual work, and nothing tells you whether you did it except handing it to a stranger.

Three runbooks

Later that week, presentations, and before them, runbooks. Murphy's Law applied to a room. Write down what you do when it breaks, before it breaks.

Setup, happy path, edge case, recovery script. I wrote three.

Slides stop advancing. The clicker dies or animations fire on their own and the deck runs ahead of whoever is talking. Stop clicking. Say "Looks like the slides want to get ahead of me, so I'll talk you through this part." Keep going off the printed outline while a teammate advances by hand.

A handoff breaks. The next speaker misses the cue and the room sits in silence. Bridge without stopping, then hand off at the next slide. The rule underneath it is the one worth keeping: every presenter knows the main point of every section, not just their own.

The laptop or projector dies. "While we get the screen sorted, let me tell you what you would be looking at." Present from the printed outline. And the line that takes actual discipline:

Do not spend more than 60 seconds troubleshooting in front of the room.

Then a preflight list. Deck in three places including a PDF with no animations. A printed outline per presenter. Cue lines agreed in advance. One teammate named as backup on the controls.

A run of show

The fourth document. Every section of the presentation with a named owner and a time block. Not "the team covers the design." A person, a topic, a number of minutes.

And scripted transitions. The exact words, written down, for every handoff between speakers. That felt excessive when I was filling it in. It is the single reason three people who are not presenters got through it sounding like one team. Nobody had to invent a sentence in front of an audience.

My sections were the solution design and the reasoning behind it, then the logic and the wireframes.

The thing I actually contributed to the product

Here is where I should be precise, because it would be easy to imply more than is true.

I did not write the retrieval code. A model wrote most of it, fast, and that is why we had something standing in four weeks instead of twelve.

What I brought was the argument about what the system is not allowed to do.

The problem, stripped of anything about the client: volunteers at a nonprofit need answers about how their organization works. The answers live in documents the organization has already approved. It looks like a search problem. Store the documents, take a question, find the best-matching passage, answer with it.

That framing is wrong in a specific way, and noticing it was my contribution.

Two documents can both apply to the same person. Both official. They disagree. One of them overrides the other, but only in some places and not others, and nothing inside either document records that.

A similarity search cannot see that, ever. It can tell you two passages are about the same subject. It has no way to express that one of them governs and the other does not. Override is a relationship between documents. Similarity is a score on a single document. You cannot get the first from the second by ranking harder.

A score lives on one document. Override lives between two.

So the design has two behaviors and no third one. Either the system answers with the document, the section, and the effective date attached, or it stops and hands a briefed summary to a person. Checks run before anything gets written and decide which of the two happens.

Two behaviors and no third one: four checks, then answer with a citation or stop and hand off

The refusal is not a fallback. It is a feature with equal standing. A confident wrong answer about a governing rule can invalidate a decision that real people already made and acted on. A refusal costs somebody ten minutes.

That is the work of a solutions analyst, which is my role on the team. Not writing the function. Deciding what the function must never be allowed to conclude.

Four documents, one idea

Four documents and one product decision, and what breaks if each does not survive the handoff

Look back at the week.

A test case is written so a stranger can reproduce it. A defect log has severity on it so somebody else can decide what to fix first. A runbook exists so the presentation survives when the person holding the clicker cannot continue. A run of show has named owners so nobody waits to find out whose turn it is.

And the product rule I argued for is the same shape: when the system cannot prove which document governs, it stops and hands the question to a human, with everything that human needs to answer it.

Every deliverable that week was about surviving a handoff. The workshops were not teaching four skills. They were teaching one, from four angles, and the fourth angle was in the product.

Where it actually stands

We presented our initial solution. Slides, in a room, fifteen minutes.

We did not demo. That was my call, and my reason at the time was that this round was a solution presentation rather than a demo, and I wanted to be sure the thing worked for this specific organization before putting it on a screen.

Then one of the judges said presentations should demonstrate, not read.

We used eight and a half minutes of our fifteen. My closing line started with the words "I guess." We did not win the round.

And the part I have not resolved: the prototype runs, the tests pass, and I have not personally verified either of those claims. A model built it and a model checked it. I wrote the rules it has to obey and I have not yet sat down and confirmed it obeys them.

Which is a strange position to argue from in a room. I could defend every design decision in that solution. I am not sure I could have answered a hard question about the code underneath it.

That gap is the actual finding from week four, and it is not one a worksheet covers.

What is coming

I am in Geek Week as I write this. In the office all week, fifteen tasks passed in sequence, three attempts each before you move on and come back. No AI tools. Notes at the table while practicing, nothing during the test.

It is the first week where nothing I produce is produced by anything but me.

If you have built something with a model recently: what did you personally verify, and what did you take on faith because it ran?

The next one covers Geek Week.

Top comments (0)