I'm going to start with the strongest argument against this post, because I'd rather you hear it from me. This month one of my open-source tools shipped with 1,558 tests passing and an authentication guard that never actually ran. An import failed silently, the suite stayed green, and an unauthenticated read path went out the door. A Docker test that ran the real thing in a real container caught it. Not my unit tests, and not my coverage number.
So when I say test evidence has become the most important thing I can show about my work, I'm not saying it as someone who has it figured out. I'm saying it as someone who got burned by his own green checkmarks and came out more convinced, and more curious, than before. That surprised me, and I've spent a good part of this year exploring why. What follows is where my thinking is right now, not a conclusion. I'm writing it partly to find out where I'm wrong.
I Used to Treat Tests as the Chore
For most of my career, tests were the thing you did so the real work could ship. I wrote them, I asked my teams to write them, and I enforced coverage gates, but if I'm honest, I never thought of tests as something that said anything about me. Code was the craft. Architecture was the craft. Tests were hygiene, like flossing. You're supposed to do it, and nobody compliments you for it.
That quietly flipped this year, and it took me a while to notice. I build a lot now with AI agents doing a large share of the typing. The code shows up fast, often faster than I can read it carefully. And I noticed that when I look at my own repos, the part I get excited about has changed. The code looks fine. It almost always looks fine, and looking fine has become cheap. What feels like mine now, the part an agent can't produce for me, is the evidence that the thing does what I claim when it meets reality. That shift has made building more fun, not less.
What I'm Learning to Point To
When I think about which of my projects I'd put in front of someone I respect, I don't reach for the cleverest code. I reach for three things, and all of them are about proof.
The first is the test count and where it runs. My agent control plane shipped v0.1.0 with 925 tests, and I've learned to report both coverage numbers, 93% on my laptop and 95.45% in CI, because the gap between them is information too. The second is the field test report: what I ran the tool against, how many cases, what broke. For one planning tool, that meant 170 real goals for $0.49, which is cheap to run and was anything but cheap to design. The third, and the one I've come to value most, is the failure list: the bugs, the misses, and the things the tool still can't do.
That last one took some getting used to. There's an old instinct, probably from years in leadership rooms, that says you lead with strengths and handle weaknesses if someone asks. Publishing a failure list goes against that. A few releases in, I realized the failure list was the part people read most closely, because it proves the rest isn't marketing. It's also where the best conversations start.
What Worked
Making my evidence checkable by strangers
I started committing field test reports, test reports, and security audits directly into the repos: public, linked, and dated. The result I didn't plan for turned into one of the best things that happened to me this year. A reader audited one of my releases in public. He wrote down what he expected to find before looking, checked it against the published artifacts, and recorded every place where my release story didn't match my own evidence.
He found real contradictions. The engine was mostly right, but the story I'd told about it was partly wrong. My first reaction was a flinch, which I think is normal. My second reaction, a day later, was gratitude, because it was the most rigorous review I've ever received, and it only happened because the evidence was sitting there for him to check. You can't audit a vibe. You can audit a report, and being auditable turns out to be a very different feeling from being right.
Treating "it passed" as a claim
AI work taught me that a confident answer and a correct answer look identical until someone checks. Somewhere along the way that lesson moved from how I treat model output to how I treat my own. A green run is now a claim I have to back up. What ran? Against what data? Where's the report? When I built the review engine where two models debate a pull request, I fixed 13 bugs before I trusted a single number it produced, including a join that silently collapsed my dataset from 2,333 rows to 359. None of those were clever architecture bugs. All of them were "you haven't looked closely enough" bugs, and they were only visible because I'd forced myself to ask for evidence instead of accepting green.
What Didn't Work
A big number still fooled me
Back to the 1,558 tests. What I found most interesting about that bug wasn't the bug. Imports fail. It's that I relaxed. I saw a large green number and felt safe, and that feeling is exactly what I've spent months telling other people not to trust. I wrote the whole story up here, but the lesson I took personally was simpler. A big test count can be theater just as easily as a big README. Now I treat a big green number as the start of a conversation, not the end of one.
Any metric I show off, I'll eventually inflate
The other thing I learned is that the moment coverage becomes something to be proud of, it becomes something that's tempting to grow for its own sake. In one of my repos, a test whose entire body was pass made it into the suite. It could never fail and would never catch anything. In the same project, a harness found zero assertions to run in 57 of 65 files and reported 0 / 0 as success. Code review caught both. My metrics didn't, which I found oddly reassuring: people reading carefully still matter here.
With AI in the loop, inflating a metric like this costs almost nothing. You can generate a thousand tests in an afternoon. If coverage alone becomes the thing we judge engineers by, we'll get exactly what we measured. The good news is that the thing underneath coverage is much harder to fake, and that's where I've landed.
Where My Thinking Is Right Now
After a few months of exploring this, my current guess is that coverage isn't the career metric. The evidence trail is. The test count is the headline. What makes it believable is everything around it: the field test that ran the real thing, the report with the bad numbers left in, the list of what still breaks, and the fact that a stranger can check all of it without asking me.
This matters to me more than a methodology debate, because it's changed what I want to be known for. Early in my career I wanted to be the engineer with the clever design. Later I wanted to be the leader with the right strategy. Now, with agents writing more of the code every month, I find myself wanting something humbler and, I think, more durable: to be someone whose claims hold up when you check them.
What I'm Still Trying to Learn
Evidence is expensive. My field tests often take longer than the features they test, and I'm not convinced that scales beyond open-source work where I control the schedule. Hiring hasn't caught up either. Most interview loops still care more about solving a puzzle live than about a public failure report, and I understand why, even though I think it's measuring the past. And I've only done this as one person. On a real team, who owns the evidence trail? Is it the author, the reviewer, or some role we haven't named yet? I don't know yet, and these feel like some of the most interesting open questions in our field right now.
Your Turn
I'd love to learn how other people are thinking about this, especially anyone whose job has shifted toward reviewing more code than they write. If you think I've got this backwards, I want to hear that too.
When you look at someone's repo, what actually makes you trust it? Stars, tests, commit history, the README, or something else entirely?
Have you ever published a bad number on purpose? How did it feel, and how did people react?
If AI writes most of your code next year, what will you point to as proof that you're good at your job?
That last one is the question I'm most excited to hear answers to. I have a feeling the comments will teach me more than this post did.
Top comments (0)