A passing test usually feels like a good result. The test ran, the assertion passed, and we get a green check. But recently I started thinking about something a little differently:
When a test passes, what did it actually prove?
It proves that the test reached the expected result for that particular scenario. But does it also prove that we tested the right behaviour? Does it prove that we covered the important part of the user flow? And does it tell us what we might still be missing?
That made me look at passing tests a little differently.
I started thinking about this more while exploring AI-assisted testing. As I looked into how AI can help with generating and maintaining tests, A test can pass because everything worked as expected. But it can also pass because the assertion isn't checking enough, the test data doesn't exercise the important case, or the scenario itself doesn't cover the risk we actually care about.
That's where I think the context around a test becomes important. At X360 AI Tech, I came across an approach where requirements, test coverage, execution, and the result are connected instead of looking at the final pass/fail status on its own.
I liked that idea because it changes the question a little.
Instead of only asking: Did the test pass?
we can also ask: What behaviour did this test actually validate?
I don't think this means every passing test needs a detailed investigation. But when AI is generating or maintaining more tests for us, I think this becomes more important.
If creating tests becomes easier, understanding what those tests actually prove may become the harder part. For me, a green test is useful. But knowing "why it is green and what it actually checked" gives me much more confidence in the result.
That's something I'm still exploring.
Top comments (1)
This is a question worth sitting with longer than the post does. A test suite full of AI generated tests with weak assertions can hit 100% pass while covering almost none of the real risk. Coverage percentage and proof of behavior aren't the same thing, and mixing them up gets easy to do once generating tests becomes cheap.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.