DEV Community

Cover image for I Stopped Trying to Prove My Code Works. I Started Trying to Prove It Wrong.
Michael Placzek
Michael Placzek

Posted on AI-assisted

I Stopped Trying to Prove My Code Works. I Started Trying to Prove It Wrong.

For a long time, my tests asked a pretty normal question:

Does this code work?

Valid input goes in.

Expected output comes out.

Green checkmark.

Ship it.

Then I started building Network Doctor, an open-source network troubleshooting tool.

And that mindset stopped being good enough.

Because network software lives in the horrible little gaps between "valid" and "invalid."

An IPv6 address can be valid and still break your assumptions.

A simulated failure can look realistic while actually describing an impossible network.

A privacy filter can correctly redact 99% of addresses and still leak the one format you forgot existed.

A test suite can be completely green while the program is still wrong.

So I changed the question.

Instead of asking:

Can I prove this works?

I started asking:

How can I prove this is wrong?

That change has probably improved the project more than any individual feature I've added.

The happy path is the easy part

Network Doctor checks multiple layers of a connection:

  • network interfaces
  • DNS
  • TCP
  • TLS
  • HTTP
  • proxies
  • routing
  • path MTU

The normal cases are surprisingly boring.

DNS resolves.

TCP connects.

TLS negotiates.

HTTP returns a response.

The interesting bugs are hiding around the edges.

So I started deliberately feeding the project things I hoped would break it.

And they did.

IPv6 has been particularly good at humbling me

Take an address like this:

fe80::1%eth0
Enter fullscreen mode Exit fullscreen mode

That %eth0 matters.

It identifies the interface scope for a link-local IPv6 address.

If you're writing code that parses, displays, compares, exports, or redacts network addresses, suddenly:

fe80::1
Enter fullscreen mode Exit fullscreen mode

and:

fe80::1%eth0
Enter fullscreen mode Exit fullscreen mode

are not necessarily interchangeable strings.

Then there are bracketed addresses:

[2001:db8::1]
Enter fullscreen mode Exit fullscreen mode

And addresses embedded inside other text.

And equivalent IPv6 spellings.

And IPv6 addresses whose final group contains hexadecimal letters instead of looking suspiciously like an IPv4 address.

Each one is perfectly reasonable.

Each one is also an opportunity for a program to quietly make a bad assumption.

I have fixed multiple bugs in Network Doctor simply by asking:

What is the weirdest valid version of this input I can think of?

That question is becoming one of my favorite testing tools.

Then I started attacking the simulator

Network Doctor has a simulator because reproducing real network failures on demand is difficult.

I can describe situations like:

DNS is slow
Enter fullscreen mode Exit fullscreen mode

or:

there is no default route
Enter fullscreen mode Exit fullscreen mode

and exercise the diagnostic logic against them.

At first, I mostly tested whether valid scenarios worked.

Then I realized the simulator itself could lie.

What if a scenario claims there is no default route, but its other configuration requires one?

What if two simulated services try to listen on the same address?

What if a scheduled DNS delay is smaller than the simulator's timing precision?

What if a timeline contains latency values that cannot actually be represented correctly?

Those are not necessarily bugs in the diagnostic engine.

They are bugs in the laboratory used to test the diagnostic engine.

That is worse.

If your test environment can represent impossible states, a passing test can give you confidence in something that was never true in the first place.

So now the simulator gets attacked too.

My new rule: prove the bug before fixing the bug

This became especially important once I started using AI coding tools heavily.

AI is very good at looking at code and saying:

I found a bug.

Sometimes it really did.

Sometimes it found something that looked suspicious but was intentional.

Sometimes it misunderstood the surrounding code.

And sometimes it confidently proposed a fix for a problem that did not exist.

So I adopted a rule:

No proof, no fix.

The workflow looks roughly like this:

  1. Find a suspicious behavior.
  2. Reproduce it.
  3. Write a test that fails because of it.
  4. Confirm the failure is actually wrong.
  5. Make the smallest reasonable change.
  6. Run the regression test.
  7. Run the rest of the test suite.
  8. Try to break the new behavior again.

For Network Doctor that can include things like:

go test ./...
go test -race ./...
./scripts/check
Enter fullscreen mode Exit fullscreen mode

along with fuzzing, cross-platform checks, simulator validation, and other project-specific tests.

The important part is not the commands.

The important part is the order.

The fix comes after the evidence.

A green checkmark does not mean "correct"

This sounds obvious.

But I think it is an incredibly easy trap to fall into.

A test suite only proves the things you asked it to prove.

If all your tests say:

Given normal input X,
does function Y return Z?
Enter fullscreen mode Exit fullscreen mode

then congratulations.

You have learned a lot about normal input X.

Production is currently preparing input Q-17 from a Windows machine connected through a VPN with a scoped IPv6 address and a proxy configuration you have never seen before.

Good luck.

So I have been trying to make tests more adversarial.

Instead of only:

Does it accept valid input?
Enter fullscreen mode Exit fullscreen mode

I want:

Which valid input looks invalid?
Enter fullscreen mode Exit fullscreen mode

Instead of:

Does redaction work?
Enter fullscreen mode Exit fullscreen mode

I want:

What representation of this address escapes redaction?
Enter fullscreen mode Exit fullscreen mode

Instead of:

Does the simulator run this scenario?
Enter fullscreen mode Exit fullscreen mode

I want:

Can I construct a scenario that should never be allowed to exist?
Enter fullscreen mode Exit fullscreen mode

Those questions find much more interesting bugs.

Strangely, this has made AI more useful

I use Claude Code, Codex, and local models during development.

The more I use them, the less interested I am in whether they can generate lots of code.

Generating code is easy now.

Verification is the expensive part.

So one of the most useful things I can ask an AI coding agent is no longer:

Implement this feature.

It is:

Try to prove this implementation is wrong.

Look for counterexamples.

Find inputs that violate assumptions.

Trace the behavior through another platform.

Construct a regression test.

Compare what the documentation promises with what the code actually guarantees.

That turns AI from a code generator into something closer to an extremely persistent adversarial reviewer.

It still gets things wrong.

But that's fine.

Because it has to prove its case too.

This feels slower

Sometimes it is.

A tiny bug can become:

investigation
→ reproduction
→ regression test
→ fix
→ full validation
Enter fullscreen mode Exit fullscreen mode

instead of:

change three lines
→ looks good
→ merge
Enter fullscreen mode Exit fullscreen mode

But the second workflow only looks faster because it stops measuring time before the bug comes back.

I would rather spend an extra hour proving a fix today than spend a day figuring out why an "obvious" fix caused something bizarre three releases later.

I don't want tests that agree with me

I want tests that are annoying.

I want fuzzers discovering inputs I didn't consider.

I want the simulator refusing contradictory scenarios.

I want privacy tests assuming I forgot another representation of an IP address.

I want CI rejecting something that looked completely harmless.

I want another developer to look at my implementation and ask:

But what happens if...?

That sentence is incredibly valuable.

Software gets interesting exactly where our assumptions stop working.

So these days, when I finish a feature, I try not to ask:

Did I make this work?

I ask:

If I wanted to embarrass the person who wrote this, how would I break it?

Unfortunately, that person is usually me.

And so far, it has been a pretty effective development strategy.


Network Doctor is open source, and I am continuing to build it in public.

What is the weirdest edge case your tests have ever caught before a user did?

I genuinely want to hear them.

Top comments (0)