DEV Community

Cover image for The Frog That Lived: Let's learn about falsification.
Sheraz Manzoor
Sheraz Manzoor

Posted on

The Frog That Lived: Let's learn about falsification.

Suppose! You're a biologist and you have a claim based on your hypothesis that:

All the frogs we put in a deep freezer will die

To prove your claim you do an experiment.

You put 10 frogs in a freezer and the next day they all are dead in the deep freezer. Your confidence in your claim will rise. Now you do the same experiment on the 90 more frogs (10 frogs in a batch) and you get the same result. You'll be even more confident in your claim.

But can you say, your claim is scientifically proven? Never!

Now imagine that one of the frog from the eleventh batch survive under same conditions in the deep freezer. What does that mean? That falsified your claim. We don't need to do 100 more experiments to prove that the claim "All frogs in a deep freezer dies" is false. That's called falsification.

That asymmetry is the important part:

A finite collection of confirming examples cannot prove a universal claim, but a single valid counterexample can falsify one.

Karl Popper made this asymmetry central to his philosophy of science. In The Logic of Scientific Discovery, he argued that universal theories cannot be derived from finite observations, while a suitable observation that contradicts a theory can refute it. The same logical structure appears in engineering every time we test a claim about software. Popper, The Logic of Scientific Discovery Stanford Encyclopedia of Philosophy: The Problem of Induction

There is an important engineering distinction here: a passing observation is usually compatible with the hypothesis, while a valid counterexample is incompatible with it.

That is why one test failure can be much more informative than thousands of passing tests.


The Problem of Induction in Software Testing

Software testing is induction with a CI pipeline. You run a finite number of cases and infer a general claim: "this function is correct." Edsger Dijkstra said it bluntly decades ago: testing can reveal the presence of bugs, never their absence.

Suppose this function is intended to sort integers:

def sort_numbers(xs):
    return sorted(set(xs))
Enter fullscreen mode Exit fullscreen mode

It works for many inputs.

It works for:

[]
[1]
[3, 1, 2]
[5, 2, 9, 7]
Enter fullscreen mode Exit fullscreen mode

A large regression suite might happily contain thousands of examples like these.

But the function silently removes duplicates.

The counterexample is tiny:

[0, 0]
Enter fullscreen mode Exit fullscreen mode

Expected:

[0, 0]
Enter fullscreen mode Exit fullscreen mode

Actual:

[0]
Enter fullscreen mode Exit fullscreen mode

One failure is enough to falsify the hypothesis:

"This function correctly sorts every list of integers."

A system with 99.999% boring behavior does not become 100% safe because the boring behavior was observed often.

The missing 0.001% may contain the incident.


Real world examples of falsification in Engineering

1. Google

Google had a large distributed MySQL environment, with many dependent services. The engineers assumed that if one the database fails, the rest of the system will continue working.

Instead of waiting for the database to crash in production, they tested their assumption. They intentionally failed one database out of hundreds. They expected a controlled failure.

But reality was different.

Within minutes of commercing test, numerous dependent services reported that both external and internal users were unable to access important systems. Some systems were intermittently or only partially accessible.

Google immediately aborted their test, restored the access and started investigating the issue.

They didn't found one issue, they found several issues:

  • Hidden dependencies existed between systems that were supposed to tolerate the database failure.
  • Their rollback procedure had not been sufficiently tested and was itself flawed.
  • A flaw existed in the database application-layer library.
  • Their incident-response process had not been sufficiently exercised.

Thw important part is what they did after all this.

Google fixed the library, improved the recovery process, and instituted periodic retesting. Their own SRE documentation explicitly describes the goal as breaking systems, observing how they fail, discovering hidden weaknesses, and then making changes to prevent those failures from recurring.

Source: Google SRE

2. Netflix

Netflix's move to AWS created a very different failure environment.

In a traditional data center, an individual server failing might be relatively infrequent and handled operationally.

In a large distributed cloud environment, individual instances can disappear, networks can fail, dependencies can become unavailable, and components can behave unpredictably. So the engineer's thought "Our services will continue working when individual infrastructure components fail."

So being the responsible engineers, instead of believing their assumption, they decided to test it. So they created an open source tool to test their assumptions called Chaos Monkey.

Chaos Monkey randomly terminates virtual machine instances and containers that run inside of your production environment. Exposing engineers to failures more frequently incentivize them to build resilient services.

But their test was even brutal.They built a Chaos KOng, which doesn't just kill a serer, it kills an entire AWS region.

During the Chaos Kong in exercise they found that traffic evacuate from the west region. The east region got a corresponding bump in traffic as it stepped up to play the role of saviour. As long as the aggregate metric followed that relatively smooth trend, Netflix knew that their system is resilient to the failover. At the end of the exercise, traffic reverted back to the west region.

AFter this reliability question changed forever. Instead of asking do we have two servers, they started asking what if we kill one?
Since then they regularly runChaos Kong exercises and they say, "it gives us confidence that even if an entire region goes down, we can still server our customers".

Source: Chaos Engineering Upgraded

3. My personal experience (A case study worth reading)

We were making a payment service in one of our project, that handled transactions through an external payment provider. The flow was very very simple:

Client -> Create Payment -> Payment Service -> Payment Provider -> Database

We ran unit tests. We had integration, duplication and timeouts tests. We had tests for successful payments and provider failure. It was a green forest. So we assumed that "If a customer retries the same payment request, we won't charge them twice".

That assumption sounded reasonable because we had an idempotency mechanism. The only problem was that we were testing known examples. We weren't actively trying to destroy the assumption.

So I changed the question. Instead of thinking does our idempotency code work, I thought what sequence of events would falsify the claim that our payment operation is idempotent.

AN interesting possibility arose.

What if:

Request 1 -> Payment Provider successfully charges the card -> Network timeout occurs -> Our API never receives the response -> Client retries -> Request 2

I deliberately produced this issue.

I mocked the payment:

Accept payment -> Record the charge -> Pretend the response was lost -> Return a timeout to our application.

Then I sent the same logical payment again. And you know what happened?
The second request generated another charge. This falsified our assumption before moving the app to production.

The solution was not simply

if retry:
    don't charge
Enter fullscreen mode Exit fullscreen mode

We redesigned the operation around a stronger idempotency contract.

The request received a stable idempotency key:

POST /payments

Idempotency-Key: 8d7e...
Enter fullscreen mode Exit fullscreen mode

The service persisted the operation state before allowing the workflow to complete:

idempotency_key
status
provider_transaction_id
result
Enter fullscreen mode Exit fullscreen mode

A repeated request with the same key would then resolve to the existing operation instead of creating another payment.

More importantly, I kept the original failure scenario as a permanent regression test.


def test_payment_is_idempotent_when_provider_succeeds_but_response_is_lost():
    provider.charge_side_effect = [
        ChargeSucceededButResponseLost(),
        ExistingChargeResponse(),
    ]

    first = create_payment(idempotency_key="abc")
    second = create_payment(idempotency_key="abc")

    assert number_of_provider_charges() == 1
    assert second.payment_id == first.payment_id

Enter fullscreen mode Exit fullscreen mode

We didn't discover the bug cuz we added another 500 happy-path tests, we discovered it because we added what would falsify our assumption.


Falsification in Practice: Property-Based Testing, Fuzzing, and Model Checking

Property-based testing: a frog hunter, not a frog counter

Example-based tests say "for this input, expect that output." Property-based testing (PBT), pioneered by QuickCheck and available in Python via Hypothesis, says "for all inputs of this shape, this property should hold" and then tries to break it.

Here are two properties any sort must satisfy:

from collections import Counter
from hypothesis import given, strategies as st

@given(st.lists(st.integers()))
def test_sort_properties(xs):
    out = my_sort(xs)

    # Property 1: output is in non-decreasing order
    assert all(a <= b for a, b in zip(out, out[1:]))

    # Property 2: output is a permutation of the input
Enter fullscreen mode Exit fullscreen mode

Our function silently drops duplicates. Hypothesis found a counterexample and then shrank it to the smallest failing input, [0, 0].

The test never asked "does [3, 1, 2] sort correctly?" It asked "can I find any input that violates these properties?"

An honest caveat is that PBT is still sampling. Hypothesis runs 100 examples per test by default, and 100 passes prove nothing about the 101st. PBT is a better frog hunter, not a proof engine. Its advantage is that it looks in places you wouldn't think to look.

Model checking: search for the counterexample trace

Formal methods embrace the asymmetry too. A model checker such as TLC (for TLA+) or SPIN takes a model of your system and an invariant and asks: is there any reachable state that violates it? If there is, you get a counterexample trace, the exact sequence of steps that leads to the violation.

Amazon has publicly described using TLA+ to find subtle concurrency and fault-tolerance bugs in distributed systems, including ones whose failing traces were dozens of steps long. No human writes a test for step 35 of an obscure interleaving, but a checker that exhaustively searches for a falsifying path will find it.

Fine print: if a checker exhausts a finite model without finding a violation, you've proven the invariant for that model. The model is not your production system, and you're still one wrong assumption away from a surviving frog.

Fuzzing: you only need one crash

You don't have to prove a parser is safe. You have to find one input that crashes it.

Fuzzers like AFL and libFuzzer do exactly that, mutating inputs and using coverage feedback to steer toward unexplored code paths. It's falsification with a compass. Barton Miller's original 1990 fuzz study threw random garbage at common UNIX utilities, and a startling fraction crashed or hung. Every crash was a frog that lived.

For anything that eats untrusted bytes, such as parsers, protocol handlers, and deserializers, fuzzing is the cheapest falsification engine you can buy.

Debugging is falsification, too

A bug report is a falsifying observation. You don't need a theory of every way your system can fail. You need one reproducible failure.

That's why the debugging workflow looks the way it does. You reproduce the failure, minimize it (shrink the input, git bisect the commit), and turn it into a test. Hypothesis's shrinking is the same instinct, automated.

One caution: a failing test falsifies the conjunction of your code, your test, your environment, and your assumptions. Before you declare victory, check that the frog was really a frog. A flaky test, a bad fixture, or a wrong expectation can all produce a false alarm. Once a failure is reproducible, though, it's decisive. Don't argue with it and don't average it against the 4,000 tests that passed.


Why This Matters for Engineering Culture

There is a difference between asking, how can we prove this works and asking what would make our belief that this works wrong?

The second question is much more productive and drives results.

Test design. A confirmation mindset writes the happy path and a couple of edge cases, then stops when the checkmarks turn green. While, a falsification mindset asks what property must always hold and how to violate it.

Code review. Instead of "does this look right?", ask "what might break this?" Reviewers who hunt counterexamples find different bugs than reviewers who skim for plausibility.

Incident response. A production incident is a falsifying observation about your mental model of the system. The right response is to update the model, not to explain why the incident was unlucky. Then you add the counterexample to the suite so that frog can never surprise you twice.

Chaos engineering makes the connection explicit. Its method is to state a hypothesis about steady-state behavior, inject failure, and look for the experiment that contradicts you.


Practical Takeaways

  • Write tests that try to falsify invariants, not just confirm happy paths. Ask "what must always be true?" and then try to violate it. In the organization, I work in, we have made several tools that inserts inputs that break the system, make yours.
  • Use property-based tests for logic-heavy code (parsers, serializers, state machines, data transformations, anything with round-trip or ordering properties).
  • Fuzz anything that handles untrusted input: parsers, protocols, file formats, input validation.
  • Treat a green suite as evidence of absence of known counterexamples, not as proof of correctness. Say it that way in your head, and in your release notes.
  • When one test fails, treat it as decisive. Stop, reproduce, fix, and add the counterexample to the suite permanently.
  • In design reviews, ask: "What would falsify this design? What's the smallest input that breaks this assumption?"
  • Define your domain precisely. Vague claims ("handles all valid input") invite arguments about whether a counterexample counts.

Top comments (0)