DEV Community

Davi Orlandi
Davi Orlandi

Posted on

Go Test Coverage Is a Smoke Alarm, Not a Quality Certificate

I used to feel proud when a Go service crossed a high coverage number. Pull requests celebrated 87.3%. Gates refused merges below 80%. The bar looked meaningful. Then production taught me what the metric never promised: that executed lines are not the same as protected behavior.

The failure that stuck with me was not exotic. A repeated request with the same idempotency key created two payment attempts. The suite had touched the code. Coverage looked healthy. The business rule that mattered most had never been asserted. That is the personal stake behind this essay. Coverage can tell you which statements ran. It cannot tell you whether the tests understood the risk.

Deceptive green

Coverage tells you which statements ran. It does not tell you whether assertions were meaningful, whether races exist, or whether the API still matches the product.

Consider a tiny function with two independent conditions. One carefully chosen input can execute every line while leaving whole combinations of behavior untested. Statement coverage reports 100%. Path coverage would still be incomplete. Go's default tooling measures statements. That is useful. It is not a proof.

Worse: a test that calls a function and ignores the result still paints lines green. Happy-path suites that never touch err != nil branches look healthy until production finds the branch. So when someone asks for an "80% gate," ask a follow-up: 80% of what risk?

What coverage is good for

Coverage earns its keep as a flashlight. It finds untouched packages before a release. It spots dead error paths nobody exercises. It guides "what should I test next?" during a refactor.

go test ./... -coverprofile=cover.out
go tool cover -html=cover.out
Enter fullscreen mode Exit fullscreen mode

Open the HTML report. Hunt red in money, auth, and data-loss code. Ignore vanity in main wiring and generated protobufs. The value is not the percentage at the top. The value is which dangerous lines never ran.

What coverage lies about

Assertion-free tests are theater: execution without checks. Happy-path bias is fragile: high coverage with zero failure cases. Concurrency is invisible: coverage does not prove race freedom; -race does. Integration reality still bites: mocks can give 100% while SQL breaks in staging.

I have seen projects with 95% coverage that still ship production bugs in the untested 5% of error handling. I have also seen quieter suites near 40% that felt more trustworthy because every tested line mattered.

Table-driven tests beat a percentage

If you copy-paste cases, you probably want a table. Tables amortize careful assertions and make gaps obvious.

func TestParseDuration(t *testing.T) {
    tests := []struct {
        name    string
        in      string
        want    time.Duration
        wantErr bool
    }{
        {"seconds", "30s", 30 * time.Second, false},
        {"invalid", "abc", 0, true},
        {"empty", "", 0, true},
        {"negative", "-1s", -time.Second, false},
    }
    for _, tt := range tests {
        t.Run(tt.name, func(t *testing.T) {
            got, err := ParseDuration(tt.in)
            if tt.wantErr {
                if err == nil {
                    t.Fatalf("expected error for %q", tt.in)
                }
                return
            }
            if err != nil {
                t.Fatalf("unexpected err: %v", err)
            }
            if got != tt.want {
                t.Fatalf("got %v want %v", got, tt.want)
            }
        })
    }
}
Enter fullscreen mode Exit fullscreen mode

Use t.Run so failures name the case. Keep error-path tables separate when assertion logic diverges. Capture loop variables correctly if you mark subtests parallel, and do not share mutable state across parallel cases.

A workflow without the cult

Write behavior tests first, edges included. Implement until they pass. Run coverage and only add tests for surprising red in risky code. Stop when residual red is glue or unreachable defense-in-depth. Chasing the last five percent often couples tests to implementation details. Those tests punish refactors without catching bugs.

Recommendations I actually follow

Identify critical packages. Business logic yes. Logger setup and server wiring usually no (cover them in integration). Aim high on the testable core, but do not worship 100%. Prefer risk conversations over global gates. Soft floors on business packages beat a blanket 95% rule that invents empty tests. Do not treat "coverage must never decrease" as sacred. Sometimes deleting bad tests is progress. Unit coverage is never enough: integration, contract tests, fuzzing, and -race catch classes of failure a percentage will not mention.

test:
  script:
    - go test ./... -race -count=1
    - go test ./... -coverprofile=cover.out -covermode=atomic
    - go tool cover -func=cover.out | tee coverage.txt
Enter fullscreen mode Exit fullscreen mode

Fail builds for selected packages when coverage drops and the diff touches those packages. Blanket gates encourage people to delete hard-to-cover defenses or add hollow tests. Also treat flaky tests as defects. A suite that is green nine times out of ten trains the team to ignore red builds, which destroys more quality than a low number ever will.

Closing

Until critical code is covered, you definitely have untested risk. Once it is covered, you might still have untested risk. Treat coverage as a smoke alarm, not a certificate. Shine it on dangerous packages, invest in sharp table-driven cases, and keep race detection in CI. Good Go testing is not about making a report look impressive. It is about enough confidence that the important behaviors survive real traffic, repeated requests, and the strange timing problems that only appear after deploy.

Top comments (0)