DEV Community

gusdmoreira
gusdmoreira

Posted on

Beyond the Green Test: A Pragmatic Framework for Evaluating AI Code Quality

You prompt your favorite AI coding assistant: "Write a service that processes bulk user telemetry, validates incoming records, and handles batch publishing."

Within seconds, pristine code streams across your screen. The structure looks clean, the syntax leverages modern language features, and it even includes neat documentation comments. You drop it into your repository, run your test suite, and watch the console turn bright green.

Ship it, right?

Not so fast.

While tools like GitHub Copilot, Claude, and specialized coding agents are incredible velocity boosters, they operate on probability, not deep contextual understanding. They excel at the happy path but frequently gloss over the messy realities of production systems: concurrency nuances, resource management under heavy load, security implications, and subtle domain edge cases.

If you blindly merge AI-generated code without a systematic evaluation framework, you are essentially outsourcing your technical debt to a non-deterministic generator.

Let's look at a practical, robust approach to auditing and validating AI-generated code, using Java for our concrete examples.

1. Shift-Left Validation: Define the Contract via TDD

The biggest trap developers fall into is letting the AI write both the implementation and the expectations. If the model hallucinates a requirement or misunderstands a business rule, its generated unit tests will happily validate its own flawed logic.

Instead, practice AI-Driven TDD: Write the test contract yourself before invoking the assistant.

Imagine you need a thread-safe rate limiter service. Before asking the AI to write the class, write your strict unit test suite covering concurrent threads, edge limits, and time windows. For instance, in Java using JUnit 5 and AssertJ:

package com.example.ratelimiter;

import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.time.Duration;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.atomic.AtomicInteger;

import static org.assertj.core.api.Assertions.assertThat;

class RateLimiterServiceTest {

    @Test
    @DisplayName("Should allow requests within limit and throttle excess under concurrency")
    void testConcurrentRateLimiting() throws InterruptedException {
        RateLimiterService limiter = new RateLimiterService(5, Duration.ofSeconds(1));
        int totalThreads = 20;
        ExecutorService executor = Executors.newFixedThreadPool(totalThreads);
        CountDownLatch latch = new CountDownLatch(totalThreads);
        AtomicInteger allowedCount = new AtomicInteger(0);

        for (int i = 0; i < totalThreads; i++) {
            executor.submit(() -> {
                try {
                    if (limiter.tryAcquire("user-123")) {
                        allowedCount.incrementAndGet();
                    }
                } finally {
                    latch.countDown();
                }
            });
        }

        latch.await();
        executor.shutdown();

        // Exactly 5 should pass out of 20 concurrent attempts
        assertThat(allowedCount.get()).isEqualTo(5);
    }
}
Enter fullscreen mode Exit fullscreen mode

The Danger of the "Fake Contract" (Make Your Test Bleed First)
Writing the contract yourself prevents the AI from validating its own flaws, but it opens the door to a quieter issue: you can write a test that passes for the wrong reasons.

For example, if you ask an AI for a deduplication function and write an assertion checking only that the output is a subset of the original elements, that test will happily pass even if the model returns the list completely untouched with duplicates intact. A test that has never failed is not a guard—it is a comment that runs.

The Golden Rule: Before you trust your test suite, break the implementation on purpose (return a dummy value or an empty list) and watch it go red. Pre-fix red is the only real evidence that a test is actually capable of failing. If a known-bad build still passes your test, your contract is an illusion.

By establishing this rigorous contract upfront, you give the AI a rigid mathematical boundary. When it returns the implementation, your test suite acts as an objective security guard.

2. Curate a Domain-Specific "Golden Dataset"

For complex data transformations, parsers, or validation rules, unit tests aren't always enough to catch structural discrepancies. This is where you build a Golden Dataset, a version-controlled reference collection containing edge cases specific to your domain.

You can maintain this as a resource file (like JSON) inside your project structure:

[
  { "inputUserId": "USR-001", "payload": "{\"tier\": \"PREMIUM\", \"amount\": 1500.00}", "expectedStatus": "ACCEPTED" },
  { "inputUserId": "USR-002", "payload": "MALFORMED_JSON_STRING", "expectedStatus": "REJECTED_MALFORMED" },
  { "inputUserId": "USR-999", "payload": "{\"tier\": \"EXPIRED\", \"amount\": 50.00}", "expectedStatus": "REJECTED_UNAUTHORIZED" },
  { "inputUserId": null, "payload": "{\"tier\": \"BASIC\", \"amount\": 10.00}", "expectedStatus": "REJECTED_NULL_ID" }
]
Enter fullscreen mode Exit fullscreen mode

Your automated evaluation test reads this dataset, passes every payload through the AI-generated service, and asserts the outcomes. When you refactor prompts or upgrade your model version, running this dataset guarantees zero regressions.

Watch Out for Rejection Bias: When building your golden dataset, it is easy to overload it with edge cases that must be rejected (nulls, empty strings, malformed syntax). If your dataset is heavily skewed toward rejections, you risk driving the AI validator toward the worst possible convergence: a service that rejects everything. A system that says "no" to every input has zero false positives on rejections, but it is completely useless. Ensure that the cases which must be accepted carry just as much weight and diversity as the ones that fail.

3. Beyond Correctness: The Production Audit Checklist

Once the code passes your functional tests, you must conduct a targeted code review focusing on dimensions that automated tests often miss:

  • Resource Management & Memory: Did the model properly handle connection closures, stream disposal, or memory lifecycle hooks? Look out for unclosed handles or collections accumulating state indefinitely.
  • Concurrency & Thread Safety: LLMs frequently default to non-thread-safe data structures in multi-threaded contexts. Ensure safe alternatives or proper synchronization mechanisms are in place.
  • Exception Hygiene: Does the code swallow broad exceptions or suppress stack traces? Verify that errors throw meaningful, domain-specific exceptions with appropriate logging context.

4. Treat Prompts as Architecture, Not Chat History

If your team relies heavily on AI workflows, automated prompt generation, or contextual system instructions, stop treating prompts like casual chat messages.

  • Version Control: Store your prompt templates in dedicated version-controlled directories rather than hardcoding them inline.
  • Consistency Metrics: Measure prompt reliability across multiple iterations before deploying changes that affect core business pipelines.

Conclusion

The mark of a senior engineer in the age of AI isn't how fast they can copy-paste generated code; it's how rigorously they can evaluate it.

Treat AI-generated code with the exact same professional skepticism you would apply to code submitted by an external contractor. Write your tests first, audit for real-world production constraints, and remember: The best developers using AI are not the ones who trust it blindly, but the ones who verify it with absolute precision.

Top comments (2)

Collapse
 
pm25coder profile image
pm25coder •

This is the right frame, and §1 is the part I'd defend hardest — but there's one step missing, and it's the step that catches everything else you've written.

"Write the contract yourself before invoking the assistant" fixes the AI-validates-its-own-flaw problem. It does not fix a quieter one: a contract you wrote yourself can still pass on the broken build. Whether a test can fail is a property of the assertion, not of who wrote it, and green tells you nothing about it.

A small measured version. Take a dedupe function the model got wrong (it returns the list unchanged):

out = dedupe_preserving_order([1, 1, 2, 3, 3])

assert set(out).issubset({1, 2, 3})   # passes on the BROKEN build
assert len(out) == 3                  # fails on the BROKEN build
Enter fullscreen mode Exit fullscreen mode

Both assertions were written by a human, against a real contract, from a real requirement. The subset assertion is satisfied by a fact other than the one it names — "the output contains only known values" is true whether or not you deduplicated. It has never once failed, and a test that has never failed is a test you haven't met. The subset check is documentation; only the length check is a guard.

So the step I'd add to §1, before you trust any of the suite:

Run every new test against the deliberately-broken code and watch it go red. The pre-fix red is the only evidence a test is capable of failing. If you can't make it fail on a known-bad implementation, you don't have a contract — you have a comment that runs.

§3's exception-hygiene line is the highest-value item on the checklist, and I'd adjust where it points. "Does the code swallow broad exceptions?" usually gets audited inside the function, but the swallow lives one layer up, at the caller that maps every failure to one verdict:

probe()           -> raises TimeoutError
caller_broad()    -> 'gone'                      # except Exception: return "gone"
caller_typed()    -> 'inconclusive(TimeoutError)'
Enter fullscreen mode Exit fullscreen mode

The broad except here turns "I could not decide" into a clean shutdown — an inconclusive measurement re-read as a confident one, which is a worse class of bug than a crash, because nothing is red. And note a test that only asserts "an exception was raised" passes on both callers, so it can't see the regression. Only an assertion naming the outcome separates them.

Which is the same disease as §1, one level up: a check that holds for both the right and the wrong answer is not evidence. The production audit checklist is four checks; the missing fifth is the one that asks, of each of the other four, "how would I make this fail?"

One thing I'd add to §2 while you're there: out of a golden dataset, the entries that matter most are the ones that must be accepted. A dataset that is all rejections can be satisfied by a service that rejects everything — and a service that rejects everything is exactly the failure mode a fail-closed AI validator converges on.

Collapse
 
gusdmoreira profile image
gusdmoreira •

Good catch on both points.

You're completely right about the test that never fails; calling it "a comment that runs" is spot on. If you can't make a test go red against a broken build, it's not actually guarding anything.

The point about rejection-heavy datasets is also a great warning. A service that just rejects everything will look great on paper until it silently drops valid traffic.

Appreciate the feedback; I'll be folding these nuances in.