DEV Community

Divyansh Kumar
Divyansh Kumar

Posted on

Why Your Testcontainers Check Is Flaky: The Quarkus Port 9000 Trap

We've all seen integration tests that fail randomly in CI. You run them on your laptop five times in a row, and they pass cleanly every single time without a hiccup. Then you push your commit, and GitHub Actions blows up during initialization with a completely unexpected connection drop. You hit retry. It goes green. It's deeply frustrating.

Drafted with AI assistance from my notes and code, then reviewed and edited by me.

I ran into this exact headache while debugging JsonSchemaRecordValidationIT in the Kroxylicious open source project. Our test suite spins up an Apicurio Registry container via Testcontainers, registers a JSON schema in @BeforeAll, and runs record validation against Kafka.

Locally, my test passed instantly. In CI, builds broke intermittently during initialization. The socket died.

java.lang.RuntimeException: could not close the reader
    at io.kiota.serialization.json.JsonParseNodeFactory.getParseNode(...)
    at io.apicurio.registry.rest.client.groups.item.artifacts.ArtifactsRequestBuilder.post(...)
    at io.kroxylicious.it.filter.validation.JsonSchemaRecordValidationIT.init(...)
Caused by: java.io.IOException: closed
Caused by: java.io.IOException: fixed content-length: 429, bytes received: 0
Caused by: java.io.EOFException: EOF reached while reading
Enter fullscreen mode Exit fullscreen mode

The error felt bizarre. Kiota expected 429 bytes back from the registry, but received zero bytes before the socket slammed shut. Why would an HTTP server drop an incoming connection right after booting?

The Deceptive Wait Strategy

I looked at RecordSchemaValidationBaseIT. Our test used a standard HTTP wait strategy on port 8080:

protected static GenericContainer<?> startRegistryContainer() {
    return new GenericContainer<>(dockerImageName)
            .withExposedPorts(8080)
            .waitingFor(Wait.forHttp("/apis/registry/v3/system/info").forStatusCode(200));
}
Enter fullscreen mode Exit fullscreen mode

On paper, that looks completely sound. Testcontainers pings /apis/registry/v3/system/info until it gets an HTTP 200 response, flags the container as healthy, and hands off control to JUnit. Our code then immediately calls client.groups().byGroupId("default").artifacts().post(...).

If the endpoint returns 200, why did our requests crash?

The Startup Race in Quarkus

Here's what I discovered under the hood. Apicurio Registry 3 runs on Quarkus. When Quarkus boots, Netty binds port 8080 right away. Any endpoint that serves static build metadata straight from memory, like /system/info, starts returning 200 in a fraction of a second.

Meanwhile, the actual database engine is still initializing on a background thread:

INFO [io.quarkus.bootstrap.runner.Timing] Listening on: http://0.0.0.0:8080. Management interface listening on http://0.0.0.0:9000.
INFO [io.apicurio.registry.storage.impl.sql.AbstractSqlRegistryStorage] Acquiring database initialization lock...
INFO [io.apicurio.registry.storage.impl.sql.AbstractSqlRegistryStorage] Database not initialized.
INFO [io.apicurio.registry.storage.impl.sql.AbstractSqlRegistryStorage] Initializing the Apicurio Registry database...
INFO [io.apicurio.registry.storage.impl.sql.AbstractSqlRegistryStorage] Database initialization lock released.
Enter fullscreen mode Exit fullscreen mode

Port 8080 accepts traffic while AbstractSqlRegistryStorage is still running schema migrations on H2.

On my workstation, the database was ready before Testcontainers even sent its first ping. But on a heavily loaded CI runner with constrained CPU, the timing inverted. Testcontainers saw /system/info return 200, thought everything was ready, and triggered our test. Because the HTTP server bound port 8080 almost immediately upon process launch, polling a static system information endpoint gave Testcontainers the false impression that the entire containerized application was completely healthy and prepared to process incoming write operations. Whenever our test runner tried to connect while the internal schema tables were still locked by the migration worker, the Quarkus runtime simply reset the underlying TCP connection without transmitting any bytes back to the waiting Kiota client. That produced the zero-byte EOFException.

Keith Wall, a maintainer on the repo, caught this in issue #5031: "We currently probe /apis/registry/v3/system/info as an indication of readiness. I wonder if this is sufficient?"

It wasn't. A metadata endpoint isn't a readiness probe.

The Fix: Quarkus Management Port 9000

In Quarkus and SmallRye Health, readiness checks don't live on the application port. In Apicurio Registry 3.x, they live on port 9000.

The /health/ready probe on port 9000 queries the Agroal datasource and storage subsystem directly. If database locks or migrations are running, it returns HTTP 503. It returns 200 only when the storage engine can safely accept writes.

I updated RecordSchemaValidationBaseIT to expose port 9000 and pointed the wait strategy to /health/ready:

protected static final int CONTAINER_PORT = 8080;
protected static final int MANAGEMENT_PORT = 9000;

protected static GenericContainer<?> startRegistryContainer() {
    return new GenericContainer<>(dockerImageName)
            .withExposedPorts(CONTAINER_PORT, MANAGEMENT_PORT)
            .waitingFor(Wait.forHttp("/health/ready").forPort(MANAGEMENT_PORT).forStatusCode(200));
}
Enter fullscreen mode Exit fullscreen mode

Our client still talks to port 8080, but Testcontainers won't let our tests run until port 9000 confirms the database is completely ready.

The flakiness disappeared completely. It worked cleanly.

What to Remember

Don't use static metadata endpoints for container readiness. If an app boots its web server before its database or cache is ready, polling an info URL creates a race condition. Always find the framework's dedicated readiness probe, and check if it lives on a separate management port.

Top comments (0)