It’s 03:00 PM on a Tuesday. Your monitoring dashboards go red.
Incoming requests are timing out, HTTP 504 errors are spiking, and users are reporting that the checkout screen hangs indefinitely.
You pull up your metrics:
- Application CPU: 7%
- PostgreSQL CPU: 4%
- JVM Heap: 30% used (GC pauses under 15ms)
- Database Deadlocks: 0
Your servers are practically idling. Yet the application is completely unresponsive.
You check the application logs and find hundreds of stack traces ending with this exact line:
org.springframework.transaction.CannotCreateTransactionException:
Could not open JPA EntityManager for transaction;
nested exception is org.hibernate.exception.JDBCConnectionException:
Unable to acquire JDBC Connection
...
Caused by: java.sql.SQLTransientConnectionException:
HikariPool-1 - Connection is not available, request timed out after 30000ms.
Your HikariCP connection pool is 100% saturated. Every single connection is checked out. New threads wait 30 seconds for a connection, time out, and die. Kubernetes health probes (/actuator/health) attempt to query the database, fail to acquire a connection, mark the pods as unready, and terminate them.
The restart cascade begins.
The culprit isn’t a slow SQL query or an unindexed table. It is an innocent-looking @Transactional annotation sitting on a method that makes a third-party HTTP call.
Table of Contents
- The Innocent Bug: An Autopsy
- What Happens Under the Hood
- The Math of Pool Exhaustion (Little’s Law)
- The Illusion of Distributed Atomicity
- The Fix: Decoupling Transactions from Network I/O
- Defensive Configuration: Catching Connection Leaks Before Production
- Summary Checklist
The Innocent Bug: An Autopsy
Here is the exact pattern that causes this outage. It exists in thousands of production Spring Boot codebases right now:
@Service
public class CheckoutService {
private final UserRepository userRepository;
private final OrderRepository orderRepository;
private final PaymentGatewayClient stripeClient; // Feign, RestClient, or WebClient
@Transactional
public OrderConfirmation processOrder(CheckoutRequest request) {
// Step 1: Read user from database (~2ms)
User user = userRepository.findById(request.userId())
.orElseThrow(() -> new UserNotFoundException(request.userId()));
// Step 2: Create initial order record (~3ms)
Order order = orderRepository.save(new Order(user, request.items()));
// Step 3: Call external payment provider over HTTPS (~1,800ms)
PaymentResult payment = stripeClient.charge(
user.getStripeCustomerId(),
request.amount()
);
// Step 4: Update order status (~2ms)
order.markPaid(payment.transactionId());
return new OrderConfirmation(order.getId(), "SUCCESS");
}
}
At first glance, this code looks clean. Developers add @Transactional because they want the entire checkout flow to be atomic: “If the payment fails, I want the order row rolled back.”
To understand why this destroys your application under load, you have to follow what happens at the JDBC connection layer.
What Happens Under the Hood
When Spring encounters @Transactional, it coordinates with JpaTransactionManager and your connection pool (HikariCP).
sequenceDiagram
autonumber
participant App as CheckoutService
participant Pool as HikariCP Pool (Size: 10)
participant DB as PostgreSQL
participant Stripe as External Payment API
App->>Pool: Borrow Connection
Pool-->>App: Connection #1 Checked Out
App->>DB: BEGIN TRANSACTION
App->>DB: SELECT * FROM users WHERE id = ...
App->>DB: INSERT INTO orders ...
Note over App,Stripe: Thread pauses in socketRead0 waiting for network packet!
App->>Stripe: POST https://api.stripe.com/v1/charges (Network I/O)
Note over Pool: Connection #1 sits 100% IDLE in DB, but HELD by thread for 1.8 seconds!
Stripe-->>App: HTTP 200 OK (After 1,800ms)
App->>DB: UPDATE orders SET status = 'PAID' ...
App->>DB: COMMIT TRANSACTION
App->>Pool: Return Connection #1
-
Step 1 (
userRepository.findById): Hibernate borrows a physicaljava.sql.Connectionfrom HikariCP and issuesBEGIN. -
Step 2 (
orderRepository.save): TheINSERTstatement executes. Total database time elapsed: 5 milliseconds. -
Step 3 (
stripeClient.charge): The thread halts its execution and waits for a TLS handshake, network routing, and Stripe’s server processing. Total time elapsed: 1,800 milliseconds. - Throughout those 1,800 milliseconds, the physical database connection remains locked to this thread. The database is doing zero work. The connection is held hostage simply because the Java thread hasn't exited the
@Transactionalmethod. -
Step 4: The
UPDATEexecutes,COMMITis sent, and the connection is finally returned to the pool.
Total transaction time: 1,807 milliseconds.
Time the database actually did work: 7 milliseconds.
Time spent holding the connection while waiting for network packets: 1,800 milliseconds (99.6%).
The Math of Pool Exhaustion (Little’s Law)
HikariCP’s default pool size is 10 connections.
Let’s calculate maximum system throughput under normal database operations versus when an external API is wrapped in a transaction:
Maximum Throughput = Pool Size / Average Connection Hold Time
Scenario A: Fast Database Transactions (No External I/O)
- Average connection hold time: 5ms
- Maximum throughput:
10 / 0.005s= 2,000 requests/second - Ten connections easily handle a high-traffic e-commerce store.
Scenario B: Database Transaction Wrapping an HTTP Call
- Average connection hold time: 1,500ms (1.5s)
- Maximum throughput:
10 / 1.5s= 6.6 requests/second
6.6 requests per second.
If just 7 concurrent users click "Checkout" at the same instant, all 10 HikariCP connections are occupied.
The 8th user’s request waits. HikariCP’s default connection-timeout is 30,000ms (30 seconds). While that 8th thread waits, subsequent requests pile up in Tomcat's thread pool. Within 10 seconds:
- Tomcat exhausts its worker threads (default 200).
- Incoming HTTP requests are rejected at the TCP socket layer.
- Actuator health checks timeout.
- Your monitoring tool alerts you that the entire service is down.
All for 7 concurrent checkout requests.
The Illusion of Distributed Atomicity
The most dangerous part of this pattern is that it does not even provide the safety developers think it does.
Consider this failure mode:
sequenceDiagram
autonumber
participant App as Spring Boot
participant DB as PostgreSQL
participant Stripe as External API
App->>DB: INSERT INTO orders (status = PENDING)
App->>Stripe: POST /charges ($100)
Stripe-->>App: 200 OK (Card Charged!)
Note over App,DB: Database network blip, disk full, or constraint violation!
App->>DB: COMMIT fails or Connection drops!
DB-->>App: ROLLBACK
Note over App,Stripe: Money taken from customer, but NO order in database!
- Spring opens a transaction.
- The order is inserted.
- Stripe charges the customer's credit card \$100.
- The database connection drops, a unique constraint fails, or the pod runs out of memory before
COMMITfinishes. - The local database transaction rolls back.
The order does not exist in your database. But Stripe successfully charged the customer's card.
A local database transaction (@Transactional) has zero control over an external HTTP service. It cannot roll back a network call that already reached another server. Wrapping third-party calls inside @Transactional gives you the illusion of distributed consistency while actively degrading your system availability.
The Fix: Decoupling Transactions from Network I/O
To eliminate this vulnerability, your database transactions must only span database operations. Network I/O must live outside transaction boundaries.
Here are the three production-proven patterns to accomplish this.
Pattern 1: Programmatic Boundaries with TransactionTemplate
Instead of annotating the entire method with @Transactional, use Spring's TransactionTemplate to create micro-transactions around only the database reads and writes.
@Service
public class SafeCheckoutService {
private static final Logger log = LoggerFactory.getLogger(SafeCheckoutService.class);
private final UserRepository userRepository;
private final OrderRepository orderRepository;
private final PaymentGatewayClient stripeClient;
private final TransactionTemplate transactionTemplate;
public SafeCheckoutService(UserRepository userRepository,
OrderRepository orderRepository,
PaymentGatewayClient stripeClient,
TransactionTemplate transactionTemplate) {
this.userRepository = userRepository;
this.orderRepository = orderRepository;
this.stripeClient = stripeClient;
this.transactionTemplate = transactionTemplate;
}
public OrderConfirmation processOrder(CheckoutRequest request) {
// Phase 1: Short local transaction to reserve and initialize (~5ms)
Order pendingOrder = transactionTemplate.execute(status -> {
User user = userRepository.findById(request.userId())
.orElseThrow(() -> new UserNotFoundException(request.userId()));
Order order = new Order(user, request.items());
order.setStatus(OrderStatus.PAYMENT_PENDING);
return orderRepository.save(order);
});
// Connection returned to HikariCP pool HERE!
// Phase 2: External HTTP call outside any database transaction (~1,500ms)
PaymentResult payment;
try {
payment = stripeClient.charge(
request.stripeCustomerId(),
request.amount(),
// Crucial: Use order ID as idempotency key for network retries
"checkout-order-" + pendingOrder.getId()
);
} catch (PaymentException e) {
log.error("Payment failed for order {}", pendingOrder.getId(), e);
markOrderFailed(pendingOrder.getId(), e.getMessage());
throw e;
}
// Phase 3: Short local transaction to finalize state (~4ms)
return transactionTemplate.execute(status -> {
Order order = orderRepository.findById(pendingOrder.getId()).orElseThrow();
order.markPaid(payment.transactionId());
orderRepository.save(order);
return new OrderConfirmation(order.getId(), "SUCCESS");
});
}
private void markOrderFailed(Long orderId, String reason) {
transactionTemplate.executeWithoutResult(status -> {
orderRepository.findById(orderId).ifPresent(order -> {
order.setStatus(OrderStatus.FAILED);
order.setFailureReason(reason);
orderRepository.save(order);
});
});
}
}
The Difference in Connection Usage
graph LR
subgraph SafeExecution["Safe Multi-Phase Execution"]
A["Phase 1: DB TX (~5ms)"] -->|Connection Released| B["Phase 2: Network I/O (~1,500ms)<br>ZERO DB Connections Held"]
B -->|Borrow Connection| C["Phase 3: DB TX (~4ms)"]
end
During the 1,500ms HTTP call to Stripe, zero database connections are held. Your pool of 10 connections remains completely open to serve other users, health checks, and read queries.
Pattern 2: Domain Events & @TransactionalEventListener
If the third-party call is a non-blocking side-effect (such as sending an email via SendGrid, dispatching a webhook, or emitting an audit log), use Spring's application events.
By setting the listener phase to AFTER_COMMIT, Spring guarantees that:
- The external API is called only after the local database transaction has safely committed.
- The database connection has already been returned to the Hikari pool before the network call starts.
The Event Definition
public record OrderPlacedEvent(Long orderId, String customerEmail, BigDecimal totalAmount) {}
The Transactional Service
@Service
public class OrderPlacementService {
private final OrderRepository orderRepository;
private final ApplicationEventPublisher eventPublisher;
public OrderPlacementService(OrderRepository orderRepository,
ApplicationEventPublisher eventPublisher) {
this.orderRepository = orderRepository;
this.eventPublisher = eventPublisher;
}
@Transactional
public Long placeOrder(CreateOrderCommand cmd) {
Order order = orderRepository.save(new Order(cmd));
// Publish event inside transaction
eventPublisher.publishEvent(new OrderPlacedEvent(
order.getId(),
cmd.customerEmail(),
cmd.totalAmount()
));
return order.getId();
} // Connection is COMMITTED and RELEASED here!
}
The Event Listener
@Component
public class OrderNotificationListener {
private static final Logger log = LoggerFactory.getLogger(OrderNotificationListener.class);
private final EmailClient emailClient; // External HTTP Client (SendGrid, Postmark)
public OrderNotificationListener(EmailClient emailClient) {
this.emailClient = emailClient;
}
@Async // Run in background thread pool to avoid blocking HTTP request
@TransactionalEventListener(phase = TransactionPhase.AFTER_COMMIT)
public void onOrderPlaced(OrderPlacedEvent event) {
log.info("Sending order confirmation email for order {}", event.orderId());
// This HTTP call runs with NO database transaction or connection attached!
emailClient.sendOrderConfirmation(event.customerEmail(), event.orderId());
}
}
If the database transaction fails and rolls back, onOrderPlaced will never execute. No customer will ever receive a "Your order is confirmed!" email for an order that failed to save.
Pattern 3: The Transactional Outbox Pattern for Mission-Critical Integrations
When an external integration must complete reliably (such as syncing with SAP, Salesforce, or dispatching an event to Kafka), use the Transactional Outbox pattern.
Instead of calling the external API synchronously:
- Write the business entity and an
outboxrecord in the same local transaction (takes < 5ms). - A separate background worker reads unprocessed outbox records and delivers them to the external API with retries and exponential backoff.
graph TD
subgraph ClientRequest["Client Request (Fast)"]
A["User Request"] --> B["Save Order + Outbox Row in 1 TX (~5ms)"]
B --> C["Return 200 OK to User immediately"]
end
subgraph BackgroundWorker["Background Worker (Resilient)"]
D[("Outbox Table")] -->|Poll / Listen| E["Outbox Poller Thread"]
E -->|Call External API| F["Stripe / Salesforce / SAP"]
F -->|200 OK| G["Mark Outbox Row as PROCESSED"]
end
This insulates your customer-facing response times from third-party outages and guarantees at-least-once delivery.
Defensive Configuration: Catching Connection Leaks Before Production
Even with careful code reviews, an unvetted library or rogue @Transactional can slip through. Add these three guardrails to your Spring Boot configuration:
1. Enable HikariCP Leak Detection
HikariCP has a built-in leak detection mechanism that logs a stack trace whenever a thread holds a connection longer than a set threshold:
# application.properties
# Log a warning if a thread holds a connection longer than 3 seconds
spring.datasource.hikari.leak-detection-threshold=3000
When a thread holds a connection for longer than 3,000ms, HikariCP outputs a warning with the full stack trace:
WARN 12940 --- [pool-thread-2] com.zaxxer.hikari.pool.ProxyLeakTask:
Connection leak detection triggered for org.postgresql.jdbc.PgConnection@4a2f8d on thread http-nio-8080-exec-4,
stack trace follows:
at com.example.service.CheckoutService.processOrder(CheckoutService.java:42)
You can pinpoint the exact line of code holding the connection.
2. Configure Strict HTTP Client Timeouts
Never use default timeouts on your HTTP clients (RestClient, WebClient, Feign, or Apache HttpClient). Default read timeouts are often infinite or 60+ seconds.
@Configuration
public class HttpClientConfig {
@Bean
public RestClient stripeRestClient() {
SimpleClientHttpRequestFactory factory = new SimpleClientHttpRequestFactory();
factory.setConnectTimeout(Duration.ofSeconds(2));
factory.setReadTimeout(Duration.ofSeconds(4)); // Fail fast if external server hangs
return RestClient.builder()
.requestFactory(factory)
.baseUrl("https://api.stripe.com")
.build();
}
}
If Stripe degrades, your client fails in 4 seconds instead of hanging for minutes and starving your thread pools.
3. Delay Connection Acquisition with Hibernate
By default, Hibernate might acquire a physical database connection as soon as a transaction begins. You can instruct Hibernate to delay physical connection acquisition until the first actual SQL query is issued:
spring.jpa.properties.hibernate.connection.provider_disables_autocommit=true
This prevents early connection checkout if your service performs non-database validation logic before its first SQL query.
Summary Checklist
| Anti-Pattern | Recommended Solution |
|---|---|
External HTTP call inside @Transactional
|
Use TransactionTemplate to split into short pre- and post-transactions |
Email, notifications, or metrics inside @Transactional
|
Use @TransactionalEventListener(phase = AFTER_COMMIT)
|
| Long-running batch or external sync | Transactional Outbox pattern with background worker |
| Undetected connection hogs | Set spring.datasource.hikari.leak-detection-threshold=3000
|
| Hanging HTTP calls | Set explicit 2–4s connect and read timeouts on all HTTP clients |
A database connection is one of the most expensive and scarce resources in your entire architecture.
Treat it like an in-memory lock: acquire it at the last possible millisecond, do your SQL work, and release it immediately. Keep your network calls out of your database transactions, and your connection pool will effortlessly handle traffic surges without breaking a sweat.
Top comments (1)
This is a useful breakdown, especially the distinction between local transaction rollback and an external side effect that already succeeded. Two caveats worth calling out: seven concurrent checkouts can’t exhaust a pool of ten by themselves, and @TransactionalEventListener(AFTER_COMMIT) isn’t durable—if the process dies after commit but before the async listener runs, the event can be lost. That’s where the outbox pattern earns its keep. One configuration detail also deserves care: provider_disables_autocommit is an assertion about the connection provider, not a generic “delay connection acquisition” switch.