I've spent the last three-plus years building and operating the microservices behind AirAsia Move — the travel super-app that handles flights, hotels, and ancillaries for millions of travelers across Southeast Asia. The system processes over 130 million requests per day, roughly 1,500 requests per second sustained, with spikes well above that during flash sales.
I own services on the Manage My Booking platform: flight changes, fare summaries, price-slash features, post-booking ancillaries. A single user action like changing a flight can fan out to eight or nine downstream calls. At this scale, the engineering problems aren't about whether your code compiles. They're about what happens when one of those downstream services is 200ms slower than usual.
Circuit Breakers with Resilience4j
The pattern that proved its weight in gold was the circuit breaker. We use Resilience4j across all Spring Boot services — if a downstream service starts failing, stop calling it instead of piling up timeouts that cascade through the system.
@CircuitBreaker(name = "pricingService", fallbackMethod = "cachedFareFallback")
public FareSummary getFareSummary(String bookingId) {
return pricingClient.fetchFare(bookingId);
}
private FareSummary cachedFareFallback(String bookingId, Throwable t) {
log.warn("Pricing service unavailable for booking {}, using cache", bookingId);
return fareCache.getLastKnown(bookingId)
.map(fare -> fare.withStaleFlag(true))
.orElseThrow(() -> new ServiceUnavailableException("No cached fare available"));
}
We pair breakers with fallbacks. If the pricing service is down, the user still sees a fare — it might be a few minutes stale, but that beats a blank screen or a 500.
We tuned thresholds through load testing with JMeter. The defaults were too aggressive and tripped breakers during normal latency variance. A 50% failure rate over a sliding window of 20 calls, with a 30-second open-state wait, matched our actual traffic pattern.
Distributed Tracing
When a request crosses ten services, debugging without distributed tracing is a lost cause. We use Spring Cloud Sleuth for trace propagation and Zipkin for visualization. We sample 10% of traces in production — at 1,500 RPS, that's 150 traces per second, more than enough to spot patterns.
The biggest tracing win wasn't debugging individual slow requests. We noticed every Tuesday between 2–4 AM UTC, latency spiked on the inventory service. A batch job was running full table scans on the same database the API was reading from. Moved the job to a read replica, spikes gone.
The Cascading Timeout
During a Diwali sale, the payment gateway started responding 300ms slower than usual. Not enough to fail — just enough to back up our thread pool. We had a fixed pool of 200 threads with a 5-second payment timeout and a bounded queue of 500. When every thread was blocked on slow payment calls, the queue filled. Upstream services calling us started timing out. Within 90 seconds, three services were effectively down.
The root cause wasn't the payment gateway. It was our thread pool and timeout configuration.
We made three changes:
- Cut the payment timeout from 5s to 2s. If it hasn't responded in 2 seconds during peak load, it's not going to.
- Added a bulkhead — isolated the payment call to its own thread pool so one slow dependency can't starve everything else.
- Switched the circuit breaker to a time-based sliding window so it reacts faster during traffic spikes.
The Memory Leak
Our flight-change service was getting OOMKilled by Kubernetes every 48 hours. Two days of profiling with JVisualVM found a ConcurrentHashMap used as an in-memory cache with no eviction. Every unique booking ID got cached and never expired. Under sustained 1,500 RPS, it grew until the JVM ran out of heap.
Five lines of code, two days of investigation:
Cache<String, FareSnapshot> fareCache = Caffeine.newBuilder()
.maximumSize(50_000)
.expireAfterWrite(Duration.ofMinutes(10))
.build();
What I'd Do Differently
Contract testing from day one. When Service A renames fare_amount to fareAmount, both services' unit tests pass. Staging explodes. Consumer-driven contract tests catch this at build time.
Structured logging earlier. Searching Kibana for a booking ID across 10 services when half log bookingId=ABC123 and the other half log Processing booking ABC123 for user xyz is painful. Consistent field names should be a service template requirement on day one.
More aggressive load shedding. At 90% capacity, return 429s for low-priority endpoints and keep booking and payment paths fully served, instead of letting everything degrade equally.
The Takeaway
Building microservices at this scale is about understanding failure modes. Every pattern — circuit breakers, tracing, async messaging, bulkheads — exists because something broke in production and we needed it to not break the same way again. The system processes 130 million requests a day not because we got the design right on the first try, but because we built the instrumentation to see what was breaking and the patterns to contain the blast radius.
Originally published on krishnakky.com — the full version has additional code examples, Kafka event-driven architecture details, and a third production war story about consumer group rebalances.
Top comments (0)