<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Krishna Yadav</title>
    <description>The latest articles on DEV Community by Krishna Yadav (@krishna_y).</description>
    <link>https://dev.to/krishna_y</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4164464%2F559c22a4-8527-4952-830b-64937fa41b56.png</url>
      <title>DEV Community: Krishna Yadav</title>
      <link>https://dev.to/krishna_y</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/krishna_y"/>
    <language>en</language>
    <item>
      <title>From Startup to Acquisition: Lessons from Building and Selling an eCommerce Platform</title>
      <dc:creator>Krishna Yadav</dc:creator>
      <pubDate>Mon, 05 Oct 2026 17:02:26 +0000</pubDate>
      <link>https://dev.to/krishna_y/from-startup-to-acquisition-lessons-from-building-and-selling-an-ecommerce-platform-40fe</link>
      <guid>https://dev.to/krishna_y/from-startup-to-acquisition-lessons-from-building-and-selling-an-ecommerce-platform-40fe</guid>
      <description>&lt;p&gt;In February 2014 I was twenty-two, halfway through an M.Tech, and convinced I could build a marketplace that would change how local vendors sold online. Three years later the platform was serving eighty-plus vendors, processing real money, and I signed the papers that handed it to someone else.&lt;/p&gt;

&lt;p&gt;This is the short version of that story, and the seven things it taught me.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Setup
&lt;/h2&gt;

&lt;p&gt;Small vendors in our region had no affordable way to sell online. Big marketplaces charged steep commissions and buried small sellers. We wanted to build a multi-vendor platform where each shop got its own storefront with shared logistics, payments, and traffic.&lt;/p&gt;

&lt;p&gt;My co-founder handled business. I handled everything that touched a terminal. We bootstrapped — no angel round, no incubator. Every hour of engineering time came directly out of the hours I was supposed to spend on coursework.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building on a Twelve-Dollar VPS
&lt;/h2&gt;

&lt;p&gt;I picked Java and Spring for the backend, MySQL for persistence, jQuery for the frontend. Not trendy, but I trusted it wouldn't collapse at midnight when I was the only person on call.&lt;/p&gt;

&lt;p&gt;The MVP took ten weeks of nights and weekends. One Spring monolith, one database, one Tomcat server on a VPS that cost twelve dollars a month. It was ugly. It worked.&lt;/p&gt;

&lt;p&gt;The first real challenge was payment integration. Two Indian payment gateways, each with its own callback format, retry logic, and definition of "success." I wrote an adapter layer that normalized both into common internal events — &lt;code&gt;PAYMENT_CAPTURED&lt;/code&gt;, &lt;code&gt;PAYMENT_FAILED&lt;/code&gt;, &lt;code&gt;PAYMENT_REFUND_INITIATED&lt;/code&gt; — so the rest of the order pipeline didn't care which gateway handled a transaction. That pattern ended up being the most reusable code in the entire system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scaling and Cracking
&lt;/h2&gt;

&lt;p&gt;Each new vendor exposed a new edge case. Weight-based items. Variant pricing. Bundle deals. Every feature request was a negotiation between "this helps one vendor" and "this complicates things for everyone."&lt;/p&gt;

&lt;p&gt;Around vendor forty, the monolith started cracking. Page loads crept past three seconds. I extracted the inventory service into its own process with its own database. Not a full microservices migration — one pragmatic extraction that relieved the biggest bottleneck.&lt;/p&gt;

&lt;p&gt;Meanwhile I was attending lectures from 9 AM to 1 PM, building the platform from 2 PM to midnight, and squeezing thesis research into whatever gaps remained. The saving grace: every distributed systems concept from the M.Tech curriculum — consistency models, fault tolerance, CAP theorem — I was seeing in production the same week I read about them in papers. The gap between theory and practice wasn't a gap. It was the same Tuesday.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Exit
&lt;/h2&gt;

&lt;p&gt;By early 2017 the platform was profitable but growth had plateaued. A regional eCommerce company approached us. They wanted our vendor network and our technology — specifically the vendor management system and the payment adapter layer, which they said would save them six to eight months of development.&lt;/p&gt;

&lt;p&gt;The negotiation took two months. We closed in May 2017. The terms were fair. The vendors kept their storefronts, the customers kept their accounts. I walked out with no regrets and a clear picture of what I wanted next: building at larger scale, which eventually led me to Oracle Financial Services, Kotak Securities, and AirAsia where I now work on systems processing 130 million daily transactions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Seven Lessons for Technical Founders
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Ship the ugly version first.&lt;/strong&gt; Our MVP was architecturally embarrassing. It was also live, collecting payments, and teaching us what users actually needed. The beautiful rewrite can happen after you've proven the business.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Pick boring technology.&lt;/strong&gt; Java and Spring in 2014 won zero hackathons. But when our payment integration broke at midnight, I could find a Stack Overflow answer in under a minute. Boring technology compounds reliability.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Solve the class, not the instance.&lt;/strong&gt; When vendor thirty-seven asks for a feature, ask what general problem it represents. Build the general solution or don't build it at all. Every special case is a maintenance liability that outlives whoever requested it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Wrap every external dependency in an adapter.&lt;/strong&gt; Payment gateways, shipping APIs, SMS providers — they will change their interfaces, go down, or get replaced. Your codebase should never know or care which vendor sits behind the adapter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Cofounder alignment beats cofounder skills.&lt;/strong&gt; We never argued about direction because we agreed on the fundamentals before we started: bootstrap, stay profitable, build something vendors actually use. Technical skills can be hired. Strategic alignment cannot.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Know your exit conditions before you need them.&lt;/strong&gt; We didn't have a formal trigger for when to sell, and that made the decision more emotional than necessary. Write down the conditions under which you'd sell, shut down, or raise money — before you launch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;7. Burnout is not a badge of honor.&lt;/strong&gt; I pushed through it because I was young and stubborn. One full day off per week would not have killed the company and would have made me a better engineer the other six days.&lt;/p&gt;

&lt;p&gt;That startup gave me scars and skills in roughly equal proportion. If you're building something right now — side project, startup, internal tool at a company that doesn't appreciate it — the work compounds. The code you're embarrassed by today teaches you to write something better tomorrow.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://krishnakky.com/blog/startup-to-acquisition-lessons" rel="noopener noreferrer"&gt;krishnakky.com&lt;/a&gt; — the full version has more detail on the technical architecture, payment integration patterns, and the acquisition negotiation.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>startup</category>
      <category>entrepreneurship</category>
      <category>career</category>
      <category>webdev</category>
    </item>
    <item>
      <title>AI-Driven Development: Running Autonomous Agents Across a Microservices Codebase</title>
      <dc:creator>Krishna Yadav</dc:creator>
      <pubDate>Mon, 05 Oct 2026 16:57:07 +0000</pubDate>
      <link>https://dev.to/krishna_y/ai-driven-development-running-autonomous-agents-across-a-microservices-codebase-1g48</link>
      <guid>https://dev.to/krishna_y/ai-driven-development-running-autonomous-agents-across-a-microservices-codebase-1g48</guid>
      <description>&lt;p&gt;At AirAsia, I own production microservices on the Manage My Booking platform — flight changes, fare summaries, price-slash, post-booking ancillaries. Over a dozen services, each with its own repo, pipeline, and accumulated quirks. The platform processes 130 million requests a day.&lt;/p&gt;

&lt;p&gt;My problem was never writing code. It was the context-switching overhead. Bouncing between a fare-calculation refactor, a circuit breaker fix three repos over, and a schema migration for the ancillary service — every switch cost 20–30 minutes of mental reload. Entire days evaporated into friction.&lt;/p&gt;

&lt;p&gt;I started using &lt;a href="https://docs.anthropic.com/en/docs/claude-code/overview" rel="noopener noreferrer"&gt;Claude Code&lt;/a&gt; not to write code for me, but to keep multiple workstreams moving in parallel. It operates directly in the terminal — reads the codebase, runs commands, writes files, stages changes. An agent, not an autocomplete engine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Permission Scopes: The Non-Negotiable First Step
&lt;/h2&gt;

&lt;p&gt;You cannot hand an AI agent unrestricted access to a production codebase. That is how you get a force-push to main at 2 AM.&lt;/p&gt;

&lt;p&gt;I set up scoped permissions per project in &lt;code&gt;.claude/settings.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"permissions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"allow"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Edit"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Write"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(npm run test*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(npm run lint*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(mvn test*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git diff*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git status*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git log*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git add -p*)"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"deny"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git push*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git checkout main*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(rm -rf*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(docker*)"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent can read, write, edit, run tests and linters, view diffs, and stage changes. It cannot push, switch to main, delete directory trees, or touch Docker. Every push and merge goes through me. Payment-adjacent services get tighter restrictions. Internal tooling repos are looser.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Three-Phase Model
&lt;/h2&gt;

&lt;p&gt;I restructured feature delivery into three phases:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inception (human only).&lt;/strong&gt; I define the scope, draw sequence diagrams, identify affected services, and write interface contracts. This produces a spec document — the agent's input. Thirty to sixty minutes here, and the quality directly determines how useful the agent is downstream.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Construction (agents + review).&lt;/strong&gt; Each agent gets a spec, a repo, and constraints. I run parallel terminal sessions — one agent per service:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Terminal 1:&lt;/strong&gt; Agent on booking-service, implementing the new endpoint&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Terminal 2:&lt;/strong&gt; Agent on pricing-engine, adding fare calculation logic&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Terminal 3:&lt;/strong&gt; Agent on ancillary-catalog, updating schema and data layer&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Me:&lt;/strong&gt; Reviewing diffs, running integration tests, handling cross-service architecture decisions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agents don't talk to each other. I am the coordination layer. When the pricing agent finishes, I review it and feed interface changes to the booking agent as context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Operations (human only).&lt;/strong&gt; Post-merge monitoring, production validation, performance profiling. I watch dashboards, check distributed traces in Zipkin, validate circuit breaker behavior. AI agents have no business in production operations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results After Six Months
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Feature delivery time dropped ~40%.&lt;/strong&gt; Cross-service features that took 4–5 days now take 2–3. The reduction comes from parallel Construction — three repos at once instead of sequential.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Unit test coverage increased.&lt;/strong&gt; Writing tests is mechanical work agents handle well. I used to skip edge-case tests when tired. The agent does not get tired.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code review quality improved.&lt;/strong&gt; When I write code, I review it with the same mental model — and miss my own blind spots. Agent-written code gets genuinely fresh-eyed review. The separation between author and reviewer becomes real.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Commit history became cleaner.&lt;/strong&gt; Selective staging and single-purpose commits made &lt;code&gt;git bisect&lt;/code&gt; actually useful.&lt;/p&gt;

&lt;p&gt;Not everything parallelizes. Tightly coupled changes where Service B depends on Service A's final interface still bottleneck. Those still see gains, but more like 15–20%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pitfalls Worth Knowing
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The agent optimizes locally.&lt;/strong&gt; It writes correct code for the service it sees while introducing interface mismatches with services it cannot see. Cross-service contracts must come from a human holding the full system model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Spec quality is the bottleneck.&lt;/strong&gt; "Add error handling to the booking endpoint" gets generic try-catch blocks. "Return HTTP 409 with BOOKING_CONFLICT when the fare class is unavailable, include available fare classes in the response body, emit a Kafka event to pricing-update" gets exactly what you need.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't let agents refactor while implementing.&lt;/strong&gt; Mixed diffs — new functionality tangled with cleanup — are impossible to review. Implement against the existing structure; refactor in a separate PR.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check in every 15–20 minutes.&lt;/strong&gt; Not because the agent will break something catastrophic (permission scopes prevent that), but because a five-minute course correction at the 20-minute mark beats throwing away 90 minutes of wrong-direction implementation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bigger Picture
&lt;/h2&gt;

&lt;p&gt;This is not about AI replacing developers. It is about restructuring the workflow so human attention goes where it matters: architecture, system-level reasoning, production reliability. The mechanical act of translating a well-defined spec into working code with tests — that is what agents handle well. At the scale we operate at AirAsia, that mechanical work was eating a disproportionate amount of my time.&lt;/p&gt;

&lt;p&gt;The spec-writing phase now takes longer than it used to. That tradeoff is worth it.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://krishnakky.com/blog/ai-driven-development-claude-code" rel="noopener noreferrer"&gt;krishnakky.com&lt;/a&gt;. Read the full version with detailed Git integration workflow, war stories, and what I'm experimenting with next.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>devtools</category>
      <category>programming</category>
    </item>
    <item>
      <title>Building Microservices at 130 Million Requests Per Day</title>
      <dc:creator>Krishna Yadav</dc:creator>
      <pubDate>Mon, 05 Oct 2026 16:56:59 +0000</pubDate>
      <link>https://dev.to/krishna_y/building-microservices-at-130-million-requests-per-day-42cl</link>
      <guid>https://dev.to/krishna_y/building-microservices-at-130-million-requests-per-day-42cl</guid>
      <description>&lt;p&gt;I've spent the last three-plus years building and operating the microservices behind AirAsia Move — the travel super-app that handles flights, hotels, and ancillaries for millions of travelers across Southeast Asia. The system processes over 130 million requests per day, roughly 1,500 requests per second sustained, with spikes well above that during flash sales.&lt;/p&gt;

&lt;p&gt;I own services on the Manage My Booking platform: flight changes, fare summaries, price-slash features, post-booking ancillaries. A single user action like changing a flight can fan out to eight or nine downstream calls. At this scale, the engineering problems aren't about whether your code compiles. They're about what happens when one of those downstream services is 200ms slower than usual.&lt;/p&gt;

&lt;h2&gt;
  
  
  Circuit Breakers with Resilience4j
&lt;/h2&gt;

&lt;p&gt;The pattern that proved its weight in gold was the circuit breaker. We use &lt;a href="https://resilience4j.readme.io/docs/getting-started" rel="noopener noreferrer"&gt;Resilience4j&lt;/a&gt; across all Spring Boot services — if a downstream service starts failing, stop calling it instead of piling up timeouts that cascade through the system.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@CircuitBreaker&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"pricingService"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fallbackMethod&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"cachedFareFallback"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;FareSummary&lt;/span&gt; &lt;span class="nf"&gt;getFareSummary&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;bookingId&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;pricingClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;fetchFare&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bookingId&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="nc"&gt;FareSummary&lt;/span&gt; &lt;span class="nf"&gt;cachedFareFallback&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;bookingId&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;Throwable&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;warn&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Pricing service unavailable for booking {}, using cache"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bookingId&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;fareCache&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getLastKnown&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bookingId&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;map&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fare&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;fare&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;withStaleFlag&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;orElseThrow&lt;/span&gt;&lt;span class="o"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ServiceUnavailableException&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"No cached fare available"&lt;/span&gt;&lt;span class="o"&gt;));&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We pair breakers with fallbacks. If the pricing service is down, the user still sees a fare — it might be a few minutes stale, but that beats a blank screen or a 500.&lt;/p&gt;

&lt;p&gt;We tuned thresholds through load testing with JMeter. The defaults were too aggressive and tripped breakers during normal latency variance. A 50% failure rate over a sliding window of 20 calls, with a 30-second open-state wait, matched our actual traffic pattern.&lt;/p&gt;

&lt;h2&gt;
  
  
  Distributed Tracing
&lt;/h2&gt;

&lt;p&gt;When a request crosses ten services, debugging without distributed tracing is a lost cause. We use Spring Cloud Sleuth for trace propagation and Zipkin for visualization. We sample 10% of traces in production — at 1,500 RPS, that's 150 traces per second, more than enough to spot patterns.&lt;/p&gt;

&lt;p&gt;The biggest tracing win wasn't debugging individual slow requests. We noticed every Tuesday between 2–4 AM UTC, latency spiked on the inventory service. A batch job was running full table scans on the same database the API was reading from. Moved the job to a read replica, spikes gone.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Cascading Timeout
&lt;/h2&gt;

&lt;p&gt;During a Diwali sale, the payment gateway started responding 300ms slower than usual. Not enough to fail — just enough to back up our thread pool. We had a fixed pool of 200 threads with a 5-second payment timeout and a bounded queue of 500. When every thread was blocked on slow payment calls, the queue filled. Upstream services calling us started timing out. Within 90 seconds, three services were effectively down.&lt;/p&gt;

&lt;p&gt;The root cause wasn't the payment gateway. It was our thread pool and timeout configuration.&lt;/p&gt;

&lt;p&gt;We made three changes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Cut the payment timeout from 5s to 2s.&lt;/strong&gt; If it hasn't responded in 2 seconds during peak load, it's not going to.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Added a bulkhead&lt;/strong&gt; — isolated the payment call to its own thread pool so one slow dependency can't starve everything else.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Switched the circuit breaker to a time-based sliding window&lt;/strong&gt; so it reacts faster during traffic spikes.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Memory Leak
&lt;/h2&gt;

&lt;p&gt;Our flight-change service was getting OOMKilled by Kubernetes every 48 hours. Two days of profiling with JVisualVM found a &lt;code&gt;ConcurrentHashMap&lt;/code&gt; used as an in-memory cache with no eviction. Every unique booking ID got cached and never expired. Under sustained 1,500 RPS, it grew until the JVM ran out of heap.&lt;/p&gt;

&lt;p&gt;Five lines of code, two days of investigation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nc"&gt;Cache&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;FareSnapshot&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;fareCache&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Caffeine&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;newBuilder&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt;
    &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;maximumSize&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;50_000&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;expireAfterWrite&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Duration&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;ofMinutes&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
    &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What I'd Do Differently
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Contract testing from day one.&lt;/strong&gt; When Service A renames &lt;code&gt;fare_amount&lt;/code&gt; to &lt;code&gt;fareAmount&lt;/code&gt;, both services' unit tests pass. Staging explodes. &lt;a href="https://martinfowler.com/articles/consumerDrivenContracts.html" rel="noopener noreferrer"&gt;Consumer-driven contract tests&lt;/a&gt; catch this at build time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Structured logging earlier.&lt;/strong&gt; Searching Kibana for a booking ID across 10 services when half log &lt;code&gt;bookingId=ABC123&lt;/code&gt; and the other half log &lt;code&gt;Processing booking ABC123 for user xyz&lt;/code&gt; is painful. Consistent field names should be a service template requirement on day one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;More aggressive load shedding.&lt;/strong&gt; At 90% capacity, return 429s for low-priority endpoints and keep booking and payment paths fully served, instead of letting everything degrade equally.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Takeaway
&lt;/h2&gt;

&lt;p&gt;Building microservices at this scale is about understanding failure modes. Every pattern — circuit breakers, tracing, async messaging, bulkheads — exists because something broke in production and we needed it to not break the same way again. The system processes 130 million requests a day not because we got the design right on the first try, but because we built the instrumentation to see what was breaking and the patterns to contain the blast radius.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://krishnakky.com/blog/building-microservices-at-scale" rel="noopener noreferrer"&gt;krishnakky.com&lt;/a&gt; — the full version has additional code examples, Kafka event-driven architecture details, and a third production war story about consumer group rebalances.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>microservices</category>
      <category>java</category>
      <category>springboot</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
