<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: tejaswipandava</title>
    <description>The latest articles on DEV Community by tejaswipandava (@tejaswipandava).</description>
    <link>https://dev.to/tejaswipandava</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F403599%2Fd88ffb6c-e246-4718-b3da-b52602587f30.png</url>
      <title>DEV Community: tejaswipandava</title>
      <link>https://dev.to/tejaswipandava</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tejaswipandava"/>
    <language>en</language>
    <item>
      <title>Kafka Interview Prep (Part 1) — Know Your Fundamentals Cold</title>
      <dc:creator>tejaswipandava</dc:creator>
      <pubDate>Wed, 07 Oct 2026 04:41:00 +0000</pubDate>
      <link>https://dev.to/tejaswipandava/kafka-interview-prep-part-1-know-your-fundamentals-cold-3m29</link>
      <guid>https://dev.to/tejaswipandava/kafka-interview-prep-part-1-know-your-fundamentals-cold-3m29</guid>
      <description>&lt;p&gt;Kafka is mandatory for any Backend / Java / Microservices interview. This is Part 1 of a&lt;br&gt;
two-part prep sheet — it covers the &lt;strong&gt;core architecture, partitioning, brokers, and delivery&lt;br&gt;
semantics&lt;/strong&gt; questions that come up in almost every interview. Part 2 moves into real-world&lt;br&gt;
design and production-failure reasoning, aimed at senior/lead-level depth.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;📎 &lt;strong&gt;Continue to Part 2&lt;/strong&gt; → &lt;em&gt;Kafka Interview Prep Part 2 — Production Scenarios &amp;amp; Real-World Design&lt;/em&gt;&lt;br&gt;
once you're comfortable with everything below.&lt;/p&gt;
&lt;/blockquote&gt;


&lt;h2&gt;
  
  
  1. Kafka Architecture
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Producer → Kafka Cluster [Broker 1, 2, 3] → Consumer Group.&lt;/strong&gt; Zookeeper (legacy) or KRaft&lt;br&gt;
(modern, Zookeeper-less mode) manages cluster metadata (broker list, partition leaders, ACLs).&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Topic&lt;/strong&gt; — a named stream of messages, split into &lt;strong&gt;partitions&lt;/strong&gt; for parallelism.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Partition&lt;/strong&gt; — an ordered, append-only log. Each partition has one &lt;strong&gt;leader&lt;/strong&gt; broker and N
&lt;strong&gt;replicas&lt;/strong&gt; (followers) for fault tolerance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Broker&lt;/strong&gt; — a Kafka server that stores partitions and serves reads/writes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consumer Group&lt;/strong&gt; — a set of consumers that jointly consume a topic; each partition is
consumed by exactly one consumer within a group at a time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zookeeper vs. KRaft&lt;/strong&gt; — Zookeeper was the original external metadata store; KRaft
(Kafka Raft) is the newer built-in consensus protocol that removes the Zookeeper dependency
entirely (default since Kafka 3.x / mandatory from Kafka 4.0).
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Producer(s) → [Broker 1 | Broker 2 | Broker 3] → Consumer Group (C1, C2, C3...)
                     ↑
              KRaft / Zookeeper (metadata, leader election)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h2&gt;
  
  
  2. acks=0, acks=1, acks=all?
&lt;/h2&gt;

&lt;p&gt;Controls how many broker acknowledgments a producer waits for before considering a write&lt;br&gt;
successful — the core producer durability vs. latency trade-off.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;acks=0&lt;/strong&gt; — Producer doesn't wait for any acknowledgment. Fastest, but data can be lost if the
broker fails before the write is even attempted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;acks=1&lt;/strong&gt; — Only the partition &lt;strong&gt;leader&lt;/strong&gt; acknowledges the write. Faster than &lt;code&gt;all&lt;/code&gt;, but if the
leader crashes before followers replicate the message, it's lost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;acks=all&lt;/strong&gt; (or &lt;code&gt;-1&lt;/code&gt;) — All &lt;strong&gt;in-sync replicas (ISR)&lt;/strong&gt; must acknowledge. Safest option — no
acknowledged write is lost as long as at least one ISR survives — but slowest, since it waits
for replication round-trips.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  3. How to prevent duplicate payment processing?
&lt;/h2&gt;

&lt;p&gt;Standard go-to answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Enable the &lt;strong&gt;idempotent producer&lt;/strong&gt;: &lt;code&gt;enable.idempotence=true&lt;/code&gt; — Kafka assigns each producer a
unique ID + sequence number per partition, so the broker can detect and drop duplicate retries
automatically. (Note: this only dedupes broker-side retries of the &lt;em&gt;same&lt;/em&gt; send — it does not
protect against an application-level duplicate, like the user double-clicking "Pay".)&lt;/li&gt;
&lt;li&gt;Store &lt;strong&gt;processed message IDs&lt;/strong&gt; in a DB/cache as an idempotency key — before processing, check
if the ID was already handled; if so, skip (or return the cached result).&lt;/li&gt;
&lt;li&gt;Use &lt;strong&gt;Kafka Transactions&lt;/strong&gt; (&lt;code&gt;transactional.id&lt;/code&gt;) combined with &lt;strong&gt;exactly-once semantics (EOS)&lt;/strong&gt;
for read-process-write workflows spanning multiple topics/partitions.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  4. What if 6 partitions and 8 consumers in one group?
&lt;/h2&gt;

&lt;p&gt;Only 6 consumers will be &lt;strong&gt;active&lt;/strong&gt; (one per partition); the remaining 2 sit &lt;strong&gt;idle&lt;/strong&gt; with no&lt;br&gt;
partitions assigned. &lt;strong&gt;Max active consumers in a group = number of partitions&lt;/strong&gt; — you cannot&lt;br&gt;
parallelize consumption beyond the partition count within a single group. (Scaling further&lt;br&gt;
requires more partitions.)&lt;/p&gt;
&lt;h2&gt;
  
  
  5. How does Kafka decide which partition a message goes to?
&lt;/h2&gt;

&lt;p&gt;Every message is routed to a partition using one of three rules, checked in this order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Explicit partition number provided&lt;/strong&gt; — if the producer explicitly sets a partition number on
the record, Kafka sends it there directly, no hashing involved (see Q6).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Key is provided, no explicit partition&lt;/strong&gt; — Kafka runs the key through a partitioner (by
default, a hash of the key) and maps it to a partition number. The same key &lt;strong&gt;always&lt;/strong&gt; maps to
the same partition, which is exactly how ordering-per-key is achieved.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No key, no explicit partition&lt;/strong&gt; — Kafka spreads messages across partitions in a
round-robin-ish / sticky-batching fashion, just to balance load evenly, with no ordering
guarantee between any two messages.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;
  
  
  6. Can a producer decide exactly which partition its message lands in?
&lt;/h2&gt;

&lt;p&gt;Yes. The producer API lets you pass a partition number directly when sending a record, which&lt;br&gt;
bypasses hashing entirely and forces that message into that exact partition. This is used&lt;br&gt;
sparingly — e.g., for manual load balancing or testing — because it bypasses the key-based&lt;br&gt;
routing that normally keeps ordering and distribution automatic and consistent.&lt;/p&gt;
&lt;h2&gt;
  
  
  7. Do all messages from one producer end up in the same partition?
&lt;/h2&gt;

&lt;p&gt;No — &lt;strong&gt;a producer is not pinned to one partition.&lt;/strong&gt; Each individual message is routed&lt;br&gt;
independently based on Q5's rules (its own key, or round-robin if keyless). A single producer&lt;br&gt;
instance can and typically does write to every partition of a topic; what stays consistent is&lt;br&gt;
only that messages &lt;em&gt;sharing the same key&lt;/em&gt; always land in the same partition, regardless of which&lt;br&gt;
producer sent them.&lt;/p&gt;
&lt;h2&gt;
  
  
  8. What exactly is a broker, and where does it fit in?
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;broker&lt;/strong&gt; is a single Kafka server process — the thing that actually stores partition data on&lt;br&gt;
disk and serves producer writes / consumer reads over the network. A &lt;strong&gt;Kafka cluster&lt;/strong&gt; is simply&lt;br&gt;
a group of brokers working together. Producers and consumers don't talk to "Kafka" as an&lt;br&gt;
abstract thing — they open connections to specific brokers, which is why the client needs&lt;br&gt;
cluster metadata (which broker currently leads which partition) before it can send or read data.&lt;/p&gt;
&lt;h2&gt;
  
  
  9. Can one broker host multiple partitions?
&lt;/h2&gt;

&lt;p&gt;Yes, routinely. A broker typically hosts partitions (as leader and/or replica) from &lt;strong&gt;many&lt;br&gt;
different topics simultaneously&lt;/strong&gt; — there's no one-partition-per-broker restriction. With 3&lt;br&gt;
brokers and a topic of 6 partitions, for example, each broker would naturally end up leading&lt;br&gt;
roughly 2 of those partitions, plus hosting replicas of others.&lt;/p&gt;
&lt;h2&gt;
  
  
  10. What's the relationship between a topic and a broker? Can one broker host multiple topics?
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;topic&lt;/strong&gt; is a logical name for a stream of data; a &lt;strong&gt;broker&lt;/strong&gt; is a physical server. They're&lt;br&gt;
independent concepts — a topic's partitions are &lt;strong&gt;spread across&lt;/strong&gt; the brokers in the cluster&lt;br&gt;
(for parallelism and fault tolerance), and conversely, a single broker commonly stores partitions&lt;br&gt;
belonging to &lt;strong&gt;many different topics&lt;/strong&gt; at once. So yes — one broker hosting multiple topics (and&lt;br&gt;
multiple partitions per topic) is the normal, expected setup, not an edge case.&lt;/p&gt;
&lt;h2&gt;
  
  
  11. If a producer targets a specific partition and that partition's leader is down, what happens?
&lt;/h2&gt;

&lt;p&gt;The producer's metadata (refreshed periodically, and on error) tells it which broker currently&lt;br&gt;
leads each partition. If the leader is down, the producer's send fails against the stale leader,&lt;br&gt;
triggers a &lt;strong&gt;metadata refresh&lt;/strong&gt;, and the cluster's controller will have already promoted an&lt;br&gt;
in-sync replica (see Part 2 for the full failover story) to be the new leader. The producer then&lt;br&gt;
retries against the new leader transparently (assuming retries are enabled, which is the&lt;br&gt;
default) — from the caller's point of view this looks like a brief delay/retry, not a hard&lt;br&gt;
failure, as long as an in-sync replica was available to take over.&lt;/p&gt;
&lt;h2&gt;
  
  
  12. Does Kafka put a hard cap on partitions, brokers, consumers, or producers?
&lt;/h2&gt;

&lt;p&gt;No single hard-coded number — these are &lt;strong&gt;practical/operational limits&lt;/strong&gt;, not protocol limits:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Partitions per topic / per cluster&lt;/strong&gt; — technically unbounded, but every partition has a
real memory/file-handle/metadata cost per broker, and very high partition counts (tens of
thousands) slow down leader elections and controller operations. Teams size this deliberately,
not infinitely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Brokers per cluster&lt;/strong&gt; — no fixed cap; production clusters range from a handful to hundreds of
brokers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consumers in a group&lt;/strong&gt; — not capped by Kafka itself, but &lt;strong&gt;only as many are ever active as
there are partitions&lt;/strong&gt; (see Q4) — any extra consumers in the same group just sit idle.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Producers&lt;/strong&gt; — effectively unlimited; any number of producers can write to the same topic
concurrently.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical takeaway for an interview: the real constraint is almost always partition count&lt;br&gt;
(it caps parallelism and has a real operational cost), not some hidden Kafka maximum.&lt;/p&gt;
&lt;h2&gt;
  
  
  13. Why did Kafka move away from Zookeeper to KRaft?
&lt;/h2&gt;

&lt;p&gt;Zookeeper was an &lt;strong&gt;external&lt;/strong&gt; system Kafka depended on purely to store cluster metadata (broker&lt;br&gt;
list, partition leaders, ACLs) and run leader elections — meaning every Kafka deployment actually&lt;br&gt;
had to run and operate &lt;em&gt;two&lt;/em&gt; distributed systems. This added real costs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Operational overhead&lt;/strong&gt; — a second system to deploy, monitor, patch, and scale separately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scalability ceiling&lt;/strong&gt; — Zookeeper itself became a bottleneck for clusters with very large
numbers of partitions, since every metadata change was a round-trip to Zookeeper.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Slower recovery&lt;/strong&gt; — controller failover (re-electing which broker manages cluster metadata)
was slower when routed through an external system.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;KRaft&lt;/strong&gt; (Kafka Raft) replaces Zookeeper with a built-in consensus protocol (Raft) run by the&lt;br&gt;
Kafka brokers themselves — metadata now lives inside Kafka as just another replicated log. This&lt;br&gt;
means &lt;strong&gt;one system instead of two&lt;/strong&gt;, faster controller failover, and the ability to scale to far&lt;br&gt;
more partitions. KRaft has been production-ready since Kafka 3.x and is mandatory (Zookeeper&lt;br&gt;
support removed) from Kafka 4.0 onward.&lt;/p&gt;
&lt;h2&gt;
  
  
  14. How to maintain ordering for same Order ID?
&lt;/h2&gt;

&lt;p&gt;Use the &lt;strong&gt;Order ID as the message key&lt;/strong&gt;. Kafka's default partitioner hashes the key to&lt;br&gt;
deterministically route all messages with the same key to the &lt;strong&gt;same partition&lt;/strong&gt;. Since a&lt;br&gt;
partition is a strictly ordered log, all events for that Order ID are processed in order.&lt;/p&gt;

&lt;p&gt;⚠️ &lt;strong&gt;Ordering is guaranteed WITHIN a partition only&lt;/strong&gt; — there's no ordering guarantee &lt;em&gt;across&lt;/em&gt;&lt;br&gt;
partitions/topics.&lt;/p&gt;
&lt;h2&gt;
  
  
  15. What is Dead Letter Topic (DLT)?
&lt;/h2&gt;

&lt;p&gt;When a message fails processing after exhausting all configured retries, instead of dropping it&lt;br&gt;
or blocking the partition indefinitely, it's routed to a separate &lt;strong&gt;Dead Letter Topic&lt;/strong&gt; for&lt;br&gt;
manual inspection, alerting, or reprocessing later. This ensures &lt;strong&gt;no message is silently lost&lt;/strong&gt;&lt;br&gt;
and a single poison-pill message doesn't stall the whole consumer.&lt;/p&gt;
&lt;h2&gt;
  
  
  16. At-Most-Once vs At-Least-Once vs Exactly-Once?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Semantic&lt;/th&gt;
&lt;th&gt;Config&lt;/th&gt;
&lt;th&gt;Behavior&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;At-Most-Once&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;acks=0&lt;/code&gt; + commit offset &lt;em&gt;before&lt;/em&gt; processing completes&lt;/td&gt;
&lt;td&gt;May lose messages, never duplicates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;At-Least-Once&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;acks=all&lt;/code&gt; + commit offset &lt;em&gt;after&lt;/em&gt; processing completes&lt;/td&gt;
&lt;td&gt;Never loses messages, may duplicate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Exactly-Once&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Idempotent producer + Kafka Transactions + manual offset commit&lt;/td&gt;
&lt;td&gt;Never loses, never duplicates&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At-Least-Once is the most common default in practice, paired with idempotent consumer logic&lt;br&gt;
(dedup by message ID) to approximate exactly-once behavior without the full transactional&lt;br&gt;
overhead.&lt;/p&gt;

&lt;p&gt;⚠️ Nuance worth mentioning if pressed: with Kafka's &lt;strong&gt;default auto-commit&lt;/strong&gt; (&lt;code&gt;enable.auto.commit=true&lt;/code&gt;),&lt;br&gt;
offsets are actually committed on a background timer (&lt;code&gt;auto.commit.interval.ms&lt;/code&gt;), not strictly&lt;br&gt;
"right before" or "right after" a given message — so depending on exactly when a crash happens&lt;br&gt;
relative to that timer, auto-commit can land you in &lt;em&gt;either&lt;/em&gt; at-most-once or at-least-once&lt;br&gt;
territory. For a predictable guarantee, disable auto-commit and commit manually at the point you&lt;br&gt;
intend (before vs. after processing).&lt;/p&gt;
&lt;h2&gt;
  
  
  17. Consumer Lag?
&lt;/h2&gt;

&lt;p&gt;The difference between the &lt;strong&gt;latest offset&lt;/strong&gt; (last message produced to a partition) and the&lt;br&gt;
&lt;strong&gt;consumer's committed offset&lt;/strong&gt; (last message it has processed). High lag means the consumer is&lt;br&gt;
falling behind the producer's rate — a key signal for scaling consumers or investigating slow&lt;br&gt;
processing.&lt;/p&gt;

&lt;p&gt;Monitor via: &lt;strong&gt;Kafka UI / AKHQ&lt;/strong&gt;, &lt;strong&gt;Grafana + Prometheus (via JMX exporter)&lt;/strong&gt;, or the CLI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kafka-consumer-groups.sh &lt;span class="nt"&gt;--bootstrap-server&lt;/span&gt; &amp;lt;broker&amp;gt; &lt;span class="nt"&gt;--describe&lt;/span&gt; &lt;span class="nt"&gt;--group&lt;/span&gt; &amp;lt;group-id&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  18. Kafka vs RabbitMQ?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Kafka&lt;/th&gt;
&lt;th&gt;RabbitMQ&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;Log-based (durable, replayable log)&lt;/td&gt;
&lt;td&gt;Queue-based (messages removed once consumed)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Throughput&lt;/td&gt;
&lt;td&gt;Very high (millions of msgs/sec)&lt;/td&gt;
&lt;td&gt;Moderate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Use case&lt;/td&gt;
&lt;td&gt;Event streaming, event sourcing, analytics pipelines&lt;/td&gt;
&lt;td&gt;Task queues, RPC-style messaging, work distribution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ordering&lt;/td&gt;
&lt;td&gt;Per-partition ordering&lt;/td&gt;
&lt;td&gt;Per-queue ordering&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Replay&lt;/td&gt;
&lt;td&gt;Yes — consumers can re-read from any offset&lt;/td&gt;
&lt;td&gt;No — once acked, message is gone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Routing&lt;/td&gt;
&lt;td&gt;Simple (topic/partition/key)&lt;/td&gt;
&lt;td&gt;Rich (exchanges: direct, topic, fanout, headers)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Rule of thumb: Kafka for high-throughput event streaming and replayable logs; RabbitMQ for&lt;br&gt;
flexible task/work queues with complex routing needs.&lt;/p&gt;

&lt;h2&gt;
  
  
  19. Real Design: Order → Payment → Inventory
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Order Service  → order-topic   → Payment Service
                                        |
                                        ▼
                                 payment-topic
                                        |
                                        ▼
                               Inventory Service
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each service &lt;strong&gt;consumes&lt;/strong&gt; events relevant to it and &lt;strong&gt;produces&lt;/strong&gt; new events downstream — this is&lt;br&gt;
the classic &lt;strong&gt;event-driven microservices / choreography&lt;/strong&gt; pattern:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Order Service places an order → publishes to &lt;code&gt;order-topic&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Payment Service consumes &lt;code&gt;order-topic&lt;/code&gt;, processes payment → publishes result to
&lt;code&gt;payment-topic&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Inventory Service consumes &lt;code&gt;payment-topic&lt;/code&gt; → reserves/deducts stock.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Benefits: &lt;strong&gt;asynchronous&lt;/strong&gt; (services don't block on each other), &lt;strong&gt;scalable&lt;/strong&gt; (each service&lt;br&gt;
scales independently by adding consumers/partitions), &lt;strong&gt;fault-tolerant&lt;/strong&gt; (a downstream service&lt;br&gt;
being briefly down doesn't lose events — they sit in the topic until it recovers).&lt;/p&gt;

&lt;p&gt;Trade-off to mention in an interview: this requires careful handling of &lt;strong&gt;failure/compensation&lt;/strong&gt;&lt;br&gt;
(e.g., what happens if Inventory can't fulfill after Payment succeeded? → Saga pattern with&lt;br&gt;
compensating transactions) and &lt;strong&gt;idempotency&lt;/strong&gt; at each consumption step (see Q3 and Q16).&lt;/p&gt;

&lt;h2&gt;
  
  
  20. What is the role of the &lt;code&gt;__consumer_offsets&lt;/code&gt; topic?
&lt;/h2&gt;

&lt;p&gt;Kafka stores committed consumer offsets in an internal, compacted topic called&lt;br&gt;
&lt;code&gt;__consumer_offsets&lt;/code&gt; (instead of an external store like Zookeeper, as in older Kafka versions).&lt;br&gt;
This lets consumers resume from their last committed position after a restart or rebalance.&lt;/p&gt;

&lt;h2&gt;
  
  
  21. What is a rebalance, and why can it be disruptive?
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;rebalance&lt;/strong&gt; occurs when consumers join/leave a group (e.g., scaling up, a consumer crash, or&lt;br&gt;
a deployment), triggering redistribution of partitions among the remaining/new consumers. During&lt;br&gt;
a rebalance, consumption pauses ("stop-the-world"), which can hurt throughput/latency if&lt;br&gt;
rebalances happen frequently. Mitigations: &lt;strong&gt;Incremental Cooperative Rebalancing&lt;/strong&gt; (&lt;code&gt;Cooperative&lt;br&gt;
Sticky&lt;/code&gt; assignor) minimizes disruption by only reassigning partitions that need to move, instead&lt;br&gt;
of revoking all partitions from all consumers.&lt;/p&gt;

&lt;h2&gt;
  
  
  22. What is ISR (In-Sync Replica)?
&lt;/h2&gt;

&lt;p&gt;The set of replicas (leader + followers) that are fully caught up with the leader's log within an&lt;br&gt;
allowed lag (&lt;code&gt;replica.lag.time.max.ms&lt;/code&gt;). &lt;code&gt;acks=all&lt;/code&gt; requires acknowledgment specifically from all&lt;br&gt;
current ISR members — if a replica falls too far behind, it's dropped from the ISR and no longer&lt;br&gt;
blocks writes.&lt;/p&gt;

&lt;h2&gt;
  
  
  23. What happens if a partition leader dies?
&lt;/h2&gt;

&lt;p&gt;Kafka's controller (elected via KRaft/Zookeeper) detects the failure and promotes one of the&lt;br&gt;
in-sync replicas (ISR) to become the new leader — &lt;strong&gt;never an out-of-sync replica&lt;/strong&gt;, since that&lt;br&gt;
could silently lose committed data (this is the &lt;code&gt;unclean.leader.election.enable&lt;/code&gt; setting, which&lt;br&gt;
defaults to &lt;code&gt;false&lt;/code&gt; for exactly this reason). Producers/consumers refresh their metadata and&lt;br&gt;
redirect requests to the new leader — this is largely transparent to clients, though it can cause&lt;br&gt;
a brief availability blip during the election.&lt;/p&gt;

&lt;h2&gt;
  
  
  24. How does Kafka achieve high throughput?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sequential disk I/O&lt;/strong&gt; — append-only log writes are much faster than random I/O.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero-copy transfer&lt;/strong&gt; — &lt;code&gt;sendfile()&lt;/code&gt; lets Kafka send data from disk to network socket without
copying through user-space memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batching &amp;amp; compression&lt;/strong&gt; — producers batch multiple messages per request and can compress
batches (gzip, snappy, lz4, zstd) to reduce network/disk overhead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Partitioning&lt;/strong&gt; — parallelism across brokers and consumers.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  25. What is log compaction?
&lt;/h2&gt;

&lt;p&gt;An alternative retention policy (vs. time/size-based deletion) where Kafka retains only the&lt;br&gt;
&lt;strong&gt;latest value for each key&lt;/strong&gt; in a topic, removing older duplicate-keyed records in the&lt;br&gt;
background. Useful for topics representing "current state" (e.g., a changelog topic backing a&lt;br&gt;
KTable) rather than an event history.&lt;/p&gt;

&lt;h2&gt;
  
  
  26. What is the difference between a Kafka Producer's &lt;code&gt;linger.ms&lt;/code&gt; and &lt;code&gt;batch.size&lt;/code&gt;?
&lt;/h2&gt;

&lt;p&gt;Both control batching behavior on the producer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;batch.size&lt;/code&gt; — max bytes to batch before sending (per partition).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;linger.ms&lt;/code&gt; — max time to wait for more messages to accumulate before sending, even if the
batch isn't full. Tuning these trades off latency (lower values) against throughput/efficiency
(higher values, larger batches).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  27. What is Kafka Streams / KSQL used for?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Kafka Streams&lt;/strong&gt; is a Java library for building stream-processing applications directly on top&lt;br&gt;
of Kafka (stateful transformations, joins, windowed aggregations) without needing a separate&lt;br&gt;
processing cluster (like Spark/Flink). &lt;strong&gt;ksqlDB&lt;/strong&gt; provides a SQL-like interface over the same&lt;br&gt;
capabilities for simpler use cases.&lt;/p&gt;

&lt;h2&gt;
  
  
  28. How do you handle schema evolution in Kafka messages?
&lt;/h2&gt;

&lt;p&gt;Use a &lt;strong&gt;Schema Registry&lt;/strong&gt; (e.g., Confluent Schema Registry) with &lt;strong&gt;Avro/Protobuf/JSON Schema&lt;/strong&gt;&lt;br&gt;
to enforce and version message schemas. Configure compatibility modes (backward, forward, full)&lt;br&gt;
so producers/consumers can evolve independently without breaking each other — critical in a&lt;br&gt;
microservices setup where teams deploy on different timelines.&lt;/p&gt;

&lt;h2&gt;
  
  
  29. What is the difference between &lt;code&gt;poll()&lt;/code&gt; timeout and &lt;code&gt;max.poll.interval.ms&lt;/code&gt;?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;poll()&lt;/code&gt; &lt;strong&gt;timeout&lt;/strong&gt; — how long a single call to &lt;code&gt;poll()&lt;/code&gt; blocks waiting for new records if none
are immediately available.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;max.poll.interval.ms&lt;/code&gt; — the max allowed time &lt;strong&gt;between&lt;/strong&gt; successive &lt;code&gt;poll()&lt;/code&gt; calls before the
consumer is considered dead and triggers a rebalance. Important when message processing itself
is slow — if processing takes longer than this interval without calling &lt;code&gt;poll()&lt;/code&gt; again, the
consumer gets kicked out of the group.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;📎 &lt;strong&gt;Up next:&lt;/strong&gt; &lt;em&gt;Kafka Interview Prep Part 2 — Production Scenarios &amp;amp; Real-World Design&lt;/em&gt;, covering&lt;br&gt;
the Zomato live-location case study and senior/lead-level production failure scenarios&lt;br&gt;
(duplicate processing, consumer lag diagnosis, multi-event-type topics).&lt;/p&gt;

</description>
      <category>techlettersbyteja</category>
      <category>systemdesign</category>
      <category>kafka</category>
    </item>
    <item>
      <title>Idempotency - "The customer clicked 'Pay' three times. How do you ensure they're charged only once?"</title>
      <dc:creator>tejaswipandava</dc:creator>
      <pubDate>Tue, 06 Oct 2026 12:45:00 +0000</pubDate>
      <link>https://dev.to/tejaswipandava/idempotency-the-customer-clicked-pay-three-times-how-do-you-ensure-theyre-charged-only-once-2bob</link>
      <guid>https://dev.to/tejaswipandava/idempotency-the-customer-clicked-pay-three-times-how-do-you-ensure-theyre-charged-only-once-2bob</guid>
      <description>&lt;p&gt;Most know what idempotency is; very few can clearly explain why it's needed and how its actually implemented.&lt;/p&gt;




&lt;h2&gt;
  
  
  🤔 The Problem
&lt;/h2&gt;

&lt;p&gt;Imagine a payment system. The user clicks &lt;strong&gt;"Pay."&lt;/strong&gt; The request reaches the server and is&lt;br&gt;
processed successfully — but the &lt;strong&gt;response never reaches the user&lt;/strong&gt; because of a network timeout,&lt;br&gt;
a dropped connection, or a slow proxy.&lt;/p&gt;

&lt;p&gt;What does the user do? They click &lt;strong&gt;"Pay" again.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Without idempotency, this retry can cause the API to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Create &lt;strong&gt;multiple orders&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Charge the customer &lt;strong&gt;duplicate payments&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Deduct &lt;strong&gt;inventory&lt;/strong&gt; more than once&lt;/li&gt;
&lt;li&gt;Produce a very &lt;strong&gt;angry customer&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The core issue: &lt;strong&gt;the client cannot tell the difference between "my request never arrived" and "my&lt;br&gt;
request succeeded but the response got lost."&lt;/strong&gt; Both look identical from the client's side — a&lt;br&gt;
timeout — but calling the API again is only safe in the first case.&lt;/p&gt;
&lt;h2&gt;
  
  
  💡 The Solution: Idempotency Keys
&lt;/h2&gt;

&lt;p&gt;Generate a unique &lt;strong&gt;Idempotency Key&lt;/strong&gt; for each &lt;em&gt;logical&lt;/em&gt; request (not each HTTP call — the same&lt;br&gt;
logical intent, e.g., "charge this cart," keeps the same key across retries).&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;First request&lt;/strong&gt; with a given key → process it normally, then store the response &lt;strong&gt;alongside&lt;/strong&gt;
the key.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retry with the same key&lt;/strong&gt; → the server recognizes it's already been handled, and simply
&lt;strong&gt;returns the previously stored response&lt;/strong&gt; instead of processing the request again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No duplicate processing&lt;/strong&gt;, no matter how many times the client retries.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The core guarantee:&lt;/strong&gt; &lt;em&gt;Same request + same key = same result, every time.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client                         Server
  │  POST /pay (Idempotency-Key: abc123)
  ├──────────────────────────────►
  │                               │  key "abc123" not seen before
  │                               │  → process payment
  │                               │  → store {key: abc123, response: {...}}
  │  ◄── 200 OK {charge_id: X}────┤
  │  (response lost in transit)   │
  │
  │  Retry: POST /pay (Idempotency-Key: abc123)
  ├──────────────────────────────►
  │                               │  key "abc123" already exists
  │                               │  → skip processing, return stored response
  │  ◄── 200 OK {charge_id: X}────┤   (same charge_id, no new charge)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  🌍 Where Idempotency Is Used
&lt;/h2&gt;

&lt;p&gt;Anywhere duplicate side effects are unacceptable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Payment APIs&lt;/li&gt;
&lt;li&gt;Order creation&lt;/li&gt;
&lt;li&gt;Money transfers&lt;/li&gt;
&lt;li&gt;Ticket booking&lt;/li&gt;
&lt;li&gt;Inventory updates&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  🎯 Common Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. What is idempotency?
&lt;/h3&gt;

&lt;p&gt;An operation is idempotent if performing it multiple times has &lt;strong&gt;the same effect as performing it&lt;br&gt;
once&lt;/strong&gt;. In HTTP/API terms: retrying an idempotent request should never change the outcome beyond&lt;br&gt;
what the first successful call already did.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Why is POST not idempotent by default?
&lt;/h3&gt;

&lt;p&gt;By HTTP semantics, &lt;code&gt;POST&lt;/code&gt; typically means &lt;strong&gt;"create a new resource"&lt;/strong&gt; — calling it twice naturally&lt;br&gt;
creates two resources (two orders, two charges), because the server has no built-in way to know&lt;br&gt;
the second call is a retry of the first rather than a genuinely new request. &lt;code&gt;GET&lt;/code&gt;, &lt;code&gt;PUT&lt;/code&gt;, and&lt;br&gt;
&lt;code&gt;DELETE&lt;/code&gt; are idempotent by convention/spec (see Q6), but &lt;code&gt;POST&lt;/code&gt; is not — which is exactly why&lt;br&gt;
payment/order-creation APIs (which are inherently &lt;code&gt;POST&lt;/code&gt;-based, since they create something) need&lt;br&gt;
an &lt;strong&gt;explicit&lt;/strong&gt; idempotency mechanism layered on top; it doesn't come for free from the HTTP verb.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. How do idempotency keys work?
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;The client generates a unique key (typically a UUID) &lt;strong&gt;per logical operation&lt;/strong&gt; — generated once
and reused across all retries of that same operation, never regenerated on retry.&lt;/li&gt;
&lt;li&gt;The client sends it with the request, usually as a header (e.g., &lt;code&gt;Idempotency-Key: &amp;lt;uuid&amp;gt;&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;The server checks a fast lookup store for that key:

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Not seen before&lt;/strong&gt; → process the request, store &lt;code&gt;{key → response}&lt;/code&gt;, return the response.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Already seen, completed&lt;/strong&gt; → skip reprocessing, return the stored response directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Already seen, still in-flight&lt;/strong&gt; → the request is currently being processed by a concurrent
call; the server should make the retry &lt;strong&gt;wait&lt;/strong&gt; (or return a &lt;code&gt;409 Conflict&lt;/code&gt; telling the client
to retry after a short delay) rather than let both proceed and race each other.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4. Where should idempotency keys be stored?
&lt;/h3&gt;

&lt;p&gt;A &lt;strong&gt;fast, shared, low-latency key-value store&lt;/strong&gt; — Redis is the typical choice, since idempotency&lt;br&gt;
checks sit directly on the hot path of every write request and need sub-millisecond lookups. Some&lt;br&gt;
systems back this with a &lt;strong&gt;database unique constraint&lt;/strong&gt; on the idempotency key column instead of&lt;br&gt;
(or in addition to) Redis, trading a bit of latency for stronger durability guarantees (see the&lt;br&gt;
Redis-crash question below for why this matters).&lt;/p&gt;

&lt;h3&gt;
  
  
  5. How long should they be retained?
&lt;/h3&gt;

&lt;p&gt;Long enough to cover the &lt;strong&gt;realistic retry/timeout window&lt;/strong&gt; of the client, but not indefinitely:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Too short → a legitimate late retry (e.g., a mobile client that was offline for a few minutes)
is treated as a brand-new request, causing a duplicate.&lt;/li&gt;
&lt;li&gt;Too long → the store grows unbounded, wasting memory/storage on keys nobody will ever reuse.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A common choice is &lt;strong&gt;24 hours&lt;/strong&gt; for payment-style APIs (covers virtually all realistic client&lt;br&gt;
retry/backoff windows), implemented via a &lt;strong&gt;TTL&lt;/strong&gt; on the stored key.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Which HTTP methods are idempotent?
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Idempotent?&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GET&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;Read-only, no side effects regardless of call count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;PUT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;Replaces a resource with a given representation — calling it N times leaves the same end state as calling it once&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;DELETE&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;Deleting an already-deleted resource still results in "resource doesn't exist"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;POST&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;❌ No (by default)&lt;/td&gt;
&lt;td&gt;Semantically "create" — repeated calls create repeated resources unless an idempotency key is explicitly layered on top&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;PATCH&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;❌ Not guaranteed&lt;/td&gt;
&lt;td&gt;Depends on the semantics of the patch — e.g., "increment balance by 10" is not idempotent, "set status to X" is&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  💡Tip
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Don't say:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;❌ "Idempotency prevents duplicate requests."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is imprecise — idempotency doesn't &lt;em&gt;prevent&lt;/em&gt; the client from sending duplicate requests at&lt;br&gt;
all (they will, especially on timeout/retry); it prevents duplicate requests from causing duplicate&lt;br&gt;
&lt;strong&gt;side effects&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Say instead:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;✅ "Idempotency ensures the same request produces the same outcome. A common implementation uses&lt;br&gt;
an Idempotency Key to identify retries and return the original response instead of processing&lt;br&gt;
the request again."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This shows you understand both &lt;strong&gt;the problem&lt;/strong&gt; (retries are inevitable and indistinguishable from&lt;br&gt;
first attempts) and &lt;strong&gt;the solution mechanism&lt;/strong&gt; (key-based deduplication with a stored response),&lt;br&gt;
not just the buzzword.&lt;/p&gt;




&lt;h2&gt;
  
  
  💬 Follow-up: What if Redis (storing the idempotency keys) crashes?
&lt;/h2&gt;

&lt;p&gt;This is the natural next question which tests your understanding of the trade-offs behind your own solution, not just the happy path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key considerations to raise:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;In-memory Redis alone is a single point of failure for a correctness-critical mechanism.&lt;/strong&gt;&lt;br&gt;
If Redis crashes and loses data (no persistence configured), a request that was already&lt;br&gt;
processed but not yet acknowledged to the client can be &lt;strong&gt;reprocessed on retry&lt;/strong&gt;, since the&lt;br&gt;
server no longer remembers seeing that key — defeating the entire point of idempotency at&lt;br&gt;
exactly the moment it matters most.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Mitigations, roughly from cheapest to strongest:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Redis persistence (AOF/RDB) + replication&lt;/strong&gt; — reduces data loss on a crash, but doesn't
eliminate the window between a write and it being durably persisted/replicated.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backing store with a durable, unique constraint&lt;/strong&gt; — write the idempotency key to a
relational DB (or DynamoDB/Cassandra with a unique key constraint) as the &lt;strong&gt;source of truth&lt;/strong&gt;,
using Redis only as a fast-path cache in front of it. If Redis is down or cold, fall back to
the DB — slower, but still correct. A DB &lt;code&gt;INSERT ... ON CONFLICT DO NOTHING&lt;/code&gt; (or equivalent
unique constraint violation) gives you atomic "claim this key or fail" semantics even without
Redis.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Idempotency at the payment provider itself, as a second line of defense&lt;/strong&gt; — many payment
processors (Stripe, etc.) support their own idempotency keys passed through to &lt;em&gt;them&lt;/em&gt;. Even if
our own store fails and we accidentally call the provider twice with the same key, the
provider's own deduplication prevents an actual double charge. This is the "defense in depth"
answer: don't rely on a single layer being perfectly durable.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Fail-closed vs. fail-open trade-off:&lt;/strong&gt; if the idempotency store is completely unavailable and&lt;br&gt;
there's no fallback, should the request be rejected (fail-closed — safe, but hurts availability)&lt;br&gt;
or allowed through with the retry risk accepted (fail-open — available, but risks a duplicate)?&lt;br&gt;
For payments specifically, I'd lean &lt;strong&gt;fail-closed&lt;/strong&gt; (reject with a clear retriable error) rather&lt;br&gt;
than risk a duplicate charge — the cost of a duplicate payment is much higher than the cost of a&lt;br&gt;
brief availability blip on the write path.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Strong answer framing:&lt;/strong&gt; "I wouldn't rely on Redis alone for something this correctness-critical.&lt;br&gt;
I'd use Redis as a fast-path cache, back it with a durable store with a unique constraint as the&lt;br&gt;
real source of truth, and pass the same idempotency key through to the downstream payment provider&lt;br&gt;
as a second independent layer of protection — so no single component failing causes an actual&lt;br&gt;
double charge."&lt;/p&gt;

</description>
      <category>techlettersbyteja</category>
      <category>systemdesign</category>
      <category>backend</category>
    </item>
    <item>
      <title>Latency Percentiles (p50/p90/p95/p99)</title>
      <dc:creator>tejaswipandava</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:11:31 +0000</pubDate>
      <link>https://dev.to/tejaswipandava/latency-percentiles-p50p90p95p99-4nh6</link>
      <guid>https://dev.to/tejaswipandava/latency-percentiles-p50p90p95p99-4nh6</guid>
      <description>&lt;p&gt;&lt;strong&gt;Latency&lt;/strong&gt; is the time an API takes to respond to a request — the delay between when the request is sent and when the response is received.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Problem
&lt;/h2&gt;

&lt;p&gt;A single "average latency" number hides the experience of your worst-served users and makes capacity/overload decisions impossible. A service can report a perfectly healthy average while a meaningful slice of real users — often the ones on slow networks, hitting a cold cache, or landing on a degraded instance — have a broken experience. You need a metric that tells you how bad the &lt;em&gt;tail&lt;/em&gt; of your traffic looks, not just the middle of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Constraint
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Latency is never normally distributed. It has a long right tail caused by GC pauses, cold starts, lock contention, cache misses, slow downstream calls, and retries. Averages are dominated by the bulk of fast requests and are mathematically insensitive to a worsening tail.&lt;/li&gt;
&lt;li&gt;Consumers of a service (other teams, external clients) need a concrete, published promise of how fast the service responds — not a vague "it's usually fast." This requires an agreed-upon number, not a feeling.&lt;/li&gt;
&lt;li&gt;In a real architecture, a single service rarely acts alone — it fans out to multiple upstream APIs, databases, and third-party services:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Upstream API 1 ──┐
Upstream API 2 ──┼──►  our service  ──┬──► DB 1
Upstream API 3 ──┘                    ├──► DB 2
                                       └──► DB 3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When a request depends on several downstream calls, the overall response is only as fast as the &lt;em&gt;slowest&lt;/em&gt; one. This means tail latency compounds across a call graph — a problem a single-hop percentile number doesn't reveal on its own (see 3.5).&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Computing an exact percentile requires sorting every recorded latency value, which doesn't scale once you're measuring millions of requests per minute across many endpoints.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. Decision
&lt;/h2&gt;

&lt;h3&gt;
  
  
  3.1 Define and publish an SLA, backed by an SLO
&lt;/h3&gt;

&lt;p&gt;An &lt;strong&gt;SLA (Service Level Agreement)&lt;/strong&gt; is the formal promise between a service provider and its consumers — it spells out what will be offered and the standard the provider commits to meeting. Underneath that promise sits an &lt;strong&gt;SLO (Service Level Objective)&lt;/strong&gt; — the measurable target (e.g., "p99 &amp;lt; 120ms") that the team actually tracks day to day. The SLA is the external contract; the SLO is the internal number engineering is held to in order to honor it.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.2 Track percentiles, not the average
&lt;/h3&gt;

&lt;p&gt;A percentile simply answers: "what's the response time for X% of requests?" A &lt;strong&gt;pXX&lt;/strong&gt; value means XX% of requests completed at or below that time — the remaining (100 − XX)% took longer. Reading left to right below, each percentile looks at a progressively smaller, slower slice of traffic:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Percentile&lt;/th&gt;
&lt;th&gt;On 100 requests...&lt;/th&gt;
&lt;th&gt;What it tells you&lt;/th&gt;
&lt;th&gt;What it represents&lt;/th&gt;
&lt;th&gt;When to use it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;p50&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;50 requests were faster than this value&lt;/td&gt;
&lt;td&gt;middle response time&lt;/td&gt;
&lt;td&gt;the typical user's experience&lt;/td&gt;
&lt;td&gt;baseline performance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;p90&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;90 requests were faster; 10 were slower&lt;/td&gt;
&lt;td&gt;early warning sign&lt;/td&gt;
&lt;td&gt;the leading edge of your slow traffic&lt;/td&gt;
&lt;td&gt;first signal that something is starting to degrade&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;p95&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;95 requests were faster; 5 were slower&lt;/td&gt;
&lt;td&gt;overall quality&lt;/td&gt;
&lt;td&gt;the bulk of your users, including the slower ones&lt;/td&gt;
&lt;td&gt;the performance budget most teams commit to&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;p99&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;99 requests were faster; only 1 was slower&lt;/td&gt;
&lt;td&gt;reliability&lt;/td&gt;
&lt;td&gt;your worst-case users&lt;/td&gt;
&lt;td&gt;debugging the tail and latency-critical paths&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Practical guidance: start by optimizing for &lt;strong&gt;p95&lt;/strong&gt; — it's achievable, actionable, and catches real degradation without you chasing every last outlier. Only once p95 is consistently met should you graduate to tightening &lt;strong&gt;p99&lt;/strong&gt;; chasing the tail before the bulk of traffic is stable is usually wasted engineering effort. (And don't stop at p50 alone — it's a useful baseline, but by definition it's blind to the slower half of your traffic, which is exactly what you need to catch before it spreads.)&lt;/p&gt;

&lt;h3&gt;
  
  
  3.3 Compute percentiles with approximate streaming structures, not sorted arrays
&lt;/h3&gt;

&lt;p&gt;At scale, monitoring tools like &lt;strong&gt;Grafana&lt;/strong&gt; and &lt;strong&gt;Prometheus&lt;/strong&gt; don't sort every raw latency value — they use streaming histogram-based approximations (e.g., Prometheus histograms, HDR Histogram, t-digest) that give a percentile estimate with a small, bounded error at a fraction of the cost.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.4 Monitor trends, not single data points
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Collect response times per endpoint.&lt;/li&gt;
&lt;li&gt;Alert on &lt;strong&gt;p95/p99 trend lines&lt;/strong&gt; rather than a single instantaneous reading — this catches genuine, sustained degradation and filters out one-off noise and spikes.&lt;/li&gt;
&lt;li&gt;Compare p95/p99 across regions to catch localized/regional performance issues that a global average would mask.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  3.5 Account for tail latency amplification in fan-out calls
&lt;/h3&gt;

&lt;p&gt;If your service calls out to multiple downstream dependencies to answer a single request (as in the diagram above), the chance that &lt;em&gt;at least one&lt;/em&gt; of them is slow grows quickly with the number of calls. For example, if each downstream call independently has a 1% chance of being "slow" (its own p99), and your service makes 50 such calls to answer one request, the probability that the overall request is slowed by at least one of them is roughly &lt;code&gt;1 - (0.99)^50 ≈ 39%&lt;/code&gt; — far worse than the 1% tail of any individual dependency. This is why a service's end-to-end p99 is often dominated by its &lt;em&gt;worst&lt;/em&gt; dependency, not its average one, and why latency budgets are usually allocated per hop rather than assumed to simply "add up."&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Trade-off
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Percentile-based SLOs require more sophisticated monitoring infrastructure (histogram-aware backends like Prometheus/Grafana) than a single average — accepted because the alternative actively hides the problem you're trying to catch.&lt;/li&gt;
&lt;li&gt;Approximate percentile structures (t-digest, HDR Histogram) introduce a small, bounded estimation error — accepted as negligible compared to the cost of exact computation at production volume.&lt;/li&gt;
&lt;li&gt;Pushing hard for a tight p99 across every dependency in a fan-out architecture is costly (redundant calls, hedged requests, over-provisioned capacity) — accepted selectively, only on the paths where the business impact of a slow tail actually justifies it. Chasing &lt;strong&gt;p99.99 everywhere, immediately&lt;/strong&gt; is the extreme version of this mistake: expensive for most endpoints, with diminishing returns that rarely justify the cost outside of genuinely latency-critical paths.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  5. Failure Mode
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Metrics blind spot&lt;/strong&gt;: dashboards show a stable p50 (and even a stable average) while p95/p99 silently climbs for days. A narrow but real degradation — e.g., one bad database replica serving 5% of traffic — can go undetected until it's wide enough to affect the median too. &lt;em&gt;Mitigation: always alert explicitly on p95/p99, never just p50.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Noisy single-sample alerting&lt;/strong&gt;: alerting on one bad percentile reading from a single window causes alert fatigue and desensitizes the team to real issues. &lt;em&gt;Mitigation: alert on trend over a rolling window, not an instantaneous spike.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tail latency amplification goes unnoticed&lt;/strong&gt;: teams monitor the service's own handler time but not each downstream call individually, so the end-to-end p99 looks "mysteriously" worse than any single dependency's reported p99, with no obvious owner to investigate. &lt;em&gt;Mitigation: track and budget latency per hop (see 3.5), not just end-to-end.&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Where to look first
&lt;/h3&gt;

&lt;p&gt;When p95/p99 degrades, these are the usual suspects:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;DNS and TLS handshake overhead.&lt;/li&gt;
&lt;li&gt;Cold starts of services/instances.&lt;/li&gt;
&lt;li&gt;Missing or cold cache.&lt;/li&gt;
&lt;li&gt;Slow database queries (missing indexes, lock contention).&lt;/li&gt;
&lt;li&gt;Connection pool exhaustion.&lt;/li&gt;
&lt;li&gt;Slow downstream dependencies.&lt;/li&gt;
&lt;li&gt;Garbage collector pauses.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>systemdesign</category>
      <category>latency</category>
      <category>architecture</category>
      <category>techlettersbyteja</category>
    </item>
  </channel>
</rss>
