Kafka is mandatory for any Backend / Java / Microservices interview. This is Part 1 of a
two-part prep sheet — it covers the core architecture, partitioning, brokers, and delivery
semantics questions that come up in almost every interview. Part 2 moves into real-world
design and production-failure reasoning, aimed at senior/lead-level depth.
📎 Continue to Part 2 → Kafka Interview Prep Part 2 — Production Scenarios & Real-World Design
once you're comfortable with everything below.
1. Kafka Architecture
Producer → Kafka Cluster [Broker 1, 2, 3] → Consumer Group. Zookeeper (legacy) or KRaft
(modern, Zookeeper-less mode) manages cluster metadata (broker list, partition leaders, ACLs).
- Topic — a named stream of messages, split into partitions for parallelism.
- Partition — an ordered, append-only log. Each partition has one leader broker and N replicas (followers) for fault tolerance.
- Broker — a Kafka server that stores partitions and serves reads/writes.
- Consumer Group — a set of consumers that jointly consume a topic; each partition is consumed by exactly one consumer within a group at a time.
- Zookeeper vs. KRaft — Zookeeper was the original external metadata store; KRaft (Kafka Raft) is the newer built-in consensus protocol that removes the Zookeeper dependency entirely (default since Kafka 3.x / mandatory from Kafka 4.0).
Producer(s) → [Broker 1 | Broker 2 | Broker 3] → Consumer Group (C1, C2, C3...)
↑
KRaft / Zookeeper (metadata, leader election)
2. acks=0, acks=1, acks=all?
Controls how many broker acknowledgments a producer waits for before considering a write
successful — the core producer durability vs. latency trade-off.
- acks=0 — Producer doesn't wait for any acknowledgment. Fastest, but data can be lost if the broker fails before the write is even attempted.
-
acks=1 — Only the partition leader acknowledges the write. Faster than
all, but if the leader crashes before followers replicate the message, it's lost. -
acks=all (or
-1) — All in-sync replicas (ISR) must acknowledge. Safest option — no acknowledged write is lost as long as at least one ISR survives — but slowest, since it waits for replication round-trips.
3. How to prevent duplicate payment processing?
Standard go-to answer:
- Enable the idempotent producer:
enable.idempotence=true— Kafka assigns each producer a unique ID + sequence number per partition, so the broker can detect and drop duplicate retries automatically. (Note: this only dedupes broker-side retries of the same send — it does not protect against an application-level duplicate, like the user double-clicking "Pay".) - Store processed message IDs in a DB/cache as an idempotency key — before processing, check if the ID was already handled; if so, skip (or return the cached result).
- Use Kafka Transactions (
transactional.id) combined with exactly-once semantics (EOS) for read-process-write workflows spanning multiple topics/partitions.
4. What if 6 partitions and 8 consumers in one group?
Only 6 consumers will be active (one per partition); the remaining 2 sit idle with no
partitions assigned. Max active consumers in a group = number of partitions — you cannot
parallelize consumption beyond the partition count within a single group. (Scaling further
requires more partitions.)
5. How does Kafka decide which partition a message goes to?
Every message is routed to a partition using one of three rules, checked in this order:
- Explicit partition number provided — if the producer explicitly sets a partition number on the record, Kafka sends it there directly, no hashing involved (see Q6).
- Key is provided, no explicit partition — Kafka runs the key through a partitioner (by default, a hash of the key) and maps it to a partition number. The same key always maps to the same partition, which is exactly how ordering-per-key is achieved.
- No key, no explicit partition — Kafka spreads messages across partitions in a round-robin-ish / sticky-batching fashion, just to balance load evenly, with no ordering guarantee between any two messages.
6. Can a producer decide exactly which partition its message lands in?
Yes. The producer API lets you pass a partition number directly when sending a record, which
bypasses hashing entirely and forces that message into that exact partition. This is used
sparingly — e.g., for manual load balancing or testing — because it bypasses the key-based
routing that normally keeps ordering and distribution automatic and consistent.
7. Do all messages from one producer end up in the same partition?
No — a producer is not pinned to one partition. Each individual message is routed
independently based on Q5's rules (its own key, or round-robin if keyless). A single producer
instance can and typically does write to every partition of a topic; what stays consistent is
only that messages sharing the same key always land in the same partition, regardless of which
producer sent them.
8. What exactly is a broker, and where does it fit in?
A broker is a single Kafka server process — the thing that actually stores partition data on
disk and serves producer writes / consumer reads over the network. A Kafka cluster is simply
a group of brokers working together. Producers and consumers don't talk to "Kafka" as an
abstract thing — they open connections to specific brokers, which is why the client needs
cluster metadata (which broker currently leads which partition) before it can send or read data.
9. Can one broker host multiple partitions?
Yes, routinely. A broker typically hosts partitions (as leader and/or replica) from many
different topics simultaneously — there's no one-partition-per-broker restriction. With 3
brokers and a topic of 6 partitions, for example, each broker would naturally end up leading
roughly 2 of those partitions, plus hosting replicas of others.
10. What's the relationship between a topic and a broker? Can one broker host multiple topics?
A topic is a logical name for a stream of data; a broker is a physical server. They're
independent concepts — a topic's partitions are spread across the brokers in the cluster
(for parallelism and fault tolerance), and conversely, a single broker commonly stores partitions
belonging to many different topics at once. So yes — one broker hosting multiple topics (and
multiple partitions per topic) is the normal, expected setup, not an edge case.
11. If a producer targets a specific partition and that partition's leader is down, what happens?
The producer's metadata (refreshed periodically, and on error) tells it which broker currently
leads each partition. If the leader is down, the producer's send fails against the stale leader,
triggers a metadata refresh, and the cluster's controller will have already promoted an
in-sync replica (see Part 2 for the full failover story) to be the new leader. The producer then
retries against the new leader transparently (assuming retries are enabled, which is the
default) — from the caller's point of view this looks like a brief delay/retry, not a hard
failure, as long as an in-sync replica was available to take over.
12. Does Kafka put a hard cap on partitions, brokers, consumers, or producers?
No single hard-coded number — these are practical/operational limits, not protocol limits:
- Partitions per topic / per cluster — technically unbounded, but every partition has a real memory/file-handle/metadata cost per broker, and very high partition counts (tens of thousands) slow down leader elections and controller operations. Teams size this deliberately, not infinitely.
- Brokers per cluster — no fixed cap; production clusters range from a handful to hundreds of brokers.
- Consumers in a group — not capped by Kafka itself, but only as many are ever active as there are partitions (see Q4) — any extra consumers in the same group just sit idle.
- Producers — effectively unlimited; any number of producers can write to the same topic concurrently.
The practical takeaway for an interview: the real constraint is almost always partition count
(it caps parallelism and has a real operational cost), not some hidden Kafka maximum.
13. Why did Kafka move away from Zookeeper to KRaft?
Zookeeper was an external system Kafka depended on purely to store cluster metadata (broker
list, partition leaders, ACLs) and run leader elections — meaning every Kafka deployment actually
had to run and operate two distributed systems. This added real costs:
- Operational overhead — a second system to deploy, monitor, patch, and scale separately.
- Scalability ceiling — Zookeeper itself became a bottleneck for clusters with very large numbers of partitions, since every metadata change was a round-trip to Zookeeper.
- Slower recovery — controller failover (re-electing which broker manages cluster metadata) was slower when routed through an external system.
KRaft (Kafka Raft) replaces Zookeeper with a built-in consensus protocol (Raft) run by the
Kafka brokers themselves — metadata now lives inside Kafka as just another replicated log. This
means one system instead of two, faster controller failover, and the ability to scale to far
more partitions. KRaft has been production-ready since Kafka 3.x and is mandatory (Zookeeper
support removed) from Kafka 4.0 onward.
14. How to maintain ordering for same Order ID?
Use the Order ID as the message key. Kafka's default partitioner hashes the key to
deterministically route all messages with the same key to the same partition. Since a
partition is a strictly ordered log, all events for that Order ID are processed in order.
⚠️ Ordering is guaranteed WITHIN a partition only — there's no ordering guarantee across
partitions/topics.
15. What is Dead Letter Topic (DLT)?
When a message fails processing after exhausting all configured retries, instead of dropping it
or blocking the partition indefinitely, it's routed to a separate Dead Letter Topic for
manual inspection, alerting, or reprocessing later. This ensures no message is silently lost
and a single poison-pill message doesn't stall the whole consumer.
16. At-Most-Once vs At-Least-Once vs Exactly-Once?
| Semantic | Config | Behavior |
|---|---|---|
| At-Most-Once |
acks=0 + commit offset before processing completes |
May lose messages, never duplicates |
| At-Least-Once |
acks=all + commit offset after processing completes |
Never loses messages, may duplicate |
| Exactly-Once | Idempotent producer + Kafka Transactions + manual offset commit | Never loses, never duplicates |
At-Least-Once is the most common default in practice, paired with idempotent consumer logic
(dedup by message ID) to approximate exactly-once behavior without the full transactional
overhead.
⚠️ Nuance worth mentioning if pressed: with Kafka's default auto-commit (enable.auto.commit=true),
offsets are actually committed on a background timer (auto.commit.interval.ms), not strictly
"right before" or "right after" a given message — so depending on exactly when a crash happens
relative to that timer, auto-commit can land you in either at-most-once or at-least-once
territory. For a predictable guarantee, disable auto-commit and commit manually at the point you
intend (before vs. after processing).
17. Consumer Lag?
The difference between the latest offset (last message produced to a partition) and the
consumer's committed offset (last message it has processed). High lag means the consumer is
falling behind the producer's rate — a key signal for scaling consumers or investigating slow
processing.
Monitor via: Kafka UI / AKHQ, Grafana + Prometheus (via JMX exporter), or the CLI:
kafka-consumer-groups.sh --bootstrap-server <broker> --describe --group <group-id>
18. Kafka vs RabbitMQ?
| Kafka | RabbitMQ | |
|---|---|---|
| Model | Log-based (durable, replayable log) | Queue-based (messages removed once consumed) |
| Throughput | Very high (millions of msgs/sec) | Moderate |
| Use case | Event streaming, event sourcing, analytics pipelines | Task queues, RPC-style messaging, work distribution |
| Ordering | Per-partition ordering | Per-queue ordering |
| Replay | Yes — consumers can re-read from any offset | No — once acked, message is gone |
| Routing | Simple (topic/partition/key) | Rich (exchanges: direct, topic, fanout, headers) |
Rule of thumb: Kafka for high-throughput event streaming and replayable logs; RabbitMQ for
flexible task/work queues with complex routing needs.
19. Real Design: Order → Payment → Inventory
Order Service → order-topic → Payment Service
|
▼
payment-topic
|
▼
Inventory Service
Each service consumes events relevant to it and produces new events downstream — this is
the classic event-driven microservices / choreography pattern:
- Order Service places an order → publishes to
order-topic. - Payment Service consumes
order-topic, processes payment → publishes result topayment-topic. - Inventory Service consumes
payment-topic→ reserves/deducts stock.
Benefits: asynchronous (services don't block on each other), scalable (each service
scales independently by adding consumers/partitions), fault-tolerant (a downstream service
being briefly down doesn't lose events — they sit in the topic until it recovers).
Trade-off to mention in an interview: this requires careful handling of failure/compensation
(e.g., what happens if Inventory can't fulfill after Payment succeeded? → Saga pattern with
compensating transactions) and idempotency at each consumption step (see Q3 and Q16).
20. What is the role of the __consumer_offsets topic?
Kafka stores committed consumer offsets in an internal, compacted topic called
__consumer_offsets (instead of an external store like Zookeeper, as in older Kafka versions).
This lets consumers resume from their last committed position after a restart or rebalance.
21. What is a rebalance, and why can it be disruptive?
A rebalance occurs when consumers join/leave a group (e.g., scaling up, a consumer crash, or
a deployment), triggering redistribution of partitions among the remaining/new consumers. During
a rebalance, consumption pauses ("stop-the-world"), which can hurt throughput/latency if
rebalances happen frequently. Mitigations: Incremental Cooperative Rebalancing (Cooperative assignor) minimizes disruption by only reassigning partitions that need to move, instead
Sticky
of revoking all partitions from all consumers.
22. What is ISR (In-Sync Replica)?
The set of replicas (leader + followers) that are fully caught up with the leader's log within an
allowed lag (replica.lag.time.max.ms). acks=all requires acknowledgment specifically from all
current ISR members — if a replica falls too far behind, it's dropped from the ISR and no longer
blocks writes.
23. What happens if a partition leader dies?
Kafka's controller (elected via KRaft/Zookeeper) detects the failure and promotes one of the
in-sync replicas (ISR) to become the new leader — never an out-of-sync replica, since that
could silently lose committed data (this is the unclean.leader.election.enable setting, which
defaults to false for exactly this reason). Producers/consumers refresh their metadata and
redirect requests to the new leader — this is largely transparent to clients, though it can cause
a brief availability blip during the election.
24. How does Kafka achieve high throughput?
- Sequential disk I/O — append-only log writes are much faster than random I/O.
-
Zero-copy transfer —
sendfile()lets Kafka send data from disk to network socket without copying through user-space memory. - Batching & compression — producers batch multiple messages per request and can compress batches (gzip, snappy, lz4, zstd) to reduce network/disk overhead.
- Partitioning — parallelism across brokers and consumers.
25. What is log compaction?
An alternative retention policy (vs. time/size-based deletion) where Kafka retains only the
latest value for each key in a topic, removing older duplicate-keyed records in the
background. Useful for topics representing "current state" (e.g., a changelog topic backing a
KTable) rather than an event history.
26. What is the difference between a Kafka Producer's linger.ms and batch.size?
Both control batching behavior on the producer:
-
batch.size— max bytes to batch before sending (per partition). -
linger.ms— max time to wait for more messages to accumulate before sending, even if the batch isn't full. Tuning these trades off latency (lower values) against throughput/efficiency (higher values, larger batches).
27. What is Kafka Streams / KSQL used for?
Kafka Streams is a Java library for building stream-processing applications directly on top
of Kafka (stateful transformations, joins, windowed aggregations) without needing a separate
processing cluster (like Spark/Flink). ksqlDB provides a SQL-like interface over the same
capabilities for simpler use cases.
28. How do you handle schema evolution in Kafka messages?
Use a Schema Registry (e.g., Confluent Schema Registry) with Avro/Protobuf/JSON Schema
to enforce and version message schemas. Configure compatibility modes (backward, forward, full)
so producers/consumers can evolve independently without breaking each other — critical in a
microservices setup where teams deploy on different timelines.
29. What is the difference between poll() timeout and max.poll.interval.ms?
-
poll()timeout — how long a single call topoll()blocks waiting for new records if none are immediately available. -
max.poll.interval.ms— the max allowed time between successivepoll()calls before the consumer is considered dead and triggers a rebalance. Important when message processing itself is slow — if processing takes longer than this interval without callingpoll()again, the consumer gets kicked out of the group.
📎 Up next: Kafka Interview Prep Part 2 — Production Scenarios & Real-World Design, covering
the Zomato live-location case study and senior/lead-level production failure scenarios
(duplicate processing, consumer lag diagnosis, multi-event-type topics).
Top comments (0)