Apache Kafka is a popular distributed event streaming platform used by organizations to process and move large amounts of data in real time. It is widely used in applications such as financial transactions, customer activity tracking, log collection, monitoring systems, recommendation engines, and data pipelines.
For beginners, Kafka can initially seem complicated because it introduces several concepts that work together. Three of the most important are topics, partitions, and offsets.
Understanding these concepts is essential because they form the foundation of how Kafka stores, organizes, and delivers messages.
A simple way to think about them is:
A topic is where messages are categorized, a partition is where those messages are stored in an ordered sequence, and an offset identifies a message's position within a partition.
This article explains these concepts step by step and shows how they work together.
1. What Is Apache Kafka?
Apache Kafka is a distributed event streaming platform designed to handle streams of data efficiently.
Instead of applications communicating directly with each other every time data needs to be transferred, Kafka can act as an intermediary.
For example, imagine an e-commerce application.
When a customer places an order, several systems may need to know about it:
- The payment system needs to process payment.
- The inventory system needs to update stock.
- The notification system needs to send a confirmation.
- The analytics system may need to record the transaction.
- The shipping system needs to prepare the order.
Rather than sending the same information separately to every system, the application can publish an event to Kafka.
Other applications can then consume that event.
A simplified flow looks like this:
Producer
|
| Order Event
↓
Apache Kafka
|
├── Payment Service
├── Inventory Service
├── Notification Service
├── Analytics Service
└── Shipping Service
This architecture makes it easier to build scalable, loosely coupled systems.
2. What Is a Kafka Topic?
A topic is a named category or stream to which messages are published.
Think of a topic as a logical container for related events.
For example, an organization might create topics such as:
orders
payments
customer-events
website-clicks
transactions
system-logs
If an application generates a new order event, it might publish that event to the orders topic.
For example:
{
"order_id": 1001,
"customer": "Alice",
"amount": 7500
}
Another order might look like:
{
"order_id": 1002,
"customer": "Brian",
"amount": 4200
}
Both messages could be published to the orders topic.
The important thing to understand is that a topic does not represent one individual message.
It represents a stream or category of related messages.
3. A Topic Is Not the Actual Storage Sequence
One common beginner mistake is imagining a Kafka topic as a single list of messages.
In reality, a topic is divided into partitions.
For example:
orders
|
├── Partition 0
├── Partition 1
└── Partition 2
Each partition is an ordered sequence of records.
This partitioning system is one of the reasons Kafka can handle very large workloads.
4. What Is a Kafka Partition?
A partition is an ordered, append-only sequence of records within a Kafka topic.
Suppose our orders topic has three partitions:
orders
├── Partition 0
├── Partition 1
└── Partition 2
Messages can be distributed across these partitions.
For example:
Partition 0:
Message A
Message D
Message G
Partition 1:
Message B
Message E
Message H
Partition 2:
Message C
Message F
Message I
Each partition maintains its own ordering.
This distinction is extremely important.
Kafka guarantees ordering within a partition, not necessarily across the entire topic.
5. Why Does Kafka Use Partitions?
Partitions provide several important benefits.
*Scalability
*
Large topics can be divided across multiple partitions.
Instead of one server handling every message, Kafka can distribute the workload across multiple brokers.
Topic
|
-----------------
| | |
P0 P1 P2
| | |
Broker A Broker B Broker C
This allows Kafka to process more data concurrently.
Parallel Processing
Different consumers can process different partitions simultaneously.
For example:
Consumer 1 → Partition 0
Consumer 2 → Partition 1
Consumer 3 → Partition 2
This makes it possible to scale message processing horizontally.
Ordering
Each partition maintains the order of its records.
For example:
Partition 0
Message 1
Message 2
Message 3
Message 4
Kafka preserves this order within the partition.
6. What Is an Offset?
An offset is a unique sequential position assigned to a record within a Kafka partition.
For example:
Partition 0
Offset 0 → Order A
Offset 1 → Order B
Offset 2 → Order C
Offset 3 → Order D
Offset 4 → Order E
The offset tells Kafka where a particular record is located within that partition.
Offsets begin at 0 for a partition and increase as records are appended.
7. Offsets Are Partition-Specific
This is another concept beginners frequently misunderstand.
An offset does not uniquely identify a message across an entire Kafka topic.
It identifies a message's position within a particular partition.
For example:
Partition 0:
Offset 0 → Message A
Offset 1 → Message B
Offset 2 → Message C
Partition 1:
Offset 0 → Message D
Offset 1 → Message E
Offset 2 → Message F
Notice that both Partition 0 and Partition 1 have an offset 0.
Therefore:
Partition + Offset
together identify a specific record.
For example:
Partition 1 + Offset 2
identifies Message F.
8. How Topics, Partitions, and Offsets Work Together
Let's put everything together.
Suppose we have an orders topic with two partitions:
Topic: orders
Partition 0
-------------------------
Offset 0 → Order 1001
Offset 1 → Order 1003
Offset 2 → Order 1005
Partition 1
-------------------------
Offset 0 → Order 1002
Offset 1 → Order 1004
Offset 2 → Order 1006
Here:
-
ordersis the topic - Partition 0 and Partition 1 are the partitions
- 0, 1, and 2 are the offsets
- Each record has an ordered position within its partition
A useful hierarchy is:
Kafka
↓
Topic
↓
Partitions
↓
Records
↓
Offsets
*9. How Does Kafka Decide Which Partition Gets a Message?
*
When a producer sends a message to a topic, Kafka needs to determine which partition should receive it.
One common way is through a message key.
For example, suppose we have:
Customer ID: 1001
The producer can use the customer ID as the key.
Kafka can then use the key to determine the partition.
Conceptually:
Customer ID
↓
Partitioning logic
↓
Partition 1
Messages with the same key are generally sent to the same partition, assuming the partition count and partitioning strategy remain unchanged.
This can be useful when ordering matters.
For example, suppose a customer performs these actions:
Order Created
Payment Completed
Order Shipped
Order Delivered
If these events use the same key and are placed in the same partition, Kafka can preserve their order within that partition.
10. What Happens If No Key Is Provided?
If a producer does not provide a key, Kafka can distribute records across partitions according to its producer partitioning behavior.
The important lesson for beginners is:
A message does not automatically belong to the entire topic as one sequential stream. It is stored in one of the topic's partitions.
Therefore, when thinking about ordering, always ask:
Which partition is this message in?
11. Producers and Partitions
A producer is an application that sends records to Kafka.
For example:
E-commerce Application
|
| Order Event
↓
Producer
|
↓
orders topic
The producer determines the topic and, directly or indirectly, the partition where the record will be written.
For example:
Producer
|
├── Order A → Partition 0
├── Order B → Partition 1
├── Order C → Partition 0
└── Order D → Partition 1
The producer does not need to know everything about how consumers will process the data.
12. Consumers and Partitions
A consumer is an application that reads records from Kafka.
For example:
orders topic
|
├── Partition 0 → Consumer
└── Partition 1 → Consumer
Consumers track their progress through partitions using offsets.
Suppose a consumer has processed:
Partition 0
Offset 0
Offset 1
Offset 2
Its next record may be:
Offset 3
This allows the consumer to continue processing from where it stopped.
13. What Are Consumer Groups?
Kafka introduces another important concept called a consumer group.
A consumer group is a collection of consumers working together to consume messages from a topic.
Suppose a topic has three partitions:
Topic
├── Partition 0
├── Partition 1
└── Partition 2
A consumer group might contain three consumers:
Consumer 1 → Partition 0
Consumer 2 → Partition 1
Consumer 3 → Partition 2
This allows the work to be distributed across consumers.
A key rule is:
Within a consumer group, a partition is assigned to only one consumer at a time.
This enables parallel processing while maintaining ordering within each partition.
*14. What If There Are More Consumers Than Partitions?
*
Suppose there are only two partitions:
Partition 0
Partition 1
But the consumer group contains four consumers:
Consumer 1
Consumer 2
Consumer 3
Consumer 4
Only two consumers can actively consume partitions at a given time:
Consumer 1 → Partition 0
Consumer 2 → Partition 1
Consumer 3 → No partition
Consumer 4 → No partition
The additional consumers remain idle until partition assignments change.
This is why the number of partitions is an important consideration when designing Kafka systems for parallel processing.
15. What Happens When a Consumer Fails?
Kafka can redistribute partitions when a consumer in a consumer group fails.
For example:
Before failure:
Consumer 1 → Partition 0
Consumer 2 → Partition 1
Consumer 3 → Partition 2
If Consumer 2 fails:
Consumer 1 → Partition 0
Consumer 3 → Partition 1
Consumer 3 → Partition 2
The exact assignment depends on Kafka's group coordination and partition assignment process.
Because consumers track offsets, the replacement consumer can continue from the appropriate position.
16. Why Are Offsets Important?
Offsets allow consumers to keep track of their progress.
Imagine a consumer processes these records:
Offset 0 → Processed
Offset 1 → Processed
Offset 2 → Processed
Offset 3 → Processed
The consumer's progress can be tracked so that if it restarts, it can continue from the appropriate position rather than automatically starting from the beginning.
This is one reason Kafka is useful for reliable event processing.
17. Kafka Does Not Delete a Message Immediately After Consumption
Another common beginner misconception is:
"Once a consumer reads a message, Kafka deletes it."
That is generally not how Kafka works.
Kafka retains records according to the topic's configured retention policies.
A consumer reading a record does not automatically remove that record from the partition.
This means multiple consumer groups can independently read the same topic.
For example:
orders topic
|
├── Analytics Consumer Group
|
├── Payment Consumer Group
|
└── Notification Consumer Group
Each group can maintain its own position in the topic.
18. Consumer Groups Have Their Own Progress
Suppose two consumer groups read the same partition.
Partition 0
Offset 0
Offset 1
Offset 2
Offset 3
Offset 4
The analytics group might have processed up to:
Offset 4
while another group might only have processed:
Offset 2
They can maintain different consumption progress independently.
This is another important reason Kafka can support multiple applications consuming the same stream.
19. Ordering in Kafka
Kafka's ordering guarantee is frequently misunderstood.
Kafka guarantees ordering within a partition.
For example:
Partition 0
Offset 0 → Event A
Offset 1 → Event B
Offset 2 → Event C
Offset 3 → Event D
The order is:
A → B → C → D
However, suppose the topic has two partitions:
Partition 0:
A → B → C
Partition 1:
D → E → F
Kafka does not provide a single global ordering such as:
A → B → C → D → E → F
across the entire topic.
Therefore, if strict ordering is required for related events, partitioning strategy becomes very important.
20. A Real-World Example: Banking Transactions
Imagine a banking application that publishes transaction events.
The topic could be:
bank-transactions
The topic might contain several partitions:
bank-transactions
├── Partition 0
├── Partition 1
├── Partition 2
└── Partition 3
Suppose a customer's account number is used as the message key.
Transactions for a particular account can then consistently map to the same partition under the partitioning strategy.
For example:
Account 1001
↓
Partition 2
↓
Offset 50 → Deposit
Offset 51 → Withdrawal
Offset 52 → Transfer
Offset 53 → Deposit
The offset allows consumers to identify the position of each transaction, while the partition preserves the order of those records within that partition.
21. Topic vs. Partition vs. Offset
Let's simplify everything into one comparison.
| Concept | What It Means | Simple Analogy |
|---|---|---|
| Topic | A category or stream of related events | A book |
| Partition | An ordered sequence within a topic | A chapter |
| Record | An individual event/message | A sentence |
| Offset | The record's position within a partition | A line number |
The analogy is not perfect, but it can make the concepts easier to visualize.
Another useful analogy is a supermarket:
Topic = Department
Partition = Aisle
Record = Product
Offset = Position of the product in the aisle
The important thing is to understand the actual Kafka behavior rather than relying entirely on analogies.
22. Common Beginner Confusions
Confusion 1: "A topic is a single queue."
Not exactly.
A topic is divided into partitions, and each partition is an ordered log.
Confusion 2: "Offsets are unique across a topic."
No.
Offsets are specific to individual partitions.
For example:
Partition 0 → Offset 10
Partition 1 → Offset 10
Both can exist at the same time.
Confusion 3: "Kafka guarantees order across the whole topic."
No.
Ordering is guaranteed within individual partitions.
Confusion 4: "Consumers delete messages."
Reading a message does not automatically delete it.
Kafka retains records according to configured retention policies.
Confusion 5: "More consumers always mean more processing power."
Not necessarily.
A consumer group cannot have more active consumers for a topic than there are partitions.
For example:
3 partitions
5 consumers
At most three consumers can actively consume those partitions at a given time within that group.
Confusion 6: "A partition is the same thing as a consumer."
No.
A partition is a storage/log unit within a topic.
A consumer is an application process that reads records.
They are different concepts.
23. Why These Concepts Matter to Data Engineers
Understanding topics, partitions, and offsets is particularly important for data engineers because Kafka is commonly used in modern data architectures.
For example:
Applications
|
↓
Kafka
|
├── Data Warehouse
├── Data Lake
├── Analytics Platform
├── Monitoring System
└── Machine Learning Pipeline
Kafka can act as a central event-streaming layer connecting applications and data systems.
Data engineers need to understand partitions to design scalable pipelines, offsets to manage processing progress, and topics to organize streams of related events.
24. Key Takeaways
The most important concepts to remember are:
A topic is a named stream or category of events.
A topic is divided into one or more partitions.
A partition is an ordered, append-only sequence of records.
An offset identifies the position of a record within a partition.
Offsets are partition-specific.
Kafka guarantees ordering within a partition, not across an entire topic.
Producers write records to topics.
Consumers read records from partitions.
Consumer groups allow multiple consumers to divide partition-processing work.
Messages are retained according to Kafka's retention configuration rather than being deleted simply because they were consumed.
The number of partitions affects the potential parallelism of consumers within a consumer group.
Partitioning strategy matters when the order of related events is important.
Kafka's topics, partitions, and offsets can seem confusing at first, but they become much easier to understand once their relationships are clear.
A topic organizes related events. A partition divides that topic into ordered sequences that Kafka can distribute and process in parallel. An offset identifies the position of a record within a particular partition.
The simplest way to remember the relationship is:
Topic
↓
Partitions
↓
Records
↓
Offsets
Or, in one sentence:
A topic contains partitions, partitions contain ordered records, and offsets identify the positions of those records.
Once these three concepts are understood, other Kafka concepts—including producers, consumers, consumer groups, replication, rebalancing, and event processing—become much easier to learn.
Top comments (0)