DEV Community

Venus-Kennedy
Venus-Kennedy

Posted on

Kafka Topics, Partitions, and Offsets Explained: The Concepts Every Beginner Confuses

Apache Kafka is a popular distributed event streaming platform used by organizations to process and move large amounts of data in real time. It is widely used in applications such as financial transactions, customer activity tracking, log collection, monitoring systems, recommendation engines, and data pipelines.

For beginners, Kafka can initially seem complicated because it introduces several concepts that work together. Three of the most important are topics, partitions, and offsets.

Understanding these concepts is essential because they form the foundation of how Kafka stores, organizes, and delivers messages.

A simple way to think about them is:

A topic is where messages are categorized, a partition is where those messages are stored in an ordered sequence, and an offset identifies a message's position within a partition.

This article explains these concepts step by step and shows how they work together.

1. What Is Apache Kafka?

Apache Kafka is a distributed event streaming platform designed to handle streams of data efficiently.

Instead of applications communicating directly with each other every time data needs to be transferred, Kafka can act as an intermediary.

For example, imagine an e-commerce application.

When a customer places an order, several systems may need to know about it:

  • The payment system needs to process payment.
  • The inventory system needs to update stock.
  • The notification system needs to send a confirmation.
  • The analytics system may need to record the transaction.
  • The shipping system needs to prepare the order.

Rather than sending the same information separately to every system, the application can publish an event to Kafka.

Other applications can then consume that event.

A simplified flow looks like this:

Producer
   |
   | Order Event
   ↓
 Apache Kafka
   |
   ├── Payment Service
   ├── Inventory Service
   ├── Notification Service
   ├── Analytics Service
   └── Shipping Service
Enter fullscreen mode Exit fullscreen mode

This architecture makes it easier to build scalable, loosely coupled systems.

2. What Is a Kafka Topic?

A topic is a named category or stream to which messages are published.

Think of a topic as a logical container for related events.

For example, an organization might create topics such as:

orders
payments
customer-events
website-clicks
transactions
system-logs
Enter fullscreen mode Exit fullscreen mode

If an application generates a new order event, it might publish that event to the orders topic.

For example:

{
  "order_id": 1001,
  "customer": "Alice",
  "amount": 7500
}
Enter fullscreen mode Exit fullscreen mode

Another order might look like:

{
  "order_id": 1002,
  "customer": "Brian",
  "amount": 4200
}
Enter fullscreen mode Exit fullscreen mode

Both messages could be published to the orders topic.

The important thing to understand is that a topic does not represent one individual message.

It represents a stream or category of related messages.

3. A Topic Is Not the Actual Storage Sequence

One common beginner mistake is imagining a Kafka topic as a single list of messages.

In reality, a topic is divided into partitions.

For example:

orders
   |
   ├── Partition 0
   ├── Partition 1
   └── Partition 2
Enter fullscreen mode Exit fullscreen mode

Each partition is an ordered sequence of records.

This partitioning system is one of the reasons Kafka can handle very large workloads.

4. What Is a Kafka Partition?

A partition is an ordered, append-only sequence of records within a Kafka topic.

Suppose our orders topic has three partitions:

orders
 ├── Partition 0
 ├── Partition 1
 └── Partition 2
Enter fullscreen mode Exit fullscreen mode

Messages can be distributed across these partitions.

For example:

Partition 0:
Message A
Message D
Message G

Partition 1:
Message B
Message E
Message H

Partition 2:
Message C
Message F
Message I
Enter fullscreen mode Exit fullscreen mode

Each partition maintains its own ordering.

This distinction is extremely important.

Kafka guarantees ordering within a partition, not necessarily across the entire topic.

5. Why Does Kafka Use Partitions?

Partitions provide several important benefits.

*Scalability
*

Large topics can be divided across multiple partitions.

Instead of one server handling every message, Kafka can distribute the workload across multiple brokers.

             Topic
               |
       -----------------
       |       |       |
       P0      P1      P2
       |       |       |
    Broker A Broker B Broker C
Enter fullscreen mode Exit fullscreen mode

This allows Kafka to process more data concurrently.

Parallel Processing

Different consumers can process different partitions simultaneously.

For example:

Consumer 1 → Partition 0
Consumer 2 → Partition 1
Consumer 3 → Partition 2
Enter fullscreen mode Exit fullscreen mode

This makes it possible to scale message processing horizontally.

Ordering

Each partition maintains the order of its records.

For example:

Partition 0

Message 1
Message 2
Message 3
Message 4
Enter fullscreen mode Exit fullscreen mode

Kafka preserves this order within the partition.

6. What Is an Offset?

An offset is a unique sequential position assigned to a record within a Kafka partition.

For example:

Partition 0

Offset 0 → Order A
Offset 1 → Order B
Offset 2 → Order C
Offset 3 → Order D
Offset 4 → Order E
Enter fullscreen mode Exit fullscreen mode

The offset tells Kafka where a particular record is located within that partition.

Offsets begin at 0 for a partition and increase as records are appended.

7. Offsets Are Partition-Specific

This is another concept beginners frequently misunderstand.

An offset does not uniquely identify a message across an entire Kafka topic.

It identifies a message's position within a particular partition.

For example:

Partition 0:
Offset 0 → Message A
Offset 1 → Message B
Offset 2 → Message C

Partition 1:
Offset 0 → Message D
Offset 1 → Message E
Offset 2 → Message F
Enter fullscreen mode Exit fullscreen mode

Notice that both Partition 0 and Partition 1 have an offset 0.

Therefore:

Partition + Offset
Enter fullscreen mode Exit fullscreen mode

together identify a specific record.

For example:

Partition 1 + Offset 2
Enter fullscreen mode Exit fullscreen mode

identifies Message F.


8. How Topics, Partitions, and Offsets Work Together

Let's put everything together.

Suppose we have an orders topic with two partitions:

Topic: orders

Partition 0
-------------------------
Offset 0 → Order 1001
Offset 1 → Order 1003
Offset 2 → Order 1005


Partition 1
-------------------------
Offset 0 → Order 1002
Offset 1 → Order 1004
Offset 2 → Order 1006
Enter fullscreen mode Exit fullscreen mode

Here:

  • orders is the topic
  • Partition 0 and Partition 1 are the partitions
  • 0, 1, and 2 are the offsets
  • Each record has an ordered position within its partition

A useful hierarchy is:

Kafka
  ↓
Topic
  ↓
Partitions
  ↓
Records
  ↓
Offsets
Enter fullscreen mode Exit fullscreen mode

*9. How Does Kafka Decide Which Partition Gets a Message?
*

When a producer sends a message to a topic, Kafka needs to determine which partition should receive it.

One common way is through a message key.

For example, suppose we have:

Customer ID: 1001
Enter fullscreen mode Exit fullscreen mode

The producer can use the customer ID as the key.

Kafka can then use the key to determine the partition.

Conceptually:

Customer ID
     ↓
Partitioning logic
     ↓
Partition 1
Enter fullscreen mode Exit fullscreen mode

Messages with the same key are generally sent to the same partition, assuming the partition count and partitioning strategy remain unchanged.

This can be useful when ordering matters.

For example, suppose a customer performs these actions:

Order Created
Payment Completed
Order Shipped
Order Delivered
Enter fullscreen mode Exit fullscreen mode

If these events use the same key and are placed in the same partition, Kafka can preserve their order within that partition.

10. What Happens If No Key Is Provided?

If a producer does not provide a key, Kafka can distribute records across partitions according to its producer partitioning behavior.

The important lesson for beginners is:

A message does not automatically belong to the entire topic as one sequential stream. It is stored in one of the topic's partitions.

Therefore, when thinking about ordering, always ask:

Which partition is this message in?

11. Producers and Partitions

A producer is an application that sends records to Kafka.

For example:

E-commerce Application
        |
        | Order Event
        ↓
     Producer
        |
        ↓
    orders topic
Enter fullscreen mode Exit fullscreen mode

The producer determines the topic and, directly or indirectly, the partition where the record will be written.

For example:

Producer
   |
   ├── Order A → Partition 0
   ├── Order B → Partition 1
   ├── Order C → Partition 0
   └── Order D → Partition 1
Enter fullscreen mode Exit fullscreen mode

The producer does not need to know everything about how consumers will process the data.

12. Consumers and Partitions

A consumer is an application that reads records from Kafka.

For example:

orders topic
     |
     ├── Partition 0 → Consumer
     └── Partition 1 → Consumer
Enter fullscreen mode Exit fullscreen mode

Consumers track their progress through partitions using offsets.

Suppose a consumer has processed:

Partition 0
Offset 0
Offset 1
Offset 2
Enter fullscreen mode Exit fullscreen mode

Its next record may be:

Offset 3
Enter fullscreen mode Exit fullscreen mode

This allows the consumer to continue processing from where it stopped.

13. What Are Consumer Groups?

Kafka introduces another important concept called a consumer group.

A consumer group is a collection of consumers working together to consume messages from a topic.

Suppose a topic has three partitions:

Topic
 ├── Partition 0
 ├── Partition 1
 └── Partition 2
Enter fullscreen mode Exit fullscreen mode

A consumer group might contain three consumers:

Consumer 1 → Partition 0
Consumer 2 → Partition 1
Consumer 3 → Partition 2
Enter fullscreen mode Exit fullscreen mode

This allows the work to be distributed across consumers.

A key rule is:

Within a consumer group, a partition is assigned to only one consumer at a time.

This enables parallel processing while maintaining ordering within each partition.

*14. What If There Are More Consumers Than Partitions?
*

Suppose there are only two partitions:

Partition 0
Partition 1
Enter fullscreen mode Exit fullscreen mode

But the consumer group contains four consumers:

Consumer 1
Consumer 2
Consumer 3
Consumer 4
Enter fullscreen mode Exit fullscreen mode

Only two consumers can actively consume partitions at a given time:

Consumer 1 → Partition 0
Consumer 2 → Partition 1

Consumer 3 → No partition
Consumer 4 → No partition
Enter fullscreen mode Exit fullscreen mode

The additional consumers remain idle until partition assignments change.

This is why the number of partitions is an important consideration when designing Kafka systems for parallel processing.

15. What Happens When a Consumer Fails?

Kafka can redistribute partitions when a consumer in a consumer group fails.

For example:

Before failure:

Consumer 1 → Partition 0
Consumer 2 → Partition 1
Consumer 3 → Partition 2
Enter fullscreen mode Exit fullscreen mode

If Consumer 2 fails:

Consumer 1 → Partition 0
Consumer 3 → Partition 1
Consumer 3 → Partition 2
Enter fullscreen mode Exit fullscreen mode

The exact assignment depends on Kafka's group coordination and partition assignment process.

Because consumers track offsets, the replacement consumer can continue from the appropriate position.

16. Why Are Offsets Important?

Offsets allow consumers to keep track of their progress.

Imagine a consumer processes these records:

Offset 0 → Processed
Offset 1 → Processed
Offset 2 → Processed
Offset 3 → Processed
Enter fullscreen mode Exit fullscreen mode

The consumer's progress can be tracked so that if it restarts, it can continue from the appropriate position rather than automatically starting from the beginning.

This is one reason Kafka is useful for reliable event processing.

17. Kafka Does Not Delete a Message Immediately After Consumption

Another common beginner misconception is:

"Once a consumer reads a message, Kafka deletes it."

That is generally not how Kafka works.

Kafka retains records according to the topic's configured retention policies.

A consumer reading a record does not automatically remove that record from the partition.

This means multiple consumer groups can independently read the same topic.

For example:

orders topic
     |
     ├── Analytics Consumer Group
     |
     ├── Payment Consumer Group
     |
     └── Notification Consumer Group
Enter fullscreen mode Exit fullscreen mode

Each group can maintain its own position in the topic.

18. Consumer Groups Have Their Own Progress

Suppose two consumer groups read the same partition.

Partition 0

Offset 0
Offset 1
Offset 2
Offset 3
Offset 4
Enter fullscreen mode Exit fullscreen mode

The analytics group might have processed up to:

Offset 4
Enter fullscreen mode Exit fullscreen mode

while another group might only have processed:

Offset 2
Enter fullscreen mode Exit fullscreen mode

They can maintain different consumption progress independently.

This is another important reason Kafka can support multiple applications consuming the same stream.

19. Ordering in Kafka

Kafka's ordering guarantee is frequently misunderstood.

Kafka guarantees ordering within a partition.

For example:

Partition 0

Offset 0 → Event A
Offset 1 → Event B
Offset 2 → Event C
Offset 3 → Event D
Enter fullscreen mode Exit fullscreen mode

The order is:

A → B → C → D
Enter fullscreen mode Exit fullscreen mode

However, suppose the topic has two partitions:

Partition 0:
A → B → C

Partition 1:
D → E → F
Enter fullscreen mode Exit fullscreen mode

Kafka does not provide a single global ordering such as:

A → B → C → D → E → F
Enter fullscreen mode Exit fullscreen mode

across the entire topic.

Therefore, if strict ordering is required for related events, partitioning strategy becomes very important.

20. A Real-World Example: Banking Transactions

Imagine a banking application that publishes transaction events.

The topic could be:

bank-transactions
Enter fullscreen mode Exit fullscreen mode

The topic might contain several partitions:

bank-transactions
 ├── Partition 0
 ├── Partition 1
 ├── Partition 2
 └── Partition 3
Enter fullscreen mode Exit fullscreen mode

Suppose a customer's account number is used as the message key.

Transactions for a particular account can then consistently map to the same partition under the partitioning strategy.

For example:

Account 1001
     ↓
Partition 2
     ↓
Offset 50 → Deposit
Offset 51 → Withdrawal
Offset 52 → Transfer
Offset 53 → Deposit
Enter fullscreen mode Exit fullscreen mode

The offset allows consumers to identify the position of each transaction, while the partition preserves the order of those records within that partition.

21. Topic vs. Partition vs. Offset

Let's simplify everything into one comparison.

Concept What It Means Simple Analogy
Topic A category or stream of related events A book
Partition An ordered sequence within a topic A chapter
Record An individual event/message A sentence
Offset The record's position within a partition A line number

The analogy is not perfect, but it can make the concepts easier to visualize.

Another useful analogy is a supermarket:

Topic = Department
Partition = Aisle
Record = Product
Offset = Position of the product in the aisle
Enter fullscreen mode Exit fullscreen mode

The important thing is to understand the actual Kafka behavior rather than relying entirely on analogies.

22. Common Beginner Confusions

Confusion 1: "A topic is a single queue."

Not exactly.

A topic is divided into partitions, and each partition is an ordered log.

Confusion 2: "Offsets are unique across a topic."

No.

Offsets are specific to individual partitions.

For example:

Partition 0 → Offset 10
Partition 1 → Offset 10
Enter fullscreen mode Exit fullscreen mode

Both can exist at the same time.

Confusion 3: "Kafka guarantees order across the whole topic."

No.

Ordering is guaranteed within individual partitions.

Confusion 4: "Consumers delete messages."

Reading a message does not automatically delete it.

Kafka retains records according to configured retention policies.

Confusion 5: "More consumers always mean more processing power."

Not necessarily.

A consumer group cannot have more active consumers for a topic than there are partitions.

For example:

3 partitions
5 consumers
Enter fullscreen mode Exit fullscreen mode

At most three consumers can actively consume those partitions at a given time within that group.

Confusion 6: "A partition is the same thing as a consumer."

No.

A partition is a storage/log unit within a topic.

A consumer is an application process that reads records.

They are different concepts.

23. Why These Concepts Matter to Data Engineers

Understanding topics, partitions, and offsets is particularly important for data engineers because Kafka is commonly used in modern data architectures.

For example:

Applications
     |
     ↓
   Kafka
     |
     ├── Data Warehouse
     ├── Data Lake
     ├── Analytics Platform
     ├── Monitoring System
     └── Machine Learning Pipeline
Enter fullscreen mode Exit fullscreen mode

Kafka can act as a central event-streaming layer connecting applications and data systems.

Data engineers need to understand partitions to design scalable pipelines, offsets to manage processing progress, and topics to organize streams of related events.

24. Key Takeaways

The most important concepts to remember are:

  1. A topic is a named stream or category of events.

  2. A topic is divided into one or more partitions.

  3. A partition is an ordered, append-only sequence of records.

  4. An offset identifies the position of a record within a partition.

  5. Offsets are partition-specific.

  6. Kafka guarantees ordering within a partition, not across an entire topic.

  7. Producers write records to topics.

  8. Consumers read records from partitions.

  9. Consumer groups allow multiple consumers to divide partition-processing work.

  10. Messages are retained according to Kafka's retention configuration rather than being deleted simply because they were consumed.

  11. The number of partitions affects the potential parallelism of consumers within a consumer group.

  12. Partitioning strategy matters when the order of related events is important.

Kafka's topics, partitions, and offsets can seem confusing at first, but they become much easier to understand once their relationships are clear.

A topic organizes related events. A partition divides that topic into ordered sequences that Kafka can distribute and process in parallel. An offset identifies the position of a record within a particular partition.

The simplest way to remember the relationship is:

Topic
  ↓
Partitions
  ↓
Records
  ↓
Offsets
Enter fullscreen mode Exit fullscreen mode

Or, in one sentence:

A topic contains partitions, partitions contain ordered records, and offsets identify the positions of those records.

Once these three concepts are understood, other Kafka concepts—including producers, consumers, consumer groups, replication, rebalancing, and event processing—become much easier to learn.

Top comments (0)