DEV Community

Cover image for What Happens After the Commit?
Rodrigo de Oliveira
Rodrigo de Oliveira

Posted on

What Happens After the Commit?

Saving a record to a database is relatively easy.

We have mature tools for that problem. We open a transaction, change some state, commit it, and roll everything back if something fails before the commit.

Things become much more interesting when the database commit is not the end of the operation.

Imagine a financial transaction. Once it is persisted, another service needs to know about it so it can update a balance projection. Maybe additional consumers will react to the same event later.

The first implementation that comes to mind seems perfectly reasonable:

save to database
publish event
Enter fullscreen mode Exit fullscreen mode

But there is an uncomfortable question between those two lines:

What if the database commits successfully and the message broker becomes unavailable immediately afterward?

The transaction exists.

The event does not.

And now two parts of the system are telling different stories.

This is one of the problems I wanted to explore in poc-arquitetura, a repository I use as an executable software architecture laboratory.

The project uses .NET, PostgreSQL, Kafka, Keycloak, OpenTelemetry, k6, and a few other tools. But the goal has never been to collect technologies or architectural patterns.

What interests me is what happens when these patterns have to work together.

Outbox solves one problem and introduces new responsibilities. At-least-once delivery forces us to think about idempotency. Eventual consistency changes how we reason about reads. A Saga requires compensation. Retry needs boundaries. Messaging requires observability beyond HTTP.

That simple question — what happens after the commit? — opens a much broader discussion about distributed systems.

GitHub logo rodri-oliveira-dev / poc-arquitetura

.NET architecture PoC for distributed services with CQRS, Kafka, PostgreSQL, Outbox, DLQ, observability and CI quality gates.

poc-arquitetura

Build Tests Quality Gate Status Security Rating Architecture Docs

POC educacional de microserviços em .NET para estudar arquitetura de software com código real: Clean Architecture, DDD, PostgreSQL, Kafka, Outbox, Inbox, JWT/JWKS com Keycloak, observabilidade, segurança, contratos e testes automatizados.

Ela demonstra um problema comum em sistemas financeiros: registrar fatos de forma transacional, publicar eventos com confiabilidade, projetar saldos em outro serviço e operar falhas sem esconder consistência eventual. O repositório também mostra contextos de identidade, transferência, pagamento externo e auditoria funcional para exercitar trade-offs de integração.

Este projeto é útil para:

  • quem está aprendendo arquitetura e quer ver os conceitos aplicados;
  • desenvolvedores .NET que querem executar, testar e alterar uma stack local;
  • arquitetos que querem avaliar decisões, limites e riscos;
  • avaliadores técnicos que querem entender a proposta rapidamente.

Visão geral

flowchart LR
    Client[Cliente ou teste] --> Keycloak[Keycloak OIDC]
    Client --> LedgerApi[LedgerService.Api]
    Client --> BalanceApi[BalanceService.Api]
    Client --> TransferApi[TransferService.Api]
    Client --> PaymentApi[PaymentService.Api]
    Client --> IdentityApi[IdentityService.Api]
    Client --> AuditApi[AuditService.Api]
    LedgerApi -->
…

Start by separating facts from projections

Two bounded contexts are especially important in this example.

LedgerService owns financial facts.

If a financial transaction happened, the Ledger is where that fact should be recorded.

BalanceService has a different responsibility. It maintains a projection optimized for balance queries.

At a high level, the architecture looks like this:

flowchart LR
    Client[Client] --> LedgerApi[Ledger API]

    LedgerApi --> LedgerDb[(PostgreSQL<br/>Ledger)]

    LedgerDb --> LedgerWorker[Ledger Worker]
    LedgerWorker --> Kafka[(Kafka)]

    Kafka --> BalanceWorker[Balance Worker]
    BalanceWorker --> BalanceDb[(PostgreSQL<br/>Balance)]

    Client --> BalanceApi[Balance API]
    BalanceApi --> BalanceDb

A transaction and a balance are not the same thing.

The transaction is a fact: something happened.

The balance is a model derived from those facts.

Could everything live in a single application and database? Absolutely. Depending on the system, that may even be the better architecture.

But this laboratory deliberately separates those responsibilities so that the consequences of the decision become visible.

One consequence appears immediately.

The balance is no longer updated inside the same database transaction as the Ledger.

We have entered the world of eventual consistency.

That does not mean the system is simply inconsistent or incorrect. It means there is a period during which the Ledger already knows about a transaction while the Balance projection has not processed it yet.

That window is not an accident.

It is part of the architecture.


The real problem is not Kafka. It is the dual write

Suppose the Ledger does something like this:

sequenceDiagram
    participant C as Client
    participant L as Ledger API
    participant DB as PostgreSQL
    participant K as Kafka

    C->>L: Create transaction
    L->>DB: INSERT LedgerEntry
    DB-->>L: COMMIT
    L->>K: Publish event
    K-->>L: Acknowledged
    L-->>C: 201 Created

    Note over L,K: What happens if the process<br/>fails after COMMIT but before publish?

When everything works, this looks fine.

Now imagine that the database commit succeeds but the process crashes before Kafka acknowledges the event.

The Ledger contains the transaction.

The event may never be published.

Publishing first and saving afterward does not solve the problem either. It simply reverses it: we can now publish an event for a transaction that later fails to commit.

This is the dual-write problem.

We need to update two independent resources — the database and the broker — and would like them to behave as if they belonged to one atomic transaction.

One option would be some form of distributed transaction.

For this project, I chose a different approach: accept that PostgreSQL and Kafka have independent lifecycles and make the intent to publish durable.

That is where the Transactional Outbox pattern becomes useful.


Outbox: do not publish now, persist the intent to publish

Instead of saving a transaction and immediately depending on Kafka, the Ledger stores both the financial fact and an Outbox message inside the same PostgreSQL transaction.

sequenceDiagram
    participant C as Client
    participant L as Ledger API
    participant DB as PostgreSQL
    participant W as Ledger Worker
    participant K as Kafka

    C->>L: Create transaction

    rect rgb(235, 235, 235)
        L->>DB: BEGIN
        L->>DB: INSERT LedgerEntry
        L->>DB: INSERT OutboxMessage
        L->>DB: COMMIT
    end

    L-->>C: Transaction confirmed

    W->>DB: Fetch pending messages
    DB-->>W: OutboxMessage
    W->>K: Publish LedgerEntryCreated
    K-->>W: Acknowledged
    W->>DB: Mark as processed

Either both records are persisted or neither is.

The HTTP request no longer needs Kafka to be available in order to preserve the intent to publish the event.

A separate worker polls pending Outbox messages and publishes them.

After Kafka confirms publication, the worker updates the Outbox entry.

If Kafka is unavailable, the financial transaction still exists together with durable information that an event still needs to be published.

If the worker crashes, processing can resume later.

This small architectural change has an important effect.

The failure is no longer a tiny invisible window between two independent writes.

It becomes state that the system can observe and recover from.

That, to me, is one of the most useful ways to understand Outbox.

It does not magically turn Kafka and PostgreSQL into a single transaction.

It does not guarantee exactly-once processing.

It turns an otherwise difficult-to-recover failure into something explicitly represented in the system.

Recoverable systems are usually far more interesting than systems designed around the assumption that failures will not happen.


Preventing message loss creates another problem: duplicates

There is a consequence to this design.

Imagine the worker publishes an event successfully, but crashes before marking the Outbox message as processed.

When it restarts, the same message may be published again.

sequenceDiagram
    participant O as Outbox
    participant W as Ledger Worker
    participant K as Kafka
    participant B as Balance Worker
    participant DB as Balance DB

    O->>W: Pending message
    W->>K: Publish event
    K-->>W: Success

    Note over W: Worker crashes before<br/>marking the message as processed

    O->>W: Same message again
    W->>K: Republish event

    K->>B: Event
    B->>DB: Was this event_id processed?
    DB-->>B: No
    B->>DB: Update projection and store event_id

    K->>B: Duplicate event
    B->>DB: Was this event_id processed?
    DB-->>B: Yes

    Note over B: Duplicate safely ignored

This is not necessarily a bug.

It is a normal consequence of at-least-once delivery.

The responsibility now moves to the consumer.

BalanceService records processed event identifiers so the same event does not modify the projection twice.

That is idempotency becoming part of the architecture.

Retries are often presented as generic resilience configuration: retry three times, use exponential backoff, done.

But retry is also a business decision.

If executing the same operation twice produces two different side effects, a retry may cause more damage than the original failure.

Before asking:

How many times should we retry?

I think a better question is:

What happens if we execute this operation again?

That question tends to uncover much more important design problems.


Eventual consistency has to exist outside the diagram too

Once Ledger and Balance are separated, this state becomes perfectly valid:

Ledger:
Transaction exists.

Balance:
Transaction has not been projected yet.
Enter fullscreen mode Exit fullscreen mode

There is nothing inherently wrong with that.

The problem appears when the architecture is asynchronous but the rest of the product behaves as though every read must immediately reflect every write.

Different domains solve this in different ways.

Some operations can expose a processing state. Some critical queries may need to consult the source of truth. Some applications can simply tolerate a small delay before projections become visible.

There is no universal answer.

The important point is that eventual consistency should be a known business and architectural property, not an accidental side effect of adding Kafka.


Events are APIs too

Once services depend on events, another problem eventually appears:

contracts change.

The project includes an evolution from events such as:

LedgerEntryCreated.v1
LedgerEntryCreated.v2
Enter fullscreen mode Exit fullscreen mode

During a migration period, consumers may need to understand both versions.

That forces us to treat event schemas as real integration contracts.

We already think carefully about HTTP APIs: OpenAPI definitions, breaking changes, versioning, client compatibility.

Events deserve similar care, perhaps even more.

An HTTP response usually exists only during a request.

An event may remain in a Kafka topic, a dead-letter queue, or a replay mechanism long after the application version that originally created it has disappeared.

That changes how we think about compatibility.

Ordering deserves attention too.

Kafka guarantees ordering within a partition, not universal ordering across the entire system.

The message key therefore becomes an architectural decision because it influences which events share a partition and which operations can be processed in parallel.

Details that look small in producer code can have significant consequences for correctness and throughput.


Retry cannot be the final answer

Some failures are temporary.

A broker can become unavailable. A network request can time out. A dependency may take longer than expected to respond.

Retry with backoff makes sense in those cases.

But a structurally invalid event will still be invalid ten seconds later.

An incompatible schema will not suddenly become compatible on attempt number 47.

Retrying forever simply turns one bad message into a permanent consumer of CPU, logs, and operational attention.

That is where a Dead Letter Queue, or DLQ, becomes useful.

A useful DLQ should preserve enough information for investigation and recovery: the original payload, event type, failure classification, source information, and correlation metadata.

But putting something in a DLQ is only half the solution.

The more important question is:

What happens next?

Can the message be safely replayed?

Was the root cause fixed?

Can a projection be rebuilt?

Should the message be discarded?

Does the operation require manual review?

The project explores requeue, replay, and projection rebuild scenarios because recovery is part of the architecture too.

A phrase I keep coming back to is:

A DLQ without a recovery strategy is just a distributed archive of unresolved problems.


Transfers make distributed failures much easier to see

Updating a read projection is relatively straightforward compared with an operation that produces multiple distributed side effects.

Consider a transfer.

At a simplified level, we need to create a debit and then a credit.

Inside one local database transaction, ACID properties solve a lot of problems for us.

Across independent components, those guarantees disappear.

The project models this flow as an orchestrated Saga:

flowchart TD
    A[Transfer Requested] --> B[Create Debit]

    B -->|Success| C[Create Credit]
    B -->|Failure| F[Transfer Failed]

    C -->|Success| D[Transfer Completed]
    C -->|Failure| E[Compensate Debit]

    E --> G[Record Reversal]
    G --> F

This makes something very explicit:

when we distribute an operation, we also distribute its failure modes.

If the debit succeeds and the credit fails, there is no global ROLLBACK statement capable of reversing time across independent services.

We have to define what compensation means.

In financial domains, compensation often should not erase the original fact. A reversal can instead create a new fact that neutralizes the previous one while preserving the history of what happened.

A Saga does not recreate ACID across services.

It models the states and compensating actions required because that local transaction boundary no longer exists.

That distinction is important.


Circuit breakers do not hide failures. They control them.

Another scenario in the laboratory intentionally takes LedgerService offline while the transfer worker is still running.

Without protection, the worker can keep calling a dependency that is already known to be unavailable.

Retries can make this even worse by multiplying calls against an unhealthy service.

A Circuit Breaker changes that behavior.

stateDiagram-v2
    [*] --> Closed

    Closed: Requests flow normally
    Open: Requests fail fast
    HalfOpen: Limited probe requests

    Closed --> Open: Consecutive failures
    Open --> HalfOpen: Wait period expires
    HalfOpen --> Closed: Dependency recovered
    HalfOpen --> Open: Dependency still unhealthy

When the circuit is open, the system has not somehow recovered.

It has acknowledged the failure and chosen not to waste resources repeatedly exercising the same broken dependency.

After a configured interval, the breaker allows a controlled probe through the half-open state.

If that succeeds, traffic can resume.

This leads to another principle that I think is worth emphasizing:

Resilience does not mean making failures invisible.

Resilience means making failure behavior predictable.


Observability becomes much more important after leaving HTTP

Following a single HTTP request through logs is usually manageable.

Following an asynchronous workflow across several processes is very different.

A single operation may travel through:

flowchart LR
    Request[HTTP Request]
        --> API[Ledger API]
        --> DB[(PostgreSQL)]
        --> Outbox[Outbox]
        --> Worker[Ledger Worker]
        --> Kafka[(Kafka)]
        --> Consumer[Balance Worker]
        --> Projection[(Balance DB)]

    Request -. Correlation ID .-> API
    API -. Trace Context .-> Outbox
    Outbox -. traceparent .-> Worker
    Worker -. Trace Context .-> Kafka
    Kafka -. traceparent .-> Consumer

If every component produces unrelated logs, diagnosing a problem becomes timestamp archaeology.

The project propagates correlation_id, and when OpenTelemetry is enabled, W3C tracing context such as traceparent and tracestate can also travel through the Outbox and Kafka messages.

The observability side can be viewed separately:

flowchart LR
    Client[Client] --> API[Ledger API]
    API --> DB[(PostgreSQL)]
    DB --> Outbox[Outbox]
    Outbox --> Worker[Ledger Worker]
    Worker --> Kafka[(Kafka)]
    Kafka --> Consumer[Balance Worker]
    Consumer --> Balance[(Balance DB)]

    API -.-> OTEL[OpenTelemetry]
    Worker -.-> OTEL
    Consumer -.-> OTEL

    OTEL --> Traces[Traces]
    OTEL --> Metrics[Metrics]

    API -. Logs .-> Logs[Centralized Logs]
    Worker -. Logs .-> Logs
    Consumer -. Logs .-> Logs

The local stack includes OpenTelemetry, Jaeger, Prometheus, Grafana, Loki, and supporting components.

But the tools are not the important part.

The important part is the questions they allow us to answer.

The API is healthy, but is the Outbox backlog growing?

Is the producer failing?

Is consumer processing slowing down?

Are duplicates increasing?

Has the DLQ started receiving messages?

Which original HTTP request produced the event that failed several minutes later in another process?

Once asynchronous processing becomes central to the system, an HTTP health endpoint tells only a very small part of the story.


Test the architecture, not only the classes

A system like this can have excellent unit-test coverage and still fail exactly where the interesting risks are: between components.

That is why the laboratory also contains integration tests and k6 scenarios that exercise complete workflows.

One scenario validates:

Ledger
  -> Outbox
  -> Kafka
  -> Balance
Enter fullscreen mode Exit fullscreen mode

Another exercises the complete Transfer Saga.

There is also a resilience scenario in which the Ledger dependency is deliberately stopped so the Circuit Breaker's behavior can be observed during failure and recovery.

I am careful not to describe these tests as production-scale benchmarks.

They run in a controlled local environment.

Their latency thresholds are regression guardrails, not production SLOs.

Running 50 requests per second on Docker Compose does not prove that an architecture can operate at banking scale.

Real capacity depends on infrastructure, partitions, database behavior, network topology, autoscaling, workload characteristics, failure modes, and many other variables.

The tests prove something narrower, but still valuable:

the architectural behavior remains testable under concurrency and controlled failure conditions.

Sometimes the best architecture test is not checking whether a certain method was called.

It is turning a dependency off and observing what the system actually does.


This is a laboratory, not a production reference architecture

Projects that demonstrate microservices, Kafka, Outbox, Sagas, and observability can easily give the impression that they represent a universal production blueprint.

This one does not.

poc-arquitetura is intentionally a laboratory for experimenting with decisions.

A real production environment would still need serious work around secrets management, workload identity, high availability, Kafka capacity and replication, disaster recovery, network security, production SLOs, infrastructure topology, deployment strategy, and many domain-specific controls.

That distinction matters.

Software architecture is not the process of fitting as many patterns as possible into a diagram.

It is the process of choosing mechanisms that are proportional to the problems the system actually has.


Quick takeaways

The database commit is often only the beginning of a distributed workflow.

Transactional Outbox does not make PostgreSQL and Kafka a single transaction; it makes the intent to publish durable.

At-least-once delivery means duplicates are expected, which makes idempotency a correctness requirement rather than a nice optimization.

Eventual consistency has to be understood by the product, not hidden behind a message broker.

Retries require operations that are safe to repeat. DLQs require recovery procedures. Sagas make states and compensations explicit when a local transaction is no longer available.

Circuit Breakers control failures instead of pretending those failures disappeared.

And observability has to cross the same boundaries that events cross.

If I had to reduce the whole experiment to one idea, it would be this:

In distributed systems, we do not design only the successful path. We also design how failures are detected, understood, and recovered.


Where to go next

For a broader view of .NET microservice architecture, Microsoft's “.NET Microservices: Architecture for Containerized .NET Applications” is a useful starting point.

For Transactional Outbox, Saga, Idempotent Consumer, and related distributed-system patterns, Chris Richardson's Microservices Patterns and the Microservices.io pattern catalog are excellent references.

For Kafka, it is worth going beyond basic producer/consumer tutorials and studying the official material on partitions, consumer groups, offsets, ordering, delivery semantics, and transactions.

For distributed tracing and context propagation, the OpenTelemetry documentation provides a vendor-neutral mental model for traces, spans, metrics, baggage, and propagation.

For load and resilience testing, Grafana k6 is approachable enough to start small while still supporting serious workloads.

And if you want to understand the deeper ideas behind transactions, replication, streams, partitioning, and distributed data, Designing Data-Intensive Applications, by Martin Kleppmann, is still one of the books I would put near the top of the list.

The repository behind this article contains the source code, ADRs, event contracts, architecture documentation, tests, and failure scenarios discussed here:

GitHub logo rodri-oliveira-dev / poc-arquitetura

.NET architecture PoC for distributed services with CQRS, Kafka, PostgreSQL, Outbox, DLQ, observability and CI quality gates.

poc-arquitetura

Build Tests Quality Gate Status Security Rating Architecture Docs

POC educacional de microserviços em .NET para estudar arquitetura de software com código real: Clean Architecture, DDD, PostgreSQL, Kafka, Outbox, Inbox, JWT/JWKS com Keycloak, observabilidade, segurança, contratos e testes automatizados.

Ela demonstra um problema comum em sistemas financeiros: registrar fatos de forma transacional, publicar eventos com confiabilidade, projetar saldos em outro serviço e operar falhas sem esconder consistência eventual. O repositório também mostra contextos de identidade, transferência, pagamento externo e auditoria funcional para exercitar trade-offs de integração.

Este projeto é útil para:

  • quem está aprendendo arquitetura e quer ver os conceitos aplicados;
  • desenvolvedores .NET que querem executar, testar e alterar uma stack local;
  • arquitetos que querem avaliar decisões, limites e riscos;
  • avaliadores técnicos que querem entender a proposta rapidamente.

Visão geral

flowchart LR
    Client[Cliente ou teste] --> Keycloak[Keycloak OIDC]
    Client --> LedgerApi[LedgerService.Api]
    Client --> BalanceApi[BalanceService.Api]
    Client --> TransferApi[TransferService.Api]
    Client --> PaymentApi[PaymentService.Api]
    Client --> IdentityApi[IdentityService.Api]
    Client --> AuditApi[AuditService.Api]
    LedgerApi -->
…

There is still plenty I want to experiment with.

And that is probably the part I enjoy most about this kind of architecture work: once failure stops being treated as an exceptional event and becomes part of the design, the questions get much more interesting.

Top comments (0)