DEV Community

Joung Park
Joung Park

Posted on

Why the Outbox Pattern, Queue, and Embedding Worker?

The previous post covered why I chose pgvector for semantic search.

But there was another question:

When should a memory be embedded?

At first, it might seem simple:

Create memory
     ▼
Generate embedding
     ▼
Save embedding
Enter fullscreen mode Exit fullscreen mode

But this makes memory creation dependent on the embedding process.

If the embedding provider is slow or unavailable, creating a memory could also fail or become slow.

I wanted to separate these two operations.

The architecture became:

Memory Service
      ▼
 PostgreSQL
      ▼
 Outbox Table
      ▼
 Queue / Job Runner
      ▼
 Embedding Worker
      ▼
 pgvector
Enter fullscreen mode Exit fullscreen mode

1. The Outbox Pattern

When a user creates a memory, the Memory Service writes the memory and an outbox event in the same database transaction.

For example:

Transaction
┌─────────────────────────────┐
│ INSERT memory               │
│ INSERT outbox event         │
└─────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

The important part is that they succeed or fail together.

This avoids a common problem with distributed systems:

Memory saved
     ▼
Embedding event lost
Enter fullscreen mode Exit fullscreen mode

The outbox gives me a durable record of the work that needs to happen.

The outbox is not the queue itself.

It is a reliable bridge between the database transaction and asynchronous processing.

2. The Queue / Job Runner

A separate process reads pending outbox events and puts jobs onto a queue.

Outbox
   ▼
Queue
   ▼
Embedding Worker
Enter fullscreen mode Exit fullscreen mode

This gives the embedding process some useful properties:

  • asynchronous processing

  • retries

  • independent scaling

  • failure isolation

  • no need to block the memory API

The user can save a memory without waiting for the embedding provider.

3. The Embedding Worker

The embedding worker has one main responsibility:

turn memory text into an embedding and store it in pgvector.

flowchart TD
    A[Get Memory] --> B[Generate embedding]
    B --> C[Store Vector]

This keeps embedding-specific logic out of the Memory Service's synchronous request path.

It also gives me a place to evolve the embedding pipeline later.

For example, I could change:

  • embedding models

  • chunking strategy

  • retry behaviour

  • batch processing

  • embedding dimensions

without changing the API used to create a memory.

4. Why asynchronous?

There is an important consequence of this design:

A newly created memory may not be immediately searchable.

There can be a small delay between:

Memory created
     │ asynchronous processing
     ▼
Embedding created
     ▼
Memory available for semantic search
Enter fullscreen mode Exit fullscreen mode

I accepted this trade-off.

For Second-Memory, immediate persistence is more important than making the embedding operation part of the user's request.

This is essentially eventual consistency between the memory record and its vector representation.

5. Failure becomes easier to handle

The asynchronous design also changes how failures work.

If the embedding provider temporarily fails:

Memory
  ▼
Outbox
  ▼
Queue
  ▼
Embedding Worker - X -> Retry
Enter fullscreen mode Exit fullscreen mode

The memory itself has already been safely stored.

The embedding job can be retried without asking the user to submit the memory again.

That separation was important to me.

The decision

The final flow became:

graph TD
    CM[Create Memory] --> PG

    subgraph PG [PostgreSQL]
            M[Memory]
            OE[Outbox Event]
    end

    OE --> Q[Queue]
    Q --> EW[Embedding Worker]
    EW --> PV[(pgvector in PostgreSQL)]

Each part has a different responsibility:

Component Responsibility
Memory Service Store the memory
Outbox Reliably record the event
Queue Deliver asynchronous work
Embedding Worker Generate embeddings
pgvector Store and search vectors

This added some complexity compared with simply generating the embedding inside the API request.

But it gave me something more important:

memory creation is no longer tightly coupled to the availability of the embedding pipeline.

That was the trade-off I wanted.

Top comments (0)