The previous post covered why I chose pgvector for semantic search.
But there was another question:
When should a memory be embedded?
At first, it might seem simple:
Create memory
▼
Generate embedding
▼
Save embedding
But this makes memory creation dependent on the embedding process.
If the embedding provider is slow or unavailable, creating a memory could also fail or become slow.
I wanted to separate these two operations.
The architecture became:
Memory Service
▼
PostgreSQL
▼
Outbox Table
▼
Queue / Job Runner
▼
Embedding Worker
▼
pgvector
1. The Outbox Pattern
When a user creates a memory, the Memory Service writes the memory and an outbox event in the same database transaction.
For example:
Transaction
┌─────────────────────────────┐
│ INSERT memory │
│ INSERT outbox event │
└─────────────────────────────┘
The important part is that they succeed or fail together.
This avoids a common problem with distributed systems:
Memory saved
▼
Embedding event lost
The outbox gives me a durable record of the work that needs to happen.
The outbox is not the queue itself.
It is a reliable bridge between the database transaction and asynchronous processing.
2. The Queue / Job Runner
A separate process reads pending outbox events and puts jobs onto a queue.
Outbox
▼
Queue
▼
Embedding Worker
This gives the embedding process some useful properties:
asynchronous processing
retries
independent scaling
failure isolation
no need to block the memory API
The user can save a memory without waiting for the embedding provider.
3. The Embedding Worker
The embedding worker has one main responsibility:
turn memory text into an embedding and store it in pgvector.
flowchart TD
A[Get Memory] --> B[Generate embedding]
B --> C[Store Vector]
This keeps embedding-specific logic out of the Memory Service's synchronous request path.
It also gives me a place to evolve the embedding pipeline later.
For example, I could change:
embedding models
chunking strategy
retry behaviour
batch processing
embedding dimensions
without changing the API used to create a memory.
4. Why asynchronous?
There is an important consequence of this design:
A newly created memory may not be immediately searchable.
There can be a small delay between:
Memory created
│ asynchronous processing
▼
Embedding created
▼
Memory available for semantic search
I accepted this trade-off.
For Second-Memory, immediate persistence is more important than making the embedding operation part of the user's request.
This is essentially eventual consistency between the memory record and its vector representation.
5. Failure becomes easier to handle
The asynchronous design also changes how failures work.
If the embedding provider temporarily fails:
Memory
▼
Outbox
▼
Queue
▼
Embedding Worker - X -> Retry
The memory itself has already been safely stored.
The embedding job can be retried without asking the user to submit the memory again.
That separation was important to me.
The decision
The final flow became:
graph TD
CM[Create Memory] --> PG
subgraph PG [PostgreSQL]
M[Memory]
OE[Outbox Event]
end
OE --> Q[Queue]
Q --> EW[Embedding Worker]
EW --> PV[(pgvector in PostgreSQL)]
Each part has a different responsibility:
| Component | Responsibility |
|---|---|
| Memory Service | Store the memory |
| Outbox | Reliably record the event |
| Queue | Deliver asynchronous work |
| Embedding Worker | Generate embeddings |
| pgvector | Store and search vectors |
This added some complexity compared with simply generating the embedding inside the API request.
But it gave me something more important:
memory creation is no longer tightly coupled to the availability of the embedding pipeline.
That was the trade-off I wanted.
Top comments (0)