DEV Community

Ashish Mishra
Ashish Mishra

Posted on Originally published at infino.ai on

Choosing a database for AI agents: memory, retrieval and state compared

· engineering

The right database for an AI agent depends on the jobs its memory has to do. An agent recalls past context by meaning, looks up exact names and identifiers, filters and counts structured state such as users, sessions and time windows, and keeps a history that only grows. Most agents need all of these at once, which is why keyword and vector search over the same records, with SQL over the results, has become the common requirement.

This page compares seven common options on those capabilities. It compares capabilities only, not prices, and it says where each option is the right choice.

What does an AI agent need from its database?

  • Recall by meaning. Vector search finds a past message or document that says the same thing in different words.
  • Recall by exact term. Names, ticket numbers, error codes and product IDs are where embeddings blur, because a rare token dissolves into its neighborhood. Keyword search with BM25 finds them. Where keyword search and vector search each fail shows both failures on real data.
  • Structured state. Memory is scoped to a user, a session, a tool or a time range, and an agent often needs a number rather than a list: how many times did this step fail this week, and for which customers.
  • Growth. An agent's history is appended to and rarely deleted, so the stored corpus grows every day while each query stays small.
  • Where it runs. A library inside the agent's process has no server to operate and no network hop. A server or hosted service can be shared by many agents and scaled on its own.

The first two are the case for hybrid search: an agent's queries mix exact identifiers with loosely worded recall, and each retriever misses what the other finds. The third is the case for SQL over the search results rather than a second system to join them in application code.

How the common options compare

system runs as keyword + vector in one query SQL over results storage format open source
Infino embedded library or hosted yes: BM25 plus vector, fused with reciprocal rank fusion in one query yes: DataFusion SQL, with search as table functions Parquet, with the indexes inside the files yes (Apache-2)
Postgres + pgvector database server partial: native full-text plus pgvector, BM25 needs another extension, ranks combined by you yes, full Postgres SQL Postgres tables yes (PostgreSQL license)
Elasticsearch / OpenSearch cluster or service yes: BM25 plus kNN, fused partial: ES|QL, a SQL plugin, PPL Lucene segments, engine-private OpenSearch Apache-2, Elasticsearch AGPL / ELv2
Qdrant server or cluster partial: dense plus sparse vectors no, retrieval API engine-private indexes yes (Apache-2)
Pinecone hosted service partial: dense plus sparse vectors no, retrieval API engine-private indexes no, hosted only
LanceDB embedded library or hosted yes: full-text index plus vector, fused, optional reranker partial: SQL filter expressions on queries Lance, an open columnar format yes (Apache-2)
Chroma embedded, server, or hosted partial: dense, sparse and full-text search, local and distributed features still converging no, retrieval API with metadata filters engine-private yes (Apache-2)

Checked against each project's documentation in September 2026. Each project moves quickly, so check the current docs for the feature you depend on. The comparison pages go system by system.

When is each one the right choice?

  • Infino when an agent needs keyword, vector and SQL over the same memory in one query, with the data kept as Parquet in your own bucket and no server to run. It is not a system of record: transactional, write-heavy state belongs in an OLTP database next to it.
  • Postgres with pgvector when the agent's application state already lives in Postgres and the memory corpus is modest. One database, transactional writes and full SQL. The cost is that relevance fusion is yours to write, and text, vectors and the application share the same instances.
  • Elasticsearch or OpenSearch when full-text relevance tuning matters most, with analyzers, synonyms and mature ranking controls, and a team already runs a search cluster.
  • Qdrant or Pinecone when the workload is dedicated vector serving at a high sustained query rate, filtered on payload metadata. Pinecone when nobody should run servers at all.
  • LanceDB when the agent works with vectors and multimodal data on files, embedded in the process, and the same tables also feed training or data pipelines.
  • Chroma for the quickest local start on a prototype, with a hosted path when it grows.

What this looks like in Infino

Infino is an open source retrieval library that runs inside the agent's process. A superfile is a search index that is a valid Parquet file: the data and the BM25 and vector indexes in one file. An agent's memory is one table of those files, on local disk or in object storage, and a single query can run a hybrid search, filter it by user and time, and group the results, because each search function is a table in SQL.

The agent memory guide walks through the pattern with code, and agents that speak MCP can use the Infino MCP server without writing any. Implementing hybrid search on Apache Parquet files is the runnable end-to-end version.

Common questions

Do AI agents need a vector database?

They need vector search, which is not the same thing. Vector search recalls past context by meaning, but it misses exact names, identifiers and error codes, and a standalone vector database adds a second system to keep in sync with the rest of the agent’s state. Engines that run keyword and vector search over the same records cover both needs.

Is hybrid search necessary for agent memory?

For most agents, yes. Their queries mix exact identifiers, such as a customer name or a ticket number, with loosely worded recall, such as what went wrong last time. Keyword search finds the first kind and vector search finds the second, and fusing the two rankings returns the rows that either one alone would miss. How hybrid search fuses the two.

Can I use Postgres for AI agent memory?

Yes. Postgres with pgvector and its native full-text search covers agent memory well when the application already runs on Postgres and the corpus is modest. BM25 ranking needs another extension, the two rankings are combined in your own SQL, and memory shares the instances that serve the application.

Should an agent’s database be embedded or a server?

Embedded when one agent process owns its memory and simplicity matters, since there is nothing to deploy and no network hop on a lookup. A server or hosted service when many agents share one store or it needs to scale apart from them. Some engines, Infino and LanceDB among them, offer both.

What storage format should agent memory use?

An open one, if other tools need to read the history. Engine-private formats are readable only through the engine that wrote them. Infino stores memory as standard Parquet, so any Parquet reader can open it for analytics or export without Infino in the path. Parquet interop.

Top comments (0)