DEV Community

HiDevs
HiDevs

Posted on

Building a Multi-Agent Load Simulator: 50 Agents, One Collection

As AI applications evolve from simple chatbots into complex multi-agent architectures, the demands placed on the vector database change fundamentally. In a single-agent loop, context retrieval is predictable and mostly sequential: an agent sends a search request, waits for a response, passes context to an LLM, and takes its next action.

When scaling to 50 independent agents operating concurrently, that neat sequential pattern breaks down.

In production agent workflows, Qdrant serves as the shared retrieval backbone supplying relevant points to dozens of agents as they reason and decide what to do next. Rather than receiving orderly, one-by-one requests, a single Qdrant collection is bombarded by overlapping, asynchronous search queries.

Standard single-query benchmarks evaluate retrieval in artificial isolation. They fail to capture the resource contention, client-side queueing, and tail-latency spikes that occur when dozens of agents query the exact same collection at the same time.

Figure 1. Multi-agent load simulator architecture.
Figure 1. Multi-agent load simulator architecture.
Fifty independent agent sessions send overlapping requests to one shared Qdrant collection. The simulator records latency percentiles, throughput, in-flight requests, and server metrics.

Overlapping Cycles and Query Interference

A single Qdrant query benchmark measures retrieval in isolation. A single-agent loop adds pauses between searches, but it still does not reproduce the shared demand created when many agents independently retrieve context from the same collection.

In an agent system, one session may be waiting for an LLM response ("think time") while another sends a vector search to Qdrant. A third may issue a payload-filtered search. Their request cycles overlap continuously, so Qdrant sees a changing mix of retrieval operations rather than one orderly sequence.

Figure 2. Sequential single-agent testing compared with concurrent agent sessions.
Figure 2. Sequential single-agent testing compared with concurrent agent sessions.
A sequential test sends the next request only after the previous response, while independent agent sessions generate overlapping requests against the same collection.

Inside the shared Qdrant collection

Every simulated session sends its searches to the same Qdrant collection. The benchmark holds the vectors, payload schema, index configuration, and search settings constant while concurrency changes. That keeps the experiment focused on how the Qdrant workload responds to more simultaneous sessions, rather than changes to the underlying collection.

The workload exercises two Qdrant search paths:

  • Unfiltered Vector Search: Uses Qdrant's HNSW index for approximate nearest neighbor (ANN) retrieval.
  • Payload-Filtered Vector Search: Combines vector retrieval with conditions on stored payload. Payload indexes support filter evaluation, while ACORN may be relevant for restrictive filtered searches depending on the collection and query configuration.

This distinction is central to the experiment. Qdrant is not just receiving more of the same request: the collection is serving both ordinary vector searches and searches constrained by payload. The benchmark varies the mix to observe whether the two query types affect one another under shared load.

How Qdrant Handles Concurrent Demand

Qdrant’s underlying architecture is designed to handle high-concurrency, mixed-query workloads against a single shared collection without dropping performance:

  • Dual Indexing Engine: Qdrant combines an HNSW index for high-throughput vector lookups with dedicated payload indexes and ACORN algorithms. This allows payload-filtered searches to execute without choking raw vector search performance.
  • Asynchronous Request Execution: Qdrant’s core engine processes overlapping operations across multiple CPU threads, insulating fast ANN lookups even when heavier filtered searches run concurrently.
  • Metric Transparency: Qdrant exposes internal telemetry via Prometheus. By comparing end-to-end client timestamps with server-side Prometheus metrics, engineering teams can verify that Qdrant’s internal execution time remains fast and flat even when client-side connection waiting accumulates under heavy load.

Empirical Verification: The Multi-Agent Load Simulator

The simulator models each agent as an independent session with its own request sequence and think time. This matters because the benchmark is intended to test Qdrant under overlapping retrieval demand, not simply to send 50 requests in a burst.

Each simulated agent runs as an independent asynchronous task. It selects a vector-search request and sends it to the shared Qdrant collection through the Qdrant Python client. The agent then records the client-observed round-trip time before waiting for a short interval and issuing its next request.

from time import perf_counter
async def timed_query(client, collection, query, agent_id, query_type):
    started = perf_counter()
    await client.query_points(
        collection_name=collection,
        **query,
)
    return {
        "agent_id": agent_id,
        "query_type": query_type,
        "latency_ms": (perf_counter() - started) * 1000,
}
Enter fullscreen mode Exit fullscreen mode

The query_points operation used above is documented in Qdrant's Search and Query documentation.

The timer measures the total round-trip time observed by the client when sending a search request to the shared Qdrant collection and receiving a response. It does not measure Qdrant's internal search execution time alone, as the total latency can also include connection-pool waiting and network transit. This distinction is important when evaluating Qdrant's search performance under concurrent agent workloads.

Figure 3. Independent agent-session request timeline.
Figure 3. Independent agent-session request timeline.
Each session follows its own request cadence and think-time intervals, producing overlapping asynchronous activity rather than one synchronized request stream.

Experiment 1: Concurrency Scaling

The first experiment increases the number of concurrent agent sessions sending requests to a shared Qdrant collection, using concurrency levels of 1, 5, 10, 25, and 50. It tracks how client-observed latency, throughput, and in-flight requests change as more agents access the same Qdrant collection.

Table 1

Figure 4. Latency percentiles across concurrency levels.
Figure 4. Latency percentiles across concurrency levels.
The chart displays p50, p95, and p99 latency at 1, 5, 10, 25, and 50 sessions. Populate it only with verified benchmark measurements.

*Experiment 2: Query-Mix Interference at Fixed Concurrency (50 Agents)
*

The second experiment keeps the workload fixed at 50 concurrent agent sessions accessing Qdrant and changes only the proportion of unfiltered vector searches and payload-filtered searches against the same collection:

Table 2
The filtered-search path is based on Qdrant's payload filtering capabilities.

Because the total remains fixed at 50 concurrent agent sessions accessing the same Qdrant collection, this experiment isolates the effect of query composition, allowing us to compare how different proportions of unfiltered vector searches and payload-filtered searches affect Qdrant under the same concurrency level.

Figure 5. Query-mix configurations at a fixed concurrency of 50 sessions.
Figure 5. Query-mix configurations at a fixed concurrency of 50 sessions.
The four runs keep total concurrency constant and vary only the proportion of unfiltered vector searches and payload-filtered searches.

Separating Qdrant Execution from Client-Side Waiting

A higher client-observed round-trip time does not necessarily mean that Qdrant's internal search execution has become slower. The total response time can include waiting for an available client connection, network transit, and Qdrant's query execution time. These components must be distinguished before attributing any increase in latency to Qdrant's search performance.

Figure 6. Client-side and server-side measurement paths.
Figure 6. Client-side and server-side measurement paths.
Client timestamps capture end-to-end round-trip time, while Qdrant Prometheus metrics provide server-side visibility. Comparing them helps distinguish client or network waiting from database execution.

Making the Comparison Reproducible

Each benchmark run follows a consistent sequence: warmup to establish the initial state of the Qdrant collection, a measured-load period to capture client requests and Qdrant server metrics, and cool-down to allow outstanding requests to complete. For a fair comparison, matched runs should begin with the same Qdrant collection state and use the same seeded request schedule and query set, changing only one variable at a time.

For filtered-search comparisons, the filter fields, Qdrant payload-index configuration, and filter selectivity should also be recorded. These factors shape how Qdrant processes payload-filtered searches alongside unfiltered vector searches and provide the necessary context for interpreting retrieval performance under concurrent workloads.

Figure 7. Reproducible benchmark execution stages.
Figure 7. Reproducible benchmark execution stages.
The benchmark proceeds through warm-up, measured load, and cool-down so runs can be compared under a consistent procedure.

Conclusion & Production Readiness

As AI applications move toward multi-agent architectures, Qdrant must handle more than isolated retrieval requests. Multiple agents can independently access the same Qdrant collection, combining unfiltered vector searches and payload-filtered searches within overlapping request cycles.

This benchmark provides a structured way to examine how Qdrant handles that shared retrieval demand. By increasing concurrency and varying the search mix, it evaluates Qdrant's HNSW-based retrieval, payload-filtered search, latency, throughput, and server-side behavior under controlled conditions.

Figure 8. Benchmark results dashboard for the 50-session workload.
Figure 8. Benchmark results dashboard for the 50-session workload.
The dashboard groups latency percentiles, throughput, in-flight requests, and server telemetry. Use only after verified measurements are available; do not populate it with estimated values.

The resulting measurements help identify how the tested Qdrant configuration behaves as concurrent demand changes. Rather than relying on single-query benchmarks, this approach provides a more representative view of Qdrant's role as the retrieval layer in multi-agent applications.

Top comments (0)