DEV Community

Silver_dev
Silver_dev

Posted on

Vector database showroom. Part 6 Milvus — The High-Speed Freight Train

🚄 Milvus — The High-Speed Freight Train

Billions of vectors aboard, uncontested on the mainline. But nobody rides before the rails are built.

Under the hood: a Go + C++ core, Apache 2.0, a graduated project of the Linux Foundation's LF AI & Data, driven by Zilliz (who also sell the managed ride, Zilliz Cloud). The distributed deployment is a real architecture: proxy, coordinator services, and worker nodes (query, data, index) — all hanging off a central log (Pulsar or Kafka), with etcd for metadata and MinIO/S3 for object storage. Your vectors live in object storage; storage and compute are separated. One train, three gauges: Milvus Lite (embedded, pip install, desk scale), Milvus Standalone (a rail bus — three containers: Milvus itself, etcd for metadata, MinIO for object storage), and Milvus Distributed (the mainline — Helm on Kubernetes).

Where it shines:

  • Billions of vectors is the home stadium: capacity scales with your object storage, not your RAM budget; horizontal scale means adding worker nodes. This is the tool that made "vector database" a category
  • GPU acceleration (since v2.4): GPU_CAGRA built on NVIDIA's RAPIDS — for when search is the bottleneck and you have the silicon
  • DiskANN (since v2.4): the index lives on SSD — vectors beyond RAM, latency traded for capacity
  • Sparse vectors since v2.4, and since v2.5 a built-in BM25 full-text search — the factory-grade hybrid, no separate sparse embedder required
  • Multi-tenancy with documented patterns (partition keys, separate databases), RBAC, TLS — and Attu, a proper admin GUI. The only vehicle in the showroom that ships with a cockpit display
  • One of the largest OSS communities in the category; SDKs for Python/Java/Go/Node; LangChain/LlamaIndex integrations; Zilliz Cloud if you'd rather rent the railway

Where it stalls:

  • The rails: Distributed means etcd + Pulsar/Kafka + MinIO/S3 + the Milvus pods themselves. Helm charts help, but the ops muscle is non-negotiable. (Lite and Standalone genuinely shrink this — see the manual below)
  • Bring your own embeddings — a database, not an embedder (same rule as the hatchback and the crossover)
  • No ACID transactions "like in Postgres": batches are not transactions — the station wagon still owns that lane
  • Tuning is part of the contract: M, efConstruction, nprobe/ef per index type — and a consistency level per query
  • Backups run through the dedicated milvus-backup tool: plan it and rehearse it. There's no "it's just my Postgres cluster" simplicity here

What breaks if you skip the manual:

  • Bought the mainline for the street corner. A composite story from community chats, details changed: a team runs semantic search over 300k documents on distributed Milvus — a small zoo of pods, of which exactly two kinds accept queries — because a big-tech talk mentioned billions. A junior asks: "Why not Postgres? We already have Postgres." Calculator: 300,000 × 1536 dims × 4 bytes ≈ 1.8 GB. The right gauge existed all along: Milvus Lite for the prototype, Standalone for small production, Distributed for the billions. Milvus isn't the mistake — the scale mismatch is
  • Benchmark trap #2 (the hatchback's cousin): fresh inserts land in growing segments that are searched by brute force until they're sealed and indexed. A naive benchmark right after loading — or on a tiny collection — measures exact search, not your index. Honest numbers: production-like volume, settled segments, warm queries
  • Consistency levels are per-query (default: bounded): strong / bounded / session / eventually. With weaker levels, "insert → immediately search" tests flake exactly like the taxi's eventual consistency. Pick session or strong for the paths that need it
  • Metric and dimension are locked at collection creation — same rule as the hatchback and the taxi. The escape hatch here is actually elegant: build the new collection and flip a collection alias — the zero-downtime re-index is a documented pattern
  • The partition key (if you use partition-key tenancy) is fixed at creation too — retrofitting tenancy later means re-shipping the data

Cost of ownership: the train is free (Apache 2.0); the railroad isn't — count the ops salary, not just the license. Zilliz Cloud (managed, serverless tier included) rents you the whole railway. Self-hosted Lite and Standalone genuinely fit a laptop and a small server respectively.

Mechanics & parts: Zilliz drives development; LF AI & Data provides neutral governance; the community is one of the largest in the category. "Milvus engineers" are rarer than Postgres ones — but this is a Kubernetes-native system your DevOps team can actually operate, with vendor support available when the stakes are high.

✅ Take it if: you're at hundreds of millions to billions of vectors, need GPU or beyond-RAM capacity, run many tenants, have (or hire) real ops muscle — or use Zilliz Cloud and skip the rails

❌ Pass if: you're under ~10M vectors on a single box (the hatchback or the wagon covers it), "we plan to grow" is your only argument, or nobody on the team wants to own a Helm chart

Test drive:

# Milvus Lite runs embedded — no containers needed (Linux/macOS)
# Production gauges: Standalone (3-container compose) or Distributed (Helm on K8s)
from pymilvus import MilvusClient, DataType

client = MilvusClient("./milvus_demo.db")  # the desk model of the train

# explicit schema: scalar fields never appear by themselves (Milvus is schema-first,
# no AutoSchema like Weaviate). auto_id=False keeps our own ids — same as Parts 3–5
schema = client.create_schema(auto_id=False, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("vector", DataType.FLOAT_VECTOR, dim=1536)
schema.add_field("content", DataType.VARCHAR, max_length=512)
schema.add_field("year", DataType.INT64)

index_params = client.prepare_index_params()
index_params.add_index(field_name="vector", index_type="AUTOINDEX", metric_type="COSINE")

client.create_collection(
    collection_name="docs",
    schema=schema,
    index_params=index_params,
)

client.insert(
    collection_name="docs",
    data=[
        {"id": 1, "vector": embed("Quarterly report 2024"), "content": "Quarterly report 2024", "year": 2024},
        {"id": 2, "vector": embed("Legal contract draft"), "content": "Legal contract draft", "year": 2024},
    ],
)
# note: fresh data sits in "growing segments" searched by brute force until they are
# sealed and indexed (benchmark trap #2). flush() seals segments — but in production
# let auto-sealing do its job; frequent manual flushes are an anti-pattern

# embed() — your embedding function (the same one from Parts 3–5 — honest comparison)
res = client.search(
    collection_name="docs",
    data=[embed("annual financial summary")],
    filter="year == 2024",
    limit=10,
    output_fields=["content", "year"],
)

for hits in res:
    for hit in hits:
        print(hit["id"], round(hit["distance"], 3), hit["entity"].get("year"))
Enter fullscreen mode Exit fullscreen mode

The train is magnificent — on the mainline, where it was built to run. Ask it at 300k vectors — Postgres wins. Ask it at 3 billion — the train wins. Buy for the scale you have; re-check at the scale you reach.

🚦 Final part of the showroom: all six vehicles side by side — the cost-of-ownership table, the decision tree, and a one-day test-drive checklist for the car you pick.

Top comments (0)