DEV Community

Alex Georgiev
Alex Georgiev

Posted on AI-assisted

MongoDB Community Edition's new search index stayed pending above 89% disk use

MongoDB Search and Vector Search for Community Edition went GA on 30 June 2026. Self-managed deployments get the same mongot search process that used to be Atlas-only, running alongside mongod instead of bolted onto a separate engine. I followed MongoDB's own Docker installation guide word for word, and the search index never left PENDING.

Setting it up

The docs describe two containers on one network: mongod (MongoDB Community Server 9.0.2) and mongot-community (mongot 1.70.5), joined by a gRPC connection and a SCRAM user with the searchCoordinator role. I ran it exactly as documented: a single-member replica set, a mongod.conf pointing at the search host, a mongot.conf pointing back at mongod, then createSearchIndex on a 50,000-document test collection.

The index came back immediately, as expected:

[ { name: 'default', status: 'PENDING', queryable: false } ]
Enter fullscreen mode Exit fullscreen mode

PENDING is normal for the first few seconds. It was not normal five minutes later, still retrying every 30 seconds with the same message in the mongot log:

"msg":"Transient error while trying to sync ... Retrying in 30000 milliseconds"
"stack_trace":"... PauseInitialSyncException: Initial syncs are paused ..."
Enter fullscreen mode Exit fullscreen mode

Finding the actual cause

mongot exposes Prometheus metrics on port 9946, and one of them settled it:

mongot_system_disk_space_data_path_free_bytes 2.8686217216E10
mongot_system_disk_space_data_path_total_bytes 2.70553174016E11
Enter fullscreen mode Exit fullscreen mode

28.7GB free out of 270.6GB, which is 89.4% used. MongoDB's own self-managed troubleshooting page states the rule directly: "Replication stops when disk usage exceeds roughly 90% and resumes after usage drops below roughly 85%. For a new index or rebuild, expect the definition to be accepted but the build to stay stuck if disk pressure is already above the protective threshold." My host had drifted into that band and the index was never going to move.

Nothing about this showed up on the client side. createSearchIndex returned normally, $listSearchIndexes just said PENDING, and df -h on the same path reported 28% used, not 89%, because df was measuring against a quota-limited allowance while mongot was dividing by the underlying device's full reported capacity. Two tools looking at the same disk gave answers 60 points apart.

I moved mongot's data directory onto tmpfs to get a mount with headroom, and the same index went from PENDING to serving queries in under five seconds:

"msg":"Finished a collection scan phase.","attr":{"numDocumentsIndexed":50000}
"msg":"Completed initial sync. Beginning first commit.","attr":{"duration":"4.850 s", ...}
"msg":"Transitioning from INITIAL_SYNC to STEADY_STATE."
Enter fullscreen mode Exit fullscreen mode

The block was real and the fix was disk headroom, nothing else.

A status field that keeps lying

After I recreated the mongot container, $listSearchIndexes kept reporting default as PENDING even though queries against it worked. The full statusDetail array explained why: it still listed the old, now-dead mongot host as PENDING alongside the new host, which was READY. The aggregated top-level status is the worst of all hosts it has ever seen, including ones that no longer exist. A vector index I created fresh, with no dead host in its history, reported READY cleanly. If you restart mongot in place, don't trust the summary status; read statusDetail per host.

What it refuses, and what it doesn't

Querying an index that exists but isn't ready yet is a hard, specific error:

OperationFailure: cannot query search index ... while in state NOT_STARTED (code 8)
Enter fullscreen mode Exit fullscreen mode

Querying an index name that doesn't exist at all is not an error. It returns zero results silently, with only a WARN "No index in catalog" line in mongot's own log, invisible to the client. A typo in an index name and a genuinely empty result set look identical from the application side.

The speed comparison I expected to favour $search

Once the index was live, I ran 15 repetitions each of five single-word and two-word queries against $search, a classic MongoDB $text index, and a plain case-insensitive regex, all on the same 50,000-document collection.

Method min median p95 max
$search (mongot) 6.25ms 9.36ms 15.43ms 56.05ms
classic $text index 0.48ms 0.67ms 0.89ms 1.57ms
regex scan 0.48ms 0.79ms 2.86ms 3.95ms

For a plain term lookup, the new search stack was 10 to 14 times slower than the mechanism it's positioned to replace. That's the gRPC round trip to a separate process rather than an in-process index scan, and it's not a reason to avoid mongot — it's a reason not to switch to it for cases the old $text index already handled well.

Where $search earns the extra hop is fuzzy matching. I queried all four methods with backpak, a one-letter typo for backpack:

Method results for "backpak"
$search with fuzzy: {maxEdits: 1} 5
$search, no fuzzy option 0
classic $text index 0
regex 0

Only the fuzzy option found anything. That's the actual capability being sold here, and it worked as documented.

Vector search against doing it by hand

I ran $vectorSearch against a 32-dimension cosine-similarity index over the same 50,000 documents, and compared it to scoring every vector in Python after fetching them once:

Method min median max
$vectorSearch 10.38ms 15.52ms 93.93ms
brute-force Python cosine scan 123.23ms 134.89ms 243.87ms

$vectorSearch was roughly 8.7 times faster at the median, and that's generous to the brute-force side: I excluded the time to fetch all 50,000 embeddings over the wire, which a real application doing this in-process would also have to pay unless it already held every vector in memory.

Concurrency and cost

Ten concurrent clients issuing $search queries finished 20 total queries in 98.6ms of wall time, against 180.7ms for one client doing the same 20 sequentially — a real throughput gain. But per-query latency degraded: the median went from 8.41ms to 28.77ms, and the worst case rose from 13.07ms to 83.40ms. mongot is a shared, single-threaded-ish gRPC service from the client's point of view, and concurrency buys you overall throughput at the cost of tail latency.

I also checked docker stats while both containers sat idle, with nothing but two small indexes over 50,000 short documents behind them. mongod held 229.5MiB resident. mongot held 1.033GiB, over four times as much, and it hadn't answered a single query yet. Anyone budgeting a self-managed box for mongod alone needs to add a second, JVM-sized line item for mongot, and that line item doesn't shrink if the collection is small.

MongoDB's launch post says $search runs on "the same query model you already know." I wanted to see if that survived contact with a real pipeline rather than a one-line example, so I chained $match, $project, $sort, {$meta: "searchScore"} and $limit after a $search stage. It worked without any special-casing, filtering on price after the text match and sorting on the filtered set, which is a normal thing to want and not something the docs show directly.

What I got wrong on the way

The sync config in the docs sets readPreference: secondaryPreferred, and when the index sat in PENDING for the first few minutes I decided that was the problem: a single-node replica set with no secondary to prefer. So I built one. I added a second mongod container, ran rs.add(), watched rs.status() report a healthy SECONDARY, and waited for the index to move. It didn't. Another five minutes of the same 30-second retry loop passed before I gave up on that theory and went looking at mongot's metrics instead, where the disk numbers had been sitting the entire time. df -h on the same host never once suggested there was a problem.

Run it yourself

This needs Docker with a working daemon, not just the CLI.

docker network create search-community

# mongod.conf: net.bindIpAll true, replication.replSetName rs0,
# plus the setParameter block pointing at mongot-community:27028
docker run -d --name mongod --network search-community \
  --network-alias mongod.search-community \
  -v "$PWD/mongod.conf:/etc/mongod.conf:ro" \
  -v "$PWD/data/db:/data/db" -p 27017:27017 \
  mongodb/mongodb-community-server:latest --config /etc/mongod.conf

docker exec mongod mongosh --port 27017 --quiet --eval '
  rs.initiate({_id:"rs0", members:[{_id:0, host:"mongod.search-community:27017"}]})'
docker exec mongod mongosh --port 27017 --quiet --eval '
  db.getSiblingDB("admin").createUser({user:"mongotUser", pwd:"testpass123", roles:["searchCoordinator"]})'

# mongot.conf: syncSource pointing at mongod, storage.dataPath /data/mongot
docker run -d --name mongot-community --network search-community \
  --network-alias mongot-community.search-community \
  -v "$PWD/data/mongot:/data/mongot" \
  -v "$PWD/mongot.conf:/mongot-community/config.default.yml:ro" \
  -v "$PWD/passwordFile:/passwordFile:ro" \
  -p 8080:8080 -p 9946:9946 \
  mongodb/mongodb-community-search:latest

curl localhost:8080/health          # expect {"status":"SERVING"}
curl localhost:9946/metrics | grep disk_space_data_path   # check your own headroom first
Enter fullscreen mode Exit fullscreen mode

Check your disk metric before you file a bug: (total - free) / total from mongot_system_disk_space_data_path_total_bytes and _free_bytes, not df -h. If you're above roughly 85%, free space before you touch the index definition.

What to do with this

Before you create a first index on a self-managed deployment, pull mongot_system_disk_space_data_path_free_bytes and do the arithmetic yourself rather than trusting df on the same box; the two disagreed by 60 percentage points on my host, and only one of them is the number mongot actually acts on. Don't rip out a working $text index or a regex query just because $search is newer — on plain term lookups mine was an order of magnitude slower. Reach for it when you need fuzzy matching or relevance scoring you can't already get, and reach for $vectorSearch once scoring vectors in application code starts to hurt, which on my 50,000-row test happened well before the collection felt large. If you restart mongot, open statusDetail before you believe whatever the top-level queryable field tells you.

Top comments (1)

Collapse
 
contentclips_st profile image
ContentClips •

The two-tools-one-disk gap (df 28% vs mongot 89%) is the part that will bite everyone: mongot enforces its own math, so the alert should too - a recording rule on mongot_system_disk_space_data_path_free_bytes / total_bytes catches the drift before any index build wedges, while df on a quota-limited overlay never will. Same asymmetry in the API: NOT_STARTED is a hard error but a typo'd index name is a silent zero, so the cheap guard is asserting the index exists via $listSearchIndexes at startup - turns 'identical from the application side' into a boot-time failure. And the worst-of-all-hosts-ever aggregation is really append-only truth without decay: a statusDetail-vs-actual-members diff is a one-line drift check on it.