The signup form that taught me to care about Bloom filters was not glamorous. People typed. We checked. Most answers were "available." Indexed database lookups are cheap one at a time. They are not free when every keystroke across a growing product becomes a read against the authoritative store. At some point the question stops being "can the database answer?" and becomes "should the database be the first system we ask every time?"
A Bloom filter sits in front of that path as a fast pre-check. It does not replace the database. It does not replace a UNIQUE constraint. It answers one useful question quickly: is this username definitely not in the set? If the filter says absent, skip the read. If it says maybe, confirm with the source of truth.
The personal stake is trust. Get the semantics wrong and you block free names, open ticket storms, and teach users that your product lies about availability.
Why direct username checks get expensive
SELECT 1 FROM users WHERE username = $1 LIMIT 1;
Caches help for popular names. They help less for random new strings nobody has typed before. That is the awkward traffic mix of a signup box: mostly negatives, often unique, always latency-sensitive in the UI.
What the filter actually promises
| Filter result | Meaning | Next step |
|---|---|---|
| Absent | Definitely not inserted | Treat as available for the read path |
| Present | Probably inserted | Confirm with the database |
| False positive | Present in filter, absent in DB | One extra lookup (acceptable) |
| False negative for inserted values | Should not happen if the filter was updated | Almost always means the filter is stale |
False positives are the trade. For username checks they are cheap. False negatives for names you already inserted are not acceptable on this path, which is why the write path must keep the filter honest.
Building the availability check as a story, not a party trick
func (s *Signup) Available(ctx context.Context, username string) (bool, error) {
u := normalize(username)
maybe, err := s.bloom.Exists(ctx, u)
if err != nil {
return s.db.UsernameFree(ctx, u) // Redis down → DB, never invent availability
}
if !maybe {
return true, nil
}
return s.db.UsernameFree(ctx, u)
}
On successful registration the sequence matters more than the data structure:
-
INSERTwith a UNIQUE constraint (the real guardian) -
BF.ADDthe username, or emit an event a worker applies
Skip step 2 and the filter drifts. Drift raises false positives (extra database hits). It does not invent false negatives for names that were correctly added earlier, unless you never added them at all.
The bug that looks like flaky signup
if bloom.Exists(username) {
return "already taken" // wrong
}
A false positive now blocks a free name. Users open tickets. Logs show Bloom hits with no matching row. Trust evaporates. The rule I write in every review is blunt: a Bloom "present" is a hint, never a verdict. The insert's unique constraint still prevents races when two clients check the same free name at once.
Production considerations that decide whether this stays elegant
Keep the database as authority. The filter speeds reads. The database decides writes. Always.
Keep the filter synchronized. Update on every successful signup. Rebuild periodically from a snapshot into a new key (bf:usernames:v2), dual-write briefly, then flip readers.
Plan for deletes and renames. Classic Bloom filters do not delete cleanly. Counting Bloom or Cuckoo filters exist, but complexity jumps. If renames are common, budget rebuilds or pick another structure.
Size for tomorrow. Estimate cardinality 18 to 24 months out. Pick a target false-positive rate (around 1% is common for pre-checks). Alert when fill factor climbs. Undersizing quietly kills the optimization as the fallback rate rises.
Normalize once. Lowercase, trim, and apply the same Unicode normalization everywhere. A filter cannot fix inconsistent identity rules.
When you should walk away
Walk away when exact answers are the product and there is no cheap confirmation path. Walk away when deletes dominate and rebuilds are not operationally owned. Walk away when the set is small enough that a Redis SET or the database alone is fine. Walk away when nobody owns metrics, rebuilds, or cutovers. Walk away when a false positive is user-visible and expensive, such as checkout lockouts or access denials.
Closing
Use a Bloom filter as a performance shield in front of a unique constraint, never as the authority. For username checks, "filter fast, confirm with DB" is the whole architecture. If your case cannot tolerate that confirmation step, choose a different tool. The structure is clever. Cleverness is not the same as correctness under concurrency.
Top comments (0)