Quick recap
This is Part 2 of a series — if you haven't read Part 1, the short version: every Go service at smritea.ai kept reinventing the same caching logic, so it got pulled into a small generic library called smartcache — a Cache[T] sitting on a swappable backend, with read-through and write-through built in. Building it surfaced three real problems, each fixed in a few lines: a cache stampede fixed with singleflight, synchronized TTL expiry fixed with jitter, and repeated database hits on keys that don't exist fixed with negative caching.
Part 1 closed on a gap none of that solved: a User that needs to be found three ways, by ID, by email, by slug. One row, three lookup paths. Cache it three times and an update has to land in three places at once — miss one and you're serving a stale email lookup for a user whose name already changed. Cache it once under the ID and the other two lookups get no benefit from caching at all. That's where this post picks up.
The trade-off at the center of this post rhymes with something more famous: CAP theorem, the idea that a distributed system can't maximize consistency, availability, and partition tolerance all at once — something has to give. What follows isn't CAP itself, but it's the same shape of problem.
Solving the three-keys gap turned into a small distributed-systems problem, and the fix ended up being two different answers depending on whether you're running one Redis instance or a cluster. Before getting into why, it's worth being precise about that CAP comparison, since it's easy to overstate.
Sounds like CAP theorem?
Loosely, yes — not literally. CAP is about consistency versus availability when a network partition splits your nodes apart. What follows here is a trade-off between two different things: keeping one entity's keys atomic under one Lua call, versus spreading those keys across a cluster so no single node carries all of one entity's traffic. There's no partition involved, and — as you'll see — the "weaker" option here doesn't actually give up correctness, only the guarantee that every key changes in one atomic step. It's the same kind of problem CAP made famous — you can't maximize two good properties in a distributed system at once, something has to give — just not CAP itself. Worth saying up front so nobody reads further expecting an actual CAP proof.
Store the value once, point at it from everywhere else
The fix for "one row, three keys" is pointer indirection: the value lives at exactly one key, and every alias is a small key that resolves to it instead of holding a copy.
bc:{user}:5 -> {"ID":"5","Name":"Ada",...} (the value, stored once)
bc:grp:{user}:email:foo@bar.com -> 5 (pointer -> primary key)
bc:grp:{user}:slug:ada -> 5 (pointer -> primary key)
A read through an alias is two hops: read the pointer, get the primary key, read the value key built from that primary key. Delete the value key and every alias breaks at the same instant — each one resolves to a key that's gone, which reads as a miss and reloads from the database. One value, no duplication, and eviction only has to happen in one place.
Why this needs a Lua script
Writing or evicting a group isn't a single command. Evict-by-alias has to: read the pointer, find the primary key it points to, read that record's list of registered aliases, then delete the value key and every one of those alias pointers. That's read, branch on what you read, then delete several keys — and Redis transactions (MULTI/EXEC) can't do that. They queue commands blind; there's no way to read a value mid-transaction and decide what to delete based on it. WATCH plus a retry loop can get you there, but it gets expensive fast under concurrent writes to the same entity.
The straightforward fix is a Lua script — the read-branch-delete sequence runs atomically on the Redis server itself, no client-side retry loop needed. Here's the whole read path for the single-instance version:
local pk = redis.call('GET', KEYS[1])
if not pk then return false end
local v = redis.call('GET', ARGV[1] .. pk)
if not v then return false end
return v
Pointer lookup, primary-key resolve, value read — one round trip, one atomic script.
The atomicity requirement creates a hotspot
To make that script atomic, every key belonging to one entity — the value, every alias pointer, the membership index — has to live on the same Redis Cluster node, because Lua scripts can only touch keys on one node at a time. Left alone, Redis Cluster hashes each key independently and scatters them across nodes by design — and a Lua call touching keys on different nodes doesn't degrade gracefully, it's rejected outright, the same CROSSSLOT error you'd get from any multi-key command spanning slots. Redis Cluster gives you a way to force keys together instead: a hash tag, the part of a key wrapped in {}, decides which slot the key lands on — everything else in the key name is ignored for hashing purposes. Tag everything for one entity with {user} and Redis Cluster puts it all on the same slot, which is exactly the constraint the atomic script needs.
// AliasColocated tags every key of an entity with {ns}: all of the entity's keys share one
// Cluster slot and every op is a single atomic Lua cascade. One slot per entity (hotspot risk
// on Cluster with a hot entity). Default.
AliasColocated AliasMode = iota
It works, and it's genuinely atomic — but it pins one entity's entire key set to one node. Redis's own clustering best practices post is direct about this: overusing hash tags gets you "an unbalanced cluster, or worse, one full node and many empty nodes," and the advice is to use them sparingly. A User with a normal amount of traffic on a cluster is fine under this scheme. A User that happens to be extremely active — or an entity type with a naturally hot member, like a popular org or a trending post — puts all of its traffic on one node no matter how many nodes the cluster has. You can verify this yourself with CLUSTER KEYSLOT: every key sharing a hash tag maps to the same slot number.
Two strategies, chosen per cache
Rather than pick one answer for every cache, smartcache exposes both and lets you choose:
type AliasMode int
const (
// AliasColocated tags every key of an entity with {ns}: all of the entity's keys share one
// Cluster slot and every op is a single atomic Lua cascade. One slot per entity (hotspot risk
// on Cluster with a hot entity). Default.
AliasColocated AliasMode = iota
// AliasSharded tags value+members with {ns:pk} (one slot per record) and reverse pointers with
// {ns:field:value} (one slot per alias), distributing load. GetByAlias is a two-hop resolve with
// validate-on-read; writes/evicts are a record-slot atomic Lua plus best-effort reverse-pointer
// ops. No hotspot, no stale reads, weaker cross-key atomicity (self-healing).
AliasSharded
)
The obvious question: if tagging everything with {user} gets you perfect atomicity in one Lua hop, why not just always do that? Because it permanently ties that entity's entire key set to whichever single node its hash tag happens to land on. Fine for most entities. Not fine for the one that turns out to be unusually busy or unusually large — there's no routing around it without changing the tag, and changing the tag is what gives up the atomicity in the first place.
Colocated is the default — one atomic call, correct everywhere, and fine as long as no single entity's traffic outgrows one node. Sharded spreads the value and each alias pointer across different slots, which means a read is now two network hops through two different nodes instead of one atomic call on one node — a small, real latency cost, paid in exchange for not having one node absorb an entire entity's traffic. That's the part that needs its own correctness argument.
Making the distributed version correct anyway
Spreading keys across slots means the atomic guarantee is gone — a pointer and the record it points to are no longer updated together. What has to not happen is a stale pointer returning the wrong value. The fix is checking the pointer's claim against the record before trusting it:
local v = redis.call('GET', KEYS[1])
if not v then return false end
if redis.call('HGET', KEYS[2], ARGV[1]) ~= ARGV[2] then return false end
return v
KEYS[1] is the value, KEYS[2] is the record's own membership index (which fields point at it, and with what value). If the record's index doesn't actually list this alias anymore, the script returns nothing — a miss, not a wrong answer. A pointer that's out of date because of a race just resolves to "not found" and gets reloaded. That's validate-on-read: it turns "the pointer might be stale" into "a stale pointer can only ever produce a safe miss."
Cleanup gets the matching treatment — a pointer is deleted only if it still points at the record being evicted:
if redis.call('GET', KEYS[1]) == ARGV[1] then
return redis.call('DEL', KEYS[1])
end
return 0
That's compare-and-delete. Without it, evicting user 5 right after their email got reassigned to user 99 could delete the pointer that now correctly belongs to user 99. With it, the delete only fires if the pointer still agrees with what's being evicted.
Worth being precise about what Sharded actually gives up, since it's less than it sounds like: not correctness — no caller ever gets a wrong value, validate-on-read rules that out — and not strong consistency of the data itself. What it gives up is atomicity across keys, and the price for that is an extra network hop per read plus, occasionally, a stale reverse-pointer sitting around until it's overwritten or compare-deleted — which shows up as a few extra cache misses while it self-heals, not as stale data reaching a caller. That's the actual trade: a small, bounded amount of latency and cache-miss noise, spent to keep one hot entity from owning a whole node.
The public interface both strategies implement is the same either way — Cache[T] never knows which one it's talking to:
type AliasOps interface {
GetValue(ctx context.Context, primary string) ([]byte, error)
PutValue(ctx context.Context, primary string, val []byte, ttl time.Duration) error
EvictByPrimary(ctx context.Context, primary string) error
GetByAlias(ctx context.Context, ref AliasRef) ([]byte, error)
PutByAlias(ctx context.Context, primary string, ref AliasRef, val []byte, ttl time.Duration) error
EvictByAlias(ctx context.Context, ref AliasRef) error
}
And picking a strategy is one field on registration:
ttl := time.Hour
sharded := smartcache.AliasSharded
cache, err := smartcache.RegisterAliasGroup[User](mgr, "user", &smartcache.EntityOptions{
TTL: &ttl,
AliasMode: &sharded, // omit this and you get AliasColocated, the default
})
Where this leaves things
Colocated for a single Redis instance, or a cluster where no entity is expected to get disproportionately hot.
Sharded once you're on a cluster and that assumption stops holding. Neither one is a half-measure.
Colocated is fully atomic by design, and Sharded is correct by construction even under races, just for a different set of guarantees.
There's one thing left on the table, and it's a real gap rather than a hypothetical one: every operation on an alias-group cache currently goes through the Lua path, including a plain lookup by primary key that was never part of any alias group to begin with. That's correct, just wasted work for the common case.
The fix on paper is a Bloom filter in front of the Lua call — check membership cheaply first, skip straight to a plain GET when a key was never registered as an alias. I have not built it yet. Because it's a pure performance optimisation with no correctness story attached to it.
And if you want to poke at the real code, smartcache is public. And stay tuned. I'm going to add more chapters to this story as well as to the actual project.
───────── ⋆⋅☆⋅⋆ ─────────
Thanks for reading this far! If it helped, I'd love for you to stick around — find me here:
Top comments (0)