Part 1: Computer basics
CPU (Central Processing Unit)
Definition: The chip that executes instructions (calculations, comparisons, running your code).
Simple meaning: The brain that does the work.
Example: Parsing a request, running a loop, or computing a sum is CPU work.
RAM (Random Access Memory)
Definition: Fast, temporary memory where a running program keeps the data it is using right now.
Simple meaning: The desk where you keep the papers you're working on.
Example: Variables, arrays and maps in your code live in RAM.
SSD (Solid State Drive)
Definition: Permanent storage with no moving parts. Data stays when power is off.
Simple meaning: The cupboard where you store papers for the long term.
Example: Your files, installed apps, and database files.
Volatile vs Persistent
Volatile: data is lost when power is off or the process restarts (RAM).
Persistent: data survives restarts and power loss (SSD).
Latency
Definition: The time one operation takes from start to finish.
Simple meaning: How long you wait for one answer.
Operation Rough latency
Read from RAM ~100 nanoseconds
Read from SSD ~100 microseconds
Network call to Redis ~0.5 to 1 millisecond
Database query ~5 to 50 milliseconds (varies a lot)
(1 millisecond = 1,000 microseconds = 1,000,000 nanoseconds)
Yes, your laptop, EC2 server and database server all have CPU, RAM and SSD.
Part 2: Why a local map works on your laptop but breaks in production
Hash map (map / dictionary)
Definition: A data structure that stores data as key → value pairs and finds a value by its key very quickly (usually O(1)).
Simple meaning: A phonebook: look up the name (key), get the number (value).
Example: cache["user:42"] = {name: "Ravi"}
Cache
Definition: A fast storage layer that keeps copies of frequently used data so you don't have to fetch it from the slow source again.
Simple meaning: A shortcut. Keep answers you need often in RAM instead of asking the database every time.
Cache hit and cache miss
Cache hit: the data is found in the cache (fast).
Cache miss: the data is not in the cache, so you go to the database (slow), then usually save it to the cache.
Hit ratio: hits ÷ total requests. A higher ratio means a better cache.
Synchronous vs Asynchronous
Synchronous: one task finishes before the next starts.
Asynchronous: a task is started, and the program continues doing other things while waiting.
A single request on your laptop is simple, so a map works. Production is different because of concurrency.
Concurrency
Definition: Many tasks in progress at overlapping times.
Simple meaning: 100 users sending requests at the same moment.
Note: Concurrency is not the same as synchronous. Your server is concurrent. The problem isn't that one request is synchronous.
Thread and Goroutine
Thread: the smallest unit of execution the OS schedules. A server uses many threads to handle many requests.
Goroutine: Go's lightweight version of a thread, managed by Go itself. A web server in Go usually runs each request in its own goroutine.
Race condition
Definition: A bug where the result depends on the timing of two or more threads touching the same data.
Example: Two goroutines both read count = 5, both add 1, and both write 6. The answer should be 7. One update is lost.
Problem 1: Concurrent map access
In Go, writing to a map from two goroutines at once causes a fatal error (concurrent map writes) and the whole process crashes. (Node.js runs JS on one thread, so it doesn't crash this way, but logical races can still occur across await calls.)
Mutex (Mutual Exclusion lock)
Definition: A lock that allows only one goroutine at a time into a protected section of code.
Simple meaning: A bathroom key. Whoever has the key goes in, and everyone else waits.
sync.RWMutex (Read-Write mutex): many readers can read together (RLock), but a writer needs exclusive access (Lock).
Thread-safe: code that behaves correctly when used by many threads at once.
Note: A mutex doesn't make requests synchronous. It makes only the small locked part sequential.
Problem 2: Memory grows forever
A map has no size limit. Keep adding keys and RAM fills up.
OOM (Out Of Memory): the OS kills your process because RAM ran out.
Memory leak: memory that keeps growing and is never released.
Eviction and LRU
Eviction: automatically removing old data from a cache when it is full.
LRU (Least Recently Used): remove the item that hasn't been used for the longest time.
Other policies: LFU (Least Frequently Used), FIFO (First In First Out), random.
Problem 3: Stale data (you missed this)
TTL (Time To Live): a timer on a key. After it ends, the key is deleted automatically.
Stale data: a cached copy that is out of date compared to the database.
Example: SET user:42 data EX 60 means the key expires after 60 seconds.
To build a real cache by hand you need Mutex + LRU eviction + TTL. This is hard and bug-prone.
Part 3: Why not use the EC2 or RDS RAM as the cache?
Typical setup
EC2 (app server): runs your Go/Node code.
RDS (database server): runs PostgreSQL/MySQL.
Each has its own CPU, RAM and SSD. RAM is never shared between machines.
Load balancer
Definition: A component that spreads incoming requests across multiple servers.
Simple meaning: A traffic police officer sending each car to a different lane.
Horizontal scaling
Definition: Handling more traffic by adding more servers (instead of making one server bigger).
Vertical scaling: making one server bigger (more CPU/RAM).
Issue A: Each server has its own separate RAM
User A's data is cached on EC2 #1.
The next request goes to EC2 #2, which has nothing, so that's a cache miss.
The same data ends up cached 10 times.
Worse, the copies can disagree, which is called inconsistency. EC2 #1 has old data and EC2 #2 has new data.
For sessions, a miss means the user is logged out. For normal data, it only means an extra database trip.
Session: data the server remembers about a logged-in user between requests.
Issue B: Restarts wipe the cache
Deploys, crashes and auto-scaling erase RAM.
Cold cache: an empty cache after a restart.
Thundering herd (cache stampede): after a cold start, thousands of requests miss at once and all hit the database together.
Why not rely on RDS's own RAM?
RDS does cache data pages in RAM automatically. But a database is built for correctness, not raw speed:
ACID: Atomicity, Consistency, Isolation, Durability. Guarantees that transactions are safe and complete.
Durability: once committed, data is saved to disk even if power fails. This forces disk writes.
Query overhead: each query is parsed, planned and executed, with joins, filters and locks costing CPU.
Connection limit: a database allows only a limited number of simultaneous connections. A connection pool reuses connections to avoid opening new ones.
Hard to scale: adding database servers is much harder than adding app servers.
Under 100,000 simple queries, RDS typically doesn't crash. It slows down or runs out of connections, and then your whole app slows down.
Conclusion: We need one shared, central, in-memory cache that all EC2 servers talk to over the network. That is Redis.
Part 4: What is Redis?
Redis (REmote DIctionary Server) is an in-memory data store that works like a hash map running on its own server, accessible over the network. It has these features built in:
Thread-safe by design: Redis executes commands one at a time (single-threaded command execution), so no race conditions on your data.
Eviction policies: LRU, LFU, etc., so memory never grows forever.
TTL: keys expire automatically.
Shared: all EC2 servers see the same data.
Optional persistence: it can save data to disk to survive restarts.
Tradeoffs to know
Network cost: Redis is a separate server, so it's slower than local RAM (~1 ms vs ~100 ns), but far faster than the database.
Single point of failure (SPOF): if the one Redis server dies, everything depending on it breaks. The fix is replication (keeping copies on other servers).
Still volatile by default, unless persistence is enabled.
Cache invalidation: deciding when to delete or update cached data after the database changes. This is famously one of the hardest problems in computing.
Top comments (0)