DEV Community

Onyedikachi Emmanuel Nnadi
Onyedikachi Emmanuel Nnadi

Posted on

Why you should get the new edition of Designing Data-Intensive Application for Distributed System

A few days ago, I picked up the new edition of Designing Data-Intensive Applications by Martin Kleppmann.

I'm still in Chapter 1.

The book has been on my reading list for a while because I want to understand distributed systems properly, not just know the buzzwords.

The problem is that I've noticed something about how I learn.

I can read about a concept, understand it while reading, and then completely fail to recognize it when it shows up in code.

So instead of waiting until I finished the book, I decided to try something different.

I built a tiny URL shortener and deliberately introduced some of the problems DDIA talks about.

Not because I needed a URL shortner.

Not because I was trying to clone Bitly.

I just wanted to see if these problems would actually happen in a real application.

Spoiler: they did.

And a lot faster than I expected.

The First Thing I Learned: Caches Don't Magically Know Anything

One of the first things I added was a simple in-memory cache.

Nothing fancy.

Just a JavaScript Map.

Immediately, the application became faster because it wasn't hitting Postgres on every request.

At that point I was feeling pretty smart.

Then I changed a URL destination in the database.

The database had the new value.

The cache didn't.

The application kept serving the old destination as if nothing had changed.

The funny thing is that the cache wasn't broken.

It was doing exactly what it was supposed to do.

I was the one assuming that somehow it would know the underlying data had changed.

That small experiment made cache invalidation click for me in a way reading about it never did.

Then My Click Counter Started Losing Clicks

The second experiment was even more interesting.

I wanted every redirect to increment a click counter.

So I wrote the most obvious code possible:

Read the current value.

Add one.

Save it back.

I honestly expected it to work.

Then I hit it with 100 concurrent requests.

The final count wasn't 100.

It wasn't even close.

Most of the updates disappeared.

That was the first time I had seen a race condition happen in something I built myself.

Not in a diagram.

Not in a blog post.

Not in a textbook.

In my own terminal.

And suddenly the whole "multiple things trying to change the same data at the same time" explanation started making sense.

The Moment It Started Feeling Like a Distributed System

The experiment that tied everything together was running two copies of the application.

Same code.

Same database.

Different ports.

I updated a URL through one instance.

One cache knew about the change.

The other didn't.

Now two parts of my application disagreed about what the correct answer was.

My cache was just a JavaScript Map, so each Node process had its own private copy. The fix was swapping it for Redis because it supports shared memory. Both instances now read and write from the same cache, so when one invalidates a key, the other sees the change too.

That's when I finally understood why people keep saying distributed systems are hard.

Not because the code is complicated.

Because keeping multiple copies of reality synchronized is complicated.

*Why I'm Glad I Did This
*

I'm still at the beginning of DDIA.

I definitely wouldn't call myself someone who understands distributed systems yet.

But this small experiment changed how I'm reading the book.

Now when I see terms like cache invalidation, consistency, stale reads, or race conditions, I can connect them to something I've actually seen happen.

It's easy to read that a race condition can lose updates.

It's much harder to forget after watching 100 requests turn into 2.

Sometimes the fastest way to learn a concept is to break something small enough that you can understand why it broke.

That's probably what I'll keep doing as I work through the rest of the book.

Top comments (0)