DEV Community

Davi Orlandi
Davi Orlandi

Posted on

Lessons I Still Use from MongoDB in Action (Without Rewriting the Book)

I am not going to retell Kyle Banker's MongoDB in Action chapter by chapter. Read the book for that. What follows are field notes: the subset of ideas that still change how I model services, especially now that Atlas makes it easy to treat MongoDB like a broken relational database with friendlier braces.

The personal stake is quieter than an outage story and just as expensive. Most MongoDB pain I see is still mismatched modeling wearing a scaling costume. Teams buy bigger tiers to paper over documents shaped like leftover ER diagrams. These rules of thumb are how I try to stop myself from doing that.

Rule 1: Model for queries, not for your ER diagram

Documents are not normalized tables. Start from the questions your API asks. If you always fetch a user with prefs, embedding often wins. If addresses are updated independently and shared, references often win. If you need both, embed a summary and reference the full document.

// Embed when the access pattern is one read
{
  _id: userId,
  email: "a@b.com",
  prefs: { locale: "pt-BR", marketingOptIn: false }
}

// Reference when the child set is unbounded
{
  _id: orderId,
  userId: userId,
  total: 1200
}
Enter fullscreen mode Exit fullscreen mode

Cardinality matters. One-to-few, one-to-many, and one-to-squillions deserve different shapes. Huge embedded arrays that grow forever invent document-size pain, and then someone discovers the 16MB limit the hard way.

Rule 2: Atomicity lives at the document boundary

Single-document updates are atomic. Multi-document workflows need transactions or an outbox you designed on purpose. Prefer invariants that fit in one document. Use multi-document transactions when you must, not on every write out of fear. Pair state changes with outbox rows in the same transaction when you publish events. That last habit has saved me from more "eventual inconsistency" debates than any schema lecture.

Rule 3: Indexes are part of the schema

Do not treat indexes as a DBA afterthought. Write the filter plus sort shape first, then the index. Avoid an index-per-field shotgun that only creates write amplification. Prefer partial or sparse indexes when only a subset of documents need the path. Watch working set versus RAM. If you cannot read an explain plan on a hot path, you are guessing, and guessing is how "flexible schema" becomes "mysterious latency."

Rule 4: Aggregation is a product feature

The pipeline is not a party trick. Match early. Project away fat fields before heavy stages. Respect memory limits; spill to disk only when needed and safe. Do not reinvent joins with application-side N+1 if a better schema removes the need. Knowing when not to recurse in the database is modeling maturity, not a lack of ambition.

Rule 5: Favor embedding until a compelling reason appears

Useful compelling reasons still look the same years later. The child must stand alone in queries. The array would grow without bound. The read/write ratio makes denormalization too expensive to maintain. Application-level joins are allowed. With correct indexes and projections they are rarely the villain beginners fear. Denormalize only fields that are read often and updated rarely.

Rule 6: Flexible schema is not "no schema"

Dumping arbitrary JSON without ownership is an application smell. Share types in the application. Add server-side validation where it pays off. Keep an explicit schemaVersion when shapes evolve in place. Own field names across services. Flexibility without ownership is just entropy with nicer query syntax.

Patterns that still surprise teammates

Cached rollups: store comment counts near the parent instead of counting on every list read. Bucketing: time-series style buckets beat millions of tiny documents when telemetry arrives fast. Hot/cold honesty: separate cold data before the cluster becomes a museum of unread history.

What I leave to the book

Deep historical storage-engine lore and exhaustive admin tours belong in the text. The durable lessons I still carry are query-first modeling, index-as-schema, and respect for document atomicity.

Closing

Design the document around the query and the invariant, not around a relational habit. If that sounds obvious, good. The hard part is keeping the habit when a sprint pressure wants a quick collection "just like the SQL table." The book taught me the vocabulary. Production taught me why the vocabulary matters.

Top comments (0)