<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Saqib Ameen Subhan</title>
    <description>The latest articles on DEV Community by Saqib Ameen Subhan (@saqibameen86).</description>
    <link>https://dev.to/saqibameen86</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4049802%2F2d78cd6c-e644-4178-86b6-6a4638bfe5a4.jpg</url>
      <title>DEV Community: Saqib Ameen Subhan</title>
      <link>https://dev.to/saqibameen86</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/saqibameen86"/>
    <language>en</language>
    <item>
      <title>They did it, and I meant it</title>
      <dc:creator>Saqib Ameen Subhan</dc:creator>
      <pubDate>Wed, 30 Sep 2026 10:24:14 +0000</pubDate>
      <link>https://dev.to/saqibameen86/they-did-it-and-i-meant-it-4ipi</link>
      <guid>https://dev.to/saqibameen86/they-did-it-and-i-meant-it-4ipi</guid>
      <description>&lt;p&gt;In an interview, I was asked about the achievement I was most proud of. I talked about a migration my team did, moving a high traffic product from RDS MySQL to Aurora Serverless with zero downtime.&lt;/p&gt;

&lt;p&gt;Then the interviewer asked what my role in it was. I gave an answer that can feel risky in an interview. I said I did not build it, my engineers did. My job was managing the stakeholders, getting the approvals, and making sure nothing got in their way.&lt;/p&gt;

&lt;p&gt;I know the usual interview advice is to take more of the credit, to say things like "I designed the approach" or "I made the key decisions." Some of that would even be partly true. But the answer I gave is how I actually think about management, and it would not mean much if I only said it when it was easy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the credit belongs to the team
&lt;/h2&gt;

&lt;p&gt;The way I see it, a manager's work is the team's results. That is the job. If a manager takes personal credit for something the team delivered, they are counting the same work twice. Once as "my team delivered this" and again as "I delivered this." Only one of those is true.&lt;/p&gt;

&lt;p&gt;The engineer who wrote the cutover script built the cutover. What I did was create the conditions for it. I got the support to do it properly instead of quickly, I made sure more than one person could run it, and I kept the stakeholders calm while the work was happening. That is real work and I am happy to take credit for exactly that. But the migration belongs to them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What giving credit actually looks like
&lt;/h2&gt;

&lt;p&gt;Saying "great job, team" in a meeting is nice, but it does not change anything for anyone. The credit that matters is the credit that reaches the people who decide promotions and careers.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Their names go in the email to leadership.&lt;/strong&gt; Not "the team delivered this," but the actual names of the people who did it. Leadership should know who the strong engineers are without having to ask me.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;They present their own work.&lt;/strong&gt; When leadership wants a walkthrough of the migration, the engineer who built it presents it. I am in the room only to handle any difficult questions, not to explain their work for them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;My name comes up when something breaks.&lt;/strong&gt; This is the other half of the same rule. Credit goes down to the team, and responsibility for problems comes up to me. You cannot do one without the other.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  You can only give credit for work you let them own
&lt;/h2&gt;

&lt;p&gt;There is one condition that people usually skip. You cannot give someone credit for work you did not let them own.&lt;/p&gt;

&lt;p&gt;If the manager makes every important decision, reviews every line and solves every hard problem before the team gets to it, then "the credit goes to my team" is just something nice to say. Everyone in the room knows who really did it. To give real credit, you first have to hand over real ownership, with real responsibility and real production risk, before you know how it will turn out. That migration was my engineers' achievement because it was genuinely theirs to get wrong.&lt;/p&gt;

&lt;p&gt;This is also why it builds over time. People who get real credit for real ownership take on more ownership next time. That is how two engineers I hired as interns went on to lead teams of their own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cost of doing it this way
&lt;/h2&gt;

&lt;p&gt;In the short term, this makes you less visible to leadership as an individual. While you are putting your engineers' names in the email, another manager might be presenting their team's work as their own, and for a while it may work for them. That can be frustrating to watch.&lt;/p&gt;

&lt;p&gt;I will not pretend it always balances out quickly. But when a manager's team keeps doing well, quarter after quarter, people start to notice who built that team. It takes longer, but it is the honest way to get there.&lt;/p&gt;

&lt;p&gt;The migration is still one of my proudest achievements, and it still belongs to them.&lt;/p&gt;

</description>
      <category>management</category>
      <category>engineering</category>
      <category>leadership</category>
      <category>career</category>
    </item>
    <item>
      <title>I stopped being the leave referee</title>
      <dc:creator>Saqib Ameen Subhan</dc:creator>
      <pubDate>Wed, 30 Sep 2026 10:20:07 +0000</pubDate>
      <link>https://dev.to/saqibameen86/i-stopped-being-the-leave-referee-3bc4</link>
      <guid>https://dev.to/saqibameen86/i-stopped-being-the-leave-referee-3bc4</guid>
      <description>&lt;p&gt;Three engineers, three festivals, and one on call rota that has to be covered every single day.&lt;/p&gt;

&lt;p&gt;Every year, Diwali, Christmas and New Year fall within a few weeks of each other. Every database team I have run has been small enough that not everyone can be on leave at the same time. For a long time I handled this the way most managers are taught to. The leave requests came to me, I looked at them, and I decided.&lt;/p&gt;

&lt;p&gt;I tried hard to be fair. I kept notes on who got which holidays last year, and I balanced it as carefully as I could. But every time, my fair and well documented decision ended the same way. Someone did not get the days they wanted, and they were unhappy with me for it.&lt;/p&gt;

&lt;p&gt;The problem was not how fair I was. The problem was that I was the one deciding. When a manager decides, someone always loses, and that person does not blame the situation. They blame the decision, because the decision has the manager's name on it. You can be as fair as you like and still end the quarter with someone on the team quietly unhappy with you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I do now
&lt;/h2&gt;

&lt;p&gt;I set the rule, and I let the team decide the calendar.&lt;/p&gt;

&lt;p&gt;The rule is about coverage, not about people, and it is the only part I keep. There must always be enough people available to cover production. Production does not stop for festivals. Who takes which days, as long as that rule holds, is not my decision anymore.&lt;/p&gt;

&lt;p&gt;The first time I did this, I expected to be pulled back in within a day. Instead, the team just talked it out between themselves. One of them said, "You take Diwali, it matters more to you than Christmas does to me. I will take Christmas. I will cover New Year this time if I can have it next year." It took about ten minutes, and the calendar was done.&lt;/p&gt;

&lt;p&gt;What is different here is that nobody lost. When I decided, one person lost. When they agreed on it themselves, both of them owned the decision. The engineer covering New Year is not unhappy with me, because he made a trade he thinks is fair, with a colleague who now owes him one. They both remember the deal better than any spreadsheet I could have kept.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it keeps working
&lt;/h2&gt;

&lt;p&gt;When I was the one making sure things were fair, it felt like I was keeping score on everyone. When the team does it themselves, it becomes a habit they follow on their own. They remember who covered last winter without anyone telling them. They sort it out before I even see the calendar. They even protect each other's leave when other teams ask for help during holidays, which is something no rule of mine ever managed.&lt;/p&gt;

&lt;p&gt;There was one more thing I did not expect. Working it out together brought the team closer. When someone gives up their days so that a colleague can be with family for a festival, it changes something between those two people. They have shown each other, not just said it, that they will cover for each other. That is exactly what you need later during an incident, when covering for each other is the whole job.&lt;/p&gt;

&lt;p&gt;The real message underneath all this is simple. I trust you to sort this out like the professionals you are. Letting them decide is itself a way of showing respect, and people notice when they are treated like adults. They give that respect back.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this does not work
&lt;/h2&gt;

&lt;p&gt;This does not work in every team, and I would not be honest if I said it did.&lt;/p&gt;

&lt;p&gt;It only works if the team already feels safe speaking up. If you hand the calendar to a team where one loud person dominates, that person will take December every year, and the quieter people will tell themselves they did not really mind. Letting the team decide does not create fairness. It makes whatever culture you already have stronger. If the team respects each other, you get the kind of conversation I described. If they do not, you have just made the unfairness automatic and removed the one person who could have stopped it.&lt;/p&gt;

&lt;p&gt;So the order matters. Build respect in the team first, and only then let them decide these things themselves. I still keep two things with me. The coverage rule, and stepping in if a deal ever breaks that rule. In all the years I have run teams this way, I have almost never had to step in.&lt;/p&gt;

&lt;p&gt;The calendar gets sorted on its own now. Nobody has missed a festival in years, and nobody thanks me for it, which is how it should be.&lt;/p&gt;

</description>
      <category>management</category>
      <category>engineering</category>
      <category>teams</category>
      <category>career</category>
    </item>
    <item>
      <title>You can buy attendance, You can buy hours, But the feeling that "this system is mine" is Not for Sale!</title>
      <dc:creator>Saqib Ameen Subhan</dc:creator>
      <pubDate>Wed, 30 Sep 2026 09:08:29 +0000</pubDate>
      <link>https://dev.to/saqibameen86/you-can-buy-attendance-you-can-buy-hours-but-the-feeling-that-this-system-is-mine-is-not-for-3k8b</link>
      <guid>https://dev.to/saqibameen86/you-can-buy-attendance-you-can-buy-hours-but-the-feeling-that-this-system-is-mine-is-not-for-3k8b</guid>
      <description>&lt;p&gt;The night an entire cloud region went down under us, I watched something that no on call policy can produce. Engineers joining the incident bridge who were not on call, had not been paged, and had every right to be asleep.&lt;/p&gt;

&lt;p&gt;Nobody summoned them. The recovery ran for thirty hours, and the people who owned those systems treated the outage the way you treat a fire in your own house. Not because a rota said so, but because somewhere along the way, the production had become &lt;em&gt;theirs&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;The lazy explanation is great work ethic, which is a way of crediting the individuals so you do not have to examine the system. I have managed long enough to know better. Work ethic is real, but it does not explain why the &lt;em&gt;same&lt;/em&gt; engineers who show up at 2 AM on one team quietly ghost their pagers on another. The variable is not the people.&lt;/p&gt;

&lt;p&gt;The 2 AM behaviour was decided months earlier, at 2 PM, when nothing was burning.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually earns it
&lt;/h2&gt;

&lt;p&gt;It is not perks and it is not speeches. In my experience it is four specific, boring, daily things.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Their leave is sacred.&lt;/strong&gt; Nobody gets called on vacation. We cross train specifically so that is true, and the team settles its own leave calendar like adults. An engineer whose rest is protected does not ration their energy defensively. They give it freely, because they have learned it will not be stolen.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Credit flows down, in public.&lt;/strong&gt; When the project lands, their names go in the email to leadership, and they present their own work upward. My name appears when something breaks. An engineer who gets the credit for a system starts to feel what is literally true. It is &lt;em&gt;their&lt;/em&gt; system.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;They are in the decisions.&lt;/strong&gt; Tooling, process, rota design. The team decides, with full context, because they live inside the consequences. You cannot hand someone the outcomes and withhold the choices and then act surprised that they feel like a bystander.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;They are never resources.&lt;/strong&gt; Not in meetings, not in planning docs, not in my head. The word matters because the word leaks into behaviour. Resources get allocated. People get respected. Only one of those shows up at 2 AM.&lt;/p&gt;

&lt;p&gt;None of this is charity. It is a chain with no skippable links. Respect makes the production &lt;em&gt;theirs&lt;/em&gt;, ownership follows, and once ownership exists, accountability stops being a conversation you ever need to have. I have never once had to ask an owner whether they will handle something. The question does not survive contact with genuine ownership.&lt;/p&gt;

&lt;p&gt;You can buy attendance. You can buy hours, and with enough money you can even buy compliance. The feeling that &lt;em&gt;this system is mine&lt;/em&gt; is not for sale. It is reciprocated or it is absent.&lt;/p&gt;

&lt;h2&gt;
  
  
  The boundary that keeps this honest
&lt;/h2&gt;

&lt;p&gt;Here is where I have to be careful, because this philosophy has an evil twin.&lt;/p&gt;

&lt;p&gt;If a manager reads "my people show up at 2 AM" and hears &lt;em&gt;free capacity&lt;/em&gt;, they have turned respect into exploitation with better branding. The willingness is a gift. The moment you treat it as a budget line, planning around it, expecting it, spending it, you are strip mining the very trust that created it, and it does not grow back at the rate you are extracting it.&lt;/p&gt;

&lt;p&gt;So the same philosophy comes with obligations that run in my direction. I answer first at 2 AM. The leader eats the first hour before anyone else is asked to. Incident nights are followed by comp time off without anyone having to request it. And I watch for the pattern where the same engineer is the hero twice in a row, because the person who always shows up is one quarter away from being the person who used to work here.&lt;/p&gt;

&lt;p&gt;Protect the gift and it compounds. Spend it and it is gone.&lt;/p&gt;

&lt;h2&gt;
  
  
  The morning after
&lt;/h2&gt;

&lt;p&gt;Somewhere past hour thirty, the estate was back, the region was healing, and everyone finally got some well deserved rest.&lt;/p&gt;

&lt;p&gt;The next morning I did not assign the postmortem. Drafts were already appearing in the channel, from the people who had been on the bridge, unprompted, because owners do not wait to be asked what happened to their own systems.&lt;/p&gt;

&lt;p&gt;That is the whole thing, really. It was never about the night.&lt;/p&gt;

&lt;p&gt;It was about all the days before it.&lt;/p&gt;

</description>
      <category>management</category>
      <category>engineering</category>
      <category>sre</category>
      <category>oncall</category>
    </item>
    <item>
      <title>A junior engineer accidentally dropped a staging database!</title>
      <dc:creator>Saqib Ameen Subhan</dc:creator>
      <pubDate>Tue, 15 Sep 2026 05:20:48 +0000</pubDate>
      <link>https://dev.to/saqibameen86/a-junior-engineer-accidentally-dropped-a-staging-database-jf0</link>
      <guid>https://dev.to/saqibameen86/a-junior-engineer-accidentally-dropped-a-staging-database-jf0</guid>
      <description>&lt;p&gt;A junior engineer on my team once ran a cleanup command on staging, believing he was on dev. The data was gone before the prompt returned.&lt;/p&gt;

&lt;p&gt;He came to me directly, shaking, honestly, and told me what happened. Which, I want to point out before anything else, is the single most important fact in this story.&lt;/p&gt;

&lt;p&gt;On a lot of teams, that engineer spends an hour trying to quietly fix it himself first, and the incident gets worse in the dark. He came straight to me. That is not a personality trait. That is a property of the team, and you build it or you destroy it long before the incident happens.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first hour
&lt;/h2&gt;

&lt;p&gt;The first hour is triage, not analysis.&lt;/p&gt;

&lt;p&gt;I took over communication with the affected dev team myself. Not because the engineer could not speak, but because the person with the least organizational armor should not be the face of an incident. We stabilized, restored from backup, verified, and closed the loop with the dev team the same day.&lt;/p&gt;

&lt;p&gt;At no point did anyone ask who did it in a public channel. The dev team was told: &lt;em&gt;we&lt;/em&gt; made a mistake, &lt;em&gt;we&lt;/em&gt; have restored it, here is what &lt;em&gt;we&lt;/em&gt; are changing.&lt;/p&gt;

&lt;p&gt;Accountability flows up. A manager who forwards blame downward is just a router.&lt;/p&gt;

&lt;h2&gt;
  
  
  The interesting question
&lt;/h2&gt;

&lt;p&gt;Once it was over, the question I sat with was not how do I make sure he is more careful.&lt;/p&gt;

&lt;p&gt;Careful is not a system. Everyone is careful until they are tired, or it is 2 AM, or two terminal windows look identical, which is the actual root cause here. He was not careless. He was a human being using a tool that made two very different worlds look exactly the same.&lt;/p&gt;

&lt;p&gt;So we changed the tools.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Terminal backgrounds now encode the environment.&lt;/strong&gt; Dev is default. Staging is blue. Production is red, a red you cannot ignore, a red that makes you sit up slightly before you type. It costs nothing and it works on the tired version of you, which is the version that makes mistakes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A pre execution checklist for destructive commands.&lt;/strong&gt; Three questions, ten seconds. &lt;em&gt;Where am I? What system is this? Am I on a read node or a write node?&lt;/em&gt; We wrote it into the runbook.&lt;/p&gt;

&lt;p&gt;I told the team the analogy I actually think in. You cut vegetables outside the vessel, then put them in. You do not chop directly into the cooking pot, because there, every slip is dinner.&lt;/p&gt;

&lt;p&gt;Neither fix is clever. That is the point. Clever fixes depend on people being sharp. Boring fixes work when they are not.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happened to the engineer
&lt;/h2&gt;

&lt;p&gt;Nothing. That is the answer, and it was deliberate and visible. No formal note, no development area in his next review, nothing.&lt;/p&gt;

&lt;p&gt;The incident was the process's fault, the process got fixed, and the fix was named after the problem, not the person.&lt;/p&gt;

&lt;p&gt;He stayed. He got better, genuinely better, the kind of better that comes from learning viscerally that mistakes are survivable here. Years later he was one of the people I trusted most in a production window.&lt;/p&gt;

&lt;p&gt;If I had made that first incident expensive for him, I would have taught him, and everyone watching, because everyone is always watching, that the smart move next time is to hide it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The general rule
&lt;/h2&gt;

&lt;p&gt;When something breaks, there are always two available explanations. The person, or the system the person was inside.&lt;/p&gt;

&lt;p&gt;Choosing the person is emotionally satisfying and fixes nothing. The next human inherits the same trap. Choosing the system is boring and permanent.&lt;/p&gt;

&lt;p&gt;We never had that class of incident again. Not because the team became more careful, but because the trap was removed.&lt;/p&gt;

&lt;p&gt;The terminal colors outlived everyone on that team, including me.&lt;/p&gt;

</description>
      <category>management</category>
      <category>engineering</category>
      <category>sre</category>
      <category>career</category>
    </item>
    <item>
      <title>Large partitions don't fail loudly. They fail everywhere at once.</title>
      <dc:creator>Saqib Ameen Subhan</dc:creator>
      <pubDate>Fri, 04 Sep 2026 17:29:00 +0000</pubDate>
      <link>https://dev.to/saqibameen86/large-partitions-dont-fail-loudly-they-fail-everywhere-at-once-2j41</link>
      <guid>https://dev.to/saqibameen86/large-partitions-dont-fail-loudly-they-fail-everywhere-at-once-2j41</guid>
      <description>&lt;p&gt;Most Cassandra problems announce themselves. A timeout, an exception, a graph going vertical.&lt;/p&gt;

&lt;p&gt;Large partitions do not. They degrade reads, compaction, repair and GC &lt;em&gt;simultaneously and gradually&lt;/em&gt;, which is why the incident channel fills with four unrelated looking symptoms and no suspect.&lt;/p&gt;

&lt;p&gt;If you have been following this series, you have already met large partitions twice without the name. They are the humongous allocations in the heap dump post, and they are the overstreaming amplifier in the repair post. This post is about the disease itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What large means and why it hurts
&lt;/h2&gt;

&lt;p&gt;A partition is the unit Cassandra reads, compacts, streams and materializes. The rule of thumb that has served me well is to &lt;strong&gt;aim under 100MB per partition&lt;/strong&gt; and treat anything past that as debt. Old timers remember the hard 2GB era limits. Modern Cassandra handles big partitions better, which mostly means they hurt you more gradually.&lt;/p&gt;

&lt;p&gt;Where the pain lands:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reads.&lt;/strong&gt; A partition read materializes structures proportional to what is scanned. In G1 terms, a multi hundred MB partition read is a parade of humongous allocations. GC pause spikes on whichever node holds it, the cluster marks it slow, and speculative retries fan the load elsewhere.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compaction.&lt;/strong&gt; Partitions are compacted as units. Giant partitions make giant, slow compaction tasks that starve neighbours, and pending tasks climb.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repair.&lt;/strong&gt; One mismatched cell in a giant partition can stream the whole thing. The overstreaming problem, concentrated.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hotspotting.&lt;/strong&gt; Big partitions are usually &lt;em&gt;hot&lt;/em&gt; partitions, the same key taking disproportionate traffic. Size and heat compound.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How partitions get big, always the same story
&lt;/h2&gt;

&lt;p&gt;The partition key models an unbounded thing.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- looks innocent in the design review&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;events_by_device&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;device_id&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;ts&lt;/span&gt; &lt;span class="nb"&gt;timestamp&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="nb"&gt;blob&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;device_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A partition key of &lt;code&gt;device_id&lt;/code&gt; means every event that device ever produces, forever, in one partition. A chatty device, a popular customer, a busy trading day. Growth has no ceiling because the model gave it none.&lt;/p&gt;

&lt;p&gt;I have seen the one big customer partition at multiple companies. It is practically a genre.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finding them
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;nodetool tablehistograms ks events_by_device
&lt;span class="c"&gt;# Percentile      Partition Size&lt;/span&gt;
&lt;span class="c"&gt;# 50%             61,214 bytes&lt;/span&gt;
&lt;span class="c"&gt;# 99%             386,857,368 bytes     &amp;lt;- there it is&lt;/span&gt;
&lt;span class="c"&gt;# Max             2,395,318,855 bytes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The p50 to p99 gap is the signature. A healthy median hiding a monster tail.&lt;/p&gt;

&lt;p&gt;To name the offenders on disk:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# per SSTable, list partitions over a threshold&lt;/span&gt;
sstablepartitions &lt;span class="nt"&gt;-t&lt;/span&gt; 100 /var/lib/cassandra/data/ks/events_by_device-&lt;span class="k"&gt;*&lt;/span&gt;/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And Cassandra logs them at write time. Grep &lt;code&gt;system.log&lt;/code&gt; for &lt;code&gt;Writing large partition&lt;/code&gt;. Those log lines are the cheapest early warning system you will ever ignore.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: give the partition a ceiling
&lt;/h2&gt;

&lt;p&gt;You cannot cap a device's lifetime events. You &lt;em&gt;can&lt;/em&gt; cap a partition, by putting time into the key.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;events_by_device_v2&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;device_id&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;day&lt;/span&gt; &lt;span class="nb"&gt;date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                       &lt;span class="c1"&gt;-- the bucket&lt;/span&gt;
  &lt;span class="n"&gt;ts&lt;/span&gt; &lt;span class="nb"&gt;timestamp&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="nb"&gt;blob&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt; &lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;device_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;day&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;ts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now a partition is one device, one day. Bounded by physics instead of hope. Two design decisions follow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bucket size is a calculation, not a vibe.&lt;/strong&gt; Estimate rows per day times row size, then pick the bucket, hour, day or week, that lands typical partitions in single digit MB and your worst realistic case under about 100MB. Do the arithmetic for your &lt;em&gt;loudest&lt;/em&gt; tenant, not your average one. Averages are how the enormous customer surprises you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multi bucket reads are now yours to orchestrate.&lt;/strong&gt; Last 6 hours spanning midnight means two partitions. Drivers make querying a known list of buckets cheap. Just design the access pattern alongside the schema, not after it. This pairs well with TWCS from the compaction post, because bucketed, TTL'd time series is the workload TWCS was born for.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Migration reality.&lt;/strong&gt; Existing giants must be rewritten into v2. Dual write new data, backfill old in throttled batches, the same discipline as the 2TB MongoDB delete earlier in this series, chunk, sleep, watch the cluster. Then cut reads over and drop v1. Weeks, calmly, not a weekend, heroically.&lt;/p&gt;

&lt;h2&gt;
  
  
  The review question that prevents all of it
&lt;/h2&gt;

&lt;p&gt;Every Cassandra schema review I run ends with one question. &lt;em&gt;What bounds this partition?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;If the answer is a business quantity, customers will not have many orders, then it is unbounded, because business quantities grow. That is their job. The only acceptable answers have units of time, or a hard modulo in the key.&lt;/p&gt;

&lt;p&gt;Large partitions are never an ops problem. They are a design review that ended one question too early.&lt;/p&gt;

</description>
      <category>cassandra</category>
      <category>datamodeling</category>
      <category>performance</category>
      <category>database</category>
    </item>
    <item>
      <title>Repair at 1,000 nodes: why we stopped running nodetool repair by hand</title>
      <dc:creator>Saqib Ameen Subhan</dc:creator>
      <pubDate>Fri, 04 Sep 2026 17:28:30 +0000</pubDate>
      <link>https://dev.to/saqibameen86/repair-at-1000-nodes-why-we-stopped-running-nodetool-repair-by-hand-289i</link>
      <guid>https://dev.to/saqibameen86/repair-at-1000-nodes-why-we-stopped-running-nodetool-repair-by-hand-289i</guid>
      <description>&lt;p&gt;At small scale, repair is a cron job. At 1,000 nodes it is a scheduling problem, and if you treat a scheduling problem like a cron job, the cluster will teach you the difference during business hours.&lt;/p&gt;

&lt;p&gt;I spent years operating Cassandra estates past the thousand node mark. Here is what repair actually does, why it is not optional, and how the operational model has to change with scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why repair is a deadline, not a chore
&lt;/h2&gt;

&lt;p&gt;Cassandra replicas drift. Hinted handoff covers short outages, but hints have a window, three hours by default. A replica down longer than that has permanently missed writes, and only &lt;strong&gt;anti entropy repair&lt;/strong&gt; reconciles it.&lt;/p&gt;

&lt;p&gt;The part that turns repair from hygiene into a deadline is tombstones. Deleted data is protected from resurrection by grave markers that compaction may purge after &lt;code&gt;gc_grace_seconds&lt;/code&gt;, ten days by default. The safety of that purge rests on one assumption. &lt;strong&gt;Every replica heard about the delete before the marker vanished.&lt;/strong&gt; Repair is how they hear.&lt;/p&gt;

&lt;p&gt;So the rule is absolute. &lt;strong&gt;Every node completes repair at least once per &lt;code&gt;gc_grace_seconds&lt;/code&gt;.&lt;/strong&gt; Miss it, and a replica that slept through a delete will happily hand the missing row back to the cluster. Zombie data, silent, and by the time anyone notices, days of writes have been made against resurrected garbage.&lt;/p&gt;

&lt;p&gt;Repair is not cleanup. It is the second half of every delete you have ever issued.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a repair costs
&lt;/h2&gt;

&lt;p&gt;Mechanically, nodes build &lt;strong&gt;Merkle trees&lt;/strong&gt; over their data ranges, exchange them, and stream any ranges whose hashes disagree. Three costs follow.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Building trees reads data.&lt;/strong&gt; Sequential I/O competing with your workload.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Streaming mismatches&lt;/strong&gt; saturates network and creates new SSTables, which means the third cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compaction debt.&lt;/strong&gt; A big repair is followed by a compaction hangover, and &lt;code&gt;nodetool compactionstats&lt;/code&gt; pending tasks tell the story.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Merkle tree resolution is finite, so one tree over a huge range means one tiny mismatch streams a disproportionately large chunk. This is &lt;strong&gt;overstreaming&lt;/strong&gt;, and it is why repairing a giant range in one shot is both slow and wasteful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Full, incremental, subrange, and what breaks at scale
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Full repair&lt;/strong&gt;, whole range, everything rehashed every time. Correct, brutal, does not scale. At 1,000 nodes, naive &lt;code&gt;nodetool repair&lt;/code&gt; across the fleet means the cluster is effectively always repairing, and repair sessions colliding on shared ranges fail in tedious ways.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Incremental repair&lt;/strong&gt; marks SSTables repaired so they are skipped next time. On paper, the fix. In practice it drags &lt;strong&gt;anticompaction&lt;/strong&gt; behind it, rewriting SSTables to segregate repaired from unrepaired data, and its operational history, particularly before Cassandra 4.0's fixes, was rocky enough that many large operators simply did not trust it. We did not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Subrange repair&lt;/strong&gt; breaks the token range into small segments repaired one at a time. Small Merkle trees, precise comparisons, minimal overstreaming, and each segment is a small retryable unit of work. This is the primitive that actually scales.&lt;/p&gt;

&lt;p&gt;But subrange repair across 1,000 nodes is thousands upon thousands of segments needing ordering, throttling, retries and collision avoidance. Congratulations, it is a scheduling problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reaper: repair as a system, not a command
&lt;/h2&gt;

&lt;p&gt;We ran &lt;strong&gt;Cassandra Reaper&lt;/strong&gt;, open source, originally Spotify's and now community maintained, as the scheduler.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Splits every table's range into &lt;strong&gt;segments&lt;/strong&gt; and runs them steadily with concurrency caps&lt;/li&gt;
&lt;li&gt;An &lt;strong&gt;intensity&lt;/strong&gt; knob throttles repair pressure so p99s do not feel it. We tuned it to stay invisible.&lt;/li&gt;
&lt;li&gt;Failed segments retry individually, so a node blip costs one segment, not the whole run&lt;/li&gt;
&lt;li&gt;A web UI that answers the only question leadership asks, which is whether we are inside the gc_grace window, yes or no&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The operating policy that came out of years of this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- every table repaired well inside gc_grace, target a complete cycle in ~7 days against a 10 day grace
- repair pressure tuned to be invisible in p99 read latency
- alert not on "repair failed" but on "time since last completed repair per table"
- pause repairs during topology changes, resume, never skip
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last alert framing matters. Failures are noise. &lt;em&gt;Staleness&lt;/em&gt; is the actual risk. Measure the deadline, not the activity.&lt;/p&gt;

&lt;h2&gt;
  
  
  The small scale corollary
&lt;/h2&gt;

&lt;p&gt;Under about 50 nodes you do not need Reaper's ceremony, but you do need the same two invariants. Every node repaired inside &lt;code&gt;gc_grace_seconds&lt;/code&gt;, and repairs that do not collide. A boring, monitored schedule beats heroics.&lt;/p&gt;

&lt;p&gt;Repair is the tax on eventual consistency. You can pay it on a schedule you chose, or all at once with interest, with zombie data as the collection notice.&lt;/p&gt;

</description>
      <category>cassandra</category>
      <category>operations</category>
      <category>distributed</category>
      <category>database</category>
    </item>
    <item>
      <title>The heap dump usually exonerates Cassandra</title>
      <dc:creator>Saqib Ameen Subhan</dc:creator>
      <pubDate>Fri, 04 Sep 2026 17:26:34 +0000</pubDate>
      <link>https://dev.to/saqibameen86/the-heap-dump-usually-exonerates-cassandra-406o</link>
      <guid>https://dev.to/saqibameen86/the-heap-dump-usually-exonerates-cassandra-406o</guid>
      <description>&lt;p&gt;When a Cassandra node OOMs, or GC pauses start stretching into seconds, the reflex is to blame Cassandra, or its favourite scapegoat, the JVM.&lt;/p&gt;

&lt;p&gt;I have read a lot of Cassandra heap dumps, and here is the uncomfortable pattern. The heap usually contains exactly what the data model asked for. The dump does not convict the database. It convicts a partition.&lt;/p&gt;

&lt;p&gt;This post is the workflow I actually use, in order, because the order saves hours.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 0: GC logs before heap dumps
&lt;/h2&gt;

&lt;p&gt;A heap dump is a biopsy. GC logs are the patient history, and they are already on disk. Two things I look for first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# gc.log (G1)
Pause Young (Normal) ... 180ms
Pause Young (Normal) ... 210ms
Pause Full (Allocation Failure) ... 8.4s      &amp;lt;- the incident
...
Humongous Allocation ... 18MB
Humongous Allocation ... 22MB
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Full GC pauses.&lt;/strong&gt; G1 doing an emergency stop the world collection because normal cycles could not keep up. On a Cassandra node, an 8 second pause does not just slow queries. The node misses gossip, gets marked down by peers, and hints start piling up elsewhere. One node's GC problem becomes the cluster's problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Humongous allocations.&lt;/strong&gt; G1 divides the heap into regions, typically 4 to 32MB. Any single allocation larger than &lt;strong&gt;half a region&lt;/strong&gt; is humongous, allocated awkwardly across contiguous regions and collected inefficiently.&lt;/p&gt;

&lt;p&gt;And what does Cassandra allocate in one giant contiguous chunk? A large partition being materialized for a read. Humongous allocation warnings in a Cassandra gc.log are large partitions announcing themselves. You can often skip the heap dump entirely at this point and go straight to &lt;code&gt;nodetool tablehistograms&lt;/code&gt; to find the table with a monster p99 partition size.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: take the dump without causing an incident
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;jmap&lt;/code&gt; on a live heap &lt;strong&gt;stops the world for the duration of the dump&lt;/strong&gt;. On a 20GB heap, that is long enough for the cluster to declare the node dead. Never dump a node that is still serving.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# isolate first, node stays alive but stops serving&lt;/span&gt;
nodetool disablebinary &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; nodetool disablegossip

jmap &lt;span class="nt"&gt;-dump&lt;/span&gt;:live,format&lt;span class="o"&gt;=&lt;/span&gt;b,file&lt;span class="o"&gt;=&lt;/span&gt;/data/dumps/cass_&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;hostname&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;_&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; +%s&lt;span class="si"&gt;)&lt;/span&gt;.hprof &amp;lt;pid&amp;gt;

&lt;span class="c"&gt;# afterwards: rejoin or just restart the node&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Better, have the JVM do it for you at the moment of truth, before any human is awake:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="py"&gt;-XX&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s"&gt;+HeapDumpOnOutOfMemoryError&lt;/span&gt;
&lt;span class="py"&gt;-XX&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s"&gt;HeapDumpPath=/data/dumps/&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those two flags belong in every production Cassandra JVM config. An OOM without a dump is an incident you get to have twice.&lt;/p&gt;

&lt;p&gt;Related discipline, keep the heap at or under 8GB and let G1 target around 200ms pauses. Bigger heaps mostly buy you longer pauses and bigger dumps, and Cassandra's real caching happens off heap and in the page cache anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Eclipse MAT, dominator tree, ten minutes
&lt;/h2&gt;

&lt;p&gt;Open the &lt;code&gt;.hprof&lt;/code&gt; in Eclipse MAT and go straight to the &lt;strong&gt;dominator tree&lt;/strong&gt;, objects ranked by retained memory, which answers what is actually holding the heap hostage. Ignore the histogram of a billion &lt;code&gt;byte[]&lt;/code&gt;. The dominator tree tells you &lt;em&gt;whose&lt;/em&gt; bytes.&lt;/p&gt;

&lt;p&gt;What the tree shows, mapped to what it means, every one of these from real incidents:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dominating the heap&lt;/th&gt;
&lt;th&gt;Actual problem&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A few enormous &lt;code&gt;byte[]&lt;/code&gt; or cell containers under a read path&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Large partition&lt;/strong&gt; being materialized. Get the key from the referencing objects, confirm with &lt;code&gt;nodetool tablehistograms&lt;/code&gt; or &lt;code&gt;sstablepartitions&lt;/code&gt;. Fix is to bucket the partition, which is a later post in this series.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memtable objects, many tables&lt;/td&gt;
&lt;td&gt;Over wide schema, or memtable flush thresholds set too generous for the heap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tombstone or range tombstone structures under a query&lt;/td&gt;
&lt;td&gt;A tombstone farm read that survived just long enough to dump&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Netty buffers or inflight requests&lt;/td&gt;
&lt;td&gt;Clients hammering with no paging. Check fetch size and unpaged full partition reads&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The punchline repeats across years and companies. &lt;strong&gt;The heap contains the workload.&lt;/strong&gt; I have almost never opened a dump and found a Cassandra bug. I have regularly opened one and found a 4GB partition someone swore was just a busy customer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: close the loop in the model
&lt;/h2&gt;

&lt;p&gt;The dump names the object. The fix lives in the schema.&lt;/p&gt;

&lt;p&gt;Large partition, time bucket the partition key. Unpaged reads, driver fetch size. Tombstone reads, the modeling fixes from two posts ago. The JVM flags and MAT are diagnosis. Cassandra data modeling is treatment.&lt;/p&gt;

&lt;p&gt;Blaming the JVM is comfortable because nobody owns the JVM. The heap dump takes that comfort away. It shows you, in retained bytes, precisely which design decision you are looking at.&lt;/p&gt;

</description>
      <category>cassandra</category>
      <category>jvm</category>
      <category>performance</category>
      <category>debugging</category>
    </item>
    <item>
      <title>A compaction strategy is a bet about your workload</title>
      <dc:creator>Saqib Ameen Subhan</dc:creator>
      <pubDate>Fri, 04 Sep 2026 17:25:24 +0000</pubDate>
      <link>https://dev.to/saqibameen86/a-compaction-strategy-is-a-bet-about-your-workload-200k</link>
      <guid>https://dev.to/saqibameen86/a-compaction-strategy-is-a-bet-about-your-workload-200k</guid>
      <description>&lt;p&gt;Choosing a Cassandra compaction strategy is placing a bet. Each strategy optimizes one thing by deliberately sacrificing another, and the honest way to choose is to decide &lt;strong&gt;which failure mode you would rather operate&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I have run all three majors across 1,000+ node estates. Here is the bet each one makes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why compaction exists at all
&lt;/h2&gt;

&lt;p&gt;Cassandra never updates in place. Writes land in a memtable, flush to immutable SSTables, and a row's fragments accumulate across many files as it is updated over time. Reads must merge every relevant fragment.&lt;/p&gt;

&lt;p&gt;Compaction is the background process that merges SSTables, consolidating fragments and purging expired tombstones after &lt;code&gt;gc_grace_seconds&lt;/code&gt;, so reads touch fewer files.&lt;/p&gt;

&lt;p&gt;The metric that keeps compaction honest:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;nodetool tablehistograms ks table
&lt;span class="c"&gt;# Percentile  SSTables     ...&lt;/span&gt;
&lt;span class="c"&gt;# 50%             1.00&lt;/span&gt;
&lt;span class="c"&gt;# 99%             4.00&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;SSTables per read at p99 is the number compaction exists to keep small.&lt;/p&gt;

&lt;h2&gt;
  
  
  STCS, SizeTiered, the write optimized default
&lt;/h2&gt;

&lt;p&gt;Merges SSTables of similar size into bigger ones, repeatedly. Minimal write amplification, maximal ingest throughput.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The bet you are making:&lt;/strong&gt; reads and disk headroom are negotiable.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A hot row's fragments can sit in many SSTables of different generations, so SSTables per read climbs.&lt;/li&gt;
&lt;li&gt;Space amplification is the operational trap. Compacting large tiers requires holding input &lt;em&gt;and&lt;/em&gt; output simultaneously. Plan for &lt;strong&gt;50% free disk&lt;/strong&gt;. I have watched teams treat 70% disk usage as plenty left, then be unable to run the very compaction that would reclaim space. That is a corner with no good exits.&lt;/li&gt;
&lt;li&gt;Tombstones buried in giant, rarely recompacted SSTables can linger far past &lt;code&gt;gc_grace_seconds&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Right when:&lt;/strong&gt; write heavy, append mostly, reads tolerate variance. It is the default because it is the least dangerous &lt;em&gt;average&lt;/em&gt; bet, not because it is right for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  LCS, Leveled, paying writes to buy reads
&lt;/h2&gt;

&lt;p&gt;Organizes SSTables into levels, L1, L2 and so on, each 10 times larger, guaranteeing non overlapping key ranges within a level. A read touches at most one SSTable per level, and in practice the vast majority of reads are satisfied by a single SSTable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The bet:&lt;/strong&gt; you will pay roughly &lt;strong&gt;10x write amplification&lt;/strong&gt;, because every row is rewritten as it migrates down levels, to make read latency tight and predictable.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;On read heavy tables with strict p99s, it is the correct bet.&lt;/li&gt;
&lt;li&gt;On write heavy tables, compaction falls behind, pending tasks pile up in &lt;code&gt;nodetool compactionstats&lt;/code&gt;, and ironically reads degrade anyway because L0 accumulates overlapping files.&lt;/li&gt;
&lt;li&gt;Easier on disk headroom than STCS, works in around 10% free, but harder on I/O and CPU, continuously.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Right when:&lt;/strong&gt; read dominated, update in place workloads. Wrong when your disks are already busy keeping up with ingest.&lt;/p&gt;

&lt;h2&gt;
  
  
  TWCS, TimeWindow, the time series specialist
&lt;/h2&gt;

&lt;p&gt;Groups SSTables by time window, say one per day. Within the current window, STCS as usual. Once a window closes, its SSTables compact once into one file and are &lt;strong&gt;never touched again&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;When TTL expires a whole window, Cassandra drops the entire file, so tombstone processing effectively vanishes for aged out data. This is the single biggest tombstone fix available for time series, and the punchline of the previous post.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The bet:&lt;/strong&gt; your data is truly time ordered and immutable once written.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Write out of order data or update old windows and you create cross window overlaps that can never compact away. The strategy's guarantees quietly die while everything still appears to work.&lt;/li&gt;
&lt;li&gt;Rule of thumb, aim for a window size that yields &lt;strong&gt;20 to 50 windows&lt;/strong&gt; over the table's TTL.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Right when:&lt;/strong&gt; metrics, events, telemetry, anything with TTL and append only semantics. IoT and monitoring tables at scale should almost always be TWCS.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bet, stated plainly
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;You optimize&lt;/th&gt;
&lt;th&gt;You pay with&lt;/th&gt;
&lt;th&gt;Operational trap&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;STCS&lt;/td&gt;
&lt;td&gt;write throughput&lt;/td&gt;
&lt;td&gt;read variance&lt;/td&gt;
&lt;td&gt;needs around 50% free disk to compact&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LCS&lt;/td&gt;
&lt;td&gt;read p99&lt;/td&gt;
&lt;td&gt;10x write amplification&lt;/td&gt;
&lt;td&gt;falls behind under heavy ingest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TWCS&lt;/td&gt;
&lt;td&gt;TTL'd time series&lt;/td&gt;
&lt;td&gt;inflexibility&lt;/td&gt;
&lt;td&gt;out of order writes break it silently&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Per table, not per cluster. A keyspace legitimately mixes all three. And changing the bet is an online operation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;ks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="n"&gt;compaction&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s1"&gt;'class'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'TimeWindowCompactionStrategy'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
   &lt;span class="s1"&gt;'compaction_window_unit'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'DAYS'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'compaction_window_size'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cassandra will re sort existing SSTables over time. On a big table, schedule the transition like the migration it is.&lt;/p&gt;

&lt;p&gt;There is no best compaction strategy. There is only the failure mode you have chosen on purpose, versus the one that chose you.&lt;/p&gt;

</description>
      <category>cassandra</category>
      <category>performance</category>
      <category>database</category>
      <category>operations</category>
    </item>
    <item>
      <title>mtools does not parse MongoDB 4.4+ JSON logs. So I built mdbkit.</title>
      <dc:creator>Saqib Ameen Subhan</dc:creator>
      <pubDate>Mon, 24 Aug 2026 19:28:11 +0000</pubDate>
      <link>https://dev.to/saqibameen86/mtools-does-not-parse-mongodb-44-json-logs-so-i-built-mdbkit-5a29</link>
      <guid>https://dev.to/saqibameen86/mtools-does-not-parse-mongodb-44-json-logs-so-i-built-mdbkit-5a29</guid>
      <description>&lt;p&gt;If you have operated MongoDB for any length of time, you know mtools. For years it was how we all read production logs. &lt;code&gt;mloginfo&lt;/code&gt;, &lt;code&gt;mplotqueries&lt;/code&gt;, and a slow query was suddenly visible.&lt;/p&gt;

&lt;p&gt;Then MongoDB 4.4 changed the log format to structured JSON, and the log tools stopped understanding what they were reading.&lt;/p&gt;

&lt;p&gt;That was six years ago. The support issue was opened in June 2020 and it is still open. We are on 8.0 now, and a lot of us have been getting by with &lt;code&gt;grep&lt;/code&gt;, &lt;code&gt;jq&lt;/code&gt; pipelines held together by hope, and a certain amount of squinting.&lt;/p&gt;

&lt;p&gt;I got tired of it and built the replacement. It is called mdbkit, it is MIT licensed, and it is public as of this week.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it in sixty seconds without touching a cluster
&lt;/h2&gt;

&lt;p&gt;This is the part I would want first if this were someone else's tool.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--upgrade&lt;/span&gt; mdbkit
&lt;span class="c"&gt;# macOS ships no pip: brew install pipx &amp;amp;&amp;amp; pipx install mdbkit&lt;/span&gt;

mdbkit demo &lt;span class="nt"&gt;-o&lt;/span&gt; demo.log
mdbkit queries demo.log
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;mdbkit demo&lt;/code&gt; generates a realistic MongoDB log containing an actual incident. No cluster, no connection string, nothing of yours involved. You get to judge the tool on a problem you can inspect before you point it at anything that matters.&lt;/p&gt;

&lt;p&gt;Here is what &lt;code&gt;queries&lt;/code&gt; gives you on that demo log:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;saqib@MacbookPro mongodb_logs % mdbkit queries demo.log
== mdbkit queries (slow query shapes) ==
&lt;/span&gt;&lt;span class="gp"&gt;lines: 1,004  parsed: 1,004  unparsed: 0  span: 2026-07-01T08:00:00+04:00 -&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;2026-07-01T09:29:44.615000+04:00
&lt;span class="go"&gt;
namespace      op         count  cumMs  mean   max    docsEx      scan     plan            shape
-------------  ---------  -----  -----  -----  -----  ----------  -------  --------------  ---------------------------------------------
shop.events    aggregate  53     5.6m   6.3s   8.9s   47,170,000  98889:1  COLLSCAN+SORT   {tenantId:eq, ts:gte} sort:{ts:-1}
shop.orders    find       89     2.4m   1.6s   2.4s   11,125,000  2976:1   COLLSCAN+SORT   {createdAt:gt, status:eq} sort:{createdAt:-1}
shop.sessions  find       58     34.7s  597ms  888ms  2,784,000   155:1    IXSCAN{active}  {active:eq, lastSeen:gt} sort:{lastSeen:-1}
shop.products  update     42     25.7s  612ms  799ms  2,268,000   54000:1  COLLSCAN        {sku:eq}
shop.users     find       56     6.7s   119ms  139ms  56          1:1      IXSCAN{email}   {email:eq}

cumMs  = total wall time accumulated across ALL occurrences (not one query)
docsEx = total documents examined across all occurrences
scan   = docsExamined per returned doc (high = index missing or weak)
&lt;/span&gt;&lt;span class="gp"&gt;plan   = most common query plan;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;COLLSCAN/+SORT &lt;span class="o"&gt;=&lt;/span&gt; index needed
&lt;span class="go"&gt;         empty plan (?) = plan not present in log (below slowms threshold)
&lt;/span&gt;&lt;span class="gp"&gt;next   : mdbkit advise &amp;lt;log&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="nt"&gt;--ns&lt;/span&gt; &amp;lt;namespace&amp;gt;] &lt;span class="k"&gt;for &lt;/span&gt;index candidates
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is not a screenshot I curated. &lt;code&gt;demo&lt;/code&gt; is seeded, so those two commands produce that exact table on your machine too, down to the last digit. It behaves the same on my laptop, on yours, and on a conference projector.&lt;/p&gt;

&lt;p&gt;The column that matters is the scan ratio. It is documents examined against documents returned, per query shape. A shape sitting at 98889:1 is reading roughly ninety nine thousand documents to hand back one. A shape at 1:1 is doing exactly the work it needs to.&lt;/p&gt;

&lt;p&gt;That single ratio is most of slow query triage. Everything else is detail.&lt;/p&gt;

&lt;h2&gt;
  
  
  The workflow I actually use
&lt;/h2&gt;

&lt;p&gt;Three commands, in order.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, find the shape that is hurting.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;mdbkit queries /var/log/mongodb/mongod.log
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Slow operations grouped by query shape rather than listed individually. This matters more than it sounds. A log with 40,000 slow query lines is usually eight or nine distinct shapes repeating, and the one costing you the most is rarely the one that appears most often.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Then, ask what would fix it.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;mdbkit advise /var/log/mongodb/mongod.log
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This proposes candidate indexes based on the shapes it observed. More on how that works below, because it is the part people should be suspicious of.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And when something has already gone wrong, build the timeline.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;mdbkit triage /var/log/mongodb/mongod.log
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Elections, rollbacks, stalls, connection storms, in order, with timestamps. When you are twenty minutes into an incident and someone senior is asking what happened, a timeline you did not have to assemble by hand is worth a great deal.&lt;/p&gt;

&lt;h2&gt;
  
  
  The log cannot tell you everything
&lt;/h2&gt;

&lt;p&gt;There is one failure mode a mongod log structurally cannot explain: the log simply stops, then starts again a minute later with no error in between.&lt;/p&gt;

&lt;p&gt;The process was killed before it could write anything. That answer lives in the system log, one line up from where you were looking.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;mdbkit triage mongod.log &lt;span class="nt"&gt;--oslog&lt;/span&gt; /var/log/syslog
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[CRIT] Process start(s) in window: mongod startup marker at 09:14:35
[CRIT] System: oom-kill: The kernel OOM killer terminated a task.
         1 occurrence (last at 09:14:03). Processes: mongod.
         next: This explains an unexplained restart in the log above.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;oslog&lt;/code&gt; reads &lt;code&gt;/var/log/syslog&lt;/code&gt; or &lt;code&gt;/var/log/messages&lt;/code&gt; and looks for the things that end database processes: OOM kills, file descriptor limits, segmentation faults, I/O errors, filesystems remounted read only, systemd service exits. Used with &lt;code&gt;triage&lt;/code&gt;, the unexplained restart gets matched to the kernel line that caused it.&lt;/p&gt;

&lt;p&gt;On journald systems there is no text log to read, and mdbkit will not run commands on your behalf, so it prints the &lt;code&gt;journalctl&lt;/code&gt; invocation for you to capture and feed back in.&lt;/p&gt;

&lt;p&gt;There is also &lt;code&gt;ftdc&lt;/code&gt;, which decodes &lt;code&gt;diagnostic.data&lt;/code&gt; offline. Every one of your nodes has been recording CPU, cache, queue depth and disk every second since the day you installed it, whether or not you run any monitoring. Plus &lt;code&gt;serverstatus&lt;/code&gt; for making sense of a &lt;code&gt;serverStatus&lt;/code&gt; dump, and &lt;code&gt;compare&lt;/code&gt; for answering whether the index you added last week actually helped.&lt;/p&gt;

&lt;h2&gt;
  
  
  It never connects to your database
&lt;/h2&gt;

&lt;p&gt;No driver. No URI. No network code anywhere in it.&lt;/p&gt;

&lt;p&gt;mdbkit reads log files you hand it. It is strictly read only. Where an action would help, it prints the command for you to review and run yourself, rather than running anything on your behalf.&lt;/p&gt;

&lt;p&gt;I built it this way because I would not run a stranger's tool on a production host, and I do not expect you to either. So rather than asking you to trust me, the README shows you how to verify it independently. &lt;code&gt;grep&lt;/code&gt; the source for socket and driver imports. Run it under &lt;code&gt;strace&lt;/code&gt; and watch it make no network calls. It takes about a minute and you should do it.&lt;/p&gt;

&lt;p&gt;Zero runtime dependencies too, nothing in the supply chain beyond the Python standard library. That was a deliberate constraint and it cost me some convenience, but a tool that ships onto database servers should not drag a dependency tree behind it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The index advice is rules, not AI
&lt;/h2&gt;

&lt;p&gt;I gave a talk at the Dubai MongoDB User Group recently about how confidently language models describe a MongoDB that no longer exists. They will happily reference &lt;code&gt;plannerVersion&lt;/code&gt;, a field removed in 5.0. They do not know about EXPRESS stages in 8.0. And more importantly, they cannot know your cardinality, your write mix, or what is currently sitting in your plan cache, because none of that is in your query text.&lt;/p&gt;

&lt;p&gt;So mdbkit's index advice is deterministic. Rules over observed query shapes. The same log always produces the same recommendation.&lt;/p&gt;

&lt;p&gt;Every suggestion comes with the evidence it reasoned from, a confidence level, and how to validate it before you act. It says candidate, never command. And it will never tell you to drop an index, because deciding an index is unused requires knowing about traffic that a log file cannot show you.&lt;/p&gt;

&lt;p&gt;That is a deliberately narrow promise. I would rather it be right about a small thing than confident about a large one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What would help most
&lt;/h2&gt;

&lt;p&gt;Two things, and they are both easy for you and hard for me.&lt;/p&gt;

&lt;p&gt;Real world log lines that parse incorrectly. Every MongoDB deployment logs slightly differently, and I have only seen the estates I have worked on. If mdbkit chokes on a line, that line is the most useful thing you can send me.&lt;/p&gt;

&lt;p&gt;Index advice that is wrong or unhelpful. If it suggests something you know is a bad idea for your workload, I want to know why, because that is a rule that needs fixing.&lt;/p&gt;

&lt;p&gt;Every release so far has come from someone doing exactly that. An FTDC decode that took twenty five minutes and pinned a CPU. A crash on a batched write. A restart with no explanation, which is why &lt;code&gt;oslog&lt;/code&gt; exists at all.&lt;/p&gt;

&lt;p&gt;Supports MongoDB 4.4 through 8.0. Python 3.8+.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Source: &lt;a href="https://github.com/saqibameen86/mdbkit" rel="noopener noreferrer"&gt;https://github.com/saqibameen86/mdbkit&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Package: &lt;a href="https://pypi.org/project/mdbkit" rel="noopener noreferrer"&gt;https://pypi.org/project/mdbkit&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Free, MIT, and built because the gap annoyed me for long enough.&lt;/p&gt;




&lt;p&gt;If you liked this, I also wrote about &lt;a href="https://dev.to/saqibameen86/how-to-delete-2tb-from-a-live-mongodb-cluster-without-anyone-noticing-1pe2"&gt;deleting 2TB from a live MongoDB cluster without anyone noticing&lt;/a&gt; and &lt;a href="https://dev.to/saqibameen86/i-take-down-a-healthy-primary-on-purpose-you-should-too-bm7"&gt;why I take down a healthy primary on purpose&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>mongodb</category>
      <category>opensource</category>
      <category>database</category>
      <category>cli</category>
    </item>
    <item>
      <title>In Cassandra, a delete is a write (and reads pay for it)</title>
      <dc:creator>Saqib Ameen Subhan</dc:creator>
      <pubDate>Sun, 23 Aug 2026 16:36:09 +0000</pubDate>
      <link>https://dev.to/saqibameen86/in-cassandra-a-delete-is-a-write-and-reads-pay-for-it-5mb</link>
      <guid>https://dev.to/saqibameen86/in-cassandra-a-delete-is-a-write-and-reads-pay-for-it-5mb</guid>
      <description>&lt;p&gt;In Cassandra, a delete is a write.&lt;/p&gt;

&lt;p&gt;That sentence sounds like trivia until it takes down a read path. &lt;code&gt;DELETE&lt;/code&gt; doesn't remove anything. It &lt;em&gt;writes&lt;/em&gt; a &lt;strong&gt;tombstone&lt;/strong&gt;, a marker saying "this data is dead as of timestamp T." The actual removal happens later, during compaction, and only after a grace period. Between those two moments, your deleted data has negative value. It's gone from your application's point of view but still physically present, and every read that touches its range must process it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the marker has to exist
&lt;/h2&gt;

&lt;p&gt;Cassandra is a distributed, eventually consistent system. Suppose a delete simply removed data, and one replica was down when the delete happened. When it comes back, it still has the row, and to the rest of the cluster, that looks like &lt;em&gt;data the others are missing&lt;/em&gt;. Repair would helpfully copy the deleted row back to everyone.&lt;/p&gt;

&lt;p&gt;Deleted data returning from the dead is called &lt;strong&gt;zombie data&lt;/strong&gt;, and tombstones exist to prevent it. The tombstone outranks the older value everywhere it's seen.&lt;/p&gt;

&lt;p&gt;Which is why tombstones must survive for &lt;code&gt;gc_grace_seconds&lt;/code&gt;, &lt;strong&gt;default 864000, ten days&lt;/strong&gt;, before compaction may purge them. Ten days is the window you have to repair a down replica. Shrink &lt;code&gt;gc_grace_seconds&lt;/code&gt; without shrinking your repair interval and you've quietly signed up for zombies. Repair at scale is its own post, and it's coming.&lt;/p&gt;

&lt;h2&gt;
  
  
  How reads pay
&lt;/h2&gt;

&lt;p&gt;A read merges data across memtable and SSTables. Every tombstone in the requested range must be read, held, and reconciled against live data before Cassandra can answer. A partition that has accumulated a million tombstones makes you &lt;em&gt;process a million dead cells to return whatever's alive&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Cassandra tells you when this is happening, in two escalating tones:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;WARN  ReadCommand - Read 812 live rows and 104,832 tombstone cells for query ...
ERROR ... Scanned over 100001 tombstones ... query aborted
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The warning fires past &lt;code&gt;tombstone_warn_threshold&lt;/code&gt;, 1,000 by default. Past &lt;code&gt;tombstone_failure_threshold&lt;/code&gt;, 100,000, the read is killed mid flight with &lt;code&gt;TombstoneOverwhelmingException&lt;/code&gt;. That's the database choosing to fail your query rather than let it OOM the node.&lt;/p&gt;

&lt;p&gt;The thresholds are guardrails, not tuning knobs. Raising them treats the symptom and keeps the disease.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where tombstone farms come from
&lt;/h2&gt;

&lt;p&gt;Every one of these I've met in production.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Queue like workloads.&lt;/strong&gt; Insert, process, delete, repeat, in the same partition. The partition becomes a graveyard the consumer must scan past on every poll. Cassandra as a queue is the canonical anti pattern for exactly this reason.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Collection overwrites.&lt;/strong&gt; &lt;code&gt;UPDATE t SET mymap = {...}&lt;/code&gt; replaces a whole collection, which writes a range tombstone over the old one first. Prefer additive updates, &lt;code&gt;mymap = mymap + {...}&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inserting NULLs.&lt;/strong&gt; Binding null in a prepared statement writes a tombstone for that cell. ORMs and "just bind every column" code generate these invisibly. Use unset values, not nulls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TTL everywhere.&lt;/strong&gt; Expired TTL cells become tombstones too. Fine &lt;em&gt;if&lt;/em&gt; your compaction strategy is built for it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Finding them, then fixing them
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# per-SSTable estimate&lt;/span&gt;
sstablemetadata /var/lib/cassandra/data/ks/table-&lt;span class="k"&gt;*&lt;/span&gt;/ &lt;span class="k"&gt;*&lt;/span&gt;&lt;span class="nt"&gt;-Data&lt;/span&gt;.db | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; droppable
&lt;span class="c"&gt;# Estimated droppable tombstones: 0.83&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;0.83 means an estimated 83% of that SSTable is purgeable tombstones.&lt;/p&gt;

&lt;p&gt;For query level visibility, &lt;code&gt;TRACING ON&lt;/code&gt; in cqlsh shows tombstone cells scanned per query. For a live table, &lt;code&gt;nodetool tablestats&lt;/code&gt; and watch average tombstones per slice.&lt;/p&gt;

&lt;p&gt;The durable fixes are modeling fixes.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Time series with TTL, use TimeWindowCompactionStrategy.&lt;/strong&gt; Whole SSTables age out and get dropped as files, so tombstone processing largely disappears. Compaction strategy trade offs are the next post.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stop deleting in place in hot partitions.&lt;/strong&gt; Partition by time bucket and let whole partitions expire instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit for null binding and collection overwrites.&lt;/strong&gt; These are the tombstones nobody meant to write.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Deletes in Cassandra are cheap to issue and expensive to have issued. Model as if every delete is a small loan against your read latency, because it is, and &lt;code&gt;gc_grace_seconds&lt;/code&gt; is the repayment schedule.&lt;/p&gt;

</description>
      <category>cassandra</category>
      <category>database</category>
      <category>performance</category>
      <category>distributed</category>
    </item>
    <item>
      <title>How to delete 2TB from a live MongoDB cluster without anyone noticing</title>
      <dc:creator>Saqib Ameen Subhan</dc:creator>
      <pubDate>Sun, 23 Aug 2026 16:04:47 +0000</pubDate>
      <link>https://dev.to/saqibameen86/how-to-delete-2tb-from-a-live-mongodb-cluster-without-anyone-noticing-1pe2</link>
      <guid>https://dev.to/saqibameen86/how-to-delete-2tb-from-a-live-mongodb-cluster-without-anyone-noticing-1pe2</guid>
      <description>&lt;p&gt;The request was simple. "There is about 2TB of expired data in this collection. Can you delete it tonight?"&lt;/p&gt;

&lt;p&gt;My answer was no, and being able to explain why is most of the job. A missing TTL index had let short lived data pile up for months in one of our busiest collections, and the team wanted to run one big deleteMany overnight. Here is what that would have actually done.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a big delete hurts, and it is not locking
&lt;/h2&gt;

&lt;p&gt;Most people think a big delete is risky because it locks the collection. In modern MongoDB it does not, because WiredTiger works at the document level, so a huge delete does not freeze the collection. The damage comes from three other places.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The oplog fills up.&lt;/strong&gt; Every deleted document becomes its own entry in the oplog (the log that secondaries copy from). Delete 500 million documents and you have written 500 million oplog entries, and every secondary has to copy and apply each one of them. Replication lag starts climbing. If a secondary falls further behind than the oplog can hold, it can no longer catch up and needs a full resync. At that point your delete has cost you a copy of your data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The cache gets flushed out.&lt;/strong&gt; To delete a document, WiredTiger first has to load it into memory. A 2TB delete pulls a lot of old data into the cache and pushes out the data your application actually uses. Queries that have nothing to do with the delete start getting slow, and whoever looks into it will probably never connect it back to the delete.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Majority writes slow down.&lt;/strong&gt; If your writes use w:"majority" (and they should, I covered why in an earlier post in this series), every write now waits on the same secondaries that are already struggling to keep up with the delete. The problems add up.&lt;/p&gt;

&lt;p&gt;I showed this to the team in a lower environment. I started the big delete, ran &lt;code&gt;rs.printSecondaryReplicationInfo()&lt;/code&gt;, and let them watch the lag grow. After that, "tonight" turned into a plan.&lt;/p&gt;

&lt;h2&gt;
  
  
  The approach: small chunks, a pause, and watching the lag
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;CHUNK&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;10000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;SLEEP_MS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cutoff&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ISODate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;2026-01-01&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;total&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ids&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;created&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;$lt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;cutoff&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
                              &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;CHUNK&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toArray&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;d&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;_id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ids&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;break&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;deleteMany&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;$in&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ids&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="nx"&gt;total&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;ids&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;SLEEP_MS&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;   &lt;span class="c1"&gt;// give the secondaries time to catch up&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;total&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="mi"&gt;1000000&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;total&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt; deleted, check the lag`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two rules we followed while running it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Only run it in off peak hours.&lt;/strong&gt; We ran it every night and stopped it before business hours. The full 2TB took around a week and a half, and nobody noticed anything, which was exactly the goal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Let replication lag decide the speed.&lt;/strong&gt; If the lag went above the level we were comfortable with, we increased &lt;code&gt;SLEEP_MS&lt;/code&gt;. The speed of the loop is decided by how healthy the cluster is, not by how quickly we want it done.&lt;/p&gt;

&lt;p&gt;The pause between chunks can look far too slow to developers. That is fine. Once the data has stopped growing, the delete has no deadline, and what stops the data from growing is the real fix below.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real fix: a TTL index, so this never happens again
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;createIndex&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;created&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;expireAfterSeconds&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2592000&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;  &lt;span class="c1"&gt;// 30 days&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few things about TTL indexes that surprise people.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The TTL process runs every 60 seconds, so documents are not deleted at the exact second they expire. That is fine for cleaning up old data, but do not rely on it for business logic that needs exact timing.&lt;/li&gt;
&lt;li&gt;It deletes in batches in the background, which is basically the same chunked loop as above, running forever. On a very large backlog it can fall behind, which is why we cleared the 2TB by hand first and then let TTL handle it from there.&lt;/li&gt;
&lt;li&gt;Put the TTL index on the field that actually decides when the data is old. A TTL index on the wrong date field will quietly delete data you wanted to keep.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We finished with a runbook covering the script, the lag limits, and the schedule, so the next time someone asks to delete 2TB, the answer is already written down.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it yourself
&lt;/h2&gt;

&lt;p&gt;If you want to see this on your laptop, mdbkit lab gives you a 3 node replica set.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;mdbkit
mdbkit lab start          &lt;span class="c"&gt;# 3 node replica set on 127.0.0.1:28110-28112&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Connect to the primary and load a million documents, about half of them older than 60 days.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// mongosh --port 28110&lt;/span&gt;
&lt;span class="nx"&gt;use&lt;/span&gt; &lt;span class="nx"&gt;shop&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;docs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;10000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;created&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;random&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;120&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;86400000&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="na"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;x&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;repeat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;insertMany&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then run the chunked delete. This version prints the time, how many documents have been deleted so far, and how far behind the secondaries are.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;secondaryLag&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;rs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;members&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;stateStr&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;PRIMARY&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;lags&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;members&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;stateStr&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;SECONDARY&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;optimeDate&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;optimeDate&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(...&lt;/span&gt;&lt;span class="nx"&gt;lags&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cutoff&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;86400000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;total&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ids&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;created&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;$lt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;cutoff&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
                       &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10000&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toArray&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;d&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ids&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;break&lt;/span&gt;
  &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;deleteMany&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;$in&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ids&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
  &lt;span class="nx"&gt;total&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;ids&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;toISOString&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;substr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;11&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;t&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;  deleted &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;total&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt; old documents   secondary lag &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nf"&gt;secondaryLag&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;s&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;done, deleted &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;total&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt; old documents&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You will see the lag stay close to zero the whole way through. If you want to compare, reload the data and run the same delete as one single &lt;code&gt;deleteMany&lt;/code&gt;, then check &lt;code&gt;rs.printSecondaryReplicationInfo()&lt;/code&gt; while it runs.&lt;/p&gt;

&lt;p&gt;When you are done:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;mdbkit lab destroy &lt;span class="nt"&gt;--yes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When you have to delete a lot of data from a live system, going slowly is the safe way to do it.&lt;/p&gt;

</description>
      <category>mongodb</category>
      <category>operations</category>
      <category>performance</category>
      <category>database</category>
    </item>
    <item>
      <title>I take down a healthy primary on purpose. You should too.</title>
      <dc:creator>Saqib Ameen Subhan</dc:creator>
      <pubDate>Sun, 23 Aug 2026 15:58:41 +0000</pubDate>
      <link>https://dev.to/saqibameen86/i-take-down-a-healthy-primary-on-purpose-you-should-too-bm7</link>
      <guid>https://dev.to/saqibameen86/i-take-down-a-healthy-primary-on-purpose-you-should-too-bm7</guid>
      <description>&lt;p&gt;In one of my previous companies, I built an automation that used to scale our MongoDB clusters automatically. When CPU crossed a threshold in CloudWatch, the automation upgraded the nodes one by one. First a secondary, then the second secondary, and the primary last. It took around 12 minutes, nobody had to log in, and the application was not impacted.&lt;/p&gt;

&lt;p&gt;The order matters here. Restarting the primary last means the cluster goes through only one election for every upscale, and the automation takes care of it. But this only works if you trust elections, and you will trust elections once you have triggered them yourself and seen how your application behaves when it happens.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens during an election?
&lt;/h2&gt;

&lt;p&gt;In simple terms, this is how it works.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Every node in a replica set checks on the other nodes every 2 seconds (heartbeats).&lt;/li&gt;
&lt;li&gt;If the primary does not respond for 10 seconds (electionTimeoutMillis), the other nodes assume it is gone.&lt;/li&gt;
&lt;li&gt;A secondary asks the other nodes to vote for it. It can only win if its data is at least as up to date as the nodes voting for it (i.e its oplog is as recent as theirs).&lt;/li&gt;
&lt;li&gt;The node that gets the majority of votes becomes the new primary, and the other secondaries start copying data from it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One thing that surprises a lot of people is the priority setting. Priority decides which nodes get preference to become primary, but it cannot make a node with old data win. A node with priority 10 that is behind will lose to a node with priority 1 that is fully up to date. Up to date data matters more than priority. Once the higher priority node catches up with the others, it calls another election and takes over as primary, as the priority dictates.&lt;/p&gt;

&lt;p&gt;In total, there can be up to 15 seconds with no primary, and during that time MongoDB cannot accept any writes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real problem is your application
&lt;/h2&gt;

&lt;p&gt;The election itself only takes a few seconds. What really matters is what your application does during those seconds. I have seen three kinds of behaviour.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Users see errors.&lt;/strong&gt; If the driver throws a NotWritablePrimary error and nobody has handled it, users start getting errors, and things like checkout start failing, which directly impacts the business.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The application retries blindly.&lt;/strong&gt; Someone wrote their own retry loop, and sometimes the same write gets saved twice. This is actually worse than failing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The driver handles it.&lt;/strong&gt; With retryWrites=true and a sensible serverSelectionTimeoutMS, the driver holds the write, waits for the new primary, and retries it exactly once. The user just sees one slow request.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You don't want to find out which one you have during an incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  The drill
&lt;/h2&gt;

&lt;p&gt;I run this every quarter. First in staging, then in production during a low traffic window. Yes, production too. Staging tells you that the drill works. Production tells you how your real application behaves.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// 1. Check which node is primary right now&lt;/span&gt;
&lt;span class="nx"&gt;rs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nx"&gt;members&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;state&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;stateStr&lt;/span&gt; &lt;span class="p"&gt;}))&lt;/span&gt;

&lt;span class="c1"&gt;// 2. Ask the primary to step down&lt;/span&gt;
&lt;span class="nx"&gt;rs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stepDown&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;// it will not try to become primary again for 60 seconds&lt;/span&gt;

&lt;span class="c1"&gt;// 3. Check from a secondary which node became the new primary&lt;/span&gt;
&lt;span class="nx"&gt;rs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;While the drill runs, I note down these numbers.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;time taken to elect a new primary : ___ seconds
errors seen by the application    : ___
p99 latency during the election   : ___ ms
any write saved twice?            : ___ (check using a unique key)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These numbers are the real output of the drill, so fill them in from your own run.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;rs.stepDown()&lt;/code&gt; is the gentle way to do this. The primary finishes what it is doing and hands over cleanly. Once you are comfortable with that, try the harder version. Kill the primary's mongod process with &lt;code&gt;kill -9&lt;/code&gt;, or cut its network. This is much closer to what happens in a real cloud outage, because the other nodes have to wait the full 10 seconds before they start an election.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I learned doing this at scale
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Always keep an odd number of voting members.&lt;/strong&gt; With four voting members, the votes can split 2 and 2, and then nobody becomes primary. If you really cannot afford a third data node, you can use an arbiter, but remember that an arbiter does not help w:"majority" writes get acknowledged, because it does not hold any data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use priority to decide which region should be primary.&lt;/strong&gt; In our multi region clusters, we gave higher priority to the nodes in the main region, so that once those nodes recovered, the primary moved back there automatically. But as I said above, priority only gives preference. The node with the most up to date data still wins.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it yourself
&lt;/h2&gt;

&lt;p&gt;If you want to try this without touching a real cluster, mdbkit lab gives you a 3 node replica set on your laptop.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;mdbkit
mdbkit lab start          &lt;span class="c"&gt;# 3 node replica set on 127.0.0.1:28110-28112&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Connect to the replica set and run a loop that saves an order every second, with the time printed next to it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// mongosh "mongodb://127.0.0.1:28110,127.0.0.1:28111,127.0.0.1:28112/"&lt;/span&gt;
&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;toISOString&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;substr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;11&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;drill&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;insertOne&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;order&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;t&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;  order &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt; saved&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;t&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;  order failed, no primary available&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In a second terminal, connect to the primary and run &lt;code&gt;rs.stepDown(30)&lt;/code&gt;. Then watch the first terminal and look at the timestamps. You will see exactly how long your orders were affected, and whether they failed or just waited.&lt;/p&gt;

&lt;p&gt;When you are done:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;mdbkit lab destroy &lt;span class="nt"&gt;--yes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An election you have practised is just a normal operational event. An election you have never practised is an outage.&lt;/p&gt;

</description>
      <category>mongodb</category>
      <category>highavailability</category>
      <category>sre</category>
      <category>database</category>
    </item>
  </channel>
</rss>
