- Critical failure state in HA & DR.
- a single cluster fractures into isolated segments due to a network partition.
- Why this happens ? --> Because the nodes can no longer communicate over their heartbeat network, multiple nodes simultaneously assume the primary/active role.
- It leads to data corruption and catastrophic system behavior.
- Eg., A scenario where both primary and secondary nodes mistakenly assume the role of the active primary server, leading to data corruption and conflicting transactions.
Root causes
- Heartbeat Network Failures
- Asymmetric Network Partitions
- Application/OS Freezes (GC Pauses)
Prevention
- To protect distributed systems, architects use a multi-layered prevention framework.
Notes
- catastrophic --> Sudden failure.






Top comments (0)