DEV Community

Cover image for Database Failover & Performance: Synchronous vs. Asynchronous Replication

Database Failover & Performance: Synchronous vs. Asynchronous Replication

Replication ensures your database remains available, consistent, and resilient—but should you choose synchronous or asynchronous replication? 🤔

In this video, Andrew Redden from Timescale breaks down the key differences:

✔ Synchronous Replication – Guarantees data consistency, but can introduce latency. Ideal for financial systems and mission-critical applications.
✔ Asynchronous Replication – Prioritizes speed and performance but may cause temporary lag. Perfect for high-traffic applications and analytics.

We’ll also cover best practices for replication setups, including high-availability configurations and failover strategies in PostgreSQL.

🛠 𝗥𝗲𝗹𝗲𝘃𝗮𝗻𝘁 𝗥𝗲𝘀𝗼𝘂𝗿𝗰𝗲𝘀
🔗 Check out Timescale’s documentation for more details

Top comments (1)

Collapse
 
scsoi profile image
疏影 •

Sync vs async is the easy decision — the harder call is what you do when the link between primary and replica goes flaky.

A few production things that rarely make it into the comparison charts:

  1. Replication lag is not symmetric. A 200ms RTT link that looks fine on SHOW REPLICA STATUS will occasionally spike to 2-5s during TCP retransmit storms. Sync mode turns every primary commit into a tail-on-the-replica latency. Async hides the cost until failover, then you eat the data loss.

  2. Semi-sync (after_sync / after_commit) helps but does not eliminate the gap. MySQL 8 semi-sync still loses the last binlog group if the primary crashes between commit and ack from the replica. People confuse semi-sync with "no loss"; it is not.

  3. Group replication / Galera adds a certification layer, which is great for write consistency but punishes long-running transactions. A single 30s UPDATE on a 10M row table will block every other writer on the cluster.

  4. Storage failover (especially in cloud) is not just about the DB. The replica's data dir lives on a mount; if that mount is an EBS / disk that takes 30-90s to attach on the new node, your failover SLA already broke before the DB even started.

  5. Backups need to be replica-aware. A snapshot taken on the primary while the async replica is mid-replay can produce a crash-recovery file that does not match the binlog position you record. Always back up from a paused replica or use the replica's relay log position.

On the storage side: for shops running MySQL with hot replicas across regions, the data dir is usually a mounted volume. We have seen teams standardise on a single mount layer (ScsDriver mounts WebDAV / Aliyun OSS / S3 / S3-compatible / NFS as a real local disk on Windows or Linux so each replica sees a stable path; failover just re-mounts at the new node) instead of juggling per-cloud volume attach scripts. The DB does not care that the underlying blob store is object storage, as long as fsync is honest.

What is everyone using for the underlay in 2025 — local NVMe with periodic snapshot, or a network mount with replication handled below the DB?