No matter how robust a single AWS availability zone or region is, enterprise production workloads require a bulletproof strategy for unexpected outages. For critical business applications, relying on a single geographic region is a significant risk.
With over 12 years in IT and 6+ years architecting scalable cloud solutions, I have designed several multi-region setups. In this post, I will share the architectural blueprints for a robust Active-Passive Disaster Recovery (DR) topology using Amazon Route 53, AWS Application Load Balancers (ALB), and Amazon RDS Cross-Region Read Replicas.
ποΈ The Multi-Region DR Architecture
We will set up our infrastructure across two distinct AWS regions:
- Primary Region (e.g., us-east-1): Handles 100% of the live production traffic.
- Secondary DR Region (e.g., us-west-2): Remains on standby, receives asynchronous database replication, and takes over immediately if the primary region goes offline.
[ Client Request ]
β
βΌ
βββββββββββββββββββββββββββ
β Amazon Route 53 β βββ (Failover Routing Policy via Health Checks)
ββββββ¬ββββββββββββββββ¬βββββ
β β
β (Active) β (Passive Standby)
βΌ βΌ
βββββββββββββββββ βββββββββββββββββ
β Primary Regionβ β DR Region β
β (us-east-1) β β (us-west-2) β
β ββββββββββ` β β βββββββββββ β
β β ALB β β β β ALB β β
β ββββββ¬βββββ β β ββββββ¬βββββ β
β βΌ β β βΌ β
β βββββββββββ β β βββββββββββ β
β β RDS Primaryβ β βRDS Replica β
β ββββββ¬βββββ β β βββββββββββ β
βββββββββΌββββββββ βββββββββββββββββ
β β²
ββ(Asynchronous)βββ
π οΈ Step 1: Configuring Global Database Replication
To minimize your Recovery Point Objective (RPO), your data must reside safely in the secondary region before disaster strikes. For relational data, we leverage Amazon RDS cross-region replication.
Here is how you define the Primary database and its encrypted, cross-region counterpart:
# Primary Database in us-east-1
resource "aws_db_instance" "primary_db" {
allocated_storage = 50
engine = "postgres"
engine_version = "15.4"
instance_class = "db.r6g.xlarge"
db_name = "production_db"
username = "cloud_admin"
password = var.db_password
backup_retention_period = 7
multi_az = true
skip_final_snapshot = false
}
# Cross-Region Read Replica in us-west-2
resource "aws_db_instance" "dr_replica" {
provider = aws.us-west-2
replicate_source_db = aws_db_instance.primary_db.arn
instance_class = "db.r6g.xlarge"
skip_final_snapshot = true
storage_encrypted = true
}
π¦ Step 2: Intelligent Routing with Route 53 DNS Failover
The heart of an automated Active-Passive setup is Amazon Route 53 Health Checks. Route 53 continuously probes the primary public-facing Application Load Balancer. If it detects a failure, it dynamically shifts DNS records to point to the DR region.
# Route 53 Health Check for the Primary Load Balancer
resource "aws_route53_health_check" "primary_health" {
fqdn = aws_lb.primary_alb.dns_name
port = 443
type = "HTTPS"
resource_path = "/healthz"
failure_threshold = 3
request_interval = 30
}
# Primary DNS Record (Active)
resource "aws_route53_record" "primary_dns" {
zone_id = var.hosted_zone_id
name = "app.myenterprise.com"
type = "A"
failover_routing_policy {
type = "PRIMARY"
}
set_identifier = "primary-active"
health_check_id = aws_route53_health_check.primary_health.id
alias {
name = aws_lb.primary_alb.dns_name
zone_id = aws_lb.primary_alb.zone_id
evaluate_target_health = true
}
}
# Secondary DNS Record (Passive Standby)
resource "aws_route53_record" "secondary_dns" {
zone_id = var.hosted_zone_id
name = "app.myenterprise.com"
type = "A"
failover_routing_policy {
type = "SECONDARY"
}
set_identifier = "secondary-passive"
alias {
name = aws_lb.dr_alb.dns_name
zone_id = aws_lb.dr_alb.zone_id
evaluate_target_health = true
}
}
π‘ Crucial Operational Metrics: RTO vs. RPO
When presenting a Disaster Recovery plan to enterprise stakeholders, you must explicitly define your core operational boundaries:
- Recovery Point Objective (RPO): The maximum targeted duration during which data might be lost due to a major incident. By using asynchronous RDS replication, our RPO is typically under a few minutes.
- Recovery Time Objective (RTO): The duration of time within which a business process must be restored. In this setup, our RTO is governed by the Route 53 TTL and health check thresholds (typically under 2 to 3 minutes for automated DNS redirection).
π Conclusion
Building a cross-region architecture requires careful planning around data serialization and DNS propagation times. However, the operational peace of mind it gives your enterprise organization is unmatched.
How does your organization handle business continuity? Do you use an Active-Passive standby, or are you utilizing a fully hot Active-Active global topology? Let's discuss in the comments below!
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.