DEV Community

Amira Abidi
Amira Abidi

Posted on

Disaster Recovery: Restoring PostgreSQL on Amazon RDS

A backup is easy to create.

The harder question is:

Can you actually recover your application when you need to?

While building LibraryCorner, my AWS project running on Amazon EKS with PostgreSQL on Amazon RDS, I wanted to look beyond simply deploying a database and enabling backups.

I wanted to understand what happens when things go wrong.

Because having a backup is not the same as having a recovery strategy.

Backup is only the starting point

For a database-backed application, disaster recovery is more than storing a copy of your data.

The real recovery flow looks more like:

Failure
   ↓
Choose recovery point
   ↓
Restore PostgreSQL
   ↓
Reconnect application
   ↓
Validate data
   ↓
Validate application
   ↓
Return to service
Enter fullscreen mode Exit fullscreen mode

Every step matters.

A restored database that the application cannot reach is not a successful recovery.

My LibraryCorner architecture

In LibraryCorner, the application runs on Amazon EKS, while PostgreSQL is hosted on Amazon RDS.

At a high level:

                  Internet
                     │
                     ▼
                  ALB
                     │
                     ▼
                Amazon EKS
                     │
                     ▼
                Amazon RDS
                 PostgreSQL
                     │
                     ▼
              Backup / Snapshot
Enter fullscreen mode Exit fullscreen mode

The production database uses a multi-AZ configuration to improve availability.

But this led me to an important distinction:

High availability is not the same as disaster recovery.

Multi-AZ helps protect against certain infrastructure failures.

Backups and snapshots provide another layer of protection when recovery is required.

What would recovery look like?

Imagine that the database needs to be replaced.

The first question isn't simply:

“Can I restore it?”

It is:

“Which recovery point should I restore?”

This is where two concepts become important.

RPO — Recovery Point Objective

How much data can I afford to lose?

A system that can tolerate an hour of data loss has very different requirements from one that needs almost zero data loss.

RTO — Recovery Time Objective

How long can the application remain unavailable?

A recovery strategy that takes several hours may be perfectly acceptable for one application and completely unacceptable for another.

RPO and RTO therefore aren't just theoretical concepts.

They influence the architecture you choose.

Restoring the database isn't the end

Suppose the PostgreSQL instance has been restored successfully.

The RDS console says:

Available ✓

Is the recovery finished?

Not yet.

The application still needs to communicate with the restored database.

The recovery process is therefore closer to:

Backup / Snapshot
        ↓
Restore RDS
        ↓
Database available
        ↓
Application reconnects
        ↓
Functional validation
        ↓
Recovery complete
Enter fullscreen mode Exit fullscreen mode

This is where my QA background strongly influences how I approach recovery.

I don't want to validate only the infrastructure.

I want to validate the complete system.

What would I validate?

First, the database:

  • Can I connect?
  • Is the expected schema present?
  • Is the expected data available?

Then the application:

  • Can it connect to PostgreSQL?
  • Can it retrieve existing data?
  • Can it create new data?
  • Do the main API operations work?

And finally, the infrastructure:

  • Is network connectivity working?
  • Are access controls still correct?
  • Is monitoring operational?

This is essentially the same mindset I used in QA:

Don't just check that the component is running. Check that the system actually works.

The lesson I took from this project

Before working on Cloud and DevOps, I thought of backups mainly as a safety mechanism.

Now I see them as one part of a larger process:

Backup → Restore → Reconnect → Validate

A resilient architecture needs both availability and recoverability.

And recovery should not be something you discover for the first time during an incident.

Ideally, the restore process should be tested.

Because a backup you have never successfully restored is ultimately an assumption.

My QA background taught me to question those assumptions.

Cloud is teaching me how to turn them into repeatable processes.

And that's probably the most important lesson from this project:

A backup gives you a recovery point. A tested restore gives you confidence.

Top comments (0)