DEV Community

Naoki
Naoki

Posted on

đźš§Where Would I Start If Asked to Find and Fix Problems in a Company's AWS Environment?

Where Would I Start If Asked to Find and Fix Problems in a Company's AWS Environment?

First, identify the production environment and the services running in it.

Then check how each service is configured. Address problems that could cause outages or expose data first.

Identify the Production Environment

Check the accounts in AWS Organizations.

Do not rely on account names alone. Compare tags, active resources, and internal documentation to identify the production accounts and Regions.

Record:

  • Production AWS accounts
  • Regions in use
  • Main services running in each account
  • The person responsible for each service

The image below is an example of the AWS Organizations account list. It is an AWS sample screenshot, not a capture of your company's environment.

Example AWS Organizations account list

Image source: AWS Cloud Operations Blog

List the Production Services

In the production accounts, check resources such as Route 53, CloudFront, ALB, ECS, EC2, and RDS.

For each customer-facing website or API, record its domain, related AWS resources, and the impact of an outage.

Service:
Member website

Domain:
example.com

Main resources:
CloudFront, ALB, ECS, Aurora

Impact of an outage:
Users cannot log in or make purchases
Enter fullscreen mode Exit fullscreen mode

This list helps you decide what to investigate first.

Start with a service whose failure would have a significant impact.

Check the Resources Used by a Service

For the selected service, find out where a user's request goes and which resources the application needs.

Route 53
  ↓
CloudFront
  ↓
ALB
  ↓
ECS
  ↓
Aurora
Enter fullscreen mode Exit fullscreen mode

Confirm each connection in the AWS settings:

  • Route 53: Where the DNS record points
  • CloudFront: Which origin the behavior uses
  • ALB: Which target group the listener rule selects
  • ECS: Which service, task definition, and tasks run the application
  • Aurora: Which database the application connects to

For example, the following AWS sample screenshot shows an ALB listener rule forwarding traffic to target groups. It does not show the service described in this article.

Example ALB listener rule and target groups

Image source: AWS Networking & Content Delivery Blog

After checking these settings, you know which AWS resources are needed to run the service.

Check for Problems

Once you understand the service, review its settings and logs.

  • CloudWatch: Are there alarms for failures, and do notifications reach someone?
  • RDS or Aurora: Are backups being taken?
  • S3: How are important files protected?
  • Security groups: Is access allowed from more sources than necessary?
  • IAM: Does the application role have excessive permissions?
  • CloudTrail: Can you investigate recent changes?

The image below is an AWS sample of CloudWatch alarm notification settings. In your own environment, check the actual alarm and its notification destination.

Example CloudWatch alarm notification settings

Image source: AWS Cloud Operations Blog

For example, Aurora backups may be enabled, but there may be no record that anyone has successfully restored one. Record that as an issue to investigate.

If the database contains important data, restore a backup into a separate environment and check whether you can read it. This verifies that the data can be recovered without changing the production database.

Fix the Problems You Find

For each problem, record the affected service, the evidence, the proposed change, and how you will verify it.

Affected service:
Member website

Setting checked:
Security group for the ECS tasks

Problem:
The application port is open to a broad range of sources

Required traffic:
Traffic from the ALB to the ECS tasks

Proposed change:
Allow traffic from the ALB security group

Verification:
Confirm that users can log in and make purchases
through the ALB
Enter fullscreen mode Exit fullscreen mode

Before changing a setting, check which current connections depend on it.

After the change, test the main application functions and check CloudWatch Logs for errors.

Move to the Next Service

Once you have investigated and addressed the selected service's problems, return to the list and choose the next service.

Identify the production environment
        ↓
List the production services
        ↓
Check the resources used by a high-impact service
        ↓
Find problems in its settings and logs
        ↓
Fix the problems and verify the results
        ↓
Move to the next service
Enter fullscreen mode Exit fullscreen mode

Following this order makes it easier to explain what you are checking and why each setting matters.

Top comments (1)