Where Would I Start If Asked to Find and Fix Problems in a Company's AWS Environment?
First, identify the production environment and the services running in it.
Then check how each service is configured. Address problems that could cause outages or expose data first.
Identify the Production Environment
Check the accounts in AWS Organizations.
Do not rely on account names alone. Compare tags, active resources, and internal documentation to identify the production accounts and Regions.
Record:
- Production AWS accounts
- Regions in use
- Main services running in each account
- The person responsible for each service
The image below is an example of the AWS Organizations account list. It is an AWS sample screenshot, not a capture of your company's environment.
Image source: AWS Cloud Operations Blog
List the Production Services
In the production accounts, check resources such as Route 53, CloudFront, ALB, ECS, EC2, and RDS.
For each customer-facing website or API, record its domain, related AWS resources, and the impact of an outage.
Service:
Member website
Domain:
example.com
Main resources:
CloudFront, ALB, ECS, Aurora
Impact of an outage:
Users cannot log in or make purchases
This list helps you decide what to investigate first.
Start with a service whose failure would have a significant impact.
Check the Resources Used by a Service
For the selected service, find out where a user's request goes and which resources the application needs.
Route 53
↓
CloudFront
↓
ALB
↓
ECS
↓
Aurora
Confirm each connection in the AWS settings:
- Route 53: Where the DNS record points
- CloudFront: Which origin the behavior uses
- ALB: Which target group the listener rule selects
- ECS: Which service, task definition, and tasks run the application
- Aurora: Which database the application connects to
For example, the following AWS sample screenshot shows an ALB listener rule forwarding traffic to target groups. It does not show the service described in this article.
Image source: AWS Networking & Content Delivery Blog
After checking these settings, you know which AWS resources are needed to run the service.
Check for Problems
Once you understand the service, review its settings and logs.
- CloudWatch: Are there alarms for failures, and do notifications reach someone?
- RDS or Aurora: Are backups being taken?
- S3: How are important files protected?
- Security groups: Is access allowed from more sources than necessary?
- IAM: Does the application role have excessive permissions?
- CloudTrail: Can you investigate recent changes?
The image below is an AWS sample of CloudWatch alarm notification settings. In your own environment, check the actual alarm and its notification destination.
Image source: AWS Cloud Operations Blog
For example, Aurora backups may be enabled, but there may be no record that anyone has successfully restored one. Record that as an issue to investigate.
If the database contains important data, restore a backup into a separate environment and check whether you can read it. This verifies that the data can be recovered without changing the production database.
Fix the Problems You Find
For each problem, record the affected service, the evidence, the proposed change, and how you will verify it.
Affected service:
Member website
Setting checked:
Security group for the ECS tasks
Problem:
The application port is open to a broad range of sources
Required traffic:
Traffic from the ALB to the ECS tasks
Proposed change:
Allow traffic from the ALB security group
Verification:
Confirm that users can log in and make purchases
through the ALB
Before changing a setting, check which current connections depend on it.
After the change, test the main application functions and check CloudWatch Logs for errors.
Move to the Next Service
Once you have investigated and addressed the selected service's problems, return to the list and choose the next service.
Identify the production environment
↓
List the production services
↓
Check the resources used by a high-impact service
↓
Find problems in its settings and logs
↓
Fix the problems and verify the results
↓
Move to the next service
Following this order makes it easier to explain what you are checking and why each setting matters.



Top comments (1)