From EC2 to Zero Trust: A Week of Building (and Breaking) AWS Infrastructure
Week 6 of the DevOps Micro Internship didn't ease me in. Three assignments, three architectures, each one building directly on the last: a highly available two-tier app behind a load balancer and Auto Scaling Group, a full three-tier production-style deployment with a completely isolated database, and finally an AI-assisted security audit that turned Claude Code into something closer to a junior security analyst than a code assistant, one that's explicitly forbidden from touching anything it finds.
None of these went in cleanly. That's the part worth writing about.
Part 1 —> High Availability, or: What "Redundant" Actually Costs You
Assignment 5 was the first time I built infrastructure meant to survive its own failure. A VPC split across two Availability Zones, public and private subnets in each, an Application Load Balancer in front, and an Auto Scaling Group keeping two to four EC2 instances alive behind it, all pointed at a private, non-publicly-accessible RDS MySQL instance.
The security model was the real lesson here: Internet → ALB → EC2 → RDS, with each layer trusting only the one directly before it. Not "the database accepts connections from the VPC"; the database accepts connections from exactly one security group, the web tier's, and nothing else. Building the groups in that specific order mattered too: you can't reference a security group that doesn't exist yet, so ALB's group had to exist before EC2's could point at it.
I didn't just build this and call it done; I tried to break it on purpose. Terminated a live EC2 instance mid-run and watched the Auto Scaling Group notice, launch a replacement, and register it with the target group automatically, with the ALB continuing to serve traffic the entire time. Then I simulated losing an entire Availability Zone by pulling an instance into Standby and confirmed the site never went down. That's the difference between infrastructure you're told is highly available and infrastructure you've actually watched survive a failure.
Part 2 —> Three Tiers, Zero Shortcuts
Assignment 6 was the capstone, and it multiplied everything from Assignment 5 by 1.5: not two tiers but three, web, app, and database, each in its own pair of subnets across two AZs, six subnets total. Two load balancers, not one: a public ALB in front of the web tier, and a completely internal ALB in front of the app tier, invisible to the internet by design. The database sat in total isolation, no NAT route, no internet gateway path, nothing. Only the app tier could reach it, and only on port 3306.
This is also where the week's real debugging happened, and I'm not going to pretend it was clean:
- A VPC/subnet mismatch on my first test EC2 launch, a security group from one VPC paired with a subnet from another, which AWS correctly refused.
- A self-referencing security group rule that couldn't be created until the group it referenced already existed, an AWS quirk, not a mistake, but one that costs you ten confused minutes the first time you hit it.
- An
EADDRINUSEcrash loop on the backend, caused by a leftover foregroundnodeprocess I'd never actually stopped before handing the same app to PM2. Two processes fighting over port 3001, restarting forever. - A path-doubling bug:
${API_URL}/api/bookswhenAPI_URLwas already/api— that took the frontend from "loading forever" to a clean404the moment I actually looked at the Network tab instead of guessing. - A CORS rejection after testing through a DNS name that wasn't yet in the backend's allowed-origins list, an easy one to fix, but only once I stopped assuming the browser was lying to me.
- And one genuinely unresolved mystery: a redirect to an unfamiliar external domain that appeared briefly during testing, disappeared after a clean dependency reinstall, and never came back, investigated thoroughly (source, dependencies, build output, auth logs, cron), never definitively explained. I wrote that up honestly rather than pretending it was solved, because "I checked and it didn't reproduce" is a different claim than "I found and fixed the cause," and conflating the two is how real incidents get closed prematurely.
Every one of these forced the same discipline: isolate the failure layer by layer. Backend locally. Frontend locally. Through the reverse proxy. Through the public IP. Through the load balancer. You don't skip a layer just because you're confident, you skip it and then spend twenty minutes debugging the wrong thing.
Part 3 —> Teaching Claude Where the Line Is
Assignment 7 was the smallest in scope and the most interesting in what it was actually testing. I built a read-only Bash script that audits five specific things against my own AWS account: S3 public-access block settings, security groups open to the world on SSH and MySQL, RDS public accessibility, and EBS volume encryption, writing a PASS/WARN/FAIL report and exiting with a different code depending on severity.
Then I wrapped that script in a Claude Code skill, /aws-audit, scoped to exactly Bash, Read, and Grep, no Write. Not "instructed not to write files." Structurally unable to. The skill reads the report, explains every finding in plain language, estimates the cost or risk impact, and recommends a remediation command, and then stops. It never runs anything that changes the account. That boundary isn't a suggestion in a markdown file; it's enforced at the tool-permission level, which is the only place a boundary like that actually holds.
I proved it the only way that counts: by running the fix myself, by hand, in a separate terminal, after reviewing what Claude recommended. revoke-security-group-ingress to close a rule, authorize-security-group-ingress scoped to my own /32 to restore access properly, then a second audit run to confirm the finding actually flipped to PASS. That's Gather → Analyze → Human Act → Verify, the same loop from earlier weeks, except this time the "Human Act" step wasn't a nice-to-have. It was the entire point of the assignment. A tool that can find a misconfigured security group and fix it unsupervised is a tool one bad recommendation away from a real outage. Keeping those two capabilities separated is what makes the recommendation trustworthy in the first place.
Worth saying plainly, since the whole exercise is about not overselling your own results: my baseline audit wasn't clean. It genuinely flagged a real S3 bucket not fully blocking public ACLs, and two unencrypted EBS volumes, findings I'm aware of and haven't remediated as part of this task, since it was scoped specifically to the security-group layer. I'd rather report that honestly than imply the account came back fully green.
What I'm Taking Away
The infrastructure part of this week, VPCs, subnets, load balancers, security group chains, is learnable from documentation. The part that actually took effort was narrower: reading an error message literally instead of pattern-matching to something familiar, checking one layer before blaming the next, and being willing to write "I don't know" in a debugging note instead of quietly smoothing over a redirect I never fully explained.
An AI-assisted workflow is genuinely useful here, but only because of what it's deliberately not allowed to do. A script that gathers evidence, an AI that analyzes it, and a human who decides what to act on: that division of labor is the actual production pattern, not a training-wheels version of one.
Live three-tier app: http://book-review-web-alb-597773926.eu-north-1.elb.amazonaws.com/
P.S. This post is part of the DevOps Micro Internship (DMI) with Agentic AI — Cohort 3 — by Pravin Mishra. You can follow my graded progress throughout the cohort. Start your DevOps journey: https://dmi.pravinmishra.com/?utm_source=student&utm_medium=ps-blog&utm_campaign=cohort3


Top comments (0)