DEV Community

Cover image for I Destroyed My Entire Cloud Environment on Purpose. One Command Brought It Back.
Vivian Chiamaka Okose
Vivian Chiamaka Okose

Posted on Originally published at vivianokose.hashnode.dev

I Destroyed My Entire Cloud Environment on Purpose. One Command Brought It Back.

I ran one command and deleted my whole AWS environment, the network, the servers, the
database, the storage, a running Kubernetes cluster, all of it. Then I ran one more command
and it came back. Same network, same servers, same security rules, identical.

That round trip is the entire point of this module, and it is the test of whether your
infrastructure genuinely lives in code or just mostly does.

The problem

A company I was modelling this on built their AWS environment by clicking through the console.
It worked, until a new engineer joined and spent two days trying to recreate the dev
environment from screenshots and a colleague's memory. Nobody could reproduce it, nobody could
review a change before it happened, and tearing it down was a hand-written list of steps you
hoped was in the right order.

I had felt every piece of that myself. An earlier module where my teardown instructions failed
because I had written the delete order from memory. Another where a whole cluster lived in a
pile of commands that only existed in my terminal history.

The fix: describe it, do not click it

Terraform lets you describe infrastructure in files, then creates, changes, or destroys it to
match. You do not click "make a VPC." You write "there should be a VPC with these subnets,"
run one command, and it builds all of it in the right order.

If that sounds like Kubernetes, it is exactly the same idea, declare the desired state and let
the tool reconcile reality to match it, pointed one level down at the cloud itself instead of
at containers.

The daily rhythm has one step that makes it genuinely safer than the console:

terraform plan shows you precisely what it would do before it does anything:

Plan: 18 to add, 0 to change, 0 to destroy.
Enter fullscreen mode Exit fullscreen mode

Every risky change I ever made blind in a console, Terraform previews first. That one feature
is most of the argument.

State, and why it goes in S3

Terraform remembers what it built, in a file called the state. It is how it tells "create new"
from "change existing." That state is precious, it holds the record of everything you own, and
it is also where secrets end up: the database password sits in it in plain text.

So the first thing I built, before any infrastructure, was a remote home for it: an S3 bucket,
versioned and encrypted and locked down, with a DynamoDB table holding a lock so two people
cannot apply at once and corrupt it.

Remote state
The state in S3, not on my laptop. The empty local listing is the point.

The architecture, as modules

I split the infrastructure into four modules, networking, compute, database, storage, each a
folder with its own inputs and outputs. The networking module hands its subnet and
security-group IDs to the others. Terraform reads those references and figures out the build
order itself; I never sequenced anything.

Applied
Eighteen resources, an entire production-shaped architecture, from one command.

The gotcha the lab did not mention

I scaled the server count from two to three just to watch the plan, and Terraform stopped me
cold: the third server tried to use a third subnet that did not exist, and plan caught it
before building anything. A real limitation of counting resources by index, and a console
would have let me walk straight into it.

Plan

A whole EKS cluster from one block

The part that made me laugh. An earlier module took me a long saga to stand up one EKS cluster
with a command-line tool, through two failures. Here it was one module block from the community
registry, fifty-two resources from about forty lines of my own.

EKS from Terraform
A full managed Kubernetes cluster, nodes ready, from a module someone else already tested.

Catching reality drifting from code

Two things that make Terraform trustworthy in a real, messy account. import adopts a
resource that already exists into management, without recreating it. And drift, someone
changing a resource by hand, gets caught by plan, which shows you the gap between your code
and what is actually there.

Drift
A manual change, surfaced. The real fix is stopping manual changes, not just re-applying.

CI that reviews the infrastructure, not just the app

Here is a part I did not expect to enjoy. Every pull request that touches a Terraform file now
runs five checks before anything merges: formatting, syntax validation, a linter, and two
security scanners (tfsec and checkov).

The first time it ran, it failed. Not on anything clever, on formatting. The scanner found my
files were not laid out in Terraform's canonical style and stopped the merge with a plain exit
code. One command fixed it (terraform fmt -recursive), but the lesson stuck: the pipeline
holds a line I would have let slide by hand.

Then the security scanners spoke up, and this is the interesting bit. They flagged ten things.
Not failures, findings, reported but not blocking. Instance metadata still allowing the older,
weaker version. A disk not explicitly set to encrypt. A community module pinned to a version
tag instead of a commit hash. I read each one and made a call: some are real hardening I would
do before this carried traffic, some are not worth it on a server that lives for an hour.

That triage is the actual skill. Running a scanner is easy. Knowing which of its findings
matter for the situation in front of you is the part that takes understanding, and it is the
difference between a report nobody reads and a decision you can defend.

CI checks
Green pipeline, security findings posted as notes. Both true at once, and that is the healthy state.

The round trip

And then the headline. Destroy everything, apply it back, get an identical environment. If it
survives that, your infrastructure truly is code. If it does not, you have drift, unmanaged
resources, or a dependency your files do not capture, things you have been pretending do not
exist.

Mine came back identical. That is the whole module in one gesture.

One last honest note, on teardown

Cleaning up had its own small trap. The state bucket had versioning switched on, which I had
set on purpose so I could recover old state. When I went to delete it, the normal "force empty
and remove" command refused, the bucket still had every old version and a pile of delete
markers underneath. Versioning keeps everything, including the things you thought you deleted.
I had to clear the versions explicitly before the bucket would go. A good reminder that the
feature protecting you is the same feature standing in your way at cleanup, and that teardown
deserves as much care as the build.

Full build, every module, the remote-state setup, the EKS block, the CI workflow, and the
discipline runbook:
https://github.com/vivianokose/nexaops-operations-lab/tree/main/12-terraform

Top comments (0)