DEV Community

Ugochukwu Oguejiofor
Ugochukwu Oguejiofor

Posted on Originally published at ugochukwuoguejiofor.com

The Terraform defaults I would lock down before running ECS Fargate in production

An ECS Fargate service can be easy to deploy and still be a poor starting point for production.

I’ve seen starters that run tasks with public IPs, keep staging online all weekend, store AWS keys in CI, and create alarms that fire whenever a non-production environment shuts down. The infrastructure works. The defaults create the problems.

When I built my ECS Fargate Terraform starter, I focused on the decisions a team would otherwise have to revisit before every deployment.

Keep the tasks private

The Application Load Balancer accepts public traffic. The ECS tasks run in private subnets across two Availability Zones, without public IPs.

That adds networking cost and setup. It also gives you a clearer boundary: traffic reaches the application through the load balancer rather than directly through a task. The starter uses VPC endpoints for ECR, CloudWatch Logs, and S3, while NAT handles other outbound access.

There is a deliberate cost trade-off. Non-production uses one NAT gateway; production uses two for better availability across zones. Neither choice should be hidden in a copied Terraform example.

Give each environment its own capacity rules

Development, staging, and production can use the same module without using the same capacity settings.

The settings I pay most attention to are:

  • allow_scale_to_zero: non-production services can shut down outside their usage window.
  • use_fargate_spot: useful for a sandbox; disabled in this starter’s production settings.
  • desired_count and min_capacity: production starts with two tasks rather than treating a single running task as sufficient.

Terraform state is separate for each environment too. A change intended for development should not share a state file with production.

These choices are explicit in the environment configuration. Someone reviewing a production change can see them in code.

Make alarms match the environment

An alarm is useful when it tells someone to act. It becomes noise when it reports expected behaviour as an incident.

In production, zero healthy targets can mean the service is down. In a development environment scheduled to scale to zero overnight, zero healthy targets can mean the schedule worked. The starter skips that particular alarm for environments allowed to sleep.

It also sets alarms for ALB 5xx responses and unhealthy hosts, with evaluation over multiple data points to reduce one-off alerts. Log groups have a defined retention period.

The test is simple: if an alarm fires, will someone know what action to take?

Deploy without stored AWS access keys

The GitHub Actions workflow uses OIDC to assume an AWS role. Its trust policy restricts which repository and branch can assume that role. The job gets temporary credentials instead of relying on long-lived AWS keys in repository secrets.

Images are tagged with a commit SHA, and deployments update the ECS task definition with a deployment circuit breaker and rollback. Production goes through a GitHub Environment where a team can require a reviewer.

OIDC does not fix an over-permissive deployment role. The role still needs permissions scoped to the resources the workflow actually manages.

Leave out infrastructure you cannot justify yet

The sample application is a small Go service with / and /health. The starter does not provision a database, Redis, or a service mesh for an application that has no demonstrated need for them.

That is the principle behind the repo: make consequential choices visible, and leave room to add components when the application requires them.

You can inspect the Terraform starter and its DECISIONS.md. I also wrote a longer explanation of the architecture choices.

Which ECS default has caused the most trouble in a production environment you’ve worked on?

Top comments (0)