DEV Community

rubendob
rubendob

Posted on

Automating SSL with Let’s Encrypt on AWS

For quite some time, the SSL certificate for my WordPress site running on AWS worked in a fairly traditional way: I purchased the certificate, renewed it manually once a year, and installed it on the server. It worked, but there was an obvious problem: it depended on periodic manual intervention.

In addition, my architecture had evolved, and this approach to certificate management no longer fitted particularly well.

In this post, we are going to see how to automate an SSL certificate using Let’s Encrypt on AWS.


Initial architecture

The blog is deployed on AWS Elastic Beanstalk, using an EC2 instance in SingleInstance mode.

In front of Elastic Beanstalk, I have a CloudFront distribution:

Internet
   |
   v
CloudFront
   |
   | HTTPS
   v
Elastic Beanstalk
   |
   v
Nginx
   |
   v
WordPress
Enter fullscreen mode Exit fullscreen mode

There are actually two different TLS connections here.

The first one:

Client -> CloudFront
Enter fullscreen mode Exit fullscreen mode

uses a certificate from AWS Certificate Manager (ACM).

That certificate was already fully automated by AWS and did not require any changes.

The second one:

CloudFront -> Elastic Beanstalk / Nginx
Enter fullscreen mode Exit fullscreen mode

used another certificate.

That was the certificate I purchased through DonDominio and that was issued by Sectigo.

Nginx loaded it from:

ssl_certificate     /etc/pki/tls/certs/ssl-bundle.crt;
ssl_certificate_key /etc/pki/tls/certs/server.key;
Enter fullscreen mode Exit fullscreen mode

The files were stored in a private S3 bucket, and during the Elastic Beanstalk deployment a hook downloaded both of them:

S3
 |
 v
Elastic Beanstalk prebuild
 |
 v
/etc/pki/tls/certs/
 |
 v
Nginx
Enter fullscreen mode Exit fullscreen mode

The mechanism worked, but every renewal required purchasing the new certificate, downloading it, preparing the certificate chain, uploading it to S3, and deploying it.

The second problem: the EC2 instance is Spot

There was another important detail. The instance used by Elastic Beanstalk is Spot.

This means that the EC2 instance must be treated as ephemeral infrastructure: AWS may replace it.

Therefore, installing Certbot directly on that instance and letting it handle renewals was not a solution I particularly liked.

It could be done, but it would require persisting Certbot's state outside the machine and restoring it whenever the instance was replaced.

I preferred to completely decouple certificate renewal from the EC2 lifecycle.

Goal

The goal was to reach an architecture where:

  • Let’s Encrypt issued the certificate.
  • Validation was completely automated.
  • No permanent AWS credentials were stored in GitHub.
  • The certificate survived Spot instance replacements.
  • A renewal automatically updated Nginx.
  • A new EC2 instance could retrieve the latest available certificate.

The resulting architecture would look like this:

GitHub Actions
      |
      | OIDC
      v
     AWS
      |
      +------> Route53
      |         DNS-01
      |
      +------> Let's Encrypt
      |
      +------> S3
      |         |
      |         +-- ACME state
      |         +-- certificate
      |
      +------> SSM Run Command
                  |
                  v
              EC2 / Nginx
Enter fullscreen mode Exit fullscreen mode

DNS-01 validation with Route53

To issue the certificate, I decided to use the ACME DNS-01 challenge.

The certificate covers:

rubenortiz.es
www.rubenortiz.es
Enter fullscreen mode Exit fullscreen mode

Since the DNS zone is hosted in Route53, Certbot can temporarily create the necessary TXT records to prove to Let’s Encrypt that I control the domain.

The workflow uses:

certbot
+
certbot-dns-route53
Enter fullscreen mode Exit fullscreen mode

This way, I do not need to expose any special URL in WordPress or modify HTTP traffic in order to perform the validation.

GitHub Actions and AWS OIDC

The renewal process runs from GitHub Actions.

Instead of storing an AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY as secrets, the workflow uses OIDC to temporarily assume an IAM Role.

The flow is:

GitHub Actions
      |
      | OIDC token
      v
AWS STS
      |
      v
IAM Role
Enter fullscreen mode Exit fullscreen mode

That role only receives the permissions required to:

  • manage the DNS challenge in Route53;
  • read and write certificates in S3;
  • locate the active EC2 instance;
  • execute commands through AWS Systems Manager.

Persisting Certbot state

GitHub Actions runners are also ephemeral.

Each execution starts on a new machine.

Certbot, however, stores important information under:

/etc/letsencrypt
Enter fullscreen mode Exit fullscreen mode

This directory contains ACME accounts, renewal configurations, previous certificates, and other required information.

To avoid losing this state, I decided to store it in S3.

Instead of synchronizing the directory directly, I package it as a tar.gz, because Certbot uses symbolic links between the live and archive directories.

The result is:

s3://mybucket/ssl/acme-state-prod/
    letsencrypt.tar.gz
Enter fullscreen mode Exit fullscreen mode

At the beginning of the workflow:

S3
 |
 v
restore ACME state
 |
 v
Certbot
Enter fullscreen mode Exit fullscreen mode

and after Certbot runs:

Certbot
 |
 v
tar.gz
 |
 v
S3
Enter fullscreen mode Exit fullscreen mode

The active certificate

The certificate that must be consumed by the instances is kept separate from Certbot's internal state:

s3://mybucket/ssl/current/
    fullchain.pem
    privkey.pem
Enter fullscreen mode Exit fullscreen mode

fullchain.pem contains both the domain certificate and the intermediate certificate chain required by Nginx.

The private key is stored as:

privkey.pem
Enter fullscreen mode Exit fullscreen mode

Updating the EC2 instance using SSM

Generating and storing the certificate was only part of the problem.

I still needed the instance running WordPress to actually start using it.

For that, I use AWS Systems Manager Run Command.

GitHub Actions first locates the active Elastic Beanstalk instance using its tags:

aws ec2 describe-instances
Enter fullscreen mode Exit fullscreen mode

It then runs:

aws ssm send-command
Enter fullscreen mode Exit fullscreen mode

using the AWS-managed document:

AWS-RunShellScript
Enter fullscreen mode Exit fullscreen mode

The remote command roughly performs the following operations:

create certificate directory
        |
        v
download fullchain.pem from S3
        |
        v
download privkey.pem from S3
        |
        v
nginx -t
        |
        v
systemctl reload nginx
Enter fullscreen mode Exit fullscreen mode

Before reloading Nginx, the following command is always executed:

nginx -t
Enter fullscreen mode Exit fullscreen mode

If the Nginx configuration is invalid, the process fails before reloading the service.

New Nginx configuration

I also changed the paths used by Nginx to make it clear that these certificates belong to the new mechanism:

ssl_certificate     /etc/pki/tls/certs/letsencrypt/fullchain.pem;
ssl_certificate_key /etc/pki/tls/certs/letsencrypt/privkey.pem;
Enter fullscreen mode Exit fullscreen mode

The first real certificate issuance produced a certificate similar to this:

subject=CN=www.rubenortiz.es
issuer=C=US, O=Let's Encrypt, CN=YE1

DNS:rubenortiz.es
DNS:www.rubenortiz.es
Enter fullscreen mode Exit fullscreen mode

After reloading Nginx, I checked directly against the local Nginx listener:

openssl s_client \
  -connect localhost:443 \
  -servername www.rubenortiz.es
Enter fullscreen mode Exit fullscreen mode

and confirmed that the origin was already serving the Let’s Encrypt certificate.

What happens if AWS replaces the Spot instance?

This was one of the most important parts of the design.

The renewal process does not depend on EC2:

GitHub Actions -> Let's Encrypt -> S3
Enter fullscreen mode Exit fullscreen mode

And S3 acts as the persistent source for the certificate.

The Elastic Beanstalk prebuild hook was also modified to download:

ssl/current/fullchain.pem
ssl/current/privkey.pem
Enter fullscreen mode Exit fullscreen mode

Therefore, if AWS removes the current Spot instance:

Old EC2
    X
    |
    v
New EC2
    |
    v
Elastic Beanstalk prebuild
    |
    v
S3 /ssl/current/
    |
    v
Nginx
Enter fullscreen mode Exit fullscreen mode

the new instance automatically retrieves the current certificate.

Automatic renewal

Finally, the workflow no longer depends on pushes and instead runs on a GitHub Actions cron schedule:

on:
  schedule:
    - cron: "0 6 * * 1"

  workflow_dispatch:
Enter fullscreen mode Exit fullscreen mode

This means it runs every Monday and can also be triggered manually.

Certbot runs with:

--keep-until-expiring
Enter fullscreen mode Exit fullscreen mode

so running the workflow weekly does not mean issuing a new certificate every week.

As long as the current certificate is still valid for long enough, Certbot keeps the existing one.

This is useful because it provides several opportunities to renew the certificate before it reaches its expiration date.

GitHub Actions recipe

In this case, in order to determine which EC2 instance currently exists, I retrieve the information using the EB-related instance tags.

name: Renew Let's Encrypt Certificate

on:
  schedule:
    - cron: "0 6 * * 1"
  workflow_dispatch:

env:
  AWS_REGION: "<AWS_REGION>"
  AWS_ACCOUNT_ID: "<AWS_ACCOUNT_ID>"
  AWS_OIDC_ROLE_NAME: "<OIDC_ROLE_NAME>"

  LETSENCRYPT_EMAIL: "<EMAIL_ADDRESS>"

  CERT_DOMAIN: "www.example.com"
  CERT_APEX_DOMAIN: "example.com"

  CERT_BUCKET: "<PRIVATE_S3_BUCKET>"
  ACME_STATE_KEY: "certificates/acme-state/state.tar.gz"
  CERT_CURRENT_PREFIX: "certificates/current"

  EB_ENV_NAME: "<ELASTIC_BEANSTALK_ENVIRONMENT>"

permissions:
  id-token: write
  contents: read

jobs:
  letsencrypt-production:
    name: Renew production certificate
    runs-on: ubuntu-latest

    steps:
      - name: Configure AWS credentials
        uses: aws-actions/configure-aws-credentials@v5
        with:
          role-to-assume: arn:aws:iam::${{ env.AWS_ACCOUNT_ID }}:role/${{ env.AWS_OIDC_ROLE_NAME }}
          aws-region: ${{ env.AWS_REGION }}

      - name: Install Certbot
        run: |
          python -m venv .venv
          .venv/bin/pip install --upgrade pip
          .venv/bin/pip install certbot certbot-dns-route53

      - name: Restore Certbot state from S3
        run: |
          mkdir -p "${{ runner.temp }}/letsencrypt"

          if aws s3 ls \
            "s3://${{ env.CERT_BUCKET }}/${{ env.ACME_STATE_KEY }}" \
            > /dev/null 2>&1; then

            aws s3 cp \
              "s3://${{ env.CERT_BUCKET }}/${{ env.ACME_STATE_KEY }}" \
              "${{ runner.temp }}/letsencrypt.tar.gz"

            tar -xzf "${{ runner.temp }}/letsencrypt.tar.gz" \
              -C "${{ runner.temp }}/letsencrypt"
          fi

      - name: Issue Let's Encrypt certificate
        run: |
          .venv/bin/certbot certonly \
            --dns-route53 \
            --config-dir "${{ runner.temp }}/letsencrypt" \
            --work-dir "${{ runner.temp }}/letsencrypt-work" \
            --logs-dir "${{ runner.temp }}/letsencrypt-logs" \
            --non-interactive \
            --agree-tos \
            --keep-until-expiring \
            --email "${{ env.LETSENCRYPT_EMAIL }}" \
            -d "${{ env.CERT_DOMAIN }}" \
            -d "${{ env.CERT_APEX_DOMAIN }}"

      - name: Validate certificate
        run: |
          openssl x509 \
            -in "${{ runner.temp }}/letsencrypt/live/${{ env.CERT_DOMAIN }}/fullchain.pem" \
            -noout \
            -subject \
            -issuer \
            -dates \
            -ext subjectAltName

      - name: Persist Certbot state to S3
        run: |
          tar -czf "${{ runner.temp }}/letsencrypt.tar.gz" \
            -C "${{ runner.temp }}/letsencrypt" .

          aws s3 cp \
            "${{ runner.temp }}/letsencrypt.tar.gz" \
            "s3://${{ env.CERT_BUCKET }}/${{ env.ACME_STATE_KEY }}"

      - name: Upload active certificate to S3
        run: |
          aws s3 cp \
            "${{ runner.temp }}/letsencrypt/live/${{ env.CERT_DOMAIN }}/fullchain.pem" \
            "s3://${{ env.CERT_BUCKET }}/${{ env.CERT_CURRENT_PREFIX }}/fullchain.pem"

          aws s3 cp \
            "${{ runner.temp }}/letsencrypt/live/${{ env.CERT_DOMAIN }}/privkey.pem" \
            "s3://${{ env.CERT_BUCKET }}/${{ env.CERT_CURRENT_PREFIX }}/privkey.pem"

      - name: Download the new SSL certificate to the EC2 instance
        run: |
          set -euo pipefail

          INSTANCE_ID=$(aws ec2 describe-instances \
            --filters \
              "Name=tag:elasticbeanstalk:environment-name,Values=${{ env.EB_ENV_NAME }}" \
              "Name=instance-state-name,Values=running" \
            --query 'Reservations[].Instances[].InstanceId' \
            --output text)

          if [ -z "$INSTANCE_ID" ]; then
            echo "No running EC2 instance found"
            exit 1
          fi

          echo "EC2 instance found"

          COMMAND_ID=$(aws ssm send-command \
            --instance-ids "$INSTANCE_ID" \
            --document-name "AWS-RunShellScript" \
            --comment "Deploy Let's Encrypt certificate" \
            --parameters 'commands=[
              "mkdir -p /etc/pki/tls/certs/example",
              "aws s3 cp s3://${{ env.CERT_BUCKET }}/${{ env.CERT_CURRENT_PREFIX }}/fullchain.pem /etc/pki/tls/certs/example/fullchain.pem",
              "aws s3 cp s3://${{ env.CERT_BUCKET }}/${{ env.CERT_CURRENT_PREFIX }}/privkey.pem /etc/pki/tls/certs/example/privkey.pem",
              "chmod 600 /etc/pki/tls/certs/example/privkey.pem",
              "nginx -t",
              "systemctl reload nginx",
              "openssl x509 -in /etc/pki/tls/certs/example/fullchain.pem -noout -subject -issuer -dates"
            ]' \
            --query 'Command.CommandId' \
            --output text)

          echo "SSM command submitted"

          set +e

          aws ssm wait command-executed \
            --command-id "$COMMAND_ID" \
            --instance-id "$INSTANCE_ID"

          WAIT_EXIT_CODE=$?

          set -e

          echo "SSM command result:"

          aws ssm get-command-invocation \
            --command-id "$COMMAND_ID" \
            --instance-id "$INSTANCE_ID" \
            --query '{
              Status:Status,
              StandardOutput:StandardOutputContent,
              StandardError:StandardErrorContent
            }'

          if [ "$WAIT_EXIT_CODE" -ne 0 ]; then
            echo "SSM command failed or timed out"
            exit 1
          fi
Enter fullscreen mode Exit fullscreen mode

Final architecture

The final result looks like this:

                     +------------------+
                     |  GitHub Actions  |
                     |   weekly cron    |
                     +--------+---------+
                              |
                             OIDC
                              |
                              v
                     +------------------+
                     |       AWS        |
                     +------------------+
                         |          |
                         |          |
                  Route53 DNS-01    |
                         |          |
                         v          |
                  Let's Encrypt     |
                         |          |
                         +----+-----+
                              |
                              v
                    +--------------------+
                    |         S3         |
                    |                    |
                    | acme-state-prod/   |
                    | current/           |
                    +---------+----------+
                              |
                             SSM
                              |
                              v
                     +------------------+
                     | Elastic Beanstalk|
                     |      EC2 Spot    |
                     +--------+---------+
                              |
                              v
                            Nginx
                              |
                              v
                          WordPress
Enter fullscreen mode Exit fullscreen mode

CloudFront continues to use ACM for the public-facing certificate:

User
  |
  | HTTPS / ACM
  |
CloudFront
  |
  | HTTPS / Let's Encrypt
  |
Nginx
Enter fullscreen mode Exit fullscreen mode

Result

With this change, I removed the manual renewal of a certificate purchased from an external provider from the normal operational workflow.

Now:

  • Let’s Encrypt issues and renews the certificate.
  • Route53 automatically handles the ACME challenge.
  • GitHub Actions runs the process periodically.
  • OIDC avoids permanent AWS credentials.
  • S3 stores both the certificate and the ACME state.
  • SSM updates the active EC2 instance.
  • Nginx validates its configuration before being reloaded.
  • A new Spot instance can rebuild itself using the certificate stored in S3.

And perhaps most importantly, certificate management no longer depends on the lifetime of a specific machine or on me remembering once a year that the certificate needs to be renewed.

Links

Top comments (0)