DEV Community

Blocking Container deployments based on vulnerability scans

I got an AWS question and implemented it to make sure that the option is correct.

A company runs containerized microservices on ECS Fargate across multiple AWS accounts. The security team has four requirements: scan images before production deployment, block images with CRITICAL vulnerabilities automatically, do it without manual intervention, and centralize all findings in a single audit account.

Two solutions cover all four requirements.


Prerequisites

Check these before running terraform apply:

Amazon Inspector must be enabled per account
Inspector is not enabled by default. It must be activated in each member account before ECR Enhanced Scanning works. The recommended approach is enabling it through AWS Organizations so new accounts get it automatically.

Security Hub must be enabled in all accounts
Security Hub also requires explicit activation per account and per region. The aggregation feature requires Security Hub to be active in both the member accounts and the central security account.

The securityhub.tf file runs in a different account
The Security Hub aggregator resources must be applied to the central security account, not to each member account. The Terraform for that file requires a separate provider alias pointing to the aggregator account. This is covered in the securityhub.tf section below.

Terraform executor permissions
The IAM principal running Terraform needs at minimum:

inspector2:Enable
inspector2:AssociateMember
securityhub:EnableSecurityHub
securityhub:CreateFindingAggregator
securityhub:UpdateFindingAggregator
ecr:PutImageScanningConfiguration
ecr:DescribeImageScanFindings
events:PutRule
events:PutTargets
lambda:CreateFunction
lambda:AddPermission
iam:CreateRole
iam:AttachRolePolicy
iam:PutRolePolicy
sns:CreateTopic
sns:Subscribe
Enter fullscreen mode Exit fullscreen mode

AdministratorAccess on each account covers all of these. Lock it down after the initial setup.

VPC not required
All resources in this solution run outside a VPC. Lambda, EventBridge, Inspector, and Security Hub are all regional services accessed via their public endpoints.


The problem

Blocking deployments automatically

A CI/CD pipeline that deploys on every image push will deploy vulnerable images unless something stops it. The gate needs to be automatic, not dependent on someone reviewing a dashboard. It also needs to be event-driven: when a scan finishes, the result should immediately trigger the enforcement logic.

The race condition

Inspector Enhanced Scanning does not complete instantly. Depending on image size and layer count, it can take two to five minutes after a push. If the CI/CD pipeline pushes an image and immediately triggers an ECS deployment, the deployment may start before Inspector finishes and before the Lambda gate has a chance to delete the image.

The safest mitigation is adding a wait step in the pipeline after the push, before the deploy step. A simple approach is polling the Inspector findings API until the scan status returns COMPLETE or a timeout is reached. If the image has been deleted by the gate Lambda during that window, the deploy step will fail when it tries to pull the image. If no CRITICAL findings exist, the image remains in ECR and the deploy proceeds normally.

A secondary mitigation is separating the ECR repository from the deploy step by requiring the deploy step to always reference an image by digest rather than by tag. This prevents a race where a new push overwrites a tag that was already scanned clean.

Centralizing findings across accounts

In a multi-account setup, each account has its own Inspector findings. The security team needs a single place to review all findings across all accounts without logging into each one individually.

Continuous scanning behavior

Enhanced Scanning is not a one-time check on push. Inspector re-evaluates images continuously as new CVEs are published. An image that passes the gate today may be flagged tomorrow if a new CRITICAL vulnerability is discovered in one of its packages. The EventBridge rule will fire again for the new finding, and the Lambda gate will delete the image from ECR even if it is already deployed.

This means deployed containers are not automatically stopped. The gate only prevents new pulls of the vulnerable image. If your ECS task uses imagePullPolicy: always equivalent behavior or if tasks restart, they will fail to pull the deleted image and the task will not start. This is the intended behavior: the image is gone, and a new clean build is required.


The solutions

ECR Enhanced Scanning with Inspector + EventBridge gate

ECR Enhanced Scanning uses Amazon Inspector instead of the built-in basic scanner. Enhanced Scanning checks OS packages and application dependencies continuously, not just on push. It cross-references findings against multiple vulnerability databases (NVD, vendor advisories) and assigns severity scores more accurately than Basic Scanning.

When Inspector finishes scanning an image, it emits an event to EventBridge. The event contains the image digest, repository, and a summary of findings by severity. An EventBridge rule filters for events where CRITICAL findings are present and routes them to a Lambda function.

The Lambda function deletes the image from ECR so it cannot be pulled, then notifies the pipeline via SNS so the deployment step fails with a clear reason.

The pipeline does not need a separate approval gate or polling loop. It pushes the image, the scan runs, and if CRITICAL vulnerabilities are found the image is deleted before the deployment step can reference it.

# inspector.tf

resource "aws_inspector2_enabler" "ecr" {
  account_ids    = [data.aws_caller_identity.current.account_id]
  resource_types = ["ECR"]
}

data "aws_caller_identity" "current" {}
Enter fullscreen mode Exit fullscreen mode

With Inspector enabled for ECR, all repositories in the account get continuous scanning automatically. No per-repository configuration is needed.

# ecr.tf

resource "aws_ecr_repository" "app" {
  for_each = toset(var.ecr_repository_names)

  name                 = each.key
  image_tag_mutability = "IMMUTABLE"

  image_scanning_configuration {
    scan_on_push = true
  }

  encryption_configuration {
    encryption_type = "AES256"
  }

  tags = {
    Name = each.key
  }
}

# Clean up untagged images left behind after the gate Lambda deletes a vulnerable image.
# Without this, deleted images accumulate as untagged layers and incur storage cost.
resource "aws_ecr_lifecycle_policy" "app" {
  for_each   = aws_ecr_repository.app
  repository = each.value.name

  policy = jsonencode({
    rules = [
      {
        rulePriority = 1
        description  = "Remove untagged images immediately"
        selection = {
          tagStatus   = "untagged"
          countType   = "sinceImagePushed"
          countUnit   = "days"
          countNumber = 1
        }
        action = {
          type = "expire"
        }
      },
      {
        rulePriority = 2
        description  = "Keep only the last 20 tagged images per repository"
        selection = {
          tagStatus     = "tagged"
          tagPrefixList = ["v"]
          countType     = "imageCountMoreThan"
          countNumber   = 20
        }
        action = {
          type = "expire"
        }
      }
    ]
  })
}
Enter fullscreen mode Exit fullscreen mode

image_tag_mutability = "IMMUTABLE" prevents overwriting an existing tag. Combined with the scan gate, this means a tag that passed scanning cannot be replaced with a vulnerable image under the same tag.

# iam.tf

resource "aws_iam_role" "scan_gate_lambda" {
  name = "ECRScanGateLambdaRole"

  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect    = "Allow"
      Principal = { Service = "lambda.amazonaws.com" }
      Action    = "sts:AssumeRole"
    }]
  })
}

resource "aws_iam_role_policy" "scan_gate_lambda" {
  name = "ECRScanGateLambdaPolicy"
  role = aws_iam_role.scan_gate_lambda.id

  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [
      {
        Sid    = "ECRDeleteImage"
        Effect = "Allow"
        Action = [
          "ecr:BatchDeleteImage",
          "ecr:DescribeImages",
          "ecr:DescribeImageScanFindings"
        ]
        Resource = [for repo in aws_ecr_repository.app : repo.arn]
      },
      {
        Sid      = "SNSPublish"
        Effect   = "Allow"
        Action   = "sns:Publish"
        Resource = aws_sns_topic.scan_alerts.arn
      },
      {
        Sid    = "Logs"
        Effect = "Allow"
        Action = [
          "logs:CreateLogGroup",
          "logs:CreateLogStream",
          "logs:PutLogEvents"
        ]
        Resource = "arn:aws:logs:*:*:*"
      }
    ]
  })
}
Enter fullscreen mode Exit fullscreen mode
# lambda.tf

resource "aws_lambda_function" "scan_gate" {
  function_name = "ecr-scan-gate"
  role          = aws_iam_role.scan_gate_lambda.arn
  handler       = "index.handler"
  runtime       = "python3.12"
  timeout       = 30

  filename         = data.archive_file.scan_gate.output_path
  source_code_hash = data.archive_file.scan_gate.output_base64sha256

  environment {
    variables = {
      SNS_TOPIC_ARN = aws_sns_topic.scan_alerts.arn
    }
  }
}

data "archive_file" "scan_gate" {
  type        = "zip"
  output_path = "${path.module}/scan_gate.zip"

  source {
    content  = <<-PYTHON
import boto3
import os
import json

ecr = boto3.client('ecr')
sns = boto3.client('sns')

def handler(event, context):
    detail = event.get('detail', {})
    repository = detail.get('repository-name')
    tag = detail.get('image-tags', ['untagged'])[0]
    digest = detail.get('image-digest')
    findings_summary = detail.get('finding-severity-counts', {})
    critical_count = findings_summary.get('CRITICAL', 0)

    print(f"Repository: {repository}, Tag: {tag}, CRITICAL: {critical_count}")

    if critical_count > 0:
        print(f"Deleting image {repository}:{tag} ({digest}) - {critical_count} CRITICAL findings")

        ecr.batch_delete_image(
            repositoryName=repository,
            imageIds=[{'imageDigest': digest}]
        )

        sns.publish(
            TopicArn=os.environ['SNS_TOPIC_ARN'],
            Subject=f"Deployment blocked: {repository}:{tag}",
            Message=json.dumps({
                'repository': repository,
                'tag': tag,
                'digest': digest,
                'critical_findings': critical_count,
                'action': 'image deleted from ECR',
                'reason': 'CRITICAL vulnerabilities found during Inspector scan'
            }, indent=2)
        )

    return {'statusCode': 200}
    PYTHON
    filename = "index.py"
  }
}

resource "aws_lambda_permission" "eventbridge" {
  statement_id  = "AllowEventBridgeInvocation"
  action        = "lambda:InvokeFunction"
  function_name = aws_lambda_function.scan_gate.function_name
  principal     = "events.amazonaws.com"
  source_arn    = aws_cloudwatch_event_rule.ecr_scan_critical.arn
}
Enter fullscreen mode Exit fullscreen mode
# eventbridge.tf

resource "aws_cloudwatch_event_rule" "ecr_scan_critical" {
  name        = "ecr-inspector-scan-complete"
  description = "Triggers when Inspector finds a CRITICAL vulnerability in an ECR image"

  event_pattern = jsonencode({
    source      = ["aws.inspector2"]
    detail-type = ["Inspector2 Finding"]
    detail = {
      resources = {
        type = ["AWS_ECR_CONTAINER_IMAGE"]
      }
      severity = ["CRITICAL"]
    }
  })
}

resource "aws_cloudwatch_event_target" "scan_gate_lambda" {
  rule      = aws_cloudwatch_event_rule.ecr_scan_critical.name
  target_id = "ScanGateLambda"
  arn       = aws_lambda_function.scan_gate.arn
}
Enter fullscreen mode Exit fullscreen mode
# sns.tf

resource "aws_sns_topic" "scan_alerts" {
  name = "ecr-scan-critical-alerts"
}

resource "aws_sns_topic_subscription" "email" {
  topic_arn = aws_sns_topic.scan_alerts.arn
  protocol  = "email"
  endpoint  = var.pipeline_notification_email
}
Enter fullscreen mode Exit fullscreen mode

Inspector findings to Security Hub with cross-account aggregation

Inspector publishes all ECR findings directly to Security Hub with no additional configuration beyond enabling both services. Security Hub normalizes findings into the ASFF (Amazon Security Finding Format), making them queryable and filterable in a consistent way regardless of which account or scanner produced them.

Security Hub's finding aggregation feature designates one account as the aggregator. All member accounts linked through Organizations forward their findings to the aggregator account automatically. The security team reviews findings in one place, sets up alerting once, and writes queries once.

No custom code. No cross-account Lambda invocations. No S3 buckets collecting logs from each account.

The securityhub.tf below must be applied to the central security account, not to member accounts. Use a separate Terraform provider alias pointing to that account.

# securityhub.tf
# Apply this with the security account provider alias, not the default provider.
# Example provider configuration in main.tf:
#
# provider "aws" {
#   alias  = "security"
#   region = var.aws_region
#   assume_role {
#     role_arn = "arn:aws:iam::${var.security_hub_aggregator_account_id}:role/TerraformRole"
#   }
# }

resource "aws_securityhub_account" "aggregator" {
  provider = aws.security
}

resource "aws_securityhub_finding_aggregator" "main" {
  provider     = aws.security
  linking_mode = "ALL_REGIONS"

  depends_on = [aws_securityhub_account.aggregator]
}

resource "aws_securityhub_member" "members" {
  provider = aws.security
  for_each = toset(var.member_account_ids)

  account_id = each.key
  invite     = true

  depends_on = [aws_securityhub_account.aggregator]
}

resource "aws_securityhub_product_subscription" "inspector" {
  provider    = aws.security
  product_arn = "arn:aws:securityhub:${var.aws_region}::product/aws/inspector"

  depends_on = [aws_securityhub_account.aggregator]
}
Enter fullscreen mode Exit fullscreen mode

Full variable list

# variables.tf

variable "aws_region" {
  description = "AWS region"
  type        = string
  default     = "us-east-1"
}

variable "ecr_repository_names" {
  description = "List of ECR repository names to protect"
  type        = list(string)
}

variable "pipeline_notification_email" {
  description = "Email address that receives notifications when a deployment is blocked"
  type        = string
}

variable "security_hub_aggregator_account_id" {
  description = "AWS account ID of the central security account that aggregates findings"
  type        = string
}

variable "member_account_ids" {
  description = "List of member account IDs to link to the Security Hub aggregator"
  type        = list(string)
}
Enter fullscreen mode Exit fullscreen mode

Outputs

# outputs.tf

output "scan_gate_lambda_arn" {
  description = "ARN of the Lambda function that enforces the scan gate"
  value       = aws_lambda_function.scan_gate.arn
}

output "scan_alerts_topic_arn" {
  description = "SNS topic ARN that receives blocked deployment notifications"
  value       = aws_sns_topic.scan_alerts.arn
}

output "ecr_repository_urls" {
  description = "ECR repository URLs for the CI/CD pipeline"
  value       = { for k, v in aws_ecr_repository.app : k => v.repository_url }
}
Enter fullscreen mode Exit fullscreen mode

Validating the deployment

Check 1: Inspector is enabled for ECR

aws inspector2 list-coverage \
  --filter-criteria '{"resourceType":[{"comparison":"EQUALS","value":"AWS_ECR_CONTAINER_IMAGE"}]}' \
  --query 'coveredResources[0].scanStatus.statusCode'
# Expected: "ACTIVE"
Enter fullscreen mode Exit fullscreen mode

Check 2: EventBridge rule is active

aws events describe-rule \
  --name ecr-inspector-scan-complete \
  --query 'State'
# Expected: "ENABLED"
Enter fullscreen mode Exit fullscreen mode

Check 3: trigger the gate with a known vulnerable image

Push an image with a known CRITICAL vulnerability to one of the protected repositories. Inspector will scan it within a few minutes and the EventBridge rule will fire. Confirm the image was deleted:

aws ecr list-images \
  --repository-name <your-repo-name> \
  --query 'imageIds[*]'
Enter fullscreen mode Exit fullscreen mode

The image should no longer appear. Check the SNS email for the blocked deployment notification.

Check 4: Security Hub is receiving Inspector findings

Run this from the aggregator account:

aws securityhub get-findings \
  --filters '{"ProductName":[{"Value":"Inspector","Comparison":"EQUALS"}]}' \
  --query 'Findings[0].{Title:Title,Severity:Severity.Label,Account:AwsAccountId}' \
  --region <aggregator-region>
Enter fullscreen mode Exit fullscreen mode

Findings from all member accounts should appear here.


Why the other options do not work?

Option 1: Basic Scanning with CodePipeline approval gate

Basic Scanning only runs on push and checks OS packages against a single vulnerability feed. It misses application-layer vulnerabilities that Inspector catches. The CodePipeline approval gate also requires per-pipeline integration: every pipeline that deploys from ECR needs its own gate, and each gate needs to know which image it is checking. This scales poorly across multiple repositories and pipelines.

Option 2: CloudWatch metric filter on ECR scan events

ECR scan result events go to EventBridge natively. Routing them through CloudWatch Logs first adds a step with no benefit. You lose the structured event payload, need a metric filter pattern to extract severity, and then need to reconnect back to a Lambda anyway. EventBridge is the direct path.

Option 3: AWS Config with ecr-image-scan-on-push rule

The Config managed rule checks whether scan-on-push is enabled for a repository. It evaluates repository configuration, not scan results. It has no visibility into whether a scan found CRITICAL vulnerabilities. Remediating non-compliance means enabling scan-on-push on repositories that have it disabled, not blocking images based on findings.


Architecture

The two flows run in parallel. Inspector feeds both the blocking gate and the centralized audit trail at the same time.

CI/CD pipeline
    |
    v
docker push --> ECR repository (IMMUTABLE tags)
                    |
              Inspector scans image (2-5 min)
                    |
          +---------+---------+
          |                   |
          v                   v
   EventBridge rule     Security Hub
   (CRITICAL finding)   (all findings, all accounts)
          |                   |
          v                   v
   Lambda gate          Aggregator account
          |             Single audit view
   CRITICAL > 0?
          |
     Yes  |  No
     |         |
     v         v
  Delete    Image stays
  image     Pipeline deploys
  from ECR  normally
     |
  SNS notification
  Pipeline fails
  (image not found)

Inspector also re-evaluates images continuously.
New CVE published --> new EventBridge event --> gate fires again.
Deployed containers are not stopped automatically,
but the image is removed from ECR so no new tasks can pull it.
Enter fullscreen mode Exit fullscreen mode

Cost to keep in mind

Amazon Inspector charges per container image scanned per month. The first 2,500 images per account per month are included in the free tier during the first 30 days. After that it charges per image. Security Hub charges per finding ingested, with a free tier of 10,000 findings per month per account for the first 12 months.

For a multi-account environment with active CI/CD pipelines, Inspector and Security Hub costs scale with the number of images pushed. The ECR lifecycle policy prevents untagged layer accumulation from inflating storage costs when the gate deletes images.

Top comments (0)