DEV Community

Varun Sharma
Varun Sharma

Posted on Originally published at github.com

Building an AWS DevOps Agent That Diagnoses but Never Deploys

Most CI/CD pipeline failures take longer to diagnose than to fix. An IAM role is missing an action after an update. A Terraform state lock didn't clear after a cancelled job. A container layer cache busted on an architecture mismatch. The fix itself is usually a two-line patch, but finding it means waiting for a massive raw terminal log to render in GitHub Actions and hunting for the error line.

I wanted to automate that triage step without adding complex machinery. When a GitHub Actions run fails, a lightweight workflow pulls the failure log, runs it through Amazon Bedrock using Claude, and comments on the pull request with a diagnosis and suggested diff. The implementation is plain Python, GitHub Actions, AWS OIDC federation, and Bedrock. The code is on GitHub at github.com/sharma-the-karma/aws-devops-agent.

How it works

The triage job only runs when a pipeline fails, so there is no persistent infrastructure or long-lived AWS keys in repository secrets.

flowchart LR
    A[GitHub Actions CI Workflow] -->|Fails| B[Triage Trigger]
    B --> C[Log Distiller Engine]
    C -->|High-Signal Context| D[AWS Bedrock Converse API]
    D -->|Claude| E[Diagnostic Report]
    E --> F[PR Comment with Diff]

    subgraph AWS Cloud
        D
        G[IAM OIDC Role] -.->|Federates| B
    end

The sequence is straightforward:

  1. A CI job fails on a pull request, triggering the triage workflow via the workflow_run event.
  2. The runner requests a temporary token from GitHub's OIDC provider and assumes an IAM role in AWS.
  3. A Python script downloads the failed step's logs, strips ANSI color codes, and extracts the lines surrounding the failure signature.
  4. The distilled snippet is sent to Bedrock's Converse API.
  5. The diagnosis is posted back to the pull request, editing the previous comment if one already exists.

What I learned

Raw logs wreck the model's reasoning

Feeding the entire raw console log into an LLM does not work well. A typical CI run prints thousands of lines of package downloads, container build steps, and environment setup before hitting the failure. In testing, sending that full output bloated the prompt to tens of thousands of tokens. The model frequently latched onto benign compiler warnings earlier in the log rather than the actual fatal error at the bottom, while adding 15 to 20 seconds of unnecessary latency.

The fix was a small distillation utility (src/log_distiller.py). It strips ANSI formatting and scans for common failure markers like AccessDeniedException, Python tracebacks, Terraform errors, and non-zero exit codes. When it finds a match, it extracts a rolling window of 15 lines before and 35 lines after:

import re
from typing import Dict, Any

ANSI_ESCAPE_PATTERN = re.compile(r"\x1B(?:[@-Z\\-_]|\[[0-?]*[ -/]*[@-~])")

ERROR_INDICATORS = [
    re.compile(r"error[:\s]", re.IGNORECASE),
    re.compile(r"traceback \(most recent call last\):", re.IGNORECASE),
    re.compile(r"failed with exit code \d+", re.IGNORECASE),
    re.compile(r"accessdeniedException", re.IGNORECASE),
    re.compile(r"terraform error[:\s]", re.IGNORECASE),
]

def extract_failure_context(raw_log: str, window_before: int = 15, window_after: int = 35) -> Dict[str, Any]:
    clean_lines = ANSI_ESCAPE_PATTERN.sub("", raw_log).splitlines()
    selected_indices = set()

    for idx, line in enumerate(clean_lines):
        for pattern in ERROR_INDICATORS:
            if pattern.search(line):
                start = max(0, idx - window_before)
                end = min(len(clean_lines), idx + window_after + 1)
                selected_indices.update(range(start, end))
                break

    sorted_indices = sorted(selected_indices)
    extracted = [f"L{i+1:04d}: {clean_lines[i]}" for i in sorted_indices]

    return {
        "distilled_log": "\n".join(extracted),
        "total_lines_analyzed": len(clean_lines),
        "lines_extracted": len(extracted)
    }
Enter fullscreen mode Exit fullscreen mode

This reduces the payload from tens of thousands of tokens down to roughly 500 to 800 tokens of high-signal context. Response times drop to a few seconds, and the model focuses on the actual failure rather than ambient build noise.

Don't give the agent permanent keys

Avoid creating static AWS access keys (AKIA...) for CI/CD runners. If a pull request modifies workflow files or runs untrusted scripts, stored secrets can be exposed.

Instead, configure an IAM OIDC Identity Provider in AWS for GitHub Actions. The runner exchanges a short-lived GitHub token for temporary STS credentials.

data "aws_iam_policy_document" "github_oidc_trust" {
  statement {
    effect  = "Allow"
    actions = ["sts:AssumeRoleWithWebIdentity"]

    principals {
      type        = "Federated"
      identifiers = [aws_iam_openid_connect_provider.github.arn]
    }

    condition {
      test     = "StringEquals"
      variable = "token.actions.githubusercontent.com:aud"
      values   = ["sts.amazonaws.com"]
    }

    condition {
      test     = "StringLike"
      variable = "token.actions.githubusercontent.com:sub"
      values   = ["repo:sharma-the-karma/aws-devops-agent:ref:refs/heads/main"]
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Because workflow_run triggers execute in the context of the default branch, the sub claim can be pinned to ref:refs/heads/main rather than using a wildcard. The role's policy grants bedrock:InvokeModel and nothing else. The agent cannot read S3, modify security groups, or alter infrastructure.

Use the Converse API

Amazon Bedrock provides the Converse API, which standardizes system prompts, messages, and inference parameters across foundation models.

import boto3

SYSTEM_PROMPT = """You are an AWS DevOps and Infrastructure Architect.
Inspect failing CI/CD logs, diagnostic traces, and infrastructure errors.
Provide a concise, practical triage assessment with actionable solutions.

Structure your response into:
1. Root Cause Breakdown (what failed and its category)
2. Immediate Fix / Patch (exact code snippet or command)
3. Preventive Action & Blast Radius (how to prevent it and potential side-effects)
"""

client = boto3.client("bedrock-runtime", region_name="us-east-1")

response = client.converse(
    modelId="anthropic.claude-3-5-sonnet-20241022-v2:0",
    messages=[
        {
            "role": "user",
            "content": [{"text": f"CI/CD Failure Trace:\n{distilled_log}"}]
        }
    ],
    system=[{"text": SYSTEM_PROMPT}],
    inferenceConfig={
        "maxTokens": 2048,
        "temperature": 0.2,
        "topP": 0.9,
    }
)

diagnostic_report = response["output"]["message"]["content"][0]["text"]
Enter fullscreen mode Exit fullscreen mode

I used anthropic.claude-3-5-sonnet-20241022-v2:0 here, but for standard CI failure triage where you just need quick stack trace interpretation, anthropic.claude-3-5-haiku-20241022-v1:0 is faster, cheaper, and handles compiler errors just as effectively. A temperature of 0.2 keeps the diagnostic output consistent across runs.

Update the PR comment instead of adding a new one

When developers push multiple commits while troubleshooting a broken branch, bot comments can easily flood the pull request.

To prevent this, the agent tags each comment with a hidden HTML marker (<!-- aws-devops-agent-triage -->) and inspects existing comments before posting. If an earlier report exists, it edits the comment in place:

COMMENT_TAG = "<!-- aws-devops-agent-triage -->"

def post_or_update_pr_triage_comment(repo, pr_number: int, analysis_markdown: str):
    body = f"{COMMENT_TAG}\n### AWS DevOps Agent Triage Report\n\n{analysis_markdown}"
    pull_request = repo.get_pull(pr_number)

    for comment in pull_request.get_issue_comments():
        if COMMENT_TAG in comment.body:
            comment.edit(body)
            return "Updated existing triage comment."

    pull_request.create_issue_comment(body)
    return "Created new triage comment."
Enter fullscreen mode Exit fullscreen mode

Note that the workflow's GITHUB_TOKEN requires pull-requests: write permissions for this step.

Never let the agent apply infrastructure changes

Letting an agent automatically commit code or run terraform apply is dangerous.

A simple example illustrates why: an ECS service update fails because a subnet is misconfigured. An agent can correctly identify the subnet mismatch and propose the right configuration. However, changing subnet attributes on certain AWS resources triggers a replacement, destroying the existing resource before creating the new one. An automated agent executing that change on its own can easily cause unintended downtime or data loss.

The agent's role is to diagnose the failure, explain the blast radius, and provide a diff. The engineer reviews the diff and decides when and how to deploy it.

What a report looks like

Here is a representative example of what the agent posts to a pull request when an ECS service update fails on an IAM permission:

AWS DevOps Agent Triage Report

1. Root Cause Breakdown
The deployment failed with AccessDeniedException while calling ecs:UpdateServicePrimaryTaskSet. The execution role assumed by the CI runner lacks this action in its attached policy. This is an infrastructure / IAM permissions issue.

2. Immediate Fix / Patch
Add the missing action to your CI deployer policy in infra/iam_deployer.tf:

 statement {
   actions = [
     "ecs:UpdateService",
     "ecs:DescribeServices",
+    "ecs:UpdateServicePrimaryTaskSet"
   ]
   resources = [aws_ecs_service.api_service.arn]
 }

3. Preventive Action & Blast Radius
Review deployment policies whenever you introduce CodeDeploy or blue/green task sets on ECS services. Applying this change is non-destructive and requires no resource recreation.

Try it yourself

Clone the repo:

git clone https://github.com/sharma-the-karma/aws-devops-agent.git
cd aws-devops-agent
Enter fullscreen mode Exit fullscreen mode

Deploy the IAM role:

cd infra
terraform init
terraform apply -var="github_org=sharma-the-karma" -var="github_repo=aws-devops-agent"
Enter fullscreen mode Exit fullscreen mode

Add the role ARN from the output as AWS_ROLE_ARN in your repository secrets (Settings > Secrets and variables > Actions).

To test locally against the sample log without calling anything for real:

cd src
pip install -r requirements.txt
python triage_agent.py --log-file sample_failure.log --dry-run
Enter fullscreen mode Exit fullscreen mode

Next steps for this setup include hooking into CloudWatch alarm webhooks so the same triage flow works for runtime ECS task crashes and Lambda errors, rather than just CI pipeline failures.

The complete code, Terraform configuration, and sample logs are available on GitHub: github.com/sharma-the-karma/aws-devops-agent.

Top comments (0)