Most CI/CD pipeline failures take longer to diagnose than to fix. An IAM role is missing an action after an update. A Terraform state lock didn't clear after a cancelled job. A container layer cache busted on an architecture mismatch. The fix itself is usually a two-line patch, but finding it means waiting for a massive raw terminal log to render in GitHub Actions and hunting for the error line.
I wanted to automate that triage step without adding complex machinery. When a GitHub Actions run fails, a lightweight workflow pulls the failure log, runs it through Amazon Bedrock using Claude, and comments on the pull request with a diagnosis and suggested diff. The implementation is plain Python, GitHub Actions, AWS OIDC federation, and Bedrock. The code is on GitHub at github.com/sharma-the-karma/aws-devops-agent.
How it works
The triage job only runs when a pipeline fails, so there is no persistent infrastructure or long-lived AWS keys in repository secrets.
flowchart LR
A[GitHub Actions CI Workflow] -->|Fails| B[Triage Trigger]
B --> C[Log Distiller Engine]
C -->|High-Signal Context| D[AWS Bedrock Converse API]
D -->|Claude| E[Diagnostic Report]
E --> F[PR Comment with Diff]
subgraph AWS Cloud
D
G[IAM OIDC Role] -.->|Federates| B
end
The sequence is straightforward:
- A CI job fails on a pull request, triggering the triage workflow via the
workflow_runevent. - The runner requests a temporary token from GitHub's OIDC provider and assumes an IAM role in AWS.
- A Python script downloads the failed step's logs, strips ANSI color codes, and extracts the lines surrounding the failure signature.
- The distilled snippet is sent to Bedrock's Converse API.
- The diagnosis is posted back to the pull request, editing the previous comment if one already exists.
What I learned
Raw logs wreck the model's reasoning
Feeding the entire raw console log into an LLM does not work well. A typical CI run prints thousands of lines of package downloads, container build steps, and environment setup before hitting the failure. In testing, sending that full output bloated the prompt to tens of thousands of tokens. The model frequently latched onto benign compiler warnings earlier in the log rather than the actual fatal error at the bottom, while adding 15 to 20 seconds of unnecessary latency.
The fix was a small distillation utility (src/log_distiller.py). It strips ANSI formatting and scans for common failure markers like AccessDeniedException, Python tracebacks, Terraform errors, and non-zero exit codes. When it finds a match, it extracts a rolling window of 15 lines before and 35 lines after:
import re
from typing import Dict, Any
ANSI_ESCAPE_PATTERN = re.compile(r"\x1B(?:[@-Z\\-_]|\[[0-?]*[ -/]*[@-~])")
ERROR_INDICATORS = [
re.compile(r"error[:\s]", re.IGNORECASE),
re.compile(r"traceback \(most recent call last\):", re.IGNORECASE),
re.compile(r"failed with exit code \d+", re.IGNORECASE),
re.compile(r"accessdeniedException", re.IGNORECASE),
re.compile(r"terraform error[:\s]", re.IGNORECASE),
]
def extract_failure_context(raw_log: str, window_before: int = 15, window_after: int = 35) -> Dict[str, Any]:
clean_lines = ANSI_ESCAPE_PATTERN.sub("", raw_log).splitlines()
selected_indices = set()
for idx, line in enumerate(clean_lines):
for pattern in ERROR_INDICATORS:
if pattern.search(line):
start = max(0, idx - window_before)
end = min(len(clean_lines), idx + window_after + 1)
selected_indices.update(range(start, end))
break
sorted_indices = sorted(selected_indices)
extracted = [f"L{i+1:04d}: {clean_lines[i]}" for i in sorted_indices]
return {
"distilled_log": "\n".join(extracted),
"total_lines_analyzed": len(clean_lines),
"lines_extracted": len(extracted)
}
This reduces the payload from tens of thousands of tokens down to roughly 500 to 800 tokens of high-signal context. Response times drop to a few seconds, and the model focuses on the actual failure rather than ambient build noise.
Don't give the agent permanent keys
Avoid creating static AWS access keys (AKIA...) for CI/CD runners. If a pull request modifies workflow files or runs untrusted scripts, stored secrets can be exposed.
Instead, configure an IAM OIDC Identity Provider in AWS for GitHub Actions. The runner exchanges a short-lived GitHub token for temporary STS credentials.
data "aws_iam_policy_document" "github_oidc_trust" {
statement {
effect = "Allow"
actions = ["sts:AssumeRoleWithWebIdentity"]
principals {
type = "Federated"
identifiers = [aws_iam_openid_connect_provider.github.arn]
}
condition {
test = "StringEquals"
variable = "token.actions.githubusercontent.com:aud"
values = ["sts.amazonaws.com"]
}
condition {
test = "StringLike"
variable = "token.actions.githubusercontent.com:sub"
values = ["repo:sharma-the-karma/aws-devops-agent:ref:refs/heads/main"]
}
}
}
Because workflow_run triggers execute in the context of the default branch, the sub claim can be pinned to ref:refs/heads/main rather than using a wildcard. The role's policy grants bedrock:InvokeModel and nothing else. The agent cannot read S3, modify security groups, or alter infrastructure.
Use the Converse API
Amazon Bedrock provides the Converse API, which standardizes system prompts, messages, and inference parameters across foundation models.
import boto3
SYSTEM_PROMPT = """You are an AWS DevOps and Infrastructure Architect.
Inspect failing CI/CD logs, diagnostic traces, and infrastructure errors.
Provide a concise, practical triage assessment with actionable solutions.
Structure your response into:
1. Root Cause Breakdown (what failed and its category)
2. Immediate Fix / Patch (exact code snippet or command)
3. Preventive Action & Blast Radius (how to prevent it and potential side-effects)
"""
client = boto3.client("bedrock-runtime", region_name="us-east-1")
response = client.converse(
modelId="anthropic.claude-3-5-sonnet-20241022-v2:0",
messages=[
{
"role": "user",
"content": [{"text": f"CI/CD Failure Trace:\n{distilled_log}"}]
}
],
system=[{"text": SYSTEM_PROMPT}],
inferenceConfig={
"maxTokens": 2048,
"temperature": 0.2,
"topP": 0.9,
}
)
diagnostic_report = response["output"]["message"]["content"][0]["text"]
I used anthropic.claude-3-5-sonnet-20241022-v2:0 here, but for standard CI failure triage where you just need quick stack trace interpretation, anthropic.claude-3-5-haiku-20241022-v1:0 is faster, cheaper, and handles compiler errors just as effectively. A temperature of 0.2 keeps the diagnostic output consistent across runs.
Update the PR comment instead of adding a new one
When developers push multiple commits while troubleshooting a broken branch, bot comments can easily flood the pull request.
To prevent this, the agent tags each comment with a hidden HTML marker (<!-- aws-devops-agent-triage -->) and inspects existing comments before posting. If an earlier report exists, it edits the comment in place:
COMMENT_TAG = "<!-- aws-devops-agent-triage -->"
def post_or_update_pr_triage_comment(repo, pr_number: int, analysis_markdown: str):
body = f"{COMMENT_TAG}\n### AWS DevOps Agent Triage Report\n\n{analysis_markdown}"
pull_request = repo.get_pull(pr_number)
for comment in pull_request.get_issue_comments():
if COMMENT_TAG in comment.body:
comment.edit(body)
return "Updated existing triage comment."
pull_request.create_issue_comment(body)
return "Created new triage comment."
Note that the workflow's GITHUB_TOKEN requires pull-requests: write permissions for this step.
Never let the agent apply infrastructure changes
Letting an agent automatically commit code or run terraform apply is dangerous.
A simple example illustrates why: an ECS service update fails because a subnet is misconfigured. An agent can correctly identify the subnet mismatch and propose the right configuration. However, changing subnet attributes on certain AWS resources triggers a replacement, destroying the existing resource before creating the new one. An automated agent executing that change on its own can easily cause unintended downtime or data loss.
The agent's role is to diagnose the failure, explain the blast radius, and provide a diff. The engineer reviews the diff and decides when and how to deploy it.
What a report looks like
Here is a representative example of what the agent posts to a pull request when an ECS service update fails on an IAM permission:
AWS DevOps Agent Triage Report
1. Root Cause Breakdown
The deployment failed withAccessDeniedExceptionwhile callingecs:UpdateServicePrimaryTaskSet. The execution role assumed by the CI runner lacks this action in its attached policy. This is an infrastructure / IAM permissions issue.2. Immediate Fix / Patch
Add the missing action to your CI deployer policy ininfra/iam_deployer.tf:statement { actions = [ "ecs:UpdateService", "ecs:DescribeServices", + "ecs:UpdateServicePrimaryTaskSet" ] resources = [aws_ecs_service.api_service.arn] }3. Preventive Action & Blast Radius
Review deployment policies whenever you introduce CodeDeploy or blue/green task sets on ECS services. Applying this change is non-destructive and requires no resource recreation.
Try it yourself
Clone the repo:
git clone https://github.com/sharma-the-karma/aws-devops-agent.git
cd aws-devops-agent
Deploy the IAM role:
cd infra
terraform init
terraform apply -var="github_org=sharma-the-karma" -var="github_repo=aws-devops-agent"
Add the role ARN from the output as AWS_ROLE_ARN in your repository secrets (Settings > Secrets and variables > Actions).
To test locally against the sample log without calling anything for real:
cd src
pip install -r requirements.txt
python triage_agent.py --log-file sample_failure.log --dry-run
Next steps for this setup include hooking into CloudWatch alarm webhooks so the same triage flow works for runtime ECS task crashes and Lambda errors, rather than just CI pipeline failures.
The complete code, Terraform configuration, and sample logs are available on GitHub: github.com/sharma-the-karma/aws-devops-agent.
Top comments (0)