In the agentic engineering world, coding agents generate a lot of code quickly. The bottleneck shifts from writing code to verifying it. I’ve written about this shift before — the need for explicit systems and multilayer verification (see my earlier posts on dev.to/remojansen, especially When Code Gets Cheap, Verification Becomes Expensive).
This post documents one concrete, useful layer of verification you can build today: a scheduled automation that periodically hunts for recurrent 4xx/5xx errors (or high-error-rate endpoints), performs root-cause analysis, and attempts a TDD-style fix. Over time this compounds into greater system resilience.
We’ll use two recent capabilities:
- Automations (Preview) in VS Code + GitHub Copilot
- The Datadog MCP Server so the agent can query real observability data
1. Automations (Preview) in VS Code GitHub Copilot
Automations let you save an agent task (prompt + session configuration) and run it on demand or on a recurring schedule (hourly, daily, or weekly). Routine work no longer has to be started manually.
From the VS Code 1.137 release notes:
- Enable the setting
chat.automations.enabled. - Open the Agents window and select Automations in the sidebar.
- Start from a template (catch up on changes, triage issues, find bugs) or write your own prompt and schedule.
- Run on demand first to review behavior, then enable the schedule.
- Automations are in Preview and rolling out gradually (off by default on Stable, on by default in Insiders).
How to enable and create one
- Open Settings and set
chat.automations.enabledtotrue. - Open the Agents window → Automations → Create Automation (or pick a template).
- Give it a clear name and a precise prompt that states the task, scope, and expected output.
- Choose workspace, agent, model, and permission level.
- Set Schedule to Manual first, create it, run it with Run now, review the session, then edit and switch to Weekly (or Daily) once you’re happy.
Important operational notes from the official docs:
- The machine / Agent Host must be available for scheduled runs.
- An automation runs one session at a time.
- Review the agent’s permission level carefully before enabling unattended schedules.
Official documentation: Create and manage agent automations
2. Enabling the Datadog MCP Server in VS Code
The Datadog MCP Server gives the agent structured access to logs, APM, metrics, error tracking, etc.
Recommended path (from the official Datadog docs):
- Install the official Datadog extension for VS Code.
- Sign in to your Datadog account inside the extension.
- Run the command Datadog: Open MCP Configuration Assistant and follow the guided setup.
Alternatively (manual):
- Open the command palette → MCP: Open User Configuration (or edit
.vscode/mcp.json/ user MCP config). - Add a server entry pointing at the Datadog MCP endpoint for your site, e.g.:
{
"servers": {
"datadog": {
"type": "http",
"url": "https://mcp.datadoghq.com/v1/mcp?toolsets=apm,error-tracking,logs"
}
}
}
(Use the correct regional endpoint for your Datadog site. Replace or expand the toolsets query parameter as needed.)
After configuration, start the server and complete the OAuth flow when prompted. Confirm the tools appear in the agent’s tool picker.
Critical warning: the 128-tool limit
GitHub Copilot (and the underlying models) enforce a hard limit of 128 tools per request. Many MCP servers (including Datadog when all toolsets are enabled, plus other MCPs you may have) easily exceed this. When the limit is hit you get errors such as “You may not include more than 128 tools in your request” and other tools become unavailable.
Best practice: load only the toolsets you actually need by using the toolsets query parameter in the MCP URL (or the equivalent filtering mechanism). Example toolsets relevant to error hunting: apm, error-tracking, logs. Avoid toolsets=all unless you are certain the total stays under the limit or you use virtual tools / selective enabling in the tools picker.
You can also deselect tools or entire MCP servers in the Chat / Agent tools picker for a given session.
3. A Verification Layer for Recurrent Errors
Agents produce a lot of code; the systems that stay healthy are the ones that continuously look for the same classes of failure and close the loop with a test + fix.
Here’s a practical pattern:
Skills (slash commands / reusable prompts) you can develop
/find-top-recurrent-errors
Query Datadog (via MCP) for the top 5 repeating 4xx/5xx errors or the endpoints with the highest error rates over the last 7 days. Surface frequency, sample traces/logs, and affected services./root-cause-analysis
Given the error data, dig into traces, logs, recent deployments, and code paths to propose a concrete root cause./tdd-bug-fix
Write a failing test that reproduces the root cause, then implement the minimal fix that makes the test pass. Prefer small, reviewable changes.
Putting it together in an Automation
Create an automation that runs, for example, every Monday morning:
You are a reliability-focused agent.
1. Use the Datadog MCP tools to identify the top 5 recurrent 4xx/5xx errors (or endpoints with the highest error rates) from the past 7 days.
2. For each, run a focused root-cause analysis using available traces, logs, and code context.
3. For the highest-impact issue that looks fixable in this codebase, follow a TDD approach:
- Write a test that reproduces the failure.
- Implement the smallest correct fix.
- Ensure the test passes and no obvious regressions are introduced.
4. Summarize findings, the proposed fix, and any remaining risks. Do not push or merge; leave changes for human review.
Prefer precision over volume. If data is insufficient, say so clearly.
Schedule it weekly (or daily if your error volume is high). Start with Manual runs until the prompt and permissions feel solid.
Over successive weeks the automation surfaces the same recurring problems, the skills improve, and the codebase gradually becomes more resilient because the verification loop is explicit and repeated.
Why this matters
When code generation is cheap, the expensive part is making sure the system keeps working under real traffic. A scheduled agent that looks at production error patterns, reasons about root cause, and produces a test + fix is one concrete multilayer verification practice you can put in place today.
It is not a replacement for human review, observability culture, or proper incident process — it is an additional, automated layer that compounds over time.
Try the automation with a conservative permission level and a carefully scoped toolset first. Iterate on the prompt and the skills. The goal is not fully autonomous fixing; it is reliable, repeatable detection and a high-quality starting point for a human engineer.
Happy (and more resilient) coding.
Top comments (2)
remo, "when code gets cheap, verification becomes expensive" is the absolute thesis statement of the ai era. this is a fantastic, actionable breakdown of how to actually build that verification layer.
your warning about the 128-tool limit is a massive, practical insight. so many devs blindly enable
toolsets=alland wonder why the agent hallucinates, gets confused, or hits context limits. strictly scoping the mcp tools (like justapm, error-tracking, logs) is a security and performance necessity, not just an optimization.this perfectly mirrors the "two-writer rule" and deterministic validation i enforce in koda. while your datadog mcp setup is brilliant for post-deployment resilience (catching and fixing prod errors via tdd), building on constrained devices means we also have to focus heavily on pre-execution verification. if an agent generates a schema-violating payload or tries to access a restricted api, a strict, deterministic validator must "fail closed" before it ever reaches the execution layer.
combining scoped mcp tools for broad system health with strict, deterministic guards at the edge is exactly how we build truly resilient agentic systems.
fantastic, deeply practical post. saving this for my own automation workflows! 🐯🛡️
tr.ee/dev-to