DEV Community

Cover image for From Vibing Chaos to Reliable Agents: The What, Why, and How of Harness Engineering
Darren "Dazbo" Lester for Google Developer Experts

Posted on Originally published at Medium on

From Vibing Chaos to Reliable Agents: The What, Why, and How of Harness Engineering

You are in Part 1 of the series “Beyond Vibe Coding: The Engineering Blueprint for Reliable AI Agents”

🗺️ The Series Roadmap: Beyond Vibe Coding

Agentic software development is drowning in fragmented buzzwords: Harness Engineering, Loop Engineering, Context Engineering, Spec-Driven Development… Are they competing frameworks? Redundant techniques? Or fundamentally different and essential?

This three-part series unifies these concepts into a single, cohesive architectural blueprint for building software with AI without writing code manually:

  • Part 1: From Vibing Chaos to Reliable Agents: The What, Why, and How of Harness Engineering (This Article) We diagnose the issues of vibe coding, unpack the paradigm shift of shifting developer interventions left, establish our bedrock formula (Agent = Model + Harness), explore a motorsport analogy, and deconstruct deterministic foundations that make up a production developer environment.
  • Part 2: Loop Engineering in the Harness: Autonomous, Self-Correcting Systems We fire up the dynamic runtime engine of the harness: deconstructing the loop components, establishing hybrid evaluation rubrics (deterministic tests + LLM goal fidelity + tool trajectory verification), and applying the Ratchet Principle.
  • Part 3: Spec-Driven Development For-the-Win: Disposable Code, Living Blueprints, and Building Like a Boss We explore Spec-Driven Development (where code is disposable and specifications are the real IP), master Google Antigravity’s unified developer harness, and deploy enterprise-ready agents to Google Cloud.

Introduction: The Chaos of Vibe Coding

From Vibing Chaos to Harness Engineering

Was your first vibe coding experience like mine? Type in a single prompt and then watch an entire application spring into existence. It felt like wizardry.

It was clear that software development had fundamentally changed. We had entered the era of “vibe coding” — that intoxicating workflow where you lean back in your chair, sip your brew, prompt an LLM with a breezy idea, and casually mash the approval button while the model writes hundreds of lines of code across half a dozen files.

For quick hobby prototypes, weekend experiments, or throwing together an interactive mock-up to impress a client on a Friday afternoon, vibe coding is both effective and brilliant fun. You watch a terminal stream lines of code at lightspeed, and for a fleeting moment you feel like an omnipotent 100x developer. You can now build things in hours that would have once taken a competent coder months.

Then Monday morning arrives.

When you attempt to take your vibe-coded solution anywhere near a real production environment, reality hits you like a cold, wet Monday in London. What felt like intoxicating velocity on Friday turns into an architectural crime scene by Monday. The code might run, but beneath the surface lies a tangled web of subtle bugs, orphaned functions, unverifiable assumptions, and utter nonsense.

You quickly uncover the dirty secret behind unconstrained vibing:

  • Hallucination City: The model confidently invents deprecated API methods and fictitious URLs, synthesises nonexistent package dependencies, and hallucinates framework features that simply don’t exist. Once an agent hallucinates a dependency, every subsequent turn doubles-down on the fiction, inventing bizarre workarounds for libraries that were never real in the first place.
  • Alert & “LGTM” Fatigue: After reading your forty-seventh sprawling diff in two hours, your eyes glaze over. You find yourself mindlessly clicking “Approve” or typing “LGTM” without actually understanding or checking the changes, silently hoping the AI knows what it’s doing. (Spoiler: it doesn’t. You haven’t eliminated manual labour; you have merely swapped creative software engineering for bleary-eyed cursory proofreading.)
  • Solutions That Don’t Do What You Wanted: The code runs without syntax errors, but the actual business logic subtly diverges from what you originally intended. Maybe because you didn’t really know what you wanted in the first place, and you just hope that the model has magically filled in all the blanks in your fluffy request.
  • Unmaintainable, Inscrutable Code: The agent generates dense blocks of code without architectural coherence. It stitches together three different design patterns across four files. You don’t understand how it works and you haven’t really checked it. And you’d probably never reveal the code to your competent peers. “What if they actually look at this?” Your code was tech-debt on day one.
  • No Documentation: Zero README updates, no Architecture Decision Records (ADRs), no API contracts, and no inline explanations of why specific trade-offs were made. The key knowledge that maybe lives in your head is completely absent from the documentation.
  • Insecure & Risky: Hardcoded secrets, wide-open IAM permissions, complete absence of rate limits, and unescaped SQL queries slipped into the codebase because the agent prioritised making the code run over making it safe. Maybe you didn’t like a warning in a log and you told the agent “Make this warning disappear.” The agent said: “Hold my beer.”
  • Inefficient Token Use & The Manual Loop Furnace: You find yourself trapped in a soul-crushing manual loop: copy terminal error → paste into chat → get half-baked fix → copy next error → watch the context window bloat → watch your token bill skyrocket. You aren’t architecting software; you are acting as an organic clipboard between a compiler and an LLM. Inevitably, unless you happen to be paying for a premium model subscription, you’re going to hit your model quota.
  • Lack of Scalability & Portability: The code works on your machine, under your specific node or python version, with zero containerisation or repeatable environment configuration. Hand it to a colleague or push it to CI, and it instantly shatters. Onboarding friction, anyone?
  • Purpose Drift: By turn eight of a messy conversation, the model has completely forgotten the initial rules and constraints you laid out in turn one. (Assuming you specified any constraints at all.) It begins re-inventing components it already built or quietly undoing your earlier bug fixes.

The conclusion is inescapable: you cannot ship vibes to production.

Yet, none of us wants to retreat to typing out thousands of lines of boilerplate by hand. The goal remains what it has always been: to create robust, scalable software solutions without writing code manually.

In an insightful Google Cloud Tech interview with Tilde Thurium, Ryan Lopopolo (Principal Engineer, Agentic Cloud Platform @ Google, who first coined the term “agent harness”) laid out his own radical ambition: to never write code by hand again. As Ryan wryly admitted: “I aspire to be a lazy programmer.” Don’t we all?

Having demonstrated that teams can build and ship massive production systems with zero manual typing by establishing that “humans steer, agents execute”, Lopopolo proved that hands-off development is genuinely achievable.

The answer isn’t to abandon AI agents or wait for newer models; it is to fundamentally change how we construct, control, and constrain them.

The Paradigm Shift: Shifting Developer Interventions Left

To understand why unconstrained vibe coding collapses, we should look at where developers are forced to spend their energy.

In naive AI interactions, developer intervention happens almost entirely on the right. That is: downstream, after the code has already been generated. You spend your afternoon scrolling through hundreds of lines of inscrutable diffs, tweaking chat prompts to fix syntax quirks, or rewriting hallucinations. You are perpetually reacting to failure after the fact, acting as an exhausted organic filter for broken output.

Harness engineering shifts developer intervention to the left.

Instead of manually reviewing broken outputs or wrestling with repetitive errors downstream, you invest your engineering effort upstream before the agent ever runs. You engineer the environment: defining rules, embedding automated linters and type-checkers, configuring deterministic sandboxes, and establishing clear evaluation rubrics. When the agent operates within an engineered environment, failures are caught and corrected autonomously inside the execution loop, long before a human ever inspects the resulting pull request.

Enter Harness Engineering.

The Paradigm Shift

The Bedrock: If It's Not the Model, It's the Harness

To understand what harness engineering is, you must start with this fundamental equation:

Agent=Model+Harness \text{Agent} = \text{Model} + \text{Harness}

And with this comes the ultimate maxim of modern AI development:

“If it’s not the model, it’s the harness.”

The Agent: Model + Harness

We have been conditioned by tech headlines to obsess over frontier models. Every week brings another benchmark, another parameter increase, and another claim of AGI just around the corner. But here is the hard reality: the harness is arguably more important than the model.

An okay, cost-effective model housed within a robust harness will consistently outperform a bleeding-edge frontier model with a poor harness. Why? Because the model is inherently probabilistic, whereas serious software engineering demands determinism. The harness is what bridges that chasm. It supplies the missing structure: living memory, automated validation, static analysis, type checking, sandboxing, and execution guardrails.

Think of it like high-performance motorsport:

The Motorsport Analogy for Agents and the Harness

  • The Model is the Driver: Incredibly talented, capable of split-second reasoning and creative problem-solving, but prone to fatigue, distraction, and occasional moments of reckless overconfidence. Left entirely to their own devices on an open airfield, they might pull off some breathtaking doughnuts, but they can’t win a grand prix.
  • The Entire Race Car IS the Harness: The complete, precision-engineered machine built to translate the driver’s intent into track-winning performance:
  • The Engine & Transmission (Loop Engineering): The internal dynamic execution engine of the harness. It delivers torque to the wheels through a disciplined cycle of intake, combustion, and telemetry. Iterative execution, continuous evaluation, and immediate feedback cycles.
  • The Chassis (Structural Harness): The living rules (AGENTS.md / GEMINI.md), directory layouts, and sandboxed runtimes that keep the car glued to the tarmac, dictating the exact physical limits of where the vehicle can steer.
  • The Brakes & Pit Telemetry (Safety & Evaluation Gates): The deterministic test suites and escalation protocols that prevent the driver from wrapping the machine around a barrier at 180 mph.

You wouldn’t put Lewis Hamilton behind the wheel of a V12 engine block with no chassis or brakes and expect him to finish a single lap at Silverstone, let alone win the race. So why do the same with agentic software development and expect robust results?

⚠️ Dazbo's Battle Scar: The Model Upgrade Fallacy

Whenever an agentic workflow breaks down in development — falling into repetitive loops, introducing subtle runtime bugs, or failing to follow architectural patterns — the instinctive developer reaction is: “We need a smarter model. Let’s switch to the latest 2-trillion parameter frontier model.”

Almost every time, this is a fallacy. Upgrading the model without improving your harness simply means your agent will generate more articulate, harder-to-spot hallucinations at four times the token cost. If your agent is failing, don’t reach for a bigger model first. Ask yourself: Did the harness enforce a test? Did the harness provide the exact schema? Did the harness restate the goal? If it’s not the model, it’s the harness!

Crucially, the harness is model-agnostic. Whether you are routing requests through a cost-effective, high-velocity workhorse like gemini-3.8-flash, or a deep-reasoning powerhouse, a well-engineered harness remains stable, repeatable, and robust. You can hot-swap the model under the hood without having to redesign your entire engineering lifecycle.

The Model is a Commodity; The Harness is the IP

This reveals a fundamental commercial reality of AI software engineering. As my friend and fellow Google Developer Expert Jaroslav Pantsjoha (JP) highlighted in his analysis of the AI-native developer experience: the model is a commodity; the harness is where your true competitive advantage and intellectual property reside.

Every organisation has access to the exact same frontier foundational models via public APIs. Simply renting access to raw model intelligence provides zero lasting competitive advantage. Your true intellectual property and your team’s operational velocity live entirely inside your harness:

  1. Your living rules and architectural guardrails that encode decades of hard-won domain knowledge.
  2. Your curated suite of tool integrations and MCP servers that securely bridge AI agents into your private internal databases and microservices.
  3. Your deterministic evaluation rubrics and automated test harnesses that guarantee safety and compliance before any code reaches production.
  4. Your version-controlled living specifications that turn disposable AI code into maintained, documented, and resilient software assets.

Foundational models will continue to be commoditised, upgraded, and replaced. But a well-architected enterprise harness compounds in value with every single engineering iteration.

Untangling the Triad: Context, Harness, and Loop Engineering

Before we deconstruct the components, it is worth clarifying how three terms frequently heard across the industry map directly back to our bedrock formula (Agent = Model + Harness):

  • Context Engineering: Determines what the agent knows. It is the cognitive payload the harness feeds to the model, including hierarchical rules (AGENTS.md or GEMINI.md that can be global, team, domain, or project-specific), active retrieval (RAG), agent skills, and other sources of knowledge.
  • Structural Harness Engineering: Determines what the agent can touch and where it runs. It provides the deterministic environment: tools, APIs, MCP servers, sandboxed execution runtimes, security guardrails, living file state, and blast-radius (damage control).
  • Loop Engineering: Determines how the agent iterates and converges over time. It is the dynamic runtime engine of the harness ; the autonomous process of observing, acting, judging against rubrics, recording memory, and knowing when to stop.

Crucially, they are not three competing systems; loop and context are the internal functional dimensions of the harness itself. If context is the agent’s briefing pack and the loop is its working cadence, the harness is the complete laboratory workbench in which all of it takes place.

And we can see that state (such as markdown files) plays a significant role in all three.

Context, Harness and Loop Engineering

Deconstructing the Harness: The Deterministic Foundations

If the model represents the non-deterministic, probabilistic core of our system, the harness represents everything else: the deterministic components that guide, restrain, and validate every single action the model takes.

So, what actually lives inside a robust developer harness?

The Foundations of the Developer Harness

Let’s break down the essential layers:

1. Hierarchical Context & Living Rules: "De-Slopping" by Design

Your agent should never have to guess your engineering standards, architectural decisions, or toolchain preferences. In an ad-hoc vibe coding setup, you might type: “Please write clean Python and test it.” The model nods, writes an unformatted script, tests it with a couple of inline assert statements, and calls it a day.

In her analysis How to De-Slop an AI-Generated Codebase, Alice Moore highlights why telling an AI to “clean this up” or “write clean code” inevitably fails: the instruction is too vague. AI “slop” can only be prevented by translating human taste into concrete, non-negotiable rules.

In an engineered harness, we establish our context in files like this:

  • Global Configuration (GEMINI.md or AGENTS.md): Cross-project standards including language preferences, mandatory package managers (e.g. using uv for Python), security baseline rules, and zero-tolerance policies for inline execution hacks.
  • Workspace-Specific Context (my-project/GEMINI.md): Architecture blueprints, directory layouts, data sources, and project goals.
  • Automated Guardrails: Explicit rules that we must always adhere to. For example, performing static analysis after each change. Or, for example, requiring the agent to practice strict Test-Driven Development (TDD). I.e. writing the failing test before implementing code.
  • Documentation Rules: Mandatory upkeep of core documentation (READMEs, Architecture Decision Records, product specs, task lists) using standardised workflows so the codebase never drifts into an undocumented quagmire.

To see what this looks like in practice, here is an authentic snippet from my own global GEMINI.md file:

# Agent Execution Guardrails & Standards

## Python Guidance & Tooling

- **Environment Source of Truth**: Respect the Python version defined in
  `pyproject.toml`. For new greenfield projects, default to `python 3.13`.
- `uv` is the preferred tool for Python package and environment management.
- **Avoid Inline `python -c` Commands**: Do NOT execute inline Python snippets
  using `python -c "..."` or `python3 -c "..."`. Because Antigravity cannot 
  safely prefix-match arbitrary inline code payloads, this triggers a manual 
  approval prompt every time. Instead, write one-off helper logic to a 
  scratch script in your scratch directory and execute it cleanly via 
  `uv run python scratch/helper.py`.
- Always include top-of-module docstrings that explain what a given .py 
  file does, why it exists (i.e. the value it provides) and an outline of 
  how it works.
- Use rich (`pip install rich`) when building console / CLI applications.
- **Zero-Tolerance for Suppression Hacks**: Never bypass linter or 
  type-checker errors using `# noqa`, `# type: ignore`, or `@ts-ignore` 
  comments. Resolve the underlying type contract or architectural root cause.
- When adding features or making changes that are anything beyond trivial, 
  **use TDD**. Create the test, verify it fails, make the change, and then verify 
  the test passes.
- To ensure Python code adheres to required standards, the following 
  commands **must** be run after creating or modifying any `.py` files:
  ```bash
  uvx codespell@latest -s # check spelling and show summary
  uvx ruff@latest check --fix . # perform static checks and auto-fix
  ```

## Agent and AI Development

- NEVER try to downgrade a Gemini model in code or variables. E.g. NEVER 
  try to change `gemini-3.8-flash` to `gemini-3-flash`.
- Unless otherwise instructed, use the latest generally available Gemini 
  models available; always check what the latest is. This also goes for 
  image generation.
- Unless I say otherwise, agents should be built using the Google ADK 
  (`google-adk`) and the Google Gen AI (`google-genai`) packages. 
- AVOID using the `google-generativeai` package. This is deprecated in 
  favour of the newer GenAI `google-genai` SDK.
Enter fullscreen mode Exit fullscreen mode

We’re not politely pleading with the model. We are creating rigid rails.

The rule banning inline python -c snippets is a classic harness engineering battle scar. In a sandboxed developer environment, the execution engine auto-approves safe binary prefixes (such as uv run python ...). When an unconstrained model starts generating dynamic inline code strings, the runtime cannot prefix-match the payload and is forced to interrupt the developer with a manual approval prompt every time it makes a minor tweak to the Python it wants to execute… Every ten seconds!

Similarly, the rule forbidding # type: ignore and # noqa suppressions closes a notorious AI loophole. When confronted with strict linting or typing rules, an unsupervised LLM will frequently take the path of least resistance by slapping suppression comments across the file to silence the compiler. By outlawing suppressions in the harness rules, the model is compelled to fix the actual defect.

2. Living State & Version-Controlled Memory

Human software engineers don’t attempt to hold an entire multi-week project in their brains; we maintain backlog boards, write technical specifications, document decisions, and keep scratchpads.

Yet developers routinely expect an AI model to hold twenty architectural considerations across a forty-turn chat window. The moment the conversation grows long, the context window compacts, and the agent quietly forgets earlier decisions. And let’s not forget the “lost-in-the-middle” problem that we frequently observe in models.

The harness solves this by providing agents with structured, file-backed living state. Although there are no rigid rules for HOW you structure this, here are some suggestions from me:

Strategic & Architectural Truth (The Blueprints):

  • PRODUCT.md: Defines the business “why”, user personas and user stories, core problem statement, and explicit non-goals (what the system is deliberately not building, preventing the agent from over-engineering features you never asked for).
  • SPEC.md: Machine-readable functional requirements, API contracts, and behavioural scenarios (Given / When / Then).
  • ARCHITECTURE.md: System topology, component boundaries, database schemas, data flow and design decisions.
  • DESIGN.md: UI design, styling rules, token definitions (colours, typography, spacing), and accessibility constraints; we can even export or synchronise this file with tools like Google Stitch via MCP.

Operational & Execution State (The Working Memory):

  • TODO.md (or active backlog): Real-time checklists where tasks are marked off as completed.
  • plan.md: Phased vertical slices and dependencies mapped out for the immediate objective.

Transient Execution State (The Scratchpad):

  • scratch/: Ephemeral execution notes, raw compiler traces, and intermediate data files that persist across model turns without polluting Git history. You would exclude this folder in your .gitignore.

(We will explore how to author, stress-test, and execute these living blueprints in depth in Part 3: Spec-Driven Development For-the-Win. Coming soon!)

Because these files are stored in standard Markdown directly in the workspace, they can be version-controlled with Git and shared seamlessly across developer sessions and different machines. If an agent session crashes, if your laptop runs out of battery, or if you hand the task over to a colleague across the globe, the harness preserves the living state. The agent simply reads the task list on turn one and resumes exactly where it left off.

3. Agent Tools: Grounding Models in the Real World

Language models in a vacuum are just predictive text generators; they interact with the outside world through tools. A robust harness equips the agent with a multi-layered execution surface:

  • Native Functions: Strongly typed functions in the code, for fast deterministic operations.
  • Sandboxed Shell and CLIs: Compilers, linters (ruff, eslint), test runners (pytest), package managers (uv, npm), and cloud CLIs (e.g. gcloud) executed inside isolated shell environments.
  • The Model Context Protocol (MCP): A standard that lets agents discover and execute tools without writing bespoke glue code for every service.

⚠️ Dazbo's Battle Scar: Tool Sprawl & Choice Paralysis

When developers first discover MCP, the temptation is to connect twenty different servers simultaneously: GitHub, Slack, Postgres, BigQuery, GKE, Firestore, Stitch, Confluence, Jira, Cloud Logging…

Don’t do it. Connecting 20 MCP servers could easily dump 150+ tool schemas into your model’s context. That’s a lot of wasted tokens before you’ve even typed an instruction. Worse, models experience tool choice paralysis and routing collisions. As I detailed in Dialling Our Agents to 11: My Favourite MCP Servers, keep your active MCP server suite lean, curated, and strictly relevant to the task at hand.

4. On-Demand Context: Progressive Disclosure via Agent Skills

Flooding an LLM’s context window with thousands of lines of reference manuals on every turn is a great way to confuse your model and increase your token costs.

Instead, the harness can use agent skills, which use a mechanism called progressive disclosure to load on-demand. Progressive disclosure operates in three levels:

  • Level 1 (Metadata): At startup, the agent only reads tiny YAML frontmatter headers (e.g. name and trigger description, typically consuming under 100 tokens).
  • Level 2 (Instructions): Only when a relevant task is initiated does the agent dynamically load the full SKILL.md instruction set.
  • Level 3 (Resources): Deep reference scripts, schemas, and templates are only pulled into context when needed.

Your agent stays lightning-fast, burns minimal tokens, and possesses instant access to all the knowledge and operational expertise it needs.

Feels a bit like cheating, doesn’t it?

⚠️ Dazbo's Battle Scar: Skills Sprawl & Context Pollution

If skills are so effective, why not install hundreds of them?

As I explored in Skills Sprawl: When Too Much of a Good Thing Confuses Your AI Agent, skills hoarding is a trap.

Too many skills — It’s a trap!

Even though Level 1 progressive disclosure reads only the YAML frontmatter, loading 150+ skills can easily inject over 15,000 tokens and routing overhead into every single model turn.

The result? Severe LLM decision fatigue, tool selection degradation, and skill routing collisions. If your agent is building a Python cloud backend, it shouldn’t be wading through metadata for React video rendering or mobile UI frameworks. Curate your active skills, prune what you don’t use, and disable non-essential sub-skills in your configuration. A disciplined harness is a clean harness. (For a deep dive into the top skills I rely on daily, check out Dialling Our Agents to 11: Agent Skills You Need to be Using!.)

5. Packaging Knowledge & Capability: The Rules, Skills, and MCP Triad

Once you begin equipping your harness with these deterministic components, an obvious engineering question arises: How do you manage all this?

In practice, we organise with three complementary layers:

  1. Rules (The Guardrails & Standards): Standing instructions that ensure compliance (e.g. safe command syntax, directory layouts, and security policies).
  2. MCP Servers (The Tools & Senses): Execution interfaces that connect the agent to external infrastructure, databases, and APIs.
  3. Agent Skills (The Procedural Playbooks): On-demand workflows and recipes that teach the agent how and when to use those tools effectively.

While open protocols like MCP have standardised the tool layer, there is not yet a single universal industry standard for how the entire capability stack should be packaged.

However, an increasingly popular architectural pattern — supported natively by developer harnesses like Google Antigravity — is the Plugin. Rather than forcing developers to manually wire up an MCP server, write separate prompt rules, and install five standalone skills, a plugin acts as an aggregated delivery vehicle:

The Plugin

A prime example of this pattern in the Google ecosystem is the google-cloud-developer plugin, which looks like this:

  • Packaging Manifest (plugin.json): The schema descriptor defining metadata, versioning, and dependencies.
  • Standing Safety Rules (rules/*.md): Markdown rules (like google-cloud-discovery.md) that establish pre-flight command validation and guardrails.
  • Modular Agent Skills (skills/*/SKILL.md): Individual skill directories with YAML frontmatter headers and procedural Markdown instructions — such as google-cloud-recipe-auth, gcloud, and dynamic discovery (finding-google-skills).
  • Tool Connectivity (mcp.json): A JSON configuration declaring active MCP endpoints — such as the streamable-HTTP connection to Google's official developer-knowledge server.
google-cloud-developer/
├── plugin.json                    # Plugin metadata & agent-plugins.org schema
├── mcp.json                       # MCP server connection endpoints
├── rules/
│   └── google-cloud-discovery.md  # Standing guardrails & catalog rules
└── skills/
    ├── finding-google-skills/
    │   └── SKILL.md               # Dynamic on-demand catalog discovery
    ├── gcloud/
    │   └── SKILL.md               # Gcloud CLI validation & execution
    └── google-cloud-recipe-auth/
        └── SKILL.md               # Authentication recipes
Enter fullscreen mode Exit fullscreen mode

Whether packaged as composite plugins or managed as standalone assets, the goal is identical: providing agents with both the architectural rules and the deep domain knowledge required to execute safely.

My Curated Production Toolkit

In my own daily engineering workflow, I rely on a focused blend of plugins and standalone skills, which include (but are not limited to):

  • google-cloud-developer plugin: Deep GCP architectural rules, safe execution guardrails, and native developer knowledge retrieval.
  • Google Agents CLI (ADK): Coding best practices, scaffolding, and deployment blueprints for building AI agents with the Google Agent Development Kit.
  • agent-platform-eval-flywheel: Google Cloud's evaluation flywheel methodology for measuring agent quality, LLM-as-judge scoring, and regression datasets.
  • spec-driven-development: Bridges intent to code by generating machine-executable specifications and behavioural contracts before coding begins (the focus of Part 3).
  • test-driven-development: Enforces strict red-green-refactor loops, guaranteeing code isn't accepted until automated test suites pass.
  • code-review-and-quality: Automated multi-axis code reviews evaluating maintainability, security, and simplicity before changes enter the main branch.
  • Vercel's find-skills: Dynamic skill discovery, allowing the agent to find and fetch installable skills from public registries on-demand without manual preloading.
  • organise-agent-skills: My own skill for auditing installed skills, pruning token overhead, and managing on-demand progressive disclosure to prevent skills sprawl.
  • maintaining-core-documentation: My own agent skill for synchronising READMEs, architectural docs, and task lists whenever code changes.

6. Sandboxing & Blast-Radius Containment

A robust harness never allows an agent unfettered access to host operating systems or production networks. Through containerised agent sandboxes, filesystem isolation, and execution policies, the harness establishes hard boundaries.

If an agent attempts an unreviewed destructive action — such as dropping a database table, executing an unsandboxed shell command that alters host configuration, or modifying .git/ internal objects — the harness halts execution and demands explicit human verification. Sandboxing transforms autonomous agents from dangerous mavericks into safe, bounded wingmen. (See what I did there?)

Maverick — can be your wingman

7. Observability & Trajectory Monitoring: Beyond Traditional Logging

In traditional software engineering, when a service fails, we inspect a deterministic stack trace and a linear log file: line 42 threw a NullPointerException. Or we run our unit tests and see exactly what test failed.

Autonomous AI agents don’t work like that. An agent doesn’t execute a single static function; it embarks on a multi-step journey toward a goal. In AI engineering, that journey is known as the agent’s trajectory.

What is an Agent Trajectory?

An agent trajectory is the complete, chronological sequence of intermediate states, reasoning steps, tool calls, environment responses, and course corrections the model takes from its initial prompt to the final output.

Think of it as a state machine where every step consists of a continuous route:

Agent Trajectory

For example, a four-turn trajectory might look like:

Turn 1: Read SPEC.md → Turn 2: Run pytest (fails) → Turn 3: Edit code → Turn 4: Re-run test (passes).

When an agent fails, it rarely crashes with a fatal error. Instead, it suffers from trajectory drift:

  • Did it misinterpret a requirement back on turn 2?
  • Did a sandboxed CLI return an ambiguous warning on turn 4 that poisoned the agent’s reasoning on turn 5?
  • Did the model get stuck in a “thrashing loop” between turns 8 and 12, fruitlessly editing the same file back and forth?
  • Did it choose a blunt command-line tool when an exact MCP endpoint was sitting right in its prompt?

Without trajectory observability, you are left staring at a vague final apology from the model, completely blind to where the journey went off the rails.

A robust developer harness treats trajectory monitoring as a first-class citizen:

  1. Step-Level Execution Traces: The harness records structured session transcripts (such as JSONL traces) capturing the exact prompt payload, the model’s internal thinking traces, the precise tool arguments sent, and the raw stdout/stderr returned by the sandbox.
  2. Context & Token Telemetry: Real-time instrumentation tracking how context size inflates turn by turn. This lets you pinpoint runaway token spikes, prompt cache hit rates, and the exact cost of each iteration before a loop drains your budget.
  3. Enterprise OpenTelemetry & Cloud Trace: When moving from local IDE harnesses to cloud-hosted agents (e.g. running Google ADK agents on Cloud Run or the Google Agent Platform), the harness exports distributed traces into Google Cloud Trace and structured events into Cloud Logging. Every tool call becomes an inspectable trace span.

When you can inspect the full trajectory, debugging ceases to be guesswork. You can conduct a forensic post-mortem on any failed run and immediately determine: Did the model experience a reasoning failure, or did my harness fail to provide the right deterministic rail?

Summary & What's Next: Firing Up the Engine

We have shifted left and engineered the deterministic foundations: living rules, version-controlled state, curated tools, progressive skill disclosure, plugin packaging, sandboxed boundaries, and execution traces. We have built the race car chassis.

Now, how does work actually get done?

In Part 2, we fire up the dynamic runtime engine inside the harness: Loop Engineering. We will explore:

  • Why autonomous feedback loops crush single-pass generation.
  • The core mechanical components of an engineered developer loop.
  • The Ratchet Principle.
  • The Diagnostic Rule of Thumb: Is it a Loop bug or a Harness bug?
  • Designing Evaluation Rubrics (combining traditional tests with LLM goal evaluation and deterministic tool trajectory verification).
  • A complete, turn-by-turn trace of an autonomous agent self-correcting in real time.
  • The loopy skill, and the Loop Library.

👉 Continue to Part 2: Loop Engineering in the Harness: Autonomous, Self-Correcting Systems (Coming Soon)

See you there!

Before You Go

  • Please share 📢 this with anyone that you think will be interested. It might help them, and it really helps me!
  • Please give me loads of claps! 👏 (Just hold down the clap button.)
  • Please leave a comment 💬. Interaction is good!
  • Add a star ⭐ on my repos!
  • Follow 👉 and subscribe 🔔, so you don’t miss my content.

References and Useful Links

The Harness & Ecosystem

Dazbo's Related Articles

Top comments (0)