DEV Community

heavy shield
heavy shield

Posted on

Your API's Newest Users Are Agents: Designing for Non-Human Clients

Introduction

Your API's newest users are agents. They don't read docs, they don't browse dashboards, and they don't file support tickets. They parse OpenAPI specs, call endpoints in loops, and expect deterministic, machine-readable responses. If your API was designed for humans clicking buttons, it's already failing this new class of client.

In this post, I'll walk through a concrete example: building a small "tool API" that an LLM-based agent can call. We'll cover the problem, a solution, and a runnable Python implementation. No hand-waving about "agentic workflows" — just code you can run and a loop with explicit termination conditions.

Problem

Human-facing APIs optimize for discoverability and forgiveness. Agents optimize for determinism and low token cost. Three specific failure modes show up when agents hit a human-designed API:

  1. Ambiguous errors. A 400 with {"error": "bad request"} forces the agent to guess. The agent will retry, hallucinate a fix, or give up.
  2. Unbounded responses. Returning 50KB of nested JSON burns context window and increases the chance the agent truncates the useful part.
  3. No idempotency. Agents retry on timeouts. Without idempotency keys, you get duplicate side effects.

Here's a minimal example of the kind of handler that causes these problems:

# bad_handler.py
from flask import Flask, request, jsonify

app = Flask(__name__)

@app.route("/create_ticket", methods=["POST"])
def create_ticket_bad():
    data = request.get_json(silent=True) or {}
    if "title" not in data:
        return jsonify({"error": "bad request"}), 400
    # ... create ticket, no idempotency, returns everything
    return jsonify({"ticket": {"id": 1, "title": data["title"], "history": [...]}})
Enter fullscreen mode Exit fullscreen mode

An agent calling this has no way to distinguish "missing field" from "malformed JSON" from "server error," and no safe retry path.

Solution

Design the API for a client that is literal, stateless between calls, and token-budgeted. Concretely:

  • Structured error codes. Every error response includes a stable code string the agent can branch on.
  • Idempotency keys. Accept an Idempotency-Key header; store the result keyed by it.
  • Bounded payloads. Return only fields the agent needs; paginate lists with explicit cursors.
  • A tool schema. Publish a JSON Schema your agent can load to construct calls. This is OpenAI-function-calling-style, not tied to any specific vendor.

Implementation

We'll build a small in-memory ticket API with the properties above, then write an agent loop that uses it. The agent loop is deliberately simple: it's a plan-act-observe loop with a hard step limit and a success predicate. No framework required.

The API

# agent_api.py
import json
import uuid
from dataclasses import dataclass, field, asdict
from typing import Optional

from flask import Flask, request, jsonify

app = Flask(__name__)

@dataclass
class Ticket:
    id: str
    title: str
    status: str = "open"

tickets: dict[str, Ticket] = {}
idempotency_store: dict[str, dict] = {}

TOOL_SCHEMA = {
    "name": "create_ticket",
    "description": "Create a support ticket. Idempotent on Idempotency-Key header.",
    "parameters": {
        "type": "object",
        "properties": {
            "title": {"type": "string", "minLength": 1, "maxLength": 200},
        },
        "required": ["title"],
        "additionalProperties": False,
    },
}

@app.route("/tools", methods=["GET"])
def list_tools():
    return jsonify({"tools": [TOOL_SCHEMA]})

@app.route("/create_ticket", methods=["POST"])
def create_ticket():
    key = request.headers.get("Idempotency-Key")
    if not key:
        return jsonify({"code": "missing_idempotency_key",
                        "message": "Provide Idempotency-Key header."}), 400

    if key in idempotency_store:
        return jsonify(idempotency_store[key]), 200

    data = request.get_json(silent=True)
    if not isinstance(data, dict):
        return jsonify({"code": "invalid_json",
                        "message": "Body must be a JSON object."}), 400

    title = data.get("title")
    if not isinstance(title, str) or not title.strip():
        return jsonify({"code": "invalid_title",
                        "message": "'title' must be a non-empty string."}), 422

    ticket = Ticket(id=str(uuid.uuid4()), title=title.strip())
    tickets[ticket.id] = ticket
    payload = {"ticket": asdict(ticket)}
    idempotency_store[key] = payload
    return jsonify(payload), 201

@app.route("/tickets/<ticket_id>", methods=["GET"])
def get_ticket(ticket_id: str):
    ticket = tickets.get(ticket_id)
    if ticket is None:
        return jsonify({"code": "not_found",
                        "message": f"No ticket {ticket_id}."}), 404
    return jsonify({"ticket": asdict(ticket)})
Enter fullscreen mode Exit fullscreen mode

Key details: every error has a stable code; the response body is small; idempotency is enforced via header, not body, so retries are safe.

The Agent Loop

The agent loop below is deliberately framework-free. It defines explicit termination conditions: it stops when the goal is met, when the model returns no tool call, or when max_steps is reached. I'm using an OpenAI-chat-completions-style interface here as a stand-in; swap in whichever client you use.

# agent_loop.py
import json
import uuid
from typing import Any

import requests

API = "http://localhost:5000"
MAX_STEPS = 6


def call_model(messages: list[dict]) -> dict:
    """Return a message dict. Replace with your model client."""
    # Placeholder: in production, call your LLM here.
    raise NotImplementedError("Wire up your model client.")


def call_tool(name: str, args: dict) -> dict:
    if name == "create_ticket":
        r = requests.post(
            f"{API}/create_ticket",
            json=args,
            headers={"Idempotency-Key": str(uuid.uuid4())},
            timeout=10,
        )
        return {"status": r.status_code, "body": r.json()}
    return {"status": 400, "body": {"code": "unknown_tool", "message": name}}


def run_agent(goal: str) -> dict:
    tools = requests.get(f"{API}/tools", timeout=10).json()["tools"]
    messages = [
        {"role": "system", "content": "You are an agent. Use tools to satisfy the goal."},
        {"role": "user", "content": goal},
    ]

    for step in range(MAX_STEPS):
        msg = call_model(messages)
        messages.append(msg)

        # Termination condition 1: model produced a final answer.
        if not msg.get("tool_calls"):
            return {"status": "done", "steps": step, "answer": msg.get("content")}

        for call in msg["tool_calls"]:
            name = call["function"]["name"]
            args = json.loads(call["function"]["arguments"])
            result = call_tool(name, args)
            messages.append({
                "role": "tool",
                "tool_call_id": call["id"],
                "content": json.dumps(result),
            })

    # Termination condition 2: step budget exhausted.
    return {"status": "max_steps_reached", "steps": MAX_STEPS}
Enter fullscreen mode Exit fullscreen mode

The loop has exactly two exit paths, both explicit. There's no "keep trying until it works" branch, which is where most agent bugs live.

Security Notes

If you extend this to execute code or shell commands, treat the tool boundary as a trust boundary. Never pass model output to eval, exec, or subprocess with shell=True. If you must run generated code, isolate it in a sandbox (a container with no network, a read-only filesystem, and a hard timeout) and validate arguments against the JSON Schema before execution. The example above avoids this entirely by only allowing a single typed tool with a bounded string parameter.

Key Takeaways

  • Agents are literal clients. Stable error code fields beat prose messages.
  • Idempotency keys belong in headers and should be required for any write endpoint an agent can retry.
  • Publish a JSON Schema for your tools so the agent can construct calls without guessing.
  • Define agent loops with explicit termination conditions — a step limit and a success predicate. "Loop until done" is not a termination condition.
  • Keep payloads small. Every byte you return is a token the agent pays for.
  • If you add code-execution tools, sandbox them. eval on model output is a remote code execution vulnerability with extra steps.

The code above is intentionally minimal so you can fork it and add your own tools. Start by auditing your existing API for the three failure modes in the Problem section — that's where the real work is.

Top comments (0)