DEV Community

Cover image for A Grok 4.7 agent in Python: allowlisted tools, a bounded loop, and fallback you can audit
Olivia Hayes
Olivia Hayes

Posted on Originally published at cometapi.com

A Grok 4.7 agent in Python: allowlisted tools, a bounded loop, and fallback you can audit

Most "agents" I ship are a while loop with an LLM call inside. That's fine. The hard part isn't the loop, it's the policy around it: which functions the model is allowed to trigger, which failures justify trying another model, and where the loop stops.

This is the version I run for a support-style agent on Grok 4.7 (grok-4.7), with GPT, Claude, Gemini and DeepSeek sitting behind the same route policy as fallbacks. A unified multi-model endpoint like CometAPI keeps the connection code from metastasizing across providers; everything else below is application-owned.

The control flow

user text -> model response -> validate tool call -> run allowlisted function
         -> append tool result -> model response
Enter fullscreen mode Exit fullscreen mode

Five parts, each independently testable:

  1. One OpenAI client on a shared base URL.
  2. Grok 4.7 as the primary model.
  3. A tool registry. The model proposes a call; only my code executes it.
  4. A bounded loop with a fixed turn cap.
  5. An ordered fallback list that advances only on classified transient failures.

The model never holds database credentials and never runs Python. It emits a structured request ("call get_order_status with this id"), and my code does the authorization, argument parsing, execution and serialization. Fallback models inherit the same tool boundary (not a wider one), and any tool result containing external content gets handled as untrusted data.

Reasoning state on multi-turn Grok 4.7

Grok 4.7 accepts low, medium, high or xhigh reasoning effort, defaulting to high. On xAI's Responses API every response includes reasoning.encrypted_content, and a client-managed multi-turn loop should pass the returned reasoning items back unchanged in the next request. Long loops can use context compaction: preserve the returned compaction item as opaque state and append new turns after it. Both are stateful, provider-specific fields, so verify the route you select returns them end to end before making them a production dependency.

Step 1: point the SDK at one base URL

pip install openai
Enter fullscreen mode Exit fullscreen mode
export COMETAPI_KEY="your-cometapi-key"
export PRIMARY_MODEL="grok-4.7"
export FALLBACK_MODEL_1="your-compatible-gpt-model-id"
export FALLBACK_MODEL_2="your-compatible-claude-model-id"
export FALLBACK_MODEL_3="your-compatible-gemini-model-id"
export FALLBACK_MODEL_4="your-compatible-deepseek-model-id"
Enter fullscreen mode Exit fullscreen mode

I use Chat Completions here because assistant tool_calls and matching tool result messages map directly onto a compact loop I can read in one screen. For longer stateful loops, evaluate /v1/responses instead; both routes are documented for the model.

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["COMETAPI_KEY"],
    base_url="https://api.cometapi.com/v1",
    max_retries=0,
    timeout=30.0,
)
Enter fullscreen mode Exit fullscreen mode

max_retries=0 and the explicit timeout are deliberate. The application classifies failures and decides whether to repeat a call or move on. Hidden SDK retries make latency, duplicated side effects and fallback behaviour much harder to reason about.

And do not paste model IDs from a blog post into production. Pull GET /api/models during deploy or startup, then confirm capabilities in the model directory.

Step 2: narrow, read-only tools

Start with reads. Writes come later, behind idempotency keys and human confirmation.

import json

TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "get_order_status",
            "description": "Read the current status of one order.",
            "parameters": {
                "type": "object",
                "properties": {
                    "order_id": {"type": "string"}
                },
                "required": ["order_id"],
                "additionalProperties": False,
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "check_inventory",
            "description": "Read available inventory for one SKU.",
            "parameters": {
                "type": "object",
                "properties": {
                    "sku": {"type": "string"}
                },
                "required": ["sku"],
                "additionalProperties": False,
            },
        },
    },
]

def get_order_status(order_id: str) -> dict:
    # Replace this demo with an authenticated, read-only service call.
    return {"order_id": order_id, "status": "in_transit"}

def check_inventory(sku: str) -> dict:
    # Replace this demo with an authenticated, read-only service call.
    return {"sku": sku, "available_units": 12}

TOOL_REGISTRY = {
    "get_order_status": get_order_status,
    "check_inventory": check_inventory,
}
Enter fullscreen mode Exit fullscreen mode

A JSON schema shapes the request. It is not authorization. Validate argument length and format, confirm the caller may see that order or SKU, and cap the size of every tool result before it goes back into the transcript.

Step 3: fallback that only fires when it should

Fallback exists to survive a flaky route, not to paper over malformed requests. The official guidance: advance to the next configured route on connection errors, timeouts, HTTP 408, HTTP 429, and temporary 5xx. Invalid credentials, unsupported parameters and invalid requests fail immediately.

from openai import APIConnectionError, APIStatusError, APITimeoutError

def configured_models() -> list[str]:
    names = [
        os.getenv("PRIMARY_MODEL", "grok-4.7"),
        os.getenv("FALLBACK_MODEL_1"),
        os.getenv("FALLBACK_MODEL_2"),
        os.getenv("FALLBACK_MODEL_3"),
        os.getenv("FALLBACK_MODEL_4"),
    ]
    return [name for name in names if name]

def is_retryable(error: Exception) -> bool:
    if isinstance(error, (APIConnectionError, APITimeoutError)):
        return True
    if isinstance(error, APIStatusError):
        return error.status_code in {408, 429} or error.status_code >= 500
    return False

def complete_with_fallback(messages: list[dict], tools: list[dict]):
    models = configured_models()
    last_error = None

    for index, model in enumerate(models):
        try:
            response = client.chat.completions.create(
                model=model,
                messages=messages,
                tools=tools,
                tool_choice="auto",
            )
            return response, model
        except Exception as error:
            last_error = error
            final_route = index == len(models) - 1
            if final_route or not is_retryable(error):
                raise

    raise RuntimeError("No configured model completed the request") from last_error
Enter fullscreen mode Exit fullscreen mode

That model list is configuration, not a quality ranking. Pick fallbacks that support the same message roles, tool schema, input modality, context requirement and response behaviour this agent needs. Log the chosen route and the failure that triggered every transition.

Step 4: the bounded loop

def execute_tool_call(tool_call) -> str:
    name = tool_call.function.name

    if name not in TOOL_REGISTRY:
        return json.dumps({"error": f"Tool not allowed: {name}"})

    try:
        arguments = json.loads(tool_call.function.arguments)
        result = TOOL_REGISTRY[name](**arguments)
        return json.dumps(result)
    except (json.JSONDecodeError, TypeError, ValueError) as error:
        return json.dumps({"error": f"Invalid tool arguments: {error}"})

def run_agent(user_text: str, max_turns: int = 4) -> dict:
    messages = [
        {
            "role": "system",
            "content": (
                "You are a support agent. Use tools only when needed. "
                "Never invent order or inventory data."
            ),
        },
        {"role": "user", "content": user_text},
    ]
    route_log = []

    for turn in range(max_turns):
        response, model = complete_with_fallback(messages, TOOLS)
        route_log.append({"turn": turn + 1, "model": model})

        assistant = response.choices[0].message
        messages.append(assistant.model_dump(exclude_none=True))

        if not assistant.tool_calls:
            return {
                "answer": assistant.content,
                "routes": route_log,
                "usage": response.usage.model_dump() if response.usage else None,
            }

        for tool_call in assistant.tool_calls:
            messages.append(
                {
                    "role": "tool",
                    "tool_call_id": tool_call.id,
                    "content": execute_tool_call(tool_call),
                }
            )

    raise RuntimeError("Agent stopped after reaching max_turns")

result = run_agent("Where is order A-104, and is SKU BLUE-42 in stock?")
print(result["answer"])
print(result["routes"])
Enter fullscreen mode Exit fullscreen mode

Multiple tool calls in one response work because every returned call gets a matching tool_call_id result appended. The moment a tool changes state (email, order, refund), add an idempotency key and a confirmation step, and never blindly re-run a turn after a timeout if the side effect may already have landed.

One transport does not mean interchangeable models

A shared base URL removes connection-layer duplication. It does not make GPT, Claude, Gemini and DeepSeek drop-in equivalents. Before a model earns a slot in the chain, check:

  • the model ID comes back from the catalog;
  • the route supports your tool schema and message roles;
  • argument shapes and parallel-call behaviour match the agent contract;
  • context window and input modalities fit the request;
  • the response can be validated before a user sees it;
  • latency and cost stay inside budget.

Provider-native features usually need a native endpoint or a separate adapter. Keep those exceptions explicit instead of forcing every capability through the common path.

Fallback is not multi-agent

Fallback picks another model after a route failure. A multi-agent system assigns different responsibilities to different agents: planner, researcher, reviewer. Different problems.

If you grow this into multi-agent, give each worker a narrow role, its own tool allowlist, a bounded budget and a structured handoff. Don't let every agent call every tool or forward an unbounded transcript. Ship one agent until your evaluation data shows role separation is worth the overhead.

Guardrails I don't skip

Validate before execution. Tool names, argument schemas, tenant ownership, permissions, rate limits. Tool descriptions are guidance for the model, never a security control.

Split read from write. Reads can usually run automatically once authorized. Writes need stronger checks, idempotency and confirmation.

Bound everything. Max turns, max tool calls, wall-clock time, prompt size, token budget. On breach, return a controlled error or escalate.

Log the decision trail. Task, policy version, selected model, fallback reason, tool name, tool latency, validation result, token usage, final status. No secrets, no unnecessary customer content.

Write contract tests instead of assuming. Same fixtures against every configured model: a normal answer, one tool call, multiple tool calls, malformed arguments, an unknown tool, a tool timeout, a primary-model 429, and an invalid API key that must not trigger fallback.

Deployment checklist

  • Fetch current model IDs and verify the Grok 4.7 route before deploying.
  • Keep the API key in a secret manager, never in source or prompts.
  • Start read-only, with explicit JSON schemas.
  • Authenticate and authorize before each tool call.
  • Allow fallback only for classified transient errors.
  • Run every fallback through the same tool-calling contract.
  • Add idempotency and confirmation before enabling writes.
  • Set loop, latency, context and cost limits.
  • Measure task success, not just API availability.

FAQ

Can Grok 4.7 call my Python functions directly?
No. It returns structured function-call requests. Your app parses, validates, executes an allowlisted function and sends back the result.

Should every error move to another model?
No. Connection failures, timeouts, 408, 429 and temporary 5xx qualify. Bad requests, auth failures and unsupported parameters are bugs to fix.

Can I reuse one tool schema across every model?
Only after testing. Shared transport does not guarantee matching argument quality, parallel-call behaviour or schema enforcement. Contract tests decide.

Is a fallback chain a multi-agent system?
No. Fallback swaps the model for a request after a route failure. Multi-agent assigns different tasks to separate agents. Separate layers, separate tests.

Sources


Originally published at cometapi.com

Top comments (0)