DEV Community

Cover image for Decoding LLMs: Building AI Agents & Multi-Step Workflows (Part 4)
Shahibur Rahman
Shahibur Rahman

Posted on

Decoding LLMs: Building AI Agents & Multi-Step Workflows (Part 4)

Welcome back to our journey into understanding Large Language Models! Moving from single API prompts to building AI agents shifts software design from static request-response calls to dynamic execution loops. Instead of generating a single completion and stopping, an agent evaluates intermediate results, invokes local functions or external APIs, updates its memory state, and repeats until a task condition is met. If you missed Part 1: Decoding LLMs: How Large Language Models Work - Fundamentals for Beginners, Part 2: Decoding LLMs: Architecture, Training, and Practical Integration, or Part 3: Decoding LLMs: Prompt Engineering, API Integration, and Structured Outputs, catch up on tokenization, KV caching, and Pydantic validation before diving into agentic mechanics.

In this fourth installment, we explore the core mechanics of agentic systems: constructing control loops, managing conversational state memory, and implementing self-correcting tool calls using Python.


Core Architectural Patterns for Building AI Agents

An AI Agent is a software system that uses an LLM as its central reasoning engine. Instead of generating a single output block, the agent evaluates state, decides which external tool to trigger, parses the returned observation, and loops until it satisfies a target goal.

+-------------------------------------------------------------------+
|                         AGENT STATE ENGINE                        |
|  +-------------------+    +------------------+    +------------+  |
|  | System Guidelines |    | Context History  |    | Tool Schemas| |
|  +-------------------+    +------------------+    +------------+  |
+----------------------------------|--------------------------------+
                                   |
                                   v
                    +------------------------------+
                    |        LLM Reasoning         |
                    |   (Decide Action / Answer)   |
                    +--------------|---------------+
                                   |
               +-------------------+-------------------+
               |                                       |
       [Tool Call Triggered]                   [Final Output]
               |                                       |
               v                                       v
     +-------------------+                       (( Return ))
     |   Execute Tool    |
     | (API/Database/OS) |
     +---------|---------+
               |
               v
     +-------------------+
     | Append Result to  |
     | Context & Loop    |
     +-------------------+
Enter fullscreen mode Exit fullscreen mode

1. The ReAct Pattern (Reasoning + Acting)

The ReAct pattern orchestrates multi-step decision-making by interlacing reasoning with tool invocation:

  1. Thought: The model evaluates current context and plans its next discrete step.
  2. Action: The model emits a structured tool call (e.g., function name and JSON arguments).
  3. Observation: The application executes the function and appends the result back into the prompt context.
  4. Loop: The cycle repeats until the model produces a final answer without triggering tools.

Key Takeaway: Running complex workflows through iterative observation-action cycles prevents error compounding compared to asking an LLM to solve a multi-step task in a single generation step.

2. Plan-and-Execute Workflows

When tasks involve high step counts, a basic ReAct loop can get stuck in repetitive action loops if an execution step fails. A Plan-and-Execute framework separates planning from execution:

  • Planner: Breaks down a complex user goal into an explicit sequential task list upfront.
  • Executor: Runs through the planned tasks individually using tools.
  • Re-Planner: Assesses intermediate observations against the master plan and updates remaining steps if unexpected runtime outputs occur.

Managing Memory and State when Building AI Agents

LLM API calls are stateless. To keep track of progress over multiple tool calls or user turns, your backend application must explicitly persist conversation history while avoiding context window overflow.

           FULL HISTORY                      TRUNCATED WINDOW
 +---------------------------------+     +-----------------------+
 | System: Diagnostic guidelines   | --> | System Instructions   |
 | User: Inspect database load     |     +-----------------------+
 | Assistant: CPU load is 95%      |     | ... (Older truncated) |
 | User: Restart service           |     +-----------------------+
 | Assistant: Service restarted    |     | User: Restart service |
 | User: What was the load earlier?| --> | Assistant: Restarted  |
 +---------------------------------+     | User: What was load?  |
                                         +-----------------------+
Enter fullscreen mode Exit fullscreen mode

Context Management Strategies

  1. Sliding Window Memory: Retains only the last N message turns. It keeps context overhead small but risks dropping critical instructions provided in initial turns.
  2. Summarized Memory: Uses a fast secondary model to condense older context turns into a concise running summary stored alongside system rules.
  3. State Ingestion: Maintains key variable states in an external database (e.g., PostgreSQL or Redis) and re-injects current state context explicitly on every invocation turn.

Practical Python Implementation: Autonomous Infrastructure Agent

Let's construct a lightweight, framework-agnostic AI agent using Python, OpenAI function calling, and Pydantic v2 validation. The agent checks system health metrics, attempts automated service restarts, and logs tickets if repairs fail.

1. Define Tools and Pydantic Schemas

import json
from typing import Any
from pydantic import BaseModel, Field

# Validation Schemas
class MetricQueryInput(BaseModel):
    node_id: str = Field(description="Target server node ID, e.g. 'node-prod-01'")

class RestartServiceInput(BaseModel):
    node_id: str = Field(description="Target server node ID")
    service_name: str = Field(description="Service to restart, e.g. 'nginx' or 'redis'")

class CreateTicketInput(BaseModel):
    node_id: str = Field(description="Target server node ID")
    description: str = Field(description="Summary of the infrastructure issue")
    severity: str = Field(description="Severity level: Low, Medium, High, or Critical")

# Tool Implementation Functions
def get_node_metrics(node_id: str) -> str:
    """Queries node health parameters."""
    mock_db = {
        "node-prod-01": {"cpu": "96%", "memory": "88%", "failed_service": "nginx"},
        "node-prod-02": {"cpu": "15%", "memory": "22%", "failed_service": None}
    }
    result = mock_db.get(node_id, {"error": f"Node '{node_id}' not found."})
    return json.dumps(result)

def restart_node_service(node_id: str, service_name: str) -> str:
    """Executes a remote service restart command."""
    if node_id == "node-prod-01" and service_name == "nginx":
        return json.dumps({"status": "success", "message": f"Service '{service_name}' on '{node_id}' restarted."})
    return json.dumps({"status": "failed", "message": f"Unable to restart '{service_name}' on '{node_id}'."})

def create_support_ticket(node_id: str, description: str, severity: str) -> str:
    """Creates an incident ticket in the tracking system."""
    return json.dumps({"ticket_id": "INC-40921", "status": "Assigned to On-Call Ops"})

# Centralized Tool Registry
TOOL_REGISTRY: dict[str, dict[str, Any]] = {
    "get_node_metrics": {
        "func": get_node_metrics,
        "schema": MetricQueryInput,
        "desc": "Query real-time health parameters and failing services for a server node."
    },
    "restart_node_service": {
        "func": restart_node_service,
        "schema": RestartServiceInput,
        "desc": "Restart a named system service on a target server node."
    },
    "create_support_ticket": {
        "func": create_support_ticket,
        "schema": CreateTicketInput,
        "desc": "Open an incident ticket when automated recovery steps fail."
    }
}
Enter fullscreen mode Exit fullscreen mode

2. Autonomous Agent Execution Loop

This control loop manages state memory, executes tools, and reflects runtime error details back into context for automatic correction:

import os
from openai import OpenAI
from pydantic import ValidationError

class DiagnosticAgent:
    def __init__(self, model: str = "gpt-4o-mini", max_steps: int = 5):
        self.client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
        self.model = model
        self.max_steps = max_steps

    def execute(self, prompt: str) -> str:
        messages: list[dict[str, Any]] = [
            {
                "role": "system",
                "content": (
                    "You are an automated infrastructure diagnostic agent. "
                    "Investigate server alerts using tools. If a service has failed, try to restart it. "
                    "If restarts fail, create a support ticket. Verify tool responses before concluding."
                )
            },
            {"role": "user", "content": prompt}
        ]

        # Prepare tool specifications for API payload
        tools_spec = [
            {
                "type": "function",
                "function": {
                    "name": name,
                    "description": details["desc"],
                    "parameters": details["schema"].model_json_schema()
                }
            }
            for name, details in TOOL_REGISTRY.items()
        ]

        # Main Control Loop
        for step in range(self.max_steps):
            print(f"\n--- Step {step + 1}/{self.max_steps}: LLM Reasoning Cycle ---")

            response = self.client.chat.completions.create(
                model=self.model,
                messages=messages,
                tools=tools_spec,
                temperature=0.1
            )

            msg = response.choices[0].message
            messages.append(msg)

            # Termination Check: Return text response when no tool is invoked
            if not msg.tool_calls:
                print("Agent completed task.")
                return msg.content or "Task complete."

            # Process Tool Invocations
            for tool_call in msg.tool_calls:
                fn_name = tool_call.function.name
                call_id = tool_call.id
                raw_args = tool_call.function.arguments

                print(f"Tool Selected: {fn_name} | Args: {raw_args}")

                if fn_name not in TOOL_REGISTRY:
                    observation = json.dumps({"error": f"Tool '{fn_name}' is not registered."})
                else:
                    tool_info = TOOL_REGISTRY[fn_name]
                    try:
                        # Enforce strict type checking with Pydantic
                        validated = tool_info["schema"].model_validate_json(raw_args)
                        observation = tool_info<a href="**validated.model_dump(">"func"</a>)
                    except ValidationError as val_err:
                        # Reflection mechanism: Send validation errors back into message history
                        observation = json.dumps({"error": f"Argument Schema Error: {val_err.errors()}"})
                    except Exception as err:
                        observation = json.dumps({"error": f"Tool Execution Error: {str(err)}"})

                print(f"Observation: {observation}")

                # Append observation to message context
                messages.append({
                    "role": "tool",
                    "tool_call_id": call_id,
                    "content": observation
                })

        return "Execution stopped: Maximum step count reached without resolution."

if __name__ == "__main__":
    agent = DiagnosticAgent()
    user_alert = "Alert: High error rates on server node node-prod-01. Check metrics and resolve active faults."
    summary = agent.execute(user_alert)
    print(f"\nFinal Result:\n{summary}")
Enter fullscreen mode Exit fullscreen mode

Production Design Principles

  • Set Iteration Thresholds: Always set a firm max_steps threshold in control loops to guard against continuous execution loops and excessive token utilization.
  • Enforce Strict Argument Parsing: Validate raw model JSON payloads against typed schemas (e.g., Pydantic) before passing arguments into internal system functions.
  • Implement Error Reflection: Return clear error details in observation messages. Models often adjust parameter structures on the next turn when provided descriptive failure messages.
  • Preserve Role Hierarchy: Maintain clear message roles (system, user, assistant, tool) to ensure predictable API response formatting.

Architectural Considerations and Next Steps

When designing agentic execution loops in production, managing state persistence and handling unexpected function failures are critical factors. Are you building agent loops using custom control flows like the Python implementation above, or relying on specialized orchestration frameworks like LangGraph, AutoGen, or CrewAI?

Share your architecture approaches and production experiences in the comments below!

References & Resources

GitHub logo openai / openai-python

The official Python library for the OpenAI API

OpenAI Python API library

PyPI version

The OpenAI Python library provides convenient access to the OpenAI REST API from any Python 3.10+ application. The library includes type definitions for all request params and response fields and offers both synchronous and asynchronous clients powered by HTTPX2.

It is generated from our OpenAPI specification.

Documentation

The REST API documentation can be found on platform.openai.com. The full API of this library can be found in api.md.

Installation

# install from PyPI
pip install openai
Enter fullscreen mode Exit fullscreen mode

Usage

The full API of this library can be found in api.md.

The primary API for interacting with OpenAI models is the Responses API. You can generate text from the model with the code below.

import os
from openai import OpenAI
client = OpenAI(
    # This is the default and can be omitted
    api_key=os.environ.get("OPENAI_API_KEY"),
)

…
Enter fullscreen mode Exit fullscreen mode

Top comments (0)