DEV Community

Cover image for ai agent python code example for FastAPI and OpenAI SDK
Ayush Kumar
Ayush Kumar

Posted on Originally published at logiclooptech.dev

ai agent python code example for FastAPI and OpenAI SDK

I’ve been running AI agents in production for over a year now. Most tutorials skip the hard parts - state, rate limits, and what happens when your agent calls a tool that times out. Here’s how I build them with FastAPI and the OpenAI SDK, the way I’d want it documented when I’m paged at 2 a.m.

What is a basic ai agent python code example with FastAPI and OpenAI SDK?

Start with a minimal FastAPI app that accepts a prompt, calls OpenAI, and returns a response. No tools, no state - just the core loop. This is your foundation.

# main.py
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
import openai
import os

app = FastAPI()
client = openai.OpenAI(api_key=os.getenv("OPENAI_API_KEY"))

class AgentRequest(BaseModel):
    prompt: str
    model: str = "gpt-4o"

class AgentResponse(BaseModel):
    response: str

@app.post("/agent", response_model=AgentResponse)
async def run_agent(request: AgentRequest):
    try:
        completion = client.chat.completions.create(
            model=request.model,
            messages=[{"role": "user", "content": request.prompt}]
        )
        return AgentResponse(response=completion.choices[0].message.content)
    except openai.RateLimitError:
        raise HTTPException(status_code=429, detail="Rate limit exceeded")
    except Exception as e:
        raise HTTPException(status_code=500, detail=str(e))
Enter fullscreen mode Exit fullscreen mode

This works for simple Q&A. But real agents need to do more than chat - they need to act.

How do you integrate tool use and function calling in python agents?

Tool use turns your agent from a talker into a doer. I define functions the agent can call - like querying a database or hitting an internal API - and let the OpenAI SDK handle the function calling loop.

# tools.py
from typing import Optional
import httpx
import os

async def get_user_info(user_id: str) -> dict:
    async with httpx.AsyncClient() as client:
        resp = await client.get(f"{os.getenv('INTERNAL_API_URL')}/users/{user_id}")
        resp.raise_for_status()
        return resp.json()

# agent_tools.py
from openai import OpenAI
from typing import List, Dict, Any
import json

client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))

TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "get_user_info",
            "description": "Fetch user details by ID",
            "parameters": {
                "type": "object",
                "properties": {
                    "user_id": {"type": "string", "description": "The user ID"}
                },
                "required": ["user_id"]
            }
        }
    }
]

async def run_agent_with_tools(prompt: str, model: str = "gpt-4o") -> str:
    messages = [{"role": "user", "content": prompt}]

    while True:
        completion = client.chat.completions.create(
            model=model,
            messages=messages,
            tools=TOOLS,
            tool_choice="auto"
        )

        message = completion.choices[0].message

        if message.tool_calls:
            # Add the assistant's message with tool calls
            messages.append(message)

            for tool_call in message.tool_calls:
                if tool_call.function.name == "get_user_info":
                    args = json.loads(tool_call.function.arguments)
                    result = await get_user_info(args["user_id"])
                    messages.append({
                        "tool_call_id": tool_call.id,
                        "role": "tool",
                        "content": json.dumps(result)
                    })
            # Continue the loop to let the agent process the tool result
            continue
        else:
            # Final response
            return message.content
Enter fullscreen mode Exit fullscreen mode

I’ve seen teams try to manage this loop manually with state machines. Don’t. Let the SDK handle the tool call/response cycle - it’s less buggy and easier to debug.

How do you handle async workflows and state management in ai agents?

Production agents aren’t stateless. They need memory across turns - chat history, user preferences, cached tool results. I use Redis for fast state and Pydantic models to keep it typed.

# state.py
import redis.asyncio as redis
from pydantic import BaseModel
from typing import Optional, Dict, Any
import json

redis_client = redis.from_url(os.getenv("REDIS_URL"), decode_responses=True)

class AgentState(BaseModel):
    session_id: str
    chat_history: List[Dict[str, str]] = []
    user_preferences: Dict[str, Any] = {}
    last_tool_result: Optional[Dict] = None

async def get_state(session_id: str) -> AgentState:
    data = await redis_client.get(f"agent_state:{session_id}")
    if data:
        return AgentState.model_validate_json(data)
    return AgentState(session_id=session_id)

async def save_state(state: AgentState):
    await redis_client.set(
        f"agent_state:{state.session_id}",
        state.model_dump_json(),
        ex=3600  # 1 hour TTL
    )
Enter fullscreen mode Exit fullscreen mode

Each request loads state, runs the agent loop (which may mutate state via tools), then saves it back. I’ve had agents corrupt state by writing concurrently - use Redis transactions or a queue if you need strong consistency.

How do you deploy ai agents to production with docker and cloud run?

Containerize it. Cloud Run scales to zero, which saves money when traffic is spiky. But cold starts hurt - keep your image small and avoid heavy imports at module level.

# Dockerfile
FROM python:3.11-slim

WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

COPY . .

ENV PORT=8080
EXPOSE 8080

CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8080"]
Enter fullscreen mode Exit fullscreen mode

I deploy with:

gcloud builds submit --tag gcr.io/my-project/ai-agent
gcloud run deploy ai-agent --image gcr.io/my-project/ai-agent --platform managed
Enter fullscreen mode Exit fullscreen mode

Watch for:

  • Large model SDKs (like langchain) bloating your image - stick to openai>=1.0.0 for minimal deps.
  • Secrets leaking into layers - use --secret in Cloud Build or Cloud Run env vars.
  • No graceful shutdown - add a SIGTERM handler to flush Redis buffers.

What are common issues in ai agent code and how do you debug them?

Rate limits are the most frequent pager. I’ve seen agents get stuck in retry loops because they didn’t back off. Token limits sneak up when chat history grows. Context overflow makes agents forget instructions.

Here’s how I defend against them:

# middleware.py
from fastapi import Request, Response
from starlette.middleware.base import BaseHTTPMiddleware
import time
from collections import defaultdict

request_counts = defaultdict(list)

class RateLimitMiddleware(BaseHTTPMiddleware):
    async def dispatch(self, request: Request, call_next):
        client_ip = request.client.host
        now = time.time()

        # Clean old requests
        request_counts[client_ip] = [t for t in request_counts[client_ip] if now - t < 60]

        if len(request_counts[client_ip]) >= 10:  # 10 req/min
            return Response("Rate limit exceeded", status_code=429)

        request_counts[client_ip].append(now)
        return await call_next(request)
Enter fullscreen mode Exit fullscreen mode

For token limits, I truncate history:

def truncate_history(messages, max_tokens=3000):
    # Rough estimate: 4 chars per token
    total_chars = sum(len(m["content"]) for m in messages)
    if total_chars > max_tokens * 4:
        # Keep system message and recent turns
        return [messages[0]] + messages[-(max_tokens//4):]  # Simplified
    return messages
Enter fullscreen mode Exit fullscreen mode

I log token usage per request:

# In your agent endpoint
usage = completion.usage
logger.info(f"Token usage: {usage.total_tokens} (prompt: {usage.prompt_tokens}, completion: {usage.completion_tokens})")
Enter fullscreen mode Exit fullscreen mode

If you see completion tokens near the limit, your agent is looping or over-explaining.

How do you evaluate agent performance with custom metrics and logging?

I track three things: success rate, latency, and tool accuracy. Success rate means the agent completed the user’s goal without human intervention. I log every turn and use a simple evaluator LLM to judge outcomes.

# metrics.py
from prometheus_client import Counter, Histogram
import time

AGENT_REQUESTS = Counter('agent_requests_total', 'Total agent requests', ['status'])
AGENT_LATENCY = Histogram('agent_latency_seconds', 'Time spent processing agent requests')
TOOL_USAGE = Counter('agent_tool_usage_total', 'Tool usage count', ['tool_name'])

async def run_agent_evaluated(prompt: str, session_id: str):
    start = time.time()
    state = await get_state(session_id)

    try:
        response = await run_agent_with_tools(prompt)
        # Simple success heuristic: no error and tool used if needed
        success = "error" not in response.lower()
        AGENT_REQUESTS.labels(status="success" if success else "failure").inc()

        # Log for human review
        logger.info({
            "session_id": session_id,
            "prompt": prompt,
            "response": response,
            "state": state.model_dump(),
            "success": success
        })

        return response
    finally:
        AGENT_LATENCY.observe(time.time() - start)
Enter fullscreen mode Exit fullscreen mode

I’ve used this to catch agents that call tools unnecessarily - wasting money and latency. If your tool usage counter is high but success rate low, your agent is confused about when to act.

FAQ

What’s the smallest viable ai agent python code example?

A FastAPI endpoint that calls OpenAI chat.completions.create with a user prompt and returns the message content. Add error handling for rate limits and you’ve got a baseline.

How do you prevent agent loops when using tool use?

Limit the number of tool call iterations per request (e.g., max 5 turns). The OpenAI SDK doesn’t do this by default - wrap your agent loop in a counter and break if exceeded.

Should I use LangChain or the OpenAI SDK directly?

For production agents needing fine-grained control over tool calls, state, and logging, use the OpenAI SDK directly. LangChain adds abstraction that obscures failure modes and makes debugging harder.

How much does running an AI agent cost per request?

At $0.005 per 1K tokens for GPT-4o, a typical agent turn (500 prompt + 150 completion tokens) costs ~$0.003. Tool calls add latency but no extra token cost - watch for iteration bloat.

Key Takeaways

  • Start simple: get the basic OpenAI loop working before adding tools or state.
  • Use the OpenAI SDK’s built-in tool calling - don’t reinvent the message loop.
  • Manage state externally (Redis) and serialize it with Pydantic for safety.
  • Deploy small containers to Cloud Run; monitor cold starts and secret handling.
  • Guard against rate limits, token limits, and context overflow with middleware and truncation.
  • Log token usage, tool calls, and outcomes - eval success rate, not just latency.
  • Avoid frameworks that hide the agent loop; transparency saves debugging time.

Top comments (0)