DEV Community

Cover image for How to build AI agent for stock analysis with FastAPI
Ayush Kumar
Ayush Kumar

Posted on • Originally published at logiclooptech.dev

How to build AI agent for stock analysis with FastAPI

I’ll show you how to build an AI agent for stock analysis that runs end-to-end in production. The recipe is simple: pull reliable market data, stitch together a LangChain (or CrewAI) workflow, let an LLM write the analysis, add memory for portfolio context, backtest the whole thing, and finally serve it behind FastAPI. Below is a practical walk-through with the code I use in my own trading bots, plus the pitfalls that kept me up at night.


Which market data APIs should I use to build an AI agent for stock analysis?

The first thing that breaks most prototypes is bad data. You need a source that offers near-real-time quotes, historical OHLCV, and a stable free tier if you’re experimenting.

Provider Real-time? Historical depth Free tier Typical cost
Alpha Vantage Delayed (15 min) 20 years 5 req/min $50/mo for premium
IEX Cloud Real-time (US) 30 years 500 req/day $9/mo for paid tier
Yahoo Finance (yfinance) Delayed 25 years Free (scrape) N/A
Polygon.io Real-time + crypto 30 years 5 req/min $199/mo

For a production-grade agent I combine a cheap real-time feed (IEX) with a free historical dump (yfinance). The real-time feed drives the decision loop; the historical data fuels backtesting and feature engineering.

import yfinance as yf
import requests
import os
from datetime import datetime, timedelta

IEX_TOKEN = os.getenv("IEX_TOKEN")
BASE_URL = "https://cloud.iexapis.com/stable"

def get_realtime_price(symbol: str) -> float:
    url = f"{BASE_URL}/stock/{symbol}/quote?token={IEX_TOKEN}"
    resp = requests.get(url, timeout=5)
    resp.raise_for_status()
    return resp.json()["latestPrice"]

def get_historical(symbol: str, start: str, end: str):
    df = yf.download(symbol, start=start, end=end, progress=False)
    df.reset_index(inplace=True)
    return df
Enter fullscreen mode Exit fullscreen mode

What can go wrong?

Rate limits – IEX will throttle you after a few hundred calls. Cache the last price in Redis and only hit the API when the cache expires (e.g., 60 seconds).

Data mismatches – IEX and yfinance use different ticker conventions; always normalise symbols before you store them.


How do I design the agent architecture with LangChain or CrewAI?

Both LangChain and CrewAI give you a composable graph of LLM calls, tools, and memory. I started with LangChain because its tooling around agents is mature, then migrated a few modules to CrewAI for better parallel tool execution.

A minimal architecture looks like this:

  1. Input layer – FastAPI endpoint receives a ticker list.
  2. Orchestrator – LangChain AgentExecutor decides whether to call a price tool, a news summariser, or the analysis chain.
  3. Tool layer – Python functions wrapped as Tools (price fetch, technical indicator, sentiment).
  4. LLM chain – Prompt template that asks the model to generate a “buy / hold / sell” recommendation with confidence.
  5. Memory – Redis vector store that stores the last N decisions and portfolio exposure.
from langchain.agents import initialize_agent, Tool
from langchain.llms import OpenAI
from langchain.prompts import PromptTemplate
import redis
import json

# Redis for short‑term memory
redis_client = redis.Redis(host="redis", port=6379, db=0)

def price_tool(symbol: str) -> str:
    price = get_realtime_price(symbol)
    return f"The current price of {symbol} is ${price:.2f}"

def sentiment_tool(symbol: str) -> str:
    # Placeholder: call a news API and summarise with LLM
    return f"Sentiment for {symbol} is neutral."

price = Tool(
    name="price",
    func=price_tool,
    description="Get the latest market price for a ticker."
)

sentiment = Tool(
    name="sentiment",
    func=sentiment_tool,
    description="Return a short sentiment summary for a ticker."
)

prompt = PromptTemplate(
    input_variables=["price", "sentiment", "history"],
    template=(
        "You are a quantitative analyst. Given the price: {price}, "
        "sentiment: {sentiment}, and the last 5 decisions: {history}, "
        "output a JSON with fields: decision (buy/hold/sell), confidence (0‑1), "
        "and a one‑sentence rationale."
    )
)

llm = OpenAI(model="gpt-4o-mini", temperature=0.2)
agent = initialize_agent(
    tools=[price, sentiment],
    llm=llm,
    agent_type="zero-shot-react-description",
    verbose=True,
)

def run_analysis(symbol: str):
    # Pull recent decisions from Redis
    past = redis_client.lrange(f"history:{symbol}", -5, -1)
    past_json = [json.loads(x) for x in past] if past else []
    history_str = ", ".join(d["decision"] for d in past_json)

    # Let the agent decide which tools to call
    result = agent.run(
        {
            "price": price_tool(symbol),
            "sentiment": sentiment_tool(symbol),
            "history": history_str,
        }
    )
    decision = json.loads(result)
    # Store decision for future memory
    redis_client.rpush(f"history:{symbol}", json.dumps(decision))
    redis_client.ltrim(f"history:{symbol}", -20, -1)  # keep last 20
    return decision
Enter fullscreen mode Exit fullscreen mode

When not to use this?

If you only need a static statistical model (e.g., a mean-reversion rule), the overhead of an LLM agent is wasted CPU and cost. Also, LangChain’s AgentExecutor adds latency (~300 ms) per call – not ideal for high-frequency strategies.

Failure modes

Tool hallucination: The LLM may request a tool that isn’t registered, causing a runtime error. Keep the verbose flag on while debugging.

State drift: Forgetting to prune Redis history leads to stale context influencing new decisions.


How can I add memory and tool use for portfolio management?

A stock-analysis agent is only useful if it knows what you already own. I store two kinds of memory:

  • Short-term – recent recommendations and positions (Redis).
  • Long-term – aggregated performance metrics (PostgreSQL).

The short-term store is already in the code above. For portfolio exposure I expose a tiny FastAPI route that reads the Redis cache and writes to a PostgreSQL table every hour.

# fastapi_app.py
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
import asyncpg
import aioredis

app = FastAPI()
redis = aioredis.from_url("redis://redis")
pg_pool = None

class Order(BaseModel):
    symbol: str
    action: str  # buy or sell
    qty: int

@app.on_event("startup")
async def startup():
    global pg_pool
    pg_pool = await asyncpg.create_pool(dsn=os.getenv("DATABASE_URL"))

@app.post("/order")
async def place_order(order: Order):
    # naive simulation of a broker call
    if order.action not in {"buy", "sell"}:
        raise HTTPException(400, "Invalid action")
    # Record in Postgres
    async with pg_pool.acquire() as conn:
        await conn.execute(
            "INSERT INTO trades (symbol, action, qty, ts) VALUES ($1,$2,$3,now())",
            order.symbol, order.action, order.qty,
        )
    # Update memory
    await redis.rpush(f"history:{order.symbol}", json.dumps({
        "decision": order.action,
        "confidence": 1.0,
        "timestamp": datetime.utcnow().isoformat()
    }))
    return {"status": "ok"}
Enter fullscreen mode Exit fullscreen mode

The agent can now query its own position before recommending a trade. Add a simple tool:

def position_tool(symbol: str) -> str:
    # Summarise current net position from Postgres
    async def fetch():
        async with pg_pool.acquire() as conn:
            row = await conn.fetchrow(
                "SELECT SUM(CASE WHEN action='buy' THEN qty ELSE -qty END) AS net "
                "FROM trades WHERE symbol=$1", symbol
            )
            return row["net"] or 0
    net = asyncio.run(fetch())
    return f"Current net position for {symbol} is {net} shares."
Enter fullscreen mode Exit fullscreen mode

Add it to the tools list and the LLM will automatically request it when needed.

Trade-offs

Consistency: If the PostgreSQL write fails, the Redis cache and DB diverge. Wrap both in a transaction or use an outbox pattern.

Cost: Each LLM call costs a few cents; adding extra tools multiplies the number of calls per analysis. Keep prompts tight.


How do I backtest and deploy the stock analysis agent?

Backtesting proves whether the agent’s logic adds value before you risk real capital. I use backtrader because it integrates nicely with pandas data frames from yfinance.

import backtrader as bt
import pandas as pd

class LLMStrategy(bt.Strategy):
    params = dict(symbol="AAPL", cash=100_000)

    def __init__(self):
        self.dataclose = self.datas[0].close
        self.order = None

    def next(self):
        if self.order:
            return  # wait for order to fill

        decision = run_analysis(self.p.symbol)  # sync call for demo
        if decision["decision"] == "buy" and not self.position:
            self.order = self.buy(size=10)
        elif decision["decision"] == "sell" and self.position:
            self.order = self.sell(size=self.position.size)

# Load data
df = get_historical("AAPL", "2020-01-01", "2023-01-01")
data = bt.feeds.PandasData(dataname=df)

cerebro = bt.Cerebro()
cerebro.addstrategy(LLMStrategy, symbol="AAPL")
cerebro.adddata(data)
cerebro.broker.setcash(100_000)
cerebro.run()
print(f"Final portfolio value: ${cerebro.broker.getvalue():,.2f}")
Enter fullscreen mode Exit fullscreen mode

Production notes

  • Speed – The backtest calls the LLM for every bar, which is far too slow. Replace the LLM with a deterministic rule derived from the backtest results (e.g., “if confidence > 0.8 and price < MA50, buy”).
  • Cost – Running a full-year backtest for 10 symbols can cost $30 in API usage. Cache the LLM outputs per day.
  • Deployment – Containerise the FastAPI service with Docker, use uvicorn --workers 4 behind an Nginx reverse proxy, and schedule a Celery beat job that triggers run_analysis nightly.
# Dockerfile
FROM python:3.11-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY . .
EXPOSE 8000
CMD ["uvicorn", "fastapi_app:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "4"]
Enter fullscreen mode Exit fullscreen mode

Deploy to a Kubernetes cluster with a HorizontalPodAutoscaler that scales on CPU. Keep the Redis and Postgres pods in the same namespace for low latency.

When to skip this pipeline – If you only need a one-off signal for a single ticker, the overhead of Docker, Kubernetes, and backtesting is overkill. A simple script run locally is enough.


FAQ

What’s the cheapest way to get real-time market data for a hobby project?

Use the free tier of IEX Cloud (500 req/day) combined with yfinance for historical data. Cache results aggressively to stay inside the limit.

Can I replace the LLM with a rule-based system?

Yes. If latency or cost is a concern, train a lightweight classifier on the LLM’s historical outputs and use that model in production.

How often should I retrain or update the prompt?

At least monthly, or whenever you notice a drift in confidence scores. Prompt engineering is an ongoing maintenance task.

Do I need a GPU for this stack?

Not for inference with OpenAI’s hosted models. If you run a local model (e.g., Llama-3) you’ll need a GPU, but that adds hardware cost and complexity.


Key Takeaways

  • Choose a reliable market-data API and cache aggressively; rate limits bite fast.
  • LangChain (or CrewAI) gives you a clean tool-based agent, but each extra tool adds latency and cost.
  • Store short-term decisions in Redis and long-term performance in Postgres to keep the agent context-aware.
  • Backtest with a deterministic surrogate of the LLM to avoid prohibitive API bills.
  • Deploy the whole stack with FastAPI, Docker, and Kubernetes; monitor Redis, Postgres, and LLM usage to catch failures early.

Top comments (0)