DEV Community

Anshul Rajpal
Anshul Rajpal

Posted on

Claude Code vs Cursor vs Codex: I Tested All Three on the Same Task

The AI coding agent space has gone from "novelty" to "daily driver" fast. Three tools — Claude Code (Anthropic), Cursor (the IDE), and OpenAI Codex CLI — all claim to save you hours. But do they actually deliver when the task gets messy?

I ran the same non-trivial feature through all three: building a URL shortener service with a REST API, Redis caching, and Docker deployment. Same repo, same constraints, same evaluation criteria. Here's what happened.

The Setup

Task: Create a Flask/FastAPI-based URL shortener with:

  • POST /shorten accepting JSON {url, custom_alias?}
  • GET /:alias redirecting with 301
  • Redis-backed cache for hot URLs
  • Docker Compose setup
  • Basic test suite

Environment: Mac M2, 16GB RAM, latest versions of each tool as of this week.

Metrics I tracked:

  • Time to first working endpoint
  • Tokens/API calls consumed
  • Code quality (did it need manual fixes?)
  • Setup friction (config, auth, context window)

What Actually Happened

Claude Code (Anthropic) won on code quality. The initial implementation was clean, type-hinted, and the Docker Compose file just worked. It asked clarifying questions when the prompt was ambiguous — which slowed it down by ~10 minutes but saved rework later.

# Example output from Claude Code — note the structured error handling
@app.post("/shorten")
def shorten_url(request: ShortenRequest, db: Session = Depends(get_db)):
    alias = request.custom_alias or generate_short_code()
    if db.query(URLMapping).filter_by(alias=alias).first():
        raise HTTPException(status_code=409, detail="Alias taken")
    # ... rest of implementation
Enter fullscreen mode Exit fullscreen mode

Cursor was fastest to first commit. Its inline acceptance model meant I could review line-by-line, which felt more controlled. But it hallucinated a Redis connection pool config that silently failed — I caught it only because I read the logs.

Codex CLI (OpenAI) was the cheapest per task (~$0.03 in API costs vs ~$0.12 for Claude). It produced working code but needed the most manual patching. The Dockerfile had a multi-stage build that broke on arm64 — classic architecture mismatch.

The Numbers

Tool Time to Working Estimated Cost Manual Fixes Context Window
Claude Code 42 min ~$0.15 2 minor ~180k tokens
Cursor 28 min ~$0.08 3 (including Redis bug) ~128k tokens
Codex CLI 35 min ~$0.03 5 (Docker, types) ~128k tokens

Where Each One Struggled

Claude Code choked on the Docker Compose networking config — kept suggesting host.docker.internal which doesn't work on Linux. Cursor over-engineered the caching layer with an unnecessary LRU decorator. Codex couldn't handle the test suite structure and wrote pytest tests that didn't actually assert anything useful.

Who Should Use What

  • Claude Code if you want fewer bugs in the first pass and can afford the slower iteration.
  • Cursor if you like staying in the IDE and reviewing changes inline.
  • Codex CLI if you're cost-sensitive and comfortable doing substantial post-editing.

None of them replaced me. All three reduced the raw typing, but the architectural decisions, the "does this actually work in production?" thinking — that's still on you.

AI coding agents comparison diagram

The honest takeaway: these tools are best treated as senior juniors — fast, mostly competent, but needing review. The real productivity gain isn't from letting them run wild; it's from using them to skip the boring parts while you stay in the driver's seat for the hard decisions.

If you've run a similar test, I'd be curious what task you threw at them — the URL shortener is boring, and the results might shift completely with something more ambiguous like "refactor this legacy codebase." What's your experience been?

Top comments (0)