DEV Community

Cover image for The AI Engineer Roadmap: From "What's a Token?" to Shipping Real AI Products
Gustavo Maple
Gustavo Maple

Posted on

The AI Engineer Roadmap: From "What's a Token?" to Shipping Real AI Products

So you opened X, saw 47 new AI tools launched before breakfast, and now you feel like everyone got a memo you didn't. Good news: there is no memo. Just a mountain of hype, a few buzzwords used wrong on purpose, and a real set of skills you can learn in a sane order.

I put this roadmap together for devs like me who already ship code and want to go from "I use ChatGPT for regex" to "I build AI features people actually use." Each phase builds on the last, and wherever real money is involved I added a cost estimate, because nothing ruins a side project like a surprise API bill at 3 AM.

Grab your coffee (or a token-efficient energy drink) and let's dive in.

TL;DR

  • AI Engineering is mostly integration and reliability, not training models from scratch.
  • The path: LLM basics → prompting → APIs → RAG → agents → evals and security → production.
  • You can do almost all of it for $0 to ~$15 USD with free tiers, local models, and cheap API models.
  • Build as you learn: there's a 5-tier project ladder at the end, from "simple API app" to "agents building agents."
  • Tools change every month, concepts don't. Learn the concepts first and treat tools as examples.

What Does an AI Engineer Actually Do?

Before we start learning things, let's answer the obvious question: what's the job? Spoiler: it's less "train a neural network in a dark room" and more "make a model behave inside a real product, reliably, without bankrupting the company."

Day to day, an AI Engineer usually does some mix of these:

  1. Work with AI APIs. Call models from code, handle streaming, errors, rate limits, and keep an eye on token costs.
  2. Design effective prompts. Write instructions that get consistent results, not just one lucky answer.
  3. Build RAG pipelines. Give a model access to your own data (docs, tickets, databases) so it stops making things up.
  4. Create agents. Let a model use tools, make decisions, and take actions instead of just chatting.
  5. Write evaluation frameworks. Build the test suites that tell you if a prompt, model, or pipeline is actually good.
  6. Fine-tune models. Adapt a model to a specific task or style when prompting and RAG aren't enough.
  7. Evaluate continuously. Measure quality in production, catch regressions, and compare models when a new one drops (which is every other Tuesday).

Notice what's not on the list: inventing new model architectures. That's research. This roadmap is about the engineering side, and if you can already build backends and frontends, you're closer than you think.

Phase 0 — The Modern AI Toolbox (Vocabulary You'll Keep Hearing)

Before the fundamentals, a quick vocabulary tour. These are terms you'll see in every AI conversation, tweet, and job description. You don't need to master them yet, just know what they are so you don't nod along like you understand (we've all done it).

  • Skills: Reusable instructions that tell an AI how to do a specific kind of task better.
  • llms.txt: A file that lets AI models read what your website is about in a direct, machine-friendly way.
  • AGENTS.md: A context file an AI agent loads every time it starts, so it always knows the rules of your project. If you've seen CLAUDE.md, it's the same idea.
  • MCP (Model Context Protocol): A standard that lets an AI connect to the services you use every day (Google Drive, Gmail, Notion...) and take actions in them.
  • Commands: Reusable shortcuts you trigger on demand to run a predefined workflow with your AI tool.

And a few buzzwords you'll bump into later. We'll come back to them, but for now, just recognize the names:

  • Harness Engineering
  • Specs-Driven Development
  • Forward Deployed Engineer

Think of Phase 0 as the map legend: you don't need to memorize it, but the rest of the trip is easier when you know what the symbols mean.

Phase 1 — LLM Fundamentals

Goal: understand how modern AI systems actually work under the hood, so you stop treating them like magic (or like a very confident intern).

  • [ ] Generative AI: AI that creates new content (text, images, code, audio) instead of just classifying or predicting.
  • [ ] Large Language Models (LLMs): Models trained on huge amounts of text to predict what comes next, which turns out to be surprisingly powerful.
  • [ ] Fine-Tuning: Further training a model on your own data to adapt its behavior or style.
  • [ ] Tokens: The chunks of text a model reads and writes. They're also how you get billed.
  • [ ] Context: Everything the model can "see" when it generates a response.
  • [ ] Memory: How information is carried across conversations, since models don't remember anything by default.
  • [ ] Context Window: The maximum amount of tokens a model can handle at once.
  • [ ] Hallucinations: When a model confidently makes things up. Very confident. Zero shame.
  • [ ] Temperature: A setting that controls how predictable or creative the output is.
  • [ ] Embeddings: Numeric representations of text that capture meaning, so you can compare things by similarity.
  • [ ] System Prompts: Hidden instructions that define how the model should behave.
  • [ ] RAG: Giving a model external knowledge at query time. We'll get to it in Phase 4.

Try it yourself: you don't need to pay anything to play with these ideas. Tools like Ollama let you run open models locally on your machine, and Hugging Face is the place to browse and discover thousands of open models.

Don't try to memorize all of this in one sitting. Come back to this list as the later phases make each concept click.

Phase 2 — Prompt Engineering (and Context Engineering)

This is the foundation for everything that comes next, and most devs underestimate it. "It's just writing text to a chatbot," they say, right before spending three days debugging a prompt that works 70% of the time.

  • [ ] Core techniques: few-shot prompting (show examples), chain-of-thought (ask the model to reason step by step), and role prompting (tell it who it is).
  • [ ] Structured outputs: JSON mode and other ways to force a specific output format, so your code can actually parse what comes back.
  • [ ] System prompts vs. user prompts: what goes where, and why the order matters.
  • [ ] Measurable prompt evaluation: test your prompts with real criteria, not just "looks good to me."
  • [ ] Context Engineering: deciding what goes into the model's context (instructions, examples, retrieved data, tool results, history), not just how you word the prompt.

A good mental shift: prompt engineering is about how you ask, while context engineering is about what the model gets to see. As your apps grow, the second one starts to matter more than the first.

Quick rule of thumb: if you change a prompt and can't tell whether it got better or worse, you don't have a prompt problem yet, you have an evaluation problem. We'll fix that in Phase 6.

Phase 3 — LLM APIs and SDKs

Time to stop chatting in a UI and start calling models from code. This is the phase where AI stops being a toy and becomes just another service in your stack (one that occasionally answers in Shakespearean English when you asked for JSON).

  • [ ] Anthropic API and OpenAI API: authentication, streaming responses, error handling, and rate limits.
  • [ ] Token costs: how to estimate them before you ship, and how to keep them under control.
  • [ ] Function calling / Tool use: letting the model request actions in your code. This is the foundation of agents.
  • [ ] Structured outputs with schemas: define the shape of the response with Pydantic (Python) or Zod (TypeScript) and validate what comes back.

💰 Estimated cost: ~$5-15 USD in API credits if you use cheap models like Haiku or GPT-4o-mini.
🆓 Free option: run open models locally with Ollama, or use the free tiers listed in the "Free APIs and Resources" section at the end of this post.

Suggested exercise: build a tiny script that takes a messy text (an email, a support ticket, a review) and returns a validated JSON object with the fields you care about. It covers auth, structured outputs, error handling, and cost tracking in one afternoon.

Phase 4 — RAG (Retrieval-Augmented Generation)

How to give an LLM "memory" or external knowledge without fine-tuning. Instead of retraining the model on your data, you fetch the relevant pieces at query time and hand them over along with the question. It's the difference between making someone memorize the whole library and letting them use the index.

  • [ ] Embeddings: what they are (no math required) and how to generate them through an API.
  • [ ] Vector databases: the core concepts, and how similarity search finds "things that mean something similar."
  • [ ] Chunking strategies: why the way you split your documents matters more than you'd expect.
  • [ ] The full pipeline: ingest → chunk → embed → store → retrieve → generate.

Stack options:

  • [ ] pgvector: a Postgres extension. If you already use Postgres, it's $0 extra.
  • [ ] 100% free alternative: ChromaDB or Qdrant, self-hosted and running locally.
  • [ ] Embeddings: text-embedding-3-small (cheap) or Ollama (free, runs locally).

💰 Estimated cost: close to $0 with the local stack, and only cents with text-embedding-3-small for a small project.

Suggested exercise: build a "chat with your docs" app over a folder of Markdown files. When it answers wrong (and it will), figure out whether the problem was retrieval or generation. That debugging habit is worth more than any tutorial.

Phase 5 — Agents and Orchestration

This is where an LLM stops answering and starts acting. It reads a goal, picks a tool, looks at the result, and decides what to do next. It's also where the hype is loudest, so let's keep our feet on the ground.

  • [ ] Real agent vs. marketing hype: a chatbot with a fancy name isn't an agent. Learn to tell the difference.
  • [ ] Agent Tools: the functions and services an agent can call (search, databases, APIs, your file system). This builds directly on the function calling you learned in Phase 3.
  • [ ] Patterns: ReAct, planning, multi-step tool use, and loops with guardrails.
  • [ ] Error handling: infinite loops, runaway costs, and what happens when a tool fails halfway through.

Agents you can use today (as examples):

  • Coding agents: Claude Code and Cursor. Use them, but also read how they work.
  • Local vs. remote agents: some run on your machine, others run in the cloud. Knowing the tradeoffs (privacy, cost, power) matters.
  • Open-source and experimental agents: OpenClaw and Hermes, if you want to poke at something less polished.
  • Cline: an agent that lives inside your editor.

Pick ONE framework so you don't scatter yourself:

  • [ ] LangGraph (Python): the most used in industry for complex agents.
  • [ ] Vercel AI SDK (TypeScript): if you'd rather stay in a web stack.

A quick word on the LangChain ecosystem. Think of building a factory:

Tool Role in the factory
LangChain The basic parts and tools
LangGraph The blueprint and the complex machinery
LangSmith The quality inspector and monitor

LangChain is great for linear flows (A → B → C) like simple RAG pipelines. LangGraph adds loops, branches, and shared state, which is what agents need to self-correct. We'll meet LangSmith in the next phase.

A simple rule: if your flow is a straight line, LangChain is enough. If your app needs to remember steps, loop, or make decisions, reach for LangGraph.

Suggested exercise: build an agent with two or three tools, a hard limit on iterations, and a cost cap. Then try to break it on purpose.

Phase 6 — Evaluation, Observability, and Security

This is what separates a prototype from something production-ready. A demo that works on your laptop is nice. A system that keeps working when real users show up (and try to break it, whether on purpose or by accident) is the actual job.

  • [ ] Systematic output evaluation: build a set of test cases and measure quality with real criteria, so you can tell whether a change made things better or worse.
  • [ ] Logging and tracing: record every LLM call (inputs, outputs, latency, tokens) so that when something breaks, you can see exactly where.
  • [ ] Prompt injection: how attackers hide instructions in user input or retrieved content to hijack your model, and how to mitigate it. This is critical if your app takes dynamic input.
  • [ ] Cost management in production: caching, and falling back to cheaper models when the task allows it.

Tools you can use (as examples):

  • [ ] Langfuse: open source, with a free cloud tier, for tracing.
  • [ ] Promptfoo: free and open source, for automated evaluation.
  • [ ] LangSmith: a platform to debug, evaluate, and monitor LLM apps. It's framework-agnostic, so it works with LangChain, LangGraph, or your own custom code. Its multi-turn evals let you judge whole conversations instead of single responses.

When does observability become non-negotiable? The moment you move from prototype to production. If you're building for real users, you need visibility into what your app is doing, and tracing is the fastest way to get it.

Suggested exercise: take the agent or RAG app you built earlier, write 15-20 test cases (including a few adversarial ones, like prompt injection attempts), and run them automatically every time you change a prompt.

Phase 7 — Integrating AI into Your Own Product

This is where everything comes together. You've got the fundamentals, prompts, APIs, RAG, agents, and evals. Now you plug it all into a real product, with real users, real latency, and real bills.

  • [ ] Structure your backend for LLM calls: keep model calls in one well-defined layer (prompts, retries, timeouts, cost tracking) instead of scattering them across your codebase.
  • [ ] Stream responses to your frontend in real time: nobody wants to stare at a spinner for 15 seconds. Show the answer as it's generated.
  • [ ] Smart caching: don't pay twice for the same (or a very similar) request, especially for repeated recommendations or common questions.
  • [ ] Fallback strategy: what happens if your main model fails or gets slow? Plan for a backup model, a graceful degraded mode, or a clear error message.

A quick checklist before you ship:

  • Do you know what one request costs, and what happens if usage 10x's overnight?
  • Are your calls traced, so you can debug a bad answer from last Tuesday?
  • Have you tested for prompt injection with real user-style input?
  • Is there a timeout and a fallback for when the model is down?

If you can answer yes to all four, congrats: you're not just playing with AI anymore, you're shipping it.

Suggested exercise: take a feature from a project you already have and add an AI-powered version of it behind a feature flag. Ship it to yourself first, watch the traces and costs for a week, then decide if it's worth opening up.

Project Ideas by Tier

Reading is great, but you learn AI engineering by building (and by watching your first agent confidently do the wrong thing). Here's a ladder of projects. Each tier maps roughly to the phases above, so you can climb as you learn.

Tier 1: Language model APIs with a simple interface

Build small apps that call an LLM API and wrap it in a basic UI or CLI. (Phases 1-3)

  1. AI Receipt Parser
    • Turn messy receipt text or photos into clean, validated JSON with vendor, date, items, and totals.
  2. AI Commit Message Writer
    • Read a git diff and suggest a clear commit message in your team's style.
  3. AI Tone Rewriter
    • Rewrite the same email as friendly, formal, or "please stop CC'ing the whole company."

Tier 2: RAG

Give the model access to your own data and make it answer with sources. (Phase 4)

  1. AI Codebase Q&A
    • Ask questions about a repository ("where do we handle auth?") and get answers that point to the relevant files.
  2. AI Personal Knowledge Search
    • Search across your notes, bookmarks, and PDFs by meaning instead of keywords.
  3. AI Support Answer Suggester
    • Draft replies to support tickets using your docs and past resolved tickets, with citations the agent can verify.

Tier 3: AI Agents

Models that use tools, make decisions, and loop until the job is done. (Phase 5)

  1. AI Research Assistant
    • Take a question, search multiple sources, cross-check them, and produce a short report with references.
  2. AI Bug Triage Agent
    • Read new issues, reproduce or classify them, label them, and suggest the likely area of the code.
  3. AI Personal Finance Organizer
    • Categorize transactions, flag unusual charges, and explain monthly spending patterns in plain language.

Tier 4: Enterprise-scale AI systems

Multi-user, observable, secure, and cost-controlled. Where prototypes go to grow up. (Phases 6-7)

  1. AI Gateway
    • A central layer that routes requests across models, caches results, enforces rate limits and budgets per team, and logs everything.
  2. AI Contract Review Pipeline
    • Process thousands of documents, extract clauses, flag risks, and route uncertain cases to humans, with full audit trails.
  3. Multi-Tenant AI Support Platform
    • Serve many customers with isolated data, per-tenant evals, prompt versioning, and safeguards against prompt injection.

Tier 5: Agents building agents by themselves

The frontier. Systems where agents create, test, and improve other agents or tools. Expect rough edges, and treat this tier as experimentation rather than a solved problem.

  1. Self-Extending Agent
    • An agent that notices it lacks a tool, writes the integration, tests it, and adds it to its own toolbox (behind human approval).
  2. Eval-Driven Prompt Optimizer
    • A system that runs your test suite, proposes prompt changes, measures the results, and keeps only the improvements.
  3. Agent Factory
    • Give it a spec ("an agent that monitors invoices"), and it generates, tests, and deploys a specialized agent with the right tools and guardrails.

How to use this ladder: don't skip tiers. Every project at Tier 3+ quietly relies on the habits you build at Tier 1 and 2: validating outputs, tracking costs, and checking whether answers are actually right.

Free APIs and Resources

You don't need a credit card to start learning this stuff. Here are places where you can get free (or nearly free) access to models. A heads-up: free tiers change often, so always check the current limits and terms before you build on top of one.

Model APIs with free access:

  • OpenRouter: one API that gives you access to many models from different providers, including some free ones. Great for comparing models without juggling accounts.
  • Google AI Studio: Google's playground and API for Gemini models, with a free tier for experimenting.
  • NVIDIA NIM: hosted, ready-to-use model endpoints from NVIDIA that you can try for development and prototyping.
  • OpenCode Zen: a curated set of models tested to work well with coding agents. Check which ones are currently free.
  • Hugging Face Inference Providers: a single interface to run open models through multiple inference providers, with some free usage included.

Also worth a look:

  • Ollama: run open models 100% locally, so there are no API costs and your data stays on your machine.
  • Groq and Cerebras: providers known for very fast inference of open models, and both have offered free tiers.
  • GitHub Models: try a catalog of models directly with your GitHub account.

How to use free tiers wisely:

  • Use them for learning, prototypes, and evals, not as the backbone of a production app.
  • Expect rate limits and occasional changes. Build your code so swapping providers is easy (this pairs nicely with the fallback strategy from Phase 7).
  • Read the data policy. Some free tiers may use your inputs for training, so don't send private or client data.

Wrapping Up

If you made it this far, you now have the full map: from "what's a token?" to shipping an AI feature with tracing, caching, and a fallback plan. That's a lot of ground, and nobody learns it in a weekend (anyone who says otherwise is selling a course).

A few things to take with you:

  • Go in order, but don't wait to be "ready." Build a small project at every tier. Reading about RAG and building a RAG app are two very different experiences.
  • Concepts first, tools second. The tools in this post will change, and some will be gone by next year. Tokens, context, retrieval, tool use, and evaluation will still be around.
  • Measure everything. Costs, quality, latency. If you can't measure it, you're just vibing (fun, but not a strategy).
  • Stay a little skeptical. Every week there's a new "revolutionary" framework. Ask what problem it solves, and whether you actually have that problem.

Pick one phase, open a terminal, and make something small this week. It will probably break in a weird way, and that's exactly how you learn.

Did I miss a tool, a concept, or a resource that helped you? Drop it in the comments, I'd love to add it. And if this roadmap helped, share it with that one friend who keeps saying "I should really learn AI." 🚀

Top comments (0)