If you already write backend code, you are closer to a GenAI developer role than you think. The job in 2026 is mostly engineering: APIs, data pipelines, reliability and scale, with an LLM in the middle.
This roadmap is ordered by what you need to ship, not by what sounds impressive.
Level 0: The foundation you already half-have
- Python (async/await, typing)
- FastAPI + Pydantic
- Postgres, Redis
- Docker and Git
If async def and dependency injection feel shaky, fix that first. Every LLM call is slow I/O, and one blocking call slows every user.
Level 1: Call an LLM like an engineer
The first skill interviewers test is reliable structured output. Not "chat with a PDF", but "return valid data every time".
from pydantic import BaseModel
from langchain.chat_models import init_chat_model
class Invoice(BaseModel):
vendor: str
gstin: str | None
total_amount: float
llm = init_chat_model("openai:gpt-4.1-mini")
extractor = llm.with_structured_output(Invoice)
result = extractor.invoke("Extract the invoice fields: ...")
print(result.total_amount)
Learn: tokens, context windows, temperature, retries, fallbacks, streaming, cost per request.
Level 2: Retrieval-Augmented Generation (RAG)
Learn embeddings, a vector database (pgvector, Qdrant or Pinecone), chunking, hybrid search, reranking and evaluation. The gap between a demo and production RAG is evaluation: build a test set of real questions and measure faithfulness and context recall before you ship.
Level 3: Agents and workflows
LangGraph: state, nodes, conditional edges, checkpointing, human-in-the-loop
MCP: expose tools to any agent through a standard protocol
A2A: let agents from different systems collaborate
Rule of thumb: most production "agents" are 80% predictable workflow and 20% LLM decision-making.
Level 4: Evals, observability, security
Tracing (LangSmith, Langfuse), quality gates in CI (RAGAS, DeepEval), prompt-injection defence, red-teaming with promptfoo.
Level 5: Cloud and fine-tuning
AWS Bedrock or Azure AI, plus LoRA/QLoRA fine-tuning with Hugging Face PEFT. Know the decision: prompt first, RAG for knowledge, fine-tune for behaviour or format.
Level 6: LLMOps at scale
vLLM, LiteLLM gateway, Redis caching and rate limiting, circuit breakers, Kubernetes, k6 load tests.
A 12-week plan if you work full-time
| Weeks | Focus | Ship this |
|---|---|---|
| 1 | FastAPI foundations | Dockerised API with auth and tests |
| 2 to 3 | LLM basics + structured output | Document extraction API |
| 4 to 5 | Production RAG | Q&A over your company docs with citations |
| 6 to 7 | LangGraph workflows | Approval workflow with human-in-the-loop |
| 8 | Agents, MCP, A2A | Multi-agent assistant with one MCP server |
| 9 | Cloud AI + fine-tuning | Same RAG on Bedrock or Azure |
| 10 to 11 | LLMOps | Gateway with caching, rate limits, load test |
| 12 | Portfolio + interview prep | READMEs, demo videos, mock interviews |
Want to do this with a cohort?
This plan mirrors the 8 sprints of Vector 2.0, a live Gen-AI developer cohort by TechSimPlus taught by Prateek Mishra. You build 5 production projects (LedgerLens, RegRadar, ClaimSense, TripPilot, TokenGrid) and do 5 mock interviews.
Next batch starts 21 Nov 2026: https://vector.techsimplus.com
What level are you at? Drop it in the comments and I will suggest your next project.
Top comments (0)