DEV Community

Cover image for I'm a backend engineer. Here's the order I'd learn AI engineering in (no math first)
Sarques
Sarques

Posted on

I'm a backend engineer. Here's the order I'd learn AI engineering in (no math first)

I spend my days on Java, Spring Boot and Kafka. When I started learning AI engineering, every resource I found had one of two problems: it opened with linear algebra, or it was a 40-hour video course I knew I'd never finish.

What actually worked was learning things in the right order, with each idea small enough to finish in one sitting. This is that order. Each step has a short explanation here, plus a link to a free 5-minute lesson if you want to go deeper.

You don't need ML experience. If you can write a function and call an API, you're ready.


Step 0: Know what the job actually is

"AI engineer" usually doesn't mean training models. It means building products on top of models that already exist: calling them, feeding them the right context, making their output reliable, and keeping them fast and cheap.

That's mostly software engineering, which is good news if you already write backends.

👉 What is AI Engineering? (7 min)


Step 1: Understand what an LLM really does

An LLM predicts the next token, one at a time, based on everything before it. That's it. Almost every "weird" behavior follows from this:

  • It hallucinates because it produces likely text, not checked facts.
  • It doesn't know recent events because its knowledge stops when its training data does.
  • It can't count the r's in "strawberry" because it never sees letters. It sees tokens.

That last one surprises people, so try it yourself: paste any sentence into this tokenizer playground and watch how it gets chopped up.

👉 What are LLMs? · Tokenization


Step 2: Make your first API call, then learn the knobs

Before any framework, call a model directly and look at the raw response. Then learn the three parameters you'll touch every day:

Parameter What it does Rule of thumb
temperature How adventurous each word choice is Low for facts and code, higher for brainstorming
max_tokens Hard cap on the answer's length Set it, or one bad prompt can cost you
top_p Limits choices to the most likely words Change this or temperature, not both

Seeing temperature beats reading about it. Drag the slider in the temperature demo and ask the same question a few times.

👉 Making Your First LLM API Call · Understanding LLM Parameters


Step 3: Make the output something your code can use

Chat replies are for humans. Your backend needs JSON that matches a schema, validated before you trust it. If you've used Pydantic or Jackson, this will feel familiar: define the shape, ask the model for it, and validate. Retry or fail loudly when it doesn't match.

This one step turns an AI demo into an AI feature.

👉 Structured Output from LLMs · Anatomy of an Effective Prompt


Step 4: Embeddings, the idea everything else is built on

An embedding turns text into a list of numbers: coordinates for meaning. Texts that mean similar things land close together, even with no words in common.

You can see the whole idea in a few lines of plain Python. These are toy 3-number vectors (real ones have hundreds of dimensions):

import math

def cosine(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))

# Toy embeddings: imagine a model produced these
reset_password = [0.9, 0.1, 0.0]
forgot_login   = [0.8, 0.2, 0.1]
pizza_recipe   = [0.0, 0.1, 0.9]

print(cosine(reset_password, forgot_login))  # ~0.98, same meaning
print(cosine(reset_password, pizza_recipe))  # ~0.01, unrelated
Enter fullscreen mode Exit fullscreen mode

"How do I reset my password?" and "I forgot my login" share no words, but a search over embeddings still matches them. A vector database is a store that does this comparison fast across millions of items.

👉 What are Embeddings? · Vector Databases in 10 Minutes


Step 5: RAG, which is just "look it up, then answer"

Retrieval-Augmented Generation sounds fancy. It's two steps:

  1. Retrieve: find the chunks of your documents closest to the question, using embeddings.
  2. Generate: give those chunks to the model and tell it to answer only from them.

That's how you get answers about your own data without fine-tuning anything. The hard parts are mostly boring engineering: how you split documents (chunking), and how you measure whether answers are correct (evals). Skip evals and you're shipping vibes.

See the whole flow animated in the RAG pipeline demo.

👉 What is RAG? · Chunking Strategies · Evals Basics


Step 6: Tools and agents (only now)

Agents are where everyone wants to start, and that's why so many agent projects fall apart. With steps 1–5 behind you, they're easy to reason about:

  • Tool calling: the model doesn't run anything. It asks your code to call a function with some arguments, your code runs it, and the result goes back to the model.
  • Agent: that loop on repeat (plan → act → observe) until the task is done.
  • MCP: a standard way to plug tools into models. Think of it as USB-C for AI tools.

The backend instincts you already have (timeouts, retries, least privilege, idempotency) are exactly what makes agents reliable.

👉 From Text Generation to Action · What is MCP? · What are AI Agents?


Step 7: Ship it like a real system

This is where backend engineers have an unfair advantage. Production AI is mostly:

  • Cost: you pay per token, so caching and shorter prompts matter.
  • Latency: streaming makes a 5-second answer feel instant.
  • Security: prompt injection is the new SQL injection. Never let model output pick which tools it's allowed to use.
  • Observability: trace every call, with its tokens and cost, or you'll never debug a chain.

👉 Understanding AI App Costs · Defending Against Prompt Injection · LLM Observability


The short version

  1. What the job is
  2. How LLMs work (tokens!)
  3. API calls and parameters
  4. Structured output
  5. Embeddings
  6. RAG and evals
  7. Tools, MCP and agents
  8. Cost, latency, security, observability

Notice what's missing from the start: calculus, training models from scratch, and choosing a framework. They can all wait, and most of them are optional for this job.


Where these lessons come from

I wrote all of the linked lessons for SproutStack, a free site I'm building on the side for people like me: engineers moving into AI, and anyone starting out in AI, ML, DSA or system design.

  • 105 lessons, 5–8 minutes each
  • Quizzes that explain why every answer is right or wrong
  • Python that runs in your browser, with nothing to install
  • No sign-up needed. Accounts are optional and only sync your progress.

The full AI path is laid out in order here: AI Engineering roadmap.

It's early, and I'd genuinely like to know: which step was hardest for you, or which one is missing? Tell me in the comments. I read all of them, and I'll write up the most requested topic next.

Top comments (0)

Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more