DEV Community

Filipe Martins
Filipe Martins

Posted on

How LLMs Work: A Journey Through One Sentence

How LLMs Work: A Journey Through Tokens, Attention, and Transformers

“My favourite rock band is…”

How does a large language model take that unfinished sentence and decide what comes next?

You’ve probably heard that LLMs “predict the next token”. But that explanation skips the interesting part: everything that happens between typing a sentence and getting a prediction.

So, let’s follow that sentence through a transformer, from text to numbers, through embeddings and attention, all the way to the next token.

I made a video walking through this journey, with vectors, matrix shapes, and probably more rock references than strictly necessary. Watch it here, or keep reading for the written walkthrough.

We’ll use a simplified GPT-style, decoder-only transformer. The example uses 512-dimensional vectors and, later, eight attention heads. These are our example settings, rather than specifications for every LLM.

1. Tokenization: turning text into IDs

Before a model can process our sentence, we need a numerical representation of it.

A tokenizer breaks text into tokens and maps those tokens to IDs in its vocabulary.

A token might represent a word, part of a word, punctuation, or text containing spaces. For readability, we’ll mostly talk about whole words, but actual token boundaries depend on the tokenizer.

That gives us a sequence of IDs the model can work with.

There’s a problem, though: an ID identifies a token, but its numerical value doesn’t describe its meaning.

If “rock” has ID 20 and “band” has ID 21, that doesn’t make them more closely related than tokens 20 and 900.

For that, we need embeddings.

2. Embeddings: giving each token a vector

An embedding layer is a lookup table containing one learned vector for each token.

The token ID tells us which row to retrieve.

In our example, every vector contains 512 numbers. That size is the model’s dimensionality, usually written as d_model.

For one sequence, our representation now has this shape:

sequence length × 512
Enter fullscreen mode Exit fullscreen mode

If we process several sequences in a batch, it becomes:

batch size × sequence length × 512
Enter fullscreen mode Exit fullscreen mode

Those embedding values start out essentially random and are adjusted during training. Over time, the model learns representations that help it predict text.

A useful mathematical operation here is the dot product: multiply matching components of two vectors, then add the results.

It measures alignment while also depending on the vectors’ lengths. If both vectors have unit length, their dot product is their cosine similarity.

We’ll meet the dot product again shortly. As I say in the video: dot products. Don’t leave home without them.

3. Position: the same words can mean different things

Consider these sentences:

The dog bit the man.

The man bit the dog.

Same words. Very different day for the dog.

Token embeddings alone don’t tell us where each token appears, so we need to supply position information too.

In the video, I use the learned positional embeddings approach associated with GPT-2. Each position has its own trainable vector, which we add to the token embedding:

input vector = token embedding + position embedding
Enter fullscreen mode Exit fullscreen mode

Our representation now carries information about both the token and its position.

That brings us to the transformer blocks.

4. Self-attention: gathering information from context

“Rock” can refer to music, a stone, or someone you’ve seen in a movie.

The surrounding text helps distinguish those meanings. Self-attention lets each position gather information from other permitted positions in the same sequence.

To calculate it, we create three learned transformations of the input:

  • Queries (Q): used to score what each position attends to.
  • Keys (K): compared with those queries.
  • Values (V): the information that gets mixed together.

If our input is X, we can express these transformations as:

Q = X @ W_Q
K = X @ W_K
V = X @ W_V
Enter fullscreen mode Exit fullscreen mode

Here, @ means matrix multiplication. The W matrices contain learned parameters.

We compare queries and keys using dot products, arranged into one matrix multiplication:

scores = Q @ Kᵀ
Enter fullscreen mode Exit fullscreen mode

Transposing K swaps its rows and columns, giving us a grid of scores between positions.

Next, we scale the scores, apply a causal mask, and use softmax to turn each row into attention weights that sum to one. Those weights determine how we mix the value vectors.

The scaled dot-product calculation comes from the transformer architecture introduced in Attention Is All You Need.

For an illustrative example, the final position in our sentence might assign 70% of its attention to “rock”, 10% to “band”, and the remaining 20% elsewhere.

These are made-up weights to explain the operation. The model calculates its actual weights from the input.

The output is a weighted mixture of value vectors: a new representation informed by context.

5. Multi-head attention: learning several ways to relate tokens

One attention calculation gives us one way of mixing information.

Multi-head attention runs several of these calculations in parallel, using different learned projections.

With our example settings:

512 dimensions ÷ 8 heads = 64 dimensions per head
Enter fullscreen mode Exit fullscreen mode

Each head works with its own 64-dimensional queries, keys, and values. The scaling inside each head therefore uses the square root of 64.

The learned projections happen before the split, so a head can draw on information from across the input vector. It doesn’t just receive 64 untouched embedding components.

We also don’t manually assign jobs such as “grammar head” or “punctuation head”. The useful patterns emerge through training.

Finally, we join the head outputs together and apply another learned transformation to mix their information.

6. Causal masking: stopping the model from seeing the answer

One of my favourite parts of this whole process is that the training text supplies its own next-token targets.

Using whole words for illustration:

“My”                       → “favourite”
“My favourite”             → “rock”
“My favourite rock”        → “band”
“My favourite rock band”   → “is”
Enter fullscreen mode Exit fullscreen mode

During training, we can calculate predictions at many positions in parallel.

But if we provide the whole sequence, what stops a position from looking ahead at the answer?

That’s the job of the causal mask.

Each position can attend to itself and earlier positions, while future positions are blocked. The mask is applied to the attention scores before softmax, so blocked positions receive zero attention weight.

Now the model has to predict using the context available up to that point.

During generation, its own newly generated tokens become part of that context. That’s why we call this process autoregressive.

7. Residual connections, normalization, and feed-forward layers

Attention is a major part of a transformer block, but there are other pieces too.

Residual connections add a sublayer’s input back to its output. Think of recording a voice, processing a separate copy with an effect, and mixing the two tracks together. These shortcuts help information and gradients travel through a deep network.

Normalization helps keep the numerical scales manageable. In the recording analogy, it’s a little like adjusting levels before further processing.

The feed-forward network transforms each position’s vector using learned transformations with a nonlinear activation between them. This lets the model learn more complicated patterns than a chain of linear transformations alone.

Attention mixes information between positions; the feed-forward network processes the resulting information at each position.

We stack transformer blocks so that one block’s output becomes the next block’s input.

8. Logits and softmax: choosing the next token

After the transformer blocks, an output projection produces a score for every token in the vocabulary.

These raw scores are called logits.

To continue our sentence, we use the predictions at the final position. Softmax converts its logits into a probability distribution over possible next tokens.

In the video, I illustrate this with a tiny pretend vocabulary containing band names and “banana”. A real tokenizer might split those names into multiple tokens, but the principle is the same: every vocabulary entry receives a score.

Temperature adjusts the logits before softmax, changing how concentrated or spread out the resulting distribution is.

One way to choose the next token is to take the most likely one, as in the video’s example. Generation can also sample from the distribution.

Once a token is chosen, we append it to the sequence and repeat.

How does the model learn?

During training, we know the target next token because it already appears in the training text.

The training process measures how well the model predicted that target and adjusts its learned parameters to reduce prediction error across examples.

Those parameters include the embeddings, attention projections, and feed-forward transformations we’ve just followed.

Repeat that process over many examples, and the model becomes better at predicting continuations.

When we generate text with the trained model, we use those learned parameters to produce one token after another.

Our unfinished sentence has travelled through quite a bit:

Text → token IDs → embeddings and position information
     → transformer blocks → logits → probabilities → next token
Enter fullscreen mode Exit fullscreen mode

If you want to follow the vectors and matrix shapes through that journey, watch the full video on YouTube. I walk through the same example step by step, including how the attention calculations fit together.

Which part would you like me to go deeper into next: embeddings, queries and keys, causal masking, or temperature? Let me know in the comments.

Top comments (0)