Large Language Models (LLMs) have changed how we interact with computers.
Tools such as ChatGPT, Claude, Gemini, and many open-source models can write code, explain complex topics, summarize documents, translate languages, and generate natural-sounding conversations.
But what is actually happening inside an LLM?
Is it searching the internet?
Is it storing every sentence it has ever seen?
Does it "understand" language like a human?
The answer is much more interesting.
In this article, we'll go step by step through the core ideas behind modern LLMs.
1. What Is an LLM?
LLM stands for Large Language Model.
Let's break that name apart.
Large
"Large" generally refers to the enormous number of parameters in the model.
A parameter is a numerical value learned during training.
Modern neural networks can contain billions or even more parameters.
You can think of parameters as the adjustable values that allow the model to transform an input sequence into a useful prediction.
Language
The model is trained primarily on sequences of language-like data.
For example:
The Linux kernel is written primarily in C.
The model learns statistical relationships between tokens in sequences.
Model
A model is a mathematical system that takes an input and produces an output.
For an autoregressive language model, a simplified description is:
Previous tokens
↓
Neural network
↓
Probability distribution
↓
Next token
For example:
The sky is
might produce probabilities roughly like:
blue 0.72
clear 0.10
dark 0.04
falling 0.01
...
The actual distribution is much larger and more complex.
2. The Most Important Idea: Predict the Next Token
One of the most important concepts for understanding LLMs is:
An autoregressive LLM generates text by repeatedly predicting what token should come next.
Consider:
I love programming
The model might predict:
because
Then the input becomes:
I love programming because
The model predicts another token:
it
Then:
I love programming because it
And so on.
This happens repeatedly.
Conceptually:
Input
│
▼
"I love programming"
│
▼
Predict next token
│
▼
"because"
│
▼
Predict next token
│
▼
"it"
│
▼
Predict next token
│
▼
"is"
│
▼
...
This simple mechanism becomes extremely powerful when the prediction system contains billions of learned parameters.
3. LLMs Don't Usually Process Text Directly
Neural networks operate on numbers.
Computers ultimately work with numerical representations.
So before an LLM can process:
Hello, world!
the text needs to be converted into something numerical.
This happens through tokenization.
4. What Is a Token?
A token is a piece of text.
Depending on the tokenizer, a token might represent:
- a complete word
- part of a word
- punctuation
- whitespace-related patterns
- symbols
- characters
For example, a sentence such as:
Programming is fun!
could be split conceptually into:
["Programming", " is", " fun", "!"]
But tokenization differs between models.
A longer word might instead be divided into multiple pieces.
For example:
unbelievable
could conceptually become:
["un", "believ", "able"]
The exact tokenization depends on the tokenizer.
5. Tokens Become Numbers
The model doesn't receive the literal token strings.
Each token is associated with an integer ID.
For example:
"Hello" → 15496
"world" → 995
"!" → 0
These numbers are only illustrative.
A real tokenizer has its own vocabulary and IDs.
So:
Hello world!
might become:
[15496, 995, 0]
Now we have something the neural network can process.
6. Token IDs Are Not Meaning
This is an important distinction.
Suppose:
cat → 1837
dog → 9281
The numbers themselves don't mean that dogs are mathematically related to cats.
The token IDs are essentially indexes into the model's vocabulary.
The next step is where things become much more interesting.
7. Embeddings
Token IDs are converted into vectors called embeddings.
Imagine:
cat → [0.21, -0.13, 0.72, ...]
dog → [0.19, -0.11, 0.69, ...]
car → [-0.42, 0.81, -0.12, ...]
Real embeddings contain many more dimensions.
A vector might contain hundreds or thousands of numerical values.
The important idea is:
Token
↓
Embedding vector
↓
Neural network
During training, the model learns useful numerical representations.
Words and tokens that occur in similar contexts can develop related representations.
8. Why Embeddings Are Powerful
Consider:
The cat chased the mouse.
and:
The dog chased the ball.
The model sees patterns between concepts and contexts.
Over enormous amounts of training data, it can learn relationships involving:
cat
dog
animal
mouse
ball
chase
eat
run
The model isn't given a dictionary saying:
cat = animal
Instead, these relationships emerge from training on many examples.
9. Position Matters
Consider:
Dog bites man.
versus:
Man bites dog.
The same words appear, but the meaning changes because their positions change.
Therefore, a transformer needs information about where tokens occur in the sequence.
This is handled using positional information.
Modern transformer architectures can use different approaches to represent position, including positional embeddings and rotary positional representations.
Conceptually:
Token embedding
+
Position information
↓
Transformer input
10. The Transformer
The architecture behind most modern LLMs is the Transformer.
The Transformer was introduced in the 2017 research paper:
"Attention Is All You Need"
The architecture became extremely influential because it provided an efficient way to model relationships between tokens.
The central mechanism is:
Attention
11. What Is Attention?
Imagine the sentence:
The programmer fixed the bug because it was causing a crash.
What does "it" refer to?
A language model needs to determine which earlier tokens are relevant.
Attention allows each token to consider other tokens in the context.
Conceptually:
The programmer fixed the bug because it was causing a crash.
↑
│
attention
│
↓
bug
The model can assign different amounts of attention to different tokens.
12. Query, Key, and Value
Self-attention is commonly described using three vectors:
Query
Key
Value
For each token, the model produces these representations.
Very simplified:
Query × Key
↓
Attention score
↓
Softmax
↓
Weighted Values
The result is a representation containing information gathered from other tokens.
13. A Simplified Attention Example
Suppose we have:
The cat sat on the mat.
When processing:
cat
the model might assign attention weights conceptually like:
The → 0.05
cat → 0.20
sat → 0.30
on → 0.10
the → 0.05
mat → 0.30
These numbers are purely illustrative.
The model calculates these relationships mathematically.
The important idea is:
Attention lets the model dynamically determine which parts of the context are useful for representing each token.
14. Multi-Head Attention
Transformers don't normally use only one attention mechanism.
They use multiple attention heads.
For example:
Input
│
┌───────────┼───────────┐
▼ ▼ ▼
Head 1 Head 2 Head 3
│ │ │
▼ ▼ ▼
Pattern Pattern Pattern
│ │ │
└───────────┼───────────┘
▼
Combine
Different heads can learn different relationships.
One head might become useful for syntactic relationships.
Another might track long-range dependencies.
Another might respond to other structural patterns.
The model doesn't receive labels saying what each head should learn. These behaviors emerge during training.
15. The Transformer Block
A simplified Transformer block looks something like:
Input
│
▼
Self-Attention
│
▼
Residual Connection
│
▼
Normalization
│
▼
Feed-Forward Network
│
▼
Residual Connection
│
▼
Normalization
│
▼
Output
A real implementation contains additional architectural details.
And an LLM contains many such blocks stacked together.
16. Feed-Forward Networks
After attention, transformer blocks typically contain a feed-forward neural network.
Conceptually:
Input vector
│
▼
Linear transformation
│
▼
Activation function
│
▼
Linear transformation
│
▼
Output
This gives the model additional capacity to transform and process representations.
17. Residual Connections
Transformers also use residual connections.
Instead of replacing a representation completely:
Input → Layer → Output
the architecture can effectively do:
Input ───────────────┐
│ │
▼ │
Layer │
│ │
└──────► Add ◄─────┘
│
▼
Output
Residual connections help information and gradients flow through deep networks.
This is one of the important engineering ideas that makes very deep neural networks practical.
18. Training an LLM
Now we reach the most important part.
How does the model actually learn?
Suppose the training text contains:
The Earth revolves around the
The expected next token might be:
Sun
Initially, the model may be terrible.
It might predict:
moon 0.20
Sun 0.10
planet 0.08
...
The training system compares the model's predicted probability distribution with the actual target.
This produces a loss.
19. Loss Function
The loss measures how wrong the model's prediction was.
For language models, a common objective is based on cross-entropy loss.
Very simplified:
Prediction
↓
Compare with correct token
↓
Calculate loss
↓
Backpropagation
↓
Update parameters
The goal is to reduce the loss over enormous amounts of training data.
20. Backpropagation
Backpropagation calculates how the model's parameters contributed to the error.
Imagine the model has:
Parameter A
Parameter B
Parameter C
...
Parameter N
The training algorithm calculates gradients such as:
∂Loss / ∂ParameterA
∂Loss / ∂ParameterB
∂Loss / ∂ParameterC
...
These gradients tell the optimizer how parameters should change.
21. Gradient Descent
An optimizer then updates the parameters.
A simplified update looks like:
new_parameter =
old_parameter - learning_rate × gradient
This happens repeatedly.
Millions, billions, or more parameter values can be updated across huge numbers of training steps.
22. The Training Loop
Conceptually:
┌───────────────┐
│ Training data │
└───────┬───────┘
│
▼
Tokenization
│
▼
LLM model
│
▼
Prediction
│
▼
Loss
│
▼
Backpropagation
│
▼
Optimizer
│
▼
Update parameters
│
└───────────┐
│
▼
Next training step
Repeat this enormous number of times.
23. What Does the Model Actually Learn?
This is one of the most fascinating questions.
The model doesn't simply store a collection of sentences.
Instead, training changes its parameters so that the network becomes better at predicting patterns in its training distribution.
Those learned patterns can include:
- syntax
- semantics
- programming
- mathematics
- facts
- reasoning patterns
- formatting
- languages
- code structures
- common conversational patterns
The knowledge is distributed throughout the network rather than being stored like ordinary database records.
24. Training vs Inference
There are two very different phases.
Training
During training:
Data
↓
Prediction
↓
Loss
↓
Backpropagation
↓
Parameter updates
The model learns.
Inference
During inference:
Prompt
↓
Tokens
↓
Transformer
↓
Next-token probabilities
↓
Select token
↓
Add token to context
↓
Repeat generation
The model normally isn't updating its parameters during ordinary inference.
It is using what it learned during training.
25. What Happens When You Send a Prompt?
Suppose you ask:
Explain how a CPU works.
A simplified pipeline is:
Your text
│
▼
Tokenizer
│
▼
Token IDs
│
▼
Embeddings
│
▼
Transformer layers
│
▼
Probability distribution
│
▼
Select next token
│
▼
Repeat
│
▼
Generated response
This process happens extremely quickly.
26. How Does the Model Choose the Next Token?
The model produces a probability distribution.
For example, imagine:
"The CPU executes"
instructions 0.65
programs 0.12
code 0.08
commands 0.05
...
The system then chooses a token according to its decoding strategy.
This is where concepts such as temperature, top-k, and top-p become important.
27. Temperature
Temperature controls how sharply or randomly probabilities are sampled.
Lower temperature
More predictable:
instructions → very likely
Higher temperature
More variation:
instructions → likely
commands → possible
operations → possible
...
Temperature does not make the underlying model "more intelligent."
It changes the sampling behavior.
28. Top-K Sampling
Top-k sampling limits the candidate tokens to the k most probable options.
For example:
Top 5 tokens:
1. instructions
2. programs
3. code
4. operations
5. commands
The model samples from this restricted set.
29. Top-P Sampling
Top-p sampling instead chooses a group of tokens whose cumulative probability reaches a specified threshold.
For example:
A
0.50
B
0.25
C
0.15
D
0.06
E
0.04
With an appropriate probability threshold, only the most likely group may be considered.
30. Why Does an LLM Sometimes Hallucinate?
An LLM is fundamentally trained to generate likely continuations.
It is not automatically a perfect fact database.
Suppose the model receives a question about something obscure.
It may generate an answer that sounds convincing even though the information is incorrect.
This behavior is commonly called a hallucination.
A useful mental model is:
Fluent language ≠ guaranteed factual accuracy
This is extremely important when using LLMs.
31. Does an LLM "Understand" Language?
This question is more philosophical and technical than it first appears.
LLMs clearly learn sophisticated representations and can perform impressive language-related tasks.
However, their internal mechanism is fundamentally a neural computation system trained through statistical optimization.
It is not a human brain.
The safest engineering perspective is:
An LLM learns complex representations and transformations that enable powerful language behavior, but that does not automatically imply human-like consciousness or understanding.
32. Why Are LLMs Called "Large"?
There are several dimensions of scale.
Parameters
The neural network may contain billions of learned parameters.
Training data
Training can involve enormous datasets.
Compute
Training large models requires substantial computational resources.
A simplified relationship is:
More data
+
More parameters
+
More computation
↓
Potentially stronger capabilities
But simply making a model larger does not guarantee unlimited improvement.
Architecture, data quality, optimization, and training methods matter enormously.
33. GPUs and Matrix Multiplication
Why do LLMs need so much computing power?
A major reason is that neural networks perform huge numbers of mathematical operations, particularly matrix operations.
For example:
A × B = C
Large matrix multiplications can contain millions or billions of arithmetic operations.
GPUs are highly suited to this kind of parallel computation.
Conceptually:
CPU
└── General-purpose computation
GPU
├── Many parallel arithmetic units
├── Matrix operations
└── High-throughput computation
Modern AI accelerators are specifically optimized for these workloads.
34. Memory Is Also Important
Running an LLM requires storing model parameters.
Suppose a model has:
70 billion parameters
If each parameter requires approximately 2 bytes:
70 billion × 2 bytes
≈ 140 GB
That's only a simplified parameter-storage calculation.
Real systems also need memory for:
- activations
- temporary tensors
- attention-related state
- runtime buffers
- other system overhead
This is why large models often require multiple GPUs or specialized hardware.
35. Context Windows
An LLM cannot necessarily process an unlimited amount of text in a single request.
The model operates over a context window.
Conceptually:
┌──────────────────────────────┐
│ System/context information │
│ User message │
│ Previous conversation │
│ Documents │
│ Current prompt │
└──────────────────────────────┘
All of this consumes tokens.
The maximum supported context depends on the model and its architecture.
36. Attention and Context
Attention allows tokens to interact with other tokens within the model's context.
For a sequence:
Token 1
Token 2
Token 3
...
Token N
the model creates relationships among these positions.
This is one reason context length matters so much.
Longer contexts generally require more computation and memory, although modern architectures and optimizations can change the scaling characteristics.
37. Pretraining Isn't the Whole Story
A modern assistant usually isn't simply a raw pretrained model.
There can be additional stages after pretraining.
A simplified pipeline is:
Large-scale pretraining
↓
Instruction tuning
↓
Preference/alignment training
↓
Deployment
The exact training pipeline varies between models.
38. Instruction Tuning
A pretrained model may be good at predicting text but not necessarily good at following instructions.
Instruction tuning trains the model on examples such as:
User:
Explain recursion.
Assistant:
Recursion is...
This helps the model become better at responding to user instructions.
39. Alignment
Additional training can encourage useful behaviors such as:
- following instructions
- refusing certain unsafe requests
- producing helpful responses
- respecting formatting requirements
- being more conversational
Different AI systems use different alignment and post-training techniques.
40. LLMs vs Traditional Programs
A traditional program might contain explicit logic:
if (temperature > 30) {
printf("Hot");
}
The programmer explicitly defines the rule.
An LLM works differently.
Its behavior emerges from learned parameters.
Conceptually:
Traditional software:
Input
↓
Explicit rules
↓
Output
LLM:
Input
↓
Learned neural network
↓
Probability distribution
↓
Output
This difference is fundamental.
41. LLMs and Databases Are Different
A database might store:
User ID: 42
Name: Farhad
Age: ...
An LLM does not generally retrieve information from its parameters in this database-like way.
Instead, information is encoded in distributed numerical representations.
This is why an LLM can sometimes produce a fact correctly, incorrectly, or inconsistently.
42. What About RAG?
One common way to improve factual access is Retrieval-Augmented Generation (RAG).
Instead of relying entirely on the model's internal learned information:
Question
↓
LLM
↓
Answer
RAG can use:
Question
↓
Search / Retrieval
↓
Relevant documents
↓
LLM
↓
Answer
For example, a company could connect an LLM to its internal documentation.
The model retrieves relevant documents and uses them as context.
43. Tools Make LLMs More Capable
Modern AI systems can also use external tools.
For example:
User
↓
LLM
↓
Decide that a tool is needed
↓
Tool
↓
Tool result
↓
LLM
↓
Final response
Tools can provide capabilities such as:
- web search
- calculators
- databases
- code execution
- file access
- APIs
This is different from the neural network itself suddenly gaining those capabilities internally.
44. A Complete Simplified Architecture
We can now combine everything:
USER
│
▼
Prompt
│
▼
Tokenization
│
▼
Token IDs
│
▼
Embeddings
│
▼
Positional Information
│
▼
┌─────────────────────────┐
│ Transformer │
│ │
│ ┌───────────────────┐ │
│ │ Self-Attention │ │
│ └─────────┬─────────┘ │
│ ▼ │
│ ┌───────────────────┐ │
│ │ Feed-Forward NN │ │
│ └─────────┬─────────┘ │
│ ▼ │
│ Repeat many │
│ layers │
└────────────┬────────────┘
│
▼
Output logits
│
▼
Probability distribution
│
▼
Token selection
│
▼
Next token
│
└───────┐
│
▼
Repeat generation
│
▼
Answer
45. The Most Important Mental Model
If you remember only one thing from this article, remember this:
Text
↓
Tokens
↓
Numbers
↓
Embeddings
↓
Transformer
↓
Attention + neural network layers
↓
Probability distribution
↓
Next token
↓
Repeat
↓
Generated text
During training:
Text
↓
Predict
↓
Calculate loss
↓
Backpropagate
↓
Update parameters
↓
Repeat billions/trillions of times
That's the core idea.
46. So Is an LLM Just "Next-Word Prediction"?
Technically, next-token prediction is at the heart of autoregressive language modeling.
But calling an advanced LLM "just autocomplete" can be misleading.
Why?
Because learning to predict tokens across enormous and diverse datasets forces the network to learn many complicated internal representations.
To predict the next token effectively, the model may need to represent:
grammar
syntax
semantics
relationships
code structure
mathematical patterns
world knowledge
long-range dependencies
So the training objective can be simple while the learned internal behavior becomes extremely sophisticated.
47. What LLMs Still Cannot Guarantee
LLMs can be extremely capable, but they have important limitations.
They can:
- produce incorrect information
- misunderstand ambiguous prompts
- make reasoning mistakes
- generate outdated information
- confidently state false claims
- struggle with some exact computations
- inherit biases from training data
Therefore:
Never confuse fluent output with guaranteed truth.
For important decisions, verify the information with appropriate primary sources or reliable tools.
48. Final Summary
An LLM is a large neural network trained to model sequences of tokens.
The basic process is:
TRAINING
Training text
↓
Tokenization
↓
Transformer
↓
Prediction
↓
Loss
↓
Backpropagation
↓
Parameter updates
↓
Repeat
INFERENCE
User prompt
↓
Tokenization
↓
Transformer
↓
Next-token probabilities
↓
Select token
↓
Add token to context
↓
Repeat
↓
Generated response
The key technologies behind this process include:
- Tokenization
- Embeddings
- Positional representations
- Transformers
- Self-attention
- Multi-head attention
- Feed-forward networks
- Residual connections
- Gradient descent
- Backpropagation
- Large-scale training
- Instruction tuning
- Alignment
- Sampling and decoding
The remarkable part isn't that the model was explicitly programmed with a giant collection of rules.
Instead, a neural network was trained to predict tokens so many times, across so much data, that it learned highly complex representations useful for language and many other tasks.
That's the fundamental idea behind modern LLMs.
What's Next?
If you want to go deeper, the natural next step is to understand the Transformer mathematically—especially how Query, Key, Value, attention scores, softmax, and matrix multiplication work together.
Once you understand that, the internal architecture of an LLM becomes much less mysterious.
Top comments (0)