DEV Community

Suresh Kumar Pallapothu
Suresh Kumar Pallapothu

Posted on Originally published at sureshpallapothu.in

Day 5: The Math Behind the Magic — Why Linear Algebra and Probability Matter

Under the Python scripts and APIs, every machine learning model is made of mathematics. To really understand how AI processes information and makes decisions, it helps to know its two foundations: linear algebra and probability. Today we look at both, in plain language, with no equations that you have to solve.

In plain terms: If a neural network were a factory, linear algebra would be the machinery that moves and transforms the materials, and probability would be the quality-control desk that decides how confident the factory can be in each product.

Linear Algebra: The Mathematics of Data

Linear algebra is the branch of mathematics that deals with vectors, matrices and the operations on them. In machine learning, almost all data and every model are stored in these forms, so algorithms can compute with them quickly.

Scalars, vectors, matrices and tensors

  • A scalar is a single number, such as a customer's age.
  • A vector is an ordered list of numbers. It usually describes one data point, such as a customer's age, income and tenure.
  • A matrix is a table of numbers in rows and columns. A whole dataset forms a matrix, where each row is a data point and each column is a feature.
  • A tensor generalizes these to more dimensions. A colour image is a 3D tensor: height, width and three colour channels. Tensors are the standard format fed into deep neural networks.

Diagram: Linear algebra's four data containers: a scalar is one number, a vector is a list, a matrix is a table, and a tensor is a stack of tables. See the animated version.

Diagram: A dataset naturally forms a matrix: each row is one data point (a vector of its features) and each column is one feature across all the data. See the animated version.

Diagram: A colour image is a 3D tensor: height by width by three colour channels (red, green, blue). Tensors are what deep networks take in. See the animated version.

Matrix operations: the engine of deep learning

Deep learning leans on huge numbers of matrix multiplications. Each layer of a neural network takes its input vector, multiplies it by a matrix of learned weights, and so produces a new vector. (As you saw on Day 3, a bias is then added and an activation function is applied, which lets the network learn more than straight-line patterns.) Each output is a weighted sum of all the inputs. This is exactly the kind of job that GPUs are built for, which is why they power modern AI.

Diagram: Each layer of a neural network multiplies its input vector by a weight matrix. Every output is a weighted sum of all the inputs, and a GPU can do millions of these at once. See the animated version.

Here is the whole idea in a few lines of Python, using the popular NumPy library:

import numpy as np

x = np.array([1.0, 2.0, 3.0])        # a vector: one data point with 3 features
W = np.array([[0.2,  0.5],
              [0.4, -0.1],
              [0.3,  0.8]])           # a weight matrix: 3 inputs -> 2 outputs

scores = x @ W                        # matrix multiplication -> [1.9, 2.7]
probs = np.exp(scores) / np.exp(scores).sum()   # softmax: scores -> probabilities
print(scores, probs.round(2))         # [1.9 2.7] [0.31 0.69]
Enter fullscreen mode Exit fullscreen mode

Embeddings: giving meaning a position

A powerful use of vectors is the embedding. An AI turns a word, a sentence or an image into a vector in such a way that things with similar meaning end up close together. Finding related ideas then becomes a simple question of measuring distance between vectors. This is the principle behind semantic search, recommendations, and the retrieval step in RAG, which we will build later in the series.

Diagram: Embeddings turn words into vectors. Words with similar meaning land close together, so an AI can find related ideas by measuring distance. This is the idea behind semantic search and RAG. See the animated version.

Dimensionality reduction

Real datasets can have hundreds of features, and many of them overlap. Techniques like Principal Component Analysis (PCA) use linear algebra, specifically eigenvectors and eigenvalues, to find the directions in which the data varies the most. The data is then compressed into fewer dimensions while keeping most of its information, which makes it easier to store, visualize and learn from.

Diagram: Principal component analysis finds the direction in which the data varies most, using eigenvectors, and projects the points onto it, turning two features into one while keeping most of the information. See the animated version.

Probability: The Language of Uncertainty

If linear algebra gives AI the structure to hold data, probability gives it the logic to reason about it. Probability measures how likely an event is. The real world is messy, so machine learning is very often probabilistic and not deterministic: instead of stating facts, models weigh the odds.

Prediction is not certainty

When a language model writes the next word, or a model predicts a price, it is not stating a fact. It gives a probability based on patterns in its training data. A language model scores every possible next word and then chooses from that distribution, and this is also why the same prompt can give different answers on different runs. The step that turns raw scores into probabilities that add up to 100% is usually a function called softmax, as in the code above.

Diagram: When an LLM writes, it does not know the next word. It assigns a probability to every possible word and picks from that distribution. See the animated version.

Conditional probability and Bayes' rule

Conditional probability is the chance of one thing given that another has already happened. Bayes' rule is the formula that updates a belief when new evidence arrives, and it sits at the heart of classifiers like Naive Bayes. A spam filter, for example, asks: given that this email contains the word "free", how likely is it to be spam?

Diagram: Conditional probability in action: out of the emails containing the word 'free', 80 are spam and 45 are not, so the chance an email is spam given that word is 64%. Naive Bayes filters combine many such clues. See the animated version.

In plain terms: The word "free" alone does not prove an email is spam. It only raises the odds. A real filter combines many small clues like this, each nudging the probability up or down.

Probability distributions

How data is spread out matters. The normal distribution (the bell curve) describes the typical value, the mean, and how far values usually spread around it. Once an algorithm knows what is normal, it can spot what is not: values far out in the tails are very unlikely, which is how many anomaly detectors flag fraud or faults.

Diagram: The normal distribution describes typical values (the mean) and how far they usually spread. Values far out in the tail are very unlikely, which is how many anomaly detectors flag them. See the animated version.

Putting the Two Together

In a single prediction, both pillars work in sequence. Linear algebra holds the data and does the heavy calculation, layer after layer. Probability then turns the final scores into a confident answer.

Diagram: In one prediction, linear algebra does the heavy calculation and probability turns the result into a confident answer. See the animated version.

Why It Matters for Implementation

You do not need to solve these equations by hand to build AI. Libraries do the heavy calculation for you. But knowing the math lets you:

  • read research papers and turn their descriptions into code,
  • understand what the model is really doing, so that you can explain and trust its results, and
  • troubleshoot when something goes wrong.

When a model gives bad predictions, the cause is often not the code syntax at all. It is more likely a mismatch in the shape of the data, features on very different scales, or a distribution in production that no longer looks like the training data (the data drift we saw on Day 4). Those are problems of linear algebra and probability.

Coming Up Next

Day 6: Data preprocessing: cleaning the messy reality of enterprise data.

ArtificialIntelligence #LinearAlgebra #Probability #MachineLearning #DataScience #MathForAI #TechEducation #Innovation


Originally published at https://sureshpallapothu.in/blog/day-5-math-behind-ai, where this post includes animated diagrams.

Top comments (0)