I am working through Mathematical Foundations of Generative AI, Prof. Prathosh AP's public lectures. The playlist is the spine. These notes are my deep dive on each section: a visual when the picture is the point, a detour when a prerequisite is doing real work, and the formula in my own words.
Section 1: The dataset and the problem
Everything later in the playlist, GANs, VAEs, diffusion, transformers, state-space models, and preference tuning, starts from the same setup. I want this part solid before I touch a model.
The setup I am carrying:
Data $D = {x_1, x_2, \dots, x_n} \overset{\text{iid}}{\sim} P_x$ (unknown), with $x_i \in \mathbb{R}^d$, where $d$ is the dimensionality of the data.
Example: $x_i$ is an image, $400 \times 400 \times 3$, so $d = 480{,}000$.
Goal: estimate $P_x$ and learn to sample from it.
Intuition
Somewhere there is a hidden recipe, a probability distribution, that produced every cat photo, every English sentence, every speech clip. I never get the recipe. I get a pile of examples it produced. Generative modeling is using that pile to reconstruct the recipe well enough to cook up new examples from the same kitchen.
An analogy I can check
Think of a production web service whose request process I cannot inspect. All I have is an access log of 1,000 requests. I want a load-testing tool that produces synthetic traffic indistinguishable from real users: same mix of endpoints, same payload sizes, same timing.
| Load-testing analogy | Notation |
|---|---|
| The real, unknown user behaviour | $P_x$, the true data distribution |
| One logged request | $x_i$, one data point |
| The fields in a log line | the $d$ coordinates of $x_i \in \mathbb{R}^d$ |
| The access log | dataset $D$ |
| The traffic generator | the generative model $P_\theta$ |
| Synthetic requests it emits | samples from $P_\theta$ |
Two consequences. I do not need to write the user-behaviour formula down. I need a generator whose output is statistically indistinguishable. That is why a GAN never writes $P_x$ down at all. And replaying the log is not the goal. I want new requests, not copies.
Visual
A hidden 2-D distribution plays the role of $P_x$. At first I only see the dots, the dataset. Reveal $P_x$ to see what is normally invisible, and drag $n$ to see how more data reveals the shape.
Each "Draw a new dataset" gives a different $D$ and the same green $P_x$ underneath. The dataset is random. The distribution is the fixed thing I am after. With $n = 5$ the shape is a guess. With $n = 1000$ the two clumps are obvious.
Detours I needed before the formula
Random variable, and random vector. A random variable is a quantity whose value is decided by chance, like a die roll $X \in {1,\dots,6}$. A random vector is the same idea with several numbers at once, such as height and weight of a randomly chosen person, $X \in \mathbb{R}^2$. Capital $X$ means the random thing in general. Lowercase $x_i$ means one specific value that actually came out.
Distribution and density. $P_x$ tells me how likely each possible value is. For a die it is a table: each face has probability 1/6. For continuous values, such as pixel intensities, it is a density $p_x(x)$, a heat map over space. High density means values land there often. In the chart above, the green shading is that density.
Two facts I will need constantly:
$$
p_x(x) \ge 0 \qquad \int_{\mathbb{R}^d} p_x(x)\, dx = 1
$$
The second says all probability mass integrates to 1. I will use $P_x$ loosely for both the distribution and its density. That is standard, and I will be explicit when the difference matters.
iid, independent and identically distributed.
- Identically distributed: every $x_i$ comes from the same $P_x$. Every log line came from the same user population, not some from production and some from a test environment.
- Independent: knowing $x_3$ tells me nothing extra about $x_7$. The first roll of a die does not influence the second.
In notation, $x_i \perp x_j$ and $x_i \sim P_x$.
The formula
$$
D = {x_1, x_2, \dots, x_n} \overset{\text{iid}}{\sim} P_x, \qquad x_i \in \mathbb{R}^d
$$
Read aloud: the dataset $D$ is a collection of $n$ points. Each point is an independent draw from the same unknown distribution $P_x$. Each point is a vector of $d$ real numbers.
| Symbol | What it is | Type / shape | Role |
|---|---|---|---|
| $D$ | the dataset | a set of $n$ vectors | the only thing I actually observe |
| $n$ | number of samples | integer | more data, a better picture of $P_x$ |
| $x_i$ | one data point | vector in $\mathbb{R}^d$ | an image, a sentence embedding, a signal |
| $d$ | dimensionality | integer | how many numbers describe one point |
| $\sim$ | "is drawn from" | relation | links data to its source |
| iid | independent and identically distributed | assumption | makes learning from samples legitimate |
| $P_x$ | true data distribution | distribution on $\mathbb{R}^d$ | the hidden target, unknown |
Why "unknown" is the whole problem. If I knew $P_x$, I would sample from it and stop. Every model in the playlist is a different trick for working with $P_x$ through samples alone.
Why iid matters. It licenses replacing an expectation, an average over $P_x$, with an average over the dataset. That is the law of large numbers:
$$
\frac{1}{n}\sum_{i=1}^n h(x_i) \;\xrightarrow{\,n\to\infty\,}\; \mathbb{E}_{x\sim P_x}[h(x)]
$$
That line is the engine behind every training loss later. The chart shows a small version: the sample mean settles as $n$ grows.
Why $d$ is so large. A 400×400 RGB image is one point in $\mathbb{R}^{480{,}000}$, one coordinate per pixel per colour channel. Real images occupy a thin region of that space. Random pixel vectors look like television static. This is the manifold hypothesis, and it comes back when naive GAN training falls over.
Two jobs hiding in one goal
Given $D$, estimate $P_x$ and learn to sample from it. Those are different jobs, and different models emphasise different ones.
| Job | Meaning | Who emphasises it |
|---|---|---|
| Density estimation | answer "how likely is this $x$?" | autoregressive models, VAEs through the ELBO |
| Sampling | produce a new $x \sim P_x$ | GANs, which never write a density down, and diffusion |
A GAN can sample well and still be unable to tell me $p(x)$ for a given image. I want to remember that split.
A worked example small enough to do by hand
Let $d = 2$ and $n = 3$. Each point is hours slept and coffees drunk for a random person:
$$
D = {x_1, x_2, x_3} = {(7, 1),\ (6, 2),\ (8, 0)}
$$
Each $x_i \in \mathbb{R}^2$, so $d = 2$. Here $x_2 = (6, 2)$ means 6 hours of sleep and 2 coffees. iid means three different people, surveyed separately, from the same population.
I want $\mathbb{E}[\text{coffees}]$ under the unknown $P_x$. I cannot integrate against a distribution I do not have. iid lets me estimate it from the sample: $\tfrac{1}{3}(1 + 2 + 0) = 1$.
The generative goal is a machine that outputs new pairs such as $(7.5, 0.5)$: plausible, not copied from $D$, and following the same pattern. More sleep tends to go with less coffee. That dependency between coordinates is what $P_x$ captures, and a per-column average does not.
The same picture with $d = 480{,}000$ and $n$ in the millions is image generation.
Where the rest of the playlist sits, as I understand it now
| Stretch | Model | How it attacks "learn $P_x$ from $D$" |
|---|---|---|
| Early | GAN / divergence minimisation | push $z \sim \mathcal{N}(0, I)$ through $g_\theta$ and shrink an f-divergence to $P_x$ |
| Next | VAE | latent-variable $P_\theta(x) = \int P_\theta(x, z)\,dz$, maximise the ELBO |
| Then | WGAN | the same game, with the Wasserstein distance |
| Then | DDPM | add noise to $x_0$, learn to reverse it |
| Then | Autoregressive / transformer | factorise $P_\theta(x) = \prod_t P_\theta(x_t \mid x_{<t})$ |
| Then | SSM / Mamba | another sequence model |
| Late | RLHF / DPO | move the learned distribution toward preferences |
The next section turns the goal into a recipe: pick a parametric family $P_\theta$, define a divergence, and minimise it using only samples.
Questions I want to be able to answer before I move on:
- If a dataset mixes 5,000 phone photos and 5,000 medical scans, which part of iid breaks, and why does that hurt?
- What is $d$ for a 28×28 grayscale digit? For a 64×64 RGB image?
- Once I have $D$, do I know $P_x$ exactly? What changes as $n \to \infty$?
Next: Section 2: The general principle of generative models.
Top comments (0)