DEV Community

Dinesh Kumar Sarangapani
Dinesh Kumar Sarangapani

Posted on Originally published at dineshkumars.dev

Section 1: The dataset and the problem

I am working through Mathematical Foundations of Generative AI, Prof. Prathosh AP's public lectures. The playlist is the spine. These notes are my deep dive on each section: a visual when the picture is the point, a detour when a prerequisite is doing real work, and the formula in my own words.

Section 1: The dataset and the problem

Everything later in the playlist, GANs, VAEs, diffusion, transformers, state-space models, and preference tuning, starts from the same setup. I want this part solid before I touch a model.

The setup I am carrying:

Data $D = {x_1, x_2, \dots, x_n} \overset{\text{iid}}{\sim} P_x$ (unknown), with $x_i \in \mathbb{R}^d$, where $d$ is the dimensionality of the data.
Example: $x_i$ is an image, $400 \times 400 \times 3$, so $d = 480{,}000$.
Goal: estimate $P_x$ and learn to sample from it.

Intuition

Somewhere there is a hidden recipe, a probability distribution, that produced every cat photo, every English sentence, every speech clip. I never get the recipe. I get a pile of examples it produced. Generative modeling is using that pile to reconstruct the recipe well enough to cook up new examples from the same kitchen.

An analogy I can check

Think of a production web service whose request process I cannot inspect. All I have is an access log of 1,000 requests. I want a load-testing tool that produces synthetic traffic indistinguishable from real users: same mix of endpoints, same payload sizes, same timing.

Load-testing analogy Notation
The real, unknown user behaviour $P_x$, the true data distribution
One logged request $x_i$, one data point
The fields in a log line the $d$ coordinates of $x_i \in \mathbb{R}^d$
The access log dataset $D$
The traffic generator the generative model $P_\theta$
Synthetic requests it emits samples from $P_\theta$

Two consequences. I do not need to write the user-behaviour formula down. I need a generator whose output is statistically indistinguishable. That is why a GAN never writes $P_x$ down at all. And replaying the log is not the goal. I want new requests, not copies.

Visual

A hidden 2-D distribution plays the role of $P_x$. At first I only see the dots, the dataset. Reveal $P_x$ to see what is normally invisible, and drag $n$ to see how more data reveals the shape.

Each "Draw a new dataset" gives a different $D$ and the same green $P_x$ underneath. The dataset is random. The distribution is the fixed thing I am after. With $n = 5$ the shape is a guess. With $n = 1000$ the two clumps are obvious.

Detours I needed before the formula

Random variable, and random vector. A random variable is a quantity whose value is decided by chance, like a die roll $X \in {1,\dots,6}$. A random vector is the same idea with several numbers at once, such as height and weight of a randomly chosen person, $X \in \mathbb{R}^2$. Capital $X$ means the random thing in general. Lowercase $x_i$ means one specific value that actually came out.

Distribution and density. $P_x$ tells me how likely each possible value is. For a die it is a table: each face has probability 1/6. For continuous values, such as pixel intensities, it is a density $p_x(x)$, a heat map over space. High density means values land there often. In the chart above, the green shading is that density.

Two facts I will need constantly:

$$
p_x(x) \ge 0 \qquad \int_{\mathbb{R}^d} p_x(x)\, dx = 1
$$

The second says all probability mass integrates to 1. I will use $P_x$ loosely for both the distribution and its density. That is standard, and I will be explicit when the difference matters.

iid, independent and identically distributed.

  • Identically distributed: every $x_i$ comes from the same $P_x$. Every log line came from the same user population, not some from production and some from a test environment.
  • Independent: knowing $x_3$ tells me nothing extra about $x_7$. The first roll of a die does not influence the second.

In notation, $x_i \perp x_j$ and $x_i \sim P_x$.

The formula

$$
D = {x_1, x_2, \dots, x_n} \overset{\text{iid}}{\sim} P_x, \qquad x_i \in \mathbb{R}^d
$$

Read aloud: the dataset $D$ is a collection of $n$ points. Each point is an independent draw from the same unknown distribution $P_x$. Each point is a vector of $d$ real numbers.

Symbol What it is Type / shape Role
$D$ the dataset a set of $n$ vectors the only thing I actually observe
$n$ number of samples integer more data, a better picture of $P_x$
$x_i$ one data point vector in $\mathbb{R}^d$ an image, a sentence embedding, a signal
$d$ dimensionality integer how many numbers describe one point
$\sim$ "is drawn from" relation links data to its source
iid independent and identically distributed assumption makes learning from samples legitimate
$P_x$ true data distribution distribution on $\mathbb{R}^d$ the hidden target, unknown

Why "unknown" is the whole problem. If I knew $P_x$, I would sample from it and stop. Every model in the playlist is a different trick for working with $P_x$ through samples alone.

Why iid matters. It licenses replacing an expectation, an average over $P_x$, with an average over the dataset. That is the law of large numbers:

$$
\frac{1}{n}\sum_{i=1}^n h(x_i) \;\xrightarrow{\,n\to\infty\,}\; \mathbb{E}_{x\sim P_x}[h(x)]
$$

That line is the engine behind every training loss later. The chart shows a small version: the sample mean settles as $n$ grows.

Why $d$ is so large. A 400×400 RGB image is one point in $\mathbb{R}^{480{,}000}$, one coordinate per pixel per colour channel. Real images occupy a thin region of that space. Random pixel vectors look like television static. This is the manifold hypothesis, and it comes back when naive GAN training falls over.

Two jobs hiding in one goal

Given $D$, estimate $P_x$ and learn to sample from it. Those are different jobs, and different models emphasise different ones.

Job Meaning Who emphasises it
Density estimation answer "how likely is this $x$?" autoregressive models, VAEs through the ELBO
Sampling produce a new $x \sim P_x$ GANs, which never write a density down, and diffusion

A GAN can sample well and still be unable to tell me $p(x)$ for a given image. I want to remember that split.

A worked example small enough to do by hand

Let $d = 2$ and $n = 3$. Each point is hours slept and coffees drunk for a random person:

$$
D = {x_1, x_2, x_3} = {(7, 1),\ (6, 2),\ (8, 0)}
$$

Each $x_i \in \mathbb{R}^2$, so $d = 2$. Here $x_2 = (6, 2)$ means 6 hours of sleep and 2 coffees. iid means three different people, surveyed separately, from the same population.

I want $\mathbb{E}[\text{coffees}]$ under the unknown $P_x$. I cannot integrate against a distribution I do not have. iid lets me estimate it from the sample: $\tfrac{1}{3}(1 + 2 + 0) = 1$.

The generative goal is a machine that outputs new pairs such as $(7.5, 0.5)$: plausible, not copied from $D$, and following the same pattern. More sleep tends to go with less coffee. That dependency between coordinates is what $P_x$ captures, and a per-column average does not.

The same picture with $d = 480{,}000$ and $n$ in the millions is image generation.

Where the rest of the playlist sits, as I understand it now

Stretch Model How it attacks "learn $P_x$ from $D$"
Early GAN / divergence minimisation push $z \sim \mathcal{N}(0, I)$ through $g_\theta$ and shrink an f-divergence to $P_x$
Next VAE latent-variable $P_\theta(x) = \int P_\theta(x, z)\,dz$, maximise the ELBO
Then WGAN the same game, with the Wasserstein distance
Then DDPM add noise to $x_0$, learn to reverse it
Then Autoregressive / transformer factorise $P_\theta(x) = \prod_t P_\theta(x_t \mid x_{<t})$
Then SSM / Mamba another sequence model
Late RLHF / DPO move the learned distribution toward preferences

The next section turns the goal into a recipe: pick a parametric family $P_\theta$, define a divergence, and minimise it using only samples.

Questions I want to be able to answer before I move on:

  1. If a dataset mixes 5,000 phone photos and 5,000 medical scans, which part of iid breaks, and why does that hurt?
  2. What is $d$ for a 28×28 grayscale digit? For a 64×64 RGB image?
  3. Once I have $D$, do I know $P_x$ exactly? What changes as $n \to \infty$?

Next: Section 2: The general principle of generative models.

Top comments (0)