From Noise to Data
A mathematical introduction to diffusion models, from discrete denoising and score matching to stochastic differential equations.
›Contents
Diffusion models revolve around a simple idea: take a sample from an unknown data distribution, add noise (usually Gaussian) until its structure disappears and only little of the original signal remain, and train a neural network to undo that corruption. Generation then starts from noise and repeatedly applies the learned reverse transformation. An image, time series, audio signal, or other high-dimensional object gradually surface.
The description is simple. The mathematics is more articulated and connects latent-variable models, variational inference, score matching, stochastic differential equations, and numerical solvers. This article aims to present these connections and core results without proving every detail.
The problem: learning a distribution
Let
denote an observation drawn from an unknown data distribution. A generative model search for a tractable approximation from which we can draw new samples.
The difficulty is that real data occupy a complicated region of a high-dimensional space. Directly transforming a simple distribution into this geometry can require architectural restrictions, adversarial training, or an intractable normalizing constant. Diffusion models take another indirect route:
- Define a fixed process that destroys the data distribution
- Learn the reverse of that process
- Generate by starting from the simple terminal distribution (prior) and moving backward
Therefore, the forward process is known. Only the reverse process must be learned.
Forward diffusion: progressively destroying structure
A denoising diffusion probabilistic model (DDPM) defines a Markov chain with noise steps:
where
The schedule controls how much noise is added at step . Define
The transition can then be written through the reparameterization trick:
Because a linear combination of independent Gaussian variables remains Gaussian, all preceding steps collapse into one closed-form marginal:
Equivalently,
This identity is the key, because it states that training does not require simulating all previous transitions, instead we can select any , draw one Gaussian noise vector, and construct directly. We are sampling directly from the conditional marginal distribution.
The signal-to-noise ratio at time is
Early steps have high SNR and retain most of the data signal. Late steps have low SNR and are dominated by noise. A useful schedule makes sufficiently small that
Reverse diffusion: learning to denoise
The forward transition gradually removes information, therefore the generation requires to "add" it back through the transition . Since this, depends on the unknown data distribution and is therefore not directly available. We approximate it with
The full generative model is
Although is unknown, conditioning additionally on the clean observation makes the posterior tractable:
with
and
This posterior tells us what the "ideal" reverse mean would be if the clean sample were known. But of course, at generation time it is not known, so the neural network must infer the missing information from and .
The variational objective
Diffusion models are latent-variable models whose latent variables are . Their negative log-likelihood is bounded by a variational objective:
Using the Markov structure, this variational upper bound can be decomposed as
The first term matches the terminal distribution to the Gaussian prior. The middle terms teach the learned reverse process to approximate the exact forward posterior. The final term reconstructs the data from the least-corrupted latent.
If the reverse variance is fixed, each middle KL divergence reduces to a weighted squared error between the true posterior mean and the learned mean .
Why the network predicts noise
The posterior mean can be reparameterized using the noise that created . Since
we can recover the clean sample, given the noise, as
Substitution into gives
The reverse mean can therefore be parameterized by a neural network that predicts the injected noise:
With a fixed reverse variance , the corresponding variational term is proportional to
Ho, Jain, and Abbeel found that removing this time-dependent coefficient often improves sample quality. This produces the now-standard simplified loss:
where is usually sampled uniformly from . A complete training observation can therefore be generated using one clean sample, one timestep, and one Gaussian noise draw.
The score: a vector field toward probable data
The DDPM derivation explains the training objective, while Score Matching (SM) explains what the network actually learns geometrically.
Suppose a density is represented as an energy-based model:
Its score is the gradient of its log-density with respect to the data:
For the energy-based model,
The unknown normalizing constant disappears. The score does not state how probable a point is, but it gives the local direction in which log-density increases the fastest.
A score network could be trained by minimizing the Fisher divergence
The true data score is unknown. Under boundary conditions, classical score matching rewrites the objective, up to a constant independent of , as
This avoids the unknown data score but introduces the divergence of the network , which is expensive in high dimensions, since it requires the computation of the Jacobian matrix. Denoising Score Matching (DSM) gives a more convenient target.
Denoising score matching
Perturb a clean sample with Gaussian noise:
The conditional perturbation kernel is known:
so its conditional score is:
DSM trains against this known target:
At its optimum, this objective estimates the score of the marginal perturbed distribution, not merely the score of a Gaussian centred on one training example:
Training over many noise scales is essential, since at low noise, the model learns fine structure near the data manifold, while at high noise, separated modes overlap and the score remains significant in regions that would otherwise contain no training data.
For the DDPM perturbation kernel, the standard deviation is . Therefore,
which gives the central equivalence
Predicting noise and estimating the score are the same task up to a known scale.
Sampling in discrete time
After training, sampling starts from
and applies the learned reverse transition for . With the noise parameterization,
where
for , and at the final step. Common fixed choices are or .
Each evaluation estimates a small amount of noise to be removed. The process is stable because no single network call must map from the prior vector directly to a clean sample. On the other side, its main weakness is computational: ancestral DDPM sampling can require hundreds or thousands of sequential network evaluations, quick is slow and expensive.
Continuous time: diffusion as an SDE
The discrete chain becomes a stochastic differential equation when the steps become infinitesimal. Let , with , follow
where is the drift, is the diffusion coefficient, and is standard Brownian motion. The forward SDE transforms the data density into a tractable terminal density .
Different choices recover important diffusion families:
| Family | Forward SDE | Behaviour |
|---|---|---|
| Variance preserving (VP) | Continuous analogue of DDPM; unit Gaussian variance | |
| Variance exploding (VE) | Adds noise without shrinking the signal; variance grows with | |
| sub-VP | Uses less diffusion than VP and is useful for likelihood-oriented models |
Here
For the VP SDE, the transition kernel remains Gaussian:
Thus
This is the continuous counterpart of the DDPM marginal, with playing the role of .
The reverse-time SDE
A forward diffusion has a corresponding reverse-time diffusion. When integrated from to , it satisfies
where during reverse integration and is Brownian motion in reverse time.
Only one quantity is unknown: the time-dependent score . Replacing it with the neural approximation results in the learned generative process:
This equation says that the reverse dynamics combine the known physical drift, a learned score correction, and stochastic noise.
The continuous denoising-score objective is
The weighting determines which noise regions dominate training. A common choice is proportional to the inverse expected squared norm of the conditional score. But also the choice has a likelihood interpretation: under regularity conditions, the weighted score error is related to the upper bound on the KL divergence between the data distribution and the distribution induced by the reverse SDE.
A deterministic path: the probability-flow ODE
Interestingly, through Fokker-Plank equation can be shown that the same marginal densities can be generated by the deterministic ordinary differential equation
or, after replacing the true score,
The reverse SDE and probability-flow ODE do not produce identical trajectories. One is stochastic and one is deterministic. They do, however, share the same marginal distribution at every time when the score is exact.
The ODE also enables likelihood evaluation through the instantaneous change-of-variables formula:
Consequently,
Computing the divergence exactly requires the trace of a large Jacobian. Hutchinson's identity replaces it with an unbiased stochastic estimate:
Automatic differentiation can evaluate the required vector-Jacobian product without materializing the full Jacobian.
What the neural network actually does
The mathematical framework does not prescribe a single architecture. The network only needs to map a noisy object and its noise level to an output of the same shape:
where is optional conditioning information.
For images, a U-Net is a natural choice because its encoder captures large-scale structure while skip connections preserve local detail. The timestep is typically transformed into a sinusoidal embedding and injected throughout the network. Self-attention adds long-range interactions. For sequences and other structured data, transformers or specialized temporal architectures can replace the U-Net backbone without changing the diffusion mathematics.
Given the equivalent parameterizations we can predict different but related targets:
| Target | Network output | Conversion to the score |
|---|---|---|
| Noise | ||
| Clean sample | ||
| Score | Direct output |
They encode the same "ideal" quantity, but at different scales across noise levels. The target, loss weighting, input scaling, noise schedule, and sampler therefore influence in practice the output. The EDM formulation makes this separation explicit by writing a denoiser as
where the coefficients keep network inputs and targets well-scaled across noise levels.
Conditioning and guidance
For a condition , such as a class label, text embedding, or observed portion of a time series, the desired distribution is . Bayes' rule gives
Classifier guidance (CG) estimates the second term with a classifier trained on noisy inputs. Classifier-free guidance (CFG) instead trains only one network both conditionally and unconditionally by sometimes replacing with a null condition . At inference, the predictions are combined as
At , this recovers the conditional prediction. Values amplify the condition, often improving adherence at the cost of diversity and sometimes realism.
Faster and cheaper diffusion
The original reverse chain is expensive because its steps are sequential. Several approaches reduce that cost:
- Non-Markovian or ODE samplers: DDIM and probability-flow methods can follow deterministic paths with fewer steps
- Higher-order solvers: methods such as Heun integration and DPM-Solver approximate the reverse dynamics more accurately per network evaluation
- Latent diffusion: an encoder maps data to a lower-dimensional representation ; diffusion operates on , and a decoder maps the result back through .
Latent diffusion reduces computation, but it changes the object being modeled. The diffusion prior learns the encoder's aggregated latent distribution, and generation quality is limited by both the diffusion model and the decoder.
The full picture
A diffusion model is easier to understand as four connected objects:
- A perturbation process defines and gradually turns data into noise
- A neural network predicts the added noise, the clean sample, or the score at each noise level
- A training objective converts reverse-distribution learning into supervised regression on synthetically corrupted data
- A numerical sampler integrates a reverse Markov chain, SDE, or ODE from noise back to data
The on key mathematical identity to keep in mind, that is that Gaussian denoising and score estimation coincide:
From the DDPM perspective, the model learns Gaussian reverse transitions by optimizing a variational bound. From the score perspective, it learns a time-dependent vector field that points noisy samples toward regions of higher probability. From the SDE perspective, that vector field is precisely the missing term required to reverse diffusion.
These are not competing explanations. They are three views of the same generative mechanism.
References
- Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., & Ganguli, S. (2015). Deep Unsupervised Learning using Nonequilibrium Thermodynamics.
- Song, Y., & Ermon, S. (2019). Generative Modeling by Estimating Gradients of the Data Distribution.
- Ho, J., Jain, A., & Abbeel, P. (2020). Denoising Diffusion Probabilistic Models.
- Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., & Poole, B. (2021). Score-Based Generative Modeling through Stochastic Differential Equations.
- Vincent, P. (2011). A Connection Between Score Matching and Denoising Autoencoders.
- Nichol, A. Q., & Dhariwal, P. (2021). Improved Denoising Diffusion Probabilistic Models.
- Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-Resolution Image Synthesis with Latent Diffusion Models.
- Karras, T., Aittala, M., Aila, T., & Laine, S. (2022). Elucidating the Design Space of Diffusion-Based Generative Models.
- Song, Y. (2021). Generative Modeling by Estimating Gradients of the Data Distribution.
- Weng, L. (2021; updated 2024). What Are Diffusion Models?.
- Diffusion_Model_Mathematics_Complete.pdf. Project mathematical notes supplied with this article.
- Master_Thesis-12-30.pdf. Background and related-work chapter supplied with this article.