← Back to Writings

From Noise to Data

A mathematical introduction to diffusion models, from discrete denoising and score matching to stochastic differential equations.

Date
June 1, 2026
›Contents

Diffusion models revolve around a simple idea: take a sample from an unknown data distribution, add noise (usually Gaussian) until its structure disappears and only little of the original signal remain, and train a neural network to undo that corruption. Generation then starts from noise and repeatedly applies the learned reverse transformation. An image, time series, audio signal, or other high-dimensional object gradually surface.

The description is simple. The mathematics is more articulated and connects latent-variable models, variational inference, score matching, stochastic differential equations, and numerical solvers. This article aims to present these connections and core results without proving every detail.

The problem: learning a distribution

Let

x0∼qdata(x)\mathbf{x}_0 \sim q_{\text{data}}(\mathbf{x})

denote an observation drawn from an unknown data distribution. A generative model search for a tractable approximation pθ(x0)p_\theta(\mathbf{x}_0) from which we can draw new samples.

The difficulty is that real data occupy a complicated region of a high-dimensional space. Directly transforming a simple distribution into this geometry can require architectural restrictions, adversarial training, or an intractable normalizing constant. Diffusion models take another indirect route:

  1. Define a fixed process that destroys the data distribution
  2. Learn the reverse of that process
  3. Generate by starting from the simple terminal distribution (prior) and moving backward

Therefore, the forward process is known. Only the reverse process must be learned.

Forward diffusion: progressively destroying structure

A denoising diffusion probabilistic model (DDPM) defines a Markov chain with TT noise steps:

q(x1:T∣x0)=∏t=1Tq(xt∣xt−1),q(\mathbf{x}_{1:T}\mid\mathbf{x}_0) = \prod_{t=1}^{T}q(\mathbf{x}_t\mid\mathbf{x}_{t-1}),

where

q(xt∣xt−1)=N ⁣(xt;1−βt xt−1,βtI).q(\mathbf{x}_t\mid\mathbf{x}_{t-1}) = \mathcal{N}\!\left( \mathbf{x}_t; \sqrt{1-\beta_t}\,\mathbf{x}_{t-1}, \beta_t\mathbf{I} \right).

The schedule βt∈(0,1)\beta_t\in(0,1) controls how much noise is added at step tt. Define

αt:=1−βt,αˉt:=∏s=1tαs.\alpha_t:=1-\beta_t, \qquad \bar\alpha_t:=\prod_{s=1}^{t}\alpha_s.

The transition can then be written through the reparameterization trick:

xt=αt xt−1+1−αt ϵt,ϵt∼N(0,I).\mathbf{x}_t = \sqrt{\alpha_t}\,\mathbf{x}_{t-1} + \sqrt{1-\alpha_t}\,\boldsymbol\epsilon_t, \qquad \boldsymbol\epsilon_t\sim\mathcal{N}(\mathbf{0},\mathbf{I}).

Because a linear combination of independent Gaussian variables remains Gaussian, all preceding steps collapse into one closed-form marginal:

q(xt∣x0)=N ⁣(xt;αˉt x0,(1−αˉt)I).q(\mathbf{x}_t\mid\mathbf{x}_0) = \mathcal{N}\!\left( \mathbf{x}_t; \sqrt{\bar\alpha_t}\,\mathbf{x}_0, (1-\bar\alpha_t)\mathbf{I} \right).

Equivalently,

xt=αˉt x0+1−αˉt ϵ,ϵ∼N(0,I)\mathbf{x}_t = \sqrt{\bar\alpha_t}\,\mathbf{x}_0 + \sqrt{1-\bar\alpha_t}\,\boldsymbol\epsilon, \qquad \boldsymbol\epsilon\sim\mathcal{N}(\mathbf{0},\mathbf{I})

This identity is the key, because it states that training does not require simulating all tt previous transitions, instead we can select any tt, draw one Gaussian noise vector, and construct xt\mathbf{x}_t directly. We are sampling directly from the conditional marginal distribution.

The signal-to-noise ratio at time tt is

SNR⁡(t)=αˉt1−αˉt.\operatorname{SNR}(t) = \frac{\bar\alpha_t}{1-\bar\alpha_t}.

Early steps have high SNR and retain most of the data signal. Late steps have low SNR and are dominated by noise. A useful schedule makes αˉT\bar\alpha_T sufficiently small that

q(xT)≈N(0,I).q(\mathbf{x}_T)\approx\mathcal{N}(\mathbf{0},\mathbf{I}).

Reverse diffusion: learning to denoise

The forward transition gradually removes information, therefore the generation requires to "add" it back through the transition q(xt−1∣xt)q(\mathbf{x}_{t-1}\mid\mathbf{x}_t). Since this, depends on the unknown data distribution and is therefore not directly available. We approximate it with

pθ(xt−1∣xt)=N ⁣(xt−1;μθ(xt,t),Σθ(xt,t)).p_\theta(\mathbf{x}_{t-1}\mid\mathbf{x}_t) = \mathcal{N}\!\left( \mathbf{x}_{t-1}; \boldsymbol\mu_\theta(\mathbf{x}_t,t), \boldsymbol\Sigma_\theta(\mathbf{x}_t,t) \right).

The full generative model is

pθ(x0:T)=p(xT)∏t=1Tpθ(xt−1∣xt),p(xT)=N(0,I).p_\theta(\mathbf{x}_{0:T}) = p(\mathbf{x}_T) \prod_{t=1}^{T} p_\theta(\mathbf{x}_{t-1}\mid\mathbf{x}_t), \qquad p(\mathbf{x}_T)=\mathcal{N}(\mathbf{0},\mathbf{I}).

Although q(xt−1∣xt)q(\mathbf{x}_{t-1}\mid\mathbf{x}_t) is unknown, conditioning additionally on the clean observation makes the posterior tractable:

q(xt−1∣xt,x0)=N ⁣(xt−1;μ~t(xt,x0),β~tI),q(\mathbf{x}_{t-1}\mid\mathbf{x}_t,\mathbf{x}_0) = \mathcal{N}\!\left( \mathbf{x}_{t-1}; \tilde{\boldsymbol\mu}_t(\mathbf{x}_t,\mathbf{x}_0), \tilde\beta_t\mathbf{I} \right),

with

β~t=1−αˉt−11−αˉtβt\tilde\beta_t = \frac{1-\bar\alpha_{t-1}}{1-\bar\alpha_t}\beta_t

and

μ~t=αˉt−1βt1−αˉtx0+αt(1−αˉt−1)1−αˉtxt.\tilde{\boldsymbol\mu}_t = \frac{\sqrt{\bar\alpha_{t-1}}\beta_t}{1-\bar\alpha_t}\mathbf{x}_0 + \frac{\sqrt{\alpha_t}(1-\bar\alpha_{t-1})}{1-\bar\alpha_t}\mathbf{x}_t.

This posterior tells us what the "ideal" reverse mean would be if the clean sample were known. But of course, at generation time it is not known, so the neural network must infer the missing information from xt\mathbf{x}_t and tt.

The variational objective

Diffusion models are latent-variable models whose latent variables are x1:T\mathbf{x}_{1:T}. Their negative log-likelihood is bounded by a variational objective:

−log⁡pθ(x0)≤Eq(x1:T∣x0)[log⁡q(x1:T∣x0)pθ(x0:T)]=:LVLB.-\log p_\theta(\mathbf{x}_0) \leq \mathbb{E}_{q(\mathbf{x}_{1:T}\mid\mathbf{x}_0)} \left[ \log \frac{q(\mathbf{x}_{1:T}\mid\mathbf{x}_0)} {p_\theta(\mathbf{x}_{0:T})} \right] =:\mathcal{L}_{\mathrm{VLB}}.

Using the Markov structure, this variational upper bound can be decomposed as

LVLB=Eq[DKL ⁣(q(xT∣x0)∥p(xT))+∑t=2TDKL ⁣(q(xt−1∣xt,x0)∥pθ(xt−1∣xt))−log⁡pθ(x0∣x1)].\begin{aligned} \mathcal{L}_{\mathrm{VLB}} =\mathbb{E}_q\Bigg[ &D_{\mathrm{KL}}\!\left( q(\mathbf{x}_T\mid\mathbf{x}_0) \parallel p(\mathbf{x}_T) \right)\\ &+\sum_{t=2}^{T} D_{\mathrm{KL}}\!\left( q(\mathbf{x}_{t-1}\mid\mathbf{x}_t,\mathbf{x}_0) \parallel p_\theta(\mathbf{x}_{t-1}\mid\mathbf{x}_t) \right)\\ &-\log p_\theta(\mathbf{x}_0\mid\mathbf{x}_1) \Bigg]. \end{aligned}

The first term matches the terminal distribution to the Gaussian prior. The middle terms teach the learned reverse process to approximate the exact forward posterior. The final term reconstructs the data from the least-corrupted latent.

If the reverse variance is fixed, each middle KL divergence reduces to a weighted squared error between the true posterior mean μ~t\tilde{\boldsymbol\mu}_t and the learned mean μθ\boldsymbol\mu_\theta.

Why the network predicts noise

The posterior mean can be reparameterized using the noise that created xt\mathbf{x}_t. Since

xt=αˉtx0+1−αˉtϵ,\mathbf{x}_t = \sqrt{\bar\alpha_t}\mathbf{x}_0 + \sqrt{1-\bar\alpha_t}\boldsymbol\epsilon,

we can recover the clean sample, given the noise, as

x0=1αˉt(xt−1−αˉtϵ).\mathbf{x}_0 = \frac{1}{\sqrt{\bar\alpha_t}} \left( \mathbf{x}_t- \sqrt{1-\bar\alpha_t}\boldsymbol\epsilon \right).

Substitution into μ~t\tilde{\boldsymbol\mu}_t gives

μ~t=1αt(xt−βt1−αˉtϵ).\tilde{\boldsymbol\mu}_t = \frac{1}{\sqrt{\alpha_t}} \left( \mathbf{x}_t - \frac{\beta_t}{\sqrt{1-\bar\alpha_t}} \boldsymbol\epsilon \right).

The reverse mean can therefore be parameterized by a neural network ϵθ\boldsymbol\epsilon_\theta that predicts the injected noise:

μθ(xt,t)=1αt(xt−βt1−αˉtϵθ(xt,t)).\boldsymbol\mu_\theta(\mathbf{x}_t,t) = \frac{1}{\sqrt{\alpha_t}} \left( \mathbf{x}_t - \frac{\beta_t}{\sqrt{1-\bar\alpha_t}} \boldsymbol\epsilon_\theta(\mathbf{x}_t,t) \right).

With a fixed reverse variance σt,rev2I\sigma_{t,\mathrm{rev}}^2\mathbf{I}, the corresponding variational term is proportional to

Lt=Ex0,ϵ[βt22σt,rev2αt(1−αˉt)∥ϵ−ϵθ(xt,t)∥22].\mathcal{L}_t = \mathbb{E}_{\mathbf{x}_0,\boldsymbol\epsilon} \left[ \frac{\beta_t^2} {2\sigma_{t,\mathrm{rev}}^2\alpha_t(1-\bar\alpha_t)} \left\| \boldsymbol\epsilon- \boldsymbol\epsilon_\theta(\mathbf{x}_t,t) \right\|_2^2 \right].

Ho, Jain, and Abbeel found that removing this time-dependent coefficient often improves sample quality. This produces the now-standard simplified loss:

Lsimple=Et,x0,ϵ[∥ϵ−ϵθ ⁣(αˉtx0+1−αˉtϵ,t)∥22]\mathcal{L}_{\mathrm{simple}} = \mathbb{E}_{ t,\mathbf{x}_0,\boldsymbol\epsilon } \left[ \left\| \boldsymbol\epsilon - \boldsymbol\epsilon_\theta\!\left( \sqrt{\bar\alpha_t}\mathbf{x}_0 + \sqrt{1-\bar\alpha_t}\boldsymbol\epsilon, t \right) \right\|_2^2 \right]

where tt is usually sampled uniformly from {1,…,T}\{1,\ldots,T\}. A complete training observation can therefore be generated using one clean sample, one timestep, and one Gaussian noise draw.

The score: a vector field toward probable data

The DDPM derivation explains the training objective, while Score Matching (SM) explains what the network actually learns geometrically.

Suppose a density is represented as an energy-based model:

pθ(x)=e−fθ(x)Zθ.p_\theta(\mathbf{x}) = \frac{e^{-f_\theta(\mathbf{x})}}{Z_\theta}.

Its score is the gradient of its log-density with respect to the data:

s(x):=∇xlog⁡p(x).\mathbf{s}(\mathbf{x}) := \nabla_\mathbf{x}\log p(\mathbf{x}).

For the energy-based model,

∇xlog⁡pθ(x)=−∇xfθ(x)−∇xlog⁡Zθ⏟0.\nabla_\mathbf{x}\log p_\theta(\mathbf{x}) = -\nabla_\mathbf{x}f_\theta(\mathbf{x}) -\underbrace{\nabla_\mathbf{x}\log Z_\theta}_{0}.

The unknown normalizing constant disappears. The score does not state how probable a point is, but it gives the local direction in which log-density increases the fastest.

A score network sθ(x)\mathbf{s}_\theta(\mathbf{x}) could be trained by minimizing the Fisher divergence

JF(θ)=12Epdata(x)[∥sθ(x)−∇xlog⁡pdata(x)∥22].\mathcal{J}_{F}(\theta) = \frac{1}{2} \mathbb{E}_{p_{\text{data}}(\mathbf{x})} \left[ \left\| \mathbf{s}_\theta(\mathbf{x}) - \nabla_\mathbf{x}\log p_{\text{data}}(\mathbf{x}) \right\|_2^2 \right].

The true data score is unknown. Under boundary conditions, classical score matching rewrites the objective, up to a constant independent of θ\theta, as

JSM(θ)=Epdata[12∥sθ(x)∥22+∇x⋅sθ(x)].\mathcal{J}_{\mathrm{SM}}(\theta) = \mathbb{E}_{p_{\text{data}}} \left[ \frac{1}{2}\|\mathbf{s}_\theta(\mathbf{x})\|_2^2 + \nabla_\mathbf{x}\cdot\mathbf{s}_\theta(\mathbf{x}) \right].

This avoids the unknown data score but introduces the divergence of the network (∇x)(\nabla_{\mathbf{x}}), which is expensive in high dimensions, since it requires the computation of the Jacobian matrix. Denoising Score Matching (DSM) gives a more convenient target.

Denoising score matching

Perturb a clean sample with Gaussian noise:

x~=x+σϵ,ϵ∼N(0,I).\tilde{\mathbf{x}} = \mathbf{x}+\sigma\boldsymbol\epsilon, \qquad \boldsymbol\epsilon\sim\mathcal{N}(\mathbf{0},\mathbf{I}).

The conditional perturbation kernel is known:

qσ(x~∣x)=N(x~;x,σ2I),q_\sigma(\tilde{\mathbf{x}}\mid\mathbf{x}) = \mathcal{N}(\tilde{\mathbf{x}};\mathbf{x},\sigma^2\mathbf{I}),

so its conditional score is:

∇x~log⁡qσ(x~∣x)=−x~−xσ2=−ϵσ.\nabla_{\tilde{\mathbf{x}}} \log q_\sigma(\tilde{\mathbf{x}}\mid\mathbf{x}) = -\frac{\tilde{\mathbf{x}}-\mathbf{x}}{\sigma^2} = -\frac{\boldsymbol\epsilon}{\sigma}.

DSM trains against this known target:

LDSM=Ex,x~[λ(σ)∥sθ(x~,σ)+x~−xσ2∥22].\mathcal{L}_{\mathrm{DSM}} = \mathbb{E}_{\mathbf{x},\tilde{\mathbf{x}}} \left[ \lambda(\sigma) \left\| \mathbf{s}_\theta(\tilde{\mathbf{x}},\sigma) + \frac{\tilde{\mathbf{x}}-\mathbf{x}}{\sigma^2} \right\|_2^2 \right].

At its optimum, this objective estimates the score of the marginal perturbed distribution, not merely the score of a Gaussian centred on one training example:

sθ(x~,σ)≈∇x~log⁡qσ(x~).\mathbf{s}_\theta(\tilde{\mathbf{x}},\sigma) \approx \nabla_{\tilde{\mathbf{x}}}\log q_\sigma(\tilde{\mathbf{x}}).

Training over many noise scales is essential, since at low noise, the model learns fine structure near the data manifold, while at high noise, separated modes overlap and the score remains significant in regions that would otherwise contain no training data.

For the DDPM perturbation kernel, the standard deviation is 1−αˉt\sqrt{1-\bar\alpha_t}. Therefore,

∇xtlog⁡q(xt∣x0)=−ϵ1−αˉt,\nabla_{\mathbf{x}_t}\log q(\mathbf{x}_t\mid\mathbf{x}_0) = -\frac{\boldsymbol\epsilon}{\sqrt{1-\bar\alpha_t}},

which gives the central equivalence

sθ(xt,t)≈−ϵθ(xt,t)1−αˉt\mathbf{s}_\theta(\mathbf{x}_t,t) \approx -\frac{ \boldsymbol\epsilon_\theta(\mathbf{x}_t,t) }{\sqrt{1-\bar\alpha_t}}

Predicting noise and estimating the score are the same task up to a known scale.

Sampling in discrete time

After training, sampling starts from

xT∼N(0,I)\mathbf{x}_T\sim\mathcal{N}(\mathbf{0},\mathbf{I})

and applies the learned reverse transition for t=T,T−1,…,1t=T,T-1,\ldots,1. With the noise parameterization,

xt−1=1αt(xt−βt1−αˉtϵθ(xt,t))+σt,revz,\mathbf{x}_{t-1} = \frac{1}{\sqrt{\alpha_t}} \left( \mathbf{x}_t - \frac{\beta_t}{\sqrt{1-\bar\alpha_t}} \boldsymbol\epsilon_\theta(\mathbf{x}_t,t) \right) + \sigma_{t,\mathrm{rev}}\mathbf{z},

where

z∼N(0,I)\mathbf{z}\sim\mathcal{N}(\mathbf{0},\mathbf{I})

for t>1t>1, and z=0\mathbf{z}=\mathbf{0} at the final step. Common fixed choices are σt,rev2=βt\sigma_{t,\mathrm{rev}}^2=\beta_t or β~t\tilde\beta_t.

Each evaluation estimates a small amount of noise to be removed. The process is stable because no single network call must map from the prior vector directly to a clean sample. On the other side, its main weakness is computational: ancestral DDPM sampling can require hundreds or thousands of sequential network evaluations, quick is slow and expensive.

Continuous time: diffusion as an SDE

The discrete chain becomes a stochastic differential equation when the steps become infinitesimal. Let x(t)∈Rd\mathbf{x}(t)\in\mathbb{R}^d, with t∈[0,T]t\in[0,T], follow

dx=f(x,t) dt+g(t) dWt,\mathrm{d}\mathbf{x} = \mathbf{f}(\mathbf{x},t)\,\mathrm{d}t + g(t)\,\mathrm{d}\mathbf{W}_t,

where f\mathbf{f} is the drift, gg is the diffusion coefficient, and Wt\mathbf{W}_t is standard Brownian motion. The forward SDE transforms the data density p0p_0 into a tractable terminal density pTp_T.

Different choices recover important diffusion families:

FamilyForward SDEBehaviour
Variance preserving (VP)dx=−12β(t)x dt+β(t) dWt\mathrm{d}\mathbf{x}=-\frac{1}{2}\beta(t)\mathbf{x}\,\mathrm{d}t+\sqrt{\beta(t)}\,\mathrm{d}\mathbf{W}_tContinuous analogue of DDPM; unit Gaussian variance
Variance exploding (VE)dx=d[σ2(t)]dt dWt\mathrm{d}\mathbf{x}=\sqrt{\frac{\mathrm{d}[\sigma^2(t)]}{\mathrm{d}t}}\,\mathrm{d}\mathbf{W}_tAdds noise without shrinking the signal; variance grows with tt
sub-VPdx=−12β(t)x dt+β(t)(1−e−2B(t)) dWt\mathrm{d}\mathbf{x}=-\frac{1}{2}\beta(t)\mathbf{x}\,\mathrm{d}t+\sqrt{\beta(t)(1-e^{-2B(t)})}\,\mathrm{d}\mathbf{W}_tUses less diffusion than VP and is useful for likelihood-oriented models

Here

B(t):=∫0tβ(s) ds.B(t):=\int_0^t\beta(s)\,\mathrm{d}s.

For the VP SDE, the transition kernel remains Gaussian:

p0t(x(t)∣x(0))=N ⁣(e−B(t)/2x(0),(1−e−B(t))I).p_{0t}(\mathbf{x}(t)\mid\mathbf{x}(0)) = \mathcal{N}\!\left( e^{-B(t)/2}\mathbf{x}(0), \left(1-e^{-B(t)}\right)\mathbf{I} \right).

Thus

x(t)=e−B(t)/2x(0)+1−e−B(t) ϵ.\mathbf{x}(t) = e^{-B(t)/2}\mathbf{x}(0) + \sqrt{1-e^{-B(t)}}\,\boldsymbol\epsilon.

This is the continuous counterpart of the DDPM marginal, with e−B(t)e^{-B(t)} playing the role of αˉt\bar\alpha_t.

The reverse-time SDE

A forward diffusion has a corresponding reverse-time diffusion. When integrated from TT to 00, it satisfies

dx=[f(x,t)−g2(t)∇xlog⁡pt(x)]dt+g(t) dWˉt\mathrm{d}\mathbf{x} = \left[ \mathbf{f}(\mathbf{x},t) - g^2(t)\nabla_\mathbf{x}\log p_t(\mathbf{x}) \right]\mathrm{d}t + g(t)\,\mathrm{d}\bar{\mathbf{W}}_t

where dt<0\mathrm{d}t<0 during reverse integration and Wˉt\bar{\mathbf{W}}_t is Brownian motion in reverse time.

Only one quantity is unknown: the time-dependent score ∇xlog⁡pt(x)\nabla_\mathbf{x}\log p_t(\mathbf{x}). Replacing it with the neural approximation sθ(x,t)\mathbf{s}_\theta(\mathbf{x},t) results in the learned generative process:

dx=[f(x,t)−g2(t)sθ(x,t)]dt+g(t) dWˉt.\mathrm{d}\mathbf{x} = \left[ \mathbf{f}(\mathbf{x},t) - g^2(t)\mathbf{s}_\theta(\mathbf{x},t) \right]\mathrm{d}t + g(t)\,\mathrm{d}\bar{\mathbf{W}}_t.

This equation says that the reverse dynamics combine the known physical drift, a learned score correction, and stochastic noise.

The continuous denoising-score objective is

L(θ)=Et∼U(0,T)Ex(0)Ex(t)∣x(0)[λ(t)∥sθ(x(t),t)−∇x(t)log⁡p0t(x(t)∣x(0))∥22].\mathcal{L}(\theta) = \mathbb{E}_{t\sim\mathcal{U}(0,T)} \mathbb{E}_{\mathbf{x}(0)} \mathbb{E}_{\mathbf{x}(t)\mid\mathbf{x}(0)} \left[ \lambda(t) \left\| \mathbf{s}_\theta(\mathbf{x}(t),t) - \nabla_{\mathbf{x}(t)} \log p_{0t}(\mathbf{x}(t)\mid\mathbf{x}(0)) \right\|_2^2 \right].

The weighting λ(t)\lambda(t) determines which noise regions dominate training. A common choice is proportional to the inverse expected squared norm of the conditional score. But also the choice λ(t)=g2(t)\lambda(t)=g^2(t) has a likelihood interpretation: under regularity conditions, the weighted score error is related to the upper bound on the KL divergence between the data distribution and the distribution induced by the reverse SDE.

A deterministic path: the probability-flow ODE

Interestingly, through Fokker-Plank equation can be shown that the same marginal densities {pt}t∈[0,T]\{p_t\}_{t\in[0,T]} can be generated by the deterministic ordinary differential equation

dxdt=f(x,t)−12g2(t)∇xlog⁡pt(x)\frac{\mathrm{d}\mathbf{x}}{\mathrm{d}t} = \mathbf{f}(\mathbf{x},t) - \frac{1}{2}g^2(t) \nabla_\mathbf{x}\log p_t(\mathbf{x})

or, after replacing the true score,

dxdt=vθ(x,t):=f(x,t)−12g2(t)sθ(x,t).\frac{\mathrm{d}\mathbf{x}}{\mathrm{d}t} = \mathbf{v}_\theta(\mathbf{x},t) := \mathbf{f}(\mathbf{x},t) - \frac{1}{2}g^2(t)\mathbf{s}_\theta(\mathbf{x},t).

The reverse SDE and probability-flow ODE do not produce identical trajectories. One is stochastic and one is deterministic. They do, however, share the same marginal distribution at every time when the score is exact.

The ODE also enables likelihood evaluation through the instantaneous change-of-variables formula:

ddtlog⁡pt(x(t))=−∇x⋅vθ(x(t),t).\frac{\mathrm{d}}{\mathrm{d}t} \log p_t(\mathbf{x}(t)) = -\nabla_\mathbf{x}\cdot \mathbf{v}_\theta(\mathbf{x}(t),t).

Consequently,

log⁡p0(x(0))=log⁡pT(x(T))+∫0T∇x⋅vθ(x(t),t) dt.\log p_0(\mathbf{x}(0)) = \log p_T(\mathbf{x}(T)) + \int_0^T \nabla_\mathbf{x}\cdot \mathbf{v}_\theta(\mathbf{x}(t),t) \,\mathrm{d}t.

Computing the divergence exactly requires the trace of a large Jacobian. Hutchinson's identity replaces it with an unbiased stochastic estimate:

Tr⁡(J)=Eξ[ξ⊤Jξ],E[ξξ⊤]=I.\operatorname{Tr}(\mathbf{J}) = \mathbb{E}_{\boldsymbol\xi} \left[ \boldsymbol\xi^\top\mathbf{J}\boldsymbol\xi \right], \qquad \mathbb{E}[\boldsymbol\xi\boldsymbol\xi^\top]=\mathbf{I}.

Automatic differentiation can evaluate the required vector-Jacobian product without materializing the full Jacobian.

What the neural network actually does

The mathematical framework does not prescribe a single architecture. The network only needs to map a noisy object and its noise level to an output of the same shape:

ϵθ:(xt,t,c)↦ϵ^,\boldsymbol\epsilon_\theta: (\mathbf{x}_t,t,\mathbf{c}) \mapsto \widehat{\boldsymbol\epsilon},

where c\mathbf{c} is optional conditioning information.

For images, a U-Net is a natural choice because its encoder captures large-scale structure while skip connections preserve local detail. The timestep is typically transformed into a sinusoidal embedding and injected throughout the network. Self-attention adds long-range interactions. For sequences and other structured data, transformers or specialized temporal architectures can replace the U-Net backbone without changing the diffusion mathematics.

Given the equivalent parameterizations we can predict different but related targets:

TargetNetwork outputConversion to the score
Noiseϵ^\widehat{\boldsymbol\epsilon}sθ=−ϵ^/1−αˉt\mathbf{s}_\theta=-\widehat{\boldsymbol\epsilon}/\sqrt{1-\bar\alpha_t}
Clean samplex^0\widehat{\mathbf{x}}_0sθ=(αˉtx^0−xt)/(1−αˉt)\mathbf{s}_\theta=(\sqrt{\bar\alpha_t}\widehat{\mathbf{x}}_0-\mathbf{x}_t)/(1-\bar\alpha_t)
Scores^\widehat{\mathbf{s}}Direct output

They encode the same "ideal" quantity, but at different scales across noise levels. The target, loss weighting, input scaling, noise schedule, and sampler therefore influence in practice the output. The EDM formulation makes this separation explicit by writing a denoiser as

Dθ(x;σ)=cskip(σ)x+cout(σ)Fθ ⁣(cin(σ)x;cnoise(σ)),D_\theta(\mathbf{x};\sigma) = c_{\mathrm{skip}}(\sigma)\mathbf{x} + c_{\mathrm{out}}(\sigma) F_\theta\!\left( c_{\mathrm{in}}(\sigma)\mathbf{x}; c_{\mathrm{noise}}(\sigma) \right),

where the coefficients keep network inputs and targets well-scaled across noise levels.

Conditioning and guidance

For a condition c\mathbf{c}, such as a class label, text embedding, or observed portion of a time series, the desired distribution is pt(x∣c)p_t(\mathbf{x}\mid\mathbf{c}). Bayes' rule gives

∇xlog⁡pt(x∣c)=∇xlog⁡pt(x)+∇xlog⁡pt(c∣x).\nabla_\mathbf{x}\log p_t(\mathbf{x}\mid\mathbf{c}) = \nabla_\mathbf{x}\log p_t(\mathbf{x}) + \nabla_\mathbf{x}\log p_t(\mathbf{c}\mid\mathbf{x}).

Classifier guidance (CG) estimates the second term with a classifier trained on noisy inputs. Classifier-free guidance (CFG) instead trains only one network both conditionally and unconditionally by sometimes replacing c\mathbf{c} with a null condition ∅\varnothing. At inference, the predictions are combined as

sCFG=sθ(xt,t,∅)+w[sθ(xt,t,c)−sθ(xt,t,∅)].\mathbf{s}_{\mathrm{CFG}} = \mathbf{s}_\theta(\mathbf{x}_t,t,\varnothing) + w\left[ \mathbf{s}_\theta(\mathbf{x}_t,t,\mathbf{c}) - \mathbf{s}_\theta(\mathbf{x}_t,t,\varnothing) \right].

At w=1w=1, this recovers the conditional prediction. Values w>1w>1 amplify the condition, often improving adherence at the cost of diversity and sometimes realism.

Faster and cheaper diffusion

The original reverse chain is expensive because its steps are sequential. Several approaches reduce that cost:

  • Non-Markovian or ODE samplers: DDIM and probability-flow methods can follow deterministic paths with fewer steps
  • Higher-order solvers: methods such as Heun integration and DPM-Solver approximate the reverse dynamics more accurately per network evaluation
  • Latent diffusion: an encoder maps data to a lower-dimensional representation z=E(x)\mathbf{z}=E(\mathbf{x}); diffusion operates on z\mathbf{z}, and a decoder maps the result back through x^=D(z)\widehat{\mathbf{x}}=D(\mathbf{z}).

Latent diffusion reduces computation, but it changes the object being modeled. The diffusion prior learns the encoder's aggregated latent distribution, and generation quality is limited by both the diffusion model and the decoder.

The full picture

A diffusion model is easier to understand as four connected objects:

  1. A perturbation process defines q(xt∣x0)q(\mathbf{x}_t\mid\mathbf{x}_0) and gradually turns data into noise
  2. A neural network predicts the added noise, the clean sample, or the score at each noise level
  3. A training objective converts reverse-distribution learning into supervised regression on synthetically corrupted data
  4. A numerical sampler integrates a reverse Markov chain, SDE, or ODE from noise back to data

The on key mathematical identity to keep in mind, that is that Gaussian denoising and score estimation coincide:

ϵθ(xt,t)⟺sθ(xt,t)=−ϵθ(xt,t)1−αˉt.\boldsymbol\epsilon_\theta(\mathbf{x}_t,t) \quad\Longleftrightarrow\quad \mathbf{s}_\theta(\mathbf{x}_t,t) = -\frac{\boldsymbol\epsilon_\theta(\mathbf{x}_t,t)} {\sqrt{1-\bar\alpha_t}}.

From the DDPM perspective, the model learns Gaussian reverse transitions by optimizing a variational bound. From the score perspective, it learns a time-dependent vector field that points noisy samples toward regions of higher probability. From the SDE perspective, that vector field is precisely the missing term required to reverse diffusion.

These are not competing explanations. They are three views of the same generative mechanism.

References

  1. Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., & Ganguli, S. (2015). Deep Unsupervised Learning using Nonequilibrium Thermodynamics.
  2. Song, Y., & Ermon, S. (2019). Generative Modeling by Estimating Gradients of the Data Distribution.
  3. Ho, J., Jain, A., & Abbeel, P. (2020). Denoising Diffusion Probabilistic Models.
  4. Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., & Poole, B. (2021). Score-Based Generative Modeling through Stochastic Differential Equations.
  5. Vincent, P. (2011). A Connection Between Score Matching and Denoising Autoencoders.
  6. Nichol, A. Q., & Dhariwal, P. (2021). Improved Denoising Diffusion Probabilistic Models.
  7. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-Resolution Image Synthesis with Latent Diffusion Models.
  8. Karras, T., Aittala, M., Aila, T., & Laine, S. (2022). Elucidating the Design Space of Diffusion-Based Generative Models.
  9. Song, Y. (2021). Generative Modeling by Estimating Gradients of the Data Distribution.
  10. Weng, L. (2021; updated 2024). What Are Diffusion Models?.
  11. Diffusion_Model_Mathematics_Complete.pdf. Project mathematical notes supplied with this article.
  12. Master_Thesis-12-30.pdf. Background and related-work chapter supplied with this article.