← Back to Projects

Conditional Latent Diffusion with SDEs

Conditional face generation with continuous-time SDEs, latent diffusion, and classifier-free guidance.

Date
February 24, 2026
Collaborators
Giacomo Negri & Marco Zimmatore
Links
›Contents

Overview

Generating an image from noise is only half of the problem. The harder question is whether the generation process can be controlled without training a separate classifier for every desired attribute.

For the Hands-on Generative AI course at the Technical University of Munich, Marco Zimmatore and I built a conditional latent diffusion model from scratch. The project combines the continuous-time score-based framework introduced by Song et al. with a pretrained Variational Autoencoder and Classifier-Free Guidance (CFG).

The model trained by us on limited resources generates faces from the CelebA dataset while conditioning the reverse diffusion process on attributes such as age, expression, hair, and facial features. Rather than diffusing full-resolution images, it learns the score of their compressed latent representations.

The project compared:

  • Variance Preserving (VP), Variance Exploding (VE), and subVP SDEs;
  • Standard score matching, Likelihood Weighting (LW), and LW with Importance Sampling (LW + IS);
  • CFG scales from 1.5 to 7.5;
  • generative quality through FID, Inception Score, and latent-space NLL.

The best FID was 55.55 with VP and LW + IS, nearly identical to the VP baseline. Increasing the CFG scale strengthened the requested attributes, but progressively damaged perceptual fidelity: for the subVP baseline, IS rose from 3.23 to 4.09, while FID deteriorated from 58.38 to 77.86.

Reverse denoising process

Reverse diffusion sampling for VP with Likelihood Weighting + Importance Sampling in latent space.

Denoised images

. Decoded output shown as a fixed reference.


The problem

Score-based diffusion models learn to reverse a process that gradually destroys the structure of the data. In pixel space, however, the network must model every visual detail at every diffusion step, making training expensive.

The project asked three related questions:

Can a compact latent diffusion model reproduce the continuous-time SDE framework, generate faces conditionally through CFG, and remain effective under limited computational resources?

Moving the process into latent space reduces the dimensionality of the learning problem, but it also creates a new dependency: the diffusion model must generate latents that remain compatible with a separately trained decoder.

This tension between computational efficiency and decoder compatibility represents one of the main findings of our project.


From images to latent diffusion

We used the pretrained stabilityai/sd-vae-ft-mse VAE to map an image xx into a lower-dimensional latent representation z0z_0:

z0∼qϕ(z∣x),z0∈Rd,d≪D.z_0 \sim q_\phi(z \mid x), \qquad z_0 \in \mathbb{R}^d, \quad d \ll D.

The diffusion process operates on z0z_0, not directly on the image. After sampling, the VAE decoder maps the denoised latent back into pixel space:

x→encoderz0→forward SDEzt→reverse SDEz^0→decoderx^.x \xrightarrow{\text{encoder}} z_0 \xrightarrow{\text{forward SDE}} z_t \xrightarrow{\text{reverse SDE}} \hat z_0 \xrightarrow{\text{decoder}} \hat x.
ComponentRole
Pretrained VAECompresses CelebA images and decodes generated latents
Continuous-time SDEDefines how latent representations are progressively noised
Conditional U-NetEstimates the time-dependent score, or equivalently the added noise
Classifier-Free GuidanceControls generation using CelebA attribute labels without a separate classifier
Euler-Maruyama solverNumerically integrates the reverse-time SDE
Latent Diffusion Model

The architecture of the latent diffusion model (LDM). Source: Rombach et al., "High-Resolution Image Synthesis with Latent Diffusion Models," 2021.


The SDE formulation

Following Song et al., the forward diffusion process is written as an Itô SDE:

dzt=f(zt,t) dt+g(t) dwt,d z_t = f(z_t,t)\,dt + g(t)\,d w_t,

where ff is the drift, gg controls the injected noise, and wtw_t is a standard Wiener process. At t=0t=0, z0z_0 follows the latent data distribution; by t=Tt=T, it approaches a tractable prior such as N(0,I)\mathcal{N}(0,I).

Generation relies on the reverse-time SDE:

dzt=[f(zt,t)−g(t)2∇zlog⁡pt(zt)]dt+g(t) dwˉt.d z_t = \left[f(z_t,t)-g(t)^2\nabla_z \log p_t(z_t)\right]dt + g(t)\,d\bar w_t.

The unknown score ∇zlog⁡pt(zt)\nabla_z \log p_t(z_t) is approximated by the U-Net sθ(zt,t,c)s_\theta(z_t,t,c). With a Gaussian perturbation kernel,

zt=μ(t)z0+σ(t)ϵ,ϵ∼N(0,I),z_t = \mu(t)z_0 + \sigma(t)\epsilon, \qquad \epsilon \sim \mathcal{N}(0,I),

score estimation can be implemented as noise prediction:

Lsimple=Ez0,ϵ,t[∥ϵ−ϵθ(zt,t,c)∥22].\mathcal{L}_{\text{simple}} = \mathbb{E}_{z_0,\epsilon,t} \left[\left\|\epsilon-\epsilon_\theta(z_t,t,c)\right\|_2^2\right].

We implemented three continuous-time processes:

SDEMain propertyObserved behaviour
VPVariance remains bounded throughout diffusionBest overall FID and stable training
subVPReduced diffusion variance relative to VPCompetitive results and strongest response to CFG
VENoise variance grows without a drift termHighest expressivity in theory, but slowest convergence in our constrained setup

Conditional generation with CFG

During training, the attribute vector cc is randomly replaced with a null condition ∅\varnothing. The same network therefore learns both the conditional and unconditional scores.

At inference, the two estimates are combined as:

sCFG(zt,t,c)=sθ(zt,t,∅)+ω[sθ(zt,t,c)−sθ(zt,t,∅)],s_{\text{CFG}}(z_t,t,c) = s_\theta(z_t,t,\varnothing) + \omega\left[s_\theta(z_t,t,c)-s_\theta(z_t,t,\varnothing)\right],

where ω\omega is the guidance scale. The difference between the conditional and unconditional scores points sampling toward latents associated with the requested attributes.

The conditioning vector and the continuous diffusion time are embedded separately, combined through an MLP, and injected into the U-Net residual blocks using Feature-wise Linear Modulation:

hcond=h⊙(1+γ(e))+β(e).h_{\text{cond}} = h \odot \left(1 + \gamma(e)\right) + \beta(e).

This allows the same denoising network to adapt both to the current noise level and to the requested facial characteristics.


Architecture and training

The score network is a symmetric encoder-decoder U-Net built from residual blocks, learned downsampling, nearest-neighbour upsampling, skip connections, and self-attention at the lowest spatial resolutions.

SettingValue
DatasetCelebA
Training imagesApproximately 60,000
Training iterationsApproximately 200,000
Batch size16
U-Net parametersApproximately 12.9 million
Reverse sampling steps1,000
Generated samples for FID and IS10,000
Held-out latents for NLL1,000

Group Normalization was used instead of Batch Normalization because GPU memory constraints required small batches. Self-attention was restricted to the deepest feature maps, where global context is useful and the quadratic attention cost is lower.

The project was implemented as a modular PyTorch Lightning pipeline with YAML experiment configurations and Weights & Biases tracking. Separate modules handle the U-Net, SDE definitions, forward and reverse processes, VAE interface, training, and evaluation.

Latent Diffusion Model

Architectural Overview of the Conditional U-Net.


Results across SDEs and objectives

FID and IS were calculated from 10,000 decoded samples. NLL was estimated on 1,000 held-out images through the probability-flow ODE, using the Hutchinson trace estimator.

Because the VAE was fixed, the reported NLL measures only the latent prior-matching term, not the full pixel-space reconstruction likelihood. It should therefore be interpreted within this project rather than compared directly with pixel-space benchmarks.

SDEObjectiveGuidanceFID ↓IS ↑Latent NLL ↓
subVPBase1.558.383.2325.00±1425.00 \pm 14
subVPLW1.567.163.1129.16±1129.16 \pm 11
subVPLW + IS1.556.702.8626.78±1226.78 \pm 12
VPBase1.555.562.9025.60±1325.60 \pm 13
VPLW + IS1.555.552.9025.60±1325.60 \pm 13
VPLW1.569.593.0030.44±1230.44 \pm 12
VEBase1.577.323.0925.20±1425.20 \pm 14

Three results stand out:

  • VP produced the best perceptual fidelity. Its baseline and LW + IS variants were effectively tied on FID.
  • Likelihood Weighting alone consistently degraded FID. In latent space, emphasizing all noise levels may force the model to spend capacity on high-frequency artifacts inherited from the fixed VAE.
  • VE underperformed despite its theoretical expressivity. Its slower convergence was a poor match for the available training budget.
Latent Diffusion Model

Denoising process with 1000 reverse steps for VP trained with LW + IS and 1.5 guidance (left), subVP trained with LW + IS and 1.5 guidance (center), and VE with 1.5 guidance (right). Reconstruction started from N (0, I) for VP and subVP, and N (0, σmax2σ^2_{max} = 50.0) for VE


The guidance trade-off

Changing only the CFG scale produced a clear and monotonic trade-off for the subVP baseline:

Guidance scale ω\omegaFID ↓IS ↑
1.558.383.23
3.562.163.54
5.067.343.78
7.577.864.09

Higher guidance made attributes more distinct, which increased IS. At the same time, FID worsened at every step, indicating that generated images moved further from the real CelebA distribution.

This is not a contradiction between the metrics. IS rewards recognizable and diverse class features, while FID penalizes a shift in the overall generated distribution. Strong guidance can improve the first while damaging the second.


Latent-decoder misalignment

One significant limitation was not simply insufficient training time. It came from separating the latent generator from the fixed VAE decoder.

Let the VAE aggregated posterior be

qϕ(z)=∫qϕ(z∣x)pdata(x) dx.q_\phi(z)=\int q_\phi(z\mid x)p_{\text{data}}(x)\,dx.

The decoder is trained primarily on latents from high-density regions of qϕ(z)q_\phi(z). The diffusion model, however, samples from an approximation pθ(z)p_\theta(z). A latent can be plausible under the learned diffusion geometry while still lying outside the region on which the decoder learned a reliable inverse mapping.

Strong CFG may makes this failure mode more likely, since guidance extrapolates the conditional score field, large values of ω\omega could push the trajectory into extreme conditional regions. When these latents are decoded, the result may contain distorted structure, incorrect colours, washed-out textures, or other artifacts.

This explains why the model could show a stronger conditional signal while receiving a worse FID score. The issue was not only what the diffusion prior learned, but whether the fixed decoder could interpret the latents it generated.


Limitations

The results should be read in the context of a course project with constrained compute.

  • Training used approximately 60,000 CelebA images, a batch size of 16, and around 200,000 iterations
  • The VAE was pretrained and fixed, preventing the encoder, decoder, and diffusion prior from adapting jointly
  • Euler-Maruyama required 1,000 reverse steps, making generation slow and allowing numerical error to accumulate
  • FID and IS do not directly measure whether every requested attribute is respected
  • The latent NLL excludes the VAE reconstruction term and is not directly comparable with full image-space likelihoods
  • Results from Song et al. use different data, architectures, and pixel-space training, so they provide context rather than a controlled benchmark

The most direct extensions would be to train or fine-tune the autoencoder jointly, evaluate explicit attribute accuracy, use higher-order samplers, test larger training budgets, and regularize guided sampling toward the typical latent region.


Conclusion

This project reproduced the core techniques and tools of score-based diffusion in latent space: continuous-time SDEs, a conditional U-Net, numerical reverse diffusion, and classifier-free guidance.

It identifies also how, with constrained training set up conditioning, CFG, and VAEs can influence the reconstruction process and its outputs.