Generative Models
8 questions
Explain GAN architecture - generator and discriminator objectives.▶
A GAN frames image generation as a two-player minimax game.
Generator G: takes random noise z ~ p(z) → generates fake sample G(z). Goal: fool D into thinking G(z) is real.
Discriminator D: binary classifier. Takes sample x → P(real). Goal: correctly classify real vs. fake.
Minimax objective:
minG maxD E[log D(x)] + E[log(1 - D(G(z)))]
- D maximizes: D(x)→1 for real, D(G(z))→0 for fake
- G minimizes: make D(G(z))→1. In practice: maximize log D(G(z)) (not minimize log(1-D(G(z)))) for stronger gradients early in training
Training: alternate D steps and G steps. Train D to optimality → update G → repeat.
At convergence (theory): G matches real data distribution, D outputs 0.5 everywhere. In practice: mode collapse and training instability prevent clean convergence.
Explain VAE - the ELBO objective and the reparameterization trick.frontier▶
VAE learns to compress data x into latent z and reconstruct it. We want to maximize log p(x) = log ∫ p(x|z)p(z)dz.
This is intractable. Instead, maximize the ELBO (Evidence Lower BOund):
L = Eq(z|x)[log p(x|z)] - KL(q(z|x) || p(z))
- Reconstruction term: decoder p(x|z) must reconstruct x from z. Maximizing log p(x|z) = minimizing reconstruction loss (MSE or BCE).
- KL term: encoder q(z|x) ~ N(μ(x), σ²(x)) must stay close to prior p(z) = N(0,I). Prevents latent space from being arbitrary. Has closed form: -½ Σ(1 + log σ² - μ² - σ²).
Reparameterization trick: can't backprop through sampling z ~ q(z|x) = N(μ, σ). Instead: z = μ + σ⊙ε where ε ~ N(0,I). Now the stochasticity is in ε (no parameters), and gradients flow through μ and σ.
What is mode collapse in GANs? How do you fix it?▶
The generator learns to produce only one or a few modes of the data distribution, even if the real data has many modes. The discriminator can't distinguish, so the generator stops diversifying.
Why it happens: the generator finds a "safe" region where the discriminator is fooled. Gradient signals stop directing toward unexplored modes.
Fixes:
- WGAN (Wasserstein GAN): replace JS divergence with Wasserstein distance. Better gradient landscapes, doesn't saturate.
- Minibatch discrimination: discriminator sees statistics across the batch, not just individual samples - penalizes low diversity.
- Mode-seeking GAN: explicitly maximize distance between generated samples (diversity loss).
- Spectral normalization: normalize discriminator weights to enforce Lipschitz constraint → stable training.
- StyleGAN / progressive growing: architectural improvements that stabilize training.
Explain diffusion models - forward process, reverse process, training.frontier▶
Forward process (fixed, no params): gradually add Gaussian noise over T steps.
q(xt|xt-1) = N(xt; √(1-βt)xt-1, βtI)
Closed form: xt = √(ᾱt)x0 + √(1-ᾱt)ε, where ε ~ N(0,I) and ᾱt = Π(1-βt). At T=1000, xT ≈ N(0,I).
Reverse process (learned): neural network εθ(xt, t) predicts the noise ε added to x0.
Training loss: L = E[||ε - εθ(xt, t)||²] - predict the noise at each timestep. Equivalent to score matching.
Generation: start from xT ~ N(0,I). Iteratively denoise: xt-1 = (1/√αt)(xt - (βt/√(1-ᾱt))εθ(xt,t)) + noise.
DDPM vs DDIM - what does DDIM improve?frontier▶
DDPM (Denoising Diffusion Probabilistic Models): stochastic reverse process. Requires T=1000 steps for high quality. Each step: add a small amount of new noise. Slow generation.
DDIM (Denoising Diffusion Implicit Models):
- Reformulates the reverse process as an ODE (deterministic), not an SDE (stochastic)
- Uses the same trained noise predictor εθ
- Allows taking large steps through the ODE (50 or 20 steps vs 1000)
- Same noise → same image: deterministic sampling enables "editing" (interpolate in noise space)
Why it works: the DDPM training objective also trains a consistent implicit probability flow ODE. DDIM uses this ODE for inference without retraining.
Speed: 50-step DDIM ≈ quality of 1000-step DDPM. 20× speedup. Almost all production diffusion models use DDIM or similar samplers (DPM-Solver++).
What is classifier-free guidance?frontier▶
In conditional generation (e.g., text-to-image), you want the model to follow the condition strongly.
Classifier guidance (older): train a separate classifier p(c|xt) to steer generation. Inconvenient - need noisy classifier.
Classifier-free guidance (Ho & Salimans 2021): train a single model that can be both conditional and unconditional. During training, randomly drop the conditioning (replace with null token) with 10-20% probability.
At inference:
ε̃ = εunconditional + w × (εconditional - εunconditional)
w is the "guidance scale" (typically 7-10 for images). Higher w = more faithful to condition but less diverse.
Intuitively: amplify the difference between conditional and unconditional predictions, steering more strongly toward the condition. Used in Stable Diffusion, DALL-E 2/3, Imagen.
What is flow matching? How does it differ from diffusion?frontier▶
Flow matching (Lipman et al. 2022, used in Stable Diffusion 3, Meta's Voicebox, Flux) is a simpler alternative to diffusion for generative modeling.
Core idea: define a probability path interpolating from noise p₀ = N(0,I) to data distribution p₁. Train a vector field vθ(x, t) that generates this path via an ODE: dx/dt = vθ(x, t).
Training (conditional flow matching): for each data point x₁, define the path as linear interpolation: xt = (1-t)x₀ + tx₁ where x₀ ~ N(0,I). Target velocity: u(xt|x₁) = x₁ - x₀. Train vθ with MSE loss to predict this velocity.
Vs. diffusion:
- Straighter paths: flow matching learns nearly straight noise→data trajectories. Diffusion has curved stochastic paths. Straight paths → fewer sampling steps (5-10 vs 50-1000).
- Simpler training: no noise schedule design, no careful βt selection. Linear interpolation is the default path.
- Same inference: both solve an ODE/SDE, but flow matching ODEs have straighter paths → fewer NFE (number of function evaluations).
SD3, Flux, Meta Voicebox all use flow matching. Becoming the preferred approach over DDPM in 2024-2025.
What is latent diffusion? How does Stable Diffusion use it?▶
Pixel-space diffusion (DDPM, DALL-E 2 first stage) applies the denoising process directly on images - expensive because images are large (512×512×3 = 786K dimensions) and the U-Net runs at full resolution.
Latent diffusion (Rombach et al., 2022 - Stable Diffusion): compresses first, then diffuse.
Three components:
- VAE encoder: compresses the image into a lower-dimensional latent z (typically 64×64×4 - 48× smaller). Trained separately with reconstruction + KL regularisation loss.
- U-Net denoiser: runs the DDPM forward/reverse process in latent space. Cross-attention layers let text conditioning (from a frozen CLIP/OpenCLIP text encoder) steer the denoising.
- VAE decoder: maps the final clean latent z₀ back to pixel space.
Why it works: the VAE learns a perceptually meaningful compact representation - nearby points in latent space are perceptually similar. The U-Net only needs to learn the distribution of these compact codes, not raw pixels.
Speed gain: inference runs ~50 denoising steps on 64×64×4 tensors rather than 512×512×3 - order-of-magnitude cheaper, enabling consumer GPU generation.