Skip to content
samit
Interviews/Generative Models

Generative Models

8 questions

8 questions
GANs & VAEs
Explain GAN architecture - generator and discriminator objectives.▶

A GAN frames image generation as a two-player minimax game.

Generator G: takes random noise z ~ p(z) → generates fake sample G(z). Goal: fool D into thinking G(z) is real.

Discriminator D: binary classifier. Takes sample x → P(real). Goal: correctly classify real vs. fake.

Minimax objective:

minG maxD E[log D(x)] + E[log(1 - D(G(z)))]

  • D maximizes: D(x)→1 for real, D(G(z))→0 for fake
  • G minimizes: make D(G(z))→1. In practice: maximize log D(G(z)) (not minimize log(1-D(G(z)))) for stronger gradients early in training

Training: alternate D steps and G steps. Train D to optimality → update G → repeat.

At convergence (theory): G matches real data distribution, D outputs 0.5 everywhere. In practice: mode collapse and training instability prevent clean convergence.

Explain VAE - the ELBO objective and the reparameterization trick.frontier▶

VAE learns to compress data x into latent z and reconstruct it. We want to maximize log p(x) = log ∫ p(x|z)p(z)dz.

This is intractable. Instead, maximize the ELBO (Evidence Lower BOund):

L = Eq(z|x)[log p(x|z)] - KL(q(z|x) || p(z))

  • Reconstruction term: decoder p(x|z) must reconstruct x from z. Maximizing log p(x|z) = minimizing reconstruction loss (MSE or BCE).
  • KL term: encoder q(z|x) ~ N(μ(x), σ²(x)) must stay close to prior p(z) = N(0,I). Prevents latent space from being arbitrary. Has closed form: -½ Σ(1 + log σ² - μ² - σ²).

Reparameterization trick: can't backprop through sampling z ~ q(z|x) = N(μ, σ). Instead: z = μ + σ⊙ε where ε ~ N(0,I). Now the stochasticity is in ε (no parameters), and gradients flow through μ and σ.

What is mode collapse in GANs? How do you fix it?▶

The generator learns to produce only one or a few modes of the data distribution, even if the real data has many modes. The discriminator can't distinguish, so the generator stops diversifying.

Why it happens: the generator finds a "safe" region where the discriminator is fooled. Gradient signals stop directing toward unexplored modes.

Fixes:

  • WGAN (Wasserstein GAN): replace JS divergence with Wasserstein distance. Better gradient landscapes, doesn't saturate.
  • Minibatch discrimination: discriminator sees statistics across the batch, not just individual samples - penalizes low diversity.
  • Mode-seeking GAN: explicitly maximize distance between generated samples (diversity loss).
  • Spectral normalization: normalize discriminator weights to enforce Lipschitz constraint → stable training.
  • StyleGAN / progressive growing: architectural improvements that stabilize training.
Diffusion Models
Explain diffusion models - forward process, reverse process, training.frontier▶

Forward process (fixed, no params): gradually add Gaussian noise over T steps.

q(xt|xt-1) = N(xt; √(1-βt)xt-1, βtI)

Closed form: xt = √(ᾱt)x0 + √(1-ᾱt)ε, where ε ~ N(0,I) and ᾱt = Π(1-βt). At T=1000, xT ≈ N(0,I).

Reverse process (learned): neural network εθ(xt, t) predicts the noise ε added to x0.

Training loss: L = E[||ε - εθ(xt, t)||²] - predict the noise at each timestep. Equivalent to score matching.

Generation: start from xT ~ N(0,I). Iteratively denoise: xt-1 = (1/√αt)(xt - (βt/√(1-ᾱt))εθ(xt,t)) + noise.

DDPM vs DDIM - what does DDIM improve?frontier▶

DDPM (Denoising Diffusion Probabilistic Models): stochastic reverse process. Requires T=1000 steps for high quality. Each step: add a small amount of new noise. Slow generation.

DDIM (Denoising Diffusion Implicit Models):

  • Reformulates the reverse process as an ODE (deterministic), not an SDE (stochastic)
  • Uses the same trained noise predictor εθ
  • Allows taking large steps through the ODE (50 or 20 steps vs 1000)
  • Same noise → same image: deterministic sampling enables "editing" (interpolate in noise space)

Why it works: the DDPM training objective also trains a consistent implicit probability flow ODE. DDIM uses this ODE for inference without retraining.

Speed: 50-step DDIM ≈ quality of 1000-step DDPM. 20× speedup. Almost all production diffusion models use DDIM or similar samplers (DPM-Solver++).

What is classifier-free guidance?frontier▶

In conditional generation (e.g., text-to-image), you want the model to follow the condition strongly.

Classifier guidance (older): train a separate classifier p(c|xt) to steer generation. Inconvenient - need noisy classifier.

Classifier-free guidance (Ho & Salimans 2021): train a single model that can be both conditional and unconditional. During training, randomly drop the conditioning (replace with null token) with 10-20% probability.

At inference:

ε̃ = εunconditional + w × (εconditional - εunconditional)

w is the "guidance scale" (typically 7-10 for images). Higher w = more faithful to condition but less diverse.

Intuitively: amplify the difference between conditional and unconditional predictions, steering more strongly toward the condition. Used in Stable Diffusion, DALL-E 2/3, Imagen.

What is flow matching? How does it differ from diffusion?frontier▶

Flow matching (Lipman et al. 2022, used in Stable Diffusion 3, Meta's Voicebox, Flux) is a simpler alternative to diffusion for generative modeling.

Core idea: define a probability path interpolating from noise p₀ = N(0,I) to data distribution p₁. Train a vector field vθ(x, t) that generates this path via an ODE: dx/dt = vθ(x, t).

Training (conditional flow matching): for each data point x₁, define the path as linear interpolation: xt = (1-t)x₀ + tx₁ where x₀ ~ N(0,I). Target velocity: u(xt|x₁) = x₁ - x₀. Train vθ with MSE loss to predict this velocity.

Vs. diffusion:

  • Straighter paths: flow matching learns nearly straight noise→data trajectories. Diffusion has curved stochastic paths. Straight paths → fewer sampling steps (5-10 vs 50-1000).
  • Simpler training: no noise schedule design, no careful βt selection. Linear interpolation is the default path.
  • Same inference: both solve an ODE/SDE, but flow matching ODEs have straighter paths → fewer NFE (number of function evaluations).

SD3, Flux, Meta Voicebox all use flow matching. Becoming the preferred approach over DDPM in 2024-2025.

What is latent diffusion? How does Stable Diffusion use it?▶

Pixel-space diffusion (DDPM, DALL-E 2 first stage) applies the denoising process directly on images - expensive because images are large (512×512×3 = 786K dimensions) and the U-Net runs at full resolution.

Latent diffusion (Rombach et al., 2022 - Stable Diffusion): compresses first, then diffuse.

Three components:

  • VAE encoder: compresses the image into a lower-dimensional latent z (typically 64×64×4 - 48× smaller). Trained separately with reconstruction + KL regularisation loss.
  • U-Net denoiser: runs the DDPM forward/reverse process in latent space. Cross-attention layers let text conditioning (from a frozen CLIP/OpenCLIP text encoder) steer the denoising.
  • VAE decoder: maps the final clean latent z₀ back to pixel space.

Why it works: the VAE learns a perceptually meaningful compact representation - nearby points in latent space are perceptually similar. The U-Net only needs to learn the distribution of these compact codes, not raw pixels.

Speed gain: inference runs ~50 denoising steps on 64×64×4 tensors rather than 512×512×3 - order-of-magnitude cheaper, enabling consumer GPU generation.