Deep Learning
18 questions
Explain backpropagation step by step.▶
Backprop efficiently computes ∂L/∂w for every weight w using the chain rule.
- Forward pass: compute activations a(l) = σ(W(l)a(l-1) + b(l)) layer by layer. Store all activations.
- Compute loss: L = loss(a(last), y)
- Backward pass: propagate error signal δ(l) = ∂L/∂z(l) backward
- Output layer: δ(L) = ∂L/∂a(L) ⊙ σ'(z(L))
- Hidden layer: δ(l) = (W(l+1)Tδ(l+1)) ⊙ σ'(z(l))
- Gradients: ∂L/∂W(l) = δ(l)(a(l-1))T, ∂L/∂b(l) = δ(l)
Why "back": we reuse computations by propagating error signals from output to input, computing gradients at each layer using already-computed upstream gradients. This is O(forward pass) - no redundant computation.
Sigmoid vs Tanh vs ReLU vs Leaky ReLU - when does each fail?▶
- Sigmoid σ(x)=1/(1+e^-x): output (0,1). Saturates at extremes (σ' → 0). Not zero-centered → zig-zag gradient updates. Only good for binary output layer.
- Tanh tanh(x): output (-1,1). Zero-centered (better than sigmoid). Still saturates. Preferred over sigmoid in hidden layers when saturation is acceptable.
- ReLU max(0,x): gradient = 1 when active (no saturation). Sparse activations (good inductive bias). Dying ReLU: neurons with negative zl for all training examples never activate, their gradients are always 0, they never update. Can freeze ~10-50% of neurons.
- Leaky ReLU max(αx,x), α=0.01: fixes dying ReLU by allowing small negative gradient. GELU (used in GPT): smooth approximation, slightly better in practice for Transformers.
Modern practice: transformers use GELU. CNNs often use ReLU or Swish. Dying ReLU rarely a problem with He init + BN + residual connections.
Explain batch normalization - what does it normalize and why does it help?▶
For a mini-batch of activations {x1,...,xm} at a layer:
μ = mean, σ² = variance computed over the batch. Normalize: x̂i = (xi-μ)/√(σ²+ε). Then: yi = γx̂i + β (learned scale and shift).
Benefits:
- Reduces internal covariate shift - each layer sees a more stable distribution
- Acts as regularizer (noise from batch statistics) - can reduce dropout need
- Allows much higher learning rates
- Reduces sensitivity to weight initialization
Problems:
- Different behavior at train vs inference (inference uses running statistics from training)
- Fails with small batches (batch size 1 or 2)
- Doesn't work well with variable-length sequences (RNNs)
- Doesn't work well in Transformers (use Layer Norm instead)
Layer norm vs batch norm - why do transformers use layer norm?▶
Both normalise activations to zero mean and unit variance, then apply learned scale γ and shift β. What they normalise over differs:
Batch Norm: normalises over the batch dimension - for each feature, compute mean/var across all examples in the batch. Works well for CNNs (fixed spatial structure, large batches). Fails at:
- Small batches (batch=1 at inference: can't estimate statistics)
- Variable-length sequences (different sequence positions have different statistics)
- Autoregressive generation (generates one token at a time, batch=1)
Layer Norm: normalises over the feature dimension - for each example, compute mean/var across all features in that layer. Independent of batch size, works at batch=1, works for variable-length sequences.
Why transformers use LN: all three BN failure cases apply directly to language models. LN has no batch-size dependency and the normalisation is per-token (each position normalised independently over its d_model features).
Pre-LN vs Post-LN: original transformer applied LN after the residual add (Post-LN). GPT-2+ switched to Pre-LN (LN before the sublayer) - training is more stable (gradients flow better through the residual path), though Pre-LN can underfit if the final LN is omitted.
Explain residual/skip connections. Why do they help with deep networks?▶
Instead of h = F(x), learn h = F(x) + x. The network only needs to learn the "residual" F(x) = h - x.
Why they help:
- Gradient highway: ∂L/∂x = ∂L/∂h · (∂F/∂x + 1). Even if ∂F/∂x ≈ 0, gradient of at least 1 passes through. Vanishing gradients solved.
- Identity is easy to learn: just set F(x) = 0 (zero out the weights). Networks can learn to skip layers that aren't useful.
- Ensemble-like behavior: a network with n residual blocks is like an implicit ensemble of O(2^n) shallow networks.
ResNet-152 became trainable (without residual connections, it performs worse than ResNet-18 due to degradation problem). Modern Transformers all use residual connections.
Explain knowledge distillation.▶
Train a smaller student to match a larger teacher's softened outputs (temperature-scaled softmax), not just hard labels. KL(student || teacher) carries relative confidence across wrong classes.
Structured pruning: remove whole heads/channels/layers - dense smaller model, runs faster on normal hardware.
Unstructured pruning: zero individual weights - higher sparsity ratio but needs sparse kernels to realize speedups.
What is contrastive learning? Give an example.▶
Learn representations by bringing similar (positive) pairs close together and pushing dissimilar (negative) pairs apart in embedding space.
SimCLR: data augmentation creates two views of same image (positive pair). All other images in batch are negatives. NT-Xent loss: maximize cosine similarity of positives, minimize for negatives.
CLIP: 400M (image, caption) pairs from internet. Image encoder + text encoder. Contrastive loss over the batch - correct (image, text) pairs should have high cosine similarity; all other cross-modal pairs should have low similarity.
InfoNCE loss: L = -log(exp(sim(q,k+)/τ) / Σ exp(sim(q,ki)/τ)). τ is temperature - lower = harder negatives.
Power: learns task-agnostic representations without labels. CLIP's representations work zero-shot on new classification tasks.
What is dropout? How does it differ at train vs. inference?▶
Dropout randomly zeroes each neuron's activation with probability p (typically 0.1-0.5) during training.
Why it works:
- Ensemble interpretation: each forward pass uses a different sub-network. At inference, the full network approximates averaging 2^n sub-networks.
- Co-adaptation: neurons can't co-adapt to rely on specific other neurons. Each must learn independently useful features.
Train vs. inference: at training, randomly zero neurons. At inference, use ALL neurons but scale activations by (1-p) to match expected training magnitude - OR (inverse dropout) scale activations by 1/(1-p) during training so inference needs no scaling. PyTorch uses inverse dropout.
Important: call model.eval before inference - this disables dropout (and BatchNorm uses running stats instead of batch stats). Forgetting this is a classic bug.
Gotcha: MC Dropout - run inference with dropout enabled N times, take mean and variance. Use variance as an uncertainty estimate. Cheap Bayesian approximation.
Weight initialization: why does it matter? Xavier vs He.frontier▶
Poor initialization causes vanishing or exploding gradients from the very first forward pass, before any training occurs.
Goal: keep the variance of activations and gradients constant across layers. If variance grows, you get exploding gradients; if it shrinks, vanishing gradients.
Xavier (Glorot) initialization: W ~ N(0, 2/(n_in + n_out)) or U(-√(6/(n_in+n_out)), √(6/(n_in+n_out))). Designed for sigmoid/tanh where the linear regime gain is 1. Keeps variance constant for linear activations.
He (Kaiming) initialization: W ~ N(0, 2/n_in). Designed for ReLU where ~half the neurons are dead (zero). Compensates for the halved variance by doubling the scale. PyTorch default for Conv and Linear layers.
Rule of thumb: He init for ReLU/GELU/SiLU, Xavier for sigmoid/tanh, identity for RNNs (avoid exploding via orthogonal init).
What is label smoothing? When does it help?▶
Instead of using hard one-hot labels (0 or 1), use soft labels: y_smooth = (1-ε)·y + ε/K where ε is smoothing factor (0.1 common) and K is number of classes.
So the target for the correct class is 1 - ε + ε/K ≈ 0.9, and each wrong class gets ε/K ≈ 0.01.
Why it helps:
- Prevents overconfidence: cross-entropy with hard labels pushes the model to output logits as extreme as possible (output → ∞ for correct class, -∞ for others). Label smoothing caps this.
- Better calibration: model probabilities align better with actual accuracy. P(correct) = 0.9 means correct ~90% of the time.
- Noisy labels: if 10% of training labels are wrong, label smoothing prevents learning with too much confidence on potentially wrong labels.
When NOT to use: knowledge distillation. The teacher's soft targets already encode uncertainty - label smoothing would corrupt this signal.
Derive the backward pass of BatchNorm.▶
Forward (per feature, over the N samples in the batch): μ = mean(x), σ² = var(x), x̂ = (x − μ)/√(σ²+ε), y = γx̂ + β.
Easy gradients: dγ = Σ dy·x̂, dβ = Σ dy, dx̂ = dy·γ.
The hard part: μ and σ² each depend on every xᵢ, so x̂ couples all N samples. Differentiate through them:
- dσ² = Σ dx̂·(x − μ)·(−½)(σ²+ε)−3/2
- dμ = Σ dx̂·(−1/√(σ²+ε)) (the σ²-path term vanishes since Σ(x − μ) = 0)
- dx = dx̂/√(σ²+ε) + dσ²·2(x − μ)/N + dμ/N
Vectorised (the clean form):
dx = (1/N)·(1/√(σ²+ε))·(N·dx̂ − Σdx̂ − x̂·Σ(dx̂·x̂))
Why BN is annoying: the batch coupling means you cannot compute stats for a single example at inference. You keep an EMA of μ and σ² during training and use those running averages at test time - a train/inference mismatch that does not exist for LayerNorm.
Derive the backward pass of LayerNorm. Why is it simpler than BatchNorm in practice?▶
Forward (per token, over the D feature dimensions): μ and σ² are computed across the D features of one sample. x̂ = (x − μ)/√(σ²+ε), y = γ⊙x̂ + β.
Backward:
- dγ = Σ dy⊙x̂, dβ = Σ dy (summed over batch and positions)
- dx̂ = dy⊙γ
- dx = (1/√(σ²+ε))·(dx̂ − mean(dx̂) − x̂·mean(dx̂⊙x̂)), means over the D feature dim
Algebraically identical shape to the BatchNorm backward - the only difference is the axis of reduction.
The whole BN↔LN distinction: same formula, different axis. BN reduces over the batch (couples samples → needs running stats). LN reduces over features within one sample → no batch coupling, identical train vs inference, works at batch=1. That last point is exactly why transformers (autoregressive decode at batch=1) use LN.
Explain the convolution operation - output size formula, stride, padding.▶
A filter (kernel) of size k×k slides over the input, computing dot products at each position.
Output size: ⌊(input - kernel + 2×padding) / stride⌋ + 1
e.g., 224×224 input, 3×3 kernel, stride=1, padding=1 → ⌊(224-3+2)/1⌋+1 = 224. Same padding preserves spatial dims.
- Stride: step size between filter applications. stride=2 halves spatial dims.
- Padding=0 (valid): output smaller than input. Padding=1 (same): output same size.
- Parameter count: (k×k×C_in + 1) × C_out - weight sharing across all spatial locations.
Why convolutions for images: parameter sharing (same filter detects edge anywhere), translation equivariance (shifting input shifts feature map), local connectivity (nearby pixels are correlated).
How does a 1×1 convolution work and why is it useful?▶
A 1×1 conv applies a linear combination across channels at each spatial location. No spatial mixing - pure channel mixing.
Uses:
- Dimensionality reduction/expansion: ResNet bottleneck uses 1×1 to reduce channels before expensive 3×3 conv, then 1×1 to expand back. Reduces FLOPs.
- Adding non-linearity without spatial cost: 1×1 → ReLU adds a non-linear channel mixing step.
- Channel projection in attention: W_Q, W_K, W_V are essentially 1×1 convolutions (or linear layers) that project to attention spaces.
- NiN (Network-in-Network): use 1×1 convs to add a mini-MLP at each spatial location.
Depthwise separable convolutions - how do they work and why are they more efficient?▶
Standard 3×3 conv over C_in channels to C_out: (3×3×C_in) × C_out parameters + C_in×C_out spatial computation.
Depthwise separable = two steps:
- Depthwise: one 3×3 filter per input channel (C_in filters total). Each filter only looks at its own channel. Parameters: 3×3×C_in.
- Pointwise: 1×1 conv across all channels → mix information across channels. Parameters: C_in × C_out.
Savings: (9C_in + C_in × C_out) vs (9 × C_in × C_out). For C_out=256, ~8-9x fewer parameters and FLOPs.
Where used: MobileNet (mobile/edge inference), EfficientNet, Xception. The accuracy-efficiency tradeoff is excellent - small accuracy drop, massive compute savings.
Intuition: separate spatial feature extraction (depthwise) from channel mixing (pointwise). Standard conv does both at once, wastefully.
What is the vanishing gradient problem in RNNs? How do LSTMs solve it?▶
Backprop through time multiplies the weight matrix Wrec at each time step. Gradient at step t depends on Wrec^(T-t). If spectral radius ρ(W) < 1: gradient vanishes exponentially. If ρ(W) > 1: gradient explodes.
LSTM cell state (highway): Ct = ft ⊙ Ct-1 + it ⊙ C̃t
- The cell state Ct is connected to Ct-1 via the forget gate ft
- If ft ≈ 1 (initialized with positive bias), gradient flows near-identically through many steps
- This is a learned shortcut connection across time - similar to residual connections across layers
Gates: forget (what to discard), input (what to write), output (what to expose). GRU simplifies to 2 gates (reset, update) - fewer params, similar performance.
Why did Transformers replace LSTMs for NLP?▶
- Parallelism: LSTMs process tokens sequentially - can't parallelize over sequence length during training. Transformers compute all positions simultaneously.
- Long-range dependencies: LSTM must carry information through many sequential steps. Information can "leak" through gates. Transformer attends directly to any position in O(1) "hops."
- Training speed: Transformer training is 10-100x faster due to parallelism.
- Scaling: Transformers scale much better with compute. No LLM at scale uses LSTMs.
LSTM still relevant: edge devices (smaller models), time series where causal sequential processing is needed, Mamba/SSMs (state space models) as modern LSTM alternatives for long sequences.
What is teacher forcing? When does it hurt?▶
During training of sequence models (LSTMs, decoder Transformers), teacher forcing uses the ground truth previous token as input at each step, rather than the model's own prediction from the previous step.
Without teacher forcing: at step t, feed in the model's prediction ŷ_{t-1}. If the prediction is wrong, all subsequent steps see wrong context - error compounds. Gradients are hard to compute due to this dependency.
With teacher forcing: at step t, feed in the true y_{t-1} regardless of what the model predicted. Training is stable and fast. Every step gets a clean gradient signal.
When it hurts - exposure bias: the model never sees its own errors during training. At inference, it must condition on its own predictions. The distribution of inputs at train vs. inference time differs - model degrades when it makes early errors.
Mitigations:
- Scheduled sampling: gradually replace ground truth tokens with model predictions as training progresses
- Professor forcing: penalize divergence between teacher-forced and free-run hidden states