Karpathy Questions
26 questions
Implement scaled dot-product attention backward pass without autograd. Given dO, compute dQ, dK, dV.frontier▶
The hardest transformer question in practice. Forward is easy. Backward reveals whether you actually understand the chain rule through softmax and matrix products.
Forward: S = QK^T / √dk, A = softmax(S), O = AV
Derive each gradient:
- dV = A^T · dO (A is constant w.r.t. V)
- dA = dO · V^T
- dS via softmax backward: dS_ij = A_ij · (dA_ij - Σ_k dA_ik·A_ik) = A ⊙ (dA - (dA·A).sum(keepdim))
- dQ = dS/√dk · K, dK = dS^T/√dk · Q
Softmax backward trick: for softmax output p and upstream gradient g: dp = p ⊙ (g - ⟨p, g⟩). Efficient - no Jacobian matrix ever materialized.
import numpy as np
def softmax(x):
x = x - x.max(-1, keepdims=True)
e = np.exp(x)
return e / e.sum(-1, keepdims=True)
def attn_forward(Q, K, V, mask=None):
dk = Q.shape[-1]
S = Q @ K.transpose(-2,-1) / np.sqrt(dk)
if mask is not None: S = np.where(mask, S, -1e9)
A = softmax(S)
return A @ V, A
def attn_backward(dO, Q, K, V, A):
"""Shapes: (B, H, L, dk)"""
dk = Q.shape[-1]
# dV: O = A @ V => dV = A^T @ dO
dV = A.transpose(-2,-1) @ dO
# dA: O = A @ V => dA = dO @ V^T
dA = dO @ V.transpose(-2,-1)
# dS via softmax backward: p.grad = p * (g - sum(p*g, keepdims=True))
dS = A * (dA - (dA * A).sum(-1, keepdims=True))
dS = dS / np.sqrt(dk)
# dQ = dS @ K, dK = dS^T @ Q
dQ = dS @ K
dK = dS.transpose(-2,-1) @ Q
return dQ, dK, dV
# Gradient check
np.random.seed(0)
Q=np.random.randn(1,2,4,8); K=np.random.randn(1,2,4,8); V=np.random.randn(1,2,4,8)
O, A = attn_forward(Q, K, V)
dQ, dK, dV = attn_backward(np.ones_like(O), Q, K, V, A)
# Compare dQ[0,0,0,0] with finite difference: (f(Q+eps)-f(Q-eps))/(2*eps)Build a minimal autograd engine. Value class with +, *, relu, backward using topological sort.frontier▶
This is the core of all deep learning frameworks. PyTorch does the same thing in C++. Understanding every line of this class means understanding backprop.
Critical correctness rule: always += for grad accumulation, never =. A value used in two downstream operations must receive gradients from both (multivariate chain rule). This is the most common bug in from-scratch autograd.
The closure trick: _bwd is defined inside __mul__ and captures self, other, out by reference. This is how local Jacobians are stored without explicit node attributes.
class Value:
def __init__(self, data, _prev=()):
self.data = float(data)
self.grad = 0.0
self._backward = lambda: None
self._prev = set(_prev)
def __add__(self, other):
other = other if isinstance(other, Value) else Value(other)
out = Value(self.data + other.data, (self, other))
def _bwd(): self.grad += out.grad; other.grad += out.grad
out._backward = _bwd; return out
def __mul__(self, other):
other = other if isinstance(other, Value) else Value(other)
out = Value(self.data * other.data, (self, other))
def _bwd():
self.grad += other.data * out.grad # product rule
other.grad += self.data * out.grad
out._backward = _bwd; return out
def relu(self):
out = Value(max(0, self.data), (self,))
def _bwd(): self.grad += (out.data > 0) * out.grad
out._backward = _bwd; return out
def __pow__(self, exp):
out = Value(self.data**exp, (self,))
def _bwd(): self.grad += exp * self.data**(exp-1) * out.grad
out._backward = _bwd; return out
def __neg__(self): return self * -1
def __sub__(self, o): return self + (-o)
def __truediv__(self, o): return self * o**-1
__rmul__ = __mul__; __radd__ = __add__
def backward(self):
topo, seen = [], set()
def build(v):
if v not in seen:
seen.add(v); [build(c) for c in v._prev]; topo.append(v)
build(self)
self.grad = 1.0
for v in reversed(topo): v._backward()
# Demo: L = relu(a*b + c), verify gradients
a,b,c = Value(2.0),Value(-3.0),Value(5.0)
L = (a*b + c).relu() # relu(-6+5) = relu(-1) = 0
L.backward()
# a.grad = b * (L.data>0) = -3*0 = 0
# b.grad = a * (L.data>0) = 2*0 = 0
# c.grad = 1 * (L.data>0) = 0
print(a.grad, b.grad, c.grad) # 0.0, 0.0, 0.0Implement AdamW from scratch. What is the precise difference between Adam and AdamW?frontier▶
Adam with L2 regularization: adds λ||w||² to loss. Gradient becomes g + λw. Adam divides this by √v̂, giving effective weight decay = lr·λw/√v̂. Inconsistent across parameters with different gradient magnitudes - parameters with small gradients get stronger decay than parameters with large gradients. Wrong.
AdamW (decoupled weight decay, Loshchilov & Hutter 2018): apply weight decay directly to weights BEFORE the Adam update. Effective decay = lr·wd·w, constant regardless of gradient history. The fix is literally adding one line before the Adam formula.
import numpy as np
class AdamW:
def __init__(self, params, lr=3e-4, betas=(0.9, 0.999), eps=1e-8, wd=0.1):
self.params = list(params)
self.lr, self.b1, self.b2, self.eps, self.wd = lr, betas[0], betas[1], eps, wd
self.m = [np.zeros_like(p) for p in self.params]
self.v = [np.zeros_like(p) for p in self.params]
self.t = 0
def step(self, grads):
self.t += 1
for i, (p, g) in enumerate(zip(self.params, grads)):
# === AdamW: DECOUPLED weight decay (applied before Adam step) ===
p -= self.lr * self.wd * p # this line makes it "W"
# Momentum (first moment)
self.m[i] = self.b1 * self.m[i] + (1 - self.b1) * g
# Variance (second moment)
self.v[i] = self.b2 * self.v[i] + (1 - self.b2) * g**2
# Bias correction: m,v start at 0 -> underestimated in early steps
m_hat = self.m[i] / (1 - self.b1**self.t)
v_hat = self.v[i] / (1 - self.b2**self.t)
# Adam update
p -= self.lr * m_hat / (np.sqrt(v_hat) + self.eps)
# Standard LLM hyperparameters: lr=3e-4, betas=(0.9, 0.95), wd=0.1 (GPT-style)
# Note: b2=0.95 (not 0.999) for LLM training - faster adaptationWhat happens when you call loss.backward in PyTorch? Trace the actual call chain.frontier▶
No magic. It is graph traversal + chain rule, implemented in C++.
loss.backward→ C++torch::autograd::backward({loss}, {ones(1)})- The autograd Engine performs reverse-topological BFS from the loss node
- Each tensor's
grad_fnstores the backward op (e.g.,AddmmBackward0,SoftmaxBackward0) - Engine calls
grad_fn(upstream_gradient)→ returns list of input gradients - Gradients accumulated into
.gradof leaf tensors (model parameters) - Non-leaf activations: their gradients are freed after use (unless
retain_gradcalled)
Key flags:
retain_graph=True: don't free the graph after backward. Needed for MAML, gradient penalties in GANs.create_graph=True: build differentiable graph through backward. Needed for higher-order gradients (Hessian computation).torch.no_grad: stops building the graph entirely. ~40% memory savings during inference.
torch.compile (PyTorch 2.0+): captures the graph via TorchDynamo (Python bytecode tracing), optimizes via TorchInductor (generates Triton kernels), fuses ops across entire forward+backward. 1.5-3× training speedup common for LLMs. Use for production training runs.
A 7B model in fp32 = 28GB. You have a 24GB GPU. All options for inference and fine-tuning.frontier▶
Memory math first - always do this in interviews: 7B params × 4 bytes (fp32) = 28GB. 7B × 2 bytes (bf16) = 14GB. 7B × 1 byte (int8) = 7GB. 7B × 0.5 bytes (int4) = 3.5GB.
Inference options (need weights + KV cache in 24GB):
- BF16: 14GB. Fits with 10GB for KV cache (~4K context at batch=1). Best quality. Default choice.
- INT8 (bitsandbytes LLM.int8): 7GB weights. Quality ≈ BF16 for most tasks. 1.5-2× slower (INT8 dequant overhead). Use when 14GB is tight.
- INT4 (GPTQ/AWQ): 3.5GB. Visible quality drop on complex reasoning/code. Consumer-hardware deployments only.
- GGUF + llama.cpp: CPU+GPU hybrid offloading. Some layers in RAM. Works on 8GB VRAM + 32GB RAM. Slow but flexible.
Fine-tuning options:
- Full fine-tuning: BF16 weights (14GB) + fp32 Adam optimizer states (28GB) + activations → 42GB minimum. Doesn't fit on 24GB.
- LoRA (r=16): freeze 14GB BF16 base. Train ~130MB adapters. Adam states = 260MB. Total ≈ 14.4GB. Fits easily. Merge at inference: W' = W + αBA, zero added latency.
- QLoRA (Dettmers 2023): base in NF4 (4-bit normalized float) = 3.5GB. LoRA adapters in BF16. Total ≈ 4GB. Fits on a 12GB RTX 3080!
- Gradient checkpointing + LoRA: recompute activations during backward instead of storing. 10-20× activation memory reduction, +30% compute. Needed for large batch sizes.
Implement top-p nucleus sampling with temperature. Why does temperature work?▶
Temperature T: divide logits by T before softmax. T<1 → peaks sharpen (more deterministic). T>1 → flattens (more random). T→0 → greedy (argmax). T→∞ → uniform distribution.
Why it works: softmax(z/T). As T→0, relative differences in logits are amplified - the highest logit gets all the mass. As T→∞, all z/T→0, softmax→1/vocab. Temperature controls the "sharpness" of the learned distribution without changing the ranking.
Top-p (nucleus): keep the smallest set of tokens whose cumulative prob ≥ p. Dynamic vocabulary size - adapts to distribution entropy. Low-entropy step (clear next word): nucleus might be 5 tokens. High-entropy step: nucleus might be 500 tokens.
import torch, torch.nn.functional as F
def sample(logits, temperature=1.0, top_p=0.9, top_k=0):
"""logits: (vocab_size,) - raw unnormalized scores from model"""
# Step 1: temperature scaling
logits = logits / max(temperature, 1e-5)
# Step 2: optional top-k filtering (before top-p)
if top_k > 0:
topk_thresh = torch.topk(logits, top_k).values[-1]
logits[logits < topk_thresh] = float('-inf')
# Step 3: softmax
probs = F.softmax(logits, dim=-1)
# Step 4: top-p nucleus filtering
sorted_probs, sorted_idx = torch.sort(probs, descending=True)
cumprobs = torch.cumsum(sorted_probs, dim=-1)
# Remove tokens AFTER the cumsum exceeds top_p (shift by 1 to include the crossing token)
to_remove = (cumprobs - sorted_probs) >= top_p
sorted_probs[to_remove] = 0.0
sorted_probs = sorted_probs / sorted_probs.sum() # renormalize
# Step 5: sample from filtered distribution
sampled = torch.multinomial(sorted_probs, num_samples=1)
return sorted_idx[sampled].item()Walk through a nanoGPT training loop line by line: data → forward → loss → backward → step → eval.▶
Data: tokenize the corpus into one flat array. get_batch samples B random offsets; for each, x = tokens[i : i+T] and y = tokens[i+1 : i+T+1]. y is x shifted by one - every position predicts the next token.
Forward: token embeddings + positional embeddings → N transformer blocks (pre-LN: LN → attention → residual, LN → MLP → residual) → final LN → logits via the LM head (weights usually tied to the token embedding).
Loss: F.cross_entropy(logits.view(-1, V), y.view(-1)) - mean negative log-likelihood across all B·T positions. Every position is a training signal, which is why a single sequence teaches T predictions.
Backward + step:
for step in range(max_steps): x, y = get_batch('train') logits, loss = model(x, y) optimizer.zero_grad(set_to_none=True) loss.backward torch.nn.utils.clip_grad_norm_(model.parameters, 1.0) lr = get_lr(step) # warmup then cosine decay for g in optimizer.param_groups: g['lr'] = lr optimizer.step # AdamW
Eval: periodically, under torch.no_grad with model.eval, average the loss over a few train and val batches, then switch back to model.train.
Details that matter: weight tying, bias-free LN/Linear, gradient accumulation for a large effective batch, mixed precision (bf16 autocast), and the warmup+cosine LR schedule.
the story: you built MicroGPT in ~218 lines. Be ready to defend every line - especially why y is x shifted by one and why the loss means over all T positions.
Implement BPE (byte-pair encoding) tokenizer from scratch.▶
Start from bytes (256 tokens). Count adjacent pairs, merge the most frequent, repeat until vocab size target.
- Encode: greedily apply merges left-to-right on bytes
- Decode: map token ids back through merge table to bytes → UTF-8 string
- Why bytes: no UNK token; any Unicode string tokenizes
Karpathy's minBPE and GPT-2 tokenizer follow this exactly. Be ready to write get_stats, merge, and the training loop.
Trace tensor shapes through one GPT-2 block forward pass (batch B, seq T, model dim C).▶
Input x: [B, T, C]
- Q,K,V projections: each
[B, T, C](with n_head splits →[B, n_h, T, d_h]where d_h = C/n_h) - Attention scores:
[B, n_h, T, T] - Attention out:
[B, n_h, T, d_h]→ merge heads →[B, T, C] - FFN:
[B, T, C] → [B, T, 4C] → [B, T, C] - Residual paths keep
[B, T, C]throughout
The only O(T²) tensor is the attention map. Everything else is O(T).
Why tie input token embeddings to the LM head weights?▶
Both map between vocab space (V) and model dim (C). Tying uses one matrix W ∈ ℝ^{V×C} for embedding lookup and logits = x @ W^T.
- Saves ~V×C parameters (for GPT-2 vocab 50K, C=768 → ~38M params)
- Regularizes: input and output representations must agree
- Standard in GPT-2/3, LLaMA, nanoGPT
Untied head can help when vocab is huge or modalities differ, but for vanilla LMs tying is the default.
Training loss is NaN after ~100 steps. Walk through the debug checklist.▶
- Print one batch: labels in range? Inputs finite?
- Lower LR 10× - exploding gradients?
clip_grad_norm_(..., 1.0)- does loss stabilize?- Overfit 1 batch: if still NaN, bug in loss/forward (wrong ignore_index, div by zero in LN)
- Check mixed precision: loss scaling, bf16 vs fp16 overflow
- Inspect per-layer grad norms - which layer blows up first?
Karpathy-style: always overfit one batch before scaling up. If you can't hit ~0 loss on 1 batch, the code is wrong.
Implement LayerNorm forward and backward in NumPy (no autograd).▶
Forward per row x ∈ ℝ^D: μ = mean(x), σ² = var(x), x̂ = (x−μ)/√(σ²+ε), y = γ⊙x̂ + β
Backward:
- dγ = Σ dy⊙x̂, dβ = Σ dy
- dx̂ = dy⊙γ
- dx = (1/√(σ²+ε)) · (dx̂ − mean(dx̂) − x̂⊙mean(dx̂⊙x̂))
Same formula as BatchNorm backward but reduction is over features, not batch - no running stats, works at batch=1.
Estimate Pi with Monte Carlo - implement it and explain the variance.▶
Sample (x,y) ~ Uniform[0,1]². Count fraction with x²+y² ≤ 1 → estimates π/4.
Variance shrinks as O(1/N). After N=10⁶ samples, std error ≈ 0.0016 on π.
Karpathy uses this to test if candidates understand probability, loops, and convergence - not memorized LeetCode.
Matrix multiply: O(n³) naive, then how do you actually make it fast on a GPU?▶
Levels:
- BLAS:
a @ bcalls cuBLAS - never write naive triple loops in production - Tiling: load blocks into shared memory to raise arithmetic intensity
- Strassen: O(n^2.807) - rarely wins except huge matrices due to constants
- Tensor Cores: WMMA ops for FP16/BF16 tile matmuls
Interview move: write naive matmul, then explain why it's memory-bound and what cuBLAS does differently.
Your network won't overfit a single batch. Walk the debug checklist.▶
Overfitting one batch to ~zero loss is the first sanity check before any real training. If it fails, something is fundamentally broken.
- Loss not decreasing at all: learning rate too low, or gradients not flowing (check
requires_grad, a detached tensor, or.dataused by accident). - Loss is NaN: LR too high, log of zero, or a bad init. Lower LR, add eps, check the loss formula.
- Loss plateaus above zero: the label and prediction are misaligned (off-by-one in targets), the output layer is wrong (logits vs probs into the loss), or a dead activation.
- Forgot
optimizer.zero_grad(): gradients accumulate across steps. - Model in eval mode (dropout/BN off) or data accidentally shuffled each step so it is never the same batch.
If you cannot drive a single batch to near-zero loss, do not move on. The model has the capacity to memorize 32 examples; failing means a wiring bug, not a learning problem.
How do you gradient-check an autograd implementation?▶
Verify analytical gradients against numerical ones. For each parameter θ, the centered finite difference:
grad_numerical ≈ (f(θ + ε) - f(θ - ε)) / (2ε), with ε ≈ 1e-5.
Compare to the backward-pass gradient with a relative error: |g_analytic - g_numeric| / (|g_analytic| + |g_numeric| + 1e-8). Below ~1e-7 is correct; above ~1e-3 means a bug.
- Use double precision (float64) so ε rounding does not dominate.
- Use the centered difference, not one-sided; the error is O(ε²) instead of O(ε).
- Watch non-differentiable kinks (ReLU at 0) - perturbation can straddle the kink and look wrong.
This is exactly how you trust a from-scratch autograd before training anything with it.
Explain tensor strides, views, and what .contiguous() does.▶
A tensor is a flat 1-D block of memory plus metadata: shape and strides (how many elements to step in memory to move one index along each dimension).
Views (transpose, slice, reshape-when-possible) return a new tensor that points at the same memory with different shape/strides - no copy. So x.T is free; it just swaps the strides.
The catch: after a view, memory order no longer matches index order. Some ops (and .view()) need a contiguous layout. .contiguous() forces a copy into row-major order so strides match the shape again.
Why it matters: understanding strides explains why .view() sometimes errors and .reshape() does not (reshape copies if needed), and why a transpose followed by view fails.
Why does weight initialization matter? What breaks with bad init?▶
Init sets the scale of activations and gradients before any learning. Get it wrong and signal either explodes or vanishes as it flows through layers.
- Too large: activations saturate (tanh/sigmoid) or blow up; gradients explode; loss goes NaN.
- Too small: activations shrink toward zero layer by layer; gradients vanish; deep layers never learn.
- All zeros: every neuron in a layer computes the identical thing and gets the identical gradient - they never differentiate. Symmetry is never broken.
The fix: keep variance roughly constant across layers. Xavier/Glorot (var ≈ 1/n) for tanh, He/Kaiming (var ≈ 2/n) for ReLU since ReLU zeros half the inputs. Modern transformers also scale residual-branch init down by depth so the residual stream does not blow up.
.data vs .detach() vs .item() vs torch.no_grad() - precise differences.▶
- .item() - pulls a single-element tensor out as a Python number. Breaks the graph because it leaves tensor-land entirely. Use for logging a scalar loss.
- .detach() - returns a new tensor sharing the same storage but cut out of the autograd graph (
requires_grad=False). The safe way to use a value without backprop through it. - .data - the old, unsafe way. Also gives a detached view but bypasses autograd's safety checks; in-place edits via
.datacan silently corrupt gradients. Avoid it. - with torch.no_grad(): - a context that disables graph building for everything inside. Use for inference and eval loops; saves memory and time by not tracking ops.
Interview trap: using .data instead of .detach() in a training loop is a classic source of "my gradients are silently wrong" bugs.
What is the residual stream? Why do residual connections make deep nets trainable?▶
A residual block computes y = x + f(x): the input is added back to the sublayer output. Stack these and there is a continuous additive path from input to output - the residual stream.
Why it trains: the gradient of y = x + f(x) is 1 + f'(x). The "+1" means gradient always flows straight back through the skip path, even if f'(x) is tiny - so no vanishing gradient through depth. The network learns a residual (a delta on the identity) instead of the full mapping, which is an easier target.
Mechanistic view (transformers): each attention and MLP block reads from and writes to the shared residual stream. The stream is the model's working memory; layers communicate by adding features into it.
This is why ResNets scaled to 100+ layers and why every transformer is residual.
What does a CUDA matmul kernel actually do? Why are Tensor Cores faster than regular CUDA cores?frontier▶
Naive CUDA matmul: each thread computes C[i,j] = Σ_k A[i,k]·B[k,j]. Reads A[i,:] and B[:,j] from global memory (HBM) per output element. HBM bandwidth ≈ 2TB/s on A100. Kernel is severely memory-bound - compute sits idle waiting for data.
Tiled matmul (what cuBLAS/Flash Attention actually do):
- Divide A, B into tiles that fit in SRAM/shared memory (~32-48KB per SM)
- Each thread block loads one tile of A + one tile of B into SRAM (fast: ~19TB/s)
- Compute partial products using SRAM data only - no HBM access during computation
- Advance tile pointer across K dimension, accumulate, repeat
- Write final C block back to HBM once
Arithmetic intensity: naive = O(1) FLOP/byte. Tiled = O(tile_size) FLOP/byte → compute-bound instead of memory-bound.
Tensor Cores: dedicated hardware units that compute a 16×16×16 matmul (D = A×B+C) in one clock cycle. Regular CUDA cores: one multiply-accumulate per clock (scalar). Tensor Cores: 256 multiply-accumulate per clock. ~8× throughput for FP16, BF16, INT8, FP8. Flash Attention uses Tensor Cores via WMMA instructions (cute/CUTLASS library).
Roofline model: LLM inference (batch=1) is memory-bound. LLM training (large batch) is compute-bound. This determines the optimization strategy - inference: reduce bytes transferred (quantization, KV cache). Training: maximize GPU utilization (batch size, kernel fusion).
Why is LLM inference memory-bandwidth bound? What optimization strategies follow from this?frontier▶
Arithmetic intensity math: for weight W ∈ ℝ^{d×d} and input x ∈ ℝ^d (batch=1):
- FLOPs: 2d² (matmul)
- Memory: d² × 2 bytes (read weights in BF16) + d × 2 bytes (read/write x) ≈ 2d² bytes
- Arithmetic intensity ≈ 2d²/(2d²) = 1 FLOP/byte
A100 ratio: 312 TFLOP/s ÷ 2 TB/s = 156 FLOP/byte. Since 1 ≪ 156, inference at batch=1 is deeply memory-bound. GPU compute sits idle ~99% of the time waiting for weights from HBM.
Optimization strategies that follow directly:
- INT4 quantization: 4× fewer bytes per weight → 4× faster. Not about fewer FLOPs - about less HBM bandwidth consumed.
- KV cache: avoids re-reading K,V for all previous tokens each step. But KV cache itself must be read → as context grows, KV cache reading dominates.
- Speculative decoding: draft model (small, bandwidth-light) proposes K tokens → big model verifies in one pass (now effective batch=K → better arithmetic intensity).
- Continuous batching (vLLM): N sequences together → arithmetic intensity scales with N → becomes compute-bound at high N → full GPU utilization.
Karpathy's rule: "If they are not batching, they are not using the GPU." The difference between batch=1 (memory-bound, 1% utilization) and batch=512 (compute-bound, 80% utilization) is the fundamental lever of LLM serving economics.
What is arithmetic intensity? Explain the roofline model.frontier▶
Arithmetic intensity = FLOPs performed ÷ bytes moved from memory. It tells one whether a kernel is limited by compute or by bandwidth.
Roofline: achievable FLOP/s = min( peak compute, bandwidth × arithmetic intensity ). Plot intensity on the x-axis: below the "ridge" they are memory-bound (the slanted bandwidth roof), above it they are compute-bound (the flat compute roof).
Where the ridge is: peak_FLOPs ÷ peak_bandwidth. For an A100 ≈ 312 TFLOP/s ÷ 2 TB/s ≈ 156 FLOP/byte. You need intensity above ~156 to saturate the compute units.
Examples:
- Large square matmul (M×K)·(K×N): 2MKN FLOPs, ~(MK+KN+MN) bytes → high intensity → compute-bound. GPUs love this.
- Matrix×vector (batch-1 decode): 2KN FLOPs, ~KN weight bytes → intensity ≈ 2 → deeply memory-bound. The GPU sits mostly idle.
This single framework explains FlashAttention, fused kernels, and why batching matters - they all push intensity up so you stop wasting the compute roof.
Estimate the tokens/sec ceiling for a 7B model on one A100. Show the arithmetic.frontier▶
Single-stream decode is memory-bound: to generate one token you read the entire model out of HBM once.
- 7B params × 2 bytes (fp16) = 14 GB read per token
- A100 HBM bandwidth ≈ 2 TB/s
- floor latency/token = 14 GB ÷ 2 TB/s = 7 ms → ceiling ≈ 143 tok/s (batch=1, ignoring KV cache and overhead)
Quantization moves the ceiling: int8 (7 GB) → ~3.5 ms → ~285 tok/s; int4 (3.5 GB) → ~570 tok/s. That is why quantization is the lever for single-stream latency - it directly cuts the bytes you read per token.
Reality check: production hits ~60-70% of this roofline. But if the measured throughput is 10× under the estimate, you have a software bug, not a hardware limit. Knowing the ceiling is how one tell the difference.
GPU memory hierarchy - the numbers that actually matter.frontier▶
Top to bottom, fast/small to slow/large (A100-class):
- Registers: ~KBs per thread, fastest, private to a thread.
- SRAM / shared memory: ~192 KB per SM (~20 MB on-chip total), ~19 TB/s. The scratchpad FlashAttention tiles Q/K/V into.
- HBM (global memory): 40-80 GB, ~2 TB/s. Where weights and activations live - and the bottleneck for almost everything.
- Host RAM over PCIe: ~64 GB/s - roughly 30× slower than HBM. CPU/disk offload is a last resort, not a plan.
One rule explains most GPU optimization: minimise HBM traffic. FlashAttention, kernel fusion, activation recomputation, and quantization are all the same idea - keep data in SRAM/registers and stop round-tripping to HBM.
What is MFU (Model FLOPs Utilization)? What caps training throughput?frontier▶
MFU = achieved FLOP/s ÷ hardware peak FLOP/s. A well-tuned large training run lands around 40-55%; above 50% is excellent.
Training FLOPs per token ≈ 6N for N parameters (≈2N forward + ≈4N backward). So tokens/sec ≈ MFU × peak_FLOPs ÷ 6N - a quick way to sanity-check a training run.
What kills MFU:
- Memory-bandwidth stalls and low arithmetic intensity (small micro-batch / short sequences)
- Communication: all-reduce / all-gather in data and tensor parallelism
- Pipeline bubbles in pipeline parallelism
- Kernel launch overhead and poor overlap of compute with comms
Levers: larger micro-batch, fused kernels (FlashAttention), the right parallelism layout, gradient-checkpoint tradeoffs, and overlapping communication with compute (FSDP/ZeRO prefetch).