Skip to content
samit
Interviews/Karpathy Questions

Karpathy Questions

26 questions

26 questions
Core Questions
Implement scaled dot-product attention backward pass without autograd. Given dO, compute dQ, dK, dV.frontier▶

The hardest transformer question in practice. Forward is easy. Backward reveals whether you actually understand the chain rule through softmax and matrix products.

Forward: S = QK^T / √dk, A = softmax(S), O = AV

Derive each gradient:

  • dV = A^T · dO (A is constant w.r.t. V)
  • dA = dO · V^T
  • dS via softmax backward: dS_ij = A_ij · (dA_ij - Σ_k dA_ik·A_ik) = A ⊙ (dA - (dA·A).sum(keepdim))
  • dQ = dS/√dk · K,   dK = dS^T/√dk · Q

Softmax backward trick: for softmax output p and upstream gradient g: dp = p ⊙ (g - ⟨p, g⟩). Efficient - no Jacobian matrix ever materialized.

python
import numpy as np

def softmax(x):
    x = x - x.max(-1, keepdims=True)
    e = np.exp(x)
    return e / e.sum(-1, keepdims=True)

def attn_forward(Q, K, V, mask=None):
    dk = Q.shape[-1]
    S = Q @ K.transpose(-2,-1) / np.sqrt(dk)
    if mask is not None: S = np.where(mask, S, -1e9)
    A = softmax(S)
    return A @ V, A

def attn_backward(dO, Q, K, V, A):
    """Shapes: (B, H, L, dk)"""
    dk = Q.shape[-1]

    # dV: O = A @ V => dV = A^T @ dO
    dV = A.transpose(-2,-1) @ dO

    # dA: O = A @ V => dA = dO @ V^T
    dA = dO @ V.transpose(-2,-1)

    # dS via softmax backward: p.grad = p * (g - sum(p*g, keepdims=True))
    dS = A * (dA - (dA * A).sum(-1, keepdims=True))
    dS = dS / np.sqrt(dk)

    # dQ = dS @ K, dK = dS^T @ Q
    dQ = dS @ K
    dK = dS.transpose(-2,-1) @ Q
    return dQ, dK, dV

# Gradient check
np.random.seed(0)
Q=np.random.randn(1,2,4,8); K=np.random.randn(1,2,4,8); V=np.random.randn(1,2,4,8)
O, A = attn_forward(Q, K, V)
dQ, dK, dV = attn_backward(np.ones_like(O), Q, K, V, A)
# Compare dQ[0,0,0,0] with finite difference: (f(Q+eps)-f(Q-eps))/(2*eps)
Build a minimal autograd engine. Value class with +, *, relu, backward using topological sort.frontier▶

This is the core of all deep learning frameworks. PyTorch does the same thing in C++. Understanding every line of this class means understanding backprop.

Critical correctness rule: always += for grad accumulation, never =. A value used in two downstream operations must receive gradients from both (multivariate chain rule). This is the most common bug in from-scratch autograd.

The closure trick: _bwd is defined inside __mul__ and captures self, other, out by reference. This is how local Jacobians are stored without explicit node attributes.

python
class Value:
    def __init__(self, data, _prev=()):
        self.data = float(data)
        self.grad = 0.0
        self._backward = lambda: None
        self._prev = set(_prev)

    def __add__(self, other):
        other = other if isinstance(other, Value) else Value(other)
        out = Value(self.data + other.data, (self, other))
        def _bwd(): self.grad += out.grad; other.grad += out.grad
        out._backward = _bwd; return out

    def __mul__(self, other):
        other = other if isinstance(other, Value) else Value(other)
        out = Value(self.data * other.data, (self, other))
        def _bwd():
            self.grad += other.data * out.grad  # product rule
            other.grad += self.data * out.grad
        out._backward = _bwd; return out

    def relu(self):
        out = Value(max(0, self.data), (self,))
        def _bwd(): self.grad += (out.data > 0) * out.grad
        out._backward = _bwd; return out

    def __pow__(self, exp):
        out = Value(self.data**exp, (self,))
        def _bwd(): self.grad += exp * self.data**(exp-1) * out.grad
        out._backward = _bwd; return out

    def __neg__(self): return self * -1
    def __sub__(self, o): return self + (-o)
    def __truediv__(self, o): return self * o**-1
    __rmul__ = __mul__; __radd__ = __add__

    def backward(self):
        topo, seen = [], set()
        def build(v):
            if v not in seen:
                seen.add(v); [build(c) for c in v._prev]; topo.append(v)
        build(self)
        self.grad = 1.0
        for v in reversed(topo): v._backward()

# Demo: L = relu(a*b + c), verify gradients
a,b,c = Value(2.0),Value(-3.0),Value(5.0)
L = (a*b + c).relu()  # relu(-6+5) = relu(-1) = 0
L.backward()
# a.grad = b * (L.data>0) = -3*0 = 0
# b.grad = a * (L.data>0) = 2*0 = 0
# c.grad = 1 * (L.data>0) = 0
print(a.grad, b.grad, c.grad)  # 0.0, 0.0, 0.0
Implement AdamW from scratch. What is the precise difference between Adam and AdamW?frontier▶

Adam with L2 regularization: adds λ||w||² to loss. Gradient becomes g + λw. Adam divides this by √v̂, giving effective weight decay = lr·λw/√v̂. Inconsistent across parameters with different gradient magnitudes - parameters with small gradients get stronger decay than parameters with large gradients. Wrong.

AdamW (decoupled weight decay, Loshchilov & Hutter 2018): apply weight decay directly to weights BEFORE the Adam update. Effective decay = lr·wd·w, constant regardless of gradient history. The fix is literally adding one line before the Adam formula.

python
import numpy as np

class AdamW:
    def __init__(self, params, lr=3e-4, betas=(0.9, 0.999), eps=1e-8, wd=0.1):
        self.params = list(params)
        self.lr, self.b1, self.b2, self.eps, self.wd = lr, betas[0], betas[1], eps, wd
        self.m = [np.zeros_like(p) for p in self.params]
        self.v = [np.zeros_like(p) for p in self.params]
        self.t = 0

    def step(self, grads):
        self.t += 1
        for i, (p, g) in enumerate(zip(self.params, grads)):
            # === AdamW: DECOUPLED weight decay (applied before Adam step) ===
            p -= self.lr * self.wd * p          # this line makes it "W"

            # Momentum (first moment)
            self.m[i] = self.b1 * self.m[i] + (1 - self.b1) * g
            # Variance (second moment)
            self.v[i] = self.b2 * self.v[i] + (1 - self.b2) * g**2

            # Bias correction: m,v start at 0 -> underestimated in early steps
            m_hat = self.m[i] / (1 - self.b1**self.t)
            v_hat = self.v[i] / (1 - self.b2**self.t)

            # Adam update
            p -= self.lr * m_hat / (np.sqrt(v_hat) + self.eps)

# Standard LLM hyperparameters: lr=3e-4, betas=(0.9, 0.95), wd=0.1 (GPT-style)
# Note: b2=0.95 (not 0.999) for LLM training - faster adaptation
What happens when you call loss.backward in PyTorch? Trace the actual call chain.frontier▶

No magic. It is graph traversal + chain rule, implemented in C++.

  1. loss.backward → C++ torch::autograd::backward({loss}, {ones(1)})
  2. The autograd Engine performs reverse-topological BFS from the loss node
  3. Each tensor's grad_fn stores the backward op (e.g., AddmmBackward0, SoftmaxBackward0)
  4. Engine calls grad_fn(upstream_gradient) → returns list of input gradients
  5. Gradients accumulated into .grad of leaf tensors (model parameters)
  6. Non-leaf activations: their gradients are freed after use (unless retain_grad called)

Key flags:

  • retain_graph=True: don't free the graph after backward. Needed for MAML, gradient penalties in GANs.
  • create_graph=True: build differentiable graph through backward. Needed for higher-order gradients (Hessian computation).
  • torch.no_grad: stops building the graph entirely. ~40% memory savings during inference.

torch.compile (PyTorch 2.0+): captures the graph via TorchDynamo (Python bytecode tracing), optimizes via TorchInductor (generates Triton kernels), fuses ops across entire forward+backward. 1.5-3× training speedup common for LLMs. Use for production training runs.

A 7B model in fp32 = 28GB. You have a 24GB GPU. All options for inference and fine-tuning.frontier▶

Memory math first - always do this in interviews: 7B params × 4 bytes (fp32) = 28GB. 7B × 2 bytes (bf16) = 14GB. 7B × 1 byte (int8) = 7GB. 7B × 0.5 bytes (int4) = 3.5GB.

Inference options (need weights + KV cache in 24GB):

  • BF16: 14GB. Fits with 10GB for KV cache (~4K context at batch=1). Best quality. Default choice.
  • INT8 (bitsandbytes LLM.int8): 7GB weights. Quality ≈ BF16 for most tasks. 1.5-2× slower (INT8 dequant overhead). Use when 14GB is tight.
  • INT4 (GPTQ/AWQ): 3.5GB. Visible quality drop on complex reasoning/code. Consumer-hardware deployments only.
  • GGUF + llama.cpp: CPU+GPU hybrid offloading. Some layers in RAM. Works on 8GB VRAM + 32GB RAM. Slow but flexible.

Fine-tuning options:

  • Full fine-tuning: BF16 weights (14GB) + fp32 Adam optimizer states (28GB) + activations → 42GB minimum. Doesn't fit on 24GB.
  • LoRA (r=16): freeze 14GB BF16 base. Train ~130MB adapters. Adam states = 260MB. Total ≈ 14.4GB. Fits easily. Merge at inference: W' = W + αBA, zero added latency.
  • QLoRA (Dettmers 2023): base in NF4 (4-bit normalized float) = 3.5GB. LoRA adapters in BF16. Total ≈ 4GB. Fits on a 12GB RTX 3080!
  • Gradient checkpointing + LoRA: recompute activations during backward instead of storing. 10-20× activation memory reduction, +30% compute. Needed for large batch sizes.
Implement top-p nucleus sampling with temperature. Why does temperature work?▶

Temperature T: divide logits by T before softmax. T<1 → peaks sharpen (more deterministic). T>1 → flattens (more random). T→0 → greedy (argmax). T→∞ → uniform distribution.

Why it works: softmax(z/T). As T→0, relative differences in logits are amplified - the highest logit gets all the mass. As T→∞, all z/T→0, softmax→1/vocab. Temperature controls the "sharpness" of the learned distribution without changing the ranking.

Top-p (nucleus): keep the smallest set of tokens whose cumulative prob ≥ p. Dynamic vocabulary size - adapts to distribution entropy. Low-entropy step (clear next word): nucleus might be 5 tokens. High-entropy step: nucleus might be 500 tokens.

python
import torch, torch.nn.functional as F

def sample(logits, temperature=1.0, top_p=0.9, top_k=0):
    """logits: (vocab_size,) - raw unnormalized scores from model"""
    # Step 1: temperature scaling
    logits = logits / max(temperature, 1e-5)

    # Step 2: optional top-k filtering (before top-p)
    if top_k > 0:
        topk_thresh = torch.topk(logits, top_k).values[-1]
        logits[logits < topk_thresh] = float('-inf')

    # Step 3: softmax
    probs = F.softmax(logits, dim=-1)

    # Step 4: top-p nucleus filtering
    sorted_probs, sorted_idx = torch.sort(probs, descending=True)
    cumprobs = torch.cumsum(sorted_probs, dim=-1)
    # Remove tokens AFTER the cumsum exceeds top_p (shift by 1 to include the crossing token)
    to_remove = (cumprobs - sorted_probs) >= top_p
    sorted_probs[to_remove] = 0.0
    sorted_probs = sorted_probs / sorted_probs.sum()  # renormalize

    # Step 5: sample from filtered distribution
    sampled = torch.multinomial(sorted_probs, num_samples=1)
    return sorted_idx[sampled].item()
Walk through a nanoGPT training loop line by line: data → forward → loss → backward → step → eval.▶

Data: tokenize the corpus into one flat array. get_batch samples B random offsets; for each, x = tokens[i : i+T] and y = tokens[i+1 : i+T+1]. y is x shifted by one - every position predicts the next token.

Forward: token embeddings + positional embeddings → N transformer blocks (pre-LN: LN → attention → residual, LN → MLP → residual) → final LN → logits via the LM head (weights usually tied to the token embedding).

Loss: F.cross_entropy(logits.view(-1, V), y.view(-1)) - mean negative log-likelihood across all B·T positions. Every position is a training signal, which is why a single sequence teaches T predictions.

Backward + step:

for step in range(max_steps): x, y = get_batch('train') logits, loss = model(x, y) optimizer.zero_grad(set_to_none=True) loss.backward torch.nn.utils.clip_grad_norm_(model.parameters, 1.0) lr = get_lr(step) # warmup then cosine decay for g in optimizer.param_groups: g['lr'] = lr optimizer.step # AdamW

Eval: periodically, under torch.no_grad with model.eval, average the loss over a few train and val batches, then switch back to model.train.

Details that matter: weight tying, bias-free LN/Linear, gradient accumulation for a large effective batch, mixed precision (bf16 autocast), and the warmup+cosine LR schedule.

the story: you built MicroGPT in ~218 lines. Be ready to defend every line - especially why y is x shifted by one and why the loss means over all T positions.

Implement BPE (byte-pair encoding) tokenizer from scratch.▶

Start from bytes (256 tokens). Count adjacent pairs, merge the most frequent, repeat until vocab size target.

  • Encode: greedily apply merges left-to-right on bytes
  • Decode: map token ids back through merge table to bytes → UTF-8 string
  • Why bytes: no UNK token; any Unicode string tokenizes

Karpathy's minBPE and GPT-2 tokenizer follow this exactly. Be ready to write get_stats, merge, and the training loop.

Trace tensor shapes through one GPT-2 block forward pass (batch B, seq T, model dim C).▶

Input x: [B, T, C]

  • Q,K,V projections: each [B, T, C] (with n_head splits → [B, n_h, T, d_h] where d_h = C/n_h)
  • Attention scores: [B, n_h, T, T]
  • Attention out: [B, n_h, T, d_h] → merge heads → [B, T, C]
  • FFN: [B, T, C] → [B, T, 4C] → [B, T, C]
  • Residual paths keep [B, T, C] throughout

The only O(T²) tensor is the attention map. Everything else is O(T).

Why tie input token embeddings to the LM head weights?▶

Both map between vocab space (V) and model dim (C). Tying uses one matrix W ∈ ℝ^{V×C} for embedding lookup and logits = x @ W^T.

  • Saves ~V×C parameters (for GPT-2 vocab 50K, C=768 → ~38M params)
  • Regularizes: input and output representations must agree
  • Standard in GPT-2/3, LLaMA, nanoGPT

Untied head can help when vocab is huge or modalities differ, but for vanilla LMs tying is the default.

Training loss is NaN after ~100 steps. Walk through the debug checklist.▶
  1. Print one batch: labels in range? Inputs finite?
  2. Lower LR 10× - exploding gradients?
  3. clip_grad_norm_(..., 1.0) - does loss stabilize?
  4. Overfit 1 batch: if still NaN, bug in loss/forward (wrong ignore_index, div by zero in LN)
  5. Check mixed precision: loss scaling, bf16 vs fp16 overflow
  6. Inspect per-layer grad norms - which layer blows up first?

Karpathy-style: always overfit one batch before scaling up. If you can't hit ~0 loss on 1 batch, the code is wrong.

Implement LayerNorm forward and backward in NumPy (no autograd).▶

Forward per row x ∈ ℝ^D: μ = mean(x), σ² = var(x), x̂ = (x−μ)/√(σ²+ε), y = γ⊙x̂ + β

Backward:

  • dγ = Σ dy⊙x̂, dβ = Σ dy
  • dx̂ = dy⊙γ
  • dx = (1/√(σ²+ε)) · (dx̂ − mean(dx̂) − x̂⊙mean(dx̂⊙x̂))

Same formula as BatchNorm backward but reduction is over features, not batch - no running stats, works at batch=1.

Estimate Pi with Monte Carlo - implement it and explain the variance.▶

Sample (x,y) ~ Uniform[0,1]². Count fraction with x²+y² ≤ 1 → estimates π/4.

Variance shrinks as O(1/N). After N=10⁶ samples, std error ≈ 0.0016 on π.

Karpathy uses this to test if candidates understand probability, loops, and convergence - not memorized LeetCode.

Matrix multiply: O(n³) naive, then how do you actually make it fast on a GPU?▶

Levels:

  1. BLAS: a @ b calls cuBLAS - never write naive triple loops in production
  2. Tiling: load blocks into shared memory to raise arithmetic intensity
  3. Strassen: O(n^2.807) - rarely wins except huge matrices due to constants
  4. Tensor Cores: WMMA ops for FP16/BF16 tile matmuls

Interview move: write naive matmul, then explain why it's memory-bound and what cuBLAS does differently.

Your network won't overfit a single batch. Walk the debug checklist.▶

Overfitting one batch to ~zero loss is the first sanity check before any real training. If it fails, something is fundamentally broken.

  • Loss not decreasing at all: learning rate too low, or gradients not flowing (check requires_grad, a detached tensor, or .data used by accident).
  • Loss is NaN: LR too high, log of zero, or a bad init. Lower LR, add eps, check the loss formula.
  • Loss plateaus above zero: the label and prediction are misaligned (off-by-one in targets), the output layer is wrong (logits vs probs into the loss), or a dead activation.
  • Forgot optimizer.zero_grad(): gradients accumulate across steps.
  • Model in eval mode (dropout/BN off) or data accidentally shuffled each step so it is never the same batch.

If you cannot drive a single batch to near-zero loss, do not move on. The model has the capacity to memorize 32 examples; failing means a wiring bug, not a learning problem.

How do you gradient-check an autograd implementation?▶

Verify analytical gradients against numerical ones. For each parameter θ, the centered finite difference:

grad_numerical ≈ (f(θ + ε) - f(θ - ε)) / (2ε), with ε ≈ 1e-5.

Compare to the backward-pass gradient with a relative error: |g_analytic - g_numeric| / (|g_analytic| + |g_numeric| + 1e-8). Below ~1e-7 is correct; above ~1e-3 means a bug.

  • Use double precision (float64) so ε rounding does not dominate.
  • Use the centered difference, not one-sided; the error is O(ε²) instead of O(ε).
  • Watch non-differentiable kinks (ReLU at 0) - perturbation can straddle the kink and look wrong.

This is exactly how you trust a from-scratch autograd before training anything with it.

Explain tensor strides, views, and what .contiguous() does.▶

A tensor is a flat 1-D block of memory plus metadata: shape and strides (how many elements to step in memory to move one index along each dimension).

Views (transpose, slice, reshape-when-possible) return a new tensor that points at the same memory with different shape/strides - no copy. So x.T is free; it just swaps the strides.

The catch: after a view, memory order no longer matches index order. Some ops (and .view()) need a contiguous layout. .contiguous() forces a copy into row-major order so strides match the shape again.

Why it matters: understanding strides explains why .view() sometimes errors and .reshape() does not (reshape copies if needed), and why a transpose followed by view fails.

Why does weight initialization matter? What breaks with bad init?▶

Init sets the scale of activations and gradients before any learning. Get it wrong and signal either explodes or vanishes as it flows through layers.

  • Too large: activations saturate (tanh/sigmoid) or blow up; gradients explode; loss goes NaN.
  • Too small: activations shrink toward zero layer by layer; gradients vanish; deep layers never learn.
  • All zeros: every neuron in a layer computes the identical thing and gets the identical gradient - they never differentiate. Symmetry is never broken.

The fix: keep variance roughly constant across layers. Xavier/Glorot (var ≈ 1/n) for tanh, He/Kaiming (var ≈ 2/n) for ReLU since ReLU zeros half the inputs. Modern transformers also scale residual-branch init down by depth so the residual stream does not blow up.

.data vs .detach() vs .item() vs torch.no_grad() - precise differences.▶
  • .item() - pulls a single-element tensor out as a Python number. Breaks the graph because it leaves tensor-land entirely. Use for logging a scalar loss.
  • .detach() - returns a new tensor sharing the same storage but cut out of the autograd graph (requires_grad=False). The safe way to use a value without backprop through it.
  • .data - the old, unsafe way. Also gives a detached view but bypasses autograd's safety checks; in-place edits via .data can silently corrupt gradients. Avoid it.
  • with torch.no_grad(): - a context that disables graph building for everything inside. Use for inference and eval loops; saves memory and time by not tracking ops.

Interview trap: using .data instead of .detach() in a training loop is a classic source of "my gradients are silently wrong" bugs.

What is the residual stream? Why do residual connections make deep nets trainable?▶

A residual block computes y = x + f(x): the input is added back to the sublayer output. Stack these and there is a continuous additive path from input to output - the residual stream.

Why it trains: the gradient of y = x + f(x) is 1 + f'(x). The "+1" means gradient always flows straight back through the skip path, even if f'(x) is tiny - so no vanishing gradient through depth. The network learns a residual (a delta on the identity) instead of the full mapping, which is an easier target.

Mechanistic view (transformers): each attention and MLP block reads from and writes to the shared residual stream. The stream is the model's working memory; layers communicate by adding features into it.

This is why ResNets scaled to 100+ layers and why every transformer is residual.

Hardware & GPUs
What does a CUDA matmul kernel actually do? Why are Tensor Cores faster than regular CUDA cores?frontier▶

Naive CUDA matmul: each thread computes C[i,j] = Σ_k A[i,k]·B[k,j]. Reads A[i,:] and B[:,j] from global memory (HBM) per output element. HBM bandwidth ≈ 2TB/s on A100. Kernel is severely memory-bound - compute sits idle waiting for data.

Tiled matmul (what cuBLAS/Flash Attention actually do):

  1. Divide A, B into tiles that fit in SRAM/shared memory (~32-48KB per SM)
  2. Each thread block loads one tile of A + one tile of B into SRAM (fast: ~19TB/s)
  3. Compute partial products using SRAM data only - no HBM access during computation
  4. Advance tile pointer across K dimension, accumulate, repeat
  5. Write final C block back to HBM once

Arithmetic intensity: naive = O(1) FLOP/byte. Tiled = O(tile_size) FLOP/byte → compute-bound instead of memory-bound.

Tensor Cores: dedicated hardware units that compute a 16×16×16 matmul (D = A×B+C) in one clock cycle. Regular CUDA cores: one multiply-accumulate per clock (scalar). Tensor Cores: 256 multiply-accumulate per clock. ~8× throughput for FP16, BF16, INT8, FP8. Flash Attention uses Tensor Cores via WMMA instructions (cute/CUTLASS library).

Roofline model: LLM inference (batch=1) is memory-bound. LLM training (large batch) is compute-bound. This determines the optimization strategy - inference: reduce bytes transferred (quantization, KV cache). Training: maximize GPU utilization (batch size, kernel fusion).

Why is LLM inference memory-bandwidth bound? What optimization strategies follow from this?frontier▶

Arithmetic intensity math: for weight W ∈ ℝ^{d×d} and input x ∈ ℝ^d (batch=1):

  • FLOPs: 2d² (matmul)
  • Memory: d² × 2 bytes (read weights in BF16) + d × 2 bytes (read/write x) ≈ 2d² bytes
  • Arithmetic intensity ≈ 2d²/(2d²) = 1 FLOP/byte

A100 ratio: 312 TFLOP/s ÷ 2 TB/s = 156 FLOP/byte. Since 1 ≪ 156, inference at batch=1 is deeply memory-bound. GPU compute sits idle ~99% of the time waiting for weights from HBM.

Optimization strategies that follow directly:

  • INT4 quantization: 4× fewer bytes per weight → 4× faster. Not about fewer FLOPs - about less HBM bandwidth consumed.
  • KV cache: avoids re-reading K,V for all previous tokens each step. But KV cache itself must be read → as context grows, KV cache reading dominates.
  • Speculative decoding: draft model (small, bandwidth-light) proposes K tokens → big model verifies in one pass (now effective batch=K → better arithmetic intensity).
  • Continuous batching (vLLM): N sequences together → arithmetic intensity scales with N → becomes compute-bound at high N → full GPU utilization.

Karpathy's rule: "If they are not batching, they are not using the GPU." The difference between batch=1 (memory-bound, 1% utilization) and batch=512 (compute-bound, 80% utilization) is the fundamental lever of LLM serving economics.

What is arithmetic intensity? Explain the roofline model.frontier▶

Arithmetic intensity = FLOPs performed ÷ bytes moved from memory. It tells one whether a kernel is limited by compute or by bandwidth.

Roofline: achievable FLOP/s = min( peak compute, bandwidth × arithmetic intensity ). Plot intensity on the x-axis: below the "ridge" they are memory-bound (the slanted bandwidth roof), above it they are compute-bound (the flat compute roof).

Where the ridge is: peak_FLOPs ÷ peak_bandwidth. For an A100 ≈ 312 TFLOP/s ÷ 2 TB/s ≈ 156 FLOP/byte. You need intensity above ~156 to saturate the compute units.

Examples:

  • Large square matmul (M×K)·(K×N): 2MKN FLOPs, ~(MK+KN+MN) bytes → high intensity → compute-bound. GPUs love this.
  • Matrix×vector (batch-1 decode): 2KN FLOPs, ~KN weight bytes → intensity ≈ 2 → deeply memory-bound. The GPU sits mostly idle.

This single framework explains FlashAttention, fused kernels, and why batching matters - they all push intensity up so you stop wasting the compute roof.

Estimate the tokens/sec ceiling for a 7B model on one A100. Show the arithmetic.frontier▶

Single-stream decode is memory-bound: to generate one token you read the entire model out of HBM once.

  • 7B params × 2 bytes (fp16) = 14 GB read per token
  • A100 HBM bandwidth ≈ 2 TB/s
  • floor latency/token = 14 GB ÷ 2 TB/s = 7 ms → ceiling ≈ 143 tok/s (batch=1, ignoring KV cache and overhead)

Quantization moves the ceiling: int8 (7 GB) → ~3.5 ms → ~285 tok/s; int4 (3.5 GB) → ~570 tok/s. That is why quantization is the lever for single-stream latency - it directly cuts the bytes you read per token.

Reality check: production hits ~60-70% of this roofline. But if the measured throughput is 10× under the estimate, you have a software bug, not a hardware limit. Knowing the ceiling is how one tell the difference.

GPU memory hierarchy - the numbers that actually matter.frontier▶

Top to bottom, fast/small to slow/large (A100-class):

  • Registers: ~KBs per thread, fastest, private to a thread.
  • SRAM / shared memory: ~192 KB per SM (~20 MB on-chip total), ~19 TB/s. The scratchpad FlashAttention tiles Q/K/V into.
  • HBM (global memory): 40-80 GB, ~2 TB/s. Where weights and activations live - and the bottleneck for almost everything.
  • Host RAM over PCIe: ~64 GB/s - roughly 30× slower than HBM. CPU/disk offload is a last resort, not a plan.

One rule explains most GPU optimization: minimise HBM traffic. FlashAttention, kernel fusion, activation recomputation, and quantization are all the same idea - keep data in SRAM/registers and stop round-tripping to HBM.

What is MFU (Model FLOPs Utilization)? What caps training throughput?frontier▶

MFU = achieved FLOP/s ÷ hardware peak FLOP/s. A well-tuned large training run lands around 40-55%; above 50% is excellent.

Training FLOPs per token ≈ 6N for N parameters (≈2N forward + ≈4N backward). So tokens/sec ≈ MFU × peak_FLOPs ÷ 6N - a quick way to sanity-check a training run.

What kills MFU:

  • Memory-bandwidth stalls and low arithmetic intensity (small micro-batch / short sequences)
  • Communication: all-reduce / all-gather in data and tensor parallelism
  • Pipeline bubbles in pipeline parallelism
  • Kernel launch overhead and poor overlap of compute with comms

Levers: larger micro-batch, fused kernels (FlashAttention), the right parallelism layout, gradient-checkpoint tradeoffs, and overlapping communication with compute (FSDP/ZeRO prefetch).