Skip to content
samit
Interviews/Reading List

Reading List

2017
Attention Is All You Need
Vaswani et al. · The Transformer paper. Know every component cold.
arXiv ↗
2020
Scaling Laws for Neural Language Models
Kaplan et al. (OpenAI) · Defines how model size, data, compute relate to loss.
arXiv ↗
2022
Training Compute-Optimal LLMs (Chinchilla)
Hoffmann et al. (DeepMind) · Corrects scaling law - 20 tokens per parameter optimal.
arXiv ↗
2017
Deep Residual Learning for Image Recognition
He et al. (MSRA) · ResNets. Why skip connections enable very deep networks.
arXiv ↗
2012
ImageNet Classification with Deep CNNs (AlexNet)
Krizhevsky et al. · Started deep learning revolution. Know the architecture.
arXiv ↗
2015
Unreasonable Effectiveness of RNNs
Karpathy · Intuition building. Read before understanding LSTMs.
arXiv ↗
2015
Understanding LSTM Networks
Olah (colah) · Best visual explanation of LSTMs. Required reading.
arXiv ↗
2021
An Image is Worth 16x16 Words (ViT)
Dosovitskiy et al. · Transformers for vision - architecture and pretraining insights.
arXiv ↗
2022
FlashAttention: Fast & Memory-Efficient Exact Attention
Dao et al. · IO-aware attention. How to actually implement Transformers efficiently.
arXiv ↗
2021
LoRA: Low-Rank Adaptation of LLMs
Hu et al. · Parameter-efficient fine-tuning. Ubiquitous in practice.
arXiv ↗
2022
Training LLMs to Follow Instructions (InstructGPT)
Ouyang et al. (OpenAI) · Original RLHF paper. How ChatGPT was made.
arXiv ↗
2023
Direct Preference Optimization (DPO)
Rafailov et al. · Simpler RLHF alternative. Used in most modern models.
arXiv ↗
2022
Constitutional AI
Bai et al. (Anthropic) · How Claude was trained. Key for Anthropic interviews.
arXiv ↗
2021
CLIP: Connecting Text and Images
Radford et al. (OpenAI) · Contrastive multimodal learning foundation. Zero-shot vision.
arXiv ↗
2022
Whisper: Robust Speech Recognition
Radford et al. (OpenAI) · Encoder-decoder Transformer for speech. Audio chunk processing.
arXiv ↗
2023
Segment Anything Model (SAM)
Kirillov et al. (Meta) · Promptable segmentation. Key CV model of 2023.
arXiv ↗
2020
Denoising Diffusion Probabilistic Models (DDPM)
Ho et al. · Diffusion model foundations. Forward/reverse process.
arXiv ↗
2020
Language Models are Few-Shot Learners (GPT-3)
Brown et al. (OpenAI) · In-context learning, scaling. The paper that changed everything.
arXiv ↗
2023
LLaMA 2: Open Foundation and Fine-Tuned Chat Models
Touvron et al. (Meta) · Open source LLM architecture. GQA, RLHF at scale.
arXiv ↗
2019
GPipe: Efficient Training using Pipeline Parallelism
Huang et al. (Google) · Ilya's reading list. Pipeline parallelism foundations.
arXiv ↗
2014
Neural Turing Machines
Graves et al. · Ilya's reading list. Memory + attention before Transformers.
arXiv ↗
2013
Auto-Encoding Variational Bayes (VAE)
Kingma & Welling · VAE foundations. ELBO, reparameterization trick.
arXiv ↗
2023
QLoRA: Efficient Finetuning of Quantized LLMs
Dettmers et al. · Fine-tune 65B model on single GPU. Critical for startup ML.
arXiv ↗
2023
Mistral 7B
Jiang et al. · GQA + sliding window attention. Dense model efficiency.
arXiv ↗
2024
DeepSeek-V2
DeepSeek AI · MLA attention, fine-grained MoE, GRPO training. Modern architecture.
arXiv ↗