Transformer & LLM Architecture
This page covers how a modern large language model is built: how text becomes tokens and vectors, what attention actually computes (shapes included), why the KV cache dominates serving cost and how MQA, GQA and MLA shrink it, how position is encoded and context extended, what is inside a decoder block, how to count parameters and FLOPs on a whiteboard, how Mixture-of-Experts works, and where hybrids with state-space and linear-attention layers stand in 2026. Interviewers use this material to tell apart people who have called an LLM from people who can reason about its memory, compute and failure modes. That reasoning drives real decisions about context length, serving hardware, fine-tuning and cost.
TL;DR: the 8–12 things to be able to say out loud
- Transformers won because attention lets every token read every earlier token in one parallel step. Training parallelizes over the sequence (RNNs can't), and the path between any two tokens has length 1. The price is \(O(n^2)\) compute and memory in sequence length.
- Attention: \(\mathrm{softmax}(QK^\top/\sqrt{d_k} + M)\,V\). The \(1/\sqrt{d_k}\) keeps logits at unit variance so the softmax doesn't saturate. \(M\) is the causal mask (\(-\infty\) above the diagonal).
- Tokenization (byte-level BPE in most LLMs) is a lossy-feeling but lossless compression step. It is behind odd failures: letter counting, arithmetic, glitch tokens, and non-English "token tax".
- Decoding is memory-bandwidth bound, and the KV cache is the main memory cost: \(2 \cdot L \cdot n_{kv} \cdot d_h \cdot \text{bytes}\) per token. For Llama-3-70B (GQA, 8 KV heads) that is about 320 KiB/token, or about 40 GiB for one 128K-token sequence.
- MQA shares one K/V head across all query heads. GQA shares one per group (the standard today). MLA (DeepSeek) caches a small low-rank latent instead of the full K/V and gets a roughly 93% cache reduction.
- RoPE rotates pairs of Q/K dimensions by an angle proportional to position, so the dot product depends only on relative offset. Context extension (PI, NTK-aware, YaRN) rescales those frequencies.
- A modern block is pre-norm (RMSNorm) → attention → residual add → RMSNorm → SwiGLU FFN → residual add. Think of the residual stream as a shared bus that every layer reads from and writes to.
- Parameters ≈ \(12 L d^2\) (plus embeddings). Forward FLOPs ≈ \(2N\) per token, training ≈ \(6N\) per token.
- Almost all frontier LLMs are decoder-only. Encoders (BERT-style) still dominate embeddings, rerankers and classifiers.
- MoE: a router sends each token to the top-k of E expert FFNs. Total parameters set capacity and memory; active parameters set FLOPs per token. Load balancing is the core difficulty (aux loss, or DeepSeek's bias-based aux-loss-free method). Shared experts are now common.
- Long context: sliding-window/local layers interleaved with global layers, attention sinks, sparse attention (DeepSeek DSA/CSA), and since 2025 hybrid linear-attention/SSM layers (Gated DeltaNet, Mamba-2) mixed with a minority of full-attention layers.
- Vision-language models: a ViT cuts the image into patches, a projector maps patch embeddings into the LLM's token space, and the LLM attends over them like text tokens.
1. Why transformers won
Before 2017, sequence modelling meant RNNs/LSTMs, sometimes with attention bolted on, and some convolutional models (ByteNet, ConvS2S). The Transformer kept the attention, dropped the recurrence, and stacked self-attention with position-wise feed-forward layers. Vaswani+ 2017 It won on three axes that matter more as scale grows.
| Property | RNN / LSTM | CNN (1-D, dilated) | Self-attention |
|---|---|---|---|
| Sequential ops per layer (training) | \(O(n)\): step t needs step t−1 | \(O(1)\) | \(O(1)\): every position at once |
| Max path length between two tokens | \(O(n)\) | \(O(\log_k n)\) (dilated) or \(O(n/k)\) | \(O(1)\) |
| Compute per layer | \(O(n d^2)\) | \(O(k n d^2)\) | \(O(n^2 d + n d^2)\) |
| Inference state per step | Fixed-size hidden state | Receptive-field window | Grows with context (KV cache) |
| Hardware fit | Poor: small sequential matmuls | Good | Excellent: big dense matmuls |
The complexity rows follow Table 1 of the original paper. Vaswani+ 2017 The deeper reasons:
- Parallel training. With teacher forcing, a causal transformer computes the loss for all \(n\) positions of a sequence in one forward pass, as a handful of large matrix multiplications. GPUs and TPUs are matmul machines, so you get far more useful FLOPs per dollar than with an LSTM's step-by-step loop.
- Gradient paths. Token 1 can affect token 4,000 through one attention hop. An RNN needs 4,000 multiplicative steps, which is where vanishing gradients come from. Residual connections plus attention give short, well-conditioned paths.
- Scaling behaviour. Transformer loss falls as a smooth power law in parameters, data and compute over many orders of magnitude. LSTMs plateau earlier, especially on later tokens in the context. Kaplan+ 2020 That predictability made it rational to spend very large budgets.
- Generality. Attention carries no built-in assumption of locality (unlike convolutions) or of order (positions are injected separately). The same block works on text, image patches, audio frames and actions.
The cost is that attention is quadratic in sequence length, and at inference you must keep K and V for every past token. Most of the architecture work since 2019 (GQA, MLA, sliding windows, FlashAttention, SSM hybrids) is about paying that cost more cheaply.
An RNN compresses the whole past into a fixed-size vector and hopes the needed detail survives. A transformer keeps the entire past around, uncompressed, and looks things up by content at every step. That costs memory, but recall is close to perfect. SSMs and linear attention go back toward a fixed-size state, which is why their weak spot is precise retrieval.
"Why did transformers replace LSTMs?" A weak answer stops at "attention is better." A strong answer gives: (1) parallelism over the sequence during training, which is what matters for GPU utilization; (2) \(O(1)\) path length for long-range dependencies; (3) clean power-law scaling. It then names the trade-off: \(O(n^2)\) attention and a KV cache that grows linearly at inference, and contrasts that with the RNN's constant-size state, which is why SSMs are back in the conversation.
- Attention Is All You Need: the original paper. Table 1 is the complexity comparison above.
- The Illustrated Transformer (Alammar): the best visual walk-through of the original architecture.
- Scaling Laws for Neural Language Models: includes the transformer vs LSTM scaling comparison.
2. Tokenization
A model works on integer IDs from a fixed vocabulary \(V\). The tokenizer maps text to IDs and back. It is trained separately, before the model, and then frozen. Changing it later means re-learning the embeddings.
2.1 BPE: the workhorse
Byte-Pair Encoding was adapted from compression to NMT subwords. Sennrich+ 2016 Training it works like this:
- Start with a base alphabet (characters, or the 256 byte values).
- Count every adjacent pair of symbols in the training corpus.
- Merge the most frequent pair into a new symbol and record the merge rule.
- Repeat until the vocabulary reaches the target size (for example 32K, 128K or 200K).
Encoding new text replays the merge rules in priority order. Frequent words become single tokens. Rare words split into pieces ("tokenization" → ["token", "ization"]). Nothing is ever out-of-vocabulary, because you can always fall back to the base symbols.
2.2 Byte-level BPE
GPT-2 ran BPE over UTF-8 bytes rather than Unicode characters, so the base vocabulary is exactly 256 symbols and any string (emoji, code, binary junk) can be encoded with no <unk>. Radford+ 2019 It also applies a regex pre-tokenizer that first splits text into chunks (words, numbers, punctuation, whitespace runs), so merges never cross those boundaries and you don't get tokens like "dog.". OpenAI's tiktoken (cl100k, o200k vocabularies) and Llama 3's 128K tokenizer follow this design. tiktoken Llama 3 2024
2.3 SentencePiece and Unigram
SentencePiece is a library, not an algorithm. It treats input as a raw character stream with whitespace encoded as a visible symbol (▁), so it needs no language-specific pre-tokenizer. That matters for Chinese and Japanese, which have no spaces. It can train either BPE or Unigram models. Kudo & Richardson 2018
Unigram LM runs BPE in reverse. It starts from a large candidate vocabulary and iteratively prunes the tokens whose removal hurts corpus likelihood least, under a model where each token has an independent probability. To encode, it finds the most probable segmentation with Viterbi. Because it defines a distribution over segmentations, you can sample different splits during training (subword regularization), which improves robustness. Kudo 2018 T5 and many earlier multilingual models used SentencePiece-Unigram. Llama 1/2 used SentencePiece-BPE with byte fallback, and Llama 3 moved to a tiktoken-style byte-level BPE.
| Scheme | Training | Base units | Notable users |
|---|---|---|---|
| WordPiece | Greedy merges chosen by likelihood gain rather than raw frequency | Characters, with a ## continuation prefix | BERT |
| BPE (char-level) | Frequency-based merges | Unicode chars | Early NMT, GPT-1 |
| Byte-level BPE | Frequency merges over bytes, with regex pre-split | 256 bytes | GPT-2/3/4, gpt-oss, Llama 3, Qwen, DeepSeek |
| SentencePiece-BPE (+byte fallback) | Merges over raw stream with ▁ | Chars, bytes for unknowns | Llama 1/2, Mistral 7B |
| SentencePiece-Unigram | Prune by likelihood loss | Chars | T5, mT5, ALBERT |
| Byte / patch (tokenizer-free) | None, or learned dynamic patching | Bytes | ByT5, Byte Latent Transformer (research) |
2.4 Vocabulary size trade-offs
Vocabulary sizes have grown: 32K (Llama 2, Mistral 7B) → 128K (Llama 3) → about 150K (Qwen) → about 200K (o200k, gpt-oss) → 256K–262K (Gemma). OpenAI 2025 Gemma 3 2025 The trade-off:
Bigger vocab
- Fewer tokens per text (better compression). That means a longer effective context, cheaper inference per character, and fewer decode steps.
- Better coverage of non-English scripts and code.
- Llama 3 reports roughly 15% fewer tokens than Llama 2's tokenizer on English. (~15%)
Costs
- Embedding and unembedding matrices are \(V \times d\) each. At \(V\)=256K and \(d\)=2,048 that is about 0.5B parameters, a big share of a small model.
- The output softmax over \(V\) costs FLOPs and memory (logits are \(n \times V\)).
- Rare tokens get few gradient updates, so they end up under-trained (glitch tokens).
2.5 Tokenization failure modes (interview favourites)
- Character-level tasks. "How many r's in strawberry?" The model sees perhaps
["str", "awberry"]and has to have memorized the spelling of each token. Reversal, anagrams and rhyme run into the same issue. - Arithmetic. Numbers get split inconsistently (
"12345"might be["123", "45"]), which misaligns place value. Llama 3 splits numbers into at most 3-digit chunks, and some tokenizers (Gemma's, for example) split every digit. Gemma 3 2025 - Glitch / under-trained tokens. Tokens that existed in the tokenizer's training corpus but almost never in the model's (Reddit usernames, for example) have near-random embeddings and can trigger bizarre outputs. There are automated ways to detect them. Land & Bartolo 2024
- Whitespace and leading-space sensitivity.
"hello"and" hello"are different tokens. Trailing spaces in prompts, odd indentation in code, and prompt templates that leave a dangling space can all hurt quality. - Token tax. Non-Latin scripts can need several times more tokens for the same content, so users pay more and get less effective context. The multiplier depends on the tokenizer and language. often 2–4×
- Prompt boundary / token healing. If a prompt ends mid-word, the most natural continuation might need a token that straddles the boundary. Constrained decoding and JSON-mode implementations have to handle partial tokens.
- Security. Special tokens such as
<|im_start|>must never be produced from user text. If they are, users can forge chat-role boundaries.
Common probes: "Why can't the model count letters?", "Why are prices different for Japanese?", "What happens if you change tokenizers on a pretrained model?" A strong answer explains that the model never sees characters, only IDs. It mentions under-trained token embeddings. For tokenizer swaps it says that embeddings are tied to IDs, so you must re-initialize them (for example as averages of the old sub-token embeddings) and continue pretraining.
Tokenizer-free models (byte-level with learned dynamic patching, such as the Byte Latent Transformer) were an active research area in 2025. Pagnoni+ 2024 As of this writing, mainstream open models still use byte-level BPE. Check whether that has changed.
- Let's build the GPT Tokenizer (Karpathy): two hours that cover every failure mode above.
- minbpe: a minimal, readable byte-level BPE implementation.
- Hugging Face tokenizer summary: BPE vs WordPiece vs Unigram side by side.
3. Embeddings and the output head
The embedding matrix \(E \in \mathbb{R}^{V \times d}\) is a lookup table: token ID \(i\) becomes row \(E_i\), a \(d\)-dimensional vector (\(d\) = 4,096 for Llama-3-8B, 7,168 for DeepSeek-V3). A batch of token IDs of shape [B, n] becomes a tensor [B, n, d], the residual stream, which flows through every layer.
At the top, the final hidden state \(h \in \mathbb{R}^d\) is projected by the unembedding (LM head) \(W_U \in \mathbb{R}^{d \times V}\) to logits, then softmaxed to a next-token distribution:
$$ p(x_{t+1} \mid x_{\le t}) = \mathrm{softmax}(h_t W_U) $$- Weight tying (\(W_U = E^\top\)) saves \(V d\) parameters. Small models usually tie (Gemma, Llama 3.2 1B/3B). Large models usually don't, because the saving is negligible and untied heads perform slightly better.
- Embeddings are learned, not pre-computed like word2vec. Early layers turn each token embedding into a context-dependent representation, so the vector for "bank" at layer 20 differs between "river bank" and "bank loan".
- Some models scale embeddings by \(\sqrt{d}\) at input (original Transformer, Gemma) so their magnitude matches the rest of the residual stream.
- LLM embeddings vs embedding models: a "sentence embedding" for RAG usually comes from a separately trained encoder (or decoder) with a pooling step and contrastive training. It is not the input embedding table. Raw input embeddings know nothing about context.
Saying "the model's embedding of a document" when you mean the input embedding table. A useful document vector comes from pooling hidden states of a model trained with a contrastive objective (for example SimCSE-style training). Input embeddings alone are just per-token lookups.
4. Scaled dot-product attention, step by step
Notation: batch \(B\), sequence length \(n\), model width \(d\), heads \(h\), head dimension \(d_h = d/h\) (usually 64–128). Take one head and one sequence (drop \(B\)), with input \(X \in \mathbb{R}^{n \times d}\).
4.1 The computation
- Project: \(Q = XW_Q,\; K = XW_K,\; V = XW_V\), with \(W_\ast \in \mathbb{R}^{d \times d_h}\). Shapes: \(Q, K, V \in \mathbb{R}^{n \times d_h}\).
Q is "what am I looking for", K is "what do I contain (for matching)", V is "what I hand over if selected". - Scores: \(S = QK^\top / \sqrt{d_h} \in \mathbb{R}^{n \times n}\). Entry \(S_{ij}\) is how much query position \(i\) matches key position \(j\).
- Mask: \(S \leftarrow S + M\), where \(M_{ij} = -\infty\) if \(j > i\) (future) and 0 otherwise. Padding positions are masked the same way.
- Normalize: \(A = \mathrm{softmax}_{\text{row}}(S)\). Each row is a probability distribution over positions \(\le i\).
- Mix: \(O = AV \in \mathbb{R}^{n \times d_h}\). Each output row is a weighted average of value vectors.
4.2 Why divide by √dh?
Suppose the components of \(q\) and \(k\) are independent with mean 0 and variance 1. Then \(q \cdot k = \sum_{i=1}^{d_h} q_i k_i\) is a sum of \(d_h\) terms, each with variance 1, so
$$ \mathrm{Var}(q\cdot k) = d_h, \qquad \mathrm{std}(q\cdot k) = \sqrt{d_h}. $$With \(d_h = 128\) the raw logits have a standard deviation of about 11. A softmax over logits that spread out is nearly one-hot, so its gradient (\(\partial a_i / \partial s_j = a_i(\delta_{ij} - a_j)\)) is close to zero almost everywhere, and learning stalls. Dividing by \(\sqrt{d_h}\) brings the logits back to unit variance, keeping the softmax in a responsive range at initialization. Vaswani+ 2017 Later, the model can learn to sharpen attention by growing \(W_Q, W_K\). Some recent models add QK-norm (RMSNorm on q and k) to keep logits bounded throughout training, a stability fix used in Gemma 3, Qwen3 and others. Gemma 3 2025 Qwen3 2025
4.3 Causal masking
A decoder predicts token \(t+1\) from tokens \(\le t\). During training we feed the whole sequence at once and get \(n\) predictions in parallel. Without a mask, position \(i\) could look at token \(i+1\) and copy the answer. Setting future logits to \(-\infty\) before the softmax gives them exactly zero weight. The same mask makes KV caching valid: past keys and values never change when new tokens arrive, so at decode time you compute Q, K and V only for the new token and attend over the cached K and V.
Attention is a differentiable, soft hash-map lookup. The query is the search key, keys are the index, values are the payload, and the softmax returns a blend of matches instead of a single hit. Multi-head attention runs several independent lookups with different "questions" (syntax, coreference, position) in parallel.
4.4 Prefill vs decode
| Prefill (process the prompt) | Decode (generate one token) | |
|---|---|---|
| Q shape per head | \([n, d_h]\) | \([1, d_h]\) |
| Work | Large matmuls, \(O(n^2 d)\) attention | Matrix-vector products, \(O(n d)\) attention per token |
| Bottleneck | Compute-bound | Memory-bandwidth-bound: every weight and the whole KV cache are read for each token |
| Metric | Time-to-first-token (TTFT) | Inter-token latency, tokens/s |
This asymmetry explains why KV-cache size, batching and quantization dominate serving economics (covered on the inference page).
"Walk me through attention with shapes" is a near-universal question. Write the shapes at every step, explain the \(\sqrt{d_h}\) argument through variance and softmax saturation (not "for numerical stability" in a vague sense), and say why the causal mask also makes KV caching possible. Bonus points for mentioning that FlashAttention never materializes the \(n\times n\) matrix in HBM: it tiles the computation in SRAM with an online softmax, which makes memory \(O(n)\) while compute stays \(O(n^2)\). Dao+ 2022
5. Multi-head attention and complexity
Rather than one attention of width \(d\), use \(h\) heads of width \(d_h = d/h\) in parallel, concatenate their outputs and mix them with \(W_O \in \mathbb{R}^{d \times d}\):
$$ \mathrm{MHA}(X) = \mathrm{Concat}(\mathrm{head}_1,\dots,\mathrm{head}_h)\,W_O,\quad \mathrm{head}_i = \mathrm{Attention}(XW_Q^{(i)}, XW_K^{(i)}, XW_V^{(i)}) $$The parameter count equals one big head (\(4d^2\) for Q, K, V and O), but each head can learn a different attention pattern. One softmax can only produce one weighted average per position, while \(h\) heads produce \(h\) different averages. Interpretability work has found heads specialized for previous-token lookups, induction (copy what followed the last occurrence of the current token), syntax and more. Elhage+ 2021
5.1 Complexity
| Component (per layer, per sequence) | FLOPs (approx.) | Memory |
|---|---|---|
| Q, K, V, O projections | \(2 \cdot 4 n d^2 = 8nd^2\) | Weights \(4d^2\) |
| \(QK^\top\) and \(AV\) | \(2 \cdot 2 n^2 d = 4 n^2 d\) | \(h n^2\) score matrix (naive); \(O(n)\) with FlashAttention |
| FFN (4× expansion) | \(16 n d^2\) | Weights \(8d^2\) |
The \(n^2\) attention term overtakes the \(nd^2\) projection and FFN terms when \(4n^2 d \gtrsim 24 n d^2\), that is when \(n \gtrsim 6d\). For \(d\) = 4,096 that is roughly 25K tokens. Below that, the matmuls dominate and "attention is quadratic" matters less in practice than people assume. Above it, as at 128K–1M context, attention dominates prefill. Decode is different: per new token, attention costs \(O(nd)\) compute but must read the whole cache, which is a memory-bandwidth problem.
Claiming "transformers are \(O(n^2)\) so they are slow" without nuance. For typical chat lengths, FFN and projection matmuls dominate FLOPs, and at decode time the real cost is bandwidth for weights and KV cache. Quote the crossover \(n \approx 6d\).
6. MQA, GQA, MLA and the KV cache
At decode time, every layer must keep K and V for all previous tokens. The memory for one sequence is:
$$ \text{KV bytes} = 2 \times L \times n_{kv} \times d_h \times n_{\text{ctx}} \times b $$Here 2 counts K and V, \(L\) is the number of layers, \(n_{kv}\) the number of K/V heads, \(d_h\) the head dimension, \(n_{\text{ctx}}\) the number of tokens, and \(b\) the bytes per element (2 for BF16, 1 for FP8). Every decode step reads the whole cache. So a smaller cache means larger batches, longer contexts, and more tokens/s per GPU.
6.1 The three sharing schemes
- MHA (multi-head): every query head has its own K and V head, so \(n_{kv} = h\).
- MQA (multi-query): all query heads share a single K/V head, so \(n_{kv}=1\). The cache shrinks by a factor of \(h\). Decoding got much faster, with some quality loss and training instability. Shazeer 2019
- GQA (grouped-query): query heads are split into \(g\) groups, and each group shares one K/V head, so \(n_{kv} = g\). Quality is close to MHA with a speed close to MQA. You can even "uptrain" an MHA checkpoint into GQA by mean-pooling K/V heads and training for about 5% of the original compute. Ainslie+ 2023 GQA with 8 KV heads is the default in Llama 2-70B and Llama 3, Mistral, Qwen, Gemma and gpt-oss.
6.2 Worked memory example
Take Llama-3-70B: \(L = 80\), \(h = 64\) query heads, \(n_{kv} = 8\), \(d_h = 128\), BF16. Llama 3 2024
$$ \text{per token} = 2 \times 80 \times 8 \times 128 \times 2\ \text{B} = 327{,}680\ \text{B} \approx 320\ \text{KiB} $$| Scheme (70B shape) | \(n_{kv}\) | Per token | 8K context | 128K context |
|---|---|---|---|---|
| MHA (hypothetical) | 64 | 2.5 MiB | 20 GiB | 320 GiB |
| GQA (actual) | 8 | 320 KiB | 2.5 GiB | 40 GiB |
| MQA (hypothetical) | 1 | 40 KiB | 0.31 GiB | 5 GiB |
| GQA + FP8 KV cache | 8 | 160 KiB | 1.25 GiB | 20 GiB |
Model weights in BF16 are about 140 GB. On 2×80 GB GPUs only about 20 GB is left after weights (less after activations and fragmentation). That is enough for one 64K-token sequence or eight 8K-token sequences. This is why the KV cache, rather than weights, often caps concurrency, and why paged attention, prefix caching, KV quantization and cache offload matter so much in production.
For a smaller model, Llama-3-8B (\(L=32, n_{kv}=8, d_h=128\)) needs \(2\cdot32\cdot8\cdot128\cdot2 = 128\) KiB per token, so 1 GiB per 8K-token sequence and 16 GiB at 128K. That is more than the 16 GB of BF16 weights.
6.3 MLA: Multi-head Latent Attention (DeepSeek)
MLA takes a different route. Instead of sharing heads, it compresses K and V into a low-rank latent. DeepSeek-V2 2024
- Down-project each token's hidden state to a small latent: \(c^{KV}_t = W^{DKV} h_t\), where \(d_c = 512\) is much smaller than \(h \cdot d_h\) (\(128 \times 128 = 16{,}384\) in DeepSeek-V2/V3).
- Per-head keys and values are up-projections of that latent: \(k_t^{(i)} = W^{UK}_{(i)} c^{KV}_t\), \(v_t^{(i)} = W^{UV}_{(i)} c^{KV}_t\). Only \(c^{KV}_t\) is cached.
- Absorption trick: \(q^\top k = q^\top W^{UK} c = (W^{UK\top} q)^\top c\). At inference you fold \(W^{UK}\) into the query projection and \(W^{UV}\) into the output projection, then attend directly against the latents without ever rebuilding full K/V.
- Decoupled RoPE: RoPE's position-dependent rotation would sit between \(W^{UK}\) and \(q\) and break the absorption. MLA therefore adds a separate small key of dimension \(d_R = 64\), shared across heads, that carries RoPE, and caches it alongside the latent.
The cache per token per layer is \(d_c + d_R = 576\) elements, against \(2 \cdot 128 \cdot 128 = 32{,}768\) for full MHA, roughly 57× smaller. DeepSeek reports a 93.3% KV-cache reduction versus their dense 67B model, with quality equal to or better than MHA. DeepSeek-V2 2024 For DeepSeek-V3 (61 layers): \(576 \times 61 \times 2\ \text{B} \approx\) 69 KiB per token, about 8.6 GiB at 128K. A 70B GQA model needs about 40 GiB, and DeepSeek-V3 is a model with 671B total parameters. DeepSeek-V3 2024
| Variant | Cached per token per layer | Quality vs MHA | Used by |
|---|---|---|---|
| MHA | \(2 h d_h\) | Baseline | GPT-2/3, Llama 1, Llama 2-7B/13B |
| MQA | \(2 d_h\) | Some degradation | PaLM, Falcon (early) |
| GQA | \(2 g d_h\) | Close to MHA | Llama 2-70B/3/4, Mistral, Qwen, Gemma, gpt-oss |
| MLA | \(d_c + d_R\) (e.g. 576) | Reported ≥ MHA | DeepSeek-V2/V3/R1/V3.2, Kimi K2, GLM-5 (2026) |
GQA says "heads don't need private copies of K/V," which is head-sharing. MLA says "the K/V across all heads lives in a low-dimensional subspace, so cache the coordinates and rebuild on demand," which is low-rank compression. MLA trades some extra compute at decode for much less memory traffic. Since decode is bandwidth-bound, that is a good trade.
Expect a back-of-envelope question like "How much memory does the KV cache take for model X at context Y, and how many concurrent users fit on this GPU?" Write the formula, plug in numbers, and convert units carefully (KiB vs KB). Follow-ups: "How would you halve it?" Possible answers: FP8 KV, fewer KV heads (GQA, which needs retraining or uptraining), MLA, sliding-window layers, prefix sharing, or evicting tokens (sinks plus recent tokens). Each changes quality in a different way.
- GQA paper: short, and includes the uptraining recipe.
- DeepSeek-V2: Section 2.1 is the clearest MLA description, including decoupled RoPE.
- Transformer Inference Arithmetic (kipply): KV-cache and FLOP arithmetic for serving.
7. Positional encodings
Self-attention is permutation-equivariant: shuffle the input tokens and the outputs shuffle the same way. Without position information, "dog bites man" and "man bites dog" look identical. (A causal mask leaks some order implicitly, which is why NoPE, no positional encoding at all, works surprisingly well in decoders.)
7.1 Sinusoidal and learned absolute
The original Transformer added a fixed vector to each input embedding: Vaswani+ 2017
$$ PE_{(pos,2i)} = \sin\!\big(pos / 10000^{2i/d}\big),\quad PE_{(pos,2i+1)} = \cos\!\big(pos / 10000^{2i/d}\big) $$Each pair of dimensions is a sinusoid at a different frequency, from fast (local position) to very slow (coarse position), like the hands of a clock. GPT-2/3 and BERT instead learned a \(n_{\max} \times d\) table. Both are absolute and generalize poorly past the trained length, because there is simply no row for position 2,049.
7.2 RoPE: rotary position embedding
RoPE is the default in essentially every modern open LLM. Su+ 2021 It does not add anything to embeddings. It rotates q and k inside each attention layer. Split the \(d_h\) dimensions into \(d_h/2\) pairs. For pair \(i\), define the frequency \(\theta_i = b^{-2i/d_h}\), with base \(b\) = 10,000 originally. At position \(m\), rotate the pair by angle \(m\theta_i\):
$$ \begin{pmatrix} q'_{2i} \\ q'_{2i+1}\end{pmatrix} = \begin{pmatrix}\cos m\theta_i & -\sin m\theta_i \\ \sin m\theta_i & \cos m\theta_i\end{pmatrix}\begin{pmatrix} q_{2i} \\ q_{2i+1}\end{pmatrix} $$The key property: rotating q by \(m\theta\) and k by \(n\theta\) gives a dot product that depends only on the angle between them, \((m-n)\theta\). In complex notation, \(\langle q e^{im\theta}, k e^{in\theta}\rangle = \mathrm{Re}[q \bar{k} e^{i(m-n)\theta}]\). So attention scores become a function of relative position, while the implementation stays a cheap element-wise operation applied to absolute positions. Nothing is added to the residual stream, and values are not rotated.
- High-frequency pairs (small \(i\)) rotate quickly. They distinguish neighbouring tokens and carry fine local order.
- Low-frequency pairs (large \(i\)) rotate slowly. The wavelength \(2\pi/\theta_i\) can exceed the training context, so those dimensions act almost as position-free semantic channels and coarse long-range position.
- Base \(b\): raising it stretches all wavelengths. Llama 3 uses 500,000 and Gemma 3's global layers use 1,000,000, which helps long-context training. Llama 3 2024 Gemma 3 2025 Increasing the base and then continuing pretraining on long documents ("ABF") is a standard context-extension recipe. Xiong+ 2023
Picture every q/k vector as a set of clock hands spinning at different speeds as you move along the sequence. Two tokens' hands differ by an angle that reflects only how far apart they are, not where they are. Fast hands tell you "adjacent or not"; slow hands tell you "same paragraph or far away."
7.3 ALiBi
ALiBi uses no embeddings at all. It adds a fixed, head-specific linear penalty to attention logits, \(S_{ij} \mathrel{+}= -m_h \cdot (i - j)\), with slopes \(m_h\) in a geometric sequence across heads. Distant tokens are penalized linearly, and some heads look far while others look locally. It extrapolated well past training length in the paper's experiments. Press+ 2021 BLOOM and MPT used it, but the field converged on RoPE plus scaling methods, partly because ALiBi's built-in recency bias can hurt retrieval of distant facts.
7.4 Context extension: PI, NTK-aware, YaRN
Suppose a RoPE model trained at length \(L\) is fed positions \(m > L\). Its low-frequency dimensions then see rotation angles never encountered in training, and quality collapses. The fixes all remap angles back into the trained range, given scale factor \(s = L'/L\):
| Method | What it does | Trade-off |
|---|---|---|
| Position Interpolation (PI) Chen+ 2023 | Divide positions by \(s\): \(m \to m/s\). All frequencies are squeezed uniformly. A short fine-tune (about 1,000 steps in the paper) recovers quality. | Also compresses the high-frequency dimensions, which blurs local resolution because neighbours look closer than they are. |
| NTK-aware scaling (community, 2023; formalized in YaRN) | Increase the base \(b\) instead (\(b' = b \cdot s^{d_h/(d_h-2)}\)). High frequencies barely change; low frequencies get interpolated. | Works reasonably even without fine-tuning; some dimensions slightly over-extrapolate. |
| YaRN Peng+ 2023 | "NTK-by-parts": leave dimensions with wavelength much shorter than \(L\) untouched, fully interpolate those longer than \(L\), and ramp between. Also adds an attention temperature (\(\sqrt{1/t} = 0.1\ln s + 1\)) to counter the entropy growth of longer contexts. | The paper reports needing far less fine-tuning data than PI. Widely used: Qwen, DeepSeek, gpt-oss (to 131K). OpenAI 2025 |
| Llama 3.1 "llama3" scaling | A by-parts frequency rescale similar in spirit to YaRN, combined with long-context continued pretraining up to 128K. | Model-specific constants. |
In practice, long context is mostly earned through continued pretraining on long data (with a scaled base or YaRN), not through the scaling trick alone. A model can "accept" 128K tokens and still use the middle poorly. Retrieval accuracy tends to dip for information placed mid-context. Liu+ 2023
Equating "context window" with "usable context." Advertised length is where the model doesn't break; effective length on multi-hop or aggregation tasks is often much shorter. Mention needle-in-a-haystack being too easy, and harder evaluations such as RULER-style multi-needle and aggregation tasks.
"Explain RoPE" separates candidates. A strong answer covers rotation of q/k pairs, relative position falling out of the dot product, the frequency spectrum, and why extension needs interpolation of the low frequencies but not the high ones (the YaRN insight). Also know that Llama 4 interleaves some layers with no positional encoding ("iRoPE") to help length generalization. Meta 2025
- Rotary Embeddings: A Relative Revolution (EleutherAI): intuition and code.
- YaRN: also the clearest write-up of PI vs NTK-aware vs NTK-by-parts.
- Effective Long-Context Scaling: the base-frequency plus continued-pretraining recipe.
8. The decoder block
Each block does two things:
$$ x \leftarrow x + \mathrm{Attn}(\mathrm{Norm}(x)), \qquad x \leftarrow x + \mathrm{FFN}(\mathrm{Norm}(x)) $$Attention is the only place where information moves between positions. The FFN works on each token independently and holds most of the parameters (about two-thirds in a dense model).
8.1 Pre-LN vs Post-LN
Post-LN (original 2017)
\(x \leftarrow \mathrm{LN}(x + \mathrm{Sublayer}(x))\)
- Normalization sits on the residual path, so gradients to early layers pass through every LN.
- Large gradients near the output at initialization need careful learning-rate warmup; deep stacks diverge easily.
- Can reach slightly better final quality when it does train.
Pre-LN (GPT-2 onward, now standard)
\(x \leftarrow x + \mathrm{Sublayer}(\mathrm{LN}(x))\)
- There is a clean identity path from output to input, so gradients are well behaved at initialization and warmup can be shortened. Xiong+ 2020
- The residual stream norm grows with depth, so later layers contribute relatively less. Needs a final norm before the head.
Variants exist: Gemma 2/3 apply RMSNorm both before and after each sublayer ("sandwich"), and some recent models (OLMo 2 reportedly among them) move the norm to the sublayer output inside the residual branch. Both target training stability at scale. QK-norm (Section 4.2) is part of the same stability toolkit.
8.2 RMSNorm
LayerNorm subtracts the mean, divides by the standard deviation, then applies a learned scale and bias. RMSNorm drops the mean-centering and the bias: Zhang & Sennrich 2019
$$ \mathrm{RMSNorm}(x) = \frac{x}{\sqrt{\tfrac{1}{d}\sum_i x_i^2 + \epsilon}} \odot g $$It is cheaper (one reduction instead of two, no bias) and empirically just as good: the re-scaling invariance seems to be what matters, not the re-centering. The paper reports speedups that vary with model and hardware. (roughly 7–64% in their tests) Nearly every modern LLM uses it.
8.3 The FFN and SwiGLU
The classic FFN is \(W_2\,\sigma(W_1 x)\) with \(W_1 \in \mathbb{R}^{d\times 4d}\) and ReLU or GELU. Modern LLMs use a gated linear unit: Shazeer 2020
$$ \mathrm{SwiGLU}(x) = W_2\big(\mathrm{SiLU}(W_1 x) \odot W_3 x\big), \qquad \mathrm{SiLU}(z) = z\,\sigma(z) $$One projection produces the "content," the other a data-dependent "gate," and they are multiplied element-wise. Because there are three matrices instead of two, the hidden width is cut to about \(\tfrac{8}{3}d\) to keep parameters equal to a \(4d\) FFN. In practice models round it: Llama-3-8B uses \(d_{ff}=14{,}336 = 3.5d\). Gated FFNs consistently gave lower perplexity in Shazeer's comparison; the paper itself (checked against the PDF) says: "We offer no explanation as to why these architectures seem to work; we attribute their success, as all else, to divine benevolence." GeGLU (GELU gate) is the same idea and is used by Gemma.
Interpretability work frames FFN layers as key-value memories: rows of \(W_1\) detect input patterns and the matching columns of \(W_2\) write associated content into the residual stream. Geva+ 2021 This is one reason factual knowledge is often localized to mid-layer FFNs.
8.4 The residual stream view
Rather than "layers transform representations," think of a \(d\)-dimensional shared bus. Embeddings write the initial state. Every attention head and every FFN reads from it (through its input projection) and adds a delta back. The unembedding reads the final state. Because everything is additive, the output decomposes into a sum of contributions from each component, which is the basis of circuit-style interpretability and the "logit lens" (unembedding intermediate layers to see predictions form). Elhage+ 2021
- It explains why layers can be skipped or pruned with modest damage, and why early exit and speculative self-drafting are possible.
- It explains why LoRA adapters and steering vectors work: small additive writes to a shared space.
- Heads in different layers compose by reading what earlier heads wrote. Induction heads, for example, combine a previous-token head with a matching head.
"Draw a transformer block" is common. Draw the pre-norm version, label it RMSNorm and SwiGLU, and say why each changed from 2017: pre-norm for stability, RMSNorm for speed, SwiGLU for quality per parameter, RoPE for relative position, and no biases in linear layers (most modern models drop them). Saying "the residual stream is the communication channel; attention moves information between tokens, the FFN processes within a token" signals real understanding.
- A Mathematical Framework for Transformer Circuits: the residual-stream view and induction heads.
- On Layer Normalization in the Transformer Architecture: Pre-LN vs Post-LN gradient analysis.
- nanoGPT: a complete GPT block in about 300 lines.
9. Counting parameters and FLOPs
9.1 The 12·L·d² rule
For a classic block with MHA and a 4× FFN:
- Attention: \(W_Q, W_K, W_V, W_O\), each \(d \times d\), giving \(4d^2\).
- FFN: \(d \times 4d\) plus \(4d \times d\), giving \(8d^2\).
- Norms and biases: \(O(d)\), negligible.
Sanity check, GPT-3 175B: \(L=96\), \(d=12{,}288\). Then \(12 \times 96 \times 12288^2 \approx 1.74\times10^{11}\), which is 174B. ✓
Modern variant, Llama-3-8B (\(L=32\), \(d=4096\), GQA with 8 KV heads of 128, SwiGLU with \(d_{ff}=14336\), \(V=128{,}256\), untied):
- Attention: \(W_Q, W_O\) are \(d^2\) each. \(W_K, W_V\) are \(d \times 1024 = 0.25d^2\) each. Total \(2.5d^2\).
- SwiGLU: \(3 \times d \times 3.5d = 10.5d^2\).
- Per layer \(13d^2 \approx 218\text{M}\). Times 32 layers gives about 6.98B.
- Embeddings plus head: \(2 \times 128256 \times 4096 \approx 1.05\)B. Total about 8.0B. ✓
GQA shrinks attention parameters and the wider SwiGLU makes up for it, so "\(12Ld^2\)" stays a good first estimate. For small models the embedding table is a large share. A 1B model with a 128K vocabulary and \(d\)=2,048 spends about 260M parameters on embeddings.
9.2 FLOPs per token: 2N and 6N
A matrix-vector product with a \(d_{in}\times d_{out}\) matrix costs \(d_{in} d_{out}\) multiply-adds, which is 2 FLOPs per parameter. So the forward pass costs about \(2N\) FLOPs per token. Backprop computes gradients with respect to both activations and weights, roughly twice the forward cost, so training costs about \(6N\) FLOPs per token. Kaplan+ 2020 The attention-score term adds roughly \(2 L n_{\text{ctx}} d\) per token forward, which is small unless \(n_{\text{ctx}}\) is comparable to \(6d\) or more, as computed earlier.
$$ C_{\text{train}} \approx 6\,N\,D \quad (D = \text{training tokens}) $$| Back-of-envelope | Computation | Result |
|---|---|---|
| Train an 8B model on 15T tokens | \(6 \times 8\!\times\!10^9 \times 1.5\!\times\!10^{13}\) | \(\approx 7.2\times10^{23}\) FLOPs |
| …on H100s at about 400 TFLOP/s effective (BF16 dense, about 40% MFU) | \(7.2\times10^{23} / 4\times10^{14}\) | \(\approx 1.8\times10^9\) GPU-s ≈ 500K GPU-hours |
| Chinchilla-optimal data for a 70B model | \(\approx 20\) tokens per parameter Hoffmann+ 2022 | ≈ 1.4T tokens (modern models train far beyond this; see the pretraining page) |
| Decode one token on a 70B dense model | \(2 \times 7\!\times\!10^{10}\) | 140 GFLOPs |
| Decode speed bound at batch 1, 70B BF16, one 3.35 TB/s GPU (if it fit) | 140 GB of weights read per token | ≈ 24 tokens/s: bandwidth-bound, with compute almost idle |
That last row is the core insight of LLM serving. At batch 1 you use about 140 GFLOPs out of hundreds of TFLOPs available per second. Batching many users shares each weight read across all of them, which is why throughput rises almost linearly with batch size until either compute or KV-cache memory runs out.
Expect "estimate the parameter count of a model with L layers and width d," "how long to train X on Y GPUs," or "why is decode slow?" Use \(12Ld^2\), \(6ND\), and "decode reads all weights per step." Show units and sanity-check against a known model (GPT-3 at 175B is the classic). Mention MFU: real training reaches roughly 35–55% of peak. (typical published range)
10. Encoder-only vs decoder-only vs encoder-decoder
| Encoder-only | Decoder-only | Encoder-decoder | |
|---|---|---|---|
| Attention mask | Bidirectional (every token sees all) | Causal | Encoder bidirectional; decoder causal plus cross-attention to encoder outputs |
| Pretraining objective | Masked LM (predict 15% masked tokens) Devlin+ 2018 | Next-token prediction | Span corruption / seq2seq Raffel+ 2019 |
| Examples | BERT, RoBERTa, DeBERTa, ModernBERT; most embedding and reranker models | GPT family, Llama, Qwen, DeepSeek, Gemma, Mistral, Claude-class models | T5, BART, original Transformer, Whisper (speech) |
| Strengths | Rich bidirectional representations for classification, NER, retrieval embeddings; cheap | One objective covers everything; generation; in-context learning; simple KV-cached inference; scales well | Natural fit for input→output tasks (translation, summarization, ASR); encoder runs once |
| Weaknesses | Cannot generate fluently | Prompt is processed causally (no bidirectional view of the input) | Two stacks, more complex serving; fell out of favour for general LLMs |
Why decoder-only won for general models: the next-token objective uses every token as a training signal (MLM uses only about 15%), every task can be phrased as text continuation, there is one homogeneous stack to scale and serve, and in-context learning emerges naturally. Encoders remain the right tool for embeddings, rerankers (cross-encoders), and cheap classifiers, which is relevant for RAG design.
"Would you use BERT or GPT for X?" For classification or retrieval at scale, use an encoder (fast, bidirectional, fine-tunable). For anything generative or instruction-following, use a decoder. A sharp candidate notes that the strongest embedding models now often come from decoder LLMs fine-tuned contrastively, so the line has blurred.
11. Mixture of Experts (MoE)
MoE replaces the single FFN in (some or all) blocks with \(E\) parallel expert FFNs plus a small router. Each token runs through only \(k\) experts, so parameter count (capacity, knowledge) is decoupled from FLOPs per token. Shazeer+ 2017 Lepikhin+ 2020
11.1 Routing
For token representation \(x\), router weights \(W_r \in \mathbb{R}^{d\times E}\):
$$ s = \mathrm{softmax}(x W_r)\ \ (\text{or sigmoid}),\quad \mathcal{T} = \mathrm{TopK}(s, k),\quad y = \sum_{i\in\mathcal{T}} \tilde{g}_i\, E_i(x) $$where \(\tilde{g}\) are the selected scores, usually renormalized to sum to 1. Switch Transformer used \(k=1\) Fedus+ 2021, Mixtral uses 8 experts with \(k=2\) Jiang+ 2024, and DeepSeek-V3 uses 256 routed experts with \(k=8\) plus 1 shared expert, with sigmoid scores. DeepSeek-V3 2024 The router is trained end-to-end through the gate values \(\tilde{g}\). The discrete top-k choice is not differentiable, but the weighting is.
11.2 Total vs active parameters
| Model | Total params | Active / token | Experts (routed / top-k / shared) |
|---|---|---|---|
| Mixtral 8×7B | 46.7B | 12.9B | 8 / 2 / 0 Jiang+ 2024 |
| DeepSeek-V3 | 671B | 37B | 256 / 8 / 1 DeepSeek 2024 |
| Qwen3-235B-A22B | 235B | 22B | 128 / 8 / 0 Qwen3 2025 |
| Llama 4 Maverick | 400B | 17B | 128 / 1 / 1, alternating dense and MoE layers Meta 2025 |
| gpt-oss-120b | 116.8B | 5.1B | 128 / 4 / 0 OpenAI 2025 |
| Kimi K2 | 1T | 32B | 384 / 8 / 1 Kimi 2025 |
- FLOPs and latency scale with active parameters: DeepSeek-V3 costs about \(2\times37\text{B}\) FLOPs per token, like a 37B dense model.
- Memory scales with total parameters: all 671B must sit somewhere (about 671 GB in FP8), so MoE needs many GPUs even though each token touches only a few experts.
- Quality sits between the two. An MoE generally beats a dense model of its active size, and falls short of a dense model of its total size at equal training tokens.
- Serving is efficient at high batch, because enough tokens arrive per expert to make matmuls efficient. At batch 1 you still read \(k\) experts' weights per token per layer, and across a sequence you read nearly all of them.
11.3 Load balancing: the central problem
Left alone, routers collapse. A few experts get slightly better early, receive more tokens, improve further, and the rest starve. With expert parallelism (experts spread over GPUs) an overloaded expert also becomes a straggler that stalls everyone.
- Auxiliary load-balancing loss (Switch-style): \(\mathcal{L}_{aux} = \alpha \cdot E \sum_{i=1}^{E} f_i P_i\), where \(f_i\) is the fraction of tokens dispatched to expert \(i\) and \(P_i\) the mean router probability for \(i\). The minimum is at a uniform distribution. \(\alpha\) is small (about 0.01). Fedus+ 2021 The downside is that this loss's gradients fight the language-modelling objective.
- Capacity factor: each expert processes at most \(C \cdot \tfrac{T k}{E}\) tokens per batch, for \(T\) tokens. Overflow tokens are dropped (they skip the FFN via the residual). \(C\) of 1.0–1.25 trades wasted padding against dropped tokens. Many recent systems avoid dropping entirely.
- Router z-loss penalizes large router logits for numerical stability. Zoph+ 2022
- Aux-loss-free balancing (DeepSeek): add a per-expert bias \(b_i\) to the scores only for the top-k selection, not for the gate weights. After each step, decrease \(b_i\) if expert \(i\) was overloaded and increase it if underloaded. Balance is enforced by a control loop instead of a gradient, so there is no interference with the main loss. Wang+ 2024 DeepSeek-V3 uses this plus a tiny sequence-level auxiliary loss, and drops no tokens. DeepSeek-V3 2024
11.4 Fine-grained and shared experts
DeepSeekMoE made two influential changes. Dai+ 2024 (1) Fine-grained experts: split each expert into smaller ones and activate proportionally more (for example 256 small experts with top-8 instead of 16 large ones with top-1). That gives combinatorially more expert combinations and sharper specialization at the same FLOPs. (2) Shared experts: one or more experts that every token uses. They absorb common knowledge (syntax, frequent patterns) so routed experts don't each have to duplicate it. Practice varies: DeepSeek-V3, Kimi K2 and Llama 4 use a shared expert, while Qwen3 and gpt-oss do not.
11.5 Systems realities
- Expert parallelism places different experts on different GPUs. Each MoE layer then needs two all-to-all exchanges (dispatch tokens to expert hosts, gather results back). Interconnect bandwidth becomes the bottleneck, which is why DeepSeek-V3 limits each token to experts on at most 4 nodes and overlaps communication with compute. DeepSeek-V3 2024
- Fine-tuning MoEs is touchier: routers can drift and experts can collapse on narrow data. Many teams freeze the router or keep balancing active during fine-tuning.
- Experts don't specialize along human-legible topics. Specialization is often syntactic or token-level.
"What is the difference between total and active params, and why would you choose MoE?" Strong answer: MoE buys knowledge capacity at constant FLOPs per token, but costs memory (all experts resident), communication (all-to-all), and engineering complexity (balancing, capacity, fine-tuning instability). It pays off when you serve at high batch on many GPUs, as cloud APIs do. For on-device or single-GPU low-batch use, a dense model of the same memory footprint can be the better choice. Know the aux-loss formula and the bias trick.
By 2026 essentially every frontier-scale open model above roughly 100B parameters is MoE, and active-parameter ratios keep falling (about 3B active of 80B in Qwen3-Next, 13B of 284B in DeepSeek-V4-Flash). Qwen 2025 HF 2026 Check current releases for expert counts and balancing methods.
- Mixture of Experts Explained (Hugging Face): history, capacity, balancing, fine-tuning.
- Switch Transformers: the classic aux loss and capacity factor.
- Auxiliary-Loss-Free Load Balancing: the bias-update method used by DeepSeek-V3.
12. Long-context techniques
Two costs grow with context: prefill compute (\(O(n^2)\)) and KV-cache memory and bandwidth (\(O(n)\) per layer). The techniques below attack one or both.
12.1 Sliding-window and local/global interleaving
Sliding-window attention (SWA) lets each token attend only to the last \(W\) tokens, so the cache is capped at \(W\) per layer. Information still propagates further through depth: after \(L\) layers the receptive field is about \(L \times W\). Mistral 7B used \(W\)=4,096 across 32 layers. Jiang+ 2023 Pure SWA can't do precise long-range retrieval, so modern models interleave local and global layers:
- Gemma 2: alternating local (4,096) and global layers, 1:1. Gemma 2 2024
- Gemma 3: 5 local : 1 global, local window 1,024. KV overhead at 32K falls from about 60% (global-only) to under 15%. Gemma 3 2025
- gpt-oss: alternating dense and banded-window layers, with a window of just 128 tokens. OpenAI 2025
12.2 Sparse attention
Fixed sparse patterns came first: strided and fixed patterns in Sparse Transformers Child+ 2019, and local plus a few global tokens in Longformer Beltagy+ 2020 and BigBird (which adds random links) Zaheer+ 2020. The modern wave uses learned, content-dependent sparsity designed for GPU kernels:
- NSA (DeepSeek, 2025): combines compressed coarse tokens, top-k selected blocks and a sliding window, and is trainable end-to-end. Yuan+ 2025
- DSA (DeepSeek-V3.2): a lightweight FP8 "lightning indexer" scores all previous tokens, then each query attends to only the top 2,048 key-value entries. Core attention drops from \(O(L^2)\) to \(O(Lk)\), although the indexer itself stays quadratic but cheap. It is built on MLA and was trained by continuing from V3.1-Terminus. DeepSeek 2025 GLM-5 adopted MLA plus DSA in 2026.
- DeepSeek-V4 (2026): interleaves Compressed Sparse Attention (KV compressed 4× along the sequence, then top-k selection) with Heavily Compressed Attention (128× compression, dense). At 1M tokens it reports about 10% of V3.2's KV memory. HF 2026
12.3 Linear attention
Replace \(\mathrm{softmax}(qk^\top)\) with a kernel \(\phi(q)\phi(k)^\top\). Matrix multiplication is associative, so \((\phi(Q)\phi(K)^\top)V = \phi(Q)(\phi(K)^\top V)\). The second form costs \(O(n d^2)\) instead of \(O(n^2 d)\). In causal form it becomes an RNN with a matrix-valued state: Katharopoulos+ 2020
$$ S_t = S_{t-1} + \phi(k_t)\, v_t^\top, \qquad o_t = S_t^\top \phi(q_t) $$The state \(S \in \mathbb{R}^{d_k \times d_v}\) has constant size, so there is no growing KV cache. The weakness: everything is summed into a fixed-capacity memory, so recall of specific earlier tokens degrades, and vanilla linear attention underperformed softmax. Modern variants add gates/decay (forget old content) and the delta rule (overwrite the value stored under key \(k\) rather than adding to it): Yang+ 2024
$$ S_t = \alpha_t\, S_{t-1}\big(I - \beta_t k_t k_t^\top\big) + \beta_t\, v_t k_t^\top \quad\text{(Gated DeltaNet, schematically)} $$12.4 Attention sinks
Trained LLMs put a surprising amount of attention mass on the first few tokens, whatever their content. The softmax must sum to 1, so heads with "nothing to attend to" dump their weight on a consistently visible position. If you evict those tokens in a sliding-window cache, perplexity explodes. StreamingLLM keeps the first ~4 "sink" tokens plus a rolling window, which allows stable generation over millions of tokens (though without recall beyond the window). Xiao+ 2023 gpt-oss builds this in as a learned per-head bias in the softmax denominator, which lets a head attend to "nothing." OpenAI 2025 Vision transformers show the same phenomenon on background patches, and adding dedicated "register" tokens fixes it. Darcet+ 2023
12.5 Systems-level tools
- FlashAttention: exact attention, tiled in SRAM, so memory is \(O(n)\). Not a sparsity method, but it made 32K–128K training practical. Dao+ 2022
- Ring / context parallelism: shard the sequence across devices and pass K/V blocks around a ring, overlapping communication with compute. Liu+ 2023
- KV-cache compression at inference: quantization (FP8/INT4), token eviction (H2O/SnapKV-style heuristics), and offload to CPU or SSD (covered on the inference page).
"How would you support 1M-token context?" A strong answer is layered: (1) architecture, meaning a mostly local/linear/compressed attention stack with a few global layers, MLA or GQA, and FP8 KV; (2) training, meaning RoPE base scaling or YaRN plus continued pretraining on long documents, and context parallelism; (3) serving, meaning paged KV, prefix caching and chunked prefill; (4) evaluation beyond needle-in-a-haystack. Also ask whether retrieval (RAG) solves the actual problem more cheaply.
- Efficient Streaming Language Models with Attention Sinks: short and surprising.
- DeepSeek-V3.2: DSA and the lightning indexer.
- The Transformer Family v2 (Lilian Weng): a survey of efficient attention variants.
13. Alternatives and hybrids: SSMs, RWKV, linear-attention hybrids
13.1 State-space models and Mamba
A (discretized) linear state-space model runs a recurrence \(h_t = \bar{A} h_{t-1} + \bar{B} x_t,\; y_t = C h_t\). Classic SSMs (S4) used fixed, input-independent \(A, B, C\), which allows training as a long convolution, but they could not decide what to remember. Mamba makes \(B, C\) and the step size \(\Delta\) functions of the input ("selective"), so the model can choose to ignore or store a token. That breaks the convolution trick, so Mamba uses a hardware-aware parallel scan. Gu & Dao 2023 Inference uses a constant-size state per layer and \(O(1)\) time per token. Mamba-2 restricts \(A\) to a scalar times the identity, which exposes a duality with (masked) linear attention and makes training matmul-friendly. Dao & Gu 2024
13.2 RWKV
RWKV is a linear-attention-style RNN with learned per-channel time decay. It trains in parallel like a transformer and runs as an RNN at inference. Peng+ 2023 Later versions (RWKV-7 "Goose") add delta-rule-like dynamic state updates, converging with the DeltaNet family. Peng+ 2025 It is community-driven and has strong efficiency results, but it is not used in frontier models.
13.3 What the evidence says
A controlled 8B comparison on the same data (up to 3.5T tokens) found pure Mamba/Mamba-2 match transformers on many tasks but lag on copying, in-context learning (5-shot MMLU, phonebook lookup) and long-context reasoning. A hybrid of 43% Mamba-2, 7% attention and 50% MLP layers beat the transformer on all 12 standard tasks evaluated. Waleffe+ 2024 This explains the field's direction: a fixed-size state is cheap but lossy, and a few full-attention layers restore exact retrieval.
13.4 Hybrid models in production
| Model | Recurrent / linear component | Mix with full attention | Notes |
|---|---|---|---|
| Jamba (AI21, 2024) | Mamba | Interleaved Transformer and Mamba blocks, plus MoE | 256K context; fits on one 80 GB GPU Lieber+ 2024 |
| MiniMax-01 (2025) | Lightning attention (linear) | Periodic softmax layers (about 1 in 8) | 456B total / 45.9B active MoE; up to 1M training context MiniMax 2025 |
| Nemotron-H (NVIDIA, 2025) | Mamba-2 | Most attention layers replaced; a small fraction kept | Up to 3× faster inference than similar-size transformers NVIDIA 2025 |
| Qwen3-Next-80B-A3B (2025) | Gated DeltaNet | 3 GDN : 1 gated attention (48 layers) | 80B total / 3B active; 262K native context Qwen 2025 |
| Kimi Linear (2025) | Kimi Delta Attention (refined Gated DeltaNet) | Layerwise hybrid with MLA (3:1) | 48B/3B; up to 75% less KV cache and up to 6× decode throughput at 1M Kimi 2025 |
| Qwen3.5 (Feb 2026) | Gated DeltaNet | Hybrid, carried over from Qwen3-Next | 397B / 17B active, multimodal Raschka 2026 |
Well established
- Pure SSM/linear models are competitive on perplexity but weaker on exact recall and in-context learning.
- Hybrids with a minority of full-attention layers close that gap and cut KV memory by large factors.
- Plain softmax-attention transformers with GQA/MLA remain the most common design.
Recent / contested (2025–2026)
- Whether hybrids beat full attention at frontier scale under equal compute (Kimi Linear claims yes in fair comparisons).
- The best ratio (3:1? 7:1?) and which linear mixer to use (Mamba-2, Gated DeltaNet, KDA, lightning).
- Whether compressed/sparse softmax attention (DeepSeek's path) or linear hybrids (Qwen/Kimi's path) wins long-context efficiency.
This is the fastest-moving part of the page. As of early-to-mid 2026, hybrid linear-attention MoE models (Qwen3.5, Ling 2.5) sat alongside MLA plus sparse attention (GLM-5, DeepSeek-V4) and "classic" GQA designs (MiniMax M2.5). Raschka 2026 One recurring observation is that data and training recipe often matter more than which of these mixers you pick. Check the latest releases before an interview.
"Will Mamba replace transformers?" A strong answer covers the constant-state advantage (memory, long-context decode speed), the recall weakness (fixed-size state compression), the empirical result that hybrids win, and the 2025–2026 production evidence (Qwen3-Next/3.5, Nemotron-H, Kimi Linear). It ends with the honest status: hybrids are mainstream in efficiency-focused open models, while pure attention plus GQA or MLA remains common, and the field hasn't settled.
- An Empirical Study of Mamba-based Language Models: the controlled 8B comparison.
- Gated Delta Networks: the linear mixer behind Qwen3-Next and Qwen3.5.
- The Big LLM Architecture Comparison (Raschka): diagrams of many 2025 open models.
14. Architectures of notable open models
The table shows the design space has converged. Nearly everything is a pre-norm decoder with RMSNorm, SwiGLU/GeGLU, RoPE and GQA or MLA. The differences are in attention sparsity, MoE configuration and vocabulary. Values come from the cited reports; cells marked with a dashed underline are from memory or secondary sources.
| Model (year) | Params (total / active) | Layers · d | Attention | Position · context | FFN / MoE | Vocab |
|---|---|---|---|---|---|---|
| Llama 3 8B (2024) Meta | 8B dense | 32 · 4096 | GQA 32Q/8KV | RoPE θ=500K · 8K→128K (3.1) | SwiGLU 14336 | 128K |
| Llama 3 70B / 405B | 70B / 405B dense | 80 · 8192 / 126 · 16384 | GQA 64Q/8KV, 128Q/8KV | RoPE θ=500K · 128K (3.1) | SwiGLU | 128K |
| Llama 4 Maverick (2025) Meta | 400B / 17B | n/a | GQA; iRoPE (some NoPE layers) | RoPE + NoPE interleave | MoE 128 routed + 1 shared, alternating dense/MoE | ~200K |
| Mistral 7B (2023) Mistral | 7.3B dense | 32 · 4096 | GQA 32Q/8KV, sliding window 4096 | RoPE · 8K (SWA) | SwiGLU 14336 | 32K |
| Mixtral 8×7B (2024) Mistral | 46.7B / 12.9B | 32 · 4096 | GQA 32Q/8KV | RoPE · 32K | MoE 8 experts, top-2 | 32K |
| DeepSeek-V3 / R1 (2024–25) DeepSeek | 671B / 37B | 61 · 7168 | MLA (128 heads, d_c=512, RoPE dim 64) | Decoupled RoPE + YaRN · 128K | DeepSeekMoE 256 routed (top-8) + 1 shared; first 3 layers dense; aux-loss-free; MTP | ~129K |
| Qwen2.5-72B (2024) Qwen | 72B dense | 80 · 8192 | GQA 64Q/8KV, QKV bias | RoPE · 128K (YaRN) | SwiGLU | ~152K |
| Qwen3-235B-A22B (2025) Qwen | 235B / 22B | 94 · 4096 | GQA 64Q/4KV, QK-norm, no QKV bias | RoPE · 32K native, 128K YaRN | MoE 128 experts top-8, no shared | ~152K |
| Qwen3-Next-80B-A3B (2025) Qwen | 80B / 3B | 48 · 2048 | Hybrid 3 Gated DeltaNet : 1 gated attention (16Q/2KV, d_h 256) | 262K native | High-sparsity MoE | ~152K |
| Gemma 2 27B (2024) Google | 27B dense | 46 · 4608 | GQA; local (4096) / global 1:1; logit soft-capping | RoPE · 8K | GeGLU; pre+post RMSNorm | 256K |
| Gemma 3 27B (2025) Google | 27B dense | n/a | GQA; local (1024) : global 5:1; QK-norm replaces soft-capping | RoPE 10K local / 1M global · 128K | GeGLU; pre+post norm | 262K |
| gpt-oss-120b (2025) OpenAI | 116.8B / 5.1B | 36 · 2880 | GQA 64Q/8KV, d_h 64; alternating dense / 128-token banded; learned sinks | RoPE + YaRN · 131K | MoE 128 experts top-4 (gated SwiGLU) | 201K (o200k_harmony) |
| Kimi K2 (2025) Moonshot | 1T / 32B | 61 · 7168 | MLA (64 heads) | RoPE · 128K | MoE 384 experts top-8 + 1 shared; MuonClip optimizer | ~160K |
| DeepSeek-V4-Pro (2026) HF | 1.6T / 49B | 61 layers | Interleaved CSA (4× compressed + sparse) / HCA (128× compressed, dense) | 1M | DeepSeekMoE; manifold-constrained hyper-connections | n/a |
Read the table as a history of the KV cache. MHA gave way to GQA (2023), then MLA and sliding-window interleaving (2024), then learned sparse and compressed attention plus linear-attention hybrids (2025–26). On the FFN side the story is dense, then coarse MoE (Mixtral), then fine-grained MoE with shared experts and ever lower active ratios.
New model families appear monthly. Rows for 2025–2026 models (Qwen3-Next, DeepSeek-V4, Kimi K2) come partly from model cards and secondary write-ups. Confirm against the latest technical reports before quoting exact numbers.
15. Multimodal input basics
15.1 Vision Transformer (ViT)
A ViT treats an image as a sequence of patches. Dosovitskiy+ 2020
- Split an \(H\times W\times 3\) image into \(P\times P\) patches. A 224×224 image with \(P\)=16 gives \(14\times14 = 196\) patches.
- Flatten each patch (\(16\cdot16\cdot3 = 768\) values) and apply a linear projection to width \(d\). This is equivalent to a stride-\(P\) convolution.
- Add position embeddings (learned 2-D, or 2-D RoPE in newer encoders) and optionally a [CLS] token.
- Run a standard bidirectional transformer encoder.
Token count grows with resolution squared. An 896×896 image at \(P\)=14 is 4,096 patches, which is why VLMs pool or compress visual tokens.
15.2 Vision encoder + projector + LLM (late fusion)
- Vision encoder: usually a CLIP or SigLIP ViT pretrained with image-text contrastive learning, so its features are already language-aligned. Radford+ 2021
- Projector: LLaVA used a single linear layer Liu+ 2023 and LLaVA-1.5 a 2-layer MLP Liu+ 2023b. Others use cross-attention resamplers that compress to a fixed number of tokens. Gemma 3 condenses SigLIP outputs to 256 tokens per image. Gemma 3 2025
- Training recipe: (1) alignment, where only the projector is trained on image-caption pairs with both towers frozen; (2) visual instruction tuning, where the projector and LLM (fully or with LoRA) are trained on multimodal conversations.
- Dynamic resolution: Qwen2-VL processes images at native resolution with variable token counts and uses M-RoPE (separate temporal, height and width rotary components). Wang+ 2024 Other models tile high-resolution images into crops plus a thumbnail.
15.3 Early fusion
Early-fusion models put image tokens (discrete VQ codes or continuous patch embeddings) into one backbone from the start of pretraining, rather than bolting a vision encoder onto a finished LLM. Examples are Chameleon Chameleon Team 2024, Llama 4 Meta 2025, and in 2026 Kimi K2.5 and Qwen3.5. Raschka 2026 The benefit is deeper cross-modal integration. The cost is training from scratch with mixed data, and the open question of how to balance modalities.
"How does a VLM see an image?" Cover patches to embeddings, a contrastively pretrained encoder, a projector into the LLM's embedding space, and the LLM attending over visual tokens as if they were words. Then discuss costs: image tokens consume context and prefill compute (hundreds to thousands per image), so resolution, pooling and caching image prefixes are real product levers. Mention OCR-heavy tasks needing high resolution as a key trade-off.
- An Image is Worth 16x16 Words: the ViT paper.
- Visual Instruction Tuning (LLaVA): the canonical encoder + projector + LLM recipe.
- Qwen2-VL: dynamic resolution and M-RoPE.
Interview question bank
1. Walk through scaled dot-product attention with tensor shapes for a batch.
Input \(X\): [B, n, d]. Project to Q, K, V with \(W_Q, W_K, W_V\) and reshape to [B, h, n, d_h], where \(d_h = d/h\). (With GQA, K and V are [B, n_kv, n, d_h] and broadcast across each group.) Scores \(QK^\top/\sqrt{d_h}\) are [B, h, n, n]. Add the causal mask (−∞ above the diagonal) and padding mask, softmax over the last axis, and multiply by V to get [B, h, n, d_h]. Transpose and reshape to [B, n, d], then apply \(W_O\). Each output token is a convex combination of value vectors from allowed positions. Mention that FlashAttention computes the same result without materializing the [n, n] matrix in HBM.
2. Why divide by √d_k? What happens if you don't?
If q and k have independent unit-variance components, \(q\cdot k\) has variance \(d_k\). Without scaling, logits for \(d_k\)=128 have a standard deviation of about 11, the softmax becomes nearly one-hot, and its Jacobian \(a_i(\delta_{ij}-a_j)\) approaches zero. Gradients vanish and training is slow or unstable at initialization. Scaling by \(1/\sqrt{d_k}\) restores unit variance. The model can still learn sharp attention by growing weight norms. QK-norm is a stronger modern version that bounds logits throughout training, not only at initialization.
3. Compute the KV-cache size for Llama-3-70B at 32K context, batch 8, BF16.
Per token: \(2 \times 80\ \text{layers} \times 8\ \text{KV heads} \times 128 \times 2\ \text{B} = 327{,}680\) B ≈ 320 KiB. Per sequence at 32K: 32,768 × 320 KiB = 10 GiB. Batch 8: about 80 GiB, on top of about 140 GB of weights. So you need roughly 3×80 GB GPUs before activations and overhead, or FP8 KV (40 GiB). With full MHA (64 KV heads) it would be 8× larger, about 640 GiB, which shows why GQA matters.
4. Compare MHA, MQA, GQA and MLA. When would you choose each?
MHA gives every query head its own K/V: the best expressivity and the largest cache. MQA shares one K/V head: the smallest cache (h× smaller) but some quality loss and instability. GQA shares one K/V per group of query heads, typically 8 KV heads: near-MHA quality at a fraction of the cache. It is the default, and an existing MHA checkpoint can be uptrained into it. MLA caches a low-rank latent (512 + 64 dims per token per layer in DeepSeek) and absorbs the up-projections into Q and O at inference, giving an even smaller cache with reported quality at or above MHA, at the cost of implementation complexity (decoupled RoPE, custom kernels). Choose GQA for simplicity and ecosystem support, and MLA when long-context serving cost dominates and you control the stack.
5. Explain RoPE and why context extension methods target low frequencies.
RoPE rotates each 2-D pair of q/k dimensions by angle \(m\theta_i\) at position \(m\), with \(\theta_i = b^{-2i/d}\). The q·k dot product then depends only on \((m-n)\theta_i\), which is relative position. High-frequency pairs complete many rotations within the training length, so the model has seen all their angles. Low-frequency pairs have wavelengths longer than the training context, so positions beyond \(L\) produce angles never seen before. That causes the failure when extrapolating. PI rescales every frequency (blurring local detail). NTK-aware scaling raises the base so mainly low frequencies are interpolated. YaRN interpolates only the dimensions whose wavelength exceeds \(L\), ramps between, and adds an attention temperature. All of these work best with some continued pretraining on long data.
6. Estimate the parameters of a decoder with L=48, d=6144, vocab 100K, untied embeddings.
Non-embedding: \(12 \times 48 \times 6144^2 = 12 \times 48 \times 37.7\text{M} \approx 21.7\)B. Embeddings plus head: \(2 \times 100\text{K} \times 6144 \approx 1.2\)B. Total about 23B. If it uses GQA and SwiGLU with \(d_{ff}=3.5d\), per layer is about \(13d^2\) instead of \(12d^2\), giving about 23.5B non-embedding. State that norms and biases are negligible.
7. How many FLOPs and GPU-hours to train a 70B dense model on 15T tokens?
\(C \approx 6ND = 6 \times 7\times10^{10} \times 1.5\times10^{13} = 6.3\times10^{24}\) FLOPs. At about 400 TFLOP/s effective per H100 (about 40% MFU of BF16 dense peak), that is \(6.3\times10^{24}/4\times10^{14} \approx 1.6\times10^{10}\) GPU-seconds, about 4.4M GPU-hours, or about 3 months on 2,000 GPUs. Name the assumptions (MFU, no restarts) and note that real runs add overhead for failures, evaluations and ablations.
8. Why is LLM decoding memory-bandwidth bound, and what does that imply?
Each decode step multiplies a single token's activations by every weight matrix (matrix-vector), so arithmetic intensity is about 1–2 FLOPs per byte loaded. A GPU needs hundreds of FLOPs per byte to be compute-bound. Time per token is therefore roughly (weight bytes + KV bytes) / HBM bandwidth. Implications: batching amortizes weight reads across users (throughput scales with batch until memory or compute runs out), weight and KV quantization speed up decode almost proportionally, smaller KV caches (GQA/MLA) allow bigger batches, and speculative decoding works because verifying several tokens costs about the same as generating one.
9. Pre-LN vs Post-LN: what's the difference and why did the field switch?
Post-LN normalizes after the residual add: \(\mathrm{LN}(x + F(x))\). Pre-LN normalizes the branch input: \(x + F(\mathrm{LN}(x))\). Pre-LN leaves an unnormalized identity path from loss to embeddings, so gradients are well scaled at initialization and training is stable without long warmup, even for very deep models. Post-LN gradients near the output are large at initialization and need careful warmup. Pre-LN's downsides are that the residual norm grows with depth (later layers matter relatively less) and a final norm is required. Variants such as sandwich norm (Gemma) and QK-norm address the remaining instabilities.
10. What is SwiGLU and why use 8/3·d hidden size?
SwiGLU is \(W_2(\mathrm{SiLU}(W_1x)\odot W_3x)\): a gated FFN where one projection modulates the other element-wise. It has three weight matrices instead of two, so to match the parameter count of a classic \(4d\) FFN (\(8d^2\)) you set \(3 \cdot d \cdot d_{ff} = 8d^2\), giving \(d_{ff} = 8d/3\). Empirically it gives better perplexity per parameter and per FLOP. Real models round \(d_{ff}\) for hardware efficiency or make it a bit larger (Llama-3-8B uses 3.5d).
11. Explain MoE routing and load balancing. What is the aux-loss-free method?
A router computes per-expert scores for each token. The top-k experts run, and their outputs are combined using the normalized scores. Without intervention routing collapses onto a few experts. The classic fix is the auxiliary loss \(\alpha E\sum_i f_i P_i\) (fraction of tokens times mean probability), minimized at uniform load, plus a capacity factor that drops overflow tokens. The aux loss adds gradients that conflict with language modelling. DeepSeek's aux-loss-free method adds a per-expert bias to the scores used only for top-k selection, and nudges it down or up after each step depending on whether the expert was over- or under-loaded. Balance is maintained by a controller rather than a gradient. DeepSeek-V3 combines this with a tiny sequence-level loss and drops no tokens.
12. A 670B-total / 37B-active MoE vs a 70B dense model: compare serving.
FLOPs per token: MoE about 74 GFLOPs vs dense about 140 GFLOPs, so the MoE is cheaper in compute. Memory: the MoE needs about 670 GB in FP8 (all experts resident), which means a multi-GPU node or several nodes, versus about 70 GB in FP8 for the dense model. At low batch the MoE still reads many experts, so its latency advantage shrinks. At high batch, with expert parallelism and good all-to-all bandwidth, it gives much better quality per FLOP. It also has an MLA KV cache about 5× smaller than the 70B model's GQA cache. Choose the MoE for high-throughput cloud serving and the dense model for constrained or single-node deployment.
13. Why do LLMs struggle to count letters or do multi-digit arithmetic? How would you mitigate it?
The model sees token IDs, not characters. "strawberry" may be two or three tokens, so counting 'r' requires memorized spelling of each token. Inconsistent number chunking breaks place-value alignment. Mitigations: digit-level or fixed 3-digit number tokenization at pretraining time, prompting the model to spell out characters first, tool use (a code interpreter or calculator), and reasoning traces. At the architecture level, byte-level or patch-based models remove the issue at higher compute cost.
14. What is the computational complexity of a transformer layer and when does the attention term dominate?
Per layer per sequence: projections and FFN cost about \(24nd^2\) FLOPs (\(8nd^2\) attention projections plus \(16nd^2\) FFN), and attention scores plus mixing cost about \(4n^2d\). The quadratic term dominates when \(n > 6d\), roughly 25K tokens for \(d\)=4,096. Memory for naive attention is \(O(hn^2)\), reduced to \(O(n)\) by FlashAttention. At decode, per token attention compute is \(O(nd)\) per layer, but the binding cost is reading the \(O(n)\) KV cache.
15. What are attention sinks and why do they matter for streaming and KV eviction?
Trained models put large attention weight on the first few tokens regardless of content, because softmax weights must sum to 1 and heads need a default place to put mass when nothing is relevant. If you evict those tokens with a pure sliding window, the distribution shifts and perplexity explodes. StreamingLLM keeps about 4 initial tokens plus a recent window for stable infinite-length generation (no long-range recall). gpt-oss builds in a learned sink logit per head. The same phenomenon appears in ViTs, where register tokens fix it. Any KV-eviction policy should pin the sink tokens.
16. Design: you must serve a 128K-context assistant on a fixed GPU budget. Which architecture choices matter?
The KV cache is the main cost, so prefer models with GQA (few KV heads) or MLA, and ideally interleaved local/global or hybrid linear layers (Gemma 3 style 5:1, Qwen3-Next style 3:1), which cut cache by large factors. Use an FP8 KV cache, paged attention, prefix caching for shared system prompts and documents, and chunked prefill to protect inter-token latency. Measure effective context quality on your own tasks, not just needle tests. Consider RAG to keep typical prompts well below 128K. Do the arithmetic per candidate: cache per token × typical context × target concurrency against free HBM.
17. Why did decoder-only models win over encoder-decoder for general LLMs?
Next-token prediction gives a training signal at every position (MLM uses about 15%), any task can be framed as continuation, which enables in-context learning and instruction following, there is one homogeneous stack that is simpler to scale, parallelize and serve with KV caching, and empirical scaling was excellent. Encoder-decoders remain good for fixed input→output tasks (translation, ASR such as Whisper). Encoders remain best for embeddings, rerankers and classifiers. Prefix-LM variants (bidirectional over the prompt) exist but didn't show a large enough win at scale.
18. Will SSMs / linear attention replace transformers? Give the current evidence.
Pure SSMs (Mamba) and linear attention have constant-size state, giving \(O(1)\) memory and time per decode step, but they compress history into fixed memory and so lag on copying, exact recall and in-context learning. A controlled 8B study found exactly this, and a hybrid with about 7% attention layers beat the transformer. Production models since 2025 follow that lesson: Qwen3-Next and Qwen3.5 (3:1 Gated DeltaNet to attention), Kimi Linear (KDA plus MLA), Nemotron-H (mostly Mamba-2), MiniMax (lightning attention). Meanwhile DeepSeek pursues compressed and sparse softmax attention instead. Honest status: hybrids are now mainstream for efficient long context, but full-attention designs remain common and the winner is unsettled.
19. What is the residual stream view and why is it useful?
The residual stream is a d-dimensional vector per token that every component reads from and adds to. Embeddings initialize it, attention heads move information across positions into it, FFNs write per-token transformations, and the unembedding reads the final state. Because contributions add, you can attribute outputs to components (circuits, logit lens), and it explains why small additive interventions (LoRA, steering vectors) and layer skipping work. It also frames attention as "routing" and FFN as "computation or memory."
20. How does a vision-language model process an image, and what does it cost in tokens?
A ViT splits the image into patches (for example 14×14 or 16×16 pixels), linearly embeds them and runs a bidirectional transformer. Encoders are usually CLIP or SigLIP, contrastively aligned with text. A projector (linear or MLP, sometimes a resampler or pooling) maps patch features into the LLM embedding space, and the resulting visual tokens are interleaved with text in the decoder's input. Cost: a 336×336 image at patch 14 is 576 tokens, and high-resolution tiling can reach thousands. Gemma 3 compresses to 256 per image. Training usually aligns the projector first (frozen towers), then instruction-tunes. Early-fusion models (Llama 4, Qwen3.5) train on mixed modalities from the start.
21. Why has vocabulary size grown from 32K to 200K+, and what are the downsides?
Larger vocabularies compress text better, especially multilingual text and code: fewer tokens per document means more effective context, fewer decode steps and lower per-character cost. Downsides: the embedding and LM-head matrices grow (\(V\times d\) each, a large share of parameters in small models), the output softmax over \(V\) costs more compute and memory, and rare tokens are under-trained (glitch tokens). There is also a data-sparsity trade-off per token. Practical sizes in 2025–26 are about 128K–262K.
22. What does MLA's "absorption" trick do and why does RoPE complicate it?
MLA caches a latent \(c\) and reconstructs per-head keys as \(k = W^{UK}c\). Since \(q^\top k = (W^{UK\top}q)^\top c\), you can precompute the product of \(W^{UK}\) with the query projection and attend directly against the cached latents, without ever materializing full K. The same holds for V with the output projection. RoPE inserts a position-dependent rotation \(R_m\) between them (\(q^\top R_{m-n} W^{UK} c\)), and because it depends on position it can't be folded into a fixed matrix. DeepSeek therefore splits off a small separate RoPE-carrying key (64 dimensions, shared across heads) that is cached alongside the latent.