Skip to content

Index

Symbols

Notation is stable from chapter 00 to 14. Each symbol is defined again where it is used. This page is the global list. is only the unembedding; the SwiGLU up-projection is .

The network: token IDs to a T×V logit matrix.

ch. object

Every learnable parameter (embeddings, projections, norms).

ch. object

Vocabulary size — number of distinct tokens.

ch. object

Sequence length of this forward pass (context window used).

ch. object

Model / residual-stream width.

ch. object

Number of Transformer layers.

ch. object

Number of attention heads (MHA: H_q = H_kv = H).

ch. attention

Query heads and key/value heads. Equal in MHA; H_kv = 1 is MQA.

ch. attention

Head dimension, usually d/H.

ch. attention

Token-ID sequence, each x_t ∈ {1,…,V}.

ch. object

Autoregressive next-token distribution of the model.

ch. object

Token embedding matrix, V×d.

ch. embeddings

Absolute position embedding matrix (GPT-2 style).

ch. embeddings

Residual stream after embedding, T×d.

ch. embeddings

RoPE base frequency ω^{−2i/d_h}.

ch. embeddings

Relative RoPE rotation. ⟨R_t q, R_s k⟩ = qᵀ R_{s−t} k.

ch. embeddings

Learned RMSNorm / LayerNorm scale, in R^d.

ch. block

Small constant for numerical stability (e.g. 10⁻⁶, 10⁻⁸).

ch. block

Query, key, value projection matrices.

ch. attention

Attention scores QKᵀ / √d_h, T×T.

ch. attention

Causal mask: 0 on and below the diagonal, −∞ above.

ch. attention

Attention weights after softmax. Rows sum to 1.

ch. attention

Output projection that merges heads.

ch. attention

diag(a) − aaᵀ. Symmetric, null space along 1.

ch. attention

z · σ(z). Swish activation used in SwiGLU.

ch. mlp

SwiGLU up-projection, d × d_ff. Not the unembedding.

ch. mlp

SwiGLU gate and down projections.

ch. mlp

Unembedding matrix, d×V. Sometimes tied to W_Eᵀ.

ch. unembed

Logits at position t, in R^V.

ch. unembed

Softmax of z_t. A categorical over the vocabulary.

ch. unembed

Token-level cross-entropy / NLL, in nats.

ch. loss

Entropy of the true next-token distribution. Irreducible.

ch. loss

Extra loss from the model being wrong.

ch. loss

Perplexity e^L. Equivalent number of equally likely tokens.

ch. loss

Bits per byte: (L / ln 2) · (T / B). Tokenizer-fair loss.

ch. loss

Fitted irreducible loss in a scaling law.

ch. loss

Parameter count (scaling laws).

ch. loss

Training tokens (scaling laws).

ch. loss

Training compute. C ≈ 6ND FLOPs.

ch. loss

Jacobian ∂v/∂u of a map g.

ch. backprop

Vector-Jacobian product. What reverse-mode AD actually computes.

ch. backprop

Gradient of CE+softmax w.r.t. logits.

ch. backprop

Learning rate (scheduled).

ch. adamw

Adam moment decays. Typical 0.9 and 0.95/0.999.

ch. adamw

Decoupled weight decay coefficient.

ch. adamw

Sampling temperature. τ → 0 is greedy.

ch. inference

Stored keys for past tokens, per layer.

ch. inference