Archive
Papers Primary sources for each step of the stack, in the order the treatise uses them. A paper earns its place by naming a map, a Jacobian, a floor, or a default this site actually writes.
What we can legally show
We do not host publisher PDFs. Open-access copies — arXiv author postings and official technical reports — are linked, and arXiv HTML can be read here. Closed works go to the DOI. That is the legal path; a local copy of Nature or IEEE would not be.
35 works
1948 A mathematical theory of communication Shannon, C. E. Entropy as the irreducible uncertainty of a discrete source. Every later claim that loss cannot hit zero on language is this paper, applied to tokens. Publisher 1951 Prediction and entropy of printed English Shannon, C. E. The floor is a measurement. Humans who have seen the preceding text still cannot guess the next character with certainty. Conditional entropy of English is low, not zero. Publisher 2006 Elements of information theory (2nd ed.) Cover, T. M., and Thomas, J. A. The identity the loss chapter writes without apology: cross-entropy = entropy + KL. Training is KL minimization in disguise. Publisher 1990 Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition Bridle, J. S. Softmax as a categorical distribution, and the collapse of softmax-plus-cross-entropy to p − y. The treatise’s most important gradient is this identification. Publisher 1986 Learning representations by back-propagating errors Rumelhart, D. E., Hinton, G. E., and Williams, R. J. The public statement that a scalar loss can be routed to every weight by the chain rule, layer by layer, in reverse. Publisher 2018 Automatic differentiation in machine learning: a survey Baydin, A. G., Pearlmutter, B. A., Radul, A. A., and Siskind, J. M. Backprop as reverse-mode automatic differentiation: vector–Jacobian products, not materialized Jacobians. The treatise’s ‘nobody forms J’ is this paper’s point. arXiv 1994 A new algorithm for data compression Gage, P. Byte-pair encoding as a compression algorithm: repeatedly merge the most frequent adjacent pair. The tokenizer is this idea, frozen, pointed at text. Publisher 2016 Neural machine translation of rare words with subword units Sennrich, R., Haddow, B., and Birch, A. BPE as the frozen map from bytes to the integers the treatise trains on. Loss, perplexity, and ‘what zero would mean’ are defined on this vocabulary. arXiv 2018 SentencePiece: a simple and language independent subword tokenizer and detokenizer Kudo, T., and Richardson, J. The other common tokenizer: unigram LM / BPE trained from raw bytes, language-agnostic, the Llama-class default alongside BPE. arXiv 2015 Neural machine translation by jointly learning to align and translate Bahdanau, D., Cho, K., and Bengio, Y. Attention as a content-based weighted sum over source states. Vaswani replaces the scoring function; the mixing primitive is already here. arXiv 2016 Deep residual learning for image recognition He, K., Zhang, X., Ren, S., and Sun, J. The skip whose Jacobian is I. Why a 32-layer Transformer trains: there is always a path that does not multiply by a layer Jacobian. arXiv 2016 Layer normalization Ba, J. L., Kiros, J. R., and Hinton, G. E. Per-token mean and variance. GPT-2’s norm. The RMSNorm Jacobian in the treatise is this operator with the mean step removed. arXiv 2017 Attention is all you need Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Scaled dot-product attention, the √d_h factor, multi-head, sinusoidal position, the original post-norm stack. The mixing equation the treatise writes is this one. arXiv 2017 Using the output embedding to improve language models Press, O., and Wolf, L. Weight tying: W_U = W_E^⊤. Two uses, one parameter. The unembedding chapter’s optional identification. arXiv 2017 SGDR: stochastic gradient descent with warm restarts Loshchilov, I., and Hutter, F. Cosine annealing of the learning rate. The schedule the treatise writes (warmup, then cosine to η_min) is this shape, usually without the restarts. arXiv 2016 Gaussian error linear units (GELUs) Hendrycks, D., and Gimpel, K. GPT-2’s nonlinearity. Smooth, not a hard gate. The older MLP in the treatise before SwiGLU. arXiv 2018 Mixed precision training Micikevicius, P., et al. BF16/FP16 forward, FP32 master weights. An implementation detail the training loop names so the math is not confused with the storage format. arXiv 2019 Language models are unsupervised multitask learners Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. The decoder-only language model this treatise writes: causal mask, next-token MLE, GELU FFN. Capabilities as a side effect of the conditionals. Author PDF 2019 Root mean square layer normalization Zhang, B., and Sennrich, R. RMSNorm: scale without mean-centering. The Llama default, and the Jacobian the backprop chapter writes. arXiv 2019 Fast Transformer decoding: one write-head is all you need Shazeer, N. Multi-query attention. The KV cache, not the matmul, bounds generation. One K/V head shared across all query heads. arXiv 2015 Adam: a method for stochastic optimization Kingma, D. P., and Ba, J. Per-parameter first and second moments. The adaptive step the treatise starts from before the ‘W’. arXiv 2019 Decoupled weight decay regularization Loshchilov, I., and Hutter, F. AdamW: λθ sits outside the second-moment rescaling. That is the entire difference, and it is the LLM default. arXiv 2020 GLU variants improve Transformer Shazeer, N. SwiGLU: the gated MLP Llama uses. Gate, up, down — three matrices, no bias. The forward and backward the treatise writes. arXiv 2020 On layer normalization in the Transformer architecture Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T. Pre-norm trains at depth; post-norm was the original and is harder. The block equation in the treatise is the pre-norm form this paper recommends. arXiv 2020 Scaling laws for neural language models Kaplan, J., et al. C ≈ 6ND. The compute unit. Power-law loss in N and D. Their N-versus-D allocation is the one Chinchilla later revised. arXiv 2020 The curious case of neural text degeneration Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. Nucleus / top-p. Inference carves a sampling set from the leftover mass, then renormalizes. No gradient. Also: typical human text is not the mode. arXiv 2021 RoFormer: enhanced Transformer with rotary position embedding Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y. RoPE. Inner products depend on t − s, not on absolute t and s. No extra residual vector. The rotation the embeddings chapter writes. arXiv 2022 FlashAttention: fast and memory-efficient exact attention with IO-awareness Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C. Same softmax Jacobian, different IO. The T×T matrix does not have to live in HBM. Exact attention, not an approximation. arXiv 2022 Training compute-optimal large language models Hoffmann, J., Borgeaud, S., Mensch, A., et al. The Chinchilla law. E is the fitted irreducible loss — the floor the treatise refuses to call zero. IsoFLOP said roughly 20 tokens per parameter. arXiv 2023 Efficiently scaling Transformer inference Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Heek, J., Xiao, K., Agrawal, S., and Dean, J. Prefill versus decode, and why the KV cache is the resource that bounds serving. The inference chapter’s operational picture. arXiv 2023 FlashAttention-2: faster attention with better parallelism and work partitioning Dao, T. The production kernel. Still exact. Still the same Jacobian. Parallelism and work partitioning, not a new formula. arXiv 2023 LLaMA: open and efficient foundation language models Touvron, H., Lavril, T., Izacard, G., et al. The modern stack this treatise writes in one notation: RoPE, RMSNorm, SwiGLU, no biases, pre-norm, decoder-only. arXiv 2023 GQA: training generalized multi-query Transformer models from multi-head checkpoints Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S. Grouped-query attention: H_kv between 1 and H_q. The Llama-3-class default. Interpolates MHA and MQA. arXiv 2024 Chinchilla Scaling: a replication attempt Besiroglu, T., Erdil, E., and Barnett, M. The printed Approach-3 constants do not recover the 20:1 policy. E is closer to 1.82 on their refit. The floor is an estimate, not a constant of nature. arXiv 2024 The Llama 3 herd of models Dubey, A., et al. (Llama Team) The current default: GQA, RoPE, RMSNorm, SwiGLU, document mask, a much larger data mix. Confirmation that the treatise’s stack is the one in production, not a 2017 museum piece. arXiv