Language Understanding and Language Models¶
Table of Contents
Activations¶
Normalisation¶
[Internal Covariate Shift][BN] Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
[LN] Layer Normalization
[RMSNorm] Root Mean Square Layer Normalization
[PreLN][Detailed Study with Mean-Field Theory] On Layer Normalization in the Transformer Architecture
Warning
For theoretical understanding of MFT and NTK, start from this MLSS video here.
Word Embeddings¶
Note
Word2Vec: Efficient Estimation of Word Representations in Vector Space
GloVe: Global Vectors forWord Representation
Evaluation methods for unsupervised word embeddings
Sequence Modeling¶
RNN¶
See also
Implementation examples live in Language Modeling Implementations.
Transformer¶
General Resources¶
Warning
Note
[harvard.edu] The Annotated Transformer
[jalammar.github.io] The Illustrated Transformer
[lilianweng.github.io] Attention? Attention!
[d2l.ai] The Transformer Architecture
[newsletter.languagemodels.co] The Illustrated DeepSeek-R1: A recipe for reasoning LLMs
Position Encoding¶
Note
[arxiv.org] Position Information in Transformers: An Overview
[arxiv.org] Rethinking Positional Encoding in Language Pre-training
[eleuther.ai] RoPE
[arxiv.org] LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
[arxiv.org] RoFormer: Enhanced Transformer with Rotary Position Embedding
Attention¶
Understanding Einsum¶
Warning
Implementation examples live in Language Modeling Implementations.
Note
Attention implementation examples live in Language Modeling Implementations.
UnitTest¶
See also
Unit tests live in Language Modeling Implementations.
Decoding¶
Beam Search, Top-K, Top-p/Nuclear, Temperature
[mlabonne.github.io] Decoding Strategies in Large Language Models
Speculative Deocding
Transformer Architecture¶
Encoder [BERT]¶
Note
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Additional Resources
[tinkerd.net] BERT Tokenization
[tinkerd.net] BERT Embeddings,
[tinkerd.net] BERT Encoder Layer
TwinBERT: Distilling Knowledge to Twin-Structured BERT Models for Efficient Retrieval
Decoder [GPT]¶
Note
[jalammar.github.io] The Illustrated GPT-2
[github.com] karpathy/nanoGPT
[cameronrwolfe.substack.com] Decoder-Only Transformers: The Workhorse of Generative LLMs
[openai.com] GPT-2: Language Models are Unsupervised Multitask Learners
[openai.com] GPT-3: Language Models are Few-Shot Learners
Encoder-Decoder [T5]¶
Autoencoder [BART]¶
Cross-Lingual¶
Note
[Encoder] XLM-R [Roberta]: Unsupervised Cross-lingual Representation Learning at Scale
[Decoder] XGLM [GPT-3]: Few-shot Learning with Multilingual Generative Language Models
[Encoder-Decoder] mT5 [T5]: A Massively Multilingual Pre-trained Text-to-Text Transformer
[Autoencoder] mBART [BART]: Multilingual Denoising Pre-training for Neural Machine Translation