
LLM Foundations: Architecture & Inference
See how large language models generate text and predict the next word from first principles.
What is Covered
- Tokenization Mechanics: Byte-Pair Encoding (BPE) and embedding vectors.
- Attention Engines: Query-Key-Value matrices, scaled dot-product attention, and multi-head projection.
- Inference Pipeline: Autoregressive decoding, KV caching, and sampling strategies.
0 lessons