LLM Foundations: Architecture & Inference

LLM Foundations: Architecture & Inference

See how large language models generate text and predict the next word from first principles.

What is Covered

  • Tokenization Mechanics: Byte-Pair Encoding (BPE) and embedding vectors.
  • Attention Engines: Query-Key-Value matrices, scaled dot-product attention, and multi-head projection.
  • Inference Pipeline: Autoregressive decoding, KV caching, and sampling strategies.
0 lessons