
From Vanilla MLPs to Modern Transformers
Synthesize core vanilla MLP principles and establish the structural bridge to token embeddings, self-attention, and large language models.
Over the course of this curriculum, we traced an unbroken mathematical progression:
A natural question arises: How does the multi-layer perceptron we built on paper connect to modern frontier AI models like GPT, Claude, and Gemini?
The answer is direct: Vanilla MLPs are the core computational engines inside every modern Transformer.
The Transformer Sublayer Architecture
A modern Large Language Model is not an alien architecture that discarded classical linear algebra. It is a deep stack of identical Transformer blocks (often 32 to 128 layers deep).
Inside every single Transformer block, there are two primary sublayers:
- The Multi-Head Attention Sublayer: Routes information between different token positions across the sequence (mixing spatial context).
- The Feed-Forward MLP Sublayer: Processes each token vector individually through a classical Multi-Layer Perceptron (storing factual knowledge and executing non-linear reasoning).
┌─────────────────────────────────────────────────────────────────────────────────────────┐
│ A SINGLE TRANSFORMER BLOCK │
│ │
│ Token Input Vector x ∈ ℝ^(d_model) │
│ │ │
│ ▼ │
│ ┌─────────────────────────────────────────────────────────────┐ │
│ │ Sublayer 1: Multi-Head Self-Attention │ │
│ │ (Routes information across token positions in sequence) │ │
│ └─────────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌─────────────────────────────────────────────────────────────┐ │
│ │ Sublayer 2: Feed-Forward MLP Network (MLP Sublayer) │ ◄── [OUR EXACT MLP] │
│ │ │ │
│ │ 1. Up-Projection: z_ff = W_up x + b_up (d_model ──► 4 d_model) │
│ │ 2. Non-Linearity: h_ff = f(z_ff) (ReLU / GeLU / SwiGLU via W_gate) │
│ │ 3. Down-Projection: y_ff = W_down h_ff (4 d_model ──► d_model) │
│ └─────────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ Updated Token Vector y ∈ ℝ^(d_model) │
└─────────────────────────────────────────────────────────────────────────────────────────┘
The Feed-Forward MLP sublayer accounts for roughly two-thirds of all trainable parameters in modern Large Language Models.
The Squeeze-and-Expand Geometry of Transformer MLPs
The MLP sublayer inside a Transformer executes the exact affine and activation equations mastered in Course 1:
- Input Representation: A single token is represented as a high-dimensional vector (in frontier models, or ).
- Up-Projection (Expansion): The vector is linearly projected from into a wider intermediate expansion dimension using weight matrix (and in modern gated architectures like SwiGLU, alongside a parallel gating projection matrix ):
- Non-Linear Activation: An element-wise non-linear activation function () is applied to create the intermediate activation state: (While we used ReLU and Sigmoid in Course 1, modern LLMs use smooth variants such as GeLU or gated activations like SwiGLU where ).
- Down-Projection (Contraction): The activated features are projected back down to the model dimension using weight matrix :
This wide intermediate layer () acts as a high-capacity key-value associative memory. The up-projection vectors detect specific concepts, features, or factual relationships; the non-linear activation isolates relevant detectors; and the down-projection writes the associated information back into the token's coordinate vector.
TEASER: Course 2 Preview: The Road to LLMs In Course 2 (LLM Foundations), we will expand from static 4D genre vectors into continuous token embeddings (), multi-head Query-Key-Value () dot-product attention, positional encodings, and autoregressive generation. Everything you learned in Course 1—matrix multiplication, affine sums, non-linear activations, loss functions, and backpropagation—forms the foundation of that architecture.
The Unbroken Foundation
You now possess the foundational mathematical literacy of artificial intelligence:
- When you read about model parameters, you see explicit weight matrices and bias vectors .
- When you read about inference, you trace dot products, affine additions, and non-linear squashing functions across a Directed Acyclic Graph.
- When you read about training and fine-tuning, you trace prediction errors back through transposed matrices via partial derivatives and multivariable gradient descent.
- When you read about interpretability and safety, you understand latent representations, linear separability, and circuit probing.
The foundations are complete. Every advanced model in deep learning is built on these exact mathematical bricks.