Interpretability of Neural Networks
Examine the boundary of transparent arithmetic and latent representations, explore mechanistic interpretability, and bridge vanilla MLPs to Transformers.
Lessons in this Topic
The True Nature of the "Black Box"
Examine the boundary of transparent arithmetic and latent representations, explore mechanistic interpretability, and bridge vanilla MLPs to Transformers.
Interpretable Weights vs. Latent Geometry
Contrast human-interpretable feature weights against high-dimensional latent coordinate geometry inside multi-layer representations.
Mechanistic Interpretability Foundations
Analyze modern scientific methods for probing, ablating, and reverse-engineering the semantic roles of hidden neurons in deep networks.
From Vanilla MLPs to Modern Transformers
Synthesize core vanilla MLP principles and establish the structural bridge to token embeddings, self-attention, and large language models.
Neural Interpretability In Practice
Master latent representation geometry, linear separability, feature ablation, and Transformer MLP sublayer calculations through manual hand traces.