
Mechanistic Interpretability Foundations
Analyze modern scientific methods for probing, ablating, and reverse-engineering the semantic roles of hidden neurons in deep networks.
Mechanistic Interpretability is the subfield of AI research dedicated to reverse-engineering neural networks into human-understandable computational algorithms.
Instead of treating the network as an inscrutable black box or relying on superficial saliency heatmaps, mechanistic interpretability treats a trained neural network like a complex physical or biological system. Researchers use controlled interventions—probing, patching, and ablating internal activations—to discover the functional sub-programs (called neural circuits) that the network synthesized during training.
┌─────────────────────────────────────────────────────────────────────────────────────────┐
│ THE MECHANISTIC INTERPRETABILITY TOOLKIT │
├──────────────────────────┬──────────────────────────────────────────────────────────────┤
│ METHOD │ HOW IT WORKS │
├──────────────────────────┼──────────────────────────────────────────────────────────────┤
│ 1. Activation Probing │ Read internal activation vectors a^(l) across diverse inputs │
│ │ to see what geometric directions align with specific traits. │
├──────────────────────────┼──────────────────────────────────────────────────────────────┤
│ 2. Feature Ablation │ Systematically zero out an activation (a_j^(l) = 0) or weight│
│ │ to measure its isolated causal impact on the final output. │
├──────────────────────────┼──────────────────────────────────────────────────────────────┤
│ 3. Activation Patching │ Swap a single activation from Run A into Run B to prove that │
│ │ a specific coordinate carries a specific piece of data. │
├──────────────────────────┼──────────────────────────────────────────────────────────────┤
│ 4. Circuit Identification│ Trace paths of high-magnitude weights across stacked layers │
│ │ to map end-to-end algorithms (e.g., edge -> shape -> object).│
└──────────────────────────┴──────────────────────────────────────────────────────────────┘
1. Activation Probing
To discover whether a hidden layer has formed a representation of a concept (such as "Startup Has Strong Revenue" or "Movie Has Happy Ending"), researchers train a simple linear probe on the hidden activation vector .
A linear probe is a vector of probe weights evaluated as:
If a simple dot product on can reliably predict whether the concept is present across thousands of test examples, the concept exists as a linearly accessible direction in that hidden layer's coordinate space.
2. Feature Ablation (Knockout Experiments)
Probing shows correlation (the activation correlates with a concept). To prove causation, researchers use Feature Ablation.
In an ablation experiment:
- We run an intact forward pass on input and record baseline prediction .
- We force a specific hidden activation coordinate to zero:
- We continue the forward pass through downstream layers to produce ablated prediction .
- We measure the causal output shift:
[ INTACT PASS ]
x ──► [ Layer 1 ] ──► a^(1) = [ 2.0, 4.0 ]^T ──► [ Layer 2 ] ──► y_hat_intact = 1.20
[ ABLATED PASS (Zeroing Neuron 2) ]
x ──► [ Layer 1 ] ──► a^(1)_ablated = [ 2.0, 0.0 ]^T ──► [ Layer 2 ] ──► y_hat_ablated = 3.20
Causal Shift: Δy_hat = 3.20 - 1.20 = +2.00
If zeroing neuron causes the network's prediction on a specific task to collapse while leaving other predictions intact, neuron is a causal component of that decision circuit.
3. Neural Circuits: Discovering Learned Sub-Programs
When we combine probing and ablation across multiple layers, we discover that deep networks structure themselves into hierarchical neural circuits:
- Early Layers (Low-Level Primitives):
- In Computer Vision: Individual neurons detect elementary orientation bars, color contrasts, and spatial edges ().
- In Natural Language: Early token embeddings and hidden states capture token identity, capitalization, punctuation patterns, and parts of speech.
- Intermediate Layers (Mid-Level Combinations):
- In Computer Vision: Edge detectors feed into intermediate neurons that activate for curves, corners, textures, and geometric junctions.
- In Natural Language: Low-level features combine into phrase-level syntactic relations, subject-verb agreement circuits, and entity boundaries.
- Late Layers (High-Level Task Representations):
- In Computer Vision: Curve and texture circuits combine into high-level detectors for object parts (wheels, dog ears, human eyes, window frames).
- In Natural Language: Intermediate circuits combine into semantic fact retrieval, sentiment analysis, context tracking, and task execution logic.
══════════════════════════════════════════════════════════════════════════════════════════════════════
HIERARCHICAL NEURAL CIRCUIT ASSEMBLY
══════════════════════════════════════════════════════════════════════════════════════════════════════
[ INPUT x ] [ LAYER 1: PRIMITIVES ] [ LAYER 2: COMBINATIONS ] [ OUTPUT DECISION ]
Raw Measurements ───► Oriented Edge Detectors ──► Curve & Corner Detectors ──► Object Classification
(Pixels / Features) (Low-Level Features) (Intermediate Shapes) (High-Level Prediction)
══════════════════════════════════════════════════════════════════════════════════════════════════════
Deep learning does not store arbitrary memorized look-up tables. Through gradient descent, it compiles an organized hierarchy of composable linear-algebraic sub-programs.