Mechanistic Interpretability Foundations hero
Lesson 3Interpretability of Neural Networks

Mechanistic Interpretability Foundations

Analyze modern scientific methods for probing, ablating, and reverse-engineering the semantic roles of hidden neurons in deep networks.

Mechanistic Interpretability is the subfield of AI research dedicated to reverse-engineering neural networks into human-understandable computational algorithms.

Instead of treating the network as an inscrutable black box or relying on superficial saliency heatmaps, mechanistic interpretability treats a trained neural network like a complex physical or biological system. Researchers use controlled interventions—probing, patching, and ablating internal activations—to discover the functional sub-programs (called neural circuits) that the network synthesized during training.

┌─────────────────────────────────────────────────────────────────────────────────────────┐
│                      THE MECHANISTIC INTERPRETABILITY TOOLKIT                           │
├──────────────────────────┬──────────────────────────────────────────────────────────────┤
│ METHOD                   │ HOW IT WORKS                                                 │
├──────────────────────────┼──────────────────────────────────────────────────────────────┤
│ 1. Activation Probing    │ Read internal activation vectors a^(l) across diverse inputs │
│                          │ to see what geometric directions align with specific traits. │
├──────────────────────────┼──────────────────────────────────────────────────────────────┤
│ 2. Feature Ablation      │ Systematically zero out an activation (a_j^(l) = 0) or weight│
│                          │ to measure its isolated causal impact on the final output.   │
├──────────────────────────┼──────────────────────────────────────────────────────────────┤
│ 3. Activation Patching   │ Swap a single activation from Run A into Run B to prove that │
│                          │ a specific coordinate carries a specific piece of data.      │
├──────────────────────────┼──────────────────────────────────────────────────────────────┤
│ 4. Circuit Identification│ Trace paths of high-magnitude weights across stacked layers  │
│                          │ to map end-to-end algorithms (e.g., edge -> shape -> object).│
└──────────────────────────┴──────────────────────────────────────────────────────────────┘

1. Activation Probing

To discover whether a hidden layer has formed a representation of a concept (such as "Startup Has Strong Revenue" or "Movie Has Happy Ending"), researchers train a simple linear probe on the hidden activation vector a(l)a^{(l)}.

A linear probe is a vector of probe weights wprobe∈Rdlw_{\text{probe}} \in \mathbb{R}^{d_l} evaluated as:

sprobe=wprobe⋅a(l)s_{\text{probe}} = w_{\text{probe}} \cdot a^{(l)}

If a simple dot product on a(l)a^{(l)} can reliably predict whether the concept is present across thousands of test examples, the concept exists as a linearly accessible direction in that hidden layer's coordinate space.


2. Feature Ablation (Knockout Experiments)

Probing shows correlation (the activation correlates with a concept). To prove causation, researchers use Feature Ablation.

In an ablation experiment:

  1. We run an intact forward pass on input xx and record baseline prediction y^intact\hat{y}_{\text{intact}}.
  2. We force a specific hidden activation coordinate to zero: aj(l)⟵0.0a_j^{(l)} \longleftarrow 0.0
  3. We continue the forward pass through downstream layers to produce ablated prediction y^ablated\hat{y}_{\text{ablated}}.
  4. We measure the causal output shift: Δy^=y^ablated−y^intact\Delta \hat{y} = \hat{y}_{\text{ablated}} - \hat{y}_{\text{intact}}
[ INTACT PASS ]
  x ──► [ Layer 1 ] ──► a^(1) = [ 2.0, 4.0 ]^T ──► [ Layer 2 ] ──► y_hat_intact = 1.20

[ ABLATED PASS (Zeroing Neuron 2) ]
  x ──► [ Layer 1 ] ──► a^(1)_ablated = [ 2.0, 0.0 ]^T ──► [ Layer 2 ] ──► y_hat_ablated = 3.20

  Causal Shift: Δy_hat = 3.20 - 1.20 = +2.00

If zeroing neuron jj causes the network's prediction on a specific task to collapse while leaving other predictions intact, neuron jj is a causal component of that decision circuit.


3. Neural Circuits: Discovering Learned Sub-Programs

When we combine probing and ablation across multiple layers, we discover that deep networks structure themselves into hierarchical neural circuits:

  1. Early Layers (Low-Level Primitives):
    • In Computer Vision: Individual neurons detect elementary orientation bars, color contrasts, and spatial edges (0∘,45∘,90∘0^\circ, 45^\circ, 90^\circ).
    • In Natural Language: Early token embeddings and hidden states capture token identity, capitalization, punctuation patterns, and parts of speech.
  2. Intermediate Layers (Mid-Level Combinations):
    • In Computer Vision: Edge detectors feed into intermediate neurons that activate for curves, corners, textures, and geometric junctions.
    • In Natural Language: Low-level features combine into phrase-level syntactic relations, subject-verb agreement circuits, and entity boundaries.
  3. Late Layers (High-Level Task Representations):
    • In Computer Vision: Curve and texture circuits combine into high-level detectors for object parts (wheels, dog ears, human eyes, window frames).
    • In Natural Language: Intermediate circuits combine into semantic fact retrieval, sentiment analysis, context tracking, and task execution logic.
══════════════════════════════════════════════════════════════════════════════════════════════════════
                               HIERARCHICAL NEURAL CIRCUIT ASSEMBLY
══════════════════════════════════════════════════════════════════════════════════════════════════════

  [ INPUT x ]               [ LAYER 1: PRIMITIVES ]      [ LAYER 2: COMBINATIONS ]    [ OUTPUT DECISION ]
  Raw Measurements   ───►   Oriented Edge Detectors ──►  Curve & Corner Detectors ──► Object Classification
  (Pixels / Features)       (Low-Level Features)         (Intermediate Shapes)        (High-Level Prediction)

══════════════════════════════════════════════════════════════════════════════════════════════════════

Deep learning does not store arbitrary memorized look-up tables. Through gradient descent, it compiles an organized hierarchy of composable linear-algebraic sub-programs.


Previous
Interpretable Weights vs. Latent Geometry