Neural Interpretability In Practice hero
LaboratoryInterpretability of Neural Networks

Neural Interpretability In Practice

Master latent representation geometry, linear separability, feature ablation, and Transformer MLP sublayer calculations through manual hand traces.

Work through every problem on paper before expanding the step-by-step solutions.


Part 1: Abstract Mechanics

Work through these 5 drill problems testing latent transformations, linear separability, feature ablation, Transformer MLP dimensions, and token projection math.


Problem 1: Latent Space Coordinate Transformation Mapping

A 2-input, 2-neuron hidden layer has the following configuration:

  • Input vector: x=[2.0−1.0]x = \begin{bmatrix} 2.0 \\ -1.0 \end{bmatrix}
  • Weight matrix: W(1)=[1.5−1.0−0.52.0]W^{(1)} = \begin{bmatrix} 1.5 & -1.0 \\ -0.5 & 2.0 \end{bmatrix}
  • Bias vector: b(1)=[−1.01.0]b^{(1)} = \begin{bmatrix} -1.0 \\ 1.0 \end{bmatrix}
  • Activation function: ReLU(z)=max⁡(0,z)\text{ReLU}(z) = \max(0, z)

Compute:

  1. The pre-activation vector z(1)=W(1)x+b(1)z^{(1)} = W^{(1)} x + b^{(1)}.
  2. The hidden latent activation vector a(1)=ReLU(z(1))a^{(1)} = \text{ReLU}(z^{(1)}).
Reveal Solution
  1. Pre-Activation Vector z(1)z^{(1)}:

    z1(1)=(1.5)(2.0)+(−1.0)(−1.0)−1.0=3.0+1.0−1.0=+3.0z_1^{(1)} = (1.5)(2.0) + (-1.0)(-1.0) - 1.0 = 3.0 + 1.0 - 1.0 = \mathbf{+3.0} z2(1)=(−0.5)(2.0)+(2.0)(−1.0)+1.0=−1.0−2.0+1.0=−2.0z_2^{(1)} = (-0.5)(2.0) + (2.0)(-1.0) + 1.0 = -1.0 - 2.0 + 1.0 = \mathbf{-2.0} z(1)=[+3.0−2.0]z^{(1)} = \begin{bmatrix} +3.0 \\ -2.0 \end{bmatrix}
  2. Hidden Activation Vector a(1)a^{(1)}:

    a(1)=ReLU([+3.0−2.0])=[max⁡(0,3.0)max⁡(0,−2.0)]=[3.00.0]a^{(1)} = \text{ReLU}\left(\begin{bmatrix} +3.0 \\ -2.0 \end{bmatrix}\right) = \begin{bmatrix} \max(0, 3.0) \\ \max(0, -2.0) \end{bmatrix} = \begin{bmatrix} \mathbf{3.0} \\ \mathbf{0.0} \end{bmatrix}

Problem 2: Linear Separability Classification Check

Consider 4 points in a 2D latent representation space R2\mathbb{R}^2:

  • Class A (y=+1y = +1): p1=[3.01.0],p2=[2.03.0]p_1 = \begin{bmatrix} 3.0 \\ 1.0 \end{bmatrix}, \quad p_2 = \begin{bmatrix} 2.0 \\ 3.0 \end{bmatrix}
  • Class B (y=−1y = -1): p3=[0.01.0],p4=[1.00.0]p_3 = \begin{bmatrix} 0.0 \\ 1.0 \end{bmatrix}, \quad p_4 = \begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix}

A candidate linear separating hyperplane has parameter vector wsep=[1.01.0]w_{\text{sep}} = \begin{bmatrix} 1.0 \\ 1.0 \end{bmatrix} and threshold bias bsep=−3.0b_{\text{sep}} = -3.0, defining the decision function:

s(p)=wsep⋅p+bseps(p) = w_{\text{sep}} \cdot p + b_{\text{sep}}

Where a point is classified as Class A if s(p)>0s(p) > 0, and Class B if s(p)<0s(p) < 0.

Compute the decision score s(pi)s(p_i) for each of the 4 points and verify whether this hyperplane cleanly separates Class A from Class B.

Reveal Solution
  1. Score for Point p1=[3.0,1.0]Tp_1 = [3.0, 1.0]^T (Class A):

    s(p1)=(1.0)(3.0)+(1.0)(1.0)−3.0=3.0+1.0−3.0=+1.0>0  ⟹  Class A (Correct)s(p_1) = (1.0)(3.0) + (1.0)(1.0) - 3.0 = 3.0 + 1.0 - 3.0 = \mathbf{+1.0} > 0 \implies \mathbf{\text{Class A (Correct)}}
  2. Score for Point p2=[2.0,3.0]Tp_2 = [2.0, 3.0]^T (Class A):

    s(p2)=(1.0)(2.0)+(1.0)(3.0)−3.0=2.0+3.0−3.0=+2.0>0  ⟹  Class A (Correct)s(p_2) = (1.0)(2.0) + (1.0)(3.0) - 3.0 = 2.0 + 3.0 - 3.0 = \mathbf{+2.0} > 0 \implies \mathbf{\text{Class A (Correct)}}
  3. Score for Point p3=[0.0,1.0]Tp_3 = [0.0, 1.0]^T (Class B):

    s(p3)=(1.0)(0.0)+(1.0)(1.0)−3.0=0.0+1.0−3.0=−2.0<0  ⟹  Class B (Correct)s(p_3) = (1.0)(0.0) + (1.0)(1.0) - 3.0 = 0.0 + 1.0 - 3.0 = \mathbf{-2.0} < 0 \implies \mathbf{\text{Class B (Correct)}}
  4. Score for Point p4=[1.0,0.0]Tp_4 = [1.0, 0.0]^T (Class B):

    s(p4)=(1.0)(1.0)+(1.0)(0.0)−3.0=1.0+0.0−3.0=−2.0<0  ⟹  Class B (Correct)s(p_4) = (1.0)(1.0) + (1.0)(0.0) - 3.0 = 1.0 + 0.0 - 3.0 = \mathbf{-2.0} < 0 \implies \mathbf{\text{Class B (Correct)}}

Conclusion: Since s(p1)>0,s(p2)>0s(p_1) > 0, s(p_2) > 0 and s(p3)<0,s(p4)<0s(p_3) < 0, s(p_4) < 0, the 4 points are strictly linearly separable by the hyperplane wsep⋅p+bsep=0w_{\text{sep}} \cdot p + b_{\text{sep}} = 0.


Problem 3: Hidden Neuron Feature Ablation Calculation

A 2-layer network with 2 hidden neurons and a linear output neuron has parameters:

  • Hidden activation vector: a(1)=[2.04.0]a^{(1)} = \begin{bmatrix} 2.0 \\ 4.0 \end{bmatrix}
  • Output weight matrix: W(2)=[1.5−0.5]W^{(2)} = \begin{bmatrix} 1.5 & -0.5 \end{bmatrix}
  • Output bias: b(2)=0.20b^{(2)} = 0.20
  • Output prediction: y^=W(2)a(1)+b(2)\hat{y} = W^{(2)} a^{(1)} + b^{(2)}

Perform an ablation experiment:

  1. Compute the intact output prediction y^intact\hat{y}_{\text{intact}}.
  2. Ablate hidden neuron 2 (set a2(1)=0.0a_2^{(1)} = 0.0) and compute the ablated output prediction y^ablated\hat{y}_{\text{ablated}}.
  3. Compute the exact causal output shift Δy^=y^ablated−y^intact\Delta \hat{y} = \hat{y}_{\text{ablated}} - \hat{y}_{\text{intact}}.
Reveal Solution
  1. Intact Output Prediction:

    y^intact=W(2)a(1)+b(2)=(1.5)(2.0)+(−0.5)(4.0)+0.20=3.0−2.0+0.20=1.20\hat{y}_{\text{intact}} = W^{(2)} a^{(1)} + b^{(2)} = (1.5)(2.0) + (-0.5)(4.0) + 0.20 = 3.0 - 2.0 + 0.20 = \mathbf{1.20}
  2. Ablated Output Prediction (a2(1)=0.0a_2^{(1)} = 0.0):

    aablated(1)=[2.00.0]a^{(1)}_{\text{ablated}} = \begin{bmatrix} 2.0 \\ 0.0 \end{bmatrix} y^ablated=(1.5)(2.0)+(−0.5)(0.0)+0.20=3.0+0.0+0.20=3.20\hat{y}_{\text{ablated}} = (1.5)(2.0) + (-0.5)(0.0) + 0.20 = 3.0 + 0.0 + 0.20 = \mathbf{3.20}
  3. Causal Output Shift Δy^\Delta \hat{y}:

    Δy^=y^ablated−y^intact=3.20−1.20=+2.00\Delta \hat{y} = \hat{y}_{\text{ablated}} - \hat{y}_{\text{intact}} = 3.20 - 1.20 = \mathbf{+2.00}

(Interpretation: Hidden neuron 2 was exerting a strong inhibitory effect of −0.5×4.0=−2.00-0.5 \times 4.0 = -2.00 on the output. Ablating it removes that inhibition, increasing the final prediction by +2.00+2.00).


Problem 4: Transformer MLP Sublayer Dimension Trace

A miniaturized Transformer model has token embedding dimension dmodel=4d_{\text{model}} = 4. Its Feed-Forward MLP sublayer expands the token vector by a factor of 4 to intermediate dimension dff=4dmodel=16d_{\text{ff}} = 4 d_{\text{model}} = 16, then projects it back to dmodel=4d_{\text{model}} = 4.

The sublayer equations are:

zff=Wupx+bup,hff=ReLU(zff),ymlp=Wdownhff+bdownz_{\text{ff}} = W_{\text{up}} x + b_{\text{up}}, \qquad h_{\text{ff}} = \text{ReLU}(z_{\text{ff}}), \qquad y_{\text{mlp}} = W_{\text{down}} h_{\text{ff}} + b_{\text{down}}

For an input token vector x∈R4×1x \in \mathbb{R}^{4 \times 1}:

  1. State the exact matrix dimensions of Wup,bup,WdownW_{\text{up}}, b_{\text{up}}, W_{\text{down}}, and bdownb_{\text{down}}.
  2. State the vector dimensions of zff,hffz_{\text{ff}}, h_{\text{ff}}, and ymlpy_{\text{mlp}}.
  3. Calculate the total number of trainable parameters (weights + biases) in this MLP sublayer.
Reveal Solution
  1. Parameter Dimensions (dout×dind_{\text{out}} \times d_{\text{in}}):

    • Wup∈R16×4W_{\text{up}} \in \mathbf{\mathbb{R}^{16 \times 4}}
    • bup∈R16×1b_{\text{up}} \in \mathbf{\mathbb{R}^{16 \times 1}}
    • Wdown∈R4×16W_{\text{down}} \in \mathbf{\mathbb{R}^{4 \times 16}}
    • bdown∈R4×1b_{\text{down}} \in \mathbf{\mathbb{R}^{4 \times 1}}
  2. Activation Vector Dimensions:

    • zff∈R16×1z_{\text{ff}} \in \mathbf{\mathbb{R}^{16 \times 1}}
    • hff∈R16×1h_{\text{ff}} \in \mathbf{\mathbb{R}^{16 \times 1}}
    • ymlp∈R4×1y_{\text{mlp}} \in \mathbf{\mathbb{R}^{4 \times 1}}
  3. Total Trainable Parameter Count:

    • Weights in WupW_{\text{up}}: 16×4=6416 \times 4 = 64
    • Biases in bupb_{\text{up}}: 1616
    • Weights in WdownW_{\text{down}}: 4×16=644 \times 16 = 64
    • Biases in bdownb_{\text{down}}: 44
    • Total Parameters: 64+16+64+4=148 parameters64 + 16 + 64 + 4 = \mathbf{148 \text{ parameters}} (or 128128 weights without biases).

Problem 5: Token Embedding Linear Projection

A language model has a vocabulary of 3 words: ["invest", "market", "risk"] (∣V∣=3|V| = 3). The model represents each word in a 2D embedding space (dmodel=2d_{\text{model}} = 2) using embedding matrix:

WE=[0.80−0.400.100.200.90−0.70]∈R2×3W_E = \begin{bmatrix} 0.80 & -0.40 & 0.10 \\ 0.20 & 0.90 & -0.70 \end{bmatrix} \in \mathbb{R}^{2 \times 3}

The word "market" corresponds to vocabulary index 2 (the second column) and is represented as one-hot vector:

xonehot=[0.01.00.0]∈R3×1x_{\text{onehot}} = \begin{bmatrix} 0.0 \\ 1.0 \\ 0.0 \end{bmatrix} \in \mathbb{R}^{3 \times 1}
  1. Compute the projected embedding coordinate vector e=WExonehote = W_E x_{\text{onehot}}.
  2. Verify that matrix multiplication by a one-hot vector is algebraically identical to selecting the corresponding column vector of WEW_E.
Reveal Solution
  1. Matrix Multiplication Calculation:

    e=WExonehot=[0.80−0.400.100.200.90−0.70][0.01.00.0]e = W_E x_{\text{onehot}} = \begin{bmatrix} 0.80 & -0.40 & 0.10 \\ 0.20 & 0.90 & -0.70 \end{bmatrix} \begin{bmatrix} 0.0 \\ 1.0 \\ 0.0 \end{bmatrix} e1=(0.80)(0.0)+(−0.40)(1.0)+(0.10)(0.0)=−0.40e_1 = (0.80)(0.0) + (-0.40)(1.0) + (0.10)(0.0) = \mathbf{-0.40} e2=(0.20)(0.0)+(0.90)(1.0)+(−0.70)(0.0)=+0.90e_2 = (0.20)(0.0) + (0.90)(1.0) + (-0.70)(0.0) = \mathbf{+0.90} e=[−0.40+0.90]e = \begin{bmatrix} \mathbf{-0.40} \\ \mathbf{+0.90} \end{bmatrix}
  2. Column Extraction Identity: Column 2 of WEW_E is [−0.400.90]\begin{bmatrix} -0.40 \\ 0.90 \end{bmatrix}. Multiplying WEW_E by the one-hot basis vector e2=[0,1,0]Te_2 = [0, 1, 0]^T acts as a linear selector that extracts column 2 exactly:

    WExonehot=WE,∗2=[−0.400.90]W_E x_{\text{onehot}} = W_{E, *2} = \begin{bmatrix} -0.40 \\ 0.90 \end{bmatrix}

Part 2: Applied Scenario: The VC Decision Circuit & Latent Representation Scenario

A venture capital firm uses a miniaturized Transformer MLP sublayer to process early-stage startup decision embeddings.

[ Startup Feature Vector x ]          [ Expanded Hidden Circuit h_ff ]          [ Updated Decision Embedding y_mlp ]
  (d_model = 2)                        (d_ff = 4 Latent Features)                 (d_model = 2)
  x_1: Team Traction                   h_1: Team Scalability                     y_1: Investment Conviction
  x_2: Market Expansion                h_2: High-Burn Risk                       y_2: Follow-On Likelihood
                                       h_3: Capital Efficiency
                                       h_4: Pure Market Velocity

System Configuration:

  • Input Decision Vector (x∈R2×1x \in \mathbb{R}^{2 \times 1}):

    x=[2.01.0](Team Traction Measurement)(Market Expansion Measurement)x = \begin{bmatrix} 2.0 \\ 1.0 \end{bmatrix} \begin{matrix} \text{(Team Traction Measurement)} \\ \text{(Market Expansion Measurement)} \end{matrix}
  • Up-Projection Parameters (Wup∈R4×2,bup∈R4×1W_{\text{up}} \in \mathbb{R}^{4 \times 2}, b_{\text{up}} \in \mathbb{R}^{4 \times 1}):

    Wup=[1.00.5−0.51.50.5−1.00.02.0],bup=[−0.5−1.00.5−1.5]W_{\text{up}} = \begin{bmatrix} 1.0 & 0.5 \\ -0.5 & 1.5 \\ 0.5 & -1.0 \\ 0.0 & 2.0 \end{bmatrix}, \qquad b_{\text{up}} = \begin{bmatrix} -0.5 \\ -1.0 \\ 0.5 \\ -1.5 \end{bmatrix}
  • Activation Function: ReLU(z)=max⁡(0,z)\text{ReLU}(z) = \max(0, z).

  • Down-Projection Parameters (Wdown∈R2×4,bdown∈R2×1W_{\text{down}} \in \mathbb{R}^{2 \times 4}, b_{\text{down}} \in \mathbb{R}^{2 \times 1}):

    Wdown=[0.51.0−0.50.51.0−0.50.50.0],bdown=[0.00.5]W_{\text{down}} = \begin{bmatrix} 0.5 & 1.0 & -0.5 & 0.5 \\ 1.0 & -0.5 & 0.5 & 0.0 \end{bmatrix}, \qquad b_{\text{down}} = \begin{bmatrix} 0.0 \\ 0.5 \end{bmatrix}

Problem 6: The VC Decision Circuit & Latent Representation Scenario

Execute the step-by-step forward trace and mechanistic ablation on paper:

  • Part A (Up-Projection & Pre-Activations): Compute the expanded pre-activation vector zff=Wupx+bup∈R4×1z_{\text{ff}} = W_{\text{up}} x + b_{\text{up}} \in \mathbb{R}^{4 \times 1}.
  • Part B (Non-Linear Latent Activations): Compute intermediate activation vector hff=ReLU(zff)h_{\text{ff}} = \text{ReLU}(z_{\text{ff}}).
  • Part C (Down-Projection to Updated Embedding): Compute final updated embedding vector ymlp=Wdownhff+bdown∈R2×1y_{\text{mlp}} = W_{\text{down}} h_{\text{ff}} + b_{\text{down}} \in \mathbb{R}^{2 \times 1}.
  • Part D (Circuit Ablation Experiment): Ablate latent neuron 4 (hff,4=0.0h_{\text{ff}, 4} = 0.0), compute ablated output ymlp,ablatedy_{\text{mlp},\text{ablated}}, and calculate the exact causal shift vector Δymlp=ymlp,ablated−ymlp\Delta y_{\text{mlp}} = y_{\text{mlp},\text{ablated}} - y_{\text{mlp}}.
Reveal Solution

Part A: Up-Projection & Pre-Activations (zffz_{\text{ff}})

zff=Wupx+bup=[1.00.5−0.51.50.5−1.00.02.0][2.01.0]+[−0.5−1.00.5−1.5]z_{\text{ff}} = W_{\text{up}} x + b_{\text{up}} = \begin{bmatrix} 1.0 & 0.5 \\ -0.5 & 1.5 \\ 0.5 & -1.0 \\ 0.0 & 2.0 \end{bmatrix} \begin{bmatrix} 2.0 \\ 1.0 \end{bmatrix} + \begin{bmatrix} -0.5 \\ -1.0 \\ 0.5 \\ -1.5 \end{bmatrix}

Compute each coordinate:

  • zff,1=(1.0)(2.0)+(0.5)(1.0)−0.5=2.0+0.5−0.5=+2.00z_{\text{ff}, 1} = (1.0)(2.0) + (0.5)(1.0) - 0.5 = 2.0 + 0.5 - 0.5 = \mathbf{+2.00}
  • zff,2=(−0.5)(2.0)+(1.5)(1.0)−1.0=−1.0+1.5−1.0=−0.50z_{\text{ff}, 2} = (-0.5)(2.0) + (1.5)(1.0) - 1.0 = -1.0 + 1.5 - 1.0 = \mathbf{-0.50}
  • zff,3=(0.5)(2.0)+(−1.0)(1.0)+0.5=1.0−1.0+0.5=+0.50z_{\text{ff}, 3} = (0.5)(2.0) + (-1.0)(1.0) + 0.5 = 1.0 - 1.0 + 0.5 = \mathbf{+0.50}
  • zff,4=(0.0)(2.0)+(2.0)(1.0)−1.5=0.0+2.0−1.5=+0.50z_{\text{ff}, 4} = (0.0)(2.0) + (2.0)(1.0) - 1.5 = 0.0 + 2.0 - 1.5 = \mathbf{+0.50}
zff=[+2.00−0.50+0.50+0.50]z_{\text{ff}} = \begin{bmatrix} +2.00 \\ -0.50 \\ +0.50 \\ +0.50 \end{bmatrix}

Part B: Non-Linear Latent Activations (hffh_{\text{ff}})

Apply ReLU(z)=max⁡(0,z)\text{ReLU}(z) = \max(0, z) element-wise:

hff=[max⁡(0,2.00)max⁡(0,−0.50)max⁡(0,0.50)max⁡(0,0.50)]=[2.000.000.500.50]h_{\text{ff}} = \begin{bmatrix} \max(0, 2.00) \\ \max(0, -0.50) \\ \max(0, 0.50) \\ \max(0, 0.50) \end{bmatrix} = \begin{bmatrix} \mathbf{2.00} \\ \mathbf{0.00} \\ \mathbf{0.50} \\ \mathbf{0.50} \end{bmatrix}

(Note: Neuron 2 receives negative pre-activation −0.50-0.50, so ReLU gates it off to 0.000.00).


Part C: Down-Projection to Updated Embedding (ymlpy_{\text{mlp}})

ymlp=Wdownhff+bdown=[0.51.0−0.50.51.0−0.50.50.0][2.000.000.500.50]+[0.000.50]y_{\text{mlp}} = W_{\text{down}} h_{\text{ff}} + b_{\text{down}} = \begin{bmatrix} 0.5 & 1.0 & -0.5 & 0.5 \\ 1.0 & -0.5 & 0.5 & 0.0 \end{bmatrix} \begin{bmatrix} 2.00 \\ 0.00 \\ 0.50 \\ 0.50 \end{bmatrix} + \begin{bmatrix} 0.00 \\ 0.50 \end{bmatrix}

Compute each output dimension:

  • Dimension 1 (Investment Conviction):

    ymlp,1=(0.5)(2.00)+(1.0)(0.00)+(−0.5)(0.50)+(0.5)(0.50)+0.00y_{\text{mlp}, 1} = (0.5)(2.00) + (1.0)(0.00) + (-0.5)(0.50) + (0.5)(0.50) + 0.00 ymlp,1=1.00+0.00−0.25+0.25+0.00=1.00y_{\text{mlp}, 1} = 1.00 + 0.00 - 0.25 + 0.25 + 0.00 = \mathbf{1.00}
  • Dimension 2 (Follow-On Likelihood):

    ymlp,2=(1.0)(2.00)+(−0.5)(0.00)+(0.5)(0.50)+(0.0)(0.50)+0.50y_{\text{mlp}, 2} = (1.0)(2.00) + (-0.5)(0.00) + (0.5)(0.50) + (0.0)(0.50) + 0.50 ymlp,2=2.00+0.00+0.25+0.00+0.50=2.75y_{\text{mlp}, 2} = 2.00 + 0.00 + 0.25 + 0.00 + 0.50 = \mathbf{2.75}
ymlp=[1.002.75]y_{\text{mlp}} = \begin{bmatrix} \mathbf{1.00} \\ \mathbf{2.75} \end{bmatrix}

Part D: Circuit Ablation Experiment

We ablate latent neuron 4 (the dedicated Market Expansion detector where Wup,4=[0.0,2.0]W_{\text{up}, 4} = [0.0, 2.0]) by forcing hff,4=0.00h_{\text{ff}, 4} = 0.00:

hff,ablated=[2.000.000.500.00]h_{\text{ff},\text{ablated}} = \begin{bmatrix} 2.00 \\ 0.00 \\ 0.50 \\ \mathbf{0.00} \end{bmatrix}

Recompute down-projection:

  • Dimension 1 (Ablated):

    ymlp,ablated,1=(0.5)(2.00)+(1.0)(0.00)+(−0.5)(0.50)+(0.5)(0.00)+0.00y_{\text{mlp},\text{ablated}, 1} = (0.5)(2.00) + (1.0)(0.00) + (-0.5)(0.50) + (0.5)(\mathbf{0.00}) + 0.00 ymlp,ablated,1=1.00+0.00−0.25+0.00+0.00=0.75y_{\text{mlp},\text{ablated}, 1} = 1.00 + 0.00 - 0.25 + 0.00 + 0.00 = \mathbf{0.75}
  • Dimension 2 (Ablated):

    ymlp,ablated,2=(1.0)(2.00)+(−0.5)(0.00)+(0.5)(0.50)+(0.0)(0.00)+0.50y_{\text{mlp},\text{ablated}, 2} = (1.0)(2.00) + (-0.5)(0.00) + (0.5)(0.50) + (0.0)(\mathbf{0.00}) + 0.50 ymlp,ablated,2=2.00+0.00+0.25+0.00+0.50=2.75y_{\text{mlp},\text{ablated}, 2} = 2.00 + 0.00 + 0.25 + 0.00 + 0.50 = \mathbf{2.75}
ymlp,ablated=[0.752.75]y_{\text{mlp},\text{ablated}} = \begin{bmatrix} \mathbf{0.75} \\ \mathbf{2.75} \end{bmatrix}

Causal Shift Vector Δymlp\Delta y_{\text{mlp}}:

Δymlp=ymlp,ablated−ymlp=[0.75−1.002.75−2.75]=[−0.250.00]\Delta y_{\text{mlp}} = y_{\text{mlp},\text{ablated}} - y_{\text{mlp}} = \begin{bmatrix} 0.75 - 1.00 \\ 2.75 - 2.75 \end{bmatrix} = \begin{bmatrix} \mathbf{-0.25} \\ \mathbf{0.00} \end{bmatrix}

Mechanistic Finding: Neuron 4 is a specialized sub-circuit that contributes exactly +0.25+0.25 to Investment Conviction (y1y_1) through weight Wdown,14=0.5W_{\text{down}, 14} = 0.5, while having zero direct causal connection (Wdown,24=0.0W_{\text{down}, 24} = 0.0) to Follow-On Likelihood (y2y_2).