Work through every problem on paper before expanding the step-by-step solutions.
Part 1: Abstract Mechanics
Work through these 5 drill problems testing latent transformations, linear separability, feature ablation, Transformer MLP dimensions, and token projection math.
Problem 1: Latent Space Coordinate Transformation Mapping
A 2-input, 2-neuron hidden layer has the following configuration:
Input vector: x = [ 2.0 − 1.0 ] x = \begin{bmatrix} 2.0 \\ -1.0 \end{bmatrix} x = [ 2.0 − 1.0 ]
Weight matrix: W ( 1 ) = [ 1.5 − 1.0 − 0.5 2.0 ] W^{(1)} = \begin{bmatrix} 1.5 & -1.0 \\ -0.5 & 2.0 \end{bmatrix} W ( 1 ) = [ 1.5 − 0.5 − 1.0 2.0 ]
Bias vector: b ( 1 ) = [ − 1.0 1.0 ] b^{(1)} = \begin{bmatrix} -1.0 \\ 1.0 \end{bmatrix} b ( 1 ) = [ − 1.0 1.0 ]
Activation function: ReLU ( z ) = max ( 0 , z ) \text{ReLU}(z) = \max(0, z) ReLU ( z ) = max ( 0 , z )
Compute:
The pre-activation vector z ( 1 ) = W ( 1 ) x + b ( 1 ) z^{(1)} = W^{(1)} x + b^{(1)} z ( 1 ) = W ( 1 ) x + b ( 1 ) .
The hidden latent activation vector a ( 1 ) = ReLU ( z ( 1 ) ) a^{(1)} = \text{ReLU}(z^{(1)}) a ( 1 ) = ReLU ( z ( 1 ) ) .
Reveal Solution
Pre-Activation Vector z ( 1 ) z^{(1)} z ( 1 ) :
z 1 ( 1 ) = ( 1.5 ) ( 2.0 ) + ( − 1.0 ) ( − 1.0 ) − 1.0 = 3.0 + 1.0 − 1.0 = + 3.0 z_1^{(1)} = (1.5)(2.0) + (-1.0)(-1.0) - 1.0 = 3.0 + 1.0 - 1.0 = \mathbf{+3.0} z 1 ( 1 ) = ( 1.5 ) ( 2.0 ) + ( − 1.0 ) ( − 1.0 ) − 1.0 = 3.0 + 1.0 − 1.0 = + 3.0
z 2 ( 1 ) = ( − 0.5 ) ( 2.0 ) + ( 2.0 ) ( − 1.0 ) + 1.0 = − 1.0 − 2.0 + 1.0 = − 2.0 z_2^{(1)} = (-0.5)(2.0) + (2.0)(-1.0) + 1.0 = -1.0 - 2.0 + 1.0 = \mathbf{-2.0} z 2 ( 1 ) = ( − 0.5 ) ( 2.0 ) + ( 2.0 ) ( − 1.0 ) + 1.0 = − 1.0 − 2.0 + 1.0 = − 2.0
z ( 1 ) = [ + 3.0 − 2.0 ] z^{(1)} = \begin{bmatrix} +3.0 \\ -2.0 \end{bmatrix} z ( 1 ) = [ + 3.0 − 2.0 ]
Hidden Activation Vector a ( 1 ) a^{(1)} a ( 1 ) :
a ( 1 ) = ReLU ( [ + 3.0 − 2.0 ] ) = [ max ( 0 , 3.0 ) max ( 0 , − 2.0 ) ] = [ 3.0 0.0 ] a^{(1)} = \text{ReLU}\left(\begin{bmatrix} +3.0 \\ -2.0 \end{bmatrix}\right) = \begin{bmatrix} \max(0, 3.0) \\ \max(0, -2.0) \end{bmatrix} = \begin{bmatrix} \mathbf{3.0} \\ \mathbf{0.0} \end{bmatrix} a ( 1 ) = ReLU ( [ + 3.0 − 2.0 ] ) = [ max ( 0 , 3.0 ) max ( 0 , − 2.0 ) ] = [ 3.0 0.0 ]
Problem 2: Linear Separability Classification Check
Consider 4 points in a 2D latent representation space R 2 \mathbb{R}^2 R 2 :
Class A (y = + 1 y = +1 y = + 1 ): p 1 = [ 3.0 1.0 ] , p 2 = [ 2.0 3.0 ] p_1 = \begin{bmatrix} 3.0 \\ 1.0 \end{bmatrix}, \quad p_2 = \begin{bmatrix} 2.0 \\ 3.0 \end{bmatrix} p 1 = [ 3.0 1.0 ] , p 2 = [ 2.0 3.0 ]
Class B (y = − 1 y = -1 y = − 1 ): p 3 = [ 0.0 1.0 ] , p 4 = [ 1.0 0.0 ] p_3 = \begin{bmatrix} 0.0 \\ 1.0 \end{bmatrix}, \quad p_4 = \begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix} p 3 = [ 0.0 1.0 ] , p 4 = [ 1.0 0.0 ]
A candidate linear separating hyperplane has parameter vector w sep = [ 1.0 1.0 ] w_{\text{sep}} = \begin{bmatrix} 1.0 \\ 1.0 \end{bmatrix} w sep = [ 1.0 1.0 ] and threshold bias b sep = − 3.0 b_{\text{sep}} = -3.0 b sep = − 3.0 , defining the decision function:
s ( p ) = w sep ⋅ p + b sep s(p) = w_{\text{sep}} \cdot p + b_{\text{sep}} s ( p ) = w sep ⋅ p + b sep
Where a point is classified as Class A if s ( p ) > 0 s(p) > 0 s ( p ) > 0 , and Class B if s ( p ) < 0 s(p) < 0 s ( p ) < 0 .
Compute the decision score s ( p i ) s(p_i) s ( p i ) for each of the 4 points and verify whether this hyperplane cleanly separates Class A from Class B.
Reveal Solution
Score for Point p 1 = [ 3.0 , 1.0 ] T p_1 = [3.0, 1.0]^T p 1 = [ 3.0 , 1.0 ] T (Class A):
s ( p 1 ) = ( 1.0 ) ( 3.0 ) + ( 1.0 ) ( 1.0 ) − 3.0 = 3.0 + 1.0 − 3.0 = + 1.0 > 0 ⟹ Class A (Correct) s(p_1) = (1.0)(3.0) + (1.0)(1.0) - 3.0 = 3.0 + 1.0 - 3.0 = \mathbf{+1.0} > 0 \implies \mathbf{\text{Class A (Correct)}} s ( p 1 ) = ( 1.0 ) ( 3.0 ) + ( 1.0 ) ( 1.0 ) − 3.0 = 3.0 + 1.0 − 3.0 = + 1.0 > 0 ⟹ Class A (Correct)
Score for Point p 2 = [ 2.0 , 3.0 ] T p_2 = [2.0, 3.0]^T p 2 = [ 2.0 , 3.0 ] T (Class A):
s ( p 2 ) = ( 1.0 ) ( 2.0 ) + ( 1.0 ) ( 3.0 ) − 3.0 = 2.0 + 3.0 − 3.0 = + 2.0 > 0 ⟹ Class A (Correct) s(p_2) = (1.0)(2.0) + (1.0)(3.0) - 3.0 = 2.0 + 3.0 - 3.0 = \mathbf{+2.0} > 0 \implies \mathbf{\text{Class A (Correct)}} s ( p 2 ) = ( 1.0 ) ( 2.0 ) + ( 1.0 ) ( 3.0 ) − 3.0 = 2.0 + 3.0 − 3.0 = + 2.0 > 0 ⟹ Class A (Correct)
Score for Point p 3 = [ 0.0 , 1.0 ] T p_3 = [0.0, 1.0]^T p 3 = [ 0.0 , 1.0 ] T (Class B):
s ( p 3 ) = ( 1.0 ) ( 0.0 ) + ( 1.0 ) ( 1.0 ) − 3.0 = 0.0 + 1.0 − 3.0 = − 2.0 < 0 ⟹ Class B (Correct) s(p_3) = (1.0)(0.0) + (1.0)(1.0) - 3.0 = 0.0 + 1.0 - 3.0 = \mathbf{-2.0} < 0 \implies \mathbf{\text{Class B (Correct)}} s ( p 3 ) = ( 1.0 ) ( 0.0 ) + ( 1.0 ) ( 1.0 ) − 3.0 = 0.0 + 1.0 − 3.0 = − 2.0 < 0 ⟹ Class B (Correct)
Score for Point p 4 = [ 1.0 , 0.0 ] T p_4 = [1.0, 0.0]^T p 4 = [ 1.0 , 0.0 ] T (Class B):
s ( p 4 ) = ( 1.0 ) ( 1.0 ) + ( 1.0 ) ( 0.0 ) − 3.0 = 1.0 + 0.0 − 3.0 = − 2.0 < 0 ⟹ Class B (Correct) s(p_4) = (1.0)(1.0) + (1.0)(0.0) - 3.0 = 1.0 + 0.0 - 3.0 = \mathbf{-2.0} < 0 \implies \mathbf{\text{Class B (Correct)}} s ( p 4 ) = ( 1.0 ) ( 1.0 ) + ( 1.0 ) ( 0.0 ) − 3.0 = 1.0 + 0.0 − 3.0 = − 2.0 < 0 ⟹ Class B (Correct)
Conclusion: Since s ( p 1 ) > 0 , s ( p 2 ) > 0 s(p_1) > 0, s(p_2) > 0 s ( p 1 ) > 0 , s ( p 2 ) > 0 and s ( p 3 ) < 0 , s ( p 4 ) < 0 s(p_3) < 0, s(p_4) < 0 s ( p 3 ) < 0 , s ( p 4 ) < 0 , the 4 points are strictly linearly separable by the hyperplane w sep ⋅ p + b sep = 0 w_{\text{sep}} \cdot p + b_{\text{sep}} = 0 w sep ⋅ p + b sep = 0 .
Problem 3: Hidden Neuron Feature Ablation Calculation
A 2-layer network with 2 hidden neurons and a linear output neuron has parameters:
Hidden activation vector: a ( 1 ) = [ 2.0 4.0 ] a^{(1)} = \begin{bmatrix} 2.0 \\ 4.0 \end{bmatrix} a ( 1 ) = [ 2.0 4.0 ]
Output weight matrix: W ( 2 ) = [ 1.5 − 0.5 ] W^{(2)} = \begin{bmatrix} 1.5 & -0.5 \end{bmatrix} W ( 2 ) = [ 1.5 − 0.5 ]
Output bias: b ( 2 ) = 0.20 b^{(2)} = 0.20 b ( 2 ) = 0.20
Output prediction: y ^ = W ( 2 ) a ( 1 ) + b ( 2 ) \hat{y} = W^{(2)} a^{(1)} + b^{(2)} y ^ = W ( 2 ) a ( 1 ) + b ( 2 )
Perform an ablation experiment:
Compute the intact output prediction y ^ intact \hat{y}_{\text{intact}} y ^ intact .
Ablate hidden neuron 2 (set a 2 ( 1 ) = 0.0 a_2^{(1)} = 0.0 a 2 ( 1 ) = 0.0 ) and compute the ablated output prediction y ^ ablated \hat{y}_{\text{ablated}} y ^ ablated .
Compute the exact causal output shift Δ y ^ = y ^ ablated − y ^ intact \Delta \hat{y} = \hat{y}_{\text{ablated}} - \hat{y}_{\text{intact}} Δ y ^ = y ^ ablated − y ^ intact .
Reveal Solution
Intact Output Prediction:
y ^ intact = W ( 2 ) a ( 1 ) + b ( 2 ) = ( 1.5 ) ( 2.0 ) + ( − 0.5 ) ( 4.0 ) + 0.20 = 3.0 − 2.0 + 0.20 = 1.20 \hat{y}_{\text{intact}} = W^{(2)} a^{(1)} + b^{(2)} = (1.5)(2.0) + (-0.5)(4.0) + 0.20 = 3.0 - 2.0 + 0.20 = \mathbf{1.20} y ^ intact = W ( 2 ) a ( 1 ) + b ( 2 ) = ( 1.5 ) ( 2.0 ) + ( − 0.5 ) ( 4.0 ) + 0.20 = 3.0 − 2.0 + 0.20 = 1.20
Ablated Output Prediction (a 2 ( 1 ) = 0.0 a_2^{(1)} = 0.0 a 2 ( 1 ) = 0.0 ):
a ablated ( 1 ) = [ 2.0 0.0 ] a^{(1)}_{\text{ablated}} = \begin{bmatrix} 2.0 \\ 0.0 \end{bmatrix} a ablated ( 1 ) = [ 2.0 0.0 ]
y ^ ablated = ( 1.5 ) ( 2.0 ) + ( − 0.5 ) ( 0.0 ) + 0.20 = 3.0 + 0.0 + 0.20 = 3.20 \hat{y}_{\text{ablated}} = (1.5)(2.0) + (-0.5)(0.0) + 0.20 = 3.0 + 0.0 + 0.20 = \mathbf{3.20} y ^ ablated = ( 1.5 ) ( 2.0 ) + ( − 0.5 ) ( 0.0 ) + 0.20 = 3.0 + 0.0 + 0.20 = 3.20
Causal Output Shift Δ y ^ \Delta \hat{y} Δ y ^ :
Δ y ^ = y ^ ablated − y ^ intact = 3.20 − 1.20 = + 2.00 \Delta \hat{y} = \hat{y}_{\text{ablated}} - \hat{y}_{\text{intact}} = 3.20 - 1.20 = \mathbf{+2.00} Δ y ^ = y ^ ablated − y ^ intact = 3.20 − 1.20 = + 2.00
(Interpretation: Hidden neuron 2 was exerting a strong inhibitory effect of − 0.5 × 4.0 = − 2.00 -0.5 \times 4.0 = -2.00 − 0.5 × 4.0 = − 2.00 on the output. Ablating it removes that inhibition, increasing the final prediction by + 2.00 +2.00 + 2.00 ).
Problem 4: Transformer MLP Sublayer Dimension Trace
A miniaturized Transformer model has token embedding dimension d model = 4 d_{\text{model}} = 4 d model = 4 . Its Feed-Forward MLP sublayer expands the token vector by a factor of 4 to intermediate dimension d ff = 4 d model = 16 d_{\text{ff}} = 4 d_{\text{model}} = 16 d ff = 4 d model = 16 , then projects it back to d model = 4 d_{\text{model}} = 4 d model = 4 .
The sublayer equations are:
z ff = W up x + b up , h ff = ReLU ( z ff ) , y mlp = W down h ff + b down z_{\text{ff}} = W_{\text{up}} x + b_{\text{up}}, \qquad h_{\text{ff}} = \text{ReLU}(z_{\text{ff}}), \qquad y_{\text{mlp}} = W_{\text{down}} h_{\text{ff}} + b_{\text{down}} z ff = W up x + b up , h ff = ReLU ( z ff ) , y mlp = W down h ff + b down
For an input token vector x ∈ R 4 × 1 x \in \mathbb{R}^{4 \times 1} x ∈ R 4 × 1 :
State the exact matrix dimensions of W up , b up , W down W_{\text{up}}, b_{\text{up}}, W_{\text{down}} W up , b up , W down , and b down b_{\text{down}} b down .
State the vector dimensions of z ff , h ff z_{\text{ff}}, h_{\text{ff}} z ff , h ff , and y mlp y_{\text{mlp}} y mlp .
Calculate the total number of trainable parameters (weights + biases) in this MLP sublayer.
Reveal Solution
Parameter Dimensions (d out × d in d_{\text{out}} \times d_{\text{in}} d out × d in ):
W up ∈ R 16 × 4 W_{\text{up}} \in \mathbf{\mathbb{R}^{16 \times 4}} W up ∈ R 16 × 4
b up ∈ R 16 × 1 b_{\text{up}} \in \mathbf{\mathbb{R}^{16 \times 1}} b up ∈ R 16 × 1
W down ∈ R 4 × 16 W_{\text{down}} \in \mathbf{\mathbb{R}^{4 \times 16}} W down ∈ R 4 × 16
b down ∈ R 4 × 1 b_{\text{down}} \in \mathbf{\mathbb{R}^{4 \times 1}} b down ∈ R 4 × 1
Activation Vector Dimensions:
z ff ∈ R 16 × 1 z_{\text{ff}} \in \mathbf{\mathbb{R}^{16 \times 1}} z ff ∈ R 16 × 1
h ff ∈ R 16 × 1 h_{\text{ff}} \in \mathbf{\mathbb{R}^{16 \times 1}} h ff ∈ R 16 × 1
y mlp ∈ R 4 × 1 y_{\text{mlp}} \in \mathbf{\mathbb{R}^{4 \times 1}} y mlp ∈ R 4 × 1
Total Trainable Parameter Count:
Weights in W up W_{\text{up}} W up : 16 × 4 = 64 16 \times 4 = 64 16 × 4 = 64
Biases in b up b_{\text{up}} b up : 16 16 16
Weights in W down W_{\text{down}} W down : 4 × 16 = 64 4 \times 16 = 64 4 × 16 = 64
Biases in b down b_{\text{down}} b down : 4 4 4
Total Parameters: 64 + 16 + 64 + 4 = 148 parameters 64 + 16 + 64 + 4 = \mathbf{148 \text{ parameters}} 64 + 16 + 64 + 4 = 148 parameters (or 128 128 128 weights without biases).
Problem 5: Token Embedding Linear Projection
A language model has a vocabulary of 3 words: ["invest", "market", "risk"] (∣ V ∣ = 3 |V| = 3 ∣ V ∣ = 3 ). The model represents each word in a 2D embedding space (d model = 2 d_{\text{model}} = 2 d model = 2 ) using embedding matrix:
W E = [ 0.80 − 0.40 0.10 0.20 0.90 − 0.70 ] ∈ R 2 × 3 W_E = \begin{bmatrix} 0.80 & -0.40 & 0.10 \\ 0.20 & 0.90 & -0.70 \end{bmatrix} \in \mathbb{R}^{2 \times 3} W E = [ 0.80 0.20 − 0.40 0.90 0.10 − 0.70 ] ∈ R 2 × 3
The word "market" corresponds to vocabulary index 2 (the second column) and is represented as one-hot vector:
x onehot = [ 0.0 1.0 0.0 ] ∈ R 3 × 1 x_{\text{onehot}} = \begin{bmatrix} 0.0 \\ 1.0 \\ 0.0 \end{bmatrix} \in \mathbb{R}^{3 \times 1} x onehot = 0.0 1.0 0.0 ∈ R 3 × 1
Compute the projected embedding coordinate vector e = W E x onehot e = W_E x_{\text{onehot}} e = W E x onehot .
Verify that matrix multiplication by a one-hot vector is algebraically identical to selecting the corresponding column vector of W E W_E W E .
Reveal Solution
Matrix Multiplication Calculation:
e = W E x onehot = [ 0.80 − 0.40 0.10 0.20 0.90 − 0.70 ] [ 0.0 1.0 0.0 ] e = W_E x_{\text{onehot}} = \begin{bmatrix} 0.80 & -0.40 & 0.10 \\ 0.20 & 0.90 & -0.70 \end{bmatrix} \begin{bmatrix} 0.0 \\ 1.0 \\ 0.0 \end{bmatrix} e = W E x onehot = [ 0.80 0.20 − 0.40 0.90 0.10 − 0.70 ] 0.0 1.0 0.0
e 1 = ( 0.80 ) ( 0.0 ) + ( − 0.40 ) ( 1.0 ) + ( 0.10 ) ( 0.0 ) = − 0.40 e_1 = (0.80)(0.0) + (-0.40)(1.0) + (0.10)(0.0) = \mathbf{-0.40} e 1 = ( 0.80 ) ( 0.0 ) + ( − 0.40 ) ( 1.0 ) + ( 0.10 ) ( 0.0 ) = − 0.40
e 2 = ( 0.20 ) ( 0.0 ) + ( 0.90 ) ( 1.0 ) + ( − 0.70 ) ( 0.0 ) = + 0.90 e_2 = (0.20)(0.0) + (0.90)(1.0) + (-0.70)(0.0) = \mathbf{+0.90} e 2 = ( 0.20 ) ( 0.0 ) + ( 0.90 ) ( 1.0 ) + ( − 0.70 ) ( 0.0 ) = + 0.90
e = [ − 0.40 + 0.90 ] e = \begin{bmatrix} \mathbf{-0.40} \\ \mathbf{+0.90} \end{bmatrix} e = [ − 0.40 + 0.90 ]
Column Extraction Identity:
Column 2 of W E W_E W E is [ − 0.40 0.90 ] \begin{bmatrix} -0.40 \\ 0.90 \end{bmatrix} [ − 0.40 0.90 ] . Multiplying W E W_E W E by the one-hot basis vector e 2 = [ 0 , 1 , 0 ] T e_2 = [0, 1, 0]^T e 2 = [ 0 , 1 , 0 ] T acts as a linear selector that extracts column 2 exactly:
W E x onehot = W E , ∗ 2 = [ − 0.40 0.90 ] W_E x_{\text{onehot}} = W_{E, *2} = \begin{bmatrix} -0.40 \\ 0.90 \end{bmatrix} W E x onehot = W E , ∗ 2 = [ − 0.40 0.90 ]
Part 2: Applied Scenario: The VC Decision Circuit & Latent Representation Scenario
A venture capital firm uses a miniaturized Transformer MLP sublayer to process early-stage startup decision embeddings.
[ Startup Feature Vector x ] [ Expanded Hidden Circuit h_ff ] [ Updated Decision Embedding y_mlp ]
(d_model = 2) (d_ff = 4 Latent Features) (d_model = 2)
x_1: Team Traction h_1: Team Scalability y_1: Investment Conviction
x_2: Market Expansion h_2: High-Burn Risk y_2: Follow-On Likelihood
h_3: Capital Efficiency
h_4: Pure Market Velocity
System Configuration:
Input Decision Vector (x ∈ R 2 × 1 x \in \mathbb{R}^{2 \times 1} x ∈ R 2 × 1 ):
x = [ 2.0 1.0 ] (Team Traction Measurement) (Market Expansion Measurement) x = \begin{bmatrix} 2.0 \\ 1.0 \end{bmatrix} \begin{matrix} \text{(Team Traction Measurement)} \\ \text{(Market Expansion Measurement)} \end{matrix} x = [ 2.0 1.0 ] (Team Traction Measurement) (Market Expansion Measurement)
Up-Projection Parameters (W up ∈ R 4 × 2 , b up ∈ R 4 × 1 W_{\text{up}} \in \mathbb{R}^{4 \times 2}, b_{\text{up}} \in \mathbb{R}^{4 \times 1} W up ∈ R 4 × 2 , b up ∈ R 4 × 1 ):
W up = [ 1.0 0.5 − 0.5 1.5 0.5 − 1.0 0.0 2.0 ] , b up = [ − 0.5 − 1.0 0.5 − 1.5 ] W_{\text{up}} = \begin{bmatrix} 1.0 & 0.5 \\ -0.5 & 1.5 \\ 0.5 & -1.0 \\ 0.0 & 2.0 \end{bmatrix}, \qquad b_{\text{up}} = \begin{bmatrix} -0.5 \\ -1.0 \\ 0.5 \\ -1.5 \end{bmatrix} W up = 1.0 − 0.5 0.5 0.0 0.5 1.5 − 1.0 2.0 , b up = − 0.5 − 1.0 0.5 − 1.5
Activation Function: ReLU ( z ) = max ( 0 , z ) \text{ReLU}(z) = \max(0, z) ReLU ( z ) = max ( 0 , z ) .
Down-Projection Parameters (W down ∈ R 2 × 4 , b down ∈ R 2 × 1 W_{\text{down}} \in \mathbb{R}^{2 \times 4}, b_{\text{down}} \in \mathbb{R}^{2 \times 1} W down ∈ R 2 × 4 , b down ∈ R 2 × 1 ):
W down = [ 0.5 1.0 − 0.5 0.5 1.0 − 0.5 0.5 0.0 ] , b down = [ 0.0 0.5 ] W_{\text{down}} = \begin{bmatrix} 0.5 & 1.0 & -0.5 & 0.5 \\ 1.0 & -0.5 & 0.5 & 0.0 \end{bmatrix}, \qquad b_{\text{down}} = \begin{bmatrix} 0.0 \\ 0.5 \end{bmatrix} W down = [ 0.5 1.0 1.0 − 0.5 − 0.5 0.5 0.5 0.0 ] , b down = [ 0.0 0.5 ]
Problem 6: The VC Decision Circuit & Latent Representation Scenario
Execute the step-by-step forward trace and mechanistic ablation on paper:
Part A (Up-Projection & Pre-Activations): Compute the expanded pre-activation vector z ff = W up x + b up ∈ R 4 × 1 z_{\text{ff}} = W_{\text{up}} x + b_{\text{up}} \in \mathbb{R}^{4 \times 1} z ff = W up x + b up ∈ R 4 × 1 .
Part B (Non-Linear Latent Activations): Compute intermediate activation vector h ff = ReLU ( z ff ) h_{\text{ff}} = \text{ReLU}(z_{\text{ff}}) h ff = ReLU ( z ff ) .
Part C (Down-Projection to Updated Embedding): Compute final updated embedding vector y mlp = W down h ff + b down ∈ R 2 × 1 y_{\text{mlp}} = W_{\text{down}} h_{\text{ff}} + b_{\text{down}} \in \mathbb{R}^{2 \times 1} y mlp = W down h ff + b down ∈ R 2 × 1 .
Part D (Circuit Ablation Experiment): Ablate latent neuron 4 (h ff , 4 = 0.0 h_{\text{ff}, 4} = 0.0 h ff , 4 = 0.0 ), compute ablated output y mlp , ablated y_{\text{mlp},\text{ablated}} y mlp , ablated , and calculate the exact causal shift vector Δ y mlp = y mlp , ablated − y mlp \Delta y_{\text{mlp}} = y_{\text{mlp},\text{ablated}} - y_{\text{mlp}} Δ y mlp = y mlp , ablated − y mlp .
Reveal Solution
Part A: Up-Projection & Pre-Activations (z ff z_{\text{ff}} z ff )
z ff = W up x + b up = [ 1.0 0.5 − 0.5 1.5 0.5 − 1.0 0.0 2.0 ] [ 2.0 1.0 ] + [ − 0.5 − 1.0 0.5 − 1.5 ] z_{\text{ff}} = W_{\text{up}} x + b_{\text{up}} = \begin{bmatrix} 1.0 & 0.5 \\ -0.5 & 1.5 \\ 0.5 & -1.0 \\ 0.0 & 2.0 \end{bmatrix} \begin{bmatrix} 2.0 \\ 1.0 \end{bmatrix} + \begin{bmatrix} -0.5 \\ -1.0 \\ 0.5 \\ -1.5 \end{bmatrix} z ff = W up x + b up = 1.0 − 0.5 0.5 0.0 0.5 1.5 − 1.0 2.0 [ 2.0 1.0 ] + − 0.5 − 1.0 0.5 − 1.5
Compute each coordinate:
z ff , 1 = ( 1.0 ) ( 2.0 ) + ( 0.5 ) ( 1.0 ) − 0.5 = 2.0 + 0.5 − 0.5 = + 2.00 z_{\text{ff}, 1} = (1.0)(2.0) + (0.5)(1.0) - 0.5 = 2.0 + 0.5 - 0.5 = \mathbf{+2.00} z ff , 1 = ( 1.0 ) ( 2.0 ) + ( 0.5 ) ( 1.0 ) − 0.5 = 2.0 + 0.5 − 0.5 = + 2.00
z ff , 2 = ( − 0.5 ) ( 2.0 ) + ( 1.5 ) ( 1.0 ) − 1.0 = − 1.0 + 1.5 − 1.0 = − 0.50 z_{\text{ff}, 2} = (-0.5)(2.0) + (1.5)(1.0) - 1.0 = -1.0 + 1.5 - 1.0 = \mathbf{-0.50} z ff , 2 = ( − 0.5 ) ( 2.0 ) + ( 1.5 ) ( 1.0 ) − 1.0 = − 1.0 + 1.5 − 1.0 = − 0.50
z ff , 3 = ( 0.5 ) ( 2.0 ) + ( − 1.0 ) ( 1.0 ) + 0.5 = 1.0 − 1.0 + 0.5 = + 0.50 z_{\text{ff}, 3} = (0.5)(2.0) + (-1.0)(1.0) + 0.5 = 1.0 - 1.0 + 0.5 = \mathbf{+0.50} z ff , 3 = ( 0.5 ) ( 2.0 ) + ( − 1.0 ) ( 1.0 ) + 0.5 = 1.0 − 1.0 + 0.5 = + 0.50
z ff , 4 = ( 0.0 ) ( 2.0 ) + ( 2.0 ) ( 1.0 ) − 1.5 = 0.0 + 2.0 − 1.5 = + 0.50 z_{\text{ff}, 4} = (0.0)(2.0) + (2.0)(1.0) - 1.5 = 0.0 + 2.0 - 1.5 = \mathbf{+0.50} z ff , 4 = ( 0.0 ) ( 2.0 ) + ( 2.0 ) ( 1.0 ) − 1.5 = 0.0 + 2.0 − 1.5 = + 0.50
z ff = [ + 2.00 − 0.50 + 0.50 + 0.50 ] z_{\text{ff}} = \begin{bmatrix} +2.00 \\ -0.50 \\ +0.50 \\ +0.50 \end{bmatrix} z ff = + 2.00 − 0.50 + 0.50 + 0.50
Part B: Non-Linear Latent Activations (h ff h_{\text{ff}} h ff )
Apply ReLU ( z ) = max ( 0 , z ) \text{ReLU}(z) = \max(0, z) ReLU ( z ) = max ( 0 , z ) element-wise:
h ff = [ max ( 0 , 2.00 ) max ( 0 , − 0.50 ) max ( 0 , 0.50 ) max ( 0 , 0.50 ) ] = [ 2.00 0.00 0.50 0.50 ] h_{\text{ff}} = \begin{bmatrix} \max(0, 2.00) \\ \max(0, -0.50) \\ \max(0, 0.50) \\ \max(0, 0.50) \end{bmatrix} = \begin{bmatrix} \mathbf{2.00} \\ \mathbf{0.00} \\ \mathbf{0.50} \\ \mathbf{0.50} \end{bmatrix} h ff = max ( 0 , 2.00 ) max ( 0 , − 0.50 ) max ( 0 , 0.50 ) max ( 0 , 0.50 ) = 2.00 0.00 0.50 0.50
(Note: Neuron 2 receives negative pre-activation − 0.50 -0.50 − 0.50 , so ReLU gates it off to 0.00 0.00 0.00 ).
Part C: Down-Projection to Updated Embedding (y mlp y_{\text{mlp}} y mlp )
y mlp = W down h ff + b down = [ 0.5 1.0 − 0.5 0.5 1.0 − 0.5 0.5 0.0 ] [ 2.00 0.00 0.50 0.50 ] + [ 0.00 0.50 ] y_{\text{mlp}} = W_{\text{down}} h_{\text{ff}} + b_{\text{down}} = \begin{bmatrix} 0.5 & 1.0 & -0.5 & 0.5 \\ 1.0 & -0.5 & 0.5 & 0.0 \end{bmatrix} \begin{bmatrix} 2.00 \\ 0.00 \\ 0.50 \\ 0.50 \end{bmatrix} + \begin{bmatrix} 0.00 \\ 0.50 \end{bmatrix} y mlp = W down h ff + b down = [ 0.5 1.0 1.0 − 0.5 − 0.5 0.5 0.5 0.0 ] 2.00 0.00 0.50 0.50 + [ 0.00 0.50 ]
Compute each output dimension:
y mlp = [ 1.00 2.75 ] y_{\text{mlp}} = \begin{bmatrix} \mathbf{1.00} \\ \mathbf{2.75} \end{bmatrix} y mlp = [ 1.00 2.75 ]
Part D: Circuit Ablation Experiment
We ablate latent neuron 4 (the dedicated Market Expansion detector where W up , 4 = [ 0.0 , 2.0 ] W_{\text{up}, 4} = [0.0, 2.0] W up , 4 = [ 0.0 , 2.0 ] ) by forcing h ff , 4 = 0.00 h_{\text{ff}, 4} = 0.00 h ff , 4 = 0.00 :
h ff , ablated = [ 2.00 0.00 0.50 0.00 ] h_{\text{ff},\text{ablated}} = \begin{bmatrix} 2.00 \\ 0.00 \\ 0.50 \\ \mathbf{0.00} \end{bmatrix} h ff , ablated = 2.00 0.00 0.50 0.00
Recompute down-projection:
Dimension 1 (Ablated):
y mlp , ablated , 1 = ( 0.5 ) ( 2.00 ) + ( 1.0 ) ( 0.00 ) + ( − 0.5 ) ( 0.50 ) + ( 0.5 ) ( 0.00 ) + 0.00 y_{\text{mlp},\text{ablated}, 1} = (0.5)(2.00) + (1.0)(0.00) + (-0.5)(0.50) + (0.5)(\mathbf{0.00}) + 0.00 y mlp , ablated , 1 = ( 0.5 ) ( 2.00 ) + ( 1.0 ) ( 0.00 ) + ( − 0.5 ) ( 0.50 ) + ( 0.5 ) ( 0.00 ) + 0.00
y mlp , ablated , 1 = 1.00 + 0.00 − 0.25 + 0.00 + 0.00 = 0.75 y_{\text{mlp},\text{ablated}, 1} = 1.00 + 0.00 - 0.25 + 0.00 + 0.00 = \mathbf{0.75} y mlp , ablated , 1 = 1.00 + 0.00 − 0.25 + 0.00 + 0.00 = 0.75
Dimension 2 (Ablated):
y mlp , ablated , 2 = ( 1.0 ) ( 2.00 ) + ( − 0.5 ) ( 0.00 ) + ( 0.5 ) ( 0.50 ) + ( 0.0 ) ( 0.00 ) + 0.50 y_{\text{mlp},\text{ablated}, 2} = (1.0)(2.00) + (-0.5)(0.00) + (0.5)(0.50) + (0.0)(\mathbf{0.00}) + 0.50 y mlp , ablated , 2 = ( 1.0 ) ( 2.00 ) + ( − 0.5 ) ( 0.00 ) + ( 0.5 ) ( 0.50 ) + ( 0.0 ) ( 0.00 ) + 0.50
y mlp , ablated , 2 = 2.00 + 0.00 + 0.25 + 0.00 + 0.50 = 2.75 y_{\text{mlp},\text{ablated}, 2} = 2.00 + 0.00 + 0.25 + 0.00 + 0.50 = \mathbf{2.75} y mlp , ablated , 2 = 2.00 + 0.00 + 0.25 + 0.00 + 0.50 = 2.75
y mlp , ablated = [ 0.75 2.75 ] y_{\text{mlp},\text{ablated}} = \begin{bmatrix} \mathbf{0.75} \\ \mathbf{2.75} \end{bmatrix} y mlp , ablated = [ 0.75 2.75 ]
Causal Shift Vector Δ y mlp \Delta y_{\text{mlp}} Δ y mlp :
Δ y mlp = y mlp , ablated − y mlp = [ 0.75 − 1.00 2.75 − 2.75 ] = [ − 0.25 0.00 ] \Delta y_{\text{mlp}} = y_{\text{mlp},\text{ablated}} - y_{\text{mlp}} = \begin{bmatrix} 0.75 - 1.00 \\ 2.75 - 2.75 \end{bmatrix} = \begin{bmatrix} \mathbf{-0.25} \\ \mathbf{0.00} \end{bmatrix} Δ y mlp = y mlp , ablated − y mlp = [ 0.75 − 1.00 2.75 − 2.75 ] = [ − 0.25 0.00 ]
Mechanistic Finding: Neuron 4 is a specialized sub-circuit that contributes exactly + 0.25 +0.25 + 0.25 to Investment Conviction (y 1 y_1 y 1 ) through weight W down , 14 = 0.5 W_{\text{down}, 14} = 0.5 W down , 14 = 0.5 , while having zero direct causal connection (W down , 24 = 0.0 W_{\text{down}, 24} = 0.0 W down , 24 = 0.0 ) to Follow-On Likelihood (y 2 y_2 y 2 ).