Interpretable Weights vs. Latent Geometry hero
Lesson 2Interpretability of Neural Networks

Interpretable Weights vs. Latent Geometry

Contrast human-interpretable feature weights against high-dimensional latent coordinate geometry inside multi-layer representations.

Why do neural networks organize their hidden layers into abstract coordinate systems instead of neat, human-named categories?

The answer lies in geometric transformation and linear separability.


The Geometric Goal of Hidden Layers

A single linear layer (z=w⋅x+bz = w \cdot x + b) can only draw a straight line (in 2D), a flat plane (in 3D), or a flat hyperplane (in nn-dimensional space) to separate classes.

If the raw input data cannot be separated by a flat boundary (such as the classic XOR problem or concentric clusters), a single linear neuron fails.

[ INPUT SPACE: NON-LINEARLY SEPARABLE ]
  x_2
   ▲
   │    Class A (o)         Class B (x)
   │        (0, 1)             (1, 1)
   │          o                  x
   │
   │          x                  o
   │        (0, 0)             (1, 0)
   │    Class B (x)         Class A (o)
   └───────────────────────────────────► x_1

No single straight line w_1 x_1 + w_2 x_2 + b = 0 can separate (o) from (x).

When input vector xx passes through a hidden layer with weights W(1)W^{(1)}, biases b(1)b^{(1)}, and non-linear activations f(z(1))f(z^{(1)}):

a(1)=f(W(1)x+b(1))a^{(1)} = f(W^{(1)} x + b^{(1)})

The hidden layer performs a geometric coordinate transformation:

  1. Affine Transformation (W(1)x+b(1)W^{(1)} x + b^{(1)}): Rotates, stretches, shears, and shifts the input coordinate space.
  2. Non-Linear Activation (ff): Bends, folds, and clips the space along coordinate axes (e.g., ReLU zeroes out negative regions, while Sigmoid squashes space into a bounded unit hypercube).

By stretching and folding the space, the hidden layer rearranges the data points into a new coordinate space—called the Latent Space Rd1\mathbb{R}^{d_1}—where the target classes become linearly separable.

[ LATENT SPACE a^(1) = ReLU(W^(1)x + b^(1)): LINEARLY SEPARABLE ]
  a_2^(1)
   ▲
   │        o  (p_1)
   │
   │             o  (p_2)
   │  - - - - - - - - - - - - - - - - - - - ◄── Hyperplane: w_sep · a^(1) + b_sep = 0
   │        x  (p_3)
   │
   │             x  (p_4)
   └───────────────────────────────────────► a_1^(1)

In latent space, a simple linear boundary cleanly separates Class A (o) from Class B (x).

The output layer (z(2)=W(2)a(1)+b(2)z^{(2)} = W^{(2)} a^{(1)} + b^{(2)}) then performs a single, simple dot product in this transformed latent space to draw the separating hyperplane.


Concrete Numerical Transformation Example

Let us trace a concrete 2D→2D2\text{D} \to 2\text{D} coordinate transformation on paper.

Consider 4 input points in input space R2\mathbb{R}^2:

  • x(A)=[0.00.0]x^{(A)} = \begin{bmatrix} 0.0 \\ 0.0 \end{bmatrix} (Class 0)
  • x(B)=[1.01.0]x^{(B)} = \begin{bmatrix} 1.0 \\ 1.0 \end{bmatrix} (Class 0)
  • x(C)=[1.00.0]x^{(C)} = \begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix} (Class 1)
  • x(D)=[0.01.0]x^{(D)} = \begin{bmatrix} 0.0 \\ 1.0 \end{bmatrix} (Class 1)

These 4 points form the classic XOR configuration: points with matching coordinates belong to Class 0; points with mismatched coordinates belong to Class 1. No single straight line w1x1+w2x2+b=0w_1 x_1 + w_2 x_2 + b = 0 can separate them.

Now pass these points through a hidden layer with parameters:

W(1)=[1.01.01.01.0],b(1)=[0.0−1.0],f(z)=ReLU(z)W^{(1)} = \begin{bmatrix} 1.0 & 1.0 \\ 1.0 & 1.0 \end{bmatrix}, \qquad b^{(1)} = \begin{bmatrix} 0.0 \\ -1.0 \end{bmatrix}, \qquad f(z) = \text{ReLU}(z)

Let us compute the latent coordinate a(1)=ReLU(W(1)x+b(1))a^{(1)} = \text{ReLU}(W^{(1)} x + b^{(1)}) for each point:

Point A (x(A)=[0.0,0.0]Tx^{(A)} = [0.0, 0.0]^T):

z(1)=[1.0(0.0)+1.0(0.0)+0.01.0(0.0)+1.0(0.0)−1.0]=[0.0−1.0]  ⟹  a(1)=[max⁡(0,0.0)max⁡(0,−1.0)]=[0.00.0]z^{(1)} = \begin{bmatrix} 1.0(0.0) + 1.0(0.0) + 0.0 \\ 1.0(0.0) + 1.0(0.0) - 1.0 \end{bmatrix} = \begin{bmatrix} 0.0 \\ -1.0 \end{bmatrix} \implies a^{(1)} = \begin{bmatrix} \max(0, 0.0) \\ \max(0, -1.0) \end{bmatrix} = \begin{bmatrix} \mathbf{0.0} \\ \mathbf{0.0} \end{bmatrix}

Point B (x(B)=[1.0,1.0]Tx^{(B)} = [1.0, 1.0]^T):

z(1)=[1.0(1.0)+1.0(1.0)+0.01.0(1.0)+1.0(1.0)−1.0]=[2.01.0]  ⟹  a(1)=[max⁡(0,2.0)max⁡(0,1.0)]=[2.01.0]z^{(1)} = \begin{bmatrix} 1.0(1.0) + 1.0(1.0) + 0.0 \\ 1.0(1.0) + 1.0(1.0) - 1.0 \end{bmatrix} = \begin{bmatrix} 2.0 \\ 1.0 \end{bmatrix} \implies a^{(1)} = \begin{bmatrix} \max(0, 2.0) \\ \max(0, 1.0) \end{bmatrix} = \begin{bmatrix} \mathbf{2.0} \\ \mathbf{1.0} \end{bmatrix}

Point C (x(C)=[1.0,0.0]Tx^{(C)} = [1.0, 0.0]^T):

z(1)=[1.0(1.0)+1.0(0.0)+0.01.0(1.0)+1.0(0.0)−1.0]=[1.00.0]  ⟹  a(1)=[max⁡(0,1.0)max⁡(0,0.0)]=[1.00.0]z^{(1)} = \begin{bmatrix} 1.0(1.0) + 1.0(0.0) + 0.0 \\ 1.0(1.0) + 1.0(0.0) - 1.0 \end{bmatrix} = \begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix} \implies a^{(1)} = \begin{bmatrix} \max(0, 1.0) \\ \max(0, 0.0) \end{bmatrix} = \begin{bmatrix} \mathbf{1.0} \\ \mathbf{0.0} \end{bmatrix}

Point D (x(D)=[0.0,1.0]Tx^{(D)} = [0.0, 1.0]^T):

z(1)=[1.0(0.0)+1.0(1.0)+0.01.0(0.0)+1.0(1.0)−1.0]=[1.00.0]  ⟹  a(1)=[max⁡(0,1.0)max⁡(0,0.0)]=[1.00.0]z^{(1)} = \begin{bmatrix} 1.0(0.0) + 1.0(1.0) + 0.0 \\ 1.0(0.0) + 1.0(1.0) - 1.0 \end{bmatrix} = \begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix} \implies a^{(1)} = \begin{bmatrix} \max(0, 1.0) \\ \max(0, 0.0) \end{bmatrix} = \begin{bmatrix} \mathbf{1.0} \\ \mathbf{0.0} \end{bmatrix}

Look at what the hidden transformation accomplished:

  • Points CC and DD (both Class 1) are mapped to the exact same latent coordinate: [1.00.0]\begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix}.
  • Point AA (Class 0) is at [0.00.0]\begin{bmatrix} 0.0 \\ 0.0 \end{bmatrix}.
  • Point BB (Class 0) is at [2.01.0]\begin{bmatrix} 2.0 \\ 1.0 \end{bmatrix}.

Now an output linear layer with weights W(2)=[1.0−2.0]W^{(2)} = \begin{bmatrix} 1.0 & -2.0 \end{bmatrix} and bias b(2)=−0.5b^{(2)} = -0.5 evaluates each latent coordinate:

  • zA(2)=1.0(0.0)−2.0(0.0)−0.5=−0.5<0  ⟹  Class 0z^{(2)}_A = 1.0(0.0) - 2.0(0.0) - 0.5 = \mathbf{-0.5} < 0 \implies \text{Class } 0
  • zB(2)=1.0(2.0)−2.0(1.0)−0.5=2.0−2.0−0.5=−0.5<0  ⟹  Class 0z^{(2)}_B = 1.0(2.0) - 2.0(1.0) - 0.5 = 2.0 - 2.0 - 0.5 = \mathbf{-0.5} < 0 \implies \text{Class } 0
  • zC(2)=1.0(1.0)−2.0(0.0)−0.5=1.0−0.0−0.5=+0.5>0  ⟹  Class 1z^{(2)}_C = 1.0(1.0) - 2.0(0.0) - 0.5 = 1.0 - 0.0 - 0.5 = \mathbf{+0.5} > 0 \implies \text{Class } 1
  • zD(2)=1.0(1.0)−2.0(0.0)−0.5=1.0−0.0−0.5=+0.5>0  ⟹  Class 1z^{(2)}_D = 1.0(1.0) - 2.0(0.0) - 0.5 = 1.0 - 0.0 - 0.5 = \mathbf{+0.5} > 0 \implies \text{Class } 1

The non-linearly separable problem in input space R2\mathbb{R}^2 became 100% linearly separable in latent space R2\mathbb{R}^2.


The Limits of Raw Weight Inspection

Why can we not simply look at a trained weight matrix and read off what it has learned?

  1. Distributed Representations: Concepts are rarely stored in a single neuron. A concept (such as "Sci-Fi Thriller" or "High-Risk Seed Stage Startup") is represented as a distributed pattern of activation across dozens or hundreds of hidden neurons simultaneously.
  2. Polysemantic Neurons: An individual neuron often responds to multiple unrelated concepts (for example, a single neuron in a vision network might activate for both cat faces and car wheels because both share curved line geometries).
  3. Rotational Invariance: In high-dimensional vector spaces, the coordinate axes themselves are arbitrary. Rotating the entire latent activation space while applying the inverse rotation to the downstream weight matrix preserves the exact output predictions while completely scrambling the individual coordinates.

Inspecting raw matrix values in a table cannot reveal how the network computes. To understand neural networks, we must use rigorous empirical tools that measure the causal impact of internal activations.


Previous
The True Nature of the "Black Box"