Hidden Layers and Network Geometry hero
Lesson 2Multi-Layer Neural Networks

Hidden Layers and Network Geometry

Insert intermediate hidden layers between inputs and outputs to create hierarchical geometric transformations across multi-layer perceptrons.

By stacking layers sequentially, we construct a Multi-Layer Perceptron (MLP).

When it comes to AI, a Multi-Layer Perceptron is a feedforward neural network consisting of at least three distinct stages:

  1. An Input Layer that receives raw feature values.
  2. One or more Hidden Layers that transform incoming signals into intermediate representations.
  3. An Output Layer that computes the final prediction.

Architectural Anatomy of an MLP

Let's inspect the structural components of a 2-layer Multi-Layer Perceptron designed for our movie prediction task:

[ Input Layer: x ]          [ Hidden Layer 1: a^(1) ]          [ Output Layer: a^(2) ]
(4 Raw Genre Features)       (2 Intermediate Concepts)           (1 Final Hit Decision)

      (Action)    x1 ───────────► ( h1: Popcorn ) ────────────┐
      (Romance)   x2 ───────────►                         ├───► ( y_hat: Greenlight )
      (Comedy)    x3 ───────────► ( h2: Rom-Com ) ────────────┘
      (Sci-Fi)    x4 ───────────►
  1. Input Layer (x∈R4x \in \mathbb{R}^4): Holds the raw feature values [Action, Romance, Comedy, Sci-Fi]. The input layer contains no weights, no biases, and performs no arithmetic; it is simply the entry vector of data.
  2. Hidden Layer (a(1)∈R2a^{(1)} \in \mathbb{R}^2): A layer of 22 parallel neurons. Each hidden neuron receives all 44 input features, computes an affine sum (z(1)=W(1)x+b(1)z^{(1)} = W^{(1)}x + b^{(1)}), and squashes the result through an activation function (a(1)=f(z(1))a^{(1)} = f(z^{(1)})).
  3. Output Layer (a(2)∈R1a^{(2)} \in \mathbb{R}^1): A single neuron that receives the 22 hidden activations (a(1)a^{(1)}), computes a second affine sum (z(2)=W(2)a(1)+b(2)z^{(2)} = W^{(2)}a^{(1)} + b^{(2)}), and applies a final Sigmoid activation to produce the predicted Hit probability (y^=a(2)\hat{y} = a^{(2)}).

Standard Layer Superscript Notation

To track parameters across multiple layers without mathematical confusion, we use parenthetical superscripts (l)(l) to designate the layer index:

  • W(l)W^{(l)}: The Weight Matrix of Layer ll.
  • b(l)b^{(l)}: The Bias Vector of Layer ll.
  • z(l)z^{(l)}: The Pre-activation Linear Sum Vector of Layer ll.
  • a(l)a^{(l)}: The Post-activation Output Vector of Layer ll.

For our 2-layer network (4→2→14 \to 2 \to 1):

Layer 1 (Hidden Layer):z(1)=W(1)x+b(1),a(1)=ReLU(z(1))\mathbf{Layer\ 1\ (Hidden\ Layer):} \qquad z^{(1)} = W^{(1)} x + b^{(1)}, \qquad a^{(1)} = \text{ReLU}\left(z^{(1)}\right)

Layer 2 (Output Layer):z(2)=W(2)a(1)+b(2),a(2)=σ(z(2))\mathbf{Layer\ 2\ (Output\ Layer):} \qquad z^{(2)} = W^{(2)} a^{(1)} + b^{(2)}, \qquad a^{(2)} = \sigma\left(z^{(2)}\right)

Let's check the matrix dimensions at each layer:

  • Layer 1 Weight Matrix (W(1)∈R2×4W^{(1)} \in \mathbb{R}^{2 \times 4}): 22 rows (hidden neurons), 44 columns (input features).
  • Layer 1 Bias Vector (b(1)∈R2×1b^{(1)} \in \mathbb{R}^{2 \times 1}): 22 baseline hurdle offsets.
  • Layer 2 Weight Matrix (W(2)∈R1×2W^{(2)} \in \mathbb{R}^{1 \times 2}): 11 row (output neuron), 22 columns (hidden activation inputs).
  • Layer 2 Bias Vector (b(2)∈R1×1b^{(2)} \in \mathbb{R}^{1 \times 1}): 11 final output hurdle offset.

Notice the fundamental rule of multi-layer compatibility: the number of columns in W(2)W^{(2)} must exactly equal the number of neurons in Layer 1 (a(1)a^{(1)}).


Why Are They Designated as "Hidden"?

The intermediate layer is called "hidden" for a simple reason: it does not interact directly with the external environment.

  • The Input Layer is visible because the human engineer supplies the raw data (xx).
  • The Output Layer is visible because the user reads the final predicted score (y^\hat{y}).
  • The Hidden Layer is internal to the model. Its pre-activations (z(1)z^{(1)}) and post-activations (a(1)a^{(1)}) exist strictly in the computer's memory registers during computation.

The network is never explicitly told what the hidden neurons must compute. During training, the optimization algorithm adjusts W(1)W^{(1)} and W(2)W^{(2)} automatically, allowing the hidden layer to discover internal representations that best help the output layer make accurate predictions.

To visualize how the 4-dimensional input features connect to the 2 hidden intermediate neurons and converge into the single output neuron, trace the structural architecture below:

Multi-Layer Perceptron Geometry and Layer Connectivity
Architectural layout of a 2-layer Multi-Layer Perceptron ($4 \to 2 \to 1$): The 4 input feature nodes connect via weight matrix $W^{(1)}$ and bias $b^{(1)}$ to 2 intermediate hidden neurons ($h_1$ and $h_2$). The hidden activations propagate through second-layer weights $W^{(2)}$ and output bias $b^{(2)}$ to produce the final scalar prediction $a^{(2)}$.

Previous
Limitations of Single-Layer Networks