
Hidden Layers and Network Geometry
Insert intermediate hidden layers between inputs and outputs to create hierarchical geometric transformations across multi-layer perceptrons.
By stacking layers sequentially, we construct a Multi-Layer Perceptron (MLP).
When it comes to AI, a Multi-Layer Perceptron is a feedforward neural network consisting of at least three distinct stages:
- An Input Layer that receives raw feature values.
- One or more Hidden Layers that transform incoming signals into intermediate representations.
- An Output Layer that computes the final prediction.
Architectural Anatomy of an MLP
Let's inspect the structural components of a 2-layer Multi-Layer Perceptron designed for our movie prediction task:
[ Input Layer: x ] [ Hidden Layer 1: a^(1) ] [ Output Layer: a^(2) ]
(4 Raw Genre Features) (2 Intermediate Concepts) (1 Final Hit Decision)
(Action) x1 ───────────► ( h1: Popcorn ) ────────────┐
(Romance) x2 ───────────► ├───► ( y_hat: Greenlight )
(Comedy) x3 ───────────► ( h2: Rom-Com ) ────────────┘
(Sci-Fi) x4 ───────────►
- Input Layer (): Holds the raw feature values
[Action, Romance, Comedy, Sci-Fi]. The input layer contains no weights, no biases, and performs no arithmetic; it is simply the entry vector of data. - Hidden Layer (): A layer of parallel neurons. Each hidden neuron receives all input features, computes an affine sum (), and squashes the result through an activation function ().
- Output Layer (): A single neuron that receives the hidden activations (), computes a second affine sum (), and applies a final Sigmoid activation to produce the predicted Hit probability ().
Standard Layer Superscript Notation
To track parameters across multiple layers without mathematical confusion, we use parenthetical superscripts to designate the layer index:
- : The Weight Matrix of Layer .
- : The Bias Vector of Layer .
- : The Pre-activation Linear Sum Vector of Layer .
- : The Post-activation Output Vector of Layer .
For our 2-layer network ():
Let's check the matrix dimensions at each layer:
- Layer 1 Weight Matrix (): rows (hidden neurons), columns (input features).
- Layer 1 Bias Vector (): baseline hurdle offsets.
- Layer 2 Weight Matrix (): row (output neuron), columns (hidden activation inputs).
- Layer 2 Bias Vector (): final output hurdle offset.
Notice the fundamental rule of multi-layer compatibility: the number of columns in must exactly equal the number of neurons in Layer 1 ().
Why Are They Designated as "Hidden"?
The intermediate layer is called "hidden" for a simple reason: it does not interact directly with the external environment.
- The Input Layer is visible because the human engineer supplies the raw data ().
- The Output Layer is visible because the user reads the final predicted score ().
- The Hidden Layer is internal to the model. Its pre-activations () and post-activations () exist strictly in the computer's memory registers during computation.
The network is never explicitly told what the hidden neurons must compute. During training, the optimization algorithm adjusts and automatically, allowing the hidden layer to discover internal representations that best help the output layer make accurate predictions.
To visualize how the 4-dimensional input features connect to the 2 hidden intermediate neurons and converge into the single output neuron, trace the structural architecture below:
