Representation Learning Foundations hero
Lesson 3Multi-Layer Neural Networks

Representation Learning Foundations

Trace how successive hidden layers automatically transform low-level input features into high-level abstract representations inside the network.

The true breakthrough of deep learning is not that multi-layer networks calculate longer chains of arithmetic. The breakthrough is Representation Learning.

When it comes to AI, representation learning is the process by which intermediate layers transform raw, low-level coordinate values into higher-level, abstract features where decision-making becomes linearly separable.


Track 1 (Underlying Mechanism): Multi-Stage Manufacturing Assembly

Consider how an automobile manufacturing plant builds a finished vehicle:

Raw Materials (Steel, Glass, Rubber)
       │
       ▼  [ Stage 1 Assembly: Sub-Components ]
Sub-Assemblies (Engine Block, Transmission, Steering Column)
       │
       ▼  [ Stage 2 Assembly: Final Synthesis ]
Finished Vehicle (Automobile)
  1. Stage 0 (Raw Ingredients): The factory starts with unrefined materials: sheets of steel, blocks of vulcanized rubber, copper wiring, and raw glass panels.
  2. Stage 1 (Sub-assembly / Hidden Representations): Workers do not bolt sheets of steel directly onto the car. First, they forge the steel into an engine block, assemble the rubber into tires, and wire the copper into an instrument cluster.
  3. Stage 2 (Final Integration): The final assembly team does not worry about individual screws or raw chemical formulas of rubber. They simply connect the pre-built sub-assemblies together to produce the finished vehicle.

A neural network operates on the exact same hierarchical principle:

  • Raw input features are the raw materials.
  • Hidden neurons build intermediate sub-assemblies (compound representations).
  • The output neuron connects the sub-assemblies into the final decision.

Track 2 (Applied Concept): Hierarchical Movie Sub-Genres

Let's look at how our 2-layer network analyzes movie screenplays.

Our input vector xx provides the raw feature values for four basic genres:

x=[1.00.00.51.0](Action)(Romance)(Comedy)(Sci-Fi)x = \begin{bmatrix} 1.0 \\ 0.0 \\ 0.5 \\ 1.0 \end{bmatrix} \begin{matrix} \text{(Action)} \\ \text{(Romance)} \\ \text{(Comedy)} \\ \text{(Sci-Fi)} \end{matrix}

Instead of forcing the network to jump directly from raw genres to a final Hit prediction, Layer 1 synthesizes two intermediate cinematic concepts:

  1. Hidden Neuron 1: The "Summer Popcorn Flick" Factor (a1(1)a_1^{(1)})

    • Learns positive weights for Action (+3.0+3.0) and Sci-Fi (+3.0+3.0).
    • Learns a negative weight for Romance (−2.0-2.0).
    • Imposes a baseline hurdle bias (b1(1)=−2.0b_1^{(1)} = -2.0).
    • Mechanical role: Activates strongly if and only if high Action and high Sci-Fi coincide.
  2. Hidden Neuron 2: The "Rom-Com" Factor (a2(1)a_2^{(1)})

    • Learns positive weights for Romance (+3.0+3.0) and Comedy (+2.0+2.0).
    • Learns negative weights for Action (−2.0-2.0) and Sci-Fi (−1.0-1.0).
    • Imposes a baseline hurdle bias (b2(1)=−1.5b_2^{(1)} = -1.5).
    • Mechanical role: Activates strongly if and only if high Romance and high Comedy coincide.

We assemble these parameters into the Layer 1 Weight Matrix (W(1)∈R2×4W^{(1)} \in \mathbb{R}^{2 \times 4}) and Bias Vector (b(1)∈R2×1b^{(1)} \in \mathbb{R}^{2 \times 1}):

W(1)=[3.0−2.00.03.0−2.03.02.0−1.0],b(1)=[−2.0−1.5]W^{(1)} = \begin{bmatrix} 3.0 & -2.0 & 0.0 & 3.0 \\ -2.0 & 3.0 & 2.0 & -1.0 \end{bmatrix}, \qquad b^{(1)} = \begin{bmatrix} -2.0 \\ -1.5 \end{bmatrix}

The Downstream Perspective: Consuming Hidden Activations

Now, examine the forward pass from the perspective of Layer 2 (The Output Layer).

The output neuron does not look at the raw 4-dimensional genre vector [Action, Romance, Comedy, Sci-Fi]. That raw data has already been digested by Layer 1.

The output neuron receives only the 2-dimensional hidden activation vector a(1)∈R2a^{(1)} \in \mathbb{R}^2:

a^{(1)} = \begin{bmatrix} a_1^{(1)} \\ a_2^{(1)} \end{bmatrix} = \begin{bmatrix} \text{Summer Popcorn Flick Activation} \\ \text{Rom-Com Activation} \end{matrix}

To Layer 2, a(1)a^{(1)} is a brand-new, refined feature vector.

The studio executive committee in Layer 2 simply sets influence weights for these two intermediate concepts:

  • Popcorn Flick Weight (W11(2)=+1.5W_{11}^{(2)} = +1.5): Popcorn blockbusters have high box office ceilings.
  • Rom-Com Weight (W12(2)=+1.0W_{12}^{(2)} = +1.0): Rom-Coms have solid, reliable commercial returns.
  • Studio Baseline Bias (b(2)=−2.0b^{(2)} = -2.0): A general market hurdle.
W(2)=[1.51.0],b(2)=[−2.0]W^{(2)} = \begin{bmatrix} 1.5 & 1.0 \end{bmatrix}, \qquad b^{(2)} = \begin{bmatrix} -2.0 \end{bmatrix}

By separating the task into two hierarchical stages, the network has turned an intractable non-linear problem into two straightforward steps:

  1. Layer 1 converts raw input values into abstract concepts.
  2. Layer 2 makes a linear decision based on those abstract concepts.

Previous
Hidden Layers and Network Geometry