Master multi-layer forward passes, hidden representations, composite activations, and dimension tracking through complete manual calculations.
Part 1: The Underlying Mechanics Drill
Work through each problem on paper before revealing the step-by-step solution.
Problem 1: Chained Matrix Evaluation & Linear Collapse
Given an input vector x=[23], a first-layer weight matrix W(1)=[23−11], and a second-layer weight matrix W(2)=[1−2] (with zero biases and no activation functions):
Part A: Calculate the intermediate vector z(1)=W(1)x.
Part B: Calculate the scalar output z(2)=W(2)z(1).
Part C: Calculate the composite product matrix Wcomposite=W(2)W(1), and compute Wcompositex directly. Explain what this reveals about multi-layer linear networks without activation functions.
Pedagogical Insight: Both calculations yield the exact same scalar (−17). Without non-linear activation functions between layers, chaining multiple weight matrices mathematically collapses into a single matrix multiplication (Wcomposite=W(2)W(1)). No matter how many linear layers are stacked, a linear network can never compute non-linear feature interactions.
Problem 2: Hidden Activation Vector Calculation with ReLU
A 5-neuron hidden layer receives an input vector and produces the following pre-activation linear sum vector z(1):
z(1)=3.5−2.40.05.1−0.8
Calculate the activated hidden representation vector a(1)=ReLU(z(1)), and state which neuron channels are active vs. dormant.
Reveal Solution
Apply the piecewise definition ReLU(zi)=max(0,zi) to each coordinate independently:
Channel Status: Neurons 1 and 4 are actively firing and transmitting signals to the downstream layer. Neurons 2, 3, and 5 are completely silenced to 0.0, generating a sparse representation where only positive evidence propagates forward.
Problem 3: Multi-Layer Dimension Compatibility & Parameter Counting
A deep learning engineer designs a 3-layer neural network with an architecture of 3→4→2→1:
Input dimension n0=3
Hidden Layer 1 width n1=4
Hidden Layer 2 width n2=2
Output Layer width n3=1
Part A: State the required matrix dimensions for weight matrices W(1), W(2), and W(3).
Part B: State the required vector dimensions for bias vectors b(1), b(2), and b(3).
Part C: State the vector dimensions of the activation states x, a(1), a(2), and a(3).
Part D: Calculate the total number of learnable parameters (weights + biases) across the entire network.
Reveal Solution
Part A: Weight Matrix Dimensions (W(l)∈Rnl×nl−1)
An output layer receives a 2-dimensional hidden activation vector a(1)=[2.01.5]. The output layer is parameterized by weight matrix W(2)=[3.0−2.0] and baseline bias b(2)=[−1.0].
Part A: Calculate the pre-activation scalar z(2)=W(2)a(1)+b(2).
Part B: Calculate the Sigmoid probability a(2)=σ(z(2)) (use e−2.0≈0.1353).
Part C: Explain how the output layer treats the hidden activations a(1) as if they were raw input features.
Part C: Downstream Invariance
To the output neuron, the vector a(1)=[2.0,1.5]T is structurally indistinguishable from an original input vector x. The output layer has no awareness of how a(1) was created, how many input features originally entered Layer 1, or whether Layer 1 used ReLU or Sigmoid. It simply executes a standard dot product on the numbers presented to it.
Problem 5: Zero Hidden State / Inactive Neuron Propagation
Suppose an input sample produces strong negative pre-activations across all hidden neurons (z(1)≤0), causing the entire hidden activation vector to be clamped to zero: a(1)=[0.00.0].
The output layer has weights W(2)=[4.05.0] and baseline bias b(2)=[−1.5].
Part A: Calculate the output pre-activation z(2)andthefinalSigmoidprobabilitya^{(2)} = \sigma(z^{(2)})(usee^{1.5} \approx 4.4817$).
Part B: Explain why the network's final output becomes completely unresponsive to the input features x or Layer 1 weights W(1) under this condition.
Part B: Mathematical Decoupling
Because a(1)=0, the dot product W(2)a(1) evaluates to exactly zero regardless of the values in W(2). The output pre-activation collapses strictly to the output bias: z(2)=0+b(2)=b(2).
When all hidden units are dormant, no information from the input vector x reaches the output layer. The network outputs a fixed default baseline probability (σ(b(2))=18.2%) determined solely by its environmental bias.
Part 2: Applied Scenario: The VC 2-Layer Representation Network
In Topics 1, 2, and 3, our venture capital firm evaluated startup pitch profiles across four locked criteria [Team Experience, Market Size, Competition, Risk]:
The investment committee evaluates the two hidden factors: it rewards Scalability (+1.0) and places high premium on Defensibility (+1.5), against a strict macro hurdle bias b(2)=−3.0:
W(2)=[1.01.5],b(2)=[−3.0],f(2)=Sigmoid
To visually trace how both startup pitch vectors propagate through intermediate latent representations into the final term sheet conversion probability, examine the comparative dataflow below:
Comparative 2-layer forward pass for OmniFlow and Solaris AI: OmniFlow's heavy competition penalty silences its Defensibility factor to 0.0 via ReLU, leaving it with insufficient signal to clear the Layer 2 hurdle ($26.9\%$ Pass). Solaris AI's uncontested market generates both Scalability ($2.8$) and Defensibility ($1.2$), securing an $83.2\%$ Term Sheet Greenlight.
Problem 6: Complete 2-Layer Startup Forward Pass
Part A: Compute the Layer 1 pre-activation vector z(1)=W(1)x+b(1) and the ReLU hidden activation vector a(1)=ReLU(z(1)) for both OmniFlow and Solaris AI.
Part B: Compute the Layer 2 scalar pre-activation z(2)=W(2)a(1)+b(2) and the final Sigmoid probability a(2)=σ(z(2)) for both startups (use e1.0≈2.7183 and e−1.6≈0.2019).
Part C: State which startup earns the term sheet greenlight (y^≥50%).
Part D: Provide an investment committee synthesis explaining how representation learning in Layer 1 exposed OmniFlow's critical vulnerability compared to Solaris AI.
Solaris AI achieves an 83.2% Term Sheet Greenlight (y^≥50%).
OmniFlow is rejected with a 26.9% Probability (y^<50%).
Part D: Strategic Representation Synthesis
The 2-layer architecture reveals why representation learning is critical for robust decisions:
OmniFlow's Hidden Vulnerability: In Layer 1, OmniFlow's veteran team (x1=1.0) generated a solid Scalability score (a1(1)=2.0). However, because OmniFlow operates in an intensely crowded market (x3=1.0), its competitive Defensibility score went negative (z2(1)=−0.3). ReLU clamped this channel strictly to 0.0. When Layer 2 evaluated OmniFlow, it had only one active factor (2.0), which was insufficient to overcome the strict −3.0 committee hurdle (z(2)=−1.0⟹26.9%).
Solaris AI's Dual-Factor Strength: Solaris AI entered with high technical execution risk (x4=0.8) and a modest team score (x1=0.3). However, its massive market (x2=1.0) and total absence of competition (x3=0.0) allowed both hidden neurons to fire strongly (a1(1)=2.8,a2(1)=1.2). When Layer 2 multiplied these intermediate concepts by their committee weights (1.0(2.8)+1.5(1.2)=4.6), it easily absorbed the −3.0 hurdle, securing a decisive 83.2% greenlight.
By utilizing hidden layers, the network avoids making simplistic decisions on raw features, structuring intermediate insights to produce nuanced, defensible predictions.