Multi-Layer Networks In Practice hero
LaboratoryMulti-Layer Neural Networks

Multi-Layer Networks In Practice

Master multi-layer forward passes, hidden representations, composite activations, and dimension tracking through complete manual calculations.

Part 1: The Underlying Mechanics Drill

Work through each problem on paper before revealing the step-by-step solution.

Problem 1: Chained Matrix Evaluation & Linear Collapse

Given an input vector x=[23]x = \begin{bmatrix} 2 \\ 3 \end{bmatrix}, a first-layer weight matrix W(1)=[2−131]W^{(1)} = \begin{bmatrix} 2 & -1 \\ 3 & 1 \end{bmatrix}, and a second-layer weight matrix W(2)=[1−2]W^{(2)} = \begin{bmatrix} 1 & -2 \end{bmatrix} (with zero biases and no activation functions):

  • Part A: Calculate the intermediate vector z(1)=W(1)xz^{(1)} = W^{(1)} x.
  • Part B: Calculate the scalar output z(2)=W(2)z(1)z^{(2)} = W^{(2)} z^{(1)}.
  • Part C: Calculate the composite product matrix Wcomposite=W(2)W(1)W_{\text{composite}} = W^{(2)} W^{(1)}, and compute WcompositexW_{\text{composite}} x directly. Explain what this reveals about multi-layer linear networks without activation functions.
Reveal Solution

Part A: Calculate z(1)=W(1)xz^{(1)} = W^{(1)} x

z(1)=[2−131][23]=[(2)(2)+(−1)(3)(3)(2)+(1)(3)]=[4−36+3]=[19]z^{(1)} = \begin{bmatrix} 2 & -1 \\ 3 & 1 \end{bmatrix} \begin{bmatrix} 2 \\ 3 \end{bmatrix} = \begin{bmatrix} (2)(2) + (-1)(3) \\ (3)(2) + (1)(3) \end{bmatrix} = \begin{bmatrix} 4 - 3 \\ 6 + 3 \end{bmatrix} = \begin{bmatrix} 1 \\ 9 \end{bmatrix}

Part B: Calculate z(2)=W(2)z(1)z^{(2)} = W^{(2)} z^{(1)}

z(2)=[1−2][19]=(1)(1)+(−2)(9)=1−18=−17z^{(2)} = \begin{bmatrix} 1 & -2 \end{bmatrix} \begin{bmatrix} 1 \\ 9 \end{bmatrix} = (1)(1) + (-2)(9) = 1 - 18 = \mathbf{-17}

Part C: Composite Matrix Multiplication & Linear Collapse

Wcomposite=W(2)W(1)=[1−2][2−131]W_{\text{composite}} = W^{(2)} W^{(1)} = \begin{bmatrix} 1 & -2 \end{bmatrix} \begin{bmatrix} 2 & -1 \\ 3 & 1 \end{bmatrix} Wcomposite=[(1)(2)+(−2)(3)(1)(−1)+(−2)(1)]=[2−6−1−2]=[−4−3]W_{\text{composite}} = \begin{bmatrix} (1)(2) + (-2)(3) & (1)(-1) + (-2)(1) \end{bmatrix} = \begin{bmatrix} 2 - 6 & -1 - 2 \end{bmatrix} = \begin{bmatrix} -4 & -3 \end{bmatrix}

Now compute WcompositexW_{\text{composite}} x:

Wcompositex=[−4−3][23]=(−4)(2)+(−3)(3)=−8−9=−17W_{\text{composite}} x = \begin{bmatrix} -4 & -3 \end{bmatrix} \begin{bmatrix} 2 \\ 3 \end{bmatrix} = (-4)(2) + (-3)(3) = -8 - 9 = \mathbf{-17}

Pedagogical Insight: Both calculations yield the exact same scalar (−17-17). Without non-linear activation functions between layers, chaining multiple weight matrices mathematically collapses into a single matrix multiplication (Wcomposite=W(2)W(1)W_{\text{composite}} = W^{(2)} W^{(1)}). No matter how many linear layers are stacked, a linear network can never compute non-linear feature interactions.


Problem 2: Hidden Activation Vector Calculation with ReLU

A 5-neuron hidden layer receives an input vector and produces the following pre-activation linear sum vector z(1)z^{(1)}:

z(1)=[3.5−2.40.05.1−0.8]z^{(1)} = \begin{bmatrix} 3.5 \\ -2.4 \\ 0.0 \\ 5.1 \\ -0.8 \end{bmatrix}

Calculate the activated hidden representation vector a(1)=ReLU(z(1))a^{(1)} = \text{ReLU}\left(z^{(1)}\right), and state which neuron channels are active vs. dormant.

Reveal Solution

Apply the piecewise definition ReLU(zi)=max⁡(0,zi)\text{ReLU}(z_i) = \max(0, z_i) to each coordinate independently:

  • a1(1)=max⁡(0,3.5)=3.5a_1^{(1)} = \max(0, 3.5) = \mathbf{3.5} (Active Channel)
  • a2(1)=max⁡(0,−2.4)=0.0a_2^{(1)} = \max(0, -2.4) = \mathbf{0.0} (Dormant / Silenced Channel)
  • a3(1)=max⁡(0,0.0)=0.0a_3^{(1)} = \max(0, 0.0) = \mathbf{0.0} (Dormant / Boundary Channel)
  • a4(1)=max⁡(0,5.1)=5.1a_4^{(1)} = \max(0, 5.1) = \mathbf{5.1} (Active Channel)
  • a5(1)=max⁡(0,−0.8)=0.0a_5^{(1)} = \max(0, -0.8) = \mathbf{0.0} (Dormant / Silenced Channel)

Assemble the resulting activation vector:

a(1)=[3.50.00.05.10.0]a^{(1)} = \begin{bmatrix} 3.5 \\ 0.0 \\ 0.0 \\ 5.1 \\ 0.0 \end{bmatrix}

Channel Status: Neurons 11 and 44 are actively firing and transmitting signals to the downstream layer. Neurons 22, 33, and 55 are completely silenced to 0.00.0, generating a sparse representation where only positive evidence propagates forward.


Problem 3: Multi-Layer Dimension Compatibility & Parameter Counting

A deep learning engineer designs a 3-layer neural network with an architecture of 3→4→2→13 \to 4 \to 2 \to 1:

  • Input dimension n0=3n_0 = 3

  • Hidden Layer 1 width n1=4n_1 = 4

  • Hidden Layer 2 width n2=2n_2 = 2

  • Output Layer width n3=1n_3 = 1

  • Part A: State the required matrix dimensions for weight matrices W(1)W^{(1)}, W(2)W^{(2)}, and W(3)W^{(3)}.

  • Part B: State the required vector dimensions for bias vectors b(1)b^{(1)}, b(2)b^{(2)}, and b(3)b^{(3)}.

  • Part C: State the vector dimensions of the activation states xx, a(1)a^{(1)}, a(2)a^{(2)}, and a(3)a^{(3)}.

  • Part D: Calculate the total number of learnable parameters (weights + biases) across the entire network.

Reveal Solution

Part A: Weight Matrix Dimensions (W(l)∈Rnl×nl−1W^{(l)} \in \mathbb{R}^{n_l \times n_{l-1}})

  • Layer 1: W(1)∈R4×3W^{(1)} \in \mathbb{R}^{4 \times 3} (44 neurons receiving 33 inputs)
  • Layer 2: W(2)∈R2×4W^{(2)} \in \mathbb{R}^{2 \times 4} (22 neurons receiving 44 hidden inputs)
  • Layer 3: W(3)∈R1×2W^{(3)} \in \mathbb{R}^{1 \times 2} (11 neuron receiving 22 hidden inputs)

Part B: Bias Vector Dimensions (b(l)∈Rnl×1b^{(l)} \in \mathbb{R}^{n_l \times 1})

  • Layer 1: b(1)∈R4×1b^{(1)} \in \mathbb{R}^{4 \times 1} (44 baseline hurdles)
  • Layer 2: b(2)∈R2×1b^{(2)} \in \mathbb{R}^{2 \times 1} (22 baseline hurdles)
  • Layer 3: b(3)∈R1×1b^{(3)} \in \mathbb{R}^{1 \times 1} (11 baseline hurdle)

Part C: Activation Vector Dimensions

  • Input Vector: x∈R3×1x \in \mathbb{R}^{3 \times 1}
  • Hidden Layer 1 Activation: a(1)∈R4×1a^{(1)} \in \mathbb{R}^{4 \times 1}
  • Hidden Layer 2 Activation: a(2)∈R2×1a^{(2)} \in \mathbb{R}^{2 \times 1}
  • Output Layer Activation: a(3)∈R1×1a^{(3)} \in \mathbb{R}^{1 \times 1}

Part D: Parameter Count Calculation

  • Layer 1: (4×3 weights)+(4 biases)=12+4=16 parameters(4 \times 3 \text{ weights}) + (4 \text{ biases}) = 12 + 4 = \mathbf{16\text{ parameters}}
  • Layer 2: (2×4 weights)+(2 biases)=8+2=10 parameters(2 \times 4 \text{ weights}) + (2 \text{ biases}) = 8 + 2 = \mathbf{10\text{ parameters}}
  • Layer 3: (1×2 weights)+(1 bias)=2+1=3 parameters(1 \times 2 \text{ weights}) + (1 \text{ bias}) = 2 + 1 = \mathbf{3\text{ parameters}}

Total Parameters=16+10+3=29 learnable parameters\text{Total Parameters} = 16 + 10 + 3 = \mathbf{29\text{ learnable parameters}}


Problem 4: Final Prediction from Hidden State

An output layer receives a 2-dimensional hidden activation vector a(1)=[2.01.5]a^{(1)} = \begin{bmatrix} 2.0 \\ 1.5 \end{bmatrix}. The output layer is parameterized by weight matrix W(2)=[3.0−2.0]W^{(2)} = \begin{bmatrix} 3.0 & -2.0 \end{bmatrix} and baseline bias b(2)=[−1.0]b^{(2)} = \begin{bmatrix} -1.0 \end{bmatrix}.

  • Part A: Calculate the pre-activation scalar z(2)=W(2)a(1)+b(2)z^{(2)} = W^{(2)} a^{(1)} + b^{(2)}.
  • Part B: Calculate the Sigmoid probability a(2)=σ(z(2))a^{(2)} = \sigma(z^{(2)}) (use e−2.0≈0.1353e^{-2.0} \approx 0.1353).
  • Part C: Explain how the output layer treats the hidden activations a(1)a^{(1)} as if they were raw input features.
Reveal Solution

Part A: Calculate z(2)z^{(2)}

z(2)=W(2)a(1)+b(2)=[3.0−2.0][2.01.5]+(−1.0)z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} = \begin{bmatrix} 3.0 & -2.0 \end{bmatrix} \begin{bmatrix} 2.0 \\ 1.5 \end{bmatrix} + (-1.0) z(2)=(3.0×2.0)+(−2.0×1.5)+(−1.0)=6.0−3.0−1.0=+2.0z^{(2)} = (3.0 \times 2.0) + (-2.0 \times 1.5) + (-1.0) = 6.0 - 3.0 - 1.0 = \mathbf{+2.0}

Part B: Calculate a(2)=σ(z(2))a^{(2)} = \sigma(z^{(2)})

a(2)=σ(2.0)=11+e−2.0=11+0.1353=11.1353≈0.8808  ⟹  88.1%a^{(2)} = \sigma(2.0) = \frac{1}{1 + e^{-2.0}} = \frac{1}{1 + 0.1353} = \frac{1}{1.1353} \approx \mathbf{0.8808} \implies \mathbf{88.1\%}

Part C: Downstream Invariance To the output neuron, the vector a(1)=[2.0,1.5]Ta^{(1)} = [2.0, 1.5]^T is structurally indistinguishable from an original input vector xx. The output layer has no awareness of how a(1)a^{(1)} was created, how many input features originally entered Layer 1, or whether Layer 1 used ReLU or Sigmoid. It simply executes a standard dot product on the numbers presented to it.


Problem 5: Zero Hidden State / Inactive Neuron Propagation

Suppose an input sample produces strong negative pre-activations across all hidden neurons (z(1)≤0z^{(1)} \le \mathbf{0}), causing the entire hidden activation vector to be clamped to zero: a(1)=[0.00.0]a^{(1)} = \begin{bmatrix} 0.0 \\ 0.0 \end{bmatrix}.

The output layer has weights W(2)=[4.05.0]W^{(2)} = \begin{bmatrix} 4.0 & 5.0 \end{bmatrix} and baseline bias b(2)=[−1.5]b^{(2)} = \begin{bmatrix} -1.5 \end{bmatrix}.

  • Part A: Calculate the output pre-activation z(2)andthefinalSigmoidprobabilityz^{(2)} and the final Sigmoid probability a^{(2)} = \sigma(z^{(2)})(use(usee^{1.5} \approx 4.4817$).
  • Part B: Explain why the network's final output becomes completely unresponsive to the input features xx or Layer 1 weights W(1)W^{(1)} under this condition.
Reveal Solution

Part A: Calculate Output Values

z(2)=W(2)a(1)+b(2)=(4.0×0.0)+(5.0×0.0)+(−1.5)=0.0+0.0−1.5=−1.5z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} = (4.0 \times 0.0) + (5.0 \times 0.0) + (-1.5) = 0.0 + 0.0 - 1.5 = \mathbf{-1.5} a(2)=σ(−1.5)=11+e−(−1.5)=11+e1.5=11+4.4817=15.4817≈0.1824  ⟹  18.2%a^{(2)} = \sigma(-1.5) = \frac{1}{1 + e^{-(-1.5)}} = \frac{1}{1 + e^{1.5}} = \frac{1}{1 + 4.4817} = \frac{1}{5.4817} \approx \mathbf{0.1824} \implies \mathbf{18.2\%}

Part B: Mathematical Decoupling Because a(1)=0a^{(1)} = \mathbf{0}, the dot product W(2)a(1)W^{(2)} a^{(1)} evaluates to exactly zero regardless of the values in W(2)W^{(2)}. The output pre-activation collapses strictly to the output bias: z(2)=0+b(2)=b(2)z^{(2)} = 0 + b^{(2)} = b^{(2)}.

When all hidden units are dormant, no information from the input vector xx reaches the output layer. The network outputs a fixed default baseline probability (σ(b(2))=18.2%\sigma(b^{(2)}) = 18.2\%) determined solely by its environmental bias.


Part 2: Applied Scenario: The VC 2-Layer Representation Network

In Topics 1, 2, and 3, our venture capital firm evaluated startup pitch profiles across four locked criteria [Team Experience, Market Size, Competition, Risk]:

xOmniFlow=[1.00.31.00.2]xSolaris=[0.31.00.00.8]x_{\text{OmniFlow}} = \begin{bmatrix} 1.0 \\ 0.3 \\ 1.0 \\ 0.2 \end{bmatrix} \qquad\qquad x_{\text{Solaris}} = \begin{bmatrix} 0.3 \\ 1.0 \\ 0.0 \\ 0.8 \end{bmatrix}

Rather than evaluating raw pitch metrics directly, the investment committee deploys a 2-layer representation network (4→2→14 \to 2 \to 1).

[ Raw Pitch Features ]         [ Layer 1: Latent Factors ]         [ Layer 2: Decision ]
(Team, Market, Comp, Risk)     (Scalability & Defensibility)      (Term Sheet Probability)

Layer 1 Configuration (Latent Investment Factors):

  1. Hidden Neuron 1 (Scalability Factor - a1(1)a_1^{(1)}): Captures total market expansion potential. Heavily rewards Market Size (+4.0+4.0) and Team Experience (+2.0+2.0), mildly penalizes technical Risk (−1.0-1.0), and ignores Competition (0.00.0), with hurdle bias b1(1)=−1.0b_1^{(1)} = -1.0.
  2. Hidden Neuron 2 (Defensibility Factor - a2(1)a_2^{(1)}): Captures competitive moats and barriers to entry. Heavily rewards Team Pedigree (+3.0+3.0), severely penalizes crowded Competition (−3.0-3.0), rewards technical barrier Risk (+1.0+1.0), and ignores Market Size (0.00.0), with hurdle bias b2(1)=−0.5b_2^{(1)} = -0.5.
W(1)=[2.04.00.0−1.03.00.0−3.01.0],b(1)=[−1.0−0.5],f(1)=ReLUW^{(1)} = \begin{bmatrix} 2.0 & 4.0 & 0.0 & -1.0 \\ 3.0 & 0.0 & -3.0 & 1.0 \end{bmatrix}, \qquad b^{(1)} = \begin{bmatrix} -1.0 \\ -0.5 \end{bmatrix}, \qquad f^{(1)} = \text{ReLU}

Layer 2 Configuration (Investment Committee Conversion):

The investment committee evaluates the two hidden factors: it rewards Scalability (+1.0+1.0) and places high premium on Defensibility (+1.5+1.5), against a strict macro hurdle bias b(2)=−3.0b^{(2)} = -3.0:

W(2)=[1.01.5],b(2)=[−3.0],f(2)=SigmoidW^{(2)} = \begin{bmatrix} 1.0 & 1.5 \end{bmatrix}, \qquad b^{(2)} = \begin{bmatrix} -3.0 \end{bmatrix}, \qquad f^{(2)} = \text{Sigmoid}

To visually trace how both startup pitch vectors propagate through intermediate latent representations into the final term sheet conversion probability, examine the comparative dataflow below:

VC 2-Layer Representation Network Dataflow
Comparative 2-layer forward pass for OmniFlow and Solaris AI: OmniFlow's heavy competition penalty silences its Defensibility factor to 0.0 via ReLU, leaving it with insufficient signal to clear the Layer 2 hurdle ($26.9\%$ Pass). Solaris AI's uncontested market generates both Scalability ($2.8$) and Defensibility ($1.2$), securing an $83.2\%$ Term Sheet Greenlight.

Problem 6: Complete 2-Layer Startup Forward Pass

  • Part A: Compute the Layer 1 pre-activation vector z(1)=W(1)x+b(1)z^{(1)} = W^{(1)} x + b^{(1)} and the ReLU hidden activation vector a(1)=ReLU(z(1))a^{(1)} = \text{ReLU}\left(z^{(1)}\right) for both OmniFlow and Solaris AI.
  • Part B: Compute the Layer 2 scalar pre-activation z(2)=W(2)a(1)+b(2)z^{(2)} = W^{(2)} a^{(1)} + b^{(2)} and the final Sigmoid probability a(2)=σ(z(2))a^{(2)} = \sigma(z^{(2)}) for both startups (use e1.0≈2.7183e^{1.0} \approx 2.7183 and e−1.6≈0.2019e^{-1.6} \approx 0.2019).
  • Part C: State which startup earns the term sheet greenlight (y^≥50%\hat{y} \ge 50\%).
  • Part D: Provide an investment committee synthesis explaining how representation learning in Layer 1 exposed OmniFlow's critical vulnerability compared to Solaris AI.
Reveal Solution

Part A: Layer 1 Forward Calculation

For OmniFlow (x=[1.0,0.3,1.0,0.2]Tx = [1.0, 0.3, 1.0, 0.2]^T):

  • Neuron 1 (Scalability): z1(1)=(2.0)(1.0)+(4.0)(0.3)+(0.0)(1.0)+(−1.0)(0.2)+(−1.0)z_1^{(1)} = (2.0)(1.0) + (4.0)(0.3) + (0.0)(1.0) + (-1.0)(0.2) + (-1.0) z1(1)=2.0+1.2+0.0−0.2−1.0=3.2−1.2=+2.0z_1^{(1)} = 2.0 + 1.2 + 0.0 - 0.2 - 1.0 = 3.2 - 1.2 = \mathbf{+2.0} a1(1)=ReLU(2.0)=2.0a_1^{(1)} = \text{ReLU}(2.0) = \mathbf{2.0}

  • Neuron 2 (Defensibility): z2(1)=(3.0)(1.0)+(0.0)(0.3)+(−3.0)(1.0)+(1.0)(0.2)+(−0.5)z_2^{(1)} = (3.0)(1.0) + (0.0)(0.3) + (-3.0)(1.0) + (1.0)(0.2) + (-0.5) z2(1)=3.0+0.0−3.0+0.2−0.5=0.2−0.5=−0.3z_2^{(1)} = 3.0 + 0.0 - 3.0 + 0.2 - 0.5 = 0.2 - 0.5 = \mathbf{-0.3} a2(1)=ReLU(−0.3)=0.0a_2^{(1)} = \text{ReLU}(-0.3) = \mathbf{0.0}

zOmniFlow(1)=[+2.0−0.3],aOmniFlow(1)=[2.00.0](Scalability Factor)(Defensibility Factor)z^{(1)}_{\text{OmniFlow}} = \begin{bmatrix} +2.0 \\ -0.3 \end{bmatrix}, \qquad a^{(1)}_{\text{OmniFlow}} = \begin{bmatrix} 2.0 \\ 0.0 \end{bmatrix} \begin{matrix} \text{(Scalability Factor)} \\ \text{(Defensibility Factor)} \end{matrix}

For Solaris AI (x=[0.3,1.0,0.0,0.8]Tx = [0.3, 1.0, 0.0, 0.8]^T):

  • Neuron 1 (Scalability): z1(1)=(2.0)(0.3)+(4.0)(1.0)+(0.0)(0.0)+(−1.0)(0.8)+(−1.0)z_1^{(1)} = (2.0)(0.3) + (4.0)(1.0) + (0.0)(0.0) + (-1.0)(0.8) + (-1.0) z1(1)=0.6+4.0+0.0−0.8−1.0=4.6−1.8=+2.8z_1^{(1)} = 0.6 + 4.0 + 0.0 - 0.8 - 1.0 = 4.6 - 1.8 = \mathbf{+2.8} a1(1)=ReLU(2.8)=2.8a_1^{(1)} = \text{ReLU}(2.8) = \mathbf{2.8}

  • Neuron 2 (Defensibility): z2(1)=(3.0)(0.3)+(0.0)(1.0)+(−3.0)(0.0)+(1.0)(0.8)+(−0.5)z_2^{(1)} = (3.0)(0.3) + (0.0)(1.0) + (-3.0)(0.0) + (1.0)(0.8) + (-0.5) z2(1)=0.9+0.0−0.0+0.8−0.5=1.7−0.5=+1.2z_2^{(1)} = 0.9 + 0.0 - 0.0 + 0.8 - 0.5 = 1.7 - 0.5 = \mathbf{+1.2} a2(1)=ReLU(1.2)=1.2a_2^{(1)} = \text{ReLU}(1.2) = \mathbf{1.2}

zSolaris(1)=[+2.8+1.2],aSolaris(1)=[2.81.2](Scalability Factor)(Defensibility Factor)z^{(1)}_{\text{Solaris}} = \begin{bmatrix} +2.8 \\ +1.2 \end{bmatrix}, \qquad a^{(1)}_{\text{Solaris}} = \begin{bmatrix} 2.8 \\ 1.2 \end{bmatrix} \begin{matrix} \text{(Scalability Factor)} \\ \text{(Defensibility Factor)} \end{matrix}

Part B: Layer 2 Forward Calculation

For OmniFlow (a(1)=[2.0,0.0]Ta^{(1)} = [2.0, 0.0]^T):

zOmniFlow(2)=W(2)a(1)+b(2)=(1.0×2.0)+(1.5×0.0)+(−3.0)=2.0+0.0−3.0=−1.0z^{(2)}_{\text{OmniFlow}} = W^{(2)} a^{(1)} + b^{(2)} = (1.0 \times 2.0) + (1.5 \times 0.0) + (-3.0) = 2.0 + 0.0 - 3.0 = \mathbf{-1.0} aOmniFlow(2)=σ(−1.0)=11+e−(−1.0)=11+e1.0=11+2.7183=13.7183≈0.2689  ⟹  26.9%a^{(2)}_{\text{OmniFlow}} = \sigma(-1.0) = \frac{1}{1 + e^{-(-1.0)}} = \frac{1}{1 + e^{1.0}} = \frac{1}{1 + 2.7183} = \frac{1}{3.7183} \approx \mathbf{0.2689} \implies \mathbf{26.9\%}

For Solaris AI (a(1)=[2.8,1.2]Ta^{(1)} = [2.8, 1.2]^T):

zSolaris(2)=W(2)a(1)+b(2)=(1.0×2.8)+(1.5×1.2)+(−3.0)=2.8+1.8−3.0=4.6−3.0=+1.6z^{(2)}_{\text{Solaris}} = W^{(2)} a^{(1)} + b^{(2)} = (1.0 \times 2.8) + (1.5 \times 1.2) + (-3.0) = 2.8 + 1.8 - 3.0 = 4.6 - 3.0 = \mathbf{+1.6} aSolaris(2)=σ(1.6)=11+e−1.6=11+0.2019=11.2019≈0.8320  ⟹  83.2%a^{(2)}_{\text{Solaris}} = \sigma(1.6) = \frac{1}{1 + e^{-1.6}} = \frac{1}{1 + 0.2019} = \frac{1}{1.2019} \approx \mathbf{0.8320} \implies \mathbf{83.2\%}

Part C: Investment Decision

  • Solaris AI achieves an 83.2%83.2\% Term Sheet Greenlight (y^≥50%\hat{y} \ge 50\%).
  • OmniFlow is rejected with a 26.9%26.9\% Probability (y^<50%\hat{y} < 50\%).

Part D: Strategic Representation Synthesis

The 2-layer architecture reveals why representation learning is critical for robust decisions:

  1. OmniFlow's Hidden Vulnerability: In Layer 1, OmniFlow's veteran team (x1=1.0x_1 = 1.0) generated a solid Scalability score (a1(1)=2.0a_1^{(1)} = 2.0). However, because OmniFlow operates in an intensely crowded market (x3=1.0x_3 = 1.0), its competitive Defensibility score went negative (z2(1)=−0.3z_2^{(1)} = -0.3). ReLU clamped this channel strictly to 0.00.0. When Layer 2 evaluated OmniFlow, it had only one active factor (2.02.0), which was insufficient to overcome the strict −3.0-3.0 committee hurdle (z(2)=−1.0  ⟹  26.9%z^{(2)} = -1.0 \implies 26.9\%).
  2. Solaris AI's Dual-Factor Strength: Solaris AI entered with high technical execution risk (x4=0.8x_4 = 0.8) and a modest team score (x1=0.3x_1 = 0.3). However, its massive market (x2=1.0x_2 = 1.0) and total absence of competition (x3=0.0x_3 = 0.0) allowed both hidden neurons to fire strongly (a1(1)=2.8,a2(1)=1.2a_1^{(1)} = 2.8, a_2^{(1)} = 1.2). When Layer 2 multiplied these intermediate concepts by their committee weights (1.0(2.8)+1.5(1.2)=4.61.0(2.8) + 1.5(1.2) = 4.6), it easily absorbed the −3.0-3.0 hurdle, securing a decisive 83.2%83.2\% greenlight.

By utilizing hidden layers, the network avoids making simplistic decisions on raw features, structuring intermediate insights to produce nuanced, defensible predictions.

Previous
The Complete MLP Forward Pass Math