Matrices and Layer Width In Practice hero
LaboratoryMatrices & Layer Width (Linear Algebra Part 2)

Matrices and Layer Width In Practice

Master matrix-vector multiplication, layer width scaling, affine layer transformations, and parallel multi-output evaluation through manual calculation.

Part 1: The Underlying Mechanics Drill

Work through each problem on paper before revealing the step-by-step solution.

Problem 1: 2×22 \times 2 Matrix-Vector Multiplication

Calculate the matrix-vector product WxWx for the given weight matrix WW and input vector xx:

W=[3−214],x=[25]W = \begin{bmatrix} 3 & -2 \\ 1 & 4 \end{bmatrix}, \qquad x = \begin{bmatrix} 2 \\ 5 \end{bmatrix}
Reveal Solution

Compute each row dot product:

  • Row 1: (3)(2)+(−2)(5)=6−10=−4(3)(2) + (-2)(5) = 6 - 10 = -4
  • Row 2: (1)(2)+(4)(5)=2+20=22(1)(2) + (4)(5) = 2 + 20 = 22

Assemble the resulting output vector:

Wx=[(3)(2)+(−2)(5)(1)(2)+(4)(5)]=[−422]Wx = \begin{bmatrix} (3)(2) + (-2)(5) \\ (1)(2) + (4)(5) \end{bmatrix} = \begin{bmatrix} -4 \\ 22 \end{bmatrix}

The 2×22 \times 2 matrix transforms the 2D input vector [2,5]T[2, 5]^T into the 2D output vector [−4,22]T[-4, 22]^T.


Problem 2: 3×23 \times 2 Transformation (Dimensional Expansion)

Calculate the matrix-vector product WxWx for the given matrix W∈R3×2W \in \mathbb{R}^{3 \times 2} and vector x∈R2×1x \in \mathbb{R}^{2 \times 1}:

W=[21−1304],x=[3−2]W = \begin{bmatrix} 2 & 1 \\ -1 & 3 \\ 0 & 4 \end{bmatrix}, \qquad x = \begin{bmatrix} 3 \\ -2 \end{bmatrix}
Reveal Solution

Compute each row dot product:

  • Row 1: (2)(3)+(1)(−2)=6−2=4(2)(3) + (1)(-2) = 6 - 2 = 4
  • Row 2: (−1)(3)+(3)(−2)=−3−6=−9(-1)(3) + (3)(-2) = -3 - 6 = -9
  • Row 3: (0)(3)+(4)(−2)=0−8=−8(0)(3) + (4)(-2) = 0 - 8 = -8

Assemble the resulting output vector:

Wx=[(2)(3)+(1)(−2)(−1)(3)+(3)(−2)(0)(3)+(4)(−2)]=[4−9−8]Wx = \begin{bmatrix} (2)(3) + (1)(-2) \\ (-1)(3) + (3)(-2) \\ (0)(3) + (4)(-2) \end{bmatrix} = \begin{bmatrix} 4 \\ -9 \\ -8 \end{bmatrix}

Dimensionality Insight: The matrix W∈R3×2W \in \mathbb{R}^{3 \times 2} accepts a 2-dimensional input vector (x∈R2x \in \mathbb{R}^2) and maps it into a 3-dimensional output coordinate space (Wx∈R3Wx \in \mathbb{R}^3).


Problem 3: Affine Layer Linear Sum (z=Wx+bz = Wx + b)

Calculate the pre-activation vector z=Wx+bz = Wx + b for the given weight matrix WW, input vector xx, and bias vector bb:

W=[1.5−0.52.01.0],x=[4.02.0],b=[−1.03.0]W = \begin{bmatrix} 1.5 & -0.5 \\ 2.0 & 1.0 \end{bmatrix}, \qquad x = \begin{bmatrix} 4.0 \\ 2.0 \end{bmatrix}, \qquad b = \begin{bmatrix} -1.0 \\ 3.0 \end{bmatrix}
Reveal Solution

Step 1: Compute matrix-vector multiplication WxWx:

  • Row 1: (1.5)(4.0)+(−0.5)(2.0)=6.0−1.0=5.0(1.5)(4.0) + (-0.5)(2.0) = 6.0 - 1.0 = 5.0
  • Row 2: (2.0)(4.0)+(1.0)(2.0)=8.0+2.0=10.0(2.0)(4.0) + (1.0)(2.0) = 8.0 + 2.0 = 10.0
Wx=[5.010.0]Wx = \begin{bmatrix} 5.0 \\ 10.0 \end{bmatrix}

Step 2: Add the bias vector bb:

z=Wx+b=[5.010.0]+[−1.03.0]=[5.0+(−1.0)10.0+3.0]=[4.013.0]z = Wx + b = \begin{bmatrix} 5.0 \\ 10.0 \end{bmatrix} + \begin{bmatrix} -1.0 \\ 3.0 \end{bmatrix} = \begin{bmatrix} 5.0 + (-1.0) \\ 10.0 + 3.0 \end{bmatrix} = \begin{bmatrix} 4.0 \\ 13.0 \end{bmatrix}

The affine linear sum vector is z=[4.0,13.0]Tz = [4.0, 13.0]^T.


Problem 4: Dimension Compatibility & Layer Sizing

Given the following matrices and vectors:

  • A∈R3×4A \in \mathbb{R}^{3 \times 4}

  • B∈R4×2B \in \mathbb{R}^{4 \times 2}

  • x∈R4×1x \in \mathbb{R}^{4 \times 1}

  • y∈R3×1y \in \mathbb{R}^{3 \times 1}

  • Part A: Is the product AxAx mathematically defined? If so, what is the dimension of the resulting vector?

  • Part B: Is the product AyAy mathematically defined? Explain why or why not.

  • Part C: Is the product BTxB^T x mathematically defined? (Recall BT∈R2×4B^T \in \mathbb{R}^{2 \times 4}). If so, what is the output dimension?

  • Part D: If an engineer designs a dense layer with layer width m=5m = 5 that accepts an input vector of n=3n = 3 features, what must the dimensions of the weight matrix WW and bias vector bb be?

Reveal Solution
  • Part A: Yes, AxAx is defined.
    • Shape check: (3×4)×(4×1)(3 \times 4) \times (4 \times 1). The inner dimensions match (4=44 = 4).
    • The resulting output vector has shape 3×13 \times 1 (Ax∈R3Ax \in \mathbb{R}^3).
  • Part B: No, AyAy is undefined.
    • Shape check: (3×4)×(3×1)(3 \times 4) \times (3 \times 1). The inner dimensions do not match (4≠34 \neq 3). Matrix AA expects an input vector with 4 features, but vector yy has only 3 entries.
  • Part C: Yes, BTxB^T x is defined.
    • Shape check: (2×4)×(4×1)(2 \times 4) \times (4 \times 1). The inner dimensions match (4=44 = 4).
    • The resulting output vector has shape 2×12 \times 1 (BTx∈R2B^T x \in \mathbb{R}^2).
  • Part D: For layer width m=5m = 5 and input dimension n=3n = 3:
    • Weight Matrix: W∈R5×3W \in \mathbb{R}^{5 \times 3} (55 rows for the 5 parallel neurons, 33 columns for the 3 input features).
    • Bias Vector: b∈R5×1b \in \mathbb{R}^{5 \times 1} (55 baseline offsets, one for each neuron).

Problem 5: Zero-Row Nullification and Parameter Isolation

Consider a 3-neuron layer parameterized by the following weight matrix WW and bias vector bb, processing an extreme input vector xx:

W=[2.0−1.03.00.00.00.01.04.0−2.0],x=[100.0500.0999.0],b=[0.0−4.01.0]W = \begin{bmatrix} 2.0 & -1.0 & 3.0 \\ 0.0 & 0.0 & 0.0 \\ 1.0 & 4.0 & -2.0 \end{bmatrix}, \qquad x = \begin{bmatrix} 100.0 \\ 500.0 \\ 999.0 \end{bmatrix}, \qquad b = \begin{bmatrix} 0.0 \\ -4.0 \\ 1.0 \end{bmatrix}
  • Part A: Calculate the pre-activation value z2z_2 for the second neuron (Row 2).
  • Part B: Explain why z2z_2 remains completely unchanged regardless of whether the input vector entries are 0.00.0, 500.0500.0, or 1,000,000.01{,}000{,}000.0.
Reveal Solution

Part A: Pre-activation Calculation for Neuron 2 (z2z_2)

z2=(W21x1+W22x2+W23x3)+b2z_2 = (W_{21} x_1 + W_{22} x_2 + W_{23} x_3) + b_2 z2=(0.0×100.0)+(0.0×500.0)+(0.0×999.0)+(−4.0)=0.0−4.0=−4.0z_2 = (0.0 \times 100.0) + (0.0 \times 500.0) + (0.0 \times 999.0) + (-4.0) = 0.0 - 4.0 = \mathbf{-4.0}

Part B: Mathematical Isolation

Because every weight in Row 2 is exactly zero (w2=[0,0,0]Tw_2 = [0, 0, 0]^T), the dot product w2⋅xw_2 \cdot x is identically zero for all possible input vectors x∈R3x \in \mathbb{R}^3.

Neuron 2 is mathematically decoupled from the input data. Its output depends exclusively on its baseline bias: z2=0+b2=−4.0z_2 = 0 + b_2 = -4.0. In neural networks, zeroing out a row of weights completely nullifies that neuron's sensitivity to the input features.


Part 2: Applied Scenario: The VC Decision Matrix

In Topics 1 and 2, we evaluated startup pitch decks for a venture capital firm using a 4-dimensional feature vector across locked features [Team Experience, Market Size, Competition, Risk]:

xOmniFlow=[1.00.31.00.2]xSolaris=[0.31.00.00.8]x_{\text{OmniFlow}} = \begin{bmatrix} 1.0 \\ 0.3 \\ 1.0 \\ 0.2 \end{bmatrix} \qquad\qquad x_{\text{Solaris}} = \begin{bmatrix} 0.3 \\ 1.0 \\ 0.0 \\ 0.8 \end{bmatrix}

Rather than evaluating a single Unicorn score, the investment committee uses a 3-channel decision matrix to evaluate every startup across three independent investment outcomes simultaneously:

  1. Channel 1 (Unicorn Potential): Looks for massive market size and exceptional team pedigree; penalizes heavy competition. w1=[4.0,2.0,−3.0,1.0]T,b1=−2.5w_1 = [4.0, 2.0, -3.0, 1.0]^T, \qquad b_1 = -2.5
  2. Channel 2 (Capital Efficiency / Cash Flow): Requires strong team execution and low capital burn; heavily penalizes high risk and execution uncertainty. w2=[2.0,1.0,−1.0,−4.0]T,b2=−0.5w_2 = [2.0, 1.0, -1.0, -4.0]^T, \qquad b_2 = -0.5
  3. Channel 3 (M&A Strategic Acquisition Target): Looks for startups operating in crowded, high-competition sectors with high risk that established tech giants frequently acquire for talent and niche market share. w3=[−1.0,0.0,4.0,1.0]T,b3=−1.0w_3 = [-1.0, 0.0, 4.0, 1.0]^T, \qquad b_3 = -1.0

We assemble these three investment channels into the Weight Matrix W∈R3×4W \in \mathbb{R}^{3 \times 4} and Bias Vector b∈R3×1b \in \mathbb{R}^{3 \times 1}:

W=[4.02.0−3.01.02.01.0−1.0−4.0−1.00.04.01.0],b=[−2.5−0.5−1.0]W = \begin{bmatrix} 4.0 & 2.0 & -3.0 & 1.0 \\ 2.0 & 1.0 & -1.0 & -4.0 \\ -1.0 & 0.0 & 4.0 & 1.0 \end{bmatrix}, \qquad b = \begin{bmatrix} -2.5 \\ -0.5 \\ -1.0 \end{bmatrix}

Problem 6: Multi-Output Startup Evaluation

  • Part A: Compute the raw matrix product WxWx for both OmniFlow and Solaris AI.
  • Part B: Compute the affine pre-activation vector z=Wx+bz = Wx + b for both startups.
  • Part C: Calculate the Sigmoid activation probability vector a=σ(z)a = \sigma(z) for both startups (use reference values: σ(−0.7)≈0.332\sigma(-0.7) \approx 0.332, σ(0.0)=0.500\sigma(0.0) = 0.500, σ(2.2)≈0.900\sigma(2.2) \approx 0.900, σ(1.5)≈0.818\sigma(1.5) \approx 0.818, σ(−2.1)≈0.109\sigma(-2.1) \approx 0.109, σ(−0.5)≈0.378\sigma(-0.5) \approx 0.378).
  • Part D: Provide a concise investment committee synthesis explaining how the multi-neuron layer reveals distinct strategic profiles for each startup.
Reveal Solution

Part A & B: Matrix-Vector Multiplication and Affine Sums

For OmniFlow (x=[1.0,0.3,1.0,0.2]Tx = [1.0, 0.3, 1.0, 0.2]^T):

  • Channel 1 (Unicorn Potential): w1⋅x=(4.0)(1.0)+(2.0)(0.3)+(−3.0)(1.0)+(1.0)(0.2)=4.0+0.6−3.0+0.2=+1.8w_1 \cdot x = (4.0)(1.0) + (2.0)(0.3) + (-3.0)(1.0) + (1.0)(0.2) = 4.0 + 0.6 - 3.0 + 0.2 = \mathbf{+1.8} z1=1.8+(−2.5)=−0.7z_1 = 1.8 + (-2.5) = \mathbf{-0.7}

  • Channel 2 (Capital Efficiency): w2⋅x=(2.0)(1.0)+(1.0)(0.3)+(−1.0)(1.0)+(−4.0)(0.2)=2.0+0.3−1.0−0.8=+0.5w_2 \cdot x = (2.0)(1.0) + (1.0)(0.3) + (-1.0)(1.0) + (-4.0)(0.2) = 2.0 + 0.3 - 1.0 - 0.8 = \mathbf{+0.5} z2=0.5+(−0.5)=0.0z_2 = 0.5 + (-0.5) = \mathbf{0.0}

  • Channel 3 (M&A Acquisition Target): w3⋅x=(−1.0)(1.0)+(0.0)(0.3)+(4.0)(1.0)+(1.0)(0.2)=−1.0+0.0+4.0+0.2=+3.2w_3 \cdot x = (-1.0)(1.0) + (0.0)(0.3) + (4.0)(1.0) + (1.0)(0.2) = -1.0 + 0.0 + 4.0 + 0.2 = \mathbf{+3.2} z3=3.2+(−1.0)=+2.2z_3 = 3.2 + (-1.0) = \mathbf{+2.2}

WxOmniFlow=[1.80.53.2],zOmniFlow=[−0.70.02.2]Wx_{\text{OmniFlow}} = \begin{bmatrix} 1.8 \\ 0.5 \\ 3.2 \end{bmatrix}, \qquad z_{\text{OmniFlow}} = \begin{bmatrix} -0.7 \\ 0.0 \\ 2.2 \end{bmatrix}

For Solaris AI (x=[0.3,1.0,0.0,0.8]Tx = [0.3, 1.0, 0.0, 0.8]^T):

  • Channel 1 (Unicorn Potential): w1⋅x=(4.0)(0.3)+(2.0)(1.0)+(−3.0)(0.0)+(1.0)(0.8)=1.2+2.0−0.0+0.8=+4.0w_1 \cdot x = (4.0)(0.3) + (2.0)(1.0) + (-3.0)(0.0) + (1.0)(0.8) = 1.2 + 2.0 - 0.0 + 0.8 = \mathbf{+4.0} z1=4.0+(−2.5)=+1.5z_1 = 4.0 + (-2.5) = \mathbf{+1.5}

  • Channel 2 (Capital Efficiency): w2⋅x=(2.0)(0.3)+(1.0)(1.0)+(−1.0)(0.0)+(−4.0)(0.8)=0.6+1.0−0.0−3.2=−1.6w_2 \cdot x = (2.0)(0.3) + (1.0)(1.0) + (-1.0)(0.0) + (-4.0)(0.8) = 0.6 + 1.0 - 0.0 - 3.2 = \mathbf{-1.6} z2=−1.6+(−0.5)=−2.1z_2 = -1.6 + (-0.5) = \mathbf{-2.1}

  • Channel 3 (M&A Acquisition Target): w3⋅x=(−1.0)(0.3)+(0.0)(1.0)+(4.0)(0.0)+(1.0)(0.8)=−0.3+0.0+0.0+0.8=+0.5w_3 \cdot x = (-1.0)(0.3) + (0.0)(1.0) + (4.0)(0.0) + (1.0)(0.8) = -0.3 + 0.0 + 0.0 + 0.8 = \mathbf{+0.5} z3=0.5+(−1.0)=−0.5z_3 = 0.5 + (-1.0) = \mathbf{-0.5}

WxSolaris=[4.0−1.60.5],zSolaris=[1.5−2.1−0.5]Wx_{\text{Solaris}} = \begin{bmatrix} 4.0 \\ -1.6 \\ 0.5 \end{bmatrix}, \qquad z_{\text{Solaris}} = \begin{bmatrix} 1.5 \\ -2.1 \\ -0.5 \end{bmatrix}

Part C: Sigmoid Activation Vectors

Applying the Sigmoid activation function ai=σ(zi)=11+e−zia_i = \sigma(z_i) = \frac{1}{1 + e^{-z_i}}:

For OmniFlow:

  • a1=σ(−0.7)=11+e0.7≈0.332  ⟹  33.2%a_1 = \sigma(-0.7) = \frac{1}{1 + e^{0.7}} \approx \mathbf{0.332} \implies \mathbf{33.2\%}
  • a2=σ(0.0)=11+e0.0=0.500  ⟹  50.0%a_2 = \sigma(0.0) = \frac{1}{1 + e^{0.0}} = \mathbf{0.500} \implies \mathbf{50.0\%}
  • a3=σ(2.2)=11+e−2.2≈0.900  ⟹  90.0%a_3 = \sigma(2.2) = \frac{1}{1 + e^{-2.2}} \approx \mathbf{0.900} \implies \mathbf{90.0\%}
a_{\text{OmniFlow}} = \begin{bmatrix} 0.332 \\ 0.500 \\ 0.900 \end{bmatrix} \begin{matrix} \text{(33.2\% Unicorn Potential)} \\ \text{(50.0\% Capital Efficiency)} \\ \text{(90.0\% M&A Target)} \end{matrix}

For Solaris AI:

  • a1=σ(1.5)=11+e−1.5≈0.818  ⟹  81.8%a_1 = \sigma(1.5) = \frac{1}{1 + e^{-1.5}} \approx \mathbf{0.818} \implies \mathbf{81.8\%}
  • a2=σ(−2.1)=11+e2.1≈0.109  ⟹  10.9%a_2 = \sigma(-2.1) = \frac{1}{1 + e^{2.1}} \approx \mathbf{0.109} \implies \mathbf{10.9\%}
  • a3=σ(−0.5)=11+e0.5≈0.378  ⟹  37.8%a_3 = \sigma(-0.5) = \frac{1}{1 + e^{0.5}} \approx \mathbf{0.378} \implies \mathbf{37.8\%}
a_{\text{Solaris}} = \begin{bmatrix} 0.818 \\ 0.109 \\ 0.378 \end{bmatrix} \begin{matrix} \text{(81.8\% Unicorn Potential)} \\ \text{(10.9\% Capital Efficiency)} \\ \text{(37.8\% M&A Target)} \end{matrix}

To visually compare how both startup pitch vectors flow through the 3-channel investment decision matrix, examine the computational flow below:

VC Decision Matrix Computational Flow
Comparative 3-channel decision matrix flow for OmniFlow and Solaris AI: OmniFlow's heavy competition penalty suppresses its Unicorn score ($33.2\%$) but elevates its M&A acquisition probability to $90.0\%$, while Solaris AI's uncontested market dominance drives an $81.8\%$ Unicorn greenlight despite high technical burn ($10.9\%$ Capital Efficiency).

Part D: Investment Committee Strategic Synthesis

The multi-output layer reveals that evaluating startups across a single dimension produces an incomplete picture:

  1. OmniFlow is rejected for a standalone Unicorn fund investment (33.2%<50%33.2\% < 50\%) because severe sector competition drags down its score. However, its veteran team and high market competition make it an ideal M&A Strategic Acquisition Target (90.0%90.0\%), representing a strong opportunity for a private equity or strategic growth fund.
  2. Solaris AI is a clear Unicorn Fund Greenlight (81.8%≥50%81.8\% \ge 50\%) driven by an uncontested, massive total addressable market (x2=1.0,x3=0.0x_2 = 1.0, x_3 = 0.0). However, its low Capital Efficiency score (10.9%10.9\%) warns the partners that the company will require substantial follow-on capital reserves to survive its high technical risk profile (x4=0.8x_4 = 0.8).

By stacking decision channels into a single matrix forward pass (a=σ(Wx+b)a = \sigma(Wx + b)), the neural network converts a single raw pitch profile into a multidimensional strategic assessment in a single computational step.

Previous
The Affine Layer Transformation Math