Matrix-Vector Multiplication hero
Lesson 2Matrices & Layer Width (Linear Algebra Part 2)

Matrix-Vector Multiplication

Matrix-vector multiplication executed as parallel row dot products to transform input feature coordinates into multi-output decision vectors.

In Lesson 1, we stacked the weights of three studio neurons into the horizontal rows of a single weight matrix WW.

When an incoming script's feature vector xx arrives at a neural network layer, the model must evaluate all of those rows against that single input simultaneously. That operation is matrix-vector multiplication.

Despite the formal name, matrix-vector multiplication is not a new branch of arithmetic. It is simply multiple dot products executed in parallel.


The Row-by-Column Dot Product Rule

To calculate the ii-th entry of the output vector zz, we compute the dot product between the ii-th row of the weight matrix WW and the input vector xx:

zi=∑j=1nWijxj=Wi1x1+Wi2x2+⋯+Winxn=wi⋅xz_i = \sum_{j=1}^n W_{ij} x_j = W_{i1} x_1 + W_{i2} x_2 + \dots + W_{in} x_n = w_i \cdot x

Writing the full operation in expanded algebraic form:

[W11W12…W1nW21W22…W2n⋮⋮⋱⋮Wm1Wm2…Wmn][x1x2⋮xn]=[W11x1+W12x2+⋯+W1nxnW21x1+W22x2+⋯+W2nxn⋮Wm1x1+Wm2x2+⋯+Wmnxn]=[w1⋅xw2⋅x⋮wm⋅x]=[z1z2⋮zm]\begin{bmatrix} W_{11} & W_{12} & \dots & W_{1n} \\ W_{21} & W_{22} & \dots & W_{2n} \\ \vdots & \vdots & \ddots & \vdots \\ W_{m1} & W_{m2} & \dots & W_{mn} \end{bmatrix} \begin{bmatrix} x_1 \\ x_2 \\ \vdots \\ x_n \end{bmatrix} = \begin{bmatrix} W_{11}x_1 + W_{12}x_2 + \dots + W_{1n}x_n \\ W_{21}x_1 + W_{22}x_2 + \dots + W_{2n}x_n \\ \vdots \\ W_{m1}x_1 + W_{m2}x_2 + \dots + W_{mn}x_n \end{bmatrix} = \begin{bmatrix} w_1 \cdot x \\ w_2 \cdot x \\ \vdots \\ w_m \cdot x \end{bmatrix} = \begin{bmatrix} z_1 \\ z_2 \\ \vdots \\ z_m \end{bmatrix}

Every row of the matrix sweeps horizontally across its columns, multiplying each weight WijW_{ij} by the corresponding feature value xjx_j, and sums the individual products into a single scalar value.


Walk-Through: Calculating Matrix Scores for 'Die Hard in Space'

Now let's execute this exact operation with concrete numbers using our running studio predictor from Lesson 1.

Recall our 3×43 \times 4 Weight Matrix (WW), where each row provides the weights for one neuron across the locked genre order [Action, Romance, Comedy, Sci-Fi]:

W=[5.0−2.01.04.0−5.05.05.00.05.00.00.04.0](Row 1: Hit, w1)(Row 2: Oscar, w2)(Row 3: Franchise, w3)W = \begin{bmatrix} 5.0 & -2.0 & 1.0 & 4.0 \\ -5.0 & 5.0 & 5.0 & 0.0 \\ 5.0 & 0.0 & 0.0 & 4.0 \end{bmatrix} \begin{matrix} \text{(Row 1: Hit, } w_1\text{)} \\ \text{(Row 2: Oscar, } w_2\text{)} \\ \text{(Row 3: Franchise, } w_3\text{)} \end{matrix}

Our incoming script feature vector for 'Die Hard in Space' is:

x=[1.00.00.51.0](Action)(Romance)(Comedy)(Sci-Fi)x = \begin{bmatrix} 1.0 \\ 0.0 \\ 0.5 \\ 1.0 \end{bmatrix} \begin{matrix} \text{(Action)} \\ \text{(Romance)} \\ \text{(Comedy)} \\ \text{(Sci-Fi)} \end{matrix}

Multiplying the weight matrix WW by the input vector xx:

Wx=[5.0−2.01.04.0−5.05.05.00.05.00.00.04.0][1.00.00.51.0]=[(5.0)(1.0)+(−2.0)(0.0)+(1.0)(0.5)+(4.0)(1.0)(−5.0)(1.0)+(5.0)(0.0)+(5.0)(0.5)+(0.0)(1.0)(5.0)(1.0)+(0.0)(0.0)+(0.0)(0.5)+(4.0)(1.0)]W x = \begin{bmatrix} 5.0 & -2.0 & 1.0 & 4.0 \\ -5.0 & 5.0 & 5.0 & 0.0 \\ 5.0 & 0.0 & 0.0 & 4.0 \end{bmatrix} \begin{bmatrix} 1.0 \\ 0.0 \\ 0.5 \\ 1.0 \end{bmatrix} = \begin{bmatrix} (5.0)(1.0) + (-2.0)(0.0) + (1.0)(0.5) + (4.0)(1.0) \\ (-5.0)(1.0) + (5.0)(0.0) + (5.0)(0.5) + (0.0)(1.0) \\ (5.0)(1.0) + (0.0)(0.0) + (0.0)(0.5) + (4.0)(1.0) \end{bmatrix}

We evaluate each row dot product step by step:

  • Row 1 (Hit): w1⋅x=(5.0×1.0)+(−2.0×0.0)+(1.0×0.5)+(4.0×1.0)=5.0+0.0+0.5+4.0=+9.5w_1 \cdot x = (5.0 \times 1.0) + (-2.0 \times 0.0) + (1.0 \times 0.5) + (4.0 \times 1.0) = 5.0 + 0.0 + 0.5 + 4.0 = \mathbf{+9.5}
  • Row 2 (Oscar): w2⋅x=(−5.0×1.0)+(5.0×0.0)+(5.0×0.5)+(0.0×1.0)=−5.0+0.0+2.5+0.0=−2.5w_2 \cdot x = (-5.0 \times 1.0) + (5.0 \times 0.0) + (5.0 \times 0.5) + (0.0 \times 1.0) = -5.0 + 0.0 + 2.5 + 0.0 = \mathbf{-2.5}
  • Row 3 (Franchise): w3⋅x=(5.0×1.0)+(0.0×0.0)+(0.0×0.5)+(4.0×1.0)=5.0+0.0+0.0+4.0=+9.0w_3 \cdot x = (5.0 \times 1.0) + (0.0 \times 0.0) + (0.0 \times 0.5) + (4.0 \times 1.0) = 5.0 + 0.0 + 0.0 + 4.0 = \mathbf{+9.0}

The resulting output is a 3-dimensional column vector holding the raw decision scores:

Wx=[+9.5−2.5+9.0](Hit Raw Score)(Oscar Raw Score)(Franchise Raw Score)W x = \begin{bmatrix} +9.5 \\ -2.5 \\ +9.0 \end{bmatrix} \begin{matrix} \text{(Hit Raw Score)} \\ \text{(Oscar Raw Score)} \\ \text{(Franchise Raw Score)} \end{matrix}

We can also inspect the arithmetic in a structured feature ledger, keeping incoming features stacked along the vertical axis:

Feature DimensionFeature Value (xx)Row 1: Hit (w1w_1)Row 2: Oscar (w2w_2)Row 3: Franchise (w3w_3)
Action (j=1j=1)1.01.01.0×5.0=+5.01.0 \times 5.0 = +5.01.0×−5.0=−5.01.0 \times -5.0 = -5.01.0×5.0=+5.01.0 \times 5.0 = +5.0
Romance (j=2j=2)0.00.00.0×−2.0=0.00.0 \times -2.0 = 0.00.0×5.0=0.00.0 \times 5.0 = 0.00.0×0.0=0.00.0 \times 0.0 = 0.0
Comedy (j=3j=3)0.50.50.5×1.0=+0.50.5 \times 1.0 = +0.50.5×5.0=+2.50.5 \times 5.0 = +2.50.5×0.0=0.00.5 \times 0.0 = 0.0
Sci-Fi (j=4j=4)1.01.01.0×4.0=+4.01.0 \times 4.0 = +4.01.0×0.0=0.01.0 \times 0.0 = 0.01.0×4.0=+4.01.0 \times 4.0 = +4.0
Raw Dot Product (wi⋅xw_i \cdot x)+9.5+9.5−2.5-2.5+9.0+9.0

Notice that Row 1 reproduces the exact raw score of +9.5+9.5 calculated for 'Die Hard in Space' by our single artificial neuron in Topics 1 and 2. Stacking vectors into a matrix did not change the arithmetic—it packaged three independent evaluation channels into a single mathematical operation.


The Dimension Compatibility Rule

Not every matrix can multiply every vector. The physical setup dictates strict dimensional alignment.

In our studio predictor, 'Die Hard in Space' supplies 4 genre features (n=4n = 4). For each row in the matrix to compute a dot product with that script, it must provide exactly 4 weights—one for each genre. If a matrix had only 3 columns, there would be no weight to evaluate Sci-Fi. If it had 5 columns, the 5th weight would have no incoming feature to pair with.

Therefore, the number of columns in the weight matrix must equal the number of rows in the input vector.

Likewise, because each row in WW represents one neuron, stacking mm rows produces exactly mm individual dot products—a column vector of length mm.

In formal linear algebra notation, the Dimension Compatibility Rule is written:

(m×n)×(n×1)⟶(m×1)(m \times n) \times (n \times 1) \longrightarrow (m \times 1)

This rule establishes two structural constraints:

  1. The Inner Dimensions Must Match (n=nn = n): The number of columns in WW (nn) must equal the number of rows in xx (nn). Every incoming feature coordinate must have a corresponding weight parameter in each row.
  2. The Outer Dimensions Determine the Output Shape (m×1m \times 1): The number of rows in WW (mm) determines the number of entries in the resulting output vector. Stacking mm neurons produces an mm-dimensional column vector.

A matrix W∈Rm×nW \in \mathbb{R}^{m \times n} acts as a dimensional transformation machine. It accepts an nn-dimensional input vector and transforms it into an mm-dimensional output vector:

Input (x∈Rn)→Weight Matrix (W∈Rm×n)Output (z∈Rm)\text{Input } (x \in \mathbb{R}^n) \quad \xrightarrow{\quad \text{Weight Matrix } (W \in \mathbb{R}^{m \times n}) \quad} \quad \text{Output } (z \in \mathbb{R}^m)

Depending on the chosen dimensions, a matrix transformation alters coordinate representations in one of three ways:

  • Dimensional Expansion (m>nm > n): Expands the representation into a higher-dimensional space with more output neurons than incoming features.
  • Dimensional Compression (m<nm < n): Compresses the representation into a lower-dimensional space, summarizing inputs into fewer coordinates.
  • Dimensional Preservation (m=nm = n): Maps features into an output space of the exact same dimensionality.

Hardware Parallelism: Why AI Runs on GPUs

Understanding matrix-vector multiplication as parallel row dot products reveals why modern deep learning relies on Graphics Processing Units (GPUs) rather than Central Processing Units (CPUs).

In the output vector:

z=Wx=[w1⋅xw2⋅x⋮wm⋅x]z = W x = \begin{bmatrix} w_1 \cdot x \\ w_2 \cdot x \\ \vdots \\ w_m \cdot x \end{bmatrix}

The calculation of Row 1 (w1⋅xw_1 \cdot x) does not rely on the result of Row 2 (w2⋅xw_2 \cdot x) or Row 3 (w3⋅xw_3 \cdot x). All mm row dot products are completely independent.

A standard CPU contains a small number of powerful, general-purpose cores engineered to execute instructions sequentially: compute Row 1, then Row 2, then Row 3. A GPU contains thousands of smaller, specialized arithmetic cores engineered to execute multiplications and additions across all rows simultaneously in the exact same physical clock cycle.

To visualize how hardware architectures execute matrix operations differently, compare sequential CPU dot products with simultaneous GPU core execution below:

Hardware Execution: CPU vs GPU Matrix Multiplication
Hardware execution comparison: A CPU executes independent row dot products sequentially one after another across general-purpose cores, whereas a GPU distributes row dot products across parallel arithmetic cores to compute the entire output vector simultaneously in a single clock cycle.

Because every row dot product in WxW x operates on the exact same input vector xx without waiting on neighboring rows, a GPU evaluates the entire matrix transformation in the time it takes to compute a single dot product. This mathematical independence is the foundational reason neural networks train and scale on GPU hardware.

TEASER: Production Scale: From 3 Neurons to Thousands of Dimensions

In our scaled-down studio predictor, the matrix has m=3m = 3 rows evaluating n=4n = 4 features. In modern foundation models, weight matrices often contain m=4,096m = 4{,}096 or m=12,288m = 12{,}288 rows multiplying vectors across thousands of latent dimensions. Because all rows evaluate in parallel, hardware accelerators process millions of parameter interactions across a layer in a single forward pass.


Previous
Matrix Dimensions and Parallel Vectors