
Matrix-Vector Multiplication
Matrix-vector multiplication executed as parallel row dot products to transform input feature coordinates into multi-output decision vectors.
In Lesson 1, we stacked the weights of three studio neurons into the horizontal rows of a single weight matrix .
When an incoming script's feature vector arrives at a neural network layer, the model must evaluate all of those rows against that single input simultaneously. That operation is matrix-vector multiplication.
Despite the formal name, matrix-vector multiplication is not a new branch of arithmetic. It is simply multiple dot products executed in parallel.
The Row-by-Column Dot Product Rule
To calculate the -th entry of the output vector , we compute the dot product between the -th row of the weight matrix and the input vector :
Writing the full operation in expanded algebraic form:
Every row of the matrix sweeps horizontally across its columns, multiplying each weight by the corresponding feature value , and sums the individual products into a single scalar value.
Walk-Through: Calculating Matrix Scores for 'Die Hard in Space'
Now let's execute this exact operation with concrete numbers using our running studio predictor from Lesson 1.
Recall our Weight Matrix (), where each row provides the weights for one neuron across the locked genre order [Action, Romance, Comedy, Sci-Fi]:
Our incoming script feature vector for 'Die Hard in Space' is:
Multiplying the weight matrix by the input vector :
We evaluate each row dot product step by step:
- Row 1 (Hit):
- Row 2 (Oscar):
- Row 3 (Franchise):
The resulting output is a 3-dimensional column vector holding the raw decision scores:
We can also inspect the arithmetic in a structured feature ledger, keeping incoming features stacked along the vertical axis:
| Feature Dimension | Feature Value () | Row 1: Hit () | Row 2: Oscar () | Row 3: Franchise () |
|---|---|---|---|---|
| Action () | ||||
| Romance () | ||||
| Comedy () | ||||
| Sci-Fi () | ||||
| Raw Dot Product () |
Notice that Row 1 reproduces the exact raw score of calculated for 'Die Hard in Space' by our single artificial neuron in Topics 1 and 2. Stacking vectors into a matrix did not change the arithmetic—it packaged three independent evaluation channels into a single mathematical operation.
The Dimension Compatibility Rule
Not every matrix can multiply every vector. The physical setup dictates strict dimensional alignment.
In our studio predictor, 'Die Hard in Space' supplies 4 genre features (). For each row in the matrix to compute a dot product with that script, it must provide exactly 4 weights—one for each genre. If a matrix had only 3 columns, there would be no weight to evaluate Sci-Fi. If it had 5 columns, the 5th weight would have no incoming feature to pair with.
Therefore, the number of columns in the weight matrix must equal the number of rows in the input vector.
Likewise, because each row in represents one neuron, stacking rows produces exactly individual dot products—a column vector of length .
In formal linear algebra notation, the Dimension Compatibility Rule is written:
This rule establishes two structural constraints:
- The Inner Dimensions Must Match (): The number of columns in () must equal the number of rows in (). Every incoming feature coordinate must have a corresponding weight parameter in each row.
- The Outer Dimensions Determine the Output Shape (): The number of rows in () determines the number of entries in the resulting output vector. Stacking neurons produces an -dimensional column vector.
A matrix acts as a dimensional transformation machine. It accepts an -dimensional input vector and transforms it into an -dimensional output vector:
Depending on the chosen dimensions, a matrix transformation alters coordinate representations in one of three ways:
- Dimensional Expansion (): Expands the representation into a higher-dimensional space with more output neurons than incoming features.
- Dimensional Compression (): Compresses the representation into a lower-dimensional space, summarizing inputs into fewer coordinates.
- Dimensional Preservation (): Maps features into an output space of the exact same dimensionality.
Hardware Parallelism: Why AI Runs on GPUs
Understanding matrix-vector multiplication as parallel row dot products reveals why modern deep learning relies on Graphics Processing Units (GPUs) rather than Central Processing Units (CPUs).
In the output vector:
The calculation of Row 1 () does not rely on the result of Row 2 () or Row 3 (). All row dot products are completely independent.
A standard CPU contains a small number of powerful, general-purpose cores engineered to execute instructions sequentially: compute Row 1, then Row 2, then Row 3. A GPU contains thousands of smaller, specialized arithmetic cores engineered to execute multiplications and additions across all rows simultaneously in the exact same physical clock cycle.
To visualize how hardware architectures execute matrix operations differently, compare sequential CPU dot products with simultaneous GPU core execution below:

Because every row dot product in operates on the exact same input vector without waiting on neighboring rows, a GPU evaluates the entire matrix transformation in the time it takes to compute a single dot product. This mathematical independence is the foundational reason neural networks train and scale on GPU hardware.
TEASER: Production Scale: From 3 Neurons to Thousands of Dimensions
In our scaled-down studio predictor, the matrix has rows evaluating features. In modern foundation models, weight matrices often contain or rows multiplying vectors across thousands of latent dimensions. Because all rows evaluate in parallel, hardware accelerators process millions of parameter interactions across a layer in a single forward pass.