
The Forward-Backward Computational Graph
Unify the forward pass, loss calculation, and backward backpropagation into a single directed acyclic graph and verify the closed-loop trace.
In Module 1, we constructed the forward pass of a Multi-Layer Perceptron. We traced how input feature vectors combine with weight matrices and bias vectors to form affine sums , which compress through non-linear activation functions and across stacked layers to produce a final scalar prediction .
In Module 2, we constructed the backward pass. We quantified prediction mistakes using Mean Squared Error loss , isolated parameter sensitivities with single-variable derivatives and multivariable gradients , propagated error deltas backward through transposed weight matrices via the composite chain rule, and adjusted parameters using a single gradient descent update step.
[ Module 1: The Forward Pass ]
Topic 1: Feature Vectors & Dot Products (x, w · x)
Topic 2: Affine Sums & Activations (z = w · x + b, a = σ(z))
Topic 3: Parallel Layers & Width (z = W x + b, a = σ(z))
Topic 4: Multi-Layer Perceptrons & Depth (a^(1) = ReLU(W^(1)x + b^(1)), a^(2) = σ(W^(2)a^(1) + b^(2)))
│
▼
[ Module 2: The Backward Pass ]
Topic 1: Prediction Error & Squared Loss (L = (y - y_hat)^2, ∂L/∂y_hat = 2(y_hat - y))
Topic 2: Single-Variable Derivatives (dy/dx, tangent slopes, power and sum rules)
Topic 3: Multivariable Gradients (∂L/∂w_i, ∇_w L, steepest descent -∇L)
Topic 4: The Chain Rule & Multilayer Backpropagation (δ^(2), (W^(2))^T δ^(2), δ^(1), W_new = W - η ∇L)
│
▼
[ Module 3: Synthesis & Interpretability ]
Topic 1: The Complete Neural Data Flow Graph ◄─── [THIS TOPIC]
Topic 2: Interpretability & The Road to LLMs
In isolation, these forward and backward steps can appear as separate algebraic procedures. In this topic, we synthesize them into a single, unified mathematical structure: the Directed Acyclic Graph (DAG).
By assembling every forward transformation and backward gradient calculation into an unbroken computational pipeline, we complete the closed learning cycle of deep learning on paper.
Map the unified computational graph connecting scalars, matrices, hidden activations, loss penalties, and backward gradient flows into a single DAG.
Every neural network computation is formalizable as a Directed Acyclic Graph , where:
- The Vertices (nodes ) represent tensors and scalar values: input features, parameter matrices, intermediate affine pre-activations, post-activation hidden states, output predictions, loss penalties, and gradient tensors.
- The Directed Edges () represent deterministic mathematical operations: matrix multiplication, vector addition, element-wise non-linear activations, difference squaring, and derivative chain multiplications.
The graph is directed because data and sensitivities flow along defined orientations. It is acyclic because computation advances in topological stages without circular recursion during a single evaluation pass.
The Unbroken Topological Sequence
When a neural network evaluates an input and updates its parameters from the resulting error, computation traverses an unbroken topological sequence of 10 core mathematical stages:
══════════════════════════════════════════════════════════════════════════════════════════════════════
THE COMPLETE NEURAL DATA FLOW GRAPH (DAG)
══════════════════════════════════════════════════════════════════════════════════════════════════════
[ INPUT STAGE ]
x (d_0 × 1) ───────────┐
│
[ LAYER 1 FORWARD ] ▼
W^(1) (d_1 × d_0) ──► [ Affine Sum: z^(1) = W^(1)x + b^(1) ] ──► [ Activation: a^(1) = ReLU(z^(1)) ]
b^(1) (d_1 × 1) ──► │
▼
[ LAYER 2 FORWARD ] [ Affine Sum: z^(2) = W^(2)a^(1) + b^(2) ]
W^(2) (d_2 × d_1) ─────────────────────────────────────────────► │
b^(2) (d_2 × 1) ─────────────────────────────────────────────► ▼
[ Activation: a^(2) = σ(z^(2)) = y_hat ]
│
[ LOSS EVALUATION ] ▼
Ground Truth y (d_2 × 1) ──────────────────────────────────────► [ Squared Loss: L = (y_hat - y)^2 ]
│
───────────────────────────────────────────────────────────────────────────────┼──────────────────────
[ LAYER 2 BACKWARD ] ▼
∂L/∂W^(2) = δ^(2)(a^(1))^T ◄── [ Output Delta: δ^(2) = 2(y_hat - y) ⊙ σ'(z^(2)) ]
∂L/∂b^(2) = δ^(2) ◄── │
▼ (Transposed Projection: (W^(2))^T δ^(2))
[ LAYER 1 BACKWARD ] │
∂L/∂W^(1) = δ^(1) x^T ◄── [ Hidden Delta: δ^(1) = ((W^(2))^T δ^(2)) ⊙ ReLU'(z^(1)) ]
∂L/∂b^(1) = δ^(1) ◄──
│
[ PARAMETER UPDATES ] ▼
W^(l)_new = W^(l)_old - η (∂L/∂W^(l))
b^(l)_new = b^(l)_old - η (∂L/∂b^(l))
══════════════════════════════════════════════════════════════════════════════════════════════════════
The Squeeze-and-Expand Geometry of the Neural Graph
Examining the graph reveals a distinct structural geometry:
1. The Forward Contraction and Expansion
- Input to Latent Space (): High-dimensional raw input measurements () are linearly projected and transformed into hidden representation coordinates ().
- Latent Space to Decision Bottleneck (): Intermediate concepts are compressed through the output layer weights into a low-dimensional decision (). For single-target classification, .
- Decision to Scalar Penalty (): The output prediction is compared against target to collapse all network behavior into a single positive scalar loss penalty .
2. The Backward Mirroring and Credit Fan-Out
- Scalar Perturbation to Output Delta (): The backward pass begins at the scalar loss , computing the initial sensitivity and gating it with output activation slope to produce error delta .
- Transposed Back-Projection (): The error delta is projected backward through the transposed output weight matrix and modulated by hidden activation slopes to form hidden error delta .
- Outer-Product Gradient Fan-Out (): Multiplying hidden delta by transposed input vector fans the error signal out into the full parameter gradient matrix .
Every parameter matrix in the forward pass has an exact companion gradient matrix in the backward pass with identical shape:
Inference vs. Learning in the Data Flow Graph
The computational graph highlights the exact mathematical distinction between running a model and training a model:
- Inference (One-Way Forward Execution): Data flows strictly downstream along the forward subgraph: . Parameters and remain static constants. No loss function is required, and no gradients are computed.
- Learning (The Closed Topological Cycle): Forward propagation produces ; loss evaluation measures mistake penalty ; backpropagation reverses through the graph to compute and ; and gradient descent updates parameters .
When the updated parameters are plugged back into the top of the graph, the network executes a new forward pass with reduced error. Learning is the execution of this closed loop.