The Forward-Backward Computational Graph hero
Lesson 1The Complete Neural Data Flow Graph

The Forward-Backward Computational Graph

Unify the forward pass, loss calculation, and backward backpropagation into a single directed acyclic graph and verify the closed-loop trace.

In Module 1, we constructed the forward pass of a Multi-Layer Perceptron. We traced how input feature vectors xx combine with weight matrices WW and bias vectors bb to form affine sums zz, which compress through non-linear activation functions ff and σ\sigma across stacked layers to produce a final scalar prediction y^\hat{y}.

In Module 2, we constructed the backward pass. We quantified prediction mistakes using Mean Squared Error loss LL, isolated parameter sensitivities with single-variable derivatives and multivariable gradients ∇L\nabla L, propagated error deltas δ\delta backward through transposed weight matrices via the composite chain rule, and adjusted parameters using a single gradient descent update step.

[ Module 1: The Forward Pass ]
  Topic 1: Feature Vectors & Dot Products (x, w · x)
  Topic 2: Affine Sums & Activations (z = w · x + b, a = σ(z))
  Topic 3: Parallel Layers & Width (z = W x + b, a = σ(z))
  Topic 4: Multi-Layer Perceptrons & Depth (a^(1) = ReLU(W^(1)x + b^(1)), a^(2) = σ(W^(2)a^(1) + b^(2)))
                                │
                                ▼
[ Module 2: The Backward Pass ]
  Topic 1: Prediction Error & Squared Loss (L = (y - y_hat)^2, ∂L/∂y_hat = 2(y_hat - y))
  Topic 2: Single-Variable Derivatives (dy/dx, tangent slopes, power and sum rules)
  Topic 3: Multivariable Gradients (∂L/∂w_i, ∇_w L, steepest descent -∇L)
  Topic 4: The Chain Rule & Multilayer Backpropagation (δ^(2), (W^(2))^T δ^(2), δ^(1), W_new = W - η ∇L)
                                │
                                ▼
[ Module 3: Synthesis & Interpretability ]
  Topic 1: The Complete Neural Data Flow Graph ◄─── [THIS TOPIC]
  Topic 2: Interpretability & The Road to LLMs

In isolation, these forward and backward steps can appear as separate algebraic procedures. In this topic, we synthesize them into a single, unified mathematical structure: the Directed Acyclic Graph (DAG).

By assembling every forward transformation and backward gradient calculation into an unbroken computational pipeline, we complete the closed learning cycle of deep learning on paper.



Map the unified computational graph connecting scalars, matrices, hidden activations, loss penalties, and backward gradient flows into a single DAG.

Every neural network computation is formalizable as a Directed Acyclic Graph G=(V,E)\mathcal{G} = (V, E), where:

  • The Vertices (nodes VV) represent tensors and scalar values: input features, parameter matrices, intermediate affine pre-activations, post-activation hidden states, output predictions, loss penalties, and gradient tensors.
  • The Directed Edges (EE) represent deterministic mathematical operations: matrix multiplication, vector addition, element-wise non-linear activations, difference squaring, and derivative chain multiplications.

The graph is directed because data and sensitivities flow along defined orientations. It is acyclic because computation advances in topological stages without circular recursion during a single evaluation pass.


The Unbroken Topological Sequence

When a neural network evaluates an input and updates its parameters from the resulting error, computation traverses an unbroken topological sequence of 10 core mathematical stages:

x⟶z(1)⟶a(1)⟶z(2)⟶y^⟶L⟶δ(2)⟶∇W(2)L⟶δ(1)⟶∇W(1)L⟶Wnewx \longrightarrow z^{(1)} \longrightarrow a^{(1)} \longrightarrow z^{(2)} \longrightarrow \hat{y} \longrightarrow L \longrightarrow \delta^{(2)} \longrightarrow \nabla_{W^{(2)}} L \longrightarrow \delta^{(1)} \longrightarrow \nabla_{W^{(1)}} L \longrightarrow W_{\text{new}}
══════════════════════════════════════════════════════════════════════════════════════════════════════
                               THE COMPLETE NEURAL DATA FLOW GRAPH (DAG)
══════════════════════════════════════════════════════════════════════════════════════════════════════

[ INPUT STAGE ]
      x (d_0 × 1) ───────────┐
                             │
[ LAYER 1 FORWARD ]          ▼
      W^(1) (d_1 × d_0) ──► [ Affine Sum: z^(1) = W^(1)x + b^(1) ] ──► [ Activation: a^(1) = ReLU(z^(1)) ]
      b^(1) (d_1 × 1)   ──►                                                    │
                                                                               ▼
[ LAYER 2 FORWARD ]                                                    [ Affine Sum: z^(2) = W^(2)a^(1) + b^(2) ]
      W^(2) (d_2 × d_1) ─────────────────────────────────────────────►         │
      b^(2) (d_2 × 1)   ─────────────────────────────────────────────►         ▼
                                                                       [ Activation: a^(2) = σ(z^(2)) = y_hat ]
                                                                               │
[ LOSS EVALUATION ]                                                            ▼
      Ground Truth y (d_2 × 1) ──────────────────────────────────────► [ Squared Loss: L = (y_hat - y)^2 ]
                                                                               │
───────────────────────────────────────────────────────────────────────────────┼──────────────────────
[ LAYER 2 BACKWARD ]                                                           ▼
      ∂L/∂W^(2) = δ^(2)(a^(1))^T ◄── [ Output Delta: δ^(2) = 2(y_hat - y) ⊙ σ'(z^(2)) ]
      ∂L/∂b^(2) = δ^(2)          ◄──                                           │
                                                                               ▼  (Transposed Projection: (W^(2))^T δ^(2))
[ LAYER 1 BACKWARD ]                                                           │
      ∂L/∂W^(1) = δ^(1) x^T      ◄── [ Hidden Delta: δ^(1) = ((W^(2))^T δ^(2)) ⊙ ReLU'(z^(1)) ]
      ∂L/∂b^(1) = δ^(1)          ◄──
                                 │
[ PARAMETER UPDATES ]            ▼
      W^(l)_new = W^(l)_old - η (∂L/∂W^(l))
      b^(l)_new = b^(l)_old - η (∂L/∂b^(l))
══════════════════════════════════════════════════════════════════════════════════════════════════════

The Squeeze-and-Expand Geometry of the Neural Graph

Examining the graph reveals a distinct structural geometry:

1. The Forward Contraction and Expansion

  • Input to Latent Space (d0→d1d_0 \to d_1): High-dimensional raw input measurements (x∈Rd0x \in \mathbb{R}^{d_0}) are linearly projected and transformed into hidden representation coordinates (a(1)∈Rd1a^{(1)} \in \mathbb{R}^{d_1}).
  • Latent Space to Decision Bottleneck (d1→d2d_1 \to d_2): Intermediate concepts are compressed through the output layer weights into a low-dimensional decision (z(2),a(2)∈Rd2z^{(2)}, a^{(2)} \in \mathbb{R}^{d_2}). For single-target classification, d2=1d_2 = 1.
  • Decision to Scalar Penalty (d2→1d_2 \to 1): The output prediction y^\hat{y} is compared against target yy to collapse all network behavior into a single positive scalar loss penalty L∈RL \in \mathbb{R}.

2. The Backward Mirroring and Credit Fan-Out

  • Scalar Perturbation to Output Delta (1→d21 \to d_2): The backward pass begins at the scalar loss LL, computing the initial sensitivity ∂L∂y^\frac{\partial L}{\partial \hat{y}} and gating it with output activation slope σ′(z(2))\sigma'(z^{(2)}) to produce error delta δ(2)∈Rd2×1\delta^{(2)} \in \mathbb{R}^{d_2 \times 1}.
  • Transposed Back-Projection (d2→d1d_2 \to d_1): The error delta is projected backward through the transposed output weight matrix (W(2))T∈Rd1×d2(W^{(2)})^T \in \mathbb{R}^{d_1 \times d_2} and modulated by hidden activation slopes f′(z(1))f'(z^{(1)}) to form hidden error delta δ(1)∈Rd1×1\delta^{(1)} \in \mathbb{R}^{d_1 \times 1}.
  • Outer-Product Gradient Fan-Out (d1→d1×d0d_1 \to d_1 \times d_0): Multiplying hidden delta δ(1)\delta^{(1)} by transposed input vector xT∈R1×d0x^T \in \mathbb{R}^{1 \times d_0} fans the error signal out into the full parameter gradient matrix ∂L∂W(1)∈Rd1×d0\frac{\partial L}{\partial W^{(1)}} \in \mathbb{R}^{d_1 \times d_0}.

Every parameter matrix in the forward pass has an exact companion gradient matrix in the backward pass with identical shape:

dim(∂L∂W(l))=dim(W(l))∈Rdout×din\text{dim}\left(\frac{\partial L}{\partial W^{(l)}}\right) = \text{dim}\left(W^{(l)}\right) \in \mathbb{R}^{d_{\text{out}} \times d_{\text{in}}} dim(∂L∂b(l))=dim(b(l))∈Rdout×1\text{dim}\left(\frac{\partial L}{\partial b^{(l)}}\right) = \text{dim}\left(b^{(l)}\right) \in \mathbb{R}^{d_{\text{out}} \times 1}

Inference vs. Learning in the Data Flow Graph

The computational graph highlights the exact mathematical distinction between running a model and training a model:

  • Inference (One-Way Forward Execution): Data flows strictly downstream along the forward subgraph: x→z(1)→a(1)→z(2)→y^x \to z^{(1)} \to a^{(1)} \to z^{(2)} \to \hat{y}. Parameters WW and bb remain static constants. No loss function is required, and no gradients are computed.
  • Learning (The Closed Topological Cycle): Forward propagation produces y^\hat{y}; loss evaluation measures mistake penalty LL; backpropagation reverses through the graph to compute ∇WL\nabla_W L and ∇bL\nabla_b L; and gradient descent updates parameters Wnew=W−η∇WLW_{\text{new}} = W - \eta \nabla_W L.

When the updated parameters are plugged back into the top of the graph, the network executes a new forward pass with reduced error. Learning is the execution of this closed loop.


Previous
The Chain Rule & Backpropagation In Practice