(Verification via expansion: y=(2x−4)3=8x3−48x2+96x−64⟹dxdy=24x2−96x+96. At x=3: 24(9)−96(3)+96=216−288+96=24.0.)
Problem 2: Sigmoid Activation Derivative Calculation
A neuron produces a post-activation output a=σ(z)=0.80. Compute the exact instantaneous activation slope σ′(z)=σ(z)(1−σ(z)).
Reveal Solution
σ′(z)=σ(z)(1−σ(z))=0.80×(1−0.80)=0.80×0.20=0.1600
Problem 3: Output Layer Error Delta δ(2) Computation
An output neuron has loss derivative ∂a(2)∂L=1.60 and activation slope σ′(z(2))=0.25. Compute the output error delta δ(2)=∂a(2)∂L⋅σ′(z(2)).
Reveal Solution
δ(2)=∂a(2)∂L⋅σ′(z(2))=1.60×0.25=0.4000
Problem 4: Hidden Layer Error Delta via Transposed Weights
An output layer error delta is δ(2)=0.40, connected via output weight W(2)=0.50 to a hidden neuron with activation slope σ′(z(1))=0.20. Compute the hidden error delta δ(1)=((W(2))Tδ(2))⋅σ′(z(1)).
Reveal Solution
Project error through transposed weight:
(W(2))Tδ(2)=0.50×0.40=0.2000
Modulate by hidden activation slope:
δ(1)=0.2000×0.20=0.0400
Problem 5: Single Parameter Update Step Calculation
A weight parameter has current value wold=0.60, computed gradient ∂w∂L=0.40, and step size factor η=0.10. Calculate the updated weight wnew=wold−η∂w∂L.
Part 2: Applied Scenario: The VC 2-Layer Backpropagation Trace
An automated venture capital risk model uses a 2→2→1 Multi-Layer Perceptron to predict whether an early-stage startup will succeed (y=1) or go bankrupt (y=0).
System Configuration:
Input Feature Vector:x=[1.02.0](Founding Team Experience)(Market Size)
Activation Functions: Standard Sigmoid activation σ(z)=1+e−z1 across all hidden and output neurons.
Ground-Truth Target:y=0 (The startup went bankrupt).
Step Size (Learning Rate):η=0.50.
Problem 6: Complete 2-Layer VC Backward Pass Audit
Part A (Forward Pass & Initial Loss): Compute z(1),a(1),z(2),a(2)=y^, and initial MSE loss L1=(a(2)−y)2.
Part B (Output Layer Error Attribution): Compute ∂a(2)∂L, σ′(z(2)), output error delta δ(2), output weight gradient matrix ∂W(2)∂L, and output bias gradient ∂b(2)∂L.
Part C (Hidden Layer Backpropagation): Compute backpropagated error signal (W(2))Tδ(2), hidden activation slopes σ′(z(1)), hidden error delta vector δ(1), hidden weight gradient matrix ∂W(1)∂L, and hidden bias gradient ∂b(1)∂L.
Part D (Single-Step Parameter Update): Compute the updated weight W11,new(1) using step size η=0.50, as well as the full updated parameter matrices Wnew(1),bnew(1),Wnew(2),bnew(2).
Part E (Second Forward Pass & Mathematical Proof): Re-evaluate the complete forward pass with all updated parameters, compute updated loss L2, and prove that L2<L1 (ΔL<0).
Part F (Venture Synthesis): In 2–3 sentences, explain why the gradient for Market Size (W12(1)→0.0440) is exactly twice as large as Team Experience (W11(1)→0.0220), and how backpropagation automatically assigns larger corrective penalties to larger input signals.
Full Parameter Updates:Wnew(1)=[0.50.50.50.5]−0.5[0.0219840.0219840.0439670.043967]=[0.4890080.4890080.4780160.478016]bnew(1)=[0.00.0]−0.5[0.0219840.021984]=[−0.010992−0.010992]Wnew(2)=[0.50.5]−0.5[0.2410150.241015]=[0.3794920.379492]bnew(2)=0.0−0.5(0.294793)=−0.147396
Part E: Second Forward Pass and Mathematical Loss Reduction Proof
Updated Loss Evaluation (L2):L2=(a(2)new−y)2=(0.614320−0.0)2≈0.377389
Proof of Strict Loss Reduction:ΔL=L2−L1=0.377389−0.481249=−0.103860<0Relative Error Reduction=0.4812490.103860×100%=21.58%
The parameter update step reduced prediction loss by 21.58%, pulling the network's failure probability estimate from 69.37% down to 61.43%.
Part F: Venture Synthesis
The gradient for Market Size (W12(1)→0.0440) is exactly twice as large as Team Experience (W11(1)→0.0220) because the input measurement for Market Size was twice as large (x2=2.0 vs x1=1.0).
Because the parameter gradient is given by ∂W(1)∂L=δ(1)xT, features with larger input magnitudes amplify the error signal and receive proportionately larger gradient penalties during backpropagation.