
Elementary Power and Sum Derivative Rules
Apply foundational power, constant, and sum rules of differential calculus to compute exact algebraic derivatives for single-variable functions.
Evaluating limits from first principles () builds foundational understanding, but doing this by hand for every equation would be slow.
Calculus provides a set of universal algebraic derivative rules that allow us to differentiate functions directly.
Core Elementary Rules of Calculus
Four fundamental rules form the basis of single-variable differentiation:
1. The Constant Rule
The derivative of any constant number is zero:
- Intuition: A constant value never changes. If , then nudging has zero effect on ().
- Examples:
2. The Linear Multiplier Rule
The derivative of a linear term is simply the constant coefficient :
- Intuition: A straight line has a fixed slope of at every point.
- Examples:
3. The Power Rule
For any variable raised to a constant power , multiply by the exponent and subtract from the power:
When the power term is scaled by a constant multiplier :
Let's step through standard applications of the power rule:
-
For :
-
For :
-
For :
-
For :
-
For :
-
For negative powers ():
4. The Sum and Difference Rule
The derivative of a sum or difference of two functions is the sum or difference of their individual derivatives:
- Intuition: The total rate of change of combined components equals the sum of their individual rates of change.
Step-by-Step Derivation of a Quadratic Polynomial
Let's apply these rules to differentiate a complete quadratic polynomial:
We differentiate term-by-term using our algebraic rules:
-
First term (): Apply the power rule:
-
Second term (): Apply the linear multiplier rule:
-
Third term (): Apply the constant rule:
Combining the results using the sum rule gives the complete derivative:
Finding the Minimum Vertex Analytically
To find the exact coordinate where this quadratic curve reaches its lowest point, set the derivative to zero ():
Evaluate the function value at this minimum:
Let's check the slopes around this vertex:
- At : (downhill slope).
- At : (flat valley floor).
- At : (uphill slope).
Notice that . This is structurally identical to a Mean Squared Error loss curve where the ground-truth target is .
Re-Deriving the Mean Squared Error Loss Derivative
In Topic 1, we calculated the derivative of Mean Squared Error with respect to the prediction for a fixed target :
Let's differentiate with respect to using our power and sum rules:
-
Differentiate with respect to :
-
Differentiate with respect to (treating as a constant target):
-
Differentiate with respect to (since target is a constant number, is also constant):
Summing the three terms yields:
The elementary power and sum rules produce the exact output loss derivative in a single algebraic step.
The Bridge to Topic 3: From Single Variables to Parameter Vectors
In this topic, we have mastered single-variable calculus: computing how an output responds to nudges in a single input ().
In an artificial neural network, decisions do not depend on a single variable. A single artificial neuron computes a linear combination of multiple inputs and weights:
When 'Die Hard in Space' flopped, the mistake was caused by multiple weights acting together in the network.
How do we measure the sensitivity of the loss to weight when weights are also changing?
In Topic 3, we will extend our calculus toolkit from single-variable derivatives to Partial Derivatives () and assemble individual sensitivities into the unified Gradient Vector ().