Elementary Power and Sum Derivative Rules hero
Lesson 4Derivatives and Sensitivity

Elementary Power and Sum Derivative Rules

Apply foundational power, constant, and sum rules of differential calculus to compute exact algebraic derivatives for single-variable functions.

Evaluating limits from first principles (lim⁡Δx→0f(x+Δx)−f(x)Δx\lim_{\Delta x \to 0} \frac{f(x + \Delta x) - f(x)}{\Delta x}) builds foundational understanding, but doing this by hand for every equation would be slow.

Calculus provides a set of universal algebraic derivative rules that allow us to differentiate functions directly.


Core Elementary Rules of Calculus

Four fundamental rules form the basis of single-variable differentiation:

1. The Constant Rule

The derivative of any constant number is zero:

ddx[c]=0(where c∈R is constant)\frac{d}{dx}[c] = 0 \quad (\text{where } c \in \mathbb{R} \text{ is constant})
  • Intuition: A constant value never changes. If y=7y = 7, then nudging xx has zero effect on yy (Δy=0\Delta y = 0).
  • Examples: ddx[5]=0,ddx[−12.8]=0,ddx[π]=0\frac{d}{dx}[5] = 0, \qquad \frac{d}{dx}[-12.8] = 0, \qquad \frac{d}{dx}[\pi] = 0

2. The Linear Multiplier Rule

The derivative of a linear term axax is simply the constant coefficient aa:

ddx[ax]=a(where a∈R)\frac{d}{dx}[ax] = a \quad (\text{where } a \in \mathbb{R})
  • Intuition: A straight line y=axy = ax has a fixed slope of aa at every point.
  • Examples: ddx[15x]=15,ddx[−4.5x]=−4.5,ddx[x]=1\frac{d}{dx}[15x] = 15, \qquad \frac{d}{dx}[-4.5x] = -4.5, \qquad \frac{d}{dx}[x] = 1

3. The Power Rule

For any variable raised to a constant power nn, multiply by the exponent nn and subtract 11 from the power:

ddx[xn]=nxn−1(n∈R)\frac{d}{dx}[x^n] = n x^{n-1} \quad (n \in \mathbb{R})

When the power term is scaled by a constant multiplier cc:

ddx[c⋅xn]=c⋅nxn−1\frac{d}{dx}[c \cdot x^n] = c \cdot n x^{n-1}

Let's step through standard applications of the power rule:

  • For x1x^1: ddx[x1]=1⋅x1−1=1⋅x0=1\frac{d}{dx}[x^1] = 1 \cdot x^{1-1} = 1 \cdot x^0 = \mathbf{1}

  • For x2x^2: ddx[x2]=2⋅x2−1=2x\frac{d}{dx}[x^2] = 2 \cdot x^{2-1} = \mathbf{2x}

  • For x3x^3: ddx[x3]=3⋅x3−1=3x2\frac{d}{dx}[x^3] = 3 \cdot x^{3-1} = \mathbf{3x^2}

  • For 4x34x^3: ddx[4x3]=4⋅(3x2)=12x2\frac{d}{dx}[4x^3] = 4 \cdot (3x^2) = \mathbf{12x^2}

  • For 0.5x20.5x^2: ddx[0.5x2]=0.5⋅(2x)=1.0x=x\frac{d}{dx}[0.5x^2] = 0.5 \cdot (2x) = \mathbf{1.0x} = x

  • For negative powers (x−1=1xx^{-1} = \frac{1}{x}): ddx[x−1]=−1⋅x−1−1=−1⋅x−2=−1x2\frac{d}{dx}[x^{-1}] = -1 \cdot x^{-1-1} = -1 \cdot x^{-2} = \mathbf{-\frac{1}{x^2}}


4. The Sum and Difference Rule

The derivative of a sum or difference of two functions is the sum or difference of their individual derivatives:

ddx[f(x)±g(x)]=ddx[f(x)]±ddx[g(x)]=f′(x)±g′(x)\frac{d}{dx}[f(x) \pm g(x)] = \frac{d}{dx}[f(x)] \pm \frac{d}{dx}[g(x)] = f'(x) \pm g'(x)
  • Intuition: The total rate of change of combined components equals the sum of their individual rates of change.

Step-by-Step Derivation of a Quadratic Polynomial

Let's apply these rules to differentiate a complete quadratic polynomial:

f(x)=x2−4x+4f(x) = x^2 - 4x + 4

We differentiate term-by-term using our algebraic rules:

  1. First term (x2x^2): Apply the power rule: ddx[x2]=2x\frac{d}{dx}[x^2] = 2x

  2. Second term (−4x-4x): Apply the linear multiplier rule: ddx[−4x]=−4\frac{d}{dx}[-4x] = -4

  3. Third term (+4+4): Apply the constant rule: ddx[4]=0\frac{d}{dx}[4] = 0

Combining the results using the sum rule gives the complete derivative:

f′(x)=2x−4+0=2x−4f'(x) = 2x - 4 + 0 = \mathbf{2x - 4}

Finding the Minimum Vertex Analytically

To find the exact coordinate where this quadratic curve reaches its lowest point, set the derivative to zero (f′(x)=0f'(x) = 0):

2x−4=0  ⟹  2x=4  ⟹  x=22x - 4 = 0 \implies 2x = 4 \implies \mathbf{x = 2}

Evaluate the function value at this minimum:

f(2)=(2)2−4(2)+4=4−8+4=0f(2) = (2)^2 - 4(2) + 4 = 4 - 8 + 4 = \mathbf{0}

Let's check the slopes around this vertex:

  • At x=0x = 0: f′(0)=2(0)−4=−4.0f'(0) = 2(0) - 4 = \mathbf{-4.0} (downhill slope).
  • At x=2x = 2: f′(2)=2(2)−4=0.0f'(2) = 2(2) - 4 = \mathbf{0.0} (flat valley floor).
  • At x=4x = 4: f′(4)=2(4)−4=+4.0f'(4) = 2(4) - 4 = \mathbf{+4.0} (uphill slope).

Notice that f(x)=x2−4x+4=(x−2)2f(x) = x^2 - 4x + 4 = (x - 2)^2. This is structurally identical to a Mean Squared Error loss curve L(y^)=(y^−2)2L(\hat{y}) = (\hat{y} - 2)^2 where the ground-truth target is y=2y = 2.


Re-Deriving the Mean Squared Error Loss Derivative

In Topic 1, we calculated the derivative of Mean Squared Error with respect to the prediction y^\hat{y} for a fixed target yy:

L(y^)=(y^−y)2=y^2−2yy^+y2L(\hat{y}) = (\hat{y} - y)^2 = \hat{y}^2 - 2y\hat{y} + y^2

Let's differentiate LL with respect to y^\hat{y} using our power and sum rules:

  1. Differentiate y^2\hat{y}^2 with respect to y^\hat{y}: ddy^[y^2]=2y^\frac{d}{d\hat{y}}[\hat{y}^2] = 2\hat{y}

  2. Differentiate −2yy^-2y\hat{y} with respect to y^\hat{y} (treating yy as a constant target): ddy^[−2yy^]=−2y\frac{d}{d\hat{y}}[-2y\hat{y}] = -2y

  3. Differentiate y2y^2 with respect to y^\hat{y} (since target yy is a constant number, y2y^2 is also constant): ddy^[y2]=0\frac{d}{d\hat{y}}[y^2] = 0

Summing the three terms yields:

dLdy^=2y^−2y=2(y^−y)\frac{dL}{d\hat{y}} = 2\hat{y} - 2y = \mathbf{2(\hat{y} - y)}

The elementary power and sum rules produce the exact output loss derivative in a single algebraic step.


The Bridge to Topic 3: From Single Variables to Parameter Vectors

In this topic, we have mastered single-variable calculus: computing how an output yy responds to nudges in a single input xx (dydx\frac{dy}{dx}).

In an artificial neural network, decisions do not depend on a single variable. A single artificial neuron computes a linear combination of multiple inputs and weights:

z=w1x1+w2x2+w3x3+w4x4+bz = w_1 x_1 + w_2 x_2 + w_3 x_3 + w_4 x_4 + b

When 'Die Hard in Space' flopped, the mistake was caused by multiple weights acting together in the network.

How do we measure the sensitivity of the loss to weight w1w_1 when weights w2,w3,w4w_2, w_3, w_4 are also changing?

In Topic 3, we will extend our calculus toolkit from single-variable derivatives to Partial Derivatives (∂L∂wi\frac{\partial L}{\partial w_i}) and assemble individual sensitivities into the unified Gradient Vector (∇L\nabla L).


Previous
The Training Wheels Transition in Calculus