The Rectified Linear Unit Activation hero
Lesson 4Multi-Layer Neural Networks

The Rectified Linear Unit Activation

Apply the piecewise linear ReLU function to introduce non-linearity while preserving identity gradients for positive activation values in neurons.

The Sigmoid function provides bounded probabilities between 00 and 11, making it ideal for the final output layer of binary classification networks.

However, inside deep neural networks with many hidden layers, Sigmoid has two major drawbacks:

  1. Computational Cost: Calculating exponents (e−ze^{-z}) and divisions for millions of neurons across billions of parameters is computationally expensive.
  2. Saturation & Vanishing Gradients: When zz is large positive (+8+8) or large negative (−8-8), the Sigmoid curve becomes nearly flat. The slope (derivative) approaches zero, which stalls learning during backpropagation.

To solve this in hidden layers, modern AI architectures predominantly rely on the Rectified Linear Unit (ReLU).


Formal Definition of ReLU

The Rectified Linear Unit is defined piecewise:

ReLU(z)=max⁡(0,z)={zif z>00if z≤0\text{ReLU}(z) = \max(0, z) = \begin{cases} z & \text{if } z > 0 \\ 0 & \text{if } z \le 0 \end{cases}

How ReLU Operates in Neural Networks

The Rectified Linear Unit functions as an exact mathematical threshold gate with three critical properties:

  1. Hard Sparsity for Negative Inputs (z≤0z \le 0): Any pre-activation score less than or equal to zero is clamped strictly to 0.00.0. In multi-layer networks, this silences inactive or unhelpful feature channels, creating sparse representations that isolate distinct patterns.
  2. Identity Pass-Through for Positive Inputs (z>0z > 0): Positive pre-activations pass through completely unscaled (ReLU(z)=z\text{ReLU}(z) = z). Because the slope for positive values is always constant (ddzReLU(z)=1\frac{d}{dz}\text{ReLU}(z) = 1), ReLU completely eliminates the vanishing gradient saturation problem that affects deep Sigmoid networks.
  3. Hardware Execution Efficiency: Evaluating max⁡(0,z)\max(0, z) requires only a single hardware comparison (if z > 0 return z else return 0), executing orders of magnitude faster on modern GPUs than exponential transcendental functions like e−ze^{-z}.

Side-by-Side Comparison: Sigmoid vs. ReLU

Architectural PropertySigmoid (σ(z)\sigma(z))ReLU (ReLU(z)\text{ReLU}(z))
Mathematical Formula11+e−z\frac{1}{1 + e^{-z}}max⁡(0,z)\max(0, z)
Output Range(0,1)(0, 1)[0,+∞)[0, +\infty)
Negative Inputs (z<0z < 0)Smooth decay toward 0.00.0Clamped strictly to 0.00.0 (hard sparsity)
Zero Input (z=0z = 0)Exactly 0.50.5Exactly 0.00.0
Positive Inputs (z>0z > 0)Saturates asymptotically at 1.01.0Linear identity pass-through (zz)
Primary Architectural RoleBinary Output Layer (Probabilities)Hidden Representation Layers

To see why modern architectures deploy different activation functions across output and hidden layers, examine the mathematical curves of Sigmoid and ReLU side-by-side:

Sigmoid vs. ReLU Activation Curves Comparison
Side-by-side comparison of Sigmoid and ReLU activation functions: Sigmoid (left) smoothly squashes unbounded inputs into calibrated probabilities between 0 and 1 but saturates at extreme values, while ReLU (right) enforces hard sparsity for negative pre-activations (clamping to 0) and provides identity pass-through with constant slope for positive signals.

Sigmoid provides bounded probabilities strictly contained in (0,1)(0, 1), making it the standard choice for binary classification output layers. However, in deep hidden layers, its flat saturation regions cause vanishing gradients. In contrast, ReLU's constant unit slope (ddz=1\frac{d}{dz} = 1) for all z>0z > 0 preserves gradient flow during backpropagation, while its hard zero clamping (z≤0  ⟹  0z \le 0 \implies 0) prunes inactive feature channels to produce sparse representations.


Previous
Representation Learning Foundations