Sigmoid Activation Function hero
Lesson 2The Artificial Neuron & Activations

Sigmoid Activation Function

The S-shaped Sigmoid activation function that squashes unbounded linear scores into standardized probabilities between zero and one.

An Artificial Neuron executes its forward pass in two consecutive computational stages:

The Two Stages of an Artificial Neuron

  1. Stage 1: The Linear Step (zz) — It calculates the dot product of input features and weights, then adds the baseline bias: z=w⋅x+bz = w \cdot x + b. (You already know how to do this from Lesson 1).
  2. Stage 2: The Activation Function — Transforms the raw linear score into a standardized format.

In Lesson 1, our linear step produced scores of z=+7.5z = +7.5 for 'Die Hard in Space' and z=−3.5z = -3.5 for 'The Notebook 2'.


What Does a Score of +7.5 Mean?

If you hand +7.5+7.5 to a studio executive, their immediate question is: "7.5 out of what? " Is +7.5+7.5 a guaranteed blockbuster? What if the score in a larger model is +10,000+10,000? What if another film script evaluated with 50 features scores +320+320?

On its own, a raw number like +7.5+7.5 or −3.5-3.5 has no standardized scale:

  • It has no upper or lower bound (z∈(−∞,+∞)z \in (-\infty, +\infty)).
  • It cannot be interpreted as a percentage or probability.
  • It cannot be compared across different models.

To solve this, the neuron passes the linear score (zz) through an Activation Function—a mathematical transformation that converts the raw score into a standardized, usable format.


The Sigmoid Activation Function

For binary classification problems—where the model must predict one of the two possible outcomes (like Hit vs. Flop)—the widely used activation function is the Sigmoid function, denoted by σ\sigma (sigma):

σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}

Where ee is the mathematical constant - Euler's number (e≈2.718e \approx 2.718).

TIP: You do not need to memorize this formula or calculate e−ze^{-z} by yourself. What matters is its core purpose: it takes any raw, unbounded score (−∞-\infty to +∞+\infty) and squashes it into a standardized probability between 00 and 11.

The Sigmoid Activation Function Curve
The Sigmoid function compresses any real number z into a value between 0 and 1. At z = 0, the output is an exact 50% midpoint (0.50). Positive scores climb toward 1.0, while negative scores fall toward 0.0.

Looking at the curve reveals four key mathematical properties:

  1. Outputs a Valid Probability (The Primary Purpose): This is Sigmoid’s core job. No matter how large or small the input score zz is, σ(z)\sigma(z) always stays strictly between 00 and 11. This makes its output directly interpretable as a probability.
  2. Symmetric Midpoint at z=0z = 0: When the input score is exactly zero, σ(0)=0.50\sigma(0) = \mathbf{0.50} (50%50\%).
  3. Directional Sensitivity: Positive inputs (z>0z > 0) yield outputs above 0.500.50 (climbing toward 1.01.0). Negative inputs (z<0z < 0) yield outputs below 0.500.50 (falling toward 0.00.0).
  4. Saturation (Diminishing Returns): Near z=0z = 0, the curve is steep—small changes in zz produce significant changes in probability. But beyond ±4\pm 4, the curve flattens into plateaus. Once ∣z∣|z| is large, extreme numbers produce diminishing returns, preventing outliers from blowing up downstream calculations.

The Binary Decision Rule

In binary classification, the model chooses between two possible outcomes: 11 (like Hit) or 00 (like Flop).

Because σ(z)\sigma(z) outputs a value between 00 and 11, its output is directly read as the predicted probability of a Hit (ranging from 0%0\% to 100%100\%). The standard decision threshold is set at 0.500.50 (50%50\%):

  • If σ(z)≥0.50\sigma(z) \ge 0.50 (which corresponds to z≥0z \ge 0), predict 11 (Hit).
  • If σ(z)<0.50\sigma(z) < 0.50 (which corresponds to z<0z < 0), predict 00 (Flop).

Applying Sigmoid to Our Movie Predictor

Now we pass our linear scores from Lesson 1 through σ(z)\sigma(z) to get probabilities:

  • 🎬 'Die Hard in Space' (z=+7.5z = +7.5):

    σ(+7.5)=11+e−7.5≈0.999(99.9%)\sigma(+7.5) = \frac{1}{1 + e^{-7.5}} \approx \mathbf{0.999} \quad (\mathbf{99.9\%})

    Because the probability is 99.9%≥50%99.9\% \ge 50\%, the neuron classifies it as a Hit.

  • 🎬 'The Notebook 2' (z=−3.5z = -3.5):

    σ(−3.5)=11+e−(−3.5)=11+e3.5≈0.029(2.9%)\sigma(-3.5) = \frac{1}{1 + e^{-(-3.5)}} = \frac{1}{1 + e^{3.5}} \approx \mathbf{0.029} \quad (\mathbf{2.9\%})

    Because the probability is 2.9%<50%2.9\% < 50\%, the neuron classifies it as a Flop. An output of 0.0290.029 doesn't mean a 2.9%2.9\% chance of a Flop—it means there is only a 2.9%2.9\% chance of a Hit, leaving a 97.1%97.1\% chance of a Flop.

ScriptDot Product
(w · x)
Industry Bias
(b)
Linear Score
(z)
Sigmoid
(σ(z))
Classification
'Die Hard in Space'+9.5−2.0+7.50.999 (99.9%)🟢 Hit
'The Notebook 2'−1.5−2.0−3.50.029 (2.9%)🔴 Flop

In two simple steps—a linear combination followed by a non-linear activation—the artificial neuron has converted raw script feature values into a Hit probability.


Previous
Linear Step and Baseline Bias