
Sigmoid Activation Function
The S-shaped Sigmoid activation function that squashes unbounded linear scores into standardized probabilities between zero and one.
An Artificial Neuron executes its forward pass in two consecutive computational stages:
The Two Stages of an Artificial Neuron
- Stage 1: The Linear Step () — It calculates the dot product of input features and weights, then adds the baseline bias: . (You already know how to do this from Lesson 1).
- Stage 2: The Activation Function — Transforms the raw linear score into a standardized format.
In Lesson 1, our linear step produced scores of for 'Die Hard in Space' and for 'The Notebook 2'.
What Does a Score of +7.5 Mean?
If you hand to a studio executive, their immediate question is: "7.5 out of what? " Is a guaranteed blockbuster? What if the score in a larger model is ? What if another film script evaluated with 50 features scores ?
On its own, a raw number like or has no standardized scale:
- It has no upper or lower bound ().
- It cannot be interpreted as a percentage or probability.
- It cannot be compared across different models.
To solve this, the neuron passes the linear score () through an Activation Function—a mathematical transformation that converts the raw score into a standardized, usable format.
The Sigmoid Activation Function
For binary classification problems—where the model must predict one of the two possible outcomes (like Hit vs. Flop)—the widely used activation function is the Sigmoid function, denoted by (sigma):
Where is the mathematical constant - Euler's number ().
TIP: You do not need to memorize this formula or calculate by yourself. What matters is its core purpose: it takes any raw, unbounded score ( to ) and squashes it into a standardized probability between and .

Looking at the curve reveals four key mathematical properties:
- Outputs a Valid Probability (The Primary Purpose): This is Sigmoid’s core job. No matter how large or small the input score is, always stays strictly between and . This makes its output directly interpretable as a probability.
- Symmetric Midpoint at : When the input score is exactly zero, ().
- Directional Sensitivity: Positive inputs () yield outputs above (climbing toward ). Negative inputs () yield outputs below (falling toward ).
- Saturation (Diminishing Returns): Near , the curve is steep—small changes in produce significant changes in probability. But beyond , the curve flattens into plateaus. Once is large, extreme numbers produce diminishing returns, preventing outliers from blowing up downstream calculations.
The Binary Decision Rule
In binary classification, the model chooses between two possible outcomes: (like Hit) or (like Flop).
Because outputs a value between and , its output is directly read as the predicted probability of a Hit (ranging from to ). The standard decision threshold is set at ():
- If (which corresponds to ), predict (Hit).
- If (which corresponds to ), predict (Flop).
Applying Sigmoid to Our Movie Predictor
Now we pass our linear scores from Lesson 1 through to get probabilities:
-
🎬 'Die Hard in Space' ():
Because the probability is , the neuron classifies it as a Hit.
-
🎬 'The Notebook 2' ():
Because the probability is , the neuron classifies it as a Flop. An output of doesn't mean a chance of a Flop—it means there is only a chance of a Hit, leaving a chance of a Flop.
| Script | Dot Product (w · x) | Industry Bias (b) | Linear Score (z) | Sigmoid (σ(z)) | Classification |
|---|---|---|---|---|---|
| 'Die Hard in Space' | +9.5 | −2.0 | +7.5 | 0.999 (99.9%) | 🟢 Hit |
| 'The Notebook 2' | −1.5 | −2.0 | −3.5 | 0.029 (2.9%) | 🔴 Flop |
In two simple steps—a linear combination followed by a non-linear activation—the artificial neuron has converted raw script feature values into a Hit probability.