Artificial Neuron Architecture hero
Lesson 3The Artificial Neuron & Activations

Artificial Neuron Architecture

The unified computational pipeline combining input features, influence weights, baseline bias offsets, and non-linear activation functions.

Over the last two lessons, we engineered the two foundational stages of an artificial neuron:

  1. The Linear Step (Lesson 1): We computed the dot product of features and weights (w⋅xw \cdot x), then added a baseline offset bias (+b+b) to produce a raw linear score: z=w⋅x+bz = w \cdot x + b.
  2. The Activation Function (Lesson 2): We passed that raw linear score through the Sigmoid function (σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}) to transform it into an interpretable probability between 00 and 11.

When you connect these two stages back-to-back, you get the complete mathematical anatomy of an Artificial Neuron.


The Unified Computational Pipeline

Every artificial neuron—from a single standalone classifier to the billions of units inside modern large language models—processes information through the exact same computational pipeline:

Input (x)→Weights (w)∑i=1nwixi→Bias (+b)z→Activation (σ)Output (σ(z))\text{Input } (x) \xrightarrow{\quad \text{Weights } (w) \quad} \sum_{i=1}^n w_i x_i \xrightarrow{\quad \text{Bias } (+b) \quad} z \xrightarrow{\quad \text{Activation } (\sigma) \quad} \text{Output } (\sigma(z))

Let's trace what happens at each stage of this journey:

  1. Input Feature Vector (x∈Rnx \in \mathbb{R}^n): The input to the neuron, containing the ordered feature values representing what we want to evaluate (e.g., our movie script specs [1.0,0.0,0.5,1.0][1.0, 0.0, 0.5, 1.0]).
  2. Weight Vector (w∈Rnw \in \mathbb{R}^n): The numerical measures of influence assigned to each feature.
  3. Dot Product (w⋅x=∑wixiw \cdot x = \sum w_i x_i): Multiplies each feature value by its corresponding weight and adds them together into a single raw score.
  4. Baseline Bias (b∈Rb \in \mathbb{R}): An independent baseline offset. It shifts the score up or down regardless of the input features.
  5. Linear Score (z=w⋅x+bz = w \cdot x + b): The combined score after adding the baseline bias to the dot product.
  6. Activation Function (σ\sigma): A mathematical transformation applied to the linear score. In our classifier, we use the Sigmoid function to compress the unbounded score zz into a probability between 00 and 11, while deeper networks use other activation functions (like ReLU) in this same slot.
  7. Neuron Output (σ(z)\sigma(z)): The final scalar activation produced by the neuron after both steps.

(In our movie classifier, this continuous activation is compared against a threshold of 0.500.50 to classify Hit vs Flop).


Tracing the End-to-End Computational Pipeline

To see how all these mathematical primitives connect into a single unified pipeline, trace how our two film scripts flow through feature weighting, dot product accumulation, baseline bias adjustment, Sigmoid compression, and threshold decision:

🎬 'Die Hard in Space': Complete Pipeline Trace

FeatureFeature Value
(x)
Weight
(w)
Product
(wi · xi)
Dot Product
(w · x)
Industry Bias
(b)
Linear Score
(z)
Sigmoid
(σ(z))
Classification
Action1.0+5.0+5.0+9.5−2.0+7.50.999
(99.9%)
🟢 Hit
(High Potential)
Romance0.0−2.00.0
Comedy0.5+1.0+0.5
Sci-Fi1.0+4.0+4.0

🎬 'The Notebook 2': Complete Pipeline Trace

FeatureFeature Value
(x)
Weight
(w)
Product
(wi · xi)
Dot Product
(w · x)
Industry Bias
(b)
Linear Score
(z)
Sigmoid
(σ(z))
Classification
Action0.0+5.00.0−1.5−2.0−3.50.029
(2.9%)
🔴 Flop
(Low Potential)
Romance1.0−2.0−2.0
Comedy0.5+1.0+0.5
Sci-Fi0.0+4.00.0

Notice the clean division of labor across every column:

  • The Weights (ww) determine which features matter.
  • The Bias (bb) determines how hard it is to activate.
  • The Activation (σ\sigma) determines the scale and range of the output.

NOTE: A Single Neuron Is Already a Complete AI Model

It takes an input vector, evaluates features, offsets a baseline, and classifies a prediction.

But a single neuron is not enough for complex real-world decisions. In Topics 3 and 4, we scale this building block by stacking multiple neurons together into layers and deep networks.


Previous
Sigmoid Activation Function