An activation affects both the forward signal and the backward gradient.
During the forward pass, an activation decides what a neuron emits. During backpropagation, its derivative decides how sensitive that output is to a change in the neuron’s weighted sum.
Think before the formulas
If every layer uses the identity function g(z) = z, can adding more layers create a nonlinear decision boundary?
Two questions for every activation
Forward question
What range and shape of signal does this function produce?
Backward question
Where is its derivative large, small, zero or undefined?
Identity: useful at a regression output
Identity passes the score through unchanged.
Its derivative is always one.
It is a natural output for an unrestricted regression target, but it cannot supply the hidden nonlinearity needed for XOR.
Sigmoid: a binary probability output
Range: (0, 1). The output can represent a binary probability.
The largest derivative is 0.25 at z = 0; it approaches zero in both saturated tails.
For a large positive or negative z, changing z barely changes the sigmoid output. Repeated multiplication by small derivatives can make early-layer gradients tiny.
Class 1 choice: sigmoid belongs at our binary output. We will not use it in the hidden layer of the main derivation.
Tanh: centred, smooth hidden activations
Range: (−1, 1). Positive and negative outputs are centred around zero.
The derivative is largest near zero and approaches zero in the tails.
Tanh still saturates, but its smooth derivative makes the first manual backpropagation example easy to follow.
ReLU: simple and effective hidden activation
Negative scores become zero; positive scores pass through.
The positive side keeps a gradient of one. At zero, software uses a chosen subgradient, commonly zero.
A neuron that remains on the negative side for every example receives zero ReLU gradient and may stop learning. This is the dying-ReLU problem.
Leaky ReLU: retain a small negative gradient
The figure uses α = 0.1 so the negative slope is visible.
Negative neurons still receive a small gradient.
Choose activation by location and task
| Location | Typical choice | Reason | Main caution |
|---|---|---|---|
| Hidden layer | ReLU or a variant | Simple, efficient and strong positive-side gradient | Inactive neurons can die |
| First teaching derivation | Tanh | Smooth derivative and signed outputs | Saturates in both tails |
| Binary output | Sigmoid | One probability in (0, 1) | Use with a numerically stable BCE implementation |
| Multiclass output | Softmax | Class probabilities summing to one | Operate on logits stably |
| Regression output | Identity | Allows unrestricted real values | May need a task-specific constrained output |