Part 2 · Functions that shape information and gradients

An activation affects both the forward signal and the backward gradient.

During the forward pass, an activation decides what a neuron emits. During backpropagation, its derivative decides how sensitive that output is to a change in the neuron’s weighted sum.

Think before the formulas

If every layer uses the identity function g(z) = z, can adding more layers create a nonlinear decision boundary?

Two questions for every activation

Forward question

What range and shape of signal does this function produce?

Backward question

Where is its derivative large, small, zero or undefined?

Identity: useful at a regression output

\[g(z)=z,\qquad g'(z)=1\]

Identity passes the score through unchanged.

Its derivative is always one.

It is a natural output for an unrestricted regression target, but it cannot supply the hidden nonlinearity needed for XOR.

Sigmoid: a binary probability output

\[\sigma(z)=\frac{1}{1+e^{-z}},\qquad \sigma'(z)=\sigma(z)\bigl(1-\sigma(z)\bigr)\]

Range: (0, 1). The output can represent a binary probability.

The largest derivative is 0.25 at z = 0; it approaches zero in both saturated tails.

For a large positive or negative z, changing z barely changes the sigmoid output. Repeated multiplication by small derivatives can make early-layer gradients tiny.

Class 1 choice: sigmoid belongs at our binary output. We will not use it in the hidden layer of the main derivation.

Tanh: centred, smooth hidden activations

\[\tanh(z)=\frac{e^z-e^{-z}}{e^z+e^{-z}},\qquad \frac{d}{dz}\tanh(z)=1-\tanh^2(z)\]

Range: (−1, 1). Positive and negative outputs are centred around zero.

The derivative is largest near zero and approaches zero in the tails.

Tanh still saturates, but its smooth derivative makes the first manual backpropagation example easy to follow.

ReLU: simple and effective hidden activation

\[\operatorname{ReLU}(z)=\max(0,z),\qquad \operatorname{ReLU}'(z)=\begin{cases}1,&z>0\\0,&z<0\end{cases}\]

Negative scores become zero; positive scores pass through.

The positive side keeps a gradient of one. At zero, software uses a chosen subgradient, commonly zero.

A neuron that remains on the negative side for every example receives zero ReLU gradient and may stop learning. This is the dying-ReLU problem.

Leaky ReLU: retain a small negative gradient

\[g(z)=\begin{cases}z,&z>0\\ \alpha z,&z\leq0\end{cases},\qquad g'(z)=\begin{cases}1,&z>0\\ \alpha,&z<0\end{cases}\]

The figure uses α = 0.1 so the negative slope is visible.

Negative neurons still receive a small gradient.

Choose activation by location and task

LocationTypical choiceReasonMain caution
Hidden layerReLU or a variantSimple, efficient and strong positive-side gradientInactive neurons can die
First teaching derivationTanhSmooth derivative and signed outputsSaturates in both tails
Binary outputSigmoidOne probability in (0, 1)Use with a numerically stable BCE implementation
Multiclass outputSoftmaxClass probabilities summing to oneOperate on logits stably
Regression outputIdentityAllows unrestricted real valuesMay need a task-specific constrained output
Previous: Why a network?Next: Architecture