Part 3 · From separate neurons to one network

A layer is a team of neurons reading the same input.

Each hidden neuron asks a different learned question about the same row. Their answers become a new representation that the output neuron can combine.

Predict before continuing

If two hidden neurons receive the same inputs, why do they not always produce the same output?

Our complete architecture

W contains trainable weights on connections; b is a trainable offset added to each neuron’s weighted sum before its activation.

The input layer merely holds the supplied features; it does not learn parameters. The hidden and output layers each perform an affine transformation followed by an activation.

Input row
x: 1 × 2
Hidden scores
z[1]: 1 × 2
Hidden features
a[1]: 1 × 2
Output score
z[2]: 1 × 1
Probability
a[2]: 1 × 1

Write each hidden neuron separately

With two input features, hidden neuron 1 and hidden neuron 2 compute:

\[z_1^{[1]}=x_1w_{11}^{[1]}+x_2w_{21}^{[1]}+b_1^{[1]},\qquad a_1^{[1]}=\tanh(z_1^{[1]})\]\[z_2^{[1]}=x_1w_{12}^{[1]}+x_2w_{22}^{[1]}+b_2^{[1]},\qquad a_2^{[1]}=\tanh(z_2^{[1]})\]

Superscript [1] means layer 1. The subscript identifies a neuron or connection. The two hidden activations now behave like two learned features.

The output neuron combines those learned features:

\[z^{[2]}=a_1^{[1]}w_{11}^{[2]}+a_2^{[1]}w_{21}^{[2]}+b^{[2]},\qquad \widehat y=a^{[2]}=\sigma(z^{[2]})\]

The same work in matrix notation

For a batch of m examples, rows are examples and columns are features or neurons:

\[Z^{[1]}=XW^{[1]}+b^{[1]},\qquad A^{[1]}=\tanh(Z^{[1]})\]\[Z^{[2]}=A^{[1]}W^{[2]}+b^{[2]},\qquad A^{[2]}=\sigma(Z^{[2]})\]
DataX: m × 2
Layer 1W[1]: 2 × 2
b[1]: 1 × 2
Layer 2W[2]: 2 × 1
b[2]: 1 × 1
OutputA[2]: m × 1

Shape rule: the inner dimensions must match. (m × 2)(2 × 2) produces m × 2; then (m × 2)(2 × 1) produces m × 1.

Biases have one row and are broadcast across all m examples. Every example gets the same learned bias for a given neuron.

How many numbers does this network learn?

\[\underbrace{2\times2}_{W^{[1]}}+\underbrace{2}_{b^{[1]}}+\underbrace{2\times1}_{W^{[2]}}+\underbrace{1}_{b^{[2]}}=9\text{ trainable parameters}\]

Width increases representational capacity, but it also increases the number of parameters, computation, and opportunities to overfit. We use only two hidden neurons here so every number remains visible.

Previous: ActivationsNext: Forward pass