A layer is a team of neurons reading the same input.
Each hidden neuron asks a different learned question about the same row. Their answers become a new representation that the output neuron can combine.
Predict before continuing
If two hidden neurons receive the same inputs, why do they not always produce the same output?
Our complete architecture
W contains trainable weights on connections; b is a trainable offset added to each neuron’s weighted sum before its activation.
The input layer merely holds the supplied features; it does not learn parameters. The hidden and output layers each perform an affine transformation followed by an activation.
x: 1 × 2
z[1]: 1 × 2
a[1]: 1 × 2
z[2]: 1 × 1
a[2]: 1 × 1
Write each hidden neuron separately
With two input features, hidden neuron 1 and hidden neuron 2 compute:
Superscript [1] means layer 1. The subscript identifies a neuron or connection. The two hidden activations now behave like two learned features.
The output neuron combines those learned features:
The same work in matrix notation
For a batch of m examples, rows are examples and columns are features or neurons:
b[1]: 1 × 2
b[2]: 1 × 1
Shape rule: the inner dimensions must match. (m × 2)(2 × 2) produces m × 2; then (m × 2)(2 × 1) produces m × 1.
Biases have one row and are broadcast across all m examples. Every example gets the same learned bias for a given neuron.
How many numbers does this network learn?
Width increases representational capacity, but it also increases the number of parameters, computation, and opportunities to overfit. We use only two hidden neurons here so every number remains visible.