How do several neurons learn together?
We already know that a neuron computes a weighted sum and applies an activation. This class connects many such computations into a single-hidden-layer network, follows one prediction forward, sends its error backward, and updates every parameter from scratch.
Opening challenge
If logistic regression already contains a weighted sum and sigmoid activation, what new capability does a hidden layer add?
The entire class in one sentence: hidden neurons learn intermediate features, the output neuron combines them, and backpropagation tells every weight how it should change to reduce the final loss.
The network we will keep throughout
A 2 → 2 → 1 binary classifier: two features, two hidden neurons, and one predicted probability.
Three-hour path
| Time | Teaching movement | Student outcome |
|---|---|---|
| 00–10 | XOR and the failure of one boundary | Explain why combining neurons is necessary |
| 10–35 | Activation functions and derivatives | Choose hidden and output activations |
| 35–65 | Single-hidden-layer architecture | Track parameters and tensor shapes |
| 65–95 | Forward pass, loss and backward pass | Follow one example mathematically |
| 95–105 | Break | |
| 105–160 | NumPy implementation | Train the network from scratch |
| 160–175 | Loss and boundary visualizations | Diagnose whether learning occurred |
| 175–180 | Retrieval summary | Reconstruct the algorithm without notes |
The story pages
Why multiple neurons?
XOR exposes the limitation of one linear boundary.
2Activation functions
Functions, derivatives, saturation and task-dependent choices.
3Network architecture
Hidden neurons as learned features, with parameters and shapes.
4Forward propagation
Calculate every intermediate value for one example and a batch.
5Measure the error
Binary cross-entropy and why confident mistakes cost more.
6Backward intuition
Credit assignment through the chain rule.
7Backward mathematics
Derive every gradient and verify its shape.
8Learn from gradients
Gradient descent, learning rate and repeated updates.
9Reconnect the whole network
The full training algorithm and NumPy map.