A neural network is repeated numerical bookkeeping with learned representations.
There is no extra learning step hidden behind the terminology. Initialize, predict, measure, differentiate, update, and repeat.
Reconstruct the order
Arrange these operations: update parameters, compute loss, initialize parameters, backward pass, forward pass.
The complete training algorithm
Minimal NumPy map
Frameworks such as PyTorch automate gradient calculation, but the mathematical sequence remains the same.
A debugging checklist
Shapes
- Does every matrix multiplication have matching inner dimensions?
- Does each gradient match its parameter’s shape?
- Are labels shaped
m × 1?
Numerics
- Are probabilities clipped for hand-written BCE?
- Is averaging by m applied exactly once?
- Are input feature scales reasonable?
Learning
- Does loss decrease on a tiny training set?
- Do parameters actually change?
- Can the model overfit a handful of examples?
Evaluation
- Is preprocessing fitted only on training data?
- Are validation examples excluded from updates?
- Is the threshold chosen for the real metric or cost?
Retrieval questions
- Why does stacking only linear transformations still produce a linear transformation?
- What two jobs does an activation function perform in training?
- What is the difference between
Z[1]andA[1]? - Why does sigmoid plus BCE produce
dZ[2] = A[2] − Y? - Why do output weights appear when error moves into the hidden layer?
- Why do bias gradients sum over rows?
- What is the difference between backpropagation and gradient descent?
- Why is all-zero hidden-weight initialization a problem?
What Class 1 has built
You can now explain and implement a one-hidden-layer binary classifier from first principles. Class 2 can extend this foundation to deeper networks, mini-batches, multiclass softmax, initialization strategies, regularization, optimizers, and framework implementations.
The durable mental model: the forward pass builds a prediction; the loss measures it; backpropagation computes sensitivities; the optimizer uses them to improve the next prediction.