Part 9 · Assemble the pieces

A neural network is repeated numerical bookkeeping with learned representations.

There is no extra learning step hidden behind the terminology. Initialize, predict, measure, differentiate, update, and repeat.

Reconstruct the order

Arrange these operations: update parameters, compute loss, initialize parameters, backward pass, forward pass.

The complete training algorithm

Prepare data: split training and validation data, then fit preprocessing on training data only.
Initialize: create small random weights and suitable biases.
Forward pass: calculate hidden scores, hidden activations, output score, and probabilities.
Loss: summarize prediction error with binary cross-entropy.
Backward pass: apply the chain rule from output to input and compute every gradient.
Update: subtract learning-rate-scaled gradients from parameters.
Monitor: repeat while tracking training and validation loss and useful classification metrics.

Minimal NumPy map

# Initialize once W1 = small_random_values((n_features, n_hidden)) b1 = np.zeros((1, n_hidden)) W2 = small_random_values((n_hidden, 1)) b2 = np.zeros((1, 1)) for epoch in range(n_epochs): # Forward Z1 = X @ W1 + b1 A1 = np.tanh(Z1) Z2 = A1 @ W2 + b2 A2 = sigmoid(Z2) # Mean BCE loss loss = binary_cross_entropy(Y, A2) # Backward: 1/m is included here exactly once dZ2 = (A2 - Y) / m dW2 = A1.T @ dZ2 db2 = np.sum(dZ2, axis=0, keepdims=True) dA1 = dZ2 @ W2.T dZ1 = dA1 * (1 - A1**2) dW1 = X.T @ dZ1 db1 = np.sum(dZ1, axis=0, keepdims=True) # Update W1 -= learning_rate * dW1 b1 -= learning_rate * db1 W2 -= learning_rate * dW2 b2 -= learning_rate * db2

Frameworks such as PyTorch automate gradient calculation, but the mathematical sequence remains the same.

A debugging checklist

Shapes

  • Does every matrix multiplication have matching inner dimensions?
  • Does each gradient match its parameter’s shape?
  • Are labels shaped m × 1?

Numerics

  • Are probabilities clipped for hand-written BCE?
  • Is averaging by m applied exactly once?
  • Are input feature scales reasonable?

Learning

  • Does loss decrease on a tiny training set?
  • Do parameters actually change?
  • Can the model overfit a handful of examples?

Evaluation

  • Is preprocessing fitted only on training data?
  • Are validation examples excluded from updates?
  • Is the threshold chosen for the real metric or cost?

Retrieval questions

  1. Why does stacking only linear transformations still produce a linear transformation?
  2. What two jobs does an activation function perform in training?
  3. What is the difference between Z[1] and A[1]?
  4. Why does sigmoid plus BCE produce dZ[2] = A[2] − Y?
  5. Why do output weights appear when error moves into the hidden layer?
  6. Why do bias gradients sum over rows?
  7. What is the difference between backpropagation and gradient descent?
  8. Why is all-zero hidden-weight initialization a problem?

What Class 1 has built

\[\text{features}\xrightarrow{\text{affine + activation}}\text{learned features}\xrightarrow{\text{affine + sigmoid}}\text{probability}\xrightarrow{\text{BCE}}\text{loss}\]

You can now explain and implement a one-hidden-layer binary classifier from first principles. Class 2 can extend this foundation to deeper networks, mini-batches, multiclass softmax, initialization strategies, regularization, optimizers, and framework implementations.

The durable mental model: the forward pass builds a prediction; the loss measures it; backpropagation computes sensitivities; the optimizer uses them to improve the next prediction.

Previous: UpdatesReturn to overview