A gradient becomes learning only after an update.
Gradient descent takes a controlled step opposite the local slope. Then the network repeats the whole forward and backward process with its new parameters.
Choose the direction
If w = 0.7, its gradient is −0.1775, and the learning rate is 0.1, will the updated weight be larger or smaller?
0.7 − 0.1(−0.1775) = 0.71775. Subtracting a negative gradient moves upward.The update rule
θ stands for any trainable weight or bias. The learning rate η controls step size. The minus sign moves against the gradient, the direction of steepest local increase.
One complete numerical update
Using learning rate η = 0.1 and the gradients from the previous page:
For this one row, these changes push the next prediction toward its target. On a batch, the update reflects the average direction suggested by all examples.
Learning rate changes the journey
Too small
Loss may fall reliably but extremely slowly. Training wastes computation.
Useful range
Loss generally falls with manageable fluctuations and meaningful progress.
Too large
Updates can overshoot low-loss regions, oscillate, or diverge.
A healthy curve need not be perfectly smooth, especially with mini-batches, but its overall direction should improve.
Why not initialize every weight to zero?
If hidden neurons start with identical weights, they produce identical outputs, receive identical gradients, and remain identical after every update. They never specialize. Small random weights break this symmetry.
Biases may begin at zero because differing incoming weights already distinguish the neurons. Initialization scale becomes more important in deep networks and returns in Class 2.
One epoch and many epochs
An epoch means the training procedure has used every training example once. Full-batch gradient descent makes one update per epoch; mini-batch gradient descent makes several. Neural-network training repeats:
Evaluation remains separate: validation data measures generalization but must not directly update the weights.