Forward propagation is an ordered chain of calculations.
Data moves from left to right. At each layer we calculate scores, apply an activation, and pass the result onward. No parameter changes yet.
Calculate first
For x = [1.0, 0.5], which shape should emerge after multiplying by a 2 × 2 hidden-weight matrix?
1 × 2 row: one pre-activation score for each of the two hidden neurons.Step 0: choose one row and initial parameters
Think of the two standardized features as usage frequency and account age. Suppose this row belongs to class 1.
Step 1: hidden weighted sums
The first hidden neuron sees positive evidence; the second receives a score just below zero.
Step 2: hidden activation
Apply tanh element by element:
Representation view: the original row [1.0, 0.5] has been re-expressed as two learned feature values [0.4219, -0.0500].
Step 3: output score and probability
The network currently assigns a probability of about 57.94% to class 1. During training we keep this probability. A threshold such as 0.5 is only needed later when converting probabilities to class decisions.
Cache the evidence for the return journey
Backpropagation will need intermediate values, especially Z[1], A[1], Z[2], and A[2]. A forward function therefore returns the prediction and stores a cache.
Forward pass question: “What does the network predict with its current parameters?” The loss asks “How bad is that prediction?” Backpropagation asks “Which parameters were responsible, and by how much?”