One neuron can draw only one linear boundary.
Logistic regression is powerful when one weighted combination of the supplied features separates the classes. XOR gives us a tiny problem where no such line exists.
Predict before revealing
Can any single straight line separate the blue XOR points from the red XOR points?
The dashed line is only one attempt. Rotating or moving a single line never solves all four observations.
Why logistic regression is limited here
Logistic regression computes a linear score and converts it to a probability:
The 0.5 decision boundary occurs when ŷ = 0.5. Since σ(0) = 0.5, the boundary is:
In two dimensions, this equation describes one straight line. The sigmoid changes the score into a probability, but it does not bend the boundary.
Important distinction: sigmoid is nonlinear as a function of the score z, yet logistic regression still has a linear boundary in the original feature space.
Let neurons divide the work
Instead of asking one neuron to solve XOR directly, hidden neurons can detect simpler conditions. Conceptually:
Hidden neuron 1
Activates for one useful region of the input space.
Hidden neuron 2
Activates for another useful region.
Output neuron
Combines those learned signals into the final probability.
The hidden activations become new features. The output neuron learns a linear rule in this learned feature space, even though the resulting boundary in the original (x₁, x₂) space can be nonlinear.
Why nonlinearity is essential
Suppose we stack two linear layers without an activation:
Substituting the first into the second:
The product and combined bias are still just another weight matrix and bias. Several linear layers collapse into one linear layer.
Therefore: multiple neurons provide capacity, but nonlinear activations prevent the entire stack from collapsing back into a single linear transformation.
What the hidden layer learns
We do not manually define “XOR detector” features. Training discovers hidden weights that make the final loss smaller. Each neuron learns:
The hidden activations a₁[1], a₂[1], ... are data-dependent learned representations. This idea scales beyond XOR: later layers operate on patterns constructed by earlier layers.