Videos

Why Gradient Descent Gets Stuck

Short · 1:03 · AI & ML · Math · Watch on YouTube

Why gradient descent gets stuck, and how it escapes.

A ball follows the slope downhill on a landscape with many valleys, just like a neural network following its loss during training. Plain gradient descent settles in the first valley it finds: a local minimum, where the gradient is zero. Step size matters too: too small and it crawls, too large and it bounces out of control. Momentum and random restarts help it reach the deepest valley.

Sources & notes

Every path in this video comes from a real optimization run in Python. Nothing is hand-animated.

gradient descentlocal minimumglobal minimumlocal minimaoptimizationmachine learningdeep learningneural network traininglearning ratemomentumstochastic gradient descentloss function