Gradient descent
Gradient descent is the step-by-step method most machine-learning systems use to learn: it measures how the error changes as each parameter changes (the gradient) and nudges every parameter a little way downhill. The learning rate sets the size of each step — too small and learning crawls, too large and it overshoots and can blow up. Variants such as stochastic gradient descent, momentum and Adam make it practical for models with millions of parameters.
From full-batch to stochastic optimisation
Machine-learning objectives are usually sums over training examples, so the full gradient costs a pass over the data. Stochastic gradient descent replaces it with an unbiased minibatch estimate, trading noise for far cheaper steps; SGD and its variants are probably the most used optimisers in machine learning.
The minibatch gradient g: the average gradient of the loss L over m randomly chosen training examples.
SGD with momentum: the velocity v is an exponentially decaying average of negative minibatch gradients.
Momentum borrows its name from mechanics — the negative gradient acts as a force on a particle in parameter space, and an extra force proportional to −v plays the role of viscous drag, so the particle gradually loses energy and comes to rest in a local minimum. Adam adapts per-parameter step sizes from running moment estimates. The learning-rate trade-off persists: too large gives violent oscillation and rising cost, too small gives slow progress or stalling at high cost.
Full explanation — the complete reference version every reading depth is based on
Walking downhill on an error landscape
Training a model means finding parameter values that make a loss (an error score) as small as possible. Picture the loss as a landscape: the height at each point is the error for those parameter values. The gradient points in the direction of steepest increase — straight uphill — so taking a small step in the opposite direction lowers the loss. Repeating that step is gradient descent, a technique Goodfellow and colleagues trace back to Cauchy in 1847.
The gradient-descent update: subtract the learning rate η times the gradient of the loss f.
Worked example: f(x) = x²
The gradient-descent simulation uses the simplest possible loss, f(x) = x², whose minimum is at x = 0 and whose gradient is f′(x) = 2x. Start at x = 10 with learning rate η = 0.1. Step 1: x = 10 − 0.1 × 20 = 8 (loss 64). Step 2: x = 8 − 0.1 × 16 = 6.4 (loss 40.96). Step 3: x = 5.12. Each step multiplies x by 1 − 2η = 0.8, so x shrinks steadily towards 0.
For f(x) = x² every step multiplies x by (1 − 2η).
So x after t steps is x₀ times (1 − 2η) to the power t.
- 0 < η < 0.5: x shrinks towards 0 without changing sign. The simulation's 'small steps' preset (η = 0.1, 20 steps) ends at x = 10 × 0.8²⁰ ≈ 0.115.
- η = 0.5: the factor 1 − 2η is 0, so the very first step lands exactly on the minimum.
- 0.5 < η < 1: x flips sign each step but still shrinks — it zig-zags into the minimum.
- η = 1: x flips between +x₀ and −x₀ for ever and never settles.
- η > 1: each step overshoots further than the last and x grows without limit — the run diverges. With η = 1.1, x goes 10, −12, 14.4, −17.28…; with η = 1.5 it goes 10, −20, 40, −80…
Choosing the learning rate
In real training the safe range is not known in advance. Goodfellow and colleagues describe the symptoms: if the learning rate is too large the learning curve shows violent oscillations and the cost often increases significantly; if it is too low, learning is slow and may get stuck at a high cost. Practitioners watch the loss curve and adjust.
Real landscapes are bumpy
Where the gradient is zero — a critical point — gradient descent has no direction to move. That point might be a local minimum (lower than all its neighbours but not necessarily the lowest overall), a maximum, or a saddle point (downhill in some directions, uphill in others). Neural-network losses are not convex, so there is no guarantee of reaching the global minimum. For many classes of high-dimensional functions, saddle points are far more common than local minima.
Making it practical
- Stochastic gradient descent (SGD): estimate the gradient from a small random minibatch of examples instead of the whole dataset — probably the most used optimisation algorithm in machine learning.
- Momentum: keep a running 'velocity' so steps build up speed along consistent directions; the name comes from a physical analogy with a particle pushed by a force under Newton's laws.
- Adam: adapt the step size for each parameter using running estimates of the gradient's moments.
- Back-propagation: the efficient way to compute the gradient for every weight in a neural network, which gradient descent then uses.
Ask ScienceVerse
Still curious about Gradient descent? Ask a question, get hints, take a short lesson or try a challenge. The tutor answers only from this concept's approved sources, and says so when it has none.
Ask the tutor about this concept on the full tutor page.
Connections
Guided learning path
See everything to learn before this, in order, with your progress:
Related concepts
- Neural networks — Applied in
- Momentum — Related to
- Forces — Related to
- Energy — Related to
Try the simulation
Explore this concept interactively, with real adjustable parameters:
- Gradient Descent (Plus)
Check your understanding
Take a quick check of two to five questions, with an explanation for every answer:
See the neighbourhood of Gradient descent in the Knowledge Galaxy
Sources and methodology
- Gradient descent reduces a function by repeatedly moving its input a small step in the direction of the negative gradient, which points directly downhill; Goodfellow and colleagues attribute the technique to Cauchy (1847). (awaiting scientific review)
- In the gradient-descent update x′ = x − ε∇f(x), the learning rate ε is a positive scalar that determines the size of each step. (awaiting scientific review)
- If the learning rate is too large, the learning curve shows violent oscillations with the cost often increasing significantly; if it is too low, learning proceeds slowly and can become stuck at a high cost. (awaiting scientific review)
- Worked calculation (author's own; it is exactly what the gradient-descent simulation computes): for f(x) = x² the gradient is 2x, so each update multiplies x by (1 − 2η); the iterates converge to the minimum at 0 when 0 < η < 1 (landing on it in a single step when η = 0.5), bounce between +x₀ and −x₀ without settling at η = 1, and diverge when η > 1. (awaiting scientific review)
- Where the derivative is zero (a critical point) it gives no information about which direction to move; a local minimum is a point where the function is lower than at all neighbouring points, so infinitesimal steps can no longer decrease it. (awaiting scientific review)
- For many classes of random functions, local minima are common in low-dimensional spaces, but in higher-dimensional spaces local minima are rare and saddle points are more common. (awaiting scientific review)
- Stochastic gradient descent (SGD) and its variants are probably the most used optimisation algorithms for machine learning; SGD follows an estimate of the gradient computed as the average over a small random minibatch of training examples. (awaiting scientific review)
- The name of the momentum optimisation algorithm derives from a physical analogy in which the negative gradient is a force moving a particle through parameter space according to Newton's laws of motion. (awaiting scientific review)
- In the physical picture behind the momentum algorithm, Goodfellow and colleagues add a force proportional to −v that corresponds to viscous drag, which makes the particle gradually lose energy over time and eventually converge to a local minimum instead of oscillating for ever. (awaiting scientific review)
- Adam (Kingma and Ba, 2014) is a first-order gradient-based method for optimising stochastic objective functions, based on adaptive estimates of lower-order moments of the gradient. (awaiting scientific review)
- Adam: A Method for Stochastic Optimization — Peer-reviewed paper
Claims marked “awaiting scientific review” cite the sources listed but have not yet been signed off by a scientific reviewer.
Content status: published 1 October 2026.
- Scientific review: this version has not yet been signed off by a scientific reviewer.
- The Advanced explanation has not yet been reviewed for age suitability.