← Back to Gradient descent
All 12 questions. After you submit, each question shows the right answer and why. Back to the quick check
What does the learning rate control in gradient descent?
In gradient descent, each step moves the parameters in the direction that makes the error go down.
On f(x) = x², starting at x = 10, which learning rate makes gradient descent land exactly on the minimum in one step?
Gradient descent on f(x) = x² (gradient 2x), starting at x = 10 with learning rate 0.1. What is x after one step?
The simulation's 'small steps' preset runs gradient descent on f(x) = x² from x = 10 with learning rate 0.1 for 20 steps. What is x at the end? (3 decimal places)
Your training loss is jumping up and down wildly and often getting bigger. What is the most likely cause?
Same loss f(x) = x², start x = 10, but learning rate 1.5. What is x after two steps?
What does stochastic gradient descent use instead of the gradient over the entire training set?
For many classes of high-dimensional functions, local minima are more common than saddle points.
What does the Adam optimiser adapt using running estimates of the gradient's lower-order moments?
Explain where the name of the 'momentum' optimiser comes from, and what plays the role of the force.
In the physical picture behind the momentum optimiser, what does the extra force proportional to −v (viscous drag) do?