Learn Before
Multiple Choice
Identifying Gradient Descent Variants from Text Descriptions
Three loss traces are described for the same model and objective:
- Trace 1 is the most jagged and takes the longest to settle.
- Trace 2 still zigzags, but less than Trace 1.
- Trace 3 is the smoothest and reaches low loss first. Match the traces to plain gradient descent, momentum with β = 0.4, and momentum with β = 0.95.
0
1
Updated 2026-08-12
Contributors are:
Who are from:
Tags
Data Science
Machine Learning Yearning @ DeepLearning.AI
Dive into Deep Learning @ D2L
Deep Learning
Machine Learning
Supervised Learning
Related
Intuition behind Gradient Descent with Momentum
Identifying Gradient Descent Variants from Text Descriptions
Adam (Deep Learning Optimization Algorithm)
Origin of the Momentum Method
Velocity Initialization in Momentum Method
Momentum Convergence on a Scalar Quadratic
Gradient Descent with Momentum Pseudocode