Skip to main content

> ML_LITERATURE // TIELEMAN-2012-LECTURE-6-5-RMSPROP-DIVIDE-GRADIENT-BY-RUNNING-AVERAGE_v1.0

Lecture 6.5 - RMSProp: Divide the gradient by a running average of its recent magnitude

Geoffrey Hinton, Nitish Srivastava, Kevin Swersky · Coursera: Neural Networks for Machine Learning (2012)

algorithm2012industry-standardthirdPartyReproduced

Principal Contribution

Resolved AdaGrad diminishing learning rate problem by maintaining an exponentially decaying moving average of squared gradients.

Operational Relevance

Serves as qualified reference for implementing task-text-generation, task-image-classification in production systems.

Assumptions

  • Underlying computational topology and mathematical bounds adhere to established convexity/smoothness guarantees

Limitations

  • Hardware runtime speedups, privacy budgets, and convergence depend on hyperparameters and network communication limits

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: