> ML_LITERATURE // TIELEMAN-2012-LECTURE-6-5-RMSPROP-DIVIDE-GRADIENT-BY-RUNNING-AVERAGE_v1.0
Lecture 6.5 - RMSProp: Divide the gradient by a running average of its recent magnitude
Geoffrey Hinton, Nitish Srivastava, Kevin Swersky · Coursera: Neural Networks for Machine Learning (2012)
algorithm2012industry-standardthirdPartyReproduced
Principal Contribution
Resolved AdaGrad diminishing learning rate problem by maintaining an exponentially decaying moving average of squared gradients.
Operational Relevance
Serves as qualified reference for implementing task-text-generation, task-image-classification in production systems.
Assumptions
- Underlying computational topology and mathematical bounds adhere to established convexity/smoothness guarantees
Limitations
- Hardware runtime speedups, privacy budgets, and convergence depend on hyperparameters and network communication limits
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
