AdaLoss: A Computationally-Efficient and Provably Convergent Adaptive Gradient Method

Wu, Xiaoxia
Xie, Yuege
Du, Simon Shaolei
Ward, Rachel

Open link

Publication date

June 2022

DOI

10.1609/aaai.v36i8.20848

Publisher

Association for the Advancement of Artificial Intelligence

Abstract

We propose a computationally-friendly adaptive learning rate schedule, ``AdaLoss", which directly uses the information of the loss function to adjust the stepsize in gradient descent methods. We prove that this schedule enjoys linear convergence in linear regression. Moreover, we extend the to the non-convex regime, in the context of two-layer over-parameterized neural networks. If the width is sufficiently large (polynomially), then AdaLoss converges robustly to the global minimum in polynomial time. We numerically verify the theoretical results and extend the scope of the numerical experiments by considering applications in LSTM models for text clarification and policy gradients for control problems

Extracted data

We use cookies to provide a better user experience.

Data Protection

AdaLoss: A Computationally-Efficient and Provably Convergent Adaptive Gradient Method

Abstract

Extracted data

AdaLoss: A Computationally-Efficient and Provably Convergent Adaptive Gradient Method

Abstract

Extracted data

Related items

Related items