Byzantine Stochastic Gradient Descent

Alistarh, Dan-Adrian
Allen-Zhu, Zeyuan
Li, Jerry
Bengio, S.
Wallach, H.
Larochelle, H.
Grauman, K.
Cesa-Bianchi, N.
Garnett, R.

Publication date

January 2018

Publisher

Neural Information Processing Systems Foundation

Abstract

This paper studies the problem of distributed stochastic optimization in an adversarial setting where, out of m machines which allegedly compute stochastic gradients every iteration, an α-fraction are Byzantine, and may behave adversarially. Our main result is a variant of stochastic gradient descent (SGD) which finds ε-approximate minimizers of convex functions in T=O~(1/ε²m+α²/ε²) iterations. In contrast, traditional mini-batch SGD needs T=O(1/ε²m) iterations, but cannot tolerate Byzantine failures. Further, we provide a lower bound showing that, up to logarithmic factors, our algorithm is information-theoretically optimal both in terms of sample complexity and time complexity

Extracted data

We use cookies to provide a better user experience.

Data Protection

Byzantine Stochastic Gradient Descent

Abstract

Extracted data

Byzantine Stochastic Gradient Descent

Abstract

Extracted data

Related items

Related items