Neural Networks with Quantization Constraints

Hounie, Ignacio
Elenter, Juan
Ribeiro, Alejandro

Publication date

October 2022

Language

English

Abstract

Enabling low precision implementations of deep learning models, without considerable performance degradation, is necessary in resource and latency constrained settings. Moreover, exploiting the differences in sensitivity to quantization across layers can allow mixed precision implementations to achieve a considerably better computation performance trade-off. However, backpropagating through the quantization operation requires introducing gradient approximations, and choosing which layers to quantize is challenging for modern architectures due to the large search space. In this work, we present a constrained learning approach to quantization aware training. We formulate low precision supervised learning as a constrained optimization problem,...

Extracted data

We use cookies to provide a better user experience.

Data Protection

Neural Networks with Quantization Constraints

Abstract

Extracted data

Neural Networks with Quantization Constraints

Abstract

Extracted data

Related items

Related items