Performance engineering for real and complex tall & skinny matrix multiplication kernels on GPUs

Ernst, Dominik
Hager, Georg
Thies, Jonas
Wellein, Gerhard

Open link

Publication date

January 2021

DOI

10.1177/1094342020965661

Publisher

SAGE Publications

Abstract

General matrix-matrix multiplications with double-precision real and complex entries (DGEMM and ZGEMM) in vendor-supplied BLAS libraries are best optimized for square matrices but often show bad performance for tall & skinny matrices, which are much taller than wide. NVIDIA’s current CUBLAS implementation delivers only a fraction of the potential performance as indicated by the roofline model in this case. We describe the challenges and key characteristics of an implementation that can achieve close to optimal performance. We further evaluate different strategies of parallelization and thread distribution and devise a flexible, configurable mapping scheme. To ensure flexibility and allow for highly tailored implementations we use code gener...

Extracted data

We use cookies to provide a better user experience.

Data Protection

Performance engineering for real and complex tall & skinny matrix multiplication kernels on GPUs

Abstract

Extracted data

Performance engineering for real and complex tall & skinny matrix multiplication kernels on GPUs

Abstract

Extracted data

Related items

Related items