Enabling Loop Fusion and Tiling for Cache Performance by Fixing Fusion-Preventing Data Dependences

Jingling Xue
Qingguang Huang
Minyi Guo

Publication date

January 2005

Publisher

SpringerVerlag

Abstract

This paper presents a new approach to enabling loop fusion and tiling for arbitrary affine loop nests. Given a set of multiple loop nests, we present techniques that automatically eliminate all the fusion-preventing dependences by means of loop tiling and ar-ray copying. Applying our techniques iteratively to multiple loop nests yields a single loop nest that can be tiled for cache locality. Our approach handles LU, QR, Cholesky and Jacobi in a unified framework. Our experimental evaluation on an SGI Octane2 sys-tem shows that the benefit from the significantly reduced L1 and L2 cache misses has far more than offset the branching and loop control overhead introduced by our approach.

Extracted data

We use cookies to provide a better user experience.

Data Protection

Enabling Loop Fusion and Tiling for Cache Performance by Fixing Fusion-Preventing Data Dependences

Abstract

Extracted data

Enabling Loop Fusion and Tiling for Cache Performance by Fixing Fusion-Preventing Data Dependences

Abstract

Extracted data

Related items

Related items