Publication View

A Unified Compiler Algorithm for Optimizing Locality, Parallelism and Communication in Out-of-Core Computations (1998)

Abstract
This paper presents compiler algorithms to optimize outof -core programs. These algorithms consider loop and data layout transformations in a unified framework. The performance of an out-of-core loop nest containing many references can be improved by a combination of restructuring the loops and file layouts. This approach considers array references one-by-one and attempts to optimize each reference for parallelism and locality. When there are references for which parallelism optimizations do not work, communication is vectorized so that data transfer can be performed before the innermost tiling loop. Preliminary results from handcompiles on IBM SP-2 and Intel Paragon show that this approach reduces the execution time, improves the bandwidth speedup and overall speedup. In addition, we extend the base algorithm to work with file layout constraints and show how it can be used for optimizing programs consisting of multiple loop nests. 1 Introduction In recent years, processor speed has ...

Publication details
Download http://citeseer.ist.psu.edu/141211.html
Source ftp://gate.ee.lsu.edu/pub/jxr/papers/conf/iopads-97.PS
Publisher unknown
Contributors The Pennsylvania State University CiteSeer Archives
Repository CiteSeer (United States)
Keywords M. Kandemir,A. Choudhary,J. Ramanujam,M. Kandaswamy A Unified Compiler Algorithm for Optimizing Locality, Parallelism and Communication in Out-of-Core Computations
Language Englisch
Relation oai:CiteSeerPSU:473746, oai:CiteSeerPSU:2700, oai:CiteSeerPSU:8253, oai:CiteSeerPSU:68182, oai:CiteSeerPSU:160404, oai:CiteSeerPSU:160724, oai:CiteSeerPSU:141211, oai:CiteSeerPSU:161159