Portable parallel Level-3 BLAS in Linda

Proceedings Scalable High Performance Computing Conference SHPCC-92. Pub Date : 1992-04-26 DOI:10.1109/SHPCC.1992.232664

B. Ghosh, M. Schultz

引用次数: 1

Abstract

Describes an approach towards providing an efficient Level-3 BLAS library over a variety of parallel architectures using C-Linda. A blocked linear algebra program calling the sequential Level-3 BLAS can now run on both shared and distributed memory environments (which support Linda) by simply replacing each call by a call to the corresponding parallel Linda Level-3 BLAS. The authors summarise some of the implementation and algorithmic issues related to the matrix multiplication subroutine. All the various matrix algorithms being block-structured, they are particularly interested in parallel computers with hierarchical memory systems. Experimental data for their implementations show substantial speedups on shared memory, disjoint memory and networked configurations of processors. The authors also present the use of their parallel subroutines in blocked dense LU decomposition and present some preliminary experimental data.<>

查看原文本刊更多论文

Linda的便携式平行3级BLAS

描述了一种使用C-Linda在各种并行架构上提供高效的Level-3 BLAS库的方法。调用顺序3级BLAS的阻塞线性代数程序现在可以在共享和分布式内存环境(支持Linda)上运行，只需将每个调用替换为对相应的并行Linda 3级BLAS的调用。作者总结了一些与矩阵乘法子程序相关的实现和算法问题。所有的矩阵算法都是块结构的，他们对具有分层存储系统的并行计算机特别感兴趣。它们实现的实验数据显示，在共享内存、分离内存和处理器的网络配置上都有显著的加速。作者还介绍了并行子程序在阻塞密集LU分解中的应用，并给出了一些初步的实验数据

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings Scalable High Performance Computing Conference SHPCC-92.

自引率

0.00%

发文量