Skip to main navigation Skip to search Skip to main content

Algorithm Design for Tensor Units

  • University of Padua
  • Free University of Bozen-Bolzano

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

3 Scopus citations

Abstract

To respond to the intense computational load of deep neural networks, a plethora of domain-specific architectures have been introduced, such as Google Tensor Processing Units and NVIDIA Tensor Cores. A common feature of these architectures is a hardware circuit for efficiently computing a dense matrix multiplication of a given small size. In order to broaden the class of algorithms that exploit these systems, we propose a computational model, named the TCU model, that captures the ability to natively multiply small matrices. We then use the TCU model for designing fast algorithms for several problems, including matrix operations (dense and sparse multiplication, Gaussian Elimination), graph algorithms (transitive closure, all pairs shortest distances), Discrete Fourier Transform, stencil computations, integer multiplication, and polynomial evaluation. We finally highlight a relation between the TCU model and the external memory model.

Original languageEnglish
Title of host publicationEuro-Par 2021
Subtitle of host publicationParallel Processing - 27th International Conference on Parallel and Distributed Computing, Proceedings
EditorsLeonel Sousa, Nuno Roma, Pedro Tomás
PublisherSpringer Science and Business Media Deutschland GmbH
Pages353-367
Number of pages15
ISBN (Print)9783030856649
DOIs
StatePublished - 2021
Event27th International European Conference on Parallel and Distributed Computing, Euro-Par 2021 - Lisbon, Portugal
Duration: Sep 1 2021Sep 3 2021

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume12820 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference27th International European Conference on Parallel and Distributed Computing, Euro-Par 2021
Country/TerritoryPortugal
CityLisbon
Period09/1/2109/3/21

Fingerprint

Dive into the research topics of 'Algorithm Design for Tensor Units'. Together they form a unique fingerprint.

Cite this