ערן טרייסטר

אקדמי בכיר

Searching for N:M Fine-grained Sparsity of Weights and Activations in Neural Networks

Ruth Akiva-Hochman, Shahaf E. Finder, Javier S. Turek, Eran Treister

Sparsity in deep neural networks has been extensively studied to compress and accelerate models for environments with limited resources. The general approach of pruning aims at enforcing sparsity on the obtained model, with minimal accuracy loss, but with a sparsity structure that enables acceleration on hardware. The sparsity can be enforced on either the weights or activations of the network, and existing works tend to focus on either one for the entire network. In this paper, we suggest a strategy based on Neural Architecture Search (NAS) to sparsify both activations and weights throughout the network, while utilizing the recent approach of N:M fine-grained structured sparsity that enables practical acceleration on dedicated GPUs. We show that a combination of weight and activation pruning is superior to each option separately. Furthermore, during the training, the choice between pruning the weights of activations can be motivated by practical inference costs (e.g., memory bandwidth). We demonstrate the efficiency of the approach on several image classification datasets.

שפת פרסום אנגלית
דפים 130-143
סטטוס פרסום פורסם - 01.01.2023

Keywords

Activation pruning
N:M fine-grained Sparsity
Neural architecture search
Weight pruning

ASJC Scopus subject areas

Theoretical Computer Science
General Computer Science
גישה למסמך
10.1007/978-3-031-25082-8_9
קבצים וקישורים אחרים
Link to publication in Scopus