Reformulating the direct convolution for high-performance deep learning inference on ARM processors

Sergio Barrachina,Adrian Castello,Manuel F. Dolz,Tze Meng Low,Hector Martinez,Enrique S. Quintana-Orti,Upasana Sridhar,Andres E. Tomas

Journal of Systems Architecture（2023）

引用 3|浏览42

暂无评分

摘要

We present two high-performance implementations of the convolution operator via the direct algorithm that outperform the so-called lowering approach based on the im2col transform plus the gemm kernel on an ARMv8-based processor. One of our methods presents the additional advantage of zero-memory overhead while the other employs an additional yet rather moderate workspace, substantially smaller than that required by the im2col+gemm solution. In contrast with a previous implementation of a similar zero-memory overhead direct convolution, this work exhibits the key advantage of preserving the conventional NHWC data layout for the input/output activations of the convolution layers.

查看译文

关键词

Convolution,Direct algorithm,Deep learning,High performance,ARMv8 architecture

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要