Efficient AI model acceleration through tensor decomposition and dataflow architecture design
File(s)
Author(s)
Yu, Zhewen
Type
Thesis
Abstract
The exponential growth of AI model size presents significant challenges in terms of computational resources and memory footprints. To address these challenges, both software and hardware engineers are exploring optimization strategies. On the software side, model compression techniques such as pruning and quantization aim to reduce the number of parameters, without significantly sacrificing accuracy. On the hardware side, specialized processors such as AI accelerators or Neural Processing Units (NPUs) are being developed, especially to optimize the execution of matrix multiplication, a key operation in AI workloads.
However, AI models are also becoming increasingly diverse, leading to varied computational and memory requirements. As such, uniform optimization strategies are no longer sufficient to handle this diversity effectively. To address this, it is crucial to expand the design space, allowing design choices to be customized based on the specific characteristics of each AI model.
This thesis proposes an expanded design space approach. On the software algorithm side, advanced tensor decomposition algorithms are introduced, allowing layerwise customization of decomposition schemes for a given AI model. This fine-grained approach significantly reduces accuracy degradation by 7 to 20 percentage points compared to prior methods. On the hardware architecture side, the thesis enhances novel dataflow designs using Field Programmable Gate Arrays (FPGAs). A novel memory management methodology intelligently allocates both on-chip and off-chip memory resources, achieving a latency reduction of 25 to 48%. Additionally, the thesis bridges the gap between software and hardware through fast and efficient design automation techniques, cutting the search time by at least 10x. This integration demonstrates significant potential for the efficient acceleration of AI models and reduces repetitive engineering efforts in the rapidly evolving AI landscape.
However, AI models are also becoming increasingly diverse, leading to varied computational and memory requirements. As such, uniform optimization strategies are no longer sufficient to handle this diversity effectively. To address this, it is crucial to expand the design space, allowing design choices to be customized based on the specific characteristics of each AI model.
This thesis proposes an expanded design space approach. On the software algorithm side, advanced tensor decomposition algorithms are introduced, allowing layerwise customization of decomposition schemes for a given AI model. This fine-grained approach significantly reduces accuracy degradation by 7 to 20 percentage points compared to prior methods. On the hardware architecture side, the thesis enhances novel dataflow designs using Field Programmable Gate Arrays (FPGAs). A novel memory management methodology intelligently allocates both on-chip and off-chip memory resources, achieving a latency reduction of 25 to 48%. Additionally, the thesis bridges the gap between software and hardware through fast and efficient design automation techniques, cutting the search time by at least 10x. This integration demonstrates significant potential for the efficient acceleration of AI models and reduces repetitive engineering efforts in the rapidly evolving AI landscape.
Version
Open Access
Date Issued
2024-10-02
Date Awarded
01/01/2025
License URL
Advisor
Bouganis, Christos-Savvas
Sponsor
Imperial College London
Publisher Department
Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
