ITERA-LLM: Boosting sub-8-bit large language model inference via iterative tensor decomposition
File(s) ITERA-LLM.pdf (2.86 MB)
Accepted version
Author(s)
Huang, Yinting
Zheng, Keran
Yu, Zhewen
Bouganis, Christos-Savvas
Type
Conference Paper
Abstract
Recent advancements in Large Language Models (LLMs) have demonstrated impressive capabilities as their scale expands to billions of parameters. Deploying these large-scale models on resource-constrained platforms presents significant challenges, with post-training fixed-point quantization often used as a model compression technique. However, quantization-only methods typically lead to significant accuracy degradation in LLMs when precision falls below 8 bits. This paper addresses this challenge through a software-hardware co-design framework,
ITERA-LLM, which integrates sub-8-bit quantization with SVD-based iterative low-rank tensor decomposition for error compensation, leading to higher compression ratios and reduced computational complexity. The proposed approach is complemented by a hardware-aware Design Space Exploration (DSE) process that optimizes accuracy, latency, and resource utilization, tailoring the configuration to the specific requirements of the targeted LLM. Our results show that ITERA-LLM achieves linear layer latency reduction of up to 41.1%, compared to quantization-only baseline approach while maintaining similar model accuracy.
ITERA-LLM, which integrates sub-8-bit quantization with SVD-based iterative low-rank tensor decomposition for error compensation, leading to higher compression ratios and reduced computational complexity. The proposed approach is complemented by a hardware-aware Design Space Exploration (DSE) process that optimizes accuracy, latency, and resource utilization, tailoring the configuration to the specific requirements of the targeted LLM. Our results show that ITERA-LLM achieves linear layer latency reduction of up to 41.1%, compared to quantization-only baseline approach while maintaining similar model accuracy.
Date Issued
2025-05-28
Date Acceptance
2025-03-11
Citation
2025 IEEE 33rd Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM), 2025, pp.114-122
ISBN
979-8-3315-0281-2
ISSN
2576-2621
Publisher
IEEE
Start Page
114
End Page
122
Journal / Book Title
2025 IEEE 33rd Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM)
Copyright Statement
Copyright © 2025 IEEE. This is the author’s accepted manuscript made available under a CC-BY licence in accordance with Imperial’s Research Publications Open Access policy (www.imperial.ac.uk/oa-policy)
License URL
Source
33rd IEEE International Symposium On Field-Programmable Custom Computing Machines
Publication Status
Published
Start Date
2025-05-04
Finish Date
2025-05-07
Coverage Spatial
Fayetteville, Arkansas, USA
Date Publish Online
2025-05-28
