Hardware-aware neural networks
File(s)
Author(s)
Andronic, Marta
Type
Thesis or dissertation
Abstract
This thesis tackles the challenge of deploying neural networks (NNs) on resource-constrained hardware, focusing on hardware-aware design for FPGA look-up table (LUT)-based NNs. These NNs offer ultra-low latency and reduced resource usage but typically rely on low-precision sparse models due to hardware constraints, leading to accuracy degradation.
Firstly, this thesis addresses the computational block problem: determining which real-valued function each LUT should model during training. Two methodologies are proposed: PolyLUT and NeuraLUT. PolyLUT expands each neuron’s feature vector with monomials up to a user-defined degree, and NeuraLUT embeds small MLPs within each LUT to maximize computational density. Both methods exploit the flexibility of LUTs to implement complex functions; however, they also expose intricate interactions among a limited number of input features, constrained by the LUT fan-in.
The second half of this thesis focuses on the connectivity problem: how to interconnect these expressive, fan-in-limited units efficiently. Two further methodologies are proposed. The first is a hardware-aware structured pruning technique using a custom group regularizer designed to guide neuron connections according to hardware constraints. The second is NeuraLUT-Assemble, a fully parametrizable framework that assembles multiple NeuraLUT neurons into tree structures with larger effective fan-in.
This thesis demonstrates that LUT-based neural networks can achieve competitive accuracy while delivering inference latencies below 10ns. PolyLUT and NeuraLUT achieve an 8.1× reduction in resource utilization and a 4.3× reduction in latency for equivalent accuracy relative to prior baselines such as LogicNets. The pruning framework further improves accuracy by up to 1.5 percentage points. Finally, NeuraLUT-Assemble bridges the remaining accuracy gap, reaching within 1 percentage point of full-precision fully connected models while achieving 3 percentage points higher accuracy and over 10× lower area-delay product than LogicNets. These results demonstrate that LUT-based networks can reconcile expressivity with scalability in efficient hardware-aware neural design.
Firstly, this thesis addresses the computational block problem: determining which real-valued function each LUT should model during training. Two methodologies are proposed: PolyLUT and NeuraLUT. PolyLUT expands each neuron’s feature vector with monomials up to a user-defined degree, and NeuraLUT embeds small MLPs within each LUT to maximize computational density. Both methods exploit the flexibility of LUTs to implement complex functions; however, they also expose intricate interactions among a limited number of input features, constrained by the LUT fan-in.
The second half of this thesis focuses on the connectivity problem: how to interconnect these expressive, fan-in-limited units efficiently. Two further methodologies are proposed. The first is a hardware-aware structured pruning technique using a custom group regularizer designed to guide neuron connections according to hardware constraints. The second is NeuraLUT-Assemble, a fully parametrizable framework that assembles multiple NeuraLUT neurons into tree structures with larger effective fan-in.
This thesis demonstrates that LUT-based neural networks can achieve competitive accuracy while delivering inference latencies below 10ns. PolyLUT and NeuraLUT achieve an 8.1× reduction in resource utilization and a 4.3× reduction in latency for equivalent accuracy relative to prior baselines such as LogicNets. The pruning framework further improves accuracy by up to 1.5 percentage points. Finally, NeuraLUT-Assemble bridges the remaining accuracy gap, reaching within 1 percentage point of full-precision fully connected models while achieving 3 percentage points higher accuracy and over 10× lower area-delay product than LogicNets. These results demonstrate that LUT-based networks can reconcile expressivity with scalability in efficient hardware-aware neural design.
Version
Open Access
Date Issued
2026-03-02
Date Awarded
2026-07-01
Copyright Statement
Attribution-NonCommercial 4.0 International Licence (CC BY-NC)
License URL
Advisor
Constantinides, George
Sponsor
Engineering and Physical Sciences Research Council
Publisher Department
Department of Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
