Reconfigurable acceleration of convolutional neural network ensemble
File(s)
Author(s)
Kwan, Pok Yee
Type
Thesis
Abstract
This thesis aims to improve the accuracy of neural networks, while reducing computation time and resources required using reconfigurable hardware. The key idea is to apply an ensemble of quantized neural networks instead of one network to enhance parallelism while boosting the output accuracy by the ensemble diversity.
The first contribution is a novel lossy multiport memory capable of high memory bandwidth, providing concurrent access to a single address space through multiple ports, facilitating memory access for every individual network in the ensemble. The lossiness introduced relaxes the memory request constraints, leading to improved performance for random memory access. Significant improvement in clock frequency and resource usage is achieved over state-of-the-art FPGA implementation of multiport memory.
The second contribution is a light-weight permutation generator for efficient data augmentation to diversify the ensemble. A novel scalable architecture is presented to automatically generate restricted local permutation efficiently, preserving the spatial correlation of the original image. The generated dataset is as good as the original dataset, with a significant speed up in generation compared to CPU and GPU implementation, with no extra memory storage or transfer.
The third contribution is a framework to utilize resources for multiple neural networks in the ensemble. The framework explores the neural network structures and determines the optimal parallelism configuration for individual parts of the ensemble, facilitating the trade-off between resources and run time. It also supports input of variable bit widths, further improving the diversity of the ensemble. The workflow for designing an integrated neural network ensemble consisting of the lossy multiport memory, light-weight permutation generator, with the optimal neural network configuration is presented with a real-life application.
The first contribution is a novel lossy multiport memory capable of high memory bandwidth, providing concurrent access to a single address space through multiple ports, facilitating memory access for every individual network in the ensemble. The lossiness introduced relaxes the memory request constraints, leading to improved performance for random memory access. Significant improvement in clock frequency and resource usage is achieved over state-of-the-art FPGA implementation of multiport memory.
The second contribution is a light-weight permutation generator for efficient data augmentation to diversify the ensemble. A novel scalable architecture is presented to automatically generate restricted local permutation efficiently, preserving the spatial correlation of the original image. The generated dataset is as good as the original dataset, with a significant speed up in generation compared to CPU and GPU implementation, with no extra memory storage or transfer.
The third contribution is a framework to utilize resources for multiple neural networks in the ensemble. The framework explores the neural network structures and determines the optimal parallelism configuration for individual parts of the ensemble, facilitating the trade-off between resources and run time. It also supports input of variable bit widths, further improving the diversity of the ensemble. The workflow for designing an integrated neural network ensemble consisting of the lossy multiport memory, light-weight permutation generator, with the optimal neural network configuration is presented with a real-life application.
Version
Open Access
Date Issued
2023-10
Date Awarded
2024-03
Copyright Statement
Creative Commons Attribution NonCommercial Licence
License URL
Advisor
Luk, Wayne
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
