Optimizing graph Neural Networks for jet tagging in particle physics on FPGAs
File(s)fpl22zq28_cr.pdf (877.96 KB)
Accepted version
Author(s)
Type
Conference Paper
Abstract
This work proposes a novel reconfigurable architecture for reducing the latency of JEDI-net, a Graph Neural
Network (GNN) based algorithm for jet tagging in particle
physics, which achieves state-of-the-art accuracy. Accelerating
JEDI-net is challenging since it requires low latency to deploy
the network for event selection at the CERN Large Hadron
Collider. This paper proposes an outer-product based matrix
multiplication approach customized for GNN-based JEDI-net,
which increases data spatial locality and reduces design latency.
It is further enhanced by code transformation with strength
reduction which exploits sparsity patterns and binary adjacency
matrices to increase hardware efficiency while reducing latency.
In addition, a customizable template for this architecture has
been designed and open-sourced, which enables the generation
of low-latency FPGA designs with efficient resource utilization
using high-level synthesis tools. Evaluation results show that our
FPGA implementation is up to 9.5 times faster and consumes up
to 6.5 times less power than a GPU implementation. Moreover,
the throughput of our FPGA design is sufficiently high to enable
deployment of JEDI-net in a sub-microsecond, real-time collider
trigger system, enabling it to benefit from improved accuracy.
Network (GNN) based algorithm for jet tagging in particle
physics, which achieves state-of-the-art accuracy. Accelerating
JEDI-net is challenging since it requires low latency to deploy
the network for event selection at the CERN Large Hadron
Collider. This paper proposes an outer-product based matrix
multiplication approach customized for GNN-based JEDI-net,
which increases data spatial locality and reduces design latency.
It is further enhanced by code transformation with strength
reduction which exploits sparsity patterns and binary adjacency
matrices to increase hardware efficiency while reducing latency.
In addition, a customizable template for this architecture has
been designed and open-sourced, which enables the generation
of low-latency FPGA designs with efficient resource utilization
using high-level synthesis tools. Evaluation results show that our
FPGA implementation is up to 9.5 times faster and consumes up
to 6.5 times less power than a GPU implementation. Moreover,
the throughput of our FPGA design is sufficiently high to enable
deployment of JEDI-net in a sub-microsecond, real-time collider
trigger system, enabling it to benefit from improved accuracy.
Date Issued
2023-02-13
Date Acceptance
2022-06-14
Citation
2022 32nd International Conference on Field-Programmable Logic and Applications (FPL), 2023, pp.327-333
Publisher
IEEE
Start Page
327
End Page
333
Journal / Book Title
2022 32nd International Conference on Field-Programmable Logic and Applications (FPL)
Copyright Statement
© 2023 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Source
International Conference on Field Programmable Logic and Applications
Publication Status
Published
Start Date
2022-08-29
Finish Date
2022-09-02
Coverage Spatial
Queens University Belfast, United Kingdom.