TY - JOUR
T1 - Frequency-Oriented HLS Accelerator Exploration via Congestion-Aware Mapping on Multi-Die FPGA
AU - Wang, Teng
AU - Gao, Yingxue
AU - Xue, Haoran
AU - Cheng, Qianyu
AU - Chen, Xianglan
AU - Gong, Lei
AU - Li, Changlong
AU - Wang, Chao
AU - Zhou, Xuehai
AU - Chen, Huaping
N1 - Publisher Copyright:
© 1982-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Researchers use high-level synthesis (HLS) to develop large-scale accelerators, leveraging extensive resources to handle intricate computational tasks. Recent Multi-Die FPGAs feature several super logic regions (SLR) within one device to meet the high resource demands of advanced computation. However, implementing a large-scale accelerator on multi-die FPGAs presents three obstacles. The first is the considerable delay penalty caused by the interconnectivity of accelerator modules across multi-die boundaries. The second is the congestion caused by modules’ excessive use of local resources. The third is that the increasing number of accelerator modules expands the design space, making it difficult to find the optimal design point. All of these pose limitations on the maximum achievable frequency for developing large-scale accelerator designs on multi-die FPGAs. Therefore, this paper proposes a frequency-oriented design exploration with the congestion-aware module mapping algorithm, FrqBooster. It combines modules’ location and resource consumption to provide coarse-grained floorplanning optimization. Firstly, we analyze the connections between accelerator modules and build optimization objectives. Secondly, we analyze the resource cost of each module in accelerator design and propose a two-stage resource-balancing strategy, combined with resource consumption, to alleviate local congestion. Finally, we design a congestion-aware algorithm to complete the overall exploration for achieving higher potential performance. In the experiments, results show that FrqBooster achieves maximum frequency improvements of 73.15% and 81.22% on the U250 and U280 platforms, respectively, compared to the default implementation in Vivado. Compared to related works on accelerator DSE scenarios, FrqBooster achieved a maximum improvement of 25.84×, resulting in better overall performance.
AB - Researchers use high-level synthesis (HLS) to develop large-scale accelerators, leveraging extensive resources to handle intricate computational tasks. Recent Multi-Die FPGAs feature several super logic regions (SLR) within one device to meet the high resource demands of advanced computation. However, implementing a large-scale accelerator on multi-die FPGAs presents three obstacles. The first is the considerable delay penalty caused by the interconnectivity of accelerator modules across multi-die boundaries. The second is the congestion caused by modules’ excessive use of local resources. The third is that the increasing number of accelerator modules expands the design space, making it difficult to find the optimal design point. All of these pose limitations on the maximum achievable frequency for developing large-scale accelerator designs on multi-die FPGAs. Therefore, this paper proposes a frequency-oriented design exploration with the congestion-aware module mapping algorithm, FrqBooster. It combines modules’ location and resource consumption to provide coarse-grained floorplanning optimization. Firstly, we analyze the connections between accelerator modules and build optimization objectives. Secondly, we analyze the resource cost of each module in accelerator design and propose a two-stage resource-balancing strategy, combined with resource consumption, to alleviate local congestion. Finally, we design a congestion-aware algorithm to complete the overall exploration for achieving higher potential performance. In the experiments, results show that FrqBooster achieves maximum frequency improvements of 73.15% and 81.22% on the U250 and U280 platforms, respectively, compared to the default implementation in Vivado. Compared to related works on accelerator DSE scenarios, FrqBooster achieved a maximum improvement of 25.84×, resulting in better overall performance.
KW - Floorplan
KW - Frequency
KW - HLS
KW - Multi-Die FPGA
UR - https://www.scopus.com/pages/publications/105040241373
U2 - 10.1109/TCAD.2026.3696756
DO - 10.1109/TCAD.2026.3696756
M3 - 文章
AN - SCOPUS:105040241373
SN - 0278-0070
JO - IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
JF - IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
ER -