跳到主要导航 跳到搜索 跳到主要内容

Frequency-Oriented HLS Accelerator Exploration via Congestion-Aware Mapping on Multi-Die FPGA

  • Teng Wang
  • , Yingxue Gao
  • , Haoran Xue
  • , Qianyu Cheng
  • , Xianglan Chen
  • , Lei Gong*
  • , Changlong Li
  • , Chao Wang*
  • , Xuehai Zhou
  • , Huaping Chen
  • *此作品的通讯作者
  • University of Science and Technology of China
  • Nanjing University of Posts and Telecommunications

科研成果: 期刊稿件文章同行评审

摘要

Researchers use high-level synthesis (HLS) to develop large-scale accelerators, leveraging extensive resources to handle intricate computational tasks. Recent Multi-Die FPGAs feature several super logic regions (SLR) within one device to meet the high resource demands of advanced computation. However, implementing a large-scale accelerator on multi-die FPGAs presents three obstacles. The first is the considerable delay penalty caused by the interconnectivity of accelerator modules across multi-die boundaries. The second is the congestion caused by modules’ excessive use of local resources. The third is that the increasing number of accelerator modules expands the design space, making it difficult to find the optimal design point. All of these pose limitations on the maximum achievable frequency for developing large-scale accelerator designs on multi-die FPGAs. Therefore, this paper proposes a frequency-oriented design exploration with the congestion-aware module mapping algorithm, FrqBooster. It combines modules’ location and resource consumption to provide coarse-grained floorplanning optimization. Firstly, we analyze the connections between accelerator modules and build optimization objectives. Secondly, we analyze the resource cost of each module in accelerator design and propose a two-stage resource-balancing strategy, combined with resource consumption, to alleviate local congestion. Finally, we design a congestion-aware algorithm to complete the overall exploration for achieving higher potential performance. In the experiments, results show that FrqBooster achieves maximum frequency improvements of 73.15% and 81.22% on the U250 and U280 platforms, respectively, compared to the default implementation in Vivado. Compared to related works on accelerator DSE scenarios, FrqBooster achieved a maximum improvement of 25.84×, resulting in better overall performance.

学术指纹

探究 'Frequency-Oriented HLS Accelerator Exploration via Congestion-Aware Mapping on Multi-Die FPGA' 的科研主题。它们共同构成独一无二的学术指纹。

引用此