跳到主要导航 跳到搜索 跳到主要内容

大数据处理系统中面向GPU 加速DNN 推理的模型共享

  • East China Normal University

科研成果: 期刊稿件文章同行评审

摘要

Big data processing is being widely used in academia and industry to handle DNN-based inference workloads for fields such as video analyses. In such cases, multiple parallel inference tasks in the big data processing system repeatedly load the same, read-only DNN model so the system does not fully utilize the GPU resources which creates a bottleneck that limits the inference performance. This paper presents a model sharing technique for single GPU cards that enables sharing of the same model among various DNN inference tasks. An allocator is used to make the model sharing technique work for each GPU in the distributed environment. This method was implemented in Spark on a GPU platform in a distributed data processing system that supports large-scale inference workloads. Tests show that for video analyses on the YOLO-v3 model, the model sharing reduces the GPU memory overhead and improves system throughput by up to 136% compared to a system without the model sharing technique.

投稿的翻译标题Model sharing for GPU-accelerated DNN inference in big data processing systems
源语言繁体中文
页(从-至)1435-1441
页数7
期刊Qinghua Daxue Xuebao/Journal of Tsinghua University
62
9
DOI
出版状态已出版 - 15 9月 2022

关键词

  • DNN inference
  • GPU
  • GPU memory
  • big data processing system
  • model sharing

指纹

探究 '大数据处理系统中面向GPU 加速DNN 推理的模型共享' 的科研主题。它们共同构成独一无二的指纹。

引用此