摘要
Big data processing is being widely used in academia and industry to handle DNN-based inference workloads for fields such as video analyses. In such cases, multiple parallel inference tasks in the big data processing system repeatedly load the same, read-only DNN model so the system does not fully utilize the GPU resources which creates a bottleneck that limits the inference performance. This paper presents a model sharing technique for single GPU cards that enables sharing of the same model among various DNN inference tasks. An allocator is used to make the model sharing technique work for each GPU in the distributed environment. This method was implemented in Spark on a GPU platform in a distributed data processing system that supports large-scale inference workloads. Tests show that for video analyses on the YOLO-v3 model, the model sharing reduces the GPU memory overhead and improves system throughput by up to 136% compared to a system without the model sharing technique.
| 投稿的翻译标题 | Model sharing for GPU-accelerated DNN inference in big data processing systems |
|---|---|
| 源语言 | 繁体中文 |
| 页(从-至) | 1435-1441 |
| 页数 | 7 |
| 期刊 | Qinghua Daxue Xuebao/Journal of Tsinghua University |
| 卷 | 62 |
| 期 | 9 |
| DOI | |
| 出版状态 | 已出版 - 15 9月 2022 |
关键词
- DNN inference
- GPU
- GPU memory
- big data processing system
- model sharing
指纹
探究 '大数据处理系统中面向GPU 加速DNN 推理的模型共享' 的科研主题。它们共同构成独一无二的指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver