跳到主要导航 跳到搜索 跳到主要内容

SemInfer: Accelerating LLM-Based Semantic Data Processing via Sparse Indexing

  • East China Normal University
  • State Grid Corporation of China

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Integrating LLMs for data processing enables semantic querying but causes GPU memory bottlenecks and redundant computations. We present SemInfer, an acceleration system for batch semantic processing that treats the KV Cache as a semantic index, offloading pre-computed caches to host storage to eliminate redundancy. To reduce the index size, we propose a pruning strategy based on last-layer aggregated attention to accurately retain critical semantic tokens. Furthermore, we employ a pipeline mechanism to enable the asynchronous overlapping of CPU-GPU transmission and inference computation. This demonstration showcases the complete workflow of SemInfer on the IMDB dataset, achieving up to a 16.3x inference speedup over direct LLM inference and a 90% reduction in index size with few semantic accuracy loss.

源语言英语
主期刊名Database Systems for Advanced Applications - 31st International Conference, DASFAA 2026, Proceedings
编辑Hyungsoo Jung, Tianzheng Wang, Masashi Toyoda, Hyuk-Yoon Kwon, Jae-woong Lee
出版商Springer Science and Business Media Deutschland GmbH
703-707
页数5
ISBN(印刷版)9789819203772
DOI
出版状态已出版 - 2026
活动31st International Conference on Database Systems for Advanced Applications, DASFAA 2026 - Jeju, 韩国
期限: 27 4月 202630 4月 2026

出版系列

姓名Lecture Notes in Computer Science
16540 LNCS
ISSN(印刷版)0302-9743
ISSN(电子版)1611-3349

会议

会议31st International Conference on Database Systems for Advanced Applications, DASFAA 2026
国家/地区韩国
Jeju
时期27/04/2630/04/26

指纹

探究 'SemInfer: Accelerating LLM-Based Semantic Data Processing via Sparse Indexing' 的科研主题。它们共同构成独一无二的指纹。

引用此