Skip to main navigation Skip to search Skip to main content

SemInfer: Accelerating LLM-Based Semantic Data Processing via Sparse Indexing

  • East China Normal University
  • State Grid Corporation of China

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Integrating LLMs for data processing enables semantic querying but causes GPU memory bottlenecks and redundant computations. We present SemInfer, an acceleration system for batch semantic processing that treats the KV Cache as a semantic index, offloading pre-computed caches to host storage to eliminate redundancy. To reduce the index size, we propose a pruning strategy based on last-layer aggregated attention to accurately retain critical semantic tokens. Furthermore, we employ a pipeline mechanism to enable the asynchronous overlapping of CPU-GPU transmission and inference computation. This demonstration showcases the complete workflow of SemInfer on the IMDB dataset, achieving up to a 16.3x inference speedup over direct LLM inference and a 90% reduction in index size with few semantic accuracy loss.

Original languageEnglish
Title of host publicationDatabase Systems for Advanced Applications - 31st International Conference, DASFAA 2026, Proceedings
EditorsHyungsoo Jung, Tianzheng Wang, Masashi Toyoda, Hyuk-Yoon Kwon, Jae-woong Lee
PublisherSpringer Science and Business Media Deutschland GmbH
Pages703-707
Number of pages5
ISBN (Print)9789819203772
DOIs
StatePublished - 2026
Event31st International Conference on Database Systems for Advanced Applications, DASFAA 2026 - Jeju, Korea, Republic of
Duration: 27 Apr 202630 Apr 2026

Publication series

NameLecture Notes in Computer Science
Volume16540 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference31st International Conference on Database Systems for Advanced Applications, DASFAA 2026
Country/TerritoryKorea, Republic of
CityJeju
Period27/04/2630/04/26

Keywords

  • Data Processing
  • KV Cache Pruning
  • LLM

Fingerprint

Dive into the research topics of 'SemInfer: Accelerating LLM-Based Semantic Data Processing via Sparse Indexing'. Together they form a unique fingerprint.

Cite this