跳到主要导航 跳到搜索 跳到主要内容

Dr.avx: A Dynamic Compilation System for Seamlessly Executing Hardware-Unsupported Vectorization Instructions

  • East China Normal University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Modern processors are breaking a fundamental rule: backward compatibility within their own ISA families. We term this Generational ISA Fragmentation (GIF), where newer processors cannot execute instructions supported by prior generations within the same ISA family. This phenomenon is exemplified by Intel's removal of AVX-512 from Alder Lake processors after years of deployment, ARM's inconsistent support for SVE across cores, and RISC-V's incompatible vector specifications. GIF causes illegal instruction crashes when running applications optimized for earlier processors on newer hardware, threatening the foundation of software portability that has underpinned decades of computing evolution.We introduce Dr.avx, a dynamic compilation system that enables seamless execution of AVX-512 instructions on hardware that lacks native support. Dr.avx addresses the most instructive GIF instance, x86 AVX-512 fragmentation, by targeting the integer and floating-point operations that dominate real workloads. Our rewrite engine performs a fine-grained classification of AVX-512 opcode-operand patterns and employs three complementary strategies: Instr Mirroring, AVX Lowering, and Scalar Fallback. Experiments show that Dr.avx incurs a geometric mean overhead of 1.44× on SPEC CINT2017 relative to native AVX-512 execution, 17.3% better than Intel's closed-source SDE. On production databases, Dr.avx sustains 75%-88% (MySQL) and 86%-99% (MongoDB) of native throughput, yielding 2.0-2.7× higher throughput than SDE. For LLM inference (llama.cpp), Dr.avx keeps 95%-99% of native tokens/s and delivers 2.5-4.8× speedup over SDE. Unlike Intel's proprietary SDE, which provides no visibility into its implementation details, Dr.avx achieves functional correctness while providing an open, extensible, near-native performance implementation. Our work offers both a remedy for AVX-512 fragmentation and a blueprint for addressing similar compatibility challenges emerging across all major ISAs.

源语言英语
主期刊名CGO 2026 - Proceedings of the 2026 IEEE/ACM International Symposium on Code Generation and Optimization
编辑Stephen M. Blackburn, Albert Cohen, Timothy M. Jones
出版商Institute of Electrical and Electronics Engineers Inc.
403-415
页数13
ISBN(电子版)9798331592882
DOI
出版状态已出版 - 2026
活动24th IEEE/ACM International Symposium on Code Generation and Optimization, CGO 2026 - Sydney, 澳大利亚
期限: 31 1月 20264 2月 2026

出版系列

姓名CGO 2026 - Proceedings of the 2026 IEEE/ACM International Symposium on Code Generation and Optimization

会议

会议24th IEEE/ACM International Symposium on Code Generation and Optimization, CGO 2026
国家/地区澳大利亚
Sydney
时期31/01/264/02/26

学术指纹

探究 'Dr.avx: A Dynamic Compilation System for Seamlessly Executing Hardware-Unsupported Vectorization Instructions' 的科研主题。它们共同构成独一无二的学术指纹。

引用此