TY - GEN
T1 - Dr.avx
T2 - 24th IEEE/ACM International Symposium on Code Generation and Optimization, CGO 2026
AU - Tang, Yue
AU - Wu, Mianzhi
AU - Li, Yufeng
AU - Liao, Haoyu
AU - Guo, Jianmei
AU - Huang, Bo
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Modern processors are breaking a fundamental rule: backward compatibility within their own ISA families. We term this Generational ISA Fragmentation (GIF), where newer processors cannot execute instructions supported by prior generations within the same ISA family. This phenomenon is exemplified by Intel's removal of AVX-512 from Alder Lake processors after years of deployment, ARM's inconsistent support for SVE across cores, and RISC-V's incompatible vector specifications. GIF causes illegal instruction crashes when running applications optimized for earlier processors on newer hardware, threatening the foundation of software portability that has underpinned decades of computing evolution.We introduce Dr.avx, a dynamic compilation system that enables seamless execution of AVX-512 instructions on hardware that lacks native support. Dr.avx addresses the most instructive GIF instance, x86 AVX-512 fragmentation, by targeting the integer and floating-point operations that dominate real workloads. Our rewrite engine performs a fine-grained classification of AVX-512 opcode-operand patterns and employs three complementary strategies: Instr Mirroring, AVX Lowering, and Scalar Fallback. Experiments show that Dr.avx incurs a geometric mean overhead of 1.44× on SPEC CINT2017 relative to native AVX-512 execution, 17.3% better than Intel's closed-source SDE. On production databases, Dr.avx sustains 75%-88% (MySQL) and 86%-99% (MongoDB) of native throughput, yielding 2.0-2.7× higher throughput than SDE. For LLM inference (llama.cpp), Dr.avx keeps 95%-99% of native tokens/s and delivers 2.5-4.8× speedup over SDE. Unlike Intel's proprietary SDE, which provides no visibility into its implementation details, Dr.avx achieves functional correctness while providing an open, extensible, near-native performance implementation. Our work offers both a remedy for AVX-512 fragmentation and a blueprint for addressing similar compatibility challenges emerging across all major ISAs.
AB - Modern processors are breaking a fundamental rule: backward compatibility within their own ISA families. We term this Generational ISA Fragmentation (GIF), where newer processors cannot execute instructions supported by prior generations within the same ISA family. This phenomenon is exemplified by Intel's removal of AVX-512 from Alder Lake processors after years of deployment, ARM's inconsistent support for SVE across cores, and RISC-V's incompatible vector specifications. GIF causes illegal instruction crashes when running applications optimized for earlier processors on newer hardware, threatening the foundation of software portability that has underpinned decades of computing evolution.We introduce Dr.avx, a dynamic compilation system that enables seamless execution of AVX-512 instructions on hardware that lacks native support. Dr.avx addresses the most instructive GIF instance, x86 AVX-512 fragmentation, by targeting the integer and floating-point operations that dominate real workloads. Our rewrite engine performs a fine-grained classification of AVX-512 opcode-operand patterns and employs three complementary strategies: Instr Mirroring, AVX Lowering, and Scalar Fallback. Experiments show that Dr.avx incurs a geometric mean overhead of 1.44× on SPEC CINT2017 relative to native AVX-512 execution, 17.3% better than Intel's closed-source SDE. On production databases, Dr.avx sustains 75%-88% (MySQL) and 86%-99% (MongoDB) of native throughput, yielding 2.0-2.7× higher throughput than SDE. For LLM inference (llama.cpp), Dr.avx keeps 95%-99% of native tokens/s and delivers 2.5-4.8× speedup over SDE. Unlike Intel's proprietary SDE, which provides no visibility into its implementation details, Dr.avx achieves functional correctness while providing an open, extensible, near-native performance implementation. Our work offers both a remedy for AVX-512 fragmentation and a blueprint for addressing similar compatibility challenges emerging across all major ISAs.
KW - AVX-512
KW - Dynamic Binary Translation
KW - ISA compatibility
KW - SIMD
UR - https://www.scopus.com/pages/publications/105041766814
U2 - 10.1109/CGO68049.2026.11394840
DO - 10.1109/CGO68049.2026.11394840
M3 - 会议稿件
AN - SCOPUS:105041766814
T3 - CGO 2026 - Proceedings of the 2026 IEEE/ACM International Symposium on Code Generation and Optimization
SP - 403
EP - 415
BT - CGO 2026 - Proceedings of the 2026 IEEE/ACM International Symposium on Code Generation and Optimization
A2 - Blackburn, Stephen M.
A2 - Cohen, Albert
A2 - Jones, Timothy M.
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 31 January 2026 through 4 February 2026
ER -