跳到主要导航 跳到搜索 跳到主要内容

CoCoder_Repair: MISRA Code Rule Repair Based on Large Language Models

  • Zhenxin Fu*
  • , Guitao Cao
  • , Min Cai
  • *此作品的通讯作者
  • East China Normal University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

In the automotive industry, code quality detection and repair is a rigorous, complex, and extremely important task. It typically involves assessing whether the code meets relevant industry standard rules, such as the MISRA C/CPP rules and AutoSAR rules, to ensure code quality and manage defect risks. In traditional approaches, professional tools are used to detect defects in code projects, and developers manually repair these defects, which requires a significant investment of human resources and time. With the rapid rise of LLMs in the field of code generation in recent years, numerous tools for code review and repair have emerged, which can greatly improve the efficiency of developers. However, according to industry research, large language models with limited parameters (32b and below) currently perform poorly in the field of code quality repair. To address this issue, this study proposes a paradigm for post-training of vertical domain models for code quality detection and repair using methods such as supervised fine-tuning and reinforcement learning. By collecting, cleaning, labeling, and defining metrics for original code data, this approach aims to achieve a certain level of accuracy for vertical domain large language models with limited parameters (below 32b) in this task. In terms of target selection, the MISRA C/CPP 2008 rules are chosen as the code quality validation standard for this project. Based on actual projects, the most frequently violated rules in the MISRA C/CPP 2008 rule checks are selected as the code rules that need to be detected and modified. Relevant metrics for code repair, such as rule recognition rate, rule repair rate, and error guidance rate, are defined to evaluate the model's capabilities in the field of code specification repair. In the post-training of large language models, the qwen2.5-coder-base-7b model is used as the base model for supervised fine-tuning and reinforcement training. A more efficient reward strategy for this task is proposed by combining the GRPO algorithm with RLAIF. As a result, the performance in this task far exceeds that of similar models with the same parameter volume and even surpasses that of flagship models with larger parameter volumes.

源语言英语
主期刊名AICCC 2025 - 2025 8th Artificial Intelligence and Cloud Computing Conference
出版商Association for Computing Machinery, Inc
460-468
页数9
ISBN(电子版)9798400718892
DOI
出版状态已出版 - 4 5月 2026
活动2025 8th Artificial Intelligence and Cloud Computing Conference, AICCC 2025 - Tokyo, 日本
期限: 20 12月 202522 12月 2025

出版系列

姓名AICCC 2025 - 2025 8th Artificial Intelligence and Cloud Computing Conference

会议

会议2025 8th Artificial Intelligence and Cloud Computing Conference, AICCC 2025
国家/地区日本
Tokyo
时期20/12/2522/12/25

指纹

探究 'CoCoder_Repair: MISRA Code Rule Repair Based on Large Language Models' 的科研主题。它们共同构成独一无二的指纹。

引用此