###
计算机系统应用英文版:2026,35(8):39-49
本文二维码信息
码上扫一扫!
CodeVulnGen: 基于LLM的代码漏洞数据集生成
(1.武汉科技大学 计算机科学与技术学院, 武汉 430065;2.武汉科技大学 智能信息处理与实时工业系统湖北省重点实验室, 武汉 430065;3.武昌理工学院 商学院, 武汉 430223)
CodeVulnGen: Code Vulnerability Datasets Generation Based on LLM
(1.School of Computer Science and Technology, Wuhan University of Science and Technology, Wuhan 430065, China;2.Hubei Province Key Laboratory of Intelligent Information Processing and Real-time Industrial System, Wuhan University of Science and Technology, Wuhan 430065, China;3.School of Business, Wuchang University of Technology, Wuhan 430223, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 38次   下载 50
Received:December 09, 2025    Revised:January 05, 2026
中文摘要: 高质量的漏洞代码数据集对于推动软件安全研究具有重要意义, 是支撑漏洞分析、漏洞检测模型训练与评测、系统对比与攻防研究的核心资源. 不同于图像等领域, 现有漏洞数据集不仅数量有限, 且质量参差不齐, 在多样性上仍存在明显不足. 为应对这一挑战, 本文提出CodeVulnGen, 一种以漏洞特征为核心、基于大语言模型(large language model, LLM)的自动化代码漏洞数据生成框架. 该框架系统化定义了代码漏洞的数据质量属性与评估方法, 在此基础上, 提出3阶段递进式提示构造策略, 将结构化描述、少样本实例、真实漏洞特征和思维链(chain-of-thought, CoT)推理逐步融合, 以生成多样、规范且更贴近真实攻击场景的高复杂度漏洞代码. 实验结果表明, 使用该数据集进行微调的Qwen-4B模型F1指标提升7.4%, 相较于基于真实数据训练提升5.5%.
Abstract:High-quality vulnerable code datasets are essential for advancing software security research. Such datasets constitute core resources for vulnerability analysis, vulnerability detection model training and evaluation, system comparison, and attack-defense research. Unlike datasets in domains such as computer vision, existing vulnerability datasets are limited in scale and often suffer from inconsistent quality and insufficient diversity. To address these limitations, this study proposes CodeVulnGen, an automated code vulnerability data generation framework based on large language models (LLMs) and centered on vulnerability characteristics. The framework systematically defines the quality attributes of vulnerable code and the corresponding evaluation methods. Based on this, a three-stage prompt construction strategy is proposed that progressively integrates structured descriptions, few-shot examples, real-world vulnerability characteristics, and chain-of-thought reasoning to generate diverse, well-structured, and highly complex vulnerable code that more closely resembles real-world attack scenarios. Experimental results show that fine-tuning the Qwen-4B model on the generated dataset improves the F1 score by 7.4% and yields a 5.5% performance improvement compared with training on real-world datasets.
文章编号:     中图分类号:    文献标志码:
基金项目:武汉市重点研发计划 (2022012202015070); 武汉东湖新技术开发区“揭榜挂帅”项目 (2022KJB126)
引用文本:
代雪晴,李涛,郭佳雄,伍章驰,何柳.CodeVulnGen: 基于LLM的代码漏洞数据集生成.计算机系统应用,2026,35(8):39-49
DAI Xue-Qing,LI Tao,GUO Jia-Xiong,WU Zhang-Chi,HE Liu.CodeVulnGen: Code Vulnerability Datasets Generation Based on LLM.COMPUTER SYSTEMS APPLICATIONS,2026,35(8):39-49