CodeVulnGen: 基于LLM的代码漏洞数据集生成
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

武汉市重点研发计划 (2022012202015070); 武汉东湖新技术开发区“揭榜挂帅”项目 (2022KJB126)


CodeVulnGen: Code Vulnerability Datasets Generation Based on LLM
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    高质量的漏洞代码数据集对于推动软件安全研究具有重要意义, 是支撑漏洞分析、漏洞检测模型训练与评测、系统对比与攻防研究的核心资源. 不同于图像等领域, 现有漏洞数据集不仅数量有限, 且质量参差不齐, 在多样性上仍存在明显不足. 为应对这一挑战, 本文提出CodeVulnGen, 一种以漏洞特征为核心、基于大语言模型(large language model, LLM)的自动化代码漏洞数据生成框架. 该框架系统化定义了代码漏洞的数据质量属性与评估方法, 在此基础上, 提出3阶段递进式提示构造策略, 将结构化描述、少样本实例、真实漏洞特征和思维链(chain-of-thought, CoT)推理逐步融合, 以生成多样、规范且更贴近真实攻击场景的高复杂度漏洞代码. 实验结果表明, 使用该数据集进行微调的Qwen-4B模型F1指标提升7.4%, 相较于基于真实数据训练提升5.5%.

    Abstract:

    High-quality vulnerable code datasets are essential for advancing software security research. Such datasets constitute core resources for vulnerability analysis, vulnerability detection model training and evaluation, system comparison, and attack-defense research. Unlike datasets in domains such as computer vision, existing vulnerability datasets are limited in scale and often suffer from inconsistent quality and insufficient diversity. To address these limitations, this study proposes CodeVulnGen, an automated code vulnerability data generation framework based on large language models (LLMs) and centered on vulnerability characteristics. The framework systematically defines the quality attributes of vulnerable code and the corresponding evaluation methods. Based on this, a three-stage prompt construction strategy is proposed that progressively integrates structured descriptions, few-shot examples, real-world vulnerability characteristics, and chain-of-thought reasoning to generate diverse, well-structured, and highly complex vulnerable code that more closely resembles real-world attack scenarios. Experimental results show that fine-tuning the Qwen-4B model on the generated dataset improves the F1 score by 7.4% and yields a 5.5% performance improvement compared with training on real-world datasets.

    参考文献
    相似文献
    引证文献
引用本文

代雪晴,李涛,郭佳雄,伍章驰,何柳. CodeVulnGen: 基于LLM的代码漏洞数据集生成.计算机系统应用,2026,35(8):39-49

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-12-09
  • 最后修改日期:2026-01-05
  • 录用日期:
  • 在线发布日期: 2026-05-22
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号