面向异构GPU的大语言模型高效流水线并行训练
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金(62172107)


Efficient Pipeline Parallelism Training for Large Language Model on Heterogeneous GPU
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    大语言模型(large language model, LLM)预训练通常依赖大规模GPU (graphics processing unit)集群, 现有流水线并行方法大多基于同构假设, 难以适应算力、显存和通信存在显著差异的异构环境. 为此, 本文将异构流水线编排建模为模型切分、拓扑构建与任务调度的联合优化问题, 并提出自适应异构流水线并行训练方法AHPipe (adaptive heterogeneous pipeline parallelism). 该方法基于性能分析获取设备的计算、显存与通信开销, 利用模拟退火搜索模型切分与流水线拓扑方案, 并通过自适应任务调度在满足显存约束的前提下压缩流水线气泡. 基于包含NVIDIA V100、H20、RTX-4090和RTX-5090的异构GPU集群, 对760M–13B参数规模GPT-3模型的实验结果表明, AHPipe在异构场景下相较现有代表性方法最高可提升39.2%的训练吞吐量, 并能在显存紧张场景下生成可行训练方案.

    Abstract:

    Large language model pre-training typically relies on large-scale graphics processing unit (GPU) clusters. Existing pipeline parallelism methods are mostly based on assumptions of homogeneity and struggle to adapt to heterogeneous environments with significant differences in computing power, memory capacity, and communication. To address this issue, this study formulates heterogeneous pipeline orchestration as a joint optimization problem of model partitioning, topology construction, and task scheduling, and proposes adaptive heterogeneous pipeline parallelism (AHPipe). This method obtains device computation, memory, and communication costs through performance analysis, uses simulated annealing to search for model partitioning and pipeline topology schemes, and employs adaptive task scheduling to reduce pipeline bubbles under memory constraints. Experimental results on a heterogeneous GPU cluster consisting of NVIDIA V100, H20, RTX-4090, and RTX-5090 GPUs and covering GPT-3 models from 760M to 13B parameters show that the proposed method improves training throughput by up to 39.2% over representative existing methods in heterogeneous scenarios and can generate feasible training schemes under memory-constrained settings.

    参考文献
    相似文献
    引证文献
引用本文

聂绍珩,孙未未.面向异构GPU的大语言模型高效流水线并行训练.计算机系统应用,,():1-16

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-03-25
  • 最后修改日期:2026-04-16
  • 录用日期:
  • 在线发布日期: 2026-07-17
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号