面向非对称并行策略的高效内存检查点
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:


Enhanced In-memory Checkpointing for Asymmetric Parallelism Strategies
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    大语言模型训练周期长、计算代价高, 检查点技术通过周期性保存训练状态以支持故障后从最近的检查点恢复. 内存检查点将训练状态存储于CPU内存并复制到对等节点以增强数据的容错性, 在对称并行策略下可以有效将检查点传输调度至通信的空闲窗口. 然而, 非对称并行策略因能够更灵活地适配模型结构与硬件资源、提升训练效率而被广泛采用, 但也带来了3个核心挑战: 参数在设备间呈现混合且细粒度的冗余分布模式, 导致冗余识别困难; 机内与跨机通信信道独立, 且进程仅具局部视图导致空闲窗口检测失效; 现有系统缺乏全局最优冗余组构建能力, 无法充分利用非对称策略下的冗余与空闲资源. 为此, 本文提出一个面向非对称并行策略的内存检查点系统Mosaic, 采用分析-运行的两阶段架构, 在分析阶段通过细粒度参数冗余分析与消除、通信流量建模与空闲窗口检测以及全局最优冗余组选择生成最优传输计划, 并在运行阶段将检查点流量穿插至空闲窗口中执行. 实验结果表明, 相较于基线系统, Mosaic在非对称并行配置下能将额外训练开销降低42.0%–50.4%.

    Abstract:

    Large language model (LLM) training is time-consuming and computationally expensive. Checkpointing periodically saves training states to enable recovery from the most recent checkpoint after failures. In-memory checkpointing stores training states in CPU memory, replicates them to peer nodes to enhance data fault tolerance, and can effectively schedule checkpoint transfers within communication idle windows under symmetric parallelism. However, asymmetric parallelism strategies are widely adopted because they can adapt more flexibly to model architectures and hardware resources and improve training efficiency, but they also introduce three core challenges: parameters exhibit mixed and fine-grained redundancy distribution patterns across devices, making redundancy identification difficult; intra-node and inter-node communication channels are independent, and each process has only a local view, rendering idle window detection ineffective; and existing systems lack the ability to construct globally optimal redundancy groups, thus failing to fully exploit redundancy and idle resources under asymmetric strategies. To this end, this study proposes Mosaic, an in-memory checkpointing system for asymmetric parallelism. Mosaic adopts a two-phase analysis-runtime architecture. In the analysis phase, the proposed system generates an optimal transfer plan through fine-grained parameter redundancy analysis and elimination, communication traffic modeling and idle window detection, and globally optimal redundancy group selection. In the runtime phase, checkpoint traffic is interleaved into idle windows for execution. Experimental results show that, compared with baseline systems, Mosaic reduces additional training overhead by 42.0% to 50.4% under asymmetric parallelism configurations.

    参考文献
    相似文献
    引证文献
引用本文

王晨晗,孙婧雯,白有辉,李诚.面向非对称并行策略的高效内存检查点.计算机系统应用,,():1-11

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-03-10
  • 最后修改日期:2026-03-30
  • 录用日期:
  • 在线发布日期: 2026-07-20
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号