本文已被:浏览 21次 下载 27次
Received:December 24, 2025 Revised:February 27, 2026
Received:December 24, 2025 Revised:February 27, 2026
中文摘要: 以最大完工时间和总能耗为目标, 建立了考虑自动引导小车(automated guided vehicle, AGV)变速和充电的绿色作业车间调度模型. 针对工序加工排序和AGV运输任务分配的协同优化调度需求, 考虑决策空间复杂性与动态关联性, 提出了一种多策略近端策略优化(multi-strategy proximal policy optimization, MSPPO)算法. 设计基于析取图和人工状态的复合状态特征, 为智能体提供更全面的环境信息. 结合多目标优化需求, 构建融合完工时间与能耗的复合奖励函数, 引导解向Pareto前沿逼近. 设计一种双策略协同网络架构, 工序策略网络负责工序加工排序, AGV策略网络负责AGV运输任务分配, 两个网络顺序执行构成完整的调度动作, 实现加工与运输的协同优化. 为增加策略网络间的协同性与适应性, 通过共享优势估计、策略反馈机制以及联合状态更新, 使两个策略网络在训练过程中能够动态适应彼此的决策行为, 提升算法的鲁棒性与收敛性. 通过拓展算例的实验分析对比, 验证算法的有效性与优越性.
Abstract:This study builds a green job-shop scheduling model considering speed-adjustable operation and charging of automated guided vehicle (AGV), aiming to minimize the makespan and total energy consumption. A multi-strategy proximal policy optimization (MSPPO) algorithm is proposed to meet the collaborative optimization scheduling requirement between operation sequencing and AGV transportation task assignment, with the complexity and dynamic correlation of decision space considered. A composite state representation is designed by integrating disjunctive graphs with artificial states, providing the agent with more comprehensive environmental information. A composite reward function combining the makespan and energy consumption is established by combining multi-objective optimization requirements to guide the solution toward the Pareto front. Furthermore, a dual-policy collaborative network architecture is designed. Specifically, the operation policy network is responsible for process sequencing, while the AGV policy network is responsible for AGV transportation task assignment. The sequential execution of the two networks constitutes a complete scheduling action, realizing collaborative optimization of processing and transportation. The two networks share advantage estimation, employ policy feedback mechanisms, and perform joint state updates to enable dynamic mutual adaptation and improve both algorithm robustness and convergence, thus enhancing inter-policy collaboration and adaptability. The effectiveness and superiority of the proposed algorithm are verified by experimental analysis and comparison on extended test cases.
keywords: job-shop scheduling speed-adjustable and rechargeable AGV green scheduling proximal policy optimization (PPO) algorithm disjunctive graph
文章编号: 中图分类号: 文献标志码:
基金项目:
引用文本:
张毅,王玉芳,葛师语,郑明波,张艺.绿色作业车间与变速充电AGV协同调度研究.计算机系统应用,2026,35(8):270-283
ZHANG Yi,WANG Yu-Fang,GE Shi-Yu,ZHENG Ming-Bo,ZHANG Yi.Research on Collaborative Scheduling of Green Job-shop and Speed-adjustable and Rechargeable AGV.COMPUTER SYSTEMS APPLICATIONS,2026,35(8):270-283
张毅,王玉芳,葛师语,郑明波,张艺.绿色作业车间与变速充电AGV协同调度研究.计算机系统应用,2026,35(8):270-283
ZHANG Yi,WANG Yu-Fang,GE Shi-Yu,ZHENG Ming-Bo,ZHANG Yi.Research on Collaborative Scheduling of Green Job-shop and Speed-adjustable and Rechargeable AGV.COMPUTER SYSTEMS APPLICATIONS,2026,35(8):270-283

