内外嵌入增强的序贯评价驱动Shapley值Q学习
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:


Sequential-evaluation-driven Shapley Value Q-learning with Internal and External Embedding Enhancement
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    当前多智能体合作强化学习(multi-agent cooperative reinforcement learning, MACRL)算法具有广泛的应用场景, 但其在信息利用方面仍存在显著缺陷, 主要体现在对冗余信息的筛选不足、多维度信息间协同机制的缺失以及对其他智能体行为影响的忽视. 针对上述问题, 本文提出一种内外嵌入增强的序贯评价驱动Shapley值Q学习模型. 受社会学习理论中交互决定论的观点, 即个人因素、外部环境和行为本身这3方面共同决定人类行为的启发, 该模型分为内嵌入模块、外嵌入模块和序贯评价驱动Shapley值Q学习模块, 分别用于学习个体因素表征、外部环境表征和评估动作价值. 具体来说, 首先通过外嵌入模块自适应捕捉关键外部信息; 其次利用内嵌入模块优化局部观测与历史经验学习, 并借助融合模块捕捉内外信息的协同效应; 最后针对信度分配问题提出序贯评价驱动Shapley值Q学习方法, 通过边际效益分解理论优化边际效益估计, 将Shapley值计算问题转化为序列建模问题, 量化外部智能体的决策对当前智能体决策的贡献. 实验结果表明, 本文提出的算法在星际争霸II的8个复杂环境中取得了0.85的平均胜率, 达到了SOTA (state-of-the-art)性能.

    Abstract:

    Multi-agent cooperative reinforcement learning (MACRL) algorithms are widely applied in various scenarios; however, significant deficiencies in information utilization remain, including insufficient filtering of redundant information, the absence of coordination mechanisms among multi-dimensional information, and inadequate consideration of the influence of other agent’ behaviors. To address these issues, this study proposes a sequential-evaluation-driven Shapley value Q-learning model with internal and external embedding enhancement. Inspired by reciprocal determinism in social learning theory, which states that individual factors, the external environment, and behaviors jointly determine outcomes, the proposed model consists of an internal embedding module, an external embedding module, and a sequential-evaluation-driven Shapley value Q-learning module, which are responsible for learning individual factor representations, external environment representations, and action value evaluation, respectively. Specifically, key external information is adaptively captured through the external embedding module, while local observations and historical experience learnings are optimized by the internal embedding module. A fusion module is further employed to model the synergistic effects between internal and external information. To address the credit assignment problem, a sequential Shapley value Q-learning method is proposed. Marginal contribution estimation is optimized based on marginal benefit decomposition theory, and the Shapley value computation is reformulated as a sequence modeling problem to quantify the contribution of other agents’ decisions to the current agent’s action selection. Experimental results show that the proposed algorithm achieves an average win rate of 0.85 across eight complex StarCraft II environments, reaching state-of-the-art (SOTA) performance.

    参考文献
    相似文献
    引证文献
引用本文

胡若凡,齐中朋,马裕博.内外嵌入增强的序贯评价驱动Shapley值Q学习.计算机系统应用,2026,35(7):304-315

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-11-19
  • 最后修改日期:2025-12-12
  • 录用日期:
  • 在线发布日期: 2026-05-22
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号