###
计算机系统应用英文版:2026,35(7):304-315
本文二维码信息
码上扫一扫!
内外嵌入增强的序贯评价驱动Shapley值Q学习
(1.大连海事大学 人工智能学院, 大连 116026;2.中国科学院 沈阳自动化研究所, 沈阳 110016)
Sequential-evaluation-driven Shapley Value Q-learning with Internal and External Embedding Enhancement
(1.College of Artificial Intelligence, Dalian Maritime University, Dalian 116026, China;2.Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 59次   下载 81
Received:November 19, 2025    Revised:December 12, 2025
中文摘要: 当前多智能体合作强化学习(multi-agent cooperative reinforcement learning, MACRL)算法具有广泛的应用场景, 但其在信息利用方面仍存在显著缺陷, 主要体现在对冗余信息的筛选不足、多维度信息间协同机制的缺失以及对其他智能体行为影响的忽视. 针对上述问题, 本文提出一种内外嵌入增强的序贯评价驱动Shapley值Q学习模型. 受社会学习理论中交互决定论的观点, 即个人因素、外部环境和行为本身这3方面共同决定人类行为的启发, 该模型分为内嵌入模块、外嵌入模块和序贯评价驱动Shapley值Q学习模块, 分别用于学习个体因素表征、外部环境表征和评估动作价值. 具体来说, 首先通过外嵌入模块自适应捕捉关键外部信息; 其次利用内嵌入模块优化局部观测与历史经验学习, 并借助融合模块捕捉内外信息的协同效应; 最后针对信度分配问题提出序贯评价驱动Shapley值Q学习方法, 通过边际效益分解理论优化边际效益估计, 将Shapley值计算问题转化为序列建模问题, 量化外部智能体的决策对当前智能体决策的贡献. 实验结果表明, 本文提出的算法在星际争霸II的8个复杂环境中取得了0.85的平均胜率, 达到了SOTA (state-of-the-art)性能.
Abstract:Multi-agent cooperative reinforcement learning (MACRL) algorithms are widely applied in various scenarios; however, significant deficiencies in information utilization remain, including insufficient filtering of redundant information, the absence of coordination mechanisms among multi-dimensional information, and inadequate consideration of the influence of other agent’ behaviors. To address these issues, this study proposes a sequential-evaluation-driven Shapley value Q-learning model with internal and external embedding enhancement. Inspired by reciprocal determinism in social learning theory, which states that individual factors, the external environment, and behaviors jointly determine outcomes, the proposed model consists of an internal embedding module, an external embedding module, and a sequential-evaluation-driven Shapley value Q-learning module, which are responsible for learning individual factor representations, external environment representations, and action value evaluation, respectively. Specifically, key external information is adaptively captured through the external embedding module, while local observations and historical experience learnings are optimized by the internal embedding module. A fusion module is further employed to model the synergistic effects between internal and external information. To address the credit assignment problem, a sequential Shapley value Q-learning method is proposed. Marginal contribution estimation is optimized based on marginal benefit decomposition theory, and the Shapley value computation is reformulated as a sequence modeling problem to quantify the contribution of other agents’ decisions to the current agent’s action selection. Experimental results show that the proposed algorithm achieves an average win rate of 0.85 across eight complex StarCraft II environments, reaching state-of-the-art (SOTA) performance.
文章编号:     中图分类号:    文献标志码:
基金项目:
引用文本:
胡若凡,齐中朋,马裕博.内外嵌入增强的序贯评价驱动Shapley值Q学习.计算机系统应用,2026,35(7):304-315
HU Ruo-Fan,QI Zhong-Peng,MA Yu-Bo.Sequential-evaluation-driven Shapley Value Q-learning with Internal and External Embedding Enhancement.COMPUTER SYSTEMS APPLICATIONS,2026,35(7):304-315