非完美博弈问题的大语言模型双模块协同微调
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:


Dual-module Collaborative Fine-tuning of LLMs for Imperfect Information Game Problems
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    近年来, 大语言模型(large language model, LLM)已广泛应用于自然语言处理、计算机视觉等领域, 但非完美信息博弈的相关研究(如掼蛋、立直麻将等)仍以强化学习为主导. 此类博弈问题存在隐藏信息推理难、数值-文本特征融合低效、多阶段策略适配性差的核心痛点. 博弈树搜索、反事实遗憾最小化等传统算法存在计算量大、收敛慢、鲁棒性不足等局限, 现有大语言模型则存在推理延迟高、动态信息处理能力欠缺等问题. 针对上述问题, 本文提出分阶段自适应LoRA/QLoRA与博弈增强型Adapter双模块微调框架, 以Qwen3系列模型为基准, 基于掼蛋和立直麻将的高质量数据集开展微调训练, 并与主流强化学习AI: DanZero+、Tenhou AI进行对抗实验. 实验结果表明, 经微调的大语言模型, 博弈胜率虽距离顶尖强化学习仍有一定差距, 但在隐藏信息推理、跨模态特征融合等核心能力维度已具备竞争力. 本文为非完美博弈的大语言模型解决方案提供了新路径, 但暴露的记忆缺失、决策片面等问题仍需后续优化.

    Abstract:

    In recent years, large language models (LLMs) have been widely applied in fields such as natural language processing and computer vision. However, research on imperfect-information games such as Guandan and Riichi Mahjong is still largely dominated by reinforcement learning approaches. These games present several core challenges, including difficulties in inferring hidden information, inefficient fusion of numerical and textual features, and limited adaptability in multi-stage strategies. Traditional algorithms such as game tree search and counterfactual regret minimization suffer from limitations including high computational complexity, slow convergence, and insufficient robustness, while existing LLMs suffer from problems such as high inference latency and inadequate dynamic information processing capability. To address these issues, this study proposes a dual-module fine-tuning framework combining stage-wise adaptive LoRA/QLoRA and a game-enhanced Adapter, with the Qwen3 series models as the baseline. Fine-tuning is conducted using high-quality datasets for Guandan and Riichi Mahjong, followed by adversarial experiments against mainstream reinforcement learning AIs (DanZero+, Tenhou AI). The experimental results show that although the fine-tuned large model still lags behind top-level reinforcement learning methods in terms of game win rate, it has become competitive in core capabilities such as hidden information inference and cross-modal feature fusion. This study provides a new approach for applying large models to imperfect-information games. However, the identified problems, including memory deficiencies and one-sided decision-making, still require further optimization.

    参考文献
    相似文献
    引证文献
引用本文

段永熙,王雷.非完美博弈问题的大语言模型双模块协同微调.计算机系统应用,,():1-11

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-01-26
  • 最后修改日期:2026-02-27
  • 录用日期:
  • 在线发布日期: 2026-06-15
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号