GET: 用于三维人体姿态估计的空间融合与自适应帧图增强Transformer
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

陕西省自然科学基金 (2025JC-YBMS-791)


GET: Spatial Fusion and Adaptive Frame Graph-enhanced Transformer for 3D Human Pose Estimation
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    现有的 Transformer 模型在建模序列依赖方面表现优异, 但在处理具有非欧几里得结构的人体骨骼关键点数据时, 却缺乏对图结构的有效建模能力. 相比之下, 图卷积网络(graph convolutional network, GCN)更擅长捕捉关键点之间的拓扑关系与局部结构信息. 基于此, 本文提出了一种Transformer架构——GET. 该方法将GCN与Transformer相融合, 构建了一个兼顾图结构建模与全局上下文建模的统一框架, 从而增强了对人体骨骼表征的学习能力, 并有效克服了现有基于Transformer的三维人体姿态估计方法的缺陷, 即其自注意力机制中QKV向量的推导均基于简单线性映射. 进一步地, GET在保持Transformer灵活性的同时, 引入图卷积来建模人体骨骼的结构先验知识, 使模型能够同时捕捉局部关节关系和全局时空依赖. 大量实验结果表明, 本文提出的GET在Human3.6M和MPI-INF-3DHP两个数据集上均取得了当前最优性能, 充分验证了其在非欧几里得结构数据建模中的有效性与通用性.

    Abstract:

    Existing Transformer models excel at modeling sequential dependencies, whereas they are less effective at capturing graph structures when dealing with non-Euclidean human skeletal keypoint data. By contrast, graph convolutional networks (GCNs) are well-suited for capturing topological relationships and local structural information among keypoints. Based on this, this study proposes a novel Transformer architecture——GET. The proposed method integrates GCN and Transformer into a unified framework that simultaneously models graph structures and global context, thus enhancing the representation learning of human skeletons and overcoming the limitation of existing Transformer-based 3D human pose estimation methods, where the Q, K, and V vectors in the self-attention mechanism are derived only through simple linear mappings. Furthermore, GET preserves the flexibility of Transformers while incorporating GCNs to encode structural priors of the human skeleton, allowing the model to jointly capture local joint relationships and global spatiotemporal dependencies. Extensive experimental results demonstrate that the proposed GET achieves state-of-the-art performance on the Human3.6M and MPI-INF-3DHP datasets, validating its effectiveness and generalization capability for modeling non-Euclidean structured data.

    参考文献
    相似文献
    引证文献
引用本文

王毓龙,王燕妮. GET: 用于三维人体姿态估计的空间融合与自适应帧图增强Transformer.计算机系统应用,2026,35(6):71-83

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-10-12
  • 最后修改日期:2025-11-03
  • 录用日期:
  • 在线发布日期: 2026-04-22
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号