本文已被:浏览 88次 下载 88次
Received:October 12, 2025 Revised:November 03, 2025
Received:October 12, 2025 Revised:November 03, 2025
中文摘要: 现有的 Transformer 模型在建模序列依赖方面表现优异, 但在处理具有非欧几里得结构的人体骨骼关键点数据时, 却缺乏对图结构的有效建模能力. 相比之下, 图卷积网络(graph convolutional network, GCN)更擅长捕捉关键点之间的拓扑关系与局部结构信息. 基于此, 本文提出了一种Transformer架构——GET. 该方法将GCN与Transformer相融合, 构建了一个兼顾图结构建模与全局上下文建模的统一框架, 从而增强了对人体骨骼表征的学习能力, 并有效克服了现有基于Transformer的三维人体姿态估计方法的缺陷, 即其自注意力机制中Q、K、V向量的推导均基于简单线性映射. 进一步地, GET在保持Transformer灵活性的同时, 引入图卷积来建模人体骨骼的结构先验知识, 使模型能够同时捕捉局部关节关系和全局时空依赖. 大量实验结果表明, 本文提出的GET在Human3.6M和MPI-INF-3DHP两个数据集上均取得了当前最优性能, 充分验证了其在非欧几里得结构数据建模中的有效性与通用性.
中文关键词: 三维人体姿态估计 图增强Transformer 空间融合 时域自适应建模
Abstract:Existing Transformer models excel at modeling sequential dependencies, whereas they are less effective at capturing graph structures when dealing with non-Euclidean human skeletal keypoint data. By contrast, graph convolutional networks (GCNs) are well-suited for capturing topological relationships and local structural information among keypoints. Based on this, this study proposes a novel Transformer architecture——GET. The proposed method integrates GCN and Transformer into a unified framework that simultaneously models graph structures and global context, thus enhancing the representation learning of human skeletons and overcoming the limitation of existing Transformer-based 3D human pose estimation methods, where the Q, K, and V vectors in the self-attention mechanism are derived only through simple linear mappings. Furthermore, GET preserves the flexibility of Transformers while incorporating GCNs to encode structural priors of the human skeleton, allowing the model to jointly capture local joint relationships and global spatiotemporal dependencies. Extensive experimental results demonstrate that the proposed GET achieves state-of-the-art performance on the Human3.6M and MPI-INF-3DHP datasets, validating its effectiveness and generalization capability for modeling non-Euclidean structured data.
keywords: 3D human pose estimation graph-enhanced Transformer spatial fusion adaptive temporal modeling
文章编号: 中图分类号: 文献标志码:
基金项目:陕西省自然科学基金 (2025JC-YBMS-791)
引用文本:
王毓龙,王燕妮.GET: 用于三维人体姿态估计的空间融合与自适应帧图增强Transformer.计算机系统应用,2026,35(6):71-83
WANG Yu-Long,WANG Yan-Ni.GET: Spatial Fusion and Adaptive Frame Graph-enhanced Transformer for 3D Human Pose Estimation.COMPUTER SYSTEMS APPLICATIONS,2026,35(6):71-83
王毓龙,王燕妮.GET: 用于三维人体姿态估计的空间融合与自适应帧图增强Transformer.计算机系统应用,2026,35(6):71-83
WANG Yu-Long,WANG Yan-Ni.GET: Spatial Fusion and Adaptive Frame Graph-enhanced Transformer for 3D Human Pose Estimation.COMPUTER SYSTEMS APPLICATIONS,2026,35(6):71-83

