ST-HGT: 基于时空超图Transformer的骨骼动作识别
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金 (62272236, 62376128); 江苏省自然科学基金 (BK20201136, BK20191401)


ST-HGT: Skeleton Action Recognition Based on Spatio-temporal Hypergraph Transformer
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    图卷积网络因其对非欧几里得数据结构的优良建模而被广泛应用于基于骨骼的动作识别. 由于图卷积具有局部聚合的特性, 它的关注点聚焦在邻域关节和短时窗口. 然而, 多数人体动作更依赖于关节全局协同作用. 目前, 基于图卷积的方法通常忽略了空间连接拓扑和高阶群体关系. 为减轻这种限制, 提出了一种集成高阶空间建模与多尺度时序建模的新型框架, 称为时空超图Transformer网络. 本文提出通过骨骼超图构建高阶空间结构, 并结合骨骼距离偏差来动态建模骨骼关节整体的相对位移和远距离骨骼特征. 同时, 为了增强骨骼边和关节之间的互补性, 利用图卷积和动态自注意机制对多帧骨骼序列中的边特征和关节特征进行联合建模. 此外, 引入多尺度时间卷积模块, 提取不同时间窗口下的骨骼时序特征. 本文在NTU-RGB+D、NTU-RGB+D 120和Northwestern-UCLA这3个具有挑战性的基准数据集上进行了广泛实验, 实验结果表明, 本文提出的ST-HGT方法在3个基准上都达到或超过了最先进方法的精度.

    Abstract:

    Graph convolutional network (GCN) have been extensively applied to skeleton-based action recognition due to their superior modeling of non-Euclidean data structures. Since graph convolution has local aggregation properties, it focuses on neighborhood joints and short-time windows. However, most human actions depend more on global joint synergy. Existing methods based on graph convolution generally ignore the spatial connection topology and higher-order group relations. To this end, this study proposes a novel framework that integrates higher-order spatial modeling with multi-scale temporal modeling, which is called the spatio-temporal hypergraph Transformer (ST-HGT) network. A skeletal hypergraph is adopted to construct high-order spatial structures, and skeletal distance deviation is combined to dynamically model the overall relative displacement of skeleton joints and distant skeletal features. Meanwhile, to enhance the complementarity between skeleton edges and joints, the study utilizes graph convolution and the dynamic self-attention mechanism to jointly model edge and joint features in multi-frame sequences. Additionally, a multi-scale temporal convolution module is introduced to extract the skeletal temporal features under different time windows. The proposed ST-HGT has been extensively experimented on the three challenging benchmark datasets of NTU-RGB+D, NTU-RGB+D 120, and Northwestern-UCLA. The experiments show that the proposed method reaches or exceeds the precision of state-of-the-art methods on all three benchmark datasets.

    参考文献
    相似文献
    引证文献
引用本文

张小瑞,周石煜,孙伟. ST-HGT: 基于时空超图Transformer的骨骼动作识别.计算机系统应用,2026,35(6):156-168

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-09-16
  • 最后修改日期:2025-10-28
  • 录用日期:
  • 在线发布日期: 2026-04-29
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号