###
计算机系统应用英文版:2026,35(6):156-168
本文二维码信息
码上扫一扫!
ST-HGT: 基于时空超图Transformer的骨骼动作识别
(1.南京信息工程大学 软件学院, 南京 210044;2.南京工业大学 计算机与信息工程学院, 南京 211816;3.南京信息工程大学 自动化学院, 南京 210044)
ST-HGT: Skeleton Action Recognition Based on Spatio-temporal Hypergraph Transformer
(1.School of Software, Nanjing University of Information Science & Technology, Nanjing 210044, China;2.College of Computer and Information Engineering, Nanjing Tech University, Nanjing 211816, China;3.School of Automation, Nanjing University of Information Science & Technology, Nanjing 210044, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 48次   下载 70
Received:September 16, 2025    Revised:October 28, 2025
中文摘要: 图卷积网络因其对非欧几里得数据结构的优良建模而被广泛应用于基于骨骼的动作识别. 由于图卷积具有局部聚合的特性, 它的关注点聚焦在邻域关节和短时窗口. 然而, 多数人体动作更依赖于关节全局协同作用. 目前, 基于图卷积的方法通常忽略了空间连接拓扑和高阶群体关系. 为减轻这种限制, 提出了一种集成高阶空间建模与多尺度时序建模的新型框架, 称为时空超图Transformer网络. 本文提出通过骨骼超图构建高阶空间结构, 并结合骨骼距离偏差来动态建模骨骼关节整体的相对位移和远距离骨骼特征. 同时, 为了增强骨骼边和关节之间的互补性, 利用图卷积和动态自注意机制对多帧骨骼序列中的边特征和关节特征进行联合建模. 此外, 引入多尺度时间卷积模块, 提取不同时间窗口下的骨骼时序特征. 本文在NTU-RGB+D、NTU-RGB+D 120和Northwestern-UCLA这3个具有挑战性的基准数据集上进行了广泛实验, 实验结果表明, 本文提出的ST-HGT方法在3个基准上都达到或超过了最先进方法的精度.
Abstract:Graph convolutional network (GCN) have been extensively applied to skeleton-based action recognition due to their superior modeling of non-Euclidean data structures. Since graph convolution has local aggregation properties, it focuses on neighborhood joints and short-time windows. However, most human actions depend more on global joint synergy. Existing methods based on graph convolution generally ignore the spatial connection topology and higher-order group relations. To this end, this study proposes a novel framework that integrates higher-order spatial modeling with multi-scale temporal modeling, which is called the spatio-temporal hypergraph Transformer (ST-HGT) network. A skeletal hypergraph is adopted to construct high-order spatial structures, and skeletal distance deviation is combined to dynamically model the overall relative displacement of skeleton joints and distant skeletal features. Meanwhile, to enhance the complementarity between skeleton edges and joints, the study utilizes graph convolution and the dynamic self-attention mechanism to jointly model edge and joint features in multi-frame sequences. Additionally, a multi-scale temporal convolution module is introduced to extract the skeletal temporal features under different time windows. The proposed ST-HGT has been extensively experimented on the three challenging benchmark datasets of NTU-RGB+D, NTU-RGB+D 120, and Northwestern-UCLA. The experiments show that the proposed method reaches or exceeds the precision of state-of-the-art methods on all three benchmark datasets.
文章编号:     中图分类号:    文献标志码:
基金项目:国家自然科学基金 (62272236, 62376128); 江苏省自然科学基金 (BK20201136, BK20191401)
引用文本:
张小瑞,周石煜,孙伟.ST-HGT: 基于时空超图Transformer的骨骼动作识别.计算机系统应用,2026,35(6):156-168
ZHANG Xiao-Rui,ZHOU Shi-Yu,SUN Wei.ST-HGT: Skeleton Action Recognition Based on Spatio-temporal Hypergraph Transformer.COMPUTER SYSTEMS APPLICATIONS,2026,35(6):156-168