本文已被:浏览 1021次 下载 1205次
Received:April 14, 2025 Revised:May 07, 2025
Received:April 14, 2025 Revised:May 07, 2025
中文摘要: 为了实现复杂场景下更精准的人体关键点定位, 本文提出一种基于多级选通的人体姿态估计网络模型MGNet. 首先, 该模型以HRNet为骨干网络并移除了HRNet中的第4阶段, 通过平行连接实现多阶段分辨率的连接, 达到平衡模型复杂度和检测能力的目的; 其次, 构建了一种新的动态上采样方法以替换双线性插值方式, 能够充分利用特征图中丰富的语义信息, 并且提高网络的特征融合能力; 然后, 设计了一种混合高效通道注意力机制, 增强模型对全局和局部通道信息, 同时提出了一种轻量化的多级选通CA注意力模块(MCAM), 该模块被用来替代HRNet中的残差块, 在确保模型性能的同时, 有效地减少了模型的复杂度, 从而实现了轻量化的目标; 最后, 设计了一种深度感知特征增强模块, 使最终输出特征在包含局部深度信息的同时还包含全局信息. 实验结果表明, MGNet在COCO2017验证集上取得了72.1%的准确率, 与骨干网络HRNet相比, 参数量减少了70%, 精度仅损失1.3%. 同样在MPII数据集上, 在减少大量计算量的同时, 取得了89.4%的平均精度. 综合COCO2017和MPII数据集上的表现, MGNet不仅在轻量化模型中实现了最优精度, 而且与大模型相比具有相当或更佳的性能.
Abstract:To achieve more accurate positioning of human key points, this study proposes a multi-stage selective pass-based human pose estimation network model MGNet. Firstly, by taking HRNet as the backbone network, this model removes the fourth stage in HRNet, and realizes the connection of multi-stage resolution via parallel connection, aiming to achieve the balance between complexity and detection ability of the model. Secondly, a new dynamic upsampling method is developed to replace the bilinear interpolation method, which can make full use of the rich semantic information in the feature map and improve the feature fusion ability of the network. Then, a hybrid efficient channel attention mechanism is designed to enhance the global and local channel information of the model, and a lightweight multi-stage selective pass-based CA attention module (MCAM) is proposed. This module is employed to replace the residual block of HRNet, which ensures the model performance, reduces the model complexity, and ultimately achieves light weight. Finally, a depth-aware feature enhancement module is designed to make the final output feature contain both local depth information and global information. Experimental results show that MGNet yields 72.1% accuracy on the COCO2017 verification set, and compared with the backbone network HRNet, the parameter amount is reduced by 70%, with the accuracy loss of 1.3%. Meanwhile, on the MPII dataset, average accuracy of 89.4% is achieved, with a large amount of computation reduced. By comprehensively considering the performance on the COCO2017 and MPII datasets, MGNet not only yields optimal accuracy in the lightweight model, but also has comparable or better performance compared to the large model.
文章编号: 中图分类号: 文献标志码:
基金项目:国家自然科学基金(62204172)
引用文本:
李小雨,刘罡,吴烨,郑泽林.基于多级选通的人体姿态估计网络.计算机系统应用,2025,34(11):172-183
LI Xiao-Yu,LIU Gang,WU Ye,ZHENG Ze-Lin.Human Posture Estimation Network Based on Multi-stage Selective Pass.COMPUTER SYSTEMS APPLICATIONS,2025,34(11):172-183
李小雨,刘罡,吴烨,郑泽林.基于多级选通的人体姿态估计网络.计算机系统应用,2025,34(11):172-183
LI Xiao-Yu,LIU Gang,WU Ye,ZHENG Ze-Lin.Human Posture Estimation Network Based on Multi-stage Selective Pass.COMPUTER SYSTEMS APPLICATIONS,2025,34(11):172-183

