本文已被:浏览 346次 下载 1118次
Received:February 27, 2025 Revised:April 08, 2025
Received:February 27, 2025 Revised:April 08, 2025
中文摘要: 针对现有的多光谱行人检测算法存在多模态相互作用不足和融合方法缺乏远程依赖性, 导致在低光照背景下小尺度行人检测性能不足的问题, 提出了一种融合自适应双重注意力和轴向注意力Transformer的多光谱小尺度行人检测算法(adaptive dual attention and axial attention Transformer network, ADATNet). 采用双分支CSPDarknet53网络分别提取可见光和红外图像中的深度特征, 充分保留两种模态的特有信息. 设计两个特征交叉融合模块: 自适应双重注意力模块(adaptive dual attention module, ADAM)和轴向注意力Transformer特征增强(axial attention Transformer feature enhancement, ATFE)模块, 其中ADAM旨在强化模型对关键特征的关注, 同时抑制不相关或冗余的信息; ATFE关联多模态特征的位置编码来融合增强的特征, 在确保计算效率的同时捕获长距离依赖关系. 将融合后的特征输入至检测头以输出最终检测结果. 实验结果表明, ADATNet在KAIST数据集上的MR–2降低至7.08%, 同时在FLIR和LLVIP数据集上的mAP50分别达到82.8%和97.6%, 较基线方法提升4.7%和1.9%, 具有良好的检测性能.
Abstract:Existing multi-spectral pedestrian detection algorithms suffer from insufficient multimodal interaction and lack of long-distance dependence of fusion algorithms, thus resulting in poor performance in small-scale pedestrian detection in low-light conditions. To this end, this study proposes an fusion adaptive dual attention and axial attention Transformer network (ADATNet) for multi-spectral small-scale pedestrian detection. The method adopts the dual-branch CSPDarknet53 network to separately extract deep features from visible and infrared images, preserving the unique information from each modality. Meanwhile, two feature fusion modules are designed, including the adaptive dual attention module (ADAM) and axial attention Transformer feature enhancement module (ATFE). ADAM aims to enhance the model’s focus on critical features and suppress irrelevant or redundant information, while ATFE correlates the positional encoding of multimodal features to fuse the enhanced features, thereby both ensuring computational efficiency and capturing long-distance dependence. The fused features are then fed into the detection head to output the final detection results. Experimental results indicate that MR–2 of ADATNet decreases to 7.08% on the KAIST dataset, while the mAP50 for the FLIR and LLVIP datasets reaches 82.8% and 97.6% respectively. They have an improvement of 4.7% and 1.9% over the baseline methods, demonstrating excellent detection performance.
keywords: multi-spectral pedestrian detection visible light and infrared adaptive dual attention axial attention
文章编号: 中图分类号: 文献标志码:
基金项目:陕西省自然科学基础研究项目(2020JM499, 2020JQ684)
引用文本:
罗建国,王燕妮,韩世鹏,吕昊,张耀荣,吴雪.融合自适应双重注意力和轴向注意力Transformer的多光谱行人检测.计算机系统应用,2025,34(10):86-100
LUO Jian-Guo,WANG Yan-Ni,HAN Shi-Peng,LYU Hao,ZHANG Yao-Rong,WU Xue.Fusion of Adaptive Dual Attention and Axial Attention Transformer for Multi-spectral Pedestrian Detection.COMPUTER SYSTEMS APPLICATIONS,2025,34(10):86-100
罗建国,王燕妮,韩世鹏,吕昊,张耀荣,吴雪.融合自适应双重注意力和轴向注意力Transformer的多光谱行人检测.计算机系统应用,2025,34(10):86-100
LUO Jian-Guo,WANG Yan-Ni,HAN Shi-Peng,LYU Hao,ZHANG Yao-Rong,WU Xue.Fusion of Adaptive Dual Attention and Axial Attention Transformer for Multi-spectral Pedestrian Detection.COMPUTER SYSTEMS APPLICATIONS,2025,34(10):86-100

