Pamamba: 基于混合注意力的复杂场景目标检测模块
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金(62166035); 中央高校基本科研业务费专项资金(31920240091, 31920250006); 甘肃省自然科学基金(25JRRA991)


Pamamba: Hybri-attention-based Module for Object Detection in Complex Scenes
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对唐卡等复杂艺术图像中目标方向性纹理显著、尺度差异大、细粒度特征相似等检测难题, 本文提出一种全局-方向混合注意力模块Pamamba. 该模块基于线性复杂度的Mamba架构, 由状态空间增强模块(enhanced state space module, ESSM)和通道注意力前馈网络(channel attention feedforward network, CAFFN)构成, 形成空间-通道协同增强机制. ESSM采用双路径并行结构: 全局路径通过深度特征提取组件(deep feature extraction module, DFEM)捕获全局上下文语义; 方向增强路径通过方向信息增强组件(directional information enhancement module, DIEM)与对角扫描机制E-SS2D, 弥补了传统Mamba在方向感知上的局限. CAFFN引入门控机制与空洞卷积, 动态增强通道间特征交互与多尺度表达. Pamamba为即插即用模块, 可嵌入YOLOv8/v10/11等主流检测器. 在COCO基准和自建唐卡数据集上的实验结果表明, 该方法能显著提升检测性能, 在唐卡数据集上将mAP@0.5最高提升至66.9%, mAP@[0.5:0.95]最高达60.6%, 尤其改善了小目标召回率, 为复杂场景下的鲁棒检测提供了有效工具.

    Abstract:

    To address the challenges of object detection in complex artistic images such as Thangka, including significant directional textures, large scale variations, and similar fine-grained features, this study proposes a global-directional hybrid attention module named Pamamba. Built upon the linear-complexity Mamba architecture, the module consists of the enhanced state space module (ESSM) and channel attention feedforward network (CAFFN), forming a spatial-channel collaborative enhancement mechanism. ESSM adopts a dual-path architecture: the global path employs the deep feature extraction module (DFEM) to capture global contextual semantics, while the directional enhancement path integrates the directional information enhancement module (DIEM) with a diagonal scanning mechanism (E-SS2D) to compensate for the directional perception limitations of conventional Mamba. CAFFN introduces a gating mechanism and dilated convolutions to dynamically enhance channel-wise feature interaction and multi-scale representation. Pamamba is a plug-and-play module that can be embedded into mainstream detectors such as YOLOv8/v10/11. Experimental results on the COCO benchmark and a self-built Thangka dataset demonstrate that the proposed method significantly improves detection performance and achieves mAP@0.5 of up to 66.9% and mAP@[0.5:0.95] of up to 60.6% on the Thangka dataset, especially improving small-object recall. It provides an effective tool for robust object detection in complex scenarios.

    参考文献
    相似文献
    引证文献
引用本文

郭晓然,周祥雨,彭昊,陈智源,王铁君,龙智磊. Pamamba: 基于混合注意力的复杂场景目标检测模块.计算机系统应用,,():1-11

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-02-11
  • 最后修改日期:2026-03-05
  • 录用日期:
  • 在线发布日期: 2026-08-21
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号