时频增强扩散模型的伪装目标检测
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:


Time-frequency Enhanced Diffusion Model for Camouflaged Object Detection
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    伪装目标检测(camouflaged object detection, COD)在计算机视觉领域占据重要地位, 其核心在于精准识别与背景在纹理、颜色等方面高度融合的目标, 这一任务因其广泛的应用前景而备受瞩目. 近年来, 扩散模型在图像生成和去噪方面表现突出, 但其在伪装目标检测任务中的应用尚面临诸多挑战. 现有基于扩散模型的COD方法主要依赖静态空间域条件特征指导去噪过程, 这种方法使得条件特征在全局特征建模能力上存在明显不足, 且无法有效感知扩散时间步的变化. 为了应对这些挑战, 本文提出了一种名为时频增强扩散模型的伪装目标检测(time-frequency enhanced diffusion model for camouflaged object detection, TFD-COD)的方法, 该方法在骨干网络中引入时间嵌入机制, 这种方式可以使得条件特征的提取过程能够敏锐感知扩散时间步的细微变化; 还设计了频域-空间双分支注意力(frequency-spatial dual branch attention, FSBA)模块, 频域分支通过傅里叶变换及自注意力机制, 精准捕捉全局结构信息, 而空间域分支则利用多尺度卷积和自注意力, 建模局部上下文. 两个分支相互协作, 实现了全局与局部特征的协同增强. 此外, TFD-COD还引入了金字塔滤波融合(pyramid filtering fusion, PFF)策略, 利用浅层特征指导深层特征的上采样, 从而实现多级特征的深度融合. 最终, 这些经过精心设计的条件特征被送入扩散模型, 引导其高效去噪. 实验结果显示, TFD-COD在伪装目标检测任务中的表现显著优于现有方法.

    Abstract:

    Camouflaged object detection (COD) plays a crucial role in computer vision, and its core is to accurately identify objects that highly resemble their background in texture and color. This task has caught extensive attention due to its broad application prospects. In recent years, diffusion models have achieved outstanding performance in image generation and denoising, but their application in COD still faces many challenges. Existing COD methods based on diffusion models mainly rely on static spatial-domain conditional features to guide the denoising process. However, this kind of method results in obvious deficiencies in the global feature modeling capacity of conditional features and fails to effectively perceive variations across diffusion time-steps. To this end, this study proposes a brand-new method called the time-frequency enhanced diffusion model for COD (TFD-COD). Specifically, the time embedding mechanism is introduced into the backbone network, enabling the sensitive perception of subtle variations of diffusion time-steps during the conditional feature extraction process. Additionally, a frequency-spatial dual branch attention (FSBA) module is designed, where the frequency branch accurately captures global structural information via Fourier transform and self-attention, while the spatial branch models local context via multi-scale convolutions and self-attention. The two branches collaborate for synergistic enhancement of global and local features. Furthermore, TFD-COD introduces a pyramid filtering fusion (PFF) strategy, which employs shallow features to guide the upsampling of deep features to achieve the deep fusion of multi-level features. Finally, these carefully designed conditional features are fed into the diffusion model to guide efficient denoising. Experimental results show that TFD-COD outperforms existing methods significantly in COD tasks.

    参考文献
    相似文献
    引证文献
引用本文

宋佳欣,张宇,王忆茹.时频增强扩散模型的伪装目标检测.计算机系统应用,2026,35(7):222-232

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-12-01
  • 最后修改日期:2025-12-24
  • 录用日期:
  • 在线发布日期: 2026-06-03
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号