基于多尺度纹理感知与类激活分块注意力的图像分类网络
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:


Image Classification Network Based on Multi-scale Texture Perception and Class-activated Block Attention
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对图像分类领域中存在的复杂纹理特征提取能力薄弱、局部信息关联性挖掘不足及聚焦关键特征欠缺等问题, 提出基于多尺度纹理感知与类激活分块注意力的图像分类网络MTCANet. MTCANet以ResNet-34残差网络为基础. 首先, 设计多尺度纹理感知卷积单元MTPC (multi-scale texture perception convolution), 基于参数化Gabor函数对卷积核进行初始化, 并结合3种尺度深度分离卷积, 实现通道独立处理与感受野多样化, 有效提升图像多尺度纹理特征的表达能力. 其次, 提出了分块类激活特征提取模块BCAF (block-wise class-activated feature extraction), 以类激活机制为基础, 并基于分块操作计算块级与块内像素级置信度, 构建双层加权机制, 强化显著区域特征并抑制背景干扰, 从而生成具有更强判别力的局部特征表示. 最后, 设计类激活分块特征交互注意力模块CAPI (class-activated block feature interaction attention), 建模分块语义交互增强局部细节一致性, 聚合局部特征生成全局语义, 并利用注意力机制实现局部-全局特征的深度融合, 有效提升模型在复杂背景中的目标感知与判别能力. 实验结果表明, MTCANet在Imagenette、Imagewoof和Flowers-102数据集上的分类准确率分别达到89.99%、80.12%和91.69%, 较现有先进方法平均提升2.37%、1.90%和1.11%, 有效提升了图像分类性能.

    Abstract:

    To address the issues in the field of image classification, including limited capability in extracting complex texture features, insufficient mining of local information correlations, and insufficient focus on key features, this study proposes an image classification network based on multi-scale texture perception and class-activated block attention, termed as MTCANet. MTCANet is built on the ResNet-34 residual network. First, a multi-scale texture perception convolution (MTPC) unit is designed, where convolution kernels are initialized using parameterized Gabor functions and combined with three-scale depthwise separable convolutions to enable channel-independent processing and diversified receptive fields, thus enhancing the representation of multi-scale texture features. Secondly, a block-wise class-activated feature extraction (BCAF) module is proposed. Based on the class activation mechanism, block-level and intra-block pixel-level confidence scores are computed via blocking operations, forming a dual-layer weighting mechanism that enhances salient region features while suppressing background interference, thus generating more discriminative local feature representations. Finally, a class-activated block feature interaction attention (CAPI) module is designed to model block-wise semantic interactions to enhance local detail consistency, aggregate local features into global semantics, and employ an attention mechanism for deep fusion of local and global features, thus improving target perception and discrimination in complex backgrounds. Experimental results demonstrate that MTCANet achieves classification accuracies of 89.99%, 80.12%, and 91.69% on the Imagenette, Imagewoof, and Flowers-102 datasets, respectively, with an average improvement of 2.37%, 1.90%, and 1.11% compared to existing state-of-the-art methods, confirming significant enhancement in image classification performance.

    参考文献
    相似文献
    引证文献
引用本文

王永贵,谷佳鸿.基于多尺度纹理感知与类激活分块注意力的图像分类网络.计算机系统应用,2026,35(5):128-142

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-10-11
  • 最后修改日期:2025-11-03
  • 录用日期:
  • 在线发布日期: 2026-04-02
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号