本文已被:浏览 65次 下载 101次
Received:October 11, 2025 Revised:November 03, 2025
Received:October 11, 2025 Revised:November 03, 2025
中文摘要: 针对图像分类领域中存在的复杂纹理特征提取能力薄弱、局部信息关联性挖掘不足及聚焦关键特征欠缺等问题, 提出基于多尺度纹理感知与类激活分块注意力的图像分类网络MTCANet. MTCANet以ResNet-34残差网络为基础. 首先, 设计多尺度纹理感知卷积单元MTPC (multi-scale texture perception convolution), 基于参数化Gabor函数对卷积核进行初始化, 并结合3种尺度深度分离卷积, 实现通道独立处理与感受野多样化, 有效提升图像多尺度纹理特征的表达能力. 其次, 提出了分块类激活特征提取模块BCAF (block-wise class-activated feature extraction), 以类激活机制为基础, 并基于分块操作计算块级与块内像素级置信度, 构建双层加权机制, 强化显著区域特征并抑制背景干扰, 从而生成具有更强判别力的局部特征表示. 最后, 设计类激活分块特征交互注意力模块CAPI (class-activated block feature interaction attention), 建模分块语义交互增强局部细节一致性, 聚合局部特征生成全局语义, 并利用注意力机制实现局部-全局特征的深度融合, 有效提升模型在复杂背景中的目标感知与判别能力. 实验结果表明, MTCANet在Imagenette、Imagewoof和Flowers-102数据集上的分类准确率分别达到89.99%、80.12%和91.69%, 较现有先进方法平均提升2.37%、1.90%和1.11%, 有效提升了图像分类性能.
Abstract:To address the issues in the field of image classification, including limited capability in extracting complex texture features, insufficient mining of local information correlations, and insufficient focus on key features, this study proposes an image classification network based on multi-scale texture perception and class-activated block attention, termed as MTCANet. MTCANet is built on the ResNet-34 residual network. First, a multi-scale texture perception convolution (MTPC) unit is designed, where convolution kernels are initialized using parameterized Gabor functions and combined with three-scale depthwise separable convolutions to enable channel-independent processing and diversified receptive fields, thus enhancing the representation of multi-scale texture features. Secondly, a block-wise class-activated feature extraction (BCAF) module is proposed. Based on the class activation mechanism, block-level and intra-block pixel-level confidence scores are computed via blocking operations, forming a dual-layer weighting mechanism that enhances salient region features while suppressing background interference, thus generating more discriminative local feature representations. Finally, a class-activated block feature interaction attention (CAPI) module is designed to model block-wise semantic interactions to enhance local detail consistency, aggregate local features into global semantics, and employ an attention mechanism for deep fusion of local and global features, thus improving target perception and discrimination in complex backgrounds. Experimental results demonstrate that MTCANet achieves classification accuracies of 89.99%, 80.12%, and 91.69% on the Imagenette, Imagewoof, and Flowers-102 datasets, respectively, with an average improvement of 2.37%, 1.90%, and 1.11% compared to existing state-of-the-art methods, confirming significant enhancement in image classification performance.
keywords: image classification multi-scale feature class-activited mapping attention mechanism residual network (ResNet)
文章编号: 中图分类号: 文献标志码:
基金项目:
引用文本:
王永贵,谷佳鸿.基于多尺度纹理感知与类激活分块注意力的图像分类网络.计算机系统应用,2026,35(5):128-142
WANG Yong-Gui,GU Jia-Hong.Image Classification Network Based on Multi-scale Texture Perception and Class-activated Block Attention.COMPUTER SYSTEMS APPLICATIONS,2026,35(5):128-142
王永贵,谷佳鸿.基于多尺度纹理感知与类激活分块注意力的图像分类网络.计算机系统应用,2026,35(5):128-142
WANG Yong-Gui,GU Jia-Hong.Image Classification Network Based on Multi-scale Texture Perception and Class-activated Block Attention.COMPUTER SYSTEMS APPLICATIONS,2026,35(5):128-142

