###
计算机系统应用英文版:2026,35(6):49-60
本文二维码信息
码上扫一扫!
融合视觉注意力模块的多视角图像字幕生成
(武汉轻工大学 数学与计算机学院, 武汉 430048)
Multi-view Image Subtitle Generation with Integrating Visual Attention Module
(School of Mathematics and Computer Science, Wuhan Polytechnic University, Wuhan 430048, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 80次   下载 87
Received:October 28, 2025    Revised:November 21, 2025
中文摘要: 近年来, 随着计算机视觉与自然语言处理领域的迅猛发展, 这一研究方向引起了学术界的广泛关注. 传统Transformer处理图像等数据时生成字幕缺乏细节, 往往忽略了图像的局部特征, 难以实现精准的语义对齐. 为了解决这个问题, 提出一种融合视觉注意模块的多视角图像字幕生成模型. 首先, 该模型将图像复制16份, 每份都加入独立的随机噪声, 每份在空间特征中都存在微小差异, 让模型拥有多个“视角”, 并添加可学习二维相对位置编码; 其次, 在两阶段特征优化阶段模拟不同区域之间的信息交互, 显式建模不同区域的上下文依赖关系, 自动生成一个更具判别性, 聚焦于显著区域的增强全局视觉表述, 提升对图像细和局部语义的捕捉能力. 文本嵌入通过两层交叉注意力机制直接查询全部视觉patch, 实现每个生成token对图像区域的选择性关注. 最后, 生成器完成图像字幕生成后, 系统会通过多维度自动化评估指标(如BLEUROUGECIDErMETEOR等)对生成的字幕进行判别, 确认生成文本是否能准确反映图像内容、具备流畅性和语义合理性, 从而衡量模型输出是否符合人类描述标准. 实验结果表明, 该方法与现有方法相比, BLEU1提升了5.17%, CIDEr提升了8.78%, ROUGE-L提升了12.52%, METEOR提升了17.24%.
中文关键词: 图像字幕  CLIP  稀疏注意力机制
Abstract:In recent years, as computer vision and natural language processing have developed rapidly, this research direction has caught widespread attention from the academic community. When traditional Transformer processes images and other data, the generated subtitles lack details and often overlook local features of the images, making it difficult to achieve accurate semantic alignment. To this end, this study proposes a multi-view image subtitle generation model with a fusion visual attention module. Firstly, the model copies the image 16 times, adds independent random noise to each copy, and each copy has slight differences in spatial features to make the model have multiple “perspectives”, with learnable two-dimensional relative position encoding added. Secondly, in the two-stage feature optimization stage, information interaction between different regions is simulated, explicitly modeling the contextual dependency between different regions. This automatically generates a more discriminative and enhanced global visual representation that focuses on salient regions, thereby improving the ability to capture subtle and local semantics of the images. Text embedding directly queries all visual patches via a two-layer cross-attention mechanism, enabling each generated token to selectively focus on image regions. Finally, after the generator completes the generation of image subtitles, the system will judge the generated subtitles via multi-dimensional automated evaluation indicators such as BLEU, ROUGE, CIDEr, and METEOR to confirm whether the generated text accurately reflects the image content, has fluency and semantic rationality, and thus measure whether the model output meets the human description standards. The experimental results show that compared with existing methods, this method improves BLEU1 by 5.17%, CIDEr by 8.78%, ROUGE-L by 12.52%, and METEOR by 17.24%.
文章编号:     中图分类号:    文献标志码:
基金项目:湖北省科技计划 (2025BBB004)
引用文本:
卢思雨,刘朔,杨洋.融合视觉注意力模块的多视角图像字幕生成.计算机系统应用,2026,35(6):49-60
LU Si-Yu,LIU Shuo,YANG Yang.Multi-view Image Subtitle Generation with Integrating Visual Attention Module.COMPUTER SYSTEMS APPLICATIONS,2026,35(6):49-60