融合视觉注意力模块的多视角图像字幕生成
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

湖北省科技计划 (2025BBB004)


Multi-view Image Subtitle Generation with Integrating Visual Attention Module
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    近年来, 随着计算机视觉与自然语言处理领域的迅猛发展, 这一研究方向引起了学术界的广泛关注. 传统Transformer处理图像等数据时生成字幕缺乏细节, 往往忽略了图像的局部特征, 难以实现精准的语义对齐. 为了解决这个问题, 提出一种融合视觉注意模块的多视角图像字幕生成模型. 首先, 该模型将图像复制16份, 每份都加入独立的随机噪声, 每份在空间特征中都存在微小差异, 让模型拥有多个“视角”, 并添加可学习二维相对位置编码; 其次, 在两阶段特征优化阶段模拟不同区域之间的信息交互, 显式建模不同区域的上下文依赖关系, 自动生成一个更具判别性, 聚焦于显著区域的增强全局视觉表述, 提升对图像细和局部语义的捕捉能力. 文本嵌入通过两层交叉注意力机制直接查询全部视觉patch, 实现每个生成token对图像区域的选择性关注. 最后, 生成器完成图像字幕生成后, 系统会通过多维度自动化评估指标(如BLEUROUGECIDErMETEOR等)对生成的字幕进行判别, 确认生成文本是否能准确反映图像内容、具备流畅性和语义合理性, 从而衡量模型输出是否符合人类描述标准. 实验结果表明, 该方法与现有方法相比, BLEU1提升了5.17%, CIDEr提升了8.78%, ROUGE-L提升了12.52%, METEOR提升了17.24%.

    Abstract:

    In recent years, as computer vision and natural language processing have developed rapidly, this research direction has caught widespread attention from the academic community. When traditional Transformer processes images and other data, the generated subtitles lack details and often overlook local features of the images, making it difficult to achieve accurate semantic alignment. To this end, this study proposes a multi-view image subtitle generation model with a fusion visual attention module. Firstly, the model copies the image 16 times, adds independent random noise to each copy, and each copy has slight differences in spatial features to make the model have multiple “perspectives”, with learnable two-dimensional relative position encoding added. Secondly, in the two-stage feature optimization stage, information interaction between different regions is simulated, explicitly modeling the contextual dependency between different regions. This automatically generates a more discriminative and enhanced global visual representation that focuses on salient regions, thereby improving the ability to capture subtle and local semantics of the images. Text embedding directly queries all visual patches via a two-layer cross-attention mechanism, enabling each generated token to selectively focus on image regions. Finally, after the generator completes the generation of image subtitles, the system will judge the generated subtitles via multi-dimensional automated evaluation indicators such as BLEU, ROUGE, CIDEr, and METEOR to confirm whether the generated text accurately reflects the image content, has fluency and semantic rationality, and thus measure whether the model output meets the human description standards. The experimental results show that compared with existing methods, this method improves BLEU1 by 5.17%, CIDEr by 8.78%, ROUGE-L by 12.52%, and METEOR by 17.24%.

    参考文献
    相似文献
    引证文献
引用本文

卢思雨,刘朔,杨洋.融合视觉注意力模块的多视角图像字幕生成.计算机系统应用,2026,35(6):49-60

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-10-28
  • 最后修改日期:2025-11-21
  • 录用日期:
  • 在线发布日期: 2026-04-30
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号