本文已被:浏览 57次 下载 74次
Received:October 26, 2025 Revised:November 21, 2025
Received:October 26, 2025 Revised:November 21, 2025
中文摘要: 视觉定位旨在依据自然语言描述确定图像中目标地物的空间位置, 对于推动人机交互技术的发展具有重要意义. 近年来, 遥感图像中的视觉定位任务逐渐受到广泛关注. 与自然图像中的视觉定位相比, 遥感视觉定位面临诸多新挑战, 例如文本描述往往更为复杂, 目标地物在图像中所占比例较小, 定位过程更依赖于全局上下文信息等. 为解决上述问题, 本文提出了一种渐进式融合多模态与多尺度的遥感视觉定位方法. 该方法通过在多模态融合过程中动态调整文本特征权重, 逐步融入上下文信息, 以增强对复杂文本描述的理解能力; 同时, 利用深层特征引导浅层特征进行多尺度融合, 在充分保留多模态语义信息的基础上, 补充细节信息以提升定位精度. 本文在RSVG与DIOR-RSVG两个公开数据集上进行了实验验证, 实验结果表明, 本方法与现有方法相比具有显著优势.
Abstract:Visual grounding aims to locate target objects in images based on natural language descriptions, which is crucial for promoting the development of human-computer interaction technology. In recent years, visual grounding tasks in remote sensing images have gradually caught attention. Compared with visual grounding in natural images, remote sensing visual grounding faces numerous new challenges, such as more complex textual descriptions, a smaller proportion occupied by target objects in the images, and greater reliance on global contextual information in localization processes. To this end, this study proposes a remote sensing visual grounding method based on progressive fusion multimodal and multi-scale. By dynamically adjusting the weights of textual features during multimodal fusion, the proposed method gradually fuses contextual information to enhance the ability to comprehend complex textual descriptions. Meanwhile, deep features are utilized to guide the multi-scale fusion of shallow features, thereby supplementing detailed information to improve localization precision on the basis of fully preserving multimodal semantic information. Experiments are conducted on two public datasets, RSVG and DIOR-RSVG, and the experimental results demonstrate that the proposed method has significant advantages over existing methods.
文章编号: 中图分类号: 文献标志码:
基金项目:
引用文本:
蔡治诚,徐偲淇,颜志雄.渐进式多模态多尺度融合的遥感视觉定位.计算机系统应用,2026,35(6):144-155
CAI Zhi-Cheng,XU Si-Qi,YAN Zhi-Xiong.Remote Sensing Visual Grounding Based on Progressive Fusion of Multimodal and Multi-scale.COMPUTER SYSTEMS APPLICATIONS,2026,35(6):144-155
蔡治诚,徐偲淇,颜志雄.渐进式多模态多尺度融合的遥感视觉定位.计算机系统应用,2026,35(6):144-155
CAI Zhi-Cheng,XU Si-Qi,YAN Zhi-Xiong.Remote Sensing Visual Grounding Based on Progressive Fusion of Multimodal and Multi-scale.COMPUTER SYSTEMS APPLICATIONS,2026,35(6):144-155

