本文已被:浏览 632次 下载 1529次
Received:August 24, 2024 Revised:September 19, 2024
Received:August 24, 2024 Revised:September 19, 2024
中文摘要: 视音频事件定位(audio-visual event localization, AVEL)任务通过观察音频信息和相对应的视觉信息来定位视频中的事件. 本文针对AVEL任务设计了一种跨模态时间对齐网络CMTAN, 该网络包含预处理、跨模态交互、时间对齐和特征融合这4个部分. 具体而言, 在预处理部分, 通过一种新的跨模态音频指导模块和一种噪音弱化模块的处理, 模态信息中的背景和噪音被减少. 随后, 在跨模态交互部分, 使用基于多头注意力机制的信息强化和信息补充模块进行跨模态交互, 单模态信息得到全局信息优化. 在时间对齐部分, 本文设计了一种聚焦于跨模态交互前后单模态全局信息的时间对齐模块, 用于执行模态信息的特征对齐. 最后, 在特征融合过程中, 通过一种多阶段融合模块, 两种模态信息被从浅入深地融合, 且融合后的模态信息最终将被用于事件定位. 大量实验表明CMTAN在弱监督和全监督AVEL任务中都具有优秀的性能.
Abstract:The audio-visual event localization (AVEL) task locates events in a video by observing audio information and corresponding visual information. In this paper, a cross-modal time alignment network named CMTAN is designed for the AVEL task. The network consists of four parts: preprocessing, cross-modal interaction, time alignment, and feature fusion. Specifically, in the preprocessing part, the background and noise in the modal information are reduced by the processing of a new cross-modal audio guidance module and a noise reduction module. Then, in the cross-modal interaction part, the information reinforcement and information complementation modules based on the multi-head attention mechanism are used for cross-modal interaction, and the unimodal information is optimized with global information. In the time alignment part, a time alignment module focusing on the unimodal global information before and after cross-modal interaction is designed to perform feature alignment of modal information. Finally, in the feature fusion process, two kinds of modal information are fused from shallow to deep by a multi-stage fusion module. The fused modal information is ultimately used for event localization. Extensive experiments demonstrate that CMTAN has excellent performance in both weakly and fully supervised AVEL tasks.
keywords: cross-modal audio-visual event localization (AVEL) weakly and fully supervised feature alignment
文章编号: 中图分类号: 文献标志码:
基金项目:重庆市教育科学规划重点课题(K22YE205098); 重庆师范大学博士启动基金(21XLB030, 21XLB029)
引用文本:
王志豪,訾玲玲.面向视音频事件定位的跨模态时间对齐网络.计算机系统应用,2025,34(3):133-142
WANG Zhi-Hao,ZI Ling-Ling.Cross-modal Time Alignment Network for Audio-visual Event Localization.COMPUTER SYSTEMS APPLICATIONS,2025,34(3):133-142
王志豪,訾玲玲.面向视音频事件定位的跨模态时间对齐网络.计算机系统应用,2025,34(3):133-142
WANG Zhi-Hao,ZI Ling-Ling.Cross-modal Time Alignment Network for Audio-visual Event Localization.COMPUTER SYSTEMS APPLICATIONS,2025,34(3):133-142

