本文已被:浏览 442次 下载 1190次
Received:March 24, 2025 Revised:May 07, 2025
Received:March 24, 2025 Revised:May 07, 2025
中文摘要: 针对现有音视频语音增强方法在复杂场景下存在的鲁棒性不足、多模态信息融合效率低下、计算复杂度高等问题, 本文提出一种双流门控视听融合(dual-stream gated audio-visual fusion, DS-GAVF)架构. 该架构通过静态与动态视觉协同建模, 以及细粒度的跨模态交互, 实现语音增强性能的提升. 在特征提取阶段, 采用U-Net编码音频时频特性, 同时设计双流视觉网络, 采用ResNet-18提取单帧面部静态特征, 使用时空图卷积网络捕捉连续面部标志点的动态运动模式. 为解决视听模态时序差异, 提出动态时间插值对齐方法, 并设计了一种门控交叉注意力机制, 通过门控单元动态调节特征融合权重, 在视觉可信度低时抑制噪声干扰. 在解码阶段, 通过跨模态跳跃连接, 将多尺度视觉线索注入U-Net解码层, 最终输出目标语音时频掩码. 实验结果表明, DS-GAVF在3个数据集上的混合噪声环境下均表现出优异性能. 与现有方法相比, DS-GAVF在保持较低计算复杂度的同时, 实现了语音质量、可懂度与鲁棒性的协同提升.
Abstract:Given the problems of insufficient robustness, low efficiency of multimodal information fusion, and high computational complexity in existing audio-visual speech enhancement methods under complex scenarios, a dual-stream gated audio-visual fusion (DS-GAVF) architecture is proposed. This architecture achieves performance improvement in speech enhancement via static-dynamic visual collaborative modeling and fine-grained cross-modal interaction. In the feature extraction stage, U-Net is employed to encode the audio’s time-frequency characteristics, and meanwhile a dual-stream visual network is designed. ResNet-18 is adopted to extract static features of single-frame facial images, and a spatiotemporal graph convolutional network is utilized to capture dynamic motion patterns of continuous facial landmarks. Additionally, a dynamic temporal interpolation alignment method is proposed to solve the time series difference of audio-visual modalities, and a novel gated cross-attention mechanism is designed, which dynamically adjusts the feature fusion weights via a gated unit and suppresses noise interference when visual reliability is low. In the decoding stage, cross-modal jump connections are leveraged to inject multi-scale visual cues into the U-Net decoding layers, finally outputting the time-frequency mask of the target speech. Experimental results show that DS-GAVF yields excellent performance under mixed noise environments on three datasets. Compared with existing methods, DS-GAVF achieves a synergistic improvement in speech quality, intelligibility, and robustness while maintaining low computational complexity.
文章编号: 中图分类号: 文献标志码:
基金项目:广东省自然科学基金面上项目(2025A1515011526)
引用文本:
彭敏轩,梁艳.基于双流门控视听融合的多模态语音增强.计算机系统应用,2025,34(11):127-138
PENG Min-Xuan,LIANG Yan.Multimodal Speech Enhancement Based on Dual-stream Gated Audio-visual Fusion.COMPUTER SYSTEMS APPLICATIONS,2025,34(11):127-138
彭敏轩,梁艳.基于双流门控视听融合的多模态语音增强.计算机系统应用,2025,34(11):127-138
PENG Min-Xuan,LIANG Yan.Multimodal Speech Enhancement Based on Dual-stream Gated Audio-visual Fusion.COMPUTER SYSTEMS APPLICATIONS,2025,34(11):127-138

