###
计算机系统应用英文版:2025,34(9):112-119
本文二维码信息
码上扫一扫!
基于KLSTM的互信息视听情绪识别
(1.西南科技大学 信息与控制工程学院, 绵阳 621000;2.四川省工业自主可控人工智能工程技术研究中心, 绵阳 621010;3.四川湖山电器股份有限公司, 绵阳 621025)
KLSTM-based Mutual Information Audiovisual Emotion Recognition
(1.School of Information and Control Engineering, Southwest University of Science and Technology, Mianyang 621000, China;2.Sichuan Engineering and Technology Research Center for Industrial Autonomous and Controllable Artificial Intelligence, Mianyang 621010, China;3.Sichuan Hushan Electrical Appliance Co. Ltd., Mianyang 621025, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 453次   下载 1058
Received:January 07, 2025    Revised:February 12, 2025
中文摘要: 针对视听情绪识别中如何高效融合音视频特征并准确提取时间依赖性情绪信息的问题, 本研究提出了一种基于KLSTM (Kolmogorov-Arnold long short-term memory)的互信息视听情绪识别模型. 利用互信息方法进行特征选择和自适应窗口处理, 从音频和视频信号中提取情绪相关的关键片段, 以减少信息冗余; 在特征提取中引入KLSTM网络, 有效捕捉视听模态信号的时间依赖特性; 在融合阶段, 通过跨模态一致性最大化确保音视频特征的协调性与互补性. 实验结果显示所提模型在CMU-MOSI和CMU-MOSEI数据集上均优于现有基准模型, 验证了其在多模态情绪识别任务中的有效性.
中文关键词: 情绪识别  视听融合  互信息
Abstract:Addressing the challenge of efficiently fusing audio and video features while accurately extracting time-dependent emotion information in audiovisual emotion recognition, a mutual information-based audiovisual emotion recognition model is proposed, incorporating Kolmogorov-Arnold long short-term memory (KLSTM). Feature selection and adaptive window processing, based on the mutual information approach, are employed to extract emotionally relevant key segments from audio and video signals, effectively reducing information redundancy. The KLSTM network is integrated into feature extraction to capture the temporal dependencies of audiovisual modal signals. In the fusion stage, cross-modal consistency maximization ensures the coordination and complementarity of audio and video features. Experimental results demonstrate that the proposed model outperforms existing benchmark models on both CMU-MOSI and CMU-MOSEI datasets, validating its effectiveness in multimodal emotion recognition tasks.
文章编号:     中图分类号:    文献标志码:
基金项目:国家自然科学基金(U21A20157)
引用文本:
罗志鑫,刘知贵,唐荣,潘志祥,李理.基于KLSTM的互信息视听情绪识别.计算机系统应用,2025,34(9):112-119
LUO Zhi-Xin,LIU Zhi-Gui,TANG Rong,PAN Zhi-Xiang,LI Li.KLSTM-based Mutual Information Audiovisual Emotion Recognition.COMPUTER SYSTEMS APPLICATIONS,2025,34(9):112-119