本文已被:浏览 525次 下载 1685次
Received:August 11, 2024 Revised:September 24, 2024
Received:August 11, 2024 Revised:September 24, 2024
中文摘要: 为了提升音频驱动人体动画生成的真实性, 对UnifiedGesture模型进行了改进研究. 首先, 通过引入编码器-解码器架构, 从音频中提取面部特征, 以弥补原模型在面部表情生成方面的不足. 其次, 结合交叉局部注意力机制和基于Transformer-XL的多头注意力机制, 以增强长序列中的时序依赖性. 同时, 利用变分量化自动编码器 (vector quantized variational autoencoder, VQVAE), 融合生成全身运动序列, 以提升生成动作的多样性和完整性. 最后, 在BEAT数据集上进行实验, 通过定量和定性分析结果表明, 改进后的UnifiedGesture-F模型在音频与人体动作同步性和整体真实感方面相比原模型有显著提升.
中文关键词: 音频驱动 人体动画生成技术 UnifiedGesture模型 VQVAE
Abstract:This study researches improving the UnifiedGesture model to enhance the realism of audio-driven human body animation generation. Firstly, an encoder-decoder architecture is introduced to extract facial features from audio, compensating for the deficiencies of the original model in facial expression generation. Secondly, the cross-local attention mechanism and the multi-head attention mechanism based on Transformer-XL are combined to enhance the temporal dependency within long sequences. Simultaneously, the vector quantized variational autoencoder (VQVAE) is utilized to integrate and generate full-body motion sequences, enhancing the diversity and integrity of the generated motions. Finally, experiments are conducted on the BEAT dataset. The quantitative and qualitative analysis results demonstrate that the improved UnifiedGesture-F model achieves a significant improvement in the synchronicity between audio and human body movements as well as in the overall realism compared to the original model.
keywords: audio-driven human animation generation technique UnifiedGesture model vector quantized variational autoencoder (VQVAE)
文章编号: 中图分类号: 文献标志码:
基金项目:陕西省科技厅重点研发计划 (2024GX-YBXM-548)
引用文本:
郭二伟,朱欣娟,高全力.基于UnifiedGesture改进模型的三维人体动画生成.计算机系统应用,2025,34(3):40-50
GUO Er-Wei,ZHU Xin-Juan,GAO Quan-Li.3D Human Animation Generation Based on Improved UnifiedGesture Model.COMPUTER SYSTEMS APPLICATIONS,2025,34(3):40-50
郭二伟,朱欣娟,高全力.基于UnifiedGesture改进模型的三维人体动画生成.计算机系统应用,2025,34(3):40-50
GUO Er-Wei,ZHU Xin-Juan,GAO Quan-Li.3D Human Animation Generation Based on Improved UnifiedGesture Model.COMPUTER SYSTEMS APPLICATIONS,2025,34(3):40-50

