###
计算机系统应用英文版:2025,34(3):40-50
本文二维码信息
码上扫一扫!
基于UnifiedGesture改进模型的三维人体动画生成
(西安工程大学 计算机科学学院, 西安 710600)
3D Human Animation Generation Based on Improved UnifiedGesture Model
(School of Computer Science, Xi’an Polytechnic University, Xi’an 710600, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 525次   下载 1685
Received:August 11, 2024    Revised:September 24, 2024
中文摘要: 为了提升音频驱动人体动画生成的真实性, 对UnifiedGesture模型进行了改进研究. 首先, 通过引入编码器-解码器架构, 从音频中提取面部特征, 以弥补原模型在面部表情生成方面的不足. 其次, 结合交叉局部注意力机制和基于Transformer-XL的多头注意力机制, 以增强长序列中的时序依赖性. 同时, 利用变分量化自动编码器 (vector quantized variational autoencoder, VQVAE), 融合生成全身运动序列, 以提升生成动作的多样性和完整性. 最后, 在BEAT数据集上进行实验, 通过定量和定性分析结果表明, 改进后的UnifiedGesture-F模型在音频与人体动作同步性和整体真实感方面相比原模型有显著提升.
Abstract:This study researches improving the UnifiedGesture model to enhance the realism of audio-driven human body animation generation. Firstly, an encoder-decoder architecture is introduced to extract facial features from audio, compensating for the deficiencies of the original model in facial expression generation. Secondly, the cross-local attention mechanism and the multi-head attention mechanism based on Transformer-XL are combined to enhance the temporal dependency within long sequences. Simultaneously, the vector quantized variational autoencoder (VQVAE) is utilized to integrate and generate full-body motion sequences, enhancing the diversity and integrity of the generated motions. Finally, experiments are conducted on the BEAT dataset. The quantitative and qualitative analysis results demonstrate that the improved UnifiedGesture-F model achieves a significant improvement in the synchronicity between audio and human body movements as well as in the overall realism compared to the original model.
文章编号:     中图分类号:    文献标志码:
基金项目:陕西省科技厅重点研发计划 (2024GX-YBXM-548)
引用文本:
郭二伟,朱欣娟,高全力.基于UnifiedGesture改进模型的三维人体动画生成.计算机系统应用,2025,34(3):40-50
GUO Er-Wei,ZHU Xin-Juan,GAO Quan-Li.3D Human Animation Generation Based on Improved UnifiedGesture Model.COMPUTER SYSTEMS APPLICATIONS,2025,34(3):40-50