基于多尺度Mamba的长视频微表情定位
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

广东省自然科学基金面上项目(2025A1515011526)


Micro-expression Spotting on Long Video Based on Multi-scale Mamba
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    基于长视频的微表情定位旨在从长视频序列中精准检测出极短时间内面部微小的肌肉运动. 针对现有方法在处理背景噪声干扰、长程时序依赖建模以及微表情尺度多样性方面的不足, 本文提出一种基于多尺度 Mamba的微表情定位方案. 首先, 通过改进特征金字塔网络, 采用双线性插值策略优化多尺度特征的融合与对齐, 采用卷积块注意力模块优化横向连接, 提取更具判别性的时序特征; 其次, 设计多尺度Mamba适配器, 利用其选择性状态空间机制, 在保持低计算复杂度的同时捕捉长程时序信息, 并结合位置编码增强模型对动作起止边界的感知力; 最后, 采用分类与回归解耦的检测头, 引入变焦损失函数和分布焦点损失函数进行协同优化. 实验结果表明, 本文方法在SAMM-LV数据集上的F1分数达到0.3025, 相较于基准模型提升了102.3%; 而在样本规模较小的CAS(ME)2数据集上亦取得了0.2188的F1分数, 展现出优于多数主流深度学习方法的竞争性能, 验证了该模型在复杂长视频场景下定位微表情的有效性与鲁棒性.

    Abstract:

    Micro-expression spotting in long videos aims to accurately detect subtle facial muscle movements of extremely short duration. To address the limitations of existing methods in coping with background noise, modeling long-range temporal dependencies, and handling the scale diversity of micro-expressions, a micro-expression spotting framework based on multi-scale Mamba is proposed. First, an improved feature pyramid network is introduced. It uses bilinear interpolation to optimize the fusion and alignment of multi-scale features and incorporates a convolutional block attention module (CBAM) to refine lateral connections and extract more discriminative temporal features. Second, a multi-scale Mamba adapter is designed to capture long-range temporal information through its selective state-space mechanism while maintaining low computational complexity. It is further combined with positional encoding to enhance the model’s ability to detect onset and offset boundaries. Finally, a decoupled detection head is employed for classification and regression, and varifocal loss and distribution focal loss are incorporated to optimize the two tasks jointly. Experimental results demonstrate that the proposed method achieves an F1-score of 0.3025 on SAMM-LV dataset, representing a 102.3% improvement over the baseline model. On the smaller CAS(ME)2 dataset, the model achieves a competitive F1-score of 0.2188 and outperforms most mainstream deep learning approaches. These results validate the effectiveness and robustness of the proposed model in spotting micro-expressions within complex long video scenarios.

    参考文献
    相似文献
    引证文献
引用本文

姚泽楠,王紫凡,梁艳.基于多尺度Mamba的长视频微表情定位.计算机系统应用,,():1-11

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-01-15
  • 最后修改日期:2026-02-09
  • 录用日期:
  • 在线发布日期: 2026-07-17
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号