伪造痕迹与语义逻辑协同建模的媒体伪造检测
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金面上项目(62572403); 国家自然科学基金企业创新发展联合基金(U22B2036); 广东省基础与应用基础研究基金(2024A1515010087)


Media Manipulation Detection via Collaborative Modeling of Forgery Traces andSemantic-logical
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    多模态媒体伪造检测旨在建模图像与文本等异构模态信息, 实现对图文内容真实性的精准判别. 然而, 现有方法在图像频域特征建模方面较为薄弱, 难以充分捕捉隐藏于频域中的细粒度篡改痕迹; 同时, 对文本内部矛盾与语义不一致等逻辑性伪造线索的挖掘能力不足, 限制了文本模态判别信息的有效利用. 为解决上述问题, 本文提出一种多模态媒体伪造检测模型TruthSeeker, 其包含3个核心部分: (1)伪造线索挖掘模块, 通过构建频域提取-基于$ k $近邻的聚类-特征压缩的层次化特征学习机制, 联合建模图像空域与频域信息, 以增强对局部篡改痕迹的感知能力; (2)文本矛盾增强模块, 利用大语言模型从事实违背、情感上下文不一致以及文本表达口语化这3个角度挖掘潜在语义冲突, 增强文本模态中的逻辑性伪造线索; (3)深层伪造推理机制, 基于跨模态交叉注意力实现图像与文本伪造特征的深度交互, 并通过多任务损失函数进行协同优化. 实验结果表明, TruthSeeker在二分类伪造识别任务中取得99.49%的AUC, 并在多项对比实验中优于现有方法, 验证了模型在多模态媒体伪造检测任务中的有效性.

    Abstract:

    Multimodal media manipulation detection aims to model heterogeneous information such as images and text to achieve accurate authenticity assessment of image-text content. However, existing approaches have limited capability in modeling image frequency-domain features, making it difficult to fully capture fine-grained tampering traces. Meanwhile, they are insufficient at exploring logical forgery cues in textual content, such as internal contradictions and semantic inconsistencies, which restricts the effective use of discriminative information in the textual modality. To address these issues, this study proposes a multimodal media manipulation detection model, TruthSeeker, which consists of three core components: (1) a forgery clue mining module, which constructs a hierarchical feature learning mechanism comprising frequency-domain extraction, k-nearest neighbor (kNN)-based clustering, and feature compression, thus jointly modeling spatial- and frequency-domain information to enhance the perception of local tampering traces; (2) a textual contradiction enhancement module, which leverages a large language model to mine potential semantic conflicts from three aspects: factual inconsistency, emotional-context inconsistency, and colloquial expression, enhancing logical forgery cues in the textual modality; and (3) a deep forgery reasoning mechanism, which is based on cross-modal cross-attention to enable deep interaction between image and text forgery features, and performs collaborative optimization through a multi-task loss function. Experimental results demonstrate that TruthSeeker achieves an AUC of 99.49% in binary forgery detection and outperforms existing methods in multiple comparative experiments, validating the effectiveness of the model in multimodal media manipulation detection tasks.

    参考文献
    相似文献
    引证文献
引用本文

王康,吴连伟,吕佳航,王博韬.伪造痕迹与语义逻辑协同建模的媒体伪造检测.计算机系统应用,,():1-13

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-01-24
  • 最后修改日期:2026-02-27
  • 录用日期:
  • 在线发布日期: 2026-06-18
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号