###
计算机系统应用英文版:,():1-13
本文二维码信息
码上扫一扫!
伪造痕迹与语义逻辑协同建模的媒体伪造检测
(西北工业大学 计算机学院, 西安 710129)
Media Manipulation Detection via Collaborative Modeling of Forgery Traces andSemantic-logical
(School of Computer Science, Northwestern Polytechnical University, Xi’an 710129, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 24次   下载 31
Received:January 24, 2026    Revised:February 27, 2026
中文摘要: 多模态媒体伪造检测旨在建模图像与文本等异构模态信息, 实现对图文内容真实性的精准判别. 然而, 现有方法在图像频域特征建模方面较为薄弱, 难以充分捕捉隐藏于频域中的细粒度篡改痕迹; 同时, 对文本内部矛盾与语义不一致等逻辑性伪造线索的挖掘能力不足, 限制了文本模态判别信息的有效利用. 为解决上述问题, 本文提出一种多模态媒体伪造检测模型TruthSeeker, 其包含3个核心部分: (1)伪造线索挖掘模块, 通过构建频域提取-基于$ k $近邻的聚类-特征压缩的层次化特征学习机制, 联合建模图像空域与频域信息, 以增强对局部篡改痕迹的感知能力; (2)文本矛盾增强模块, 利用大语言模型从事实违背、情感上下文不一致以及文本表达口语化这3个角度挖掘潜在语义冲突, 增强文本模态中的逻辑性伪造线索; (3)深层伪造推理机制, 基于跨模态交叉注意力实现图像与文本伪造特征的深度交互, 并通过多任务损失函数进行协同优化. 实验结果表明, TruthSeeker在二分类伪造识别任务中取得99.49%的AUC, 并在多项对比实验中优于现有方法, 验证了模型在多模态媒体伪造检测任务中的有效性.
Abstract:Multimodal media manipulation detection aims to model heterogeneous information such as images and text to achieve accurate authenticity assessment of image-text content. However, existing approaches have limited capability in modeling image frequency-domain features, making it difficult to fully capture fine-grained tampering traces. Meanwhile, they are insufficient at exploring logical forgery cues in textual content, such as internal contradictions and semantic inconsistencies, which restricts the effective use of discriminative information in the textual modality. To address these issues, this study proposes a multimodal media manipulation detection model, TruthSeeker, which consists of three core components: (1) a forgery clue mining module, which constructs a hierarchical feature learning mechanism comprising frequency-domain extraction, k-nearest neighbor (kNN)-based clustering, and feature compression, thus jointly modeling spatial- and frequency-domain information to enhance the perception of local tampering traces; (2) a textual contradiction enhancement module, which leverages a large language model to mine potential semantic conflicts from three aspects: factual inconsistency, emotional-context inconsistency, and colloquial expression, enhancing logical forgery cues in the textual modality; and (3) a deep forgery reasoning mechanism, which is based on cross-modal cross-attention to enable deep interaction between image and text forgery features, and performs collaborative optimization through a multi-task loss function. Experimental results demonstrate that TruthSeeker achieves an AUC of 99.49% in binary forgery detection and outperforms existing methods in multiple comparative experiments, validating the effectiveness of the model in multimodal media manipulation detection tasks.
文章编号:     中图分类号:    文献标志码:
基金项目:国家自然科学基金面上项目(62572403); 国家自然科学基金企业创新发展联合基金(U22B2036); 广东省基础与应用基础研究基金(2024A1515010087)
引用文本:
王康,吴连伟,吕佳航,王博韬.伪造痕迹与语义逻辑协同建模的媒体伪造检测.计算机系统应用,,():1-13
WANG Kang,WU Lian-Wei,LYU Jia-Hang,WANG Bo-Tao.Media Manipulation Detection via Collaborative Modeling of Forgery Traces andSemantic-logical.COMPUTER SYSTEMS APPLICATIONS,,():1-13