论文多源确认

CCDF:面向真实监控场景的深度伪造视频检测基准数据集

CCDF: A Benchmark Dataset for Deepfake Detection in Real-World Surveillance Footage

精选理由

有人用 VEO 3.1、Sora 2 这些工具伪造监控视频做了一份测试集,10 个主流检测器全都没扛住,研究造假检测的可以看看。

研究团队发布了视频伪造检测数据集 CCDF,包含 1840 段视频(460 段真实、1380 段生成),覆盖 16 类犯罪与事故场景。生成内容来自 Grok Imagine、Google VEO 3.1 和 OpenAI Sora 2 三款商用视频工具。团队还提供去除元数据线索的清洗版和模拟低门槛后期攻击的变体版。用 10 个最新检测器评估后发现,它们在现有基准上表现良好,却无法可靠区分 CCDF 中的生成视频与真实视频。

原文 · arXiv: OpenAI

CCDF: A Benchmark Dataset for Deepfake Detection in Real-World Surveillance Footage

Due to rapid advances in Generative AI, commercial video generation tools can be used to produce fabricated surveillance footage that can fool both human viewers and automated synthetic video detectors. Since these tools are so widely accessible, a malicious user can create a harmful video clip at minimal cost. The production and dissemination of such videos in high-stakes settings, such as crime reporting and elections, can misdirect emergency response efforts or distort political discourse. Existing deepfake video datasets, used by the research community to develop deepfake detection algorithms, exhibit two limitations: (1) they emphasize benign web content rather than footage of possibly malicious activity, and (2) they rely on older or open-source generators that do not represent recent advances in generative systems. We assemble CCtv DeepFakes (CCDF), a video deepfake dataset, to address both gaps. CCDF contains 1840 videos (460 real and 1380 generated) spanning 16 crime and accident categories, with generated content produced using three leading commercial systems: Grok Imagine, Google VEO 3.1, and OpenAI Sora 2. CCDF is a highly realistic, small-scale, manually annotated dataset targeting evaluation of detection models. We release three versions of the dataset: the raw generated data, a cleaned version in which video metadata are standardized between real and synthetic samples to prevent detectors from exploiting trivial cues, and an altered version simulating low-effort post-processing attacks. We evaluate CCDF with ten recent state-of-the-art detectors covering different detection approaches. Our results suggest that these approaches do not reliably distinguish CCDF's generated videos from real ones, despite their strong reported performance on existing datasets. These results further confirm that existing datasets are not well-suited to evaluating certain threats.

  • IT之家10-06 02:41原文
  • Thomas Wolf10-06 23:15原文
  • VTV Công nghệ10-04 12:03原文
  • TestingCatalog10-05 07:37原文
  • Rappler: Technology10-05 07:57原文
  • Tibor Blaho10-05 15:16原文
  • Andrew Curran10-05 16:21原文
  • theverge10-05 18:08原文
  • lmarena.ai10-05 19:02原文
  • Dylan Patel (SemiAnalysis)10-05 23:58原文