Scoop: Mitigation of Recapture Attacks on Provenance-Based Media Authentication

Yuxin (Myles) Liu (UC Irvine)

34th USENIX Security Symposium (USENIX Security '25) · Day 2 · System Security 3: Mobile Platforms

Overview

In an era dominated by rapidly spreading digital information and the proliferation of sophisticated generative AI, distinguishing authentic content from fabricated material has become an increasingly critical challenge. This talk, "Scoop: Mitigation of Recapture Attacks on Provenance-Based Media Authentication," by Yuxin (Myles) Liu and collaborators from UC Irvine and Microsoft, addresses a significant vulnerability in emerging provenance-based media authentication systems. While these systems cryptographically prove the origin and history of digital content, they are inherently blind to manipulations that occur within the content itself, particularly through recapture attacks.

Watch on YouTube · Slides

Visual summary for Scoop: Mitigation of Recapture Attacks on Provenance-Based Media Authentication by Yuxin (Myles) Liu
Visual summary for Scoop: Mitigation of Recapture Attacks on Provenance-Based Media Authentication by Yuxin (Myles) Liu

Key moments

  1. 0:00 Introduction to provenance-based media authentication
  2. 2:15 Challenge 1: Provenance doesn't authenticate content
  3. 2:40 Detailed example of a recapture attack scenario
  4. 4:20 Why recapture attacks are a renewed threat
  5. 7:20 User study: Humans fail to detect recaptured images
  6. 8:00 Introducing depth-enabled provenance as a solution
  7. 9:00 Challenge 2: Not all flat surfaces are recaptures
  8. 10:15 SCOOP: Leveraging human perceptual depth for detection

Scoop: Mitigation of Recapture Attacks on Provenance-Based Media Authentication

Speakers: Yuxin (Myles) Liu, UC Irvine; Habiba, Ardelon, Jin, UC Irvine; Sherat, Microsoft

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=Nd3WqvKWmwI

Overview

In an era dominated by rapidly spreading digital information and the proliferation of sophisticated generative AI, distinguishing authentic content from fabricated material has become an increasingly critical challenge. This talk, "Scoop: Mitigation of Recapture Attacks on Provenance-Based Media Authentication," by Yuxin (Myles) Liu and collaborators from UC Irvine and Microsoft, addresses a significant vulnerability in emerging provenance-based media authentication systems. While these systems cryptographically prove the origin and history of digital content, they are inherently blind to manipulations that occur within the content itself, particularly through recapture attacks.

The presentation introduces Scoop, an innovative countermeasure that leverages depth-enabled provenance to effectively detect and highlight regions of an image or video that have been subjected to recapture. By integrating time-of-flight (ToF) depth sensors into the content capture process and comparing the resulting real-world depth data with estimated perceptual depths, Scoop provides a robust mechanism to restore trust in digital media. This work is crucial for safeguarding against the deceptive potential of deepfakes and other forms of visual content manipulation, offering a proactive defense where traditional detection methods have repeatedly failed.

The significance of Scoop extends beyond academic interest, offering practical implications for industries reliant on trustworthy visual evidence, such as insurance, journalism, and law enforcement. As provenance systems like the Content Authenticity Initiative (CAI) and C2PA gain traction, understanding and mitigating recapture attacks becomes paramount to prevent a false sense of security among users. Scoop represents a vital step towards building a more resilient and trustworthy digital information ecosystem.

Background

▶ Watch: Introduction to provenance-based media authentication (0:00)

The digital age, characterized by the rapid dissemination of information across social media platforms, has unfortunately also accelerated the spread of misinformation and fraudulent content. The advent of advanced technologies like deepfakes and generative AI has further exacerbated this problem, making it easier than ever to create highly convincing yet entirely fabricated visual media. Traditional detection-based approaches, which attempt to identify manipulated content after the fact, have proven to be an "endless race" that is increasingly difficult to win, as attackers continuously evolve their techniques.

In response to this challenge, a new paradigm known as provenance-based media authentication has emerged. This approach aims to provide a cryptographic proof for visual content, recording its entire history from capture to publication. In a provenance-enabled system, a camera (e.g., a smartphone) generates a provenance alongside the visual content upon capture. Any subsequent modifications, such as post-processing or sharing on a platform, result in updates to this provenance information, creating an auditable trail. Industry giants like Adobe and Microsoft, alongside many others, are investing heavily in initiatives like the Content Authenticity Initiative (CAI) and C2PA (Coalition for Content Provenance and Authenticity) to standardize and implement these systems.

Despite their promise, provenance systems face a critical limitation: they authenticate the origin and history of content, but not the information within the content itself. This vulnerability is exploited by recapture attacks, where an attacker displays manipulated content on a screen (or prints it) and then re-photographs it with a provenance-enabled camera, effectively "re-originating" the fraudulent content with a new, seemingly legitimate provenance. The talk illustrates this with a compelling example: a facility manager photographs an empty fire cabinet, uses generative AI to digitally insert an extinguisher, and then recaptures the image of the screen with their provenance-enabled camera, creating a "verified" photo of a properly equipped cabinet. More dangerous scenarios include insurance fraud or evidence tampering.

Recapture attacks are not new, but their threat is renewed due to several factors:

  1. Ease of Manipulation: Deepfake and generative AI technologies make content fabrication simpler and more realistic.
  2. Illusion of Reality: Advanced displays (e.g., windowless buses using screens as windows) blur the lines between reality and digital representation.
  3. False Sense of Trust: The widespread adoption of provenance systems will lead users to implicitly trust all content bearing provenance, making them more susceptible to sophisticated recapture attacks.

Existing countermeasures against recapture attacks, broadly categorized into screen-based and print-based methods, primarily focus on detecting tiny, detailed patterns inherent in the recapture medium (e.g., screen pixel grids, print artifacts). However, these methods are fundamentally flawed because they depend on prior knowledge of the recapture medium and the camera used. Attackers, being in control, can easily bypass these detections by manipulating either the display, the printer, or the camera settings. A related paper presented at USENIX Security 2025 explicitly demonstrates how such bypasses are achieved. Furthermore, human detection of recaptured images is remarkably poor; a user study involving 43 participants showed an average correct classification rate of only 50%, akin to "pure randomness." This underscores the urgent need for an automated, robust solution.

Key Findings

▶ Watch: Detailed example of a recapture attack scenario (2:40)

The talk highlights several critical findings that underscore the problem and present Scoop as a viable solution:

  1. Provenance's Blind Spot: While provenance-based media authentication is a crucial step forward, it fundamentally authenticates the origin and history of content, not the fidelity of the information contained within the visual data. This leaves a significant gap that sophisticated attackers can exploit.
  2. The Renewed Threat of Recapture Attacks: Recapture attacks, though not novel, are gaining unprecedented potency due to advancements in generative AI and deepfake technologies, which facilitate highly realistic content manipulation, and the proliferation of high-fidelity displays. The impending widespread adoption of provenance systems will create a dangerous "false sense of trust," making these attacks more effective than ever.
  3. Failure of Existing Detection Methods: Current academic and industry efforts to detect recapture attacks are largely ineffective. These methods rely on identifying subtle artifacts (e.g., screen patterns, print characteristics) that are specific to the recapture medium and camera. Attackers, having control over both, can easily bypass these detections, rendering them obsolete.
  4. Human Inability to Detect Recaptures: A user study demonstrated that human observers are no better than random chance (50% accuracy) at distinguishing original photos from skillfully recaptured ones. This confirms that human intuition cannot serve as a reliable defense mechanism against these attacks.
  5. The Power of Depth-Enabled Provenance: The core insight is that depth information provides an immutable physical property of the captured scene that is fundamentally altered during a recapture event. A real-world scene has varying depths, whereas a recaptured image, taken from a screen or print, will appear as a flat surface in terms of depth.
  6. Scoop as an Effective Countermeasure: The proposed system, Scoop, leverages this principle by integrating time-of-flight (ToF) depth sensors into the camera to capture depth-enabled provenance. By comparing this ground-truth depth data with an estimated "perceptual depth" (derived from learning-based monocular depth estimation algorithms), Scoop can reliably identify and highlight regions of an image that exhibit anomalous flatness, indicating a recapture. This approach moves beyond an "endless race" of detection to a more fundamental authentication of content integrity.

Technical Deep Dive

▶ Watch: User study: Humans fail to detect recaptured images (7:20)

Scoop's technical innovation lies in its novel use of depth information to counter recapture attacks, shifting the focus from post-facto image analysis to integrating an additional layer of authentication at the point of capture. The system is built upon the principle that a real-world 3D scene, when re-photographed from a 2D display (or print), will exhibit a distinct and measurable change in its depth profile.

The foundation of Scoop is depth-enabled provenance. This requires cameras, particularly modern smartphones, to be equipped with time-of-flight (ToF) depth sensors. These sensors, which can be either direct time-of-flight or indirect time-of-flight, measure the absolute distance from the camera to various points in the scene. When a photo or video is captured, the ToF sensor simultaneously records this depth information, which is then cryptographically bound to the visual content as part of its provenance. This creates a ground-truth depth point cloud for the captured scene.

The initial intuition might be to simply flag any content exhibiting a flat depth profile as a recapture. However, this approach is insufficient because many legitimate scenes, such as a photo of a wall, are inherently flat. The challenge, therefore, is to distinguish between benign flat surfaces and maliciously flattened recapture surfaces. Scoop addresses this by comparing the physically measured provenance depths with what a human would perceive as the depth of the scene.

Since directly comparing human perception is impractical, Scoop leverages learning-based monocular depth estimation algorithms. These algorithms, trained on vast datasets, can infer the 3D structure and depth information of a scene from a single 2D RGB image, effectively approximating human perceptual depth. This process generates a learning-based point cloud from the visual content.

The core workflow of Scoop involves the following steps:

  1. Input: Scoop receives a digital content (photo or video frame) that includes depth-enabled provenance. This means it contains both the standard RGB visual data and the associated depth map captured by the ToF sensor.
  2. Learning-Based Depth Estimation: Using only the RGB content of the digital media, a learning-based monocular depth estimation algorithm is applied. This algorithm processes the 2D image to predict the depth of various objects and regions within the scene, creating a learning-based point cloud that represents the estimated "human perceptual depths."
  3. Ground-Truth Depth Generation: Simultaneously, the raw depth data from the depth-enabled provenance (captured by the ToF sensor) is used to generate a ground-truth point cloud. This represents the actual, measured depths of the scene at the time of original capture.
  4. Regional Segmentation: Both the learning-based and ground-truth point clouds are then subjected to regional segmentation. This process divides the image into distinct regions based on their depth characteristics and visual content.
  5. Regional Comparison: The segmented regions from the learning-based point cloud are compared against their corresponding regions in the ground-truth point cloud. The key here is to identify discrepancies in depth profiles. For instance, if the learning-based algorithm detects a 3D object (like a fire extinguisher in a cabinet), but the ground-truth depth data for that specific region shows a uniformly flat surface, it indicates a strong disagreement.
  6. Highlighting Discrepancies: When significant disagreements are detected between the estimated perceptual depths and the actual provenance depths within a specific region, Scoop highlights these regions. This visual cue alerts the user or an automated system that a recapture attack has likely occurred in that particular area of the content.

The speaker emphasizes that while the presentation provides a high-level conceptual overview, the actual technical implementation involves intricate details that are fully elaborated in their research paper. The system is designed to be robust against variations in content, differentiating between legitimately flat surfaces (like a wall, where both point clouds would agree on flatness) and deceptively flattened surfaces resulting from a recapture (where the learning-based cloud might infer 3D structure, but the ground-truth cloud reveals flatness).

Demo / Proof of Concept

▶ Watch: Introducing depth-enabled provenance as a solution (8:00)

To demonstrate the feasibility and effectiveness of Scoop, the research team developed several fully functional prototypes using commodity devices. These prototypes serve as a proof-of-concept for integrating depth-enabled provenance and the Scoop detection mechanism into real-world applications.

The prototypes include:

  1. Desktop Viewer Application: This application allows users to view visual content that has been captured with depth-enabled provenance. It processes the content, performs the regional comparison between learning-based and ground-truth depths, and visually highlights any detected recapture regions. This provides a clear interface for authentication and anomaly detection.
  2. iOS Camera Application: A dedicated camera application for iOS devices was developed. This app leverages the built-in time-of-flight (ToF) depth sensors present in modern iPhones (e.g., LiDAR scanners) to capture both the RGB visual content and its associated depth map simultaneously. This data is then packaged with the provenance information.
  3. Android Camera Application: Similarly, an Android camera application was built to utilize the depth sensors available on compatible Android smartphones. This ensures cross-platform applicability for capturing the necessary depth-enabled provenance.

These camera applications allow users to capture any visual content, including both photos and videos, with the required depth provenance information embedded. The captured content can then be viewed and analyzed using the desktop viewer or potentially within the camera apps themselves, enabling on-device or remote authentication.

A critical aspect of evaluating Scoop was the absence of a suitable dataset for recapture attack detection using depth information. To address this, the researchers meticulously compiled a first-of-its-kind dataset. This dataset was captured using their two smartphone prototypes (iOS and Android) and comprises a total of 122 unique data points. Crucially, it includes 78 distinct recapture scenarios, utilizing various recapture mediums such as TVs, cardboard cutouts, projectors, and mixed scenarios. The dataset contains both photos and videos, providing a comprehensive resource for evaluating systems like Scoop and setting a new cornerstone for future research in this domain.

Evaluation results from this custom dataset demonstrated Scoop's effectiveness:

  • True Positive Rate: Both the iOS and Android prototypes achieved "relatively high" true positive rates, meaning they correctly identified misleading recaptures with high accuracy.
  • False Positive Rate: Conversely, both prototypes exhibited "relatively low" false positive rates, indicating that they rarely misclassified benign, non-recaptured photos as fraudulent.

The talk also addressed the overhead incurred by Scoop in terms of runtime and storage. While the prototypes showed some overhead, the speaker clarified that these were proof-of-concept implementations not optimized for performance. For instance, computations were performed using a single CPU thread, whereas significant performance gains could be achieved by offloading tasks to multi-CPU threading or GPU computation. For storage, reducing the depth resolution for the depth provenance was shown to have minimal impact on system performance, as detailed in their paper, offering a practical optimization. This highlights that while there are performance considerations, Scoop is a "doable" and practical system.

Defensive Implications

▶ Watch: SCOOP: Leveraging human perceptual depth for detection (10:15)

The introduction of Scoop and the concept of depth-enabled provenance carries significant implications for defenders in the fight against digital misinformation and fraud. It represents a paradigm shift from reactive detection to proactive authentication based on immutable physical properties.

Here's what defenders should do with this information:

  1. Embrace Depth-Enabled Capture: The most critical defensive implication is the urgent need for camera manufacturers and software developers to integrate depth-enabled provenance as a standard feature. This means ensuring that cameras, particularly those on smartphones, utilize their existing or newly integrated time-of-flight (ToF) depth sensors to capture and securely bind depth data with every photo and video. This move would provide the foundational data layer necessary for systems like Scoop to operate.
  2. Advocate for C2PA and Provenance Standards with Depth: Organizations involved in content authenticity initiatives, such as the C2PA (Coalition for Content Provenance and Authenticity), should consider incorporating depth data as an essential component of provenance metadata. Standardizing the capture and embedding of depth information within content provenance will ensure interoperability and widespread adoption of robust anti-recapture mechanisms.
  3. Implement Scoop-like Authentication at Ingestion Points: Platforms that ingest and display user-generated content (e.g., social media, news outlets, insurance claim portals) should implement authentication systems similar to Scoop. Before content is published or accepted as evidence, it should be automatically analyzed to detect signs of recapture. This shifts the burden of detection from the end-user to the platform, enhancing overall trust.
  4. Educate Users on the Limitations of Provenance: While provenance systems are valuable, defenders must proactively educate users about their inherent limitations, particularly regarding recapture attacks. Users should understand that "provenance-authenticated" does not automatically mean "content-verified" and that tools like Scoop are necessary to bridge this gap. This prevents the "false sense of trust" that attackers aim to exploit.
  5. Utilize the Open-Source Prototypes and Dataset: The researchers have made their Scoop prototypes and the first-of-its-kind dataset open source. Defenders, researchers, and developers should leverage these resources to further develop, test, and integrate depth-based recapture detection into their own systems. The dataset, in particular, is invaluable for training and evaluating new models.
  6. Focus on Proactive, Multi-Layered Authentication: Scoop demonstrates that relying solely on post-facto detection or human judgment is insufficient. A multi-layered defensive strategy should include:
  • Secure Capture: Ensuring content is captured with tamper-resistant provenance, including depth data.
  • Automated Validation: Deploying systems like Scoop to validate the integrity of content against known manipulation techniques like recapture.
  • User Education: Empowering users with knowledge about digital media authenticity.
  1. Consider Hardware-Level Security: The reliance on time-of-flight depth sensors suggests that future hardware security modules could play a role in attesting to the integrity of depth data, making it even harder for sophisticated attackers to spoof.

By adopting these defensive strategies, organizations and individuals can significantly bolster their defenses against the growing threat of visually deceptive content, restoring a much-needed layer of trust in digital media.

Key Takeaways

  • Recapture attacks exploit a critical blind spot in provenance-based media authentication, which only verifies origin and history, not content integrity.
  • Existing recapture detection methods are ineffective due to their reliance on prior knowledge of the recapture medium and camera, which attackers can easily manipulate.
  • Human detection of recaptured content is unreliable, with studies showing performance no better than random chance.
  • Scoop introduces depth-enabled provenance, utilizing time-of-flight (ToF) depth sensors to capture ground-truth depth data alongside visual content.
  • Scoop detects recaptures by comparing ground-truth depths with learning-based perceptual depths, highlighting regions where these two conflict, indicating a flat, recaptured surface masquerading as a 3D scene.
  • The project includes open-source prototypes (iOS/Android camera apps, desktop viewer) and a first-of-its-kind dataset for evaluating depth-based recapture detection systems.

About the Speaker(s)

Yuxin (Myles) Liu is a researcher who presented this work, "Scoop: Mitigation of Recapture Attacks on Provenance-Based Media Authentication," a joint effort with colleagues Habiba, Ardelon, and Jin from UC Irvine, and collaborator Sherat from Microsoft. The research reflects expertise in security and digital media forensics, particularly in addressing vulnerabilities within emerging authentication systems.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Scoop attacks a real and underappreciated gap in provenance systems — the fact that C2PA/CAI authenticate the chain of custody, not the content itself — and proposes a physically grounded countermeasure that doesn't play the eternal cat-and-mouse game of artifact detection. The depth-discrepancy approach is elegant: ToF gives you ground-truth geometry; monocular depth estimation gives you what a 3D scene should look like; disagreement exposes the flat screen masquerading as reality. That's a clean, novel framing.

Heather Calloway (CISO) — SOLID

Technically sound research that identifies a real and underappreciated gap in provenance-based authentication systems. The mechanism is clever, the proof-of-concept is functional, and the threat framing around false trust in C2PA/CAI is the most important thing said — but the talk never crosses from research contribution to institutional decision. A CISO walks out understanding the problem better, not knowing what to do Monday morning.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)