RollingEvidence: Autoregressive Video Evidence via Rolling Shutter Effect
Feng Qian
34th USENIX Security Symposium (USENIX Security '25) · Day 1 · System Security 1: Threat Detection, Exploitation, and Adaptive Defenses
Overview
In an era increasingly defined by sophisticated AI-driven manipulations, the integrity of video evidence has become a critical concern. The "RollingEvidence" system, presented by Feng Qian from Ant Group at USENIX Security, offers a novel and robust solution to this burgeoning problem. This talk introduces a system designed to create and verify authentic videos by leveraging a unique property of camera sensors: the rolling shutter effect. By embedding invisible, tamper-resistant probes directly into video frames during recording, RollingEvidence aims to restore transparency and trustworthiness to visual documentation, which is vital for legal, security, and justice applications.
Watch on YouTube · Read the paper · Download the PDF (PDF) · Slides
Paper abstract
Advanced Persistent Threats (APTs) pose critical security challenges to institutions and enterprises through sophisticated, long-duration attack campaigns. While recent APT detection methods primarily leverage provenance graphs constructed from kernel-level audit logs to reveal attack patterns, they face severe scalability limitations in production environments. The provenance graphs grow rapidly (several GB per day) and require long-term maintenance to capture APT campaigns that span months, creating prohibitive storage and computational overhead for real-time detection. To address these challenges, we propose TAPAS, an efficient online APT detection framework that reduces graph dimensionality in both spatial and temporal spaces. For spatial dimensionality, TAPAS focuses on the backbone of the provenance graph, which is often large-scale but sparse. Specifically, TAPAS constructs stacked LSTM-GRU models that iteratively update the representations of the backbone nodes based on relevant redundant nodes, replacing direct storage and computation of these redundancies. For temporal dimensionality, TAPAS designs a task-guided backbone graph segmentation algorithm that identifies active subgraphs as objects to be detected in real-time, reducing structural redundancy in the temporal space. Evaluation in widely used benchmark datasets, DARPA TC and OpTC, demonstrates TAPAS's effectiveness in providing fast, low-overhead online detection while maintaining similar detection accuracy to state-of-the-art methods. Our results show that TAPAS reduces storage requirements by up to 1806× and achieves 99.99% accuracy with an average detection time of 12.78 seconds per GB of audit data, validating its practicality for enterprise deployment with throughputs well above the enterprise requirement of 10^4KB/s.

Key moments
- 0:00 Introduction: Addressing AI-driven video manipulation with RollingEvidence.
- 0:40 Core concept: Using rolling shutter effect to embed probes.
- 2:00 Probe encoding details: FSK, permutations, and pattern standardization.
- 4:00 Verification process: Deep networks for strip extraction and decoding.
- 6:00 Security analysis: Theoretical model and attack vector resistance.
- 7:00 Prototype implementation and experimental setup details.
- 7:40 Experimental results: High accuracy in temper detection and dstripping.
- 9:40 Summary of key contributions and system capabilities.
RollingEvidence: Autoregressive Video Evidence via Rolling Shutter Effect
Speakers: Feng Qian, Ant Group
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=KxTJwKwIYdg
Overview
In an era increasingly defined by sophisticated AI-driven manipulations, the integrity of video evidence has become a critical concern. The "RollingEvidence" system, presented by Feng Qian from Ant Group at USENIX Security, offers a novel and robust solution to this burgeoning problem. This talk introduces a system designed to create and verify authentic videos by leveraging a unique property of camera sensors: the rolling shutter effect. By embedding invisible, tamper-resistant probes directly into video frames during recording, RollingEvidence aims to restore transparency and trustworthiness to visual documentation, which is vital for legal, security, and justice applications.
The core innovation lies in using a modulated LED to generate specific, high-frequency light patterns that are imperceptible to the human eye. When captured by a camera's CMOS sensor, which exposes rows of pixels sequentially, these patterns manifest as unique strip formations within each frame. During verification, a specialized deep neural network extracts these subtle strips, decodes the embedded probes, and employs an exponential mean implication method to detect any temporal tampering. This system directly addresses the threat posed by advanced AI tools like deepfakes and other generative models, ensuring that video evidence remains reliable and untampered.
RollingEvidence distinguishes itself by focusing specifically on the robust detection of altered frames, rather than merely data integrity or transmission speed, as is common in visible light communication (VLC) systems. The prototype, backed by extensive experimentation, demonstrates its utility across various applications, including video notarization, authentication, and forensic analysis. By transforming an inherent camera artifact into a security feature, RollingEvidence provides a powerful new tool in the fight against video manipulation.
Background
▶ Watch: Introduction: Addressing AI-driven video manipulation with RollingEvidence. (0:00)
The pervasive integration of cameras into modern life has made video an indispensable form of evidence across numerous domains, from surveillance and law enforcement to personal documentation. However, the rapid advancement of artificial intelligence, particularly in areas like computer vision and generative adversarial networks (GANs), has given rise to highly sophisticated video manipulation techniques. Tools capable of generating hyper-realistic synthetic media, often termed deepfakes or AI-driven manipulations, can seamlessly alter or create video content, making it nearly impossible for humans to discern authenticity. This capability poses a severe threat to the credibility of video evidence, potentially undermining legal processes, national security, and public trust.
Traditional methods for video authentication often rely on cryptographic hashes or digital watermarks applied after recording, making them vulnerable to tampering if the original source material is compromised. Furthermore, these methods might be visible or require significant computational overhead, impacting usability or video quality. The challenge is to embed integrity checks directly into the video at the point of capture, in a manner that is both invisible to the human eye and resilient to sophisticated digital alteration.
RollingEvidence addresses this by exploiting the rolling shutter effect, a common phenomenon in CMOS sensors. Unlike global shutter sensors that expose all pixels simultaneously, rolling shutter sensors expose and read out pixels row by row. This sequential exposure can cause distortions in images of fast-moving objects or create peculiar visual artifacts when recording under flickering lights. This temporal aliasing, usually considered a drawback, is precisely what RollingEvidence leverages. By carefully modulating an LED light source at frequencies unseen by the human eye, the system intentionally generates unique, invisible strip patterns within video frames due to the asynchronous interaction between the flickering light and the camera's row-by-row exposure mechanism. This transforms a camera's inherent characteristic into a mechanism for embedding hidden, autoregressive proofs, ensuring robust temper detection across video frames. Unlike visible light communication (VLC) systems, which prioritize high data throughput and integrity for communication, RollingEvidence's primary goal is the robust and compact encoding of evidence for temper detection.
Key Findings
▶ Watch: Probe encoding details: FSK, permutations, and pattern standardization. (2:00)
The research behind RollingEvidence yielded several critical findings that underscore its effectiveness and potential as a robust video authentication system:
Firstly, the system successfully demonstrates the feasibility of embedding invisible, tamper-resistant probes directly into video frames by exploiting the camera's rolling shutter effect. These probes, generated through modulated LED frequencies, manifest as unique strip patterns imperceptible to human viewers but detectable by specialized algorithms. This innovative approach provides an intrinsic layer of security at the point of capture, making it fundamentally more robust than post-processing authentication methods.
Secondly, RollingEvidence achieves highly accurate and reliable temper detection. Through its multi-task deep neural network, the system can extract subtle strip patterns, decode complex probes, and identify manipulated frames with high precision. Rigorous experiments, including scenarios involving insertion, removal, alteration, face swapping, and lip-syncing manipulations, showed that the system accurately identified manipulations in almost all cases without misclassifying untempered videos. This performance is crucial for real-world applications where both false positives and false negatives carry significant consequences.
Thirdly, the research established specific parameters for optimal system performance. An SNR (Signal-to-Noise Ratio) above -18.19 decibels was found to achieve a zero false acceptance rate (FAR) and less than a 0.6% false rejection rate (FRR). Furthermore, a window size of 30 to 150 frames and a 4,096-dimension probe optimized detection, maintaining a 0% FAR. Smaller windows or dimensions below 512 significantly increased error rates, highlighting the importance of these configurations. With a support threshold of 0.5, RollingEvidence demonstrated a 0% FRR and less than a 0.5% FAR for 15 to 120 windows, achieving 99.5% accuracy in detecting tempered videos. These metrics provide clear guidance for implementing and configuring the system for maximum reliability.
Finally, the security analysis confirmed strong resistance against sophisticated attackers. The system's cryptographic design ensures that forging valid strip patterns or probes requires computational effort akin to SHA 256 pre-image attacks, offering 128-bit resistance for certain attack vectors (A-C) and 256-bit security for others (D-E). The mitigation strategies for replay attacks, involving server-side timestamping, further bolster the system's resilience against coordinated attacks where an attacker might compromise both the camera and LED.
Technical Deep Dive
▶ Watch: Security analysis: Theoretical model and attack vector resistance. (6:00)
The technical foundation of RollingEvidence rests on the ingenious exploitation of the rolling shutter effect inherent in CMOS sensors. Unlike global shutter cameras, CMOS sensors capture an image by sequentially exposing and reading out rows of pixels. When a camera with a rolling shutter captures light from a rapidly flickering source, such as a modulated LED, the asynchronous exposure of individual pixel rows creates distinct, fixed-frequency strip patterns across the video frame. These patterns are imperceptible to the human eye because the LED's modulation frequency is carefully chosen to be outside the visible spectrum or occurs too rapidly for human perception.
To embed information, RollingEvidence employs Frequency Shift Keying (FSK). This modulation technique encodes data by shifting the carrier frequency among a set of discrete frequencies. The system utilizes 13 distinct frequencies to create 4,096 unique probe permutations, derived from combinations of one to four frequencies. These probes are intelligently surrounded by a splitter frequency, which enhances detection accuracy by providing clear boundaries between probe segments. This design ensures that at least one complete pattern fits within a frame, mitigating concerns about packet loss or idle gaps – the mandatory delays between consecutive frame exposures in CMOS cameras.
Standardizing strip patterns across different cameras is crucial for robust verification. This is achieved by fixing the camera's exposure time and carefully calibrating the LED's modulation frequency. Analysis showed that strip intensities scale with exposure time, necessitating frequencies below half the exposure time for high contrast. Specifically, a splitter frequency of 34 pixels with enhanced contrast (exposure to one over three exposure time) ensures accurate probe extraction. The system accounts for the bounds of strip pixel intensity, from a hyper-intensity limit where the sensor's exposure window aligns perfectly with the LED's active phase (maximum photon accumulation) to a lower bound representing minimal signal during LED on/off intervals.
The system employs stochastic probe encoding, a sophisticated mechanism that links each video frame to prior frames and the device's unique cryptographic keys. This creates an autoregressive chain of evidence, where the integrity of any frame depends on the integrity of its predecessors. Video segments are divided into overlapping windows, each associated with a random sequence derived from the device's cryptographic keys. An exponential mean sampling method, coupled with an auxiliary lambda system, is used to select probes that maximize randomness and resistance to temporal manipulation. This prioritization of larger random values is critical for effective temper detection.
Verification is performed by a multi-task deep neural network designed for three primary functions: strip extraction, probe decoding, and temper detection. From three consecutive frames, the network extracts strip intensity curves. A specialized row attention module enhances focus on brighter rows, improving the precision of strip extraction and the clarity of the resulting strip-free videos generated for clear viewing. The rapid LED switching ensures that these strips do not cause persistent frame occlusion.
Once intensity curves are extracted, they are segmented by the splitter patterns, and a pre-trained neural network decodes the embedded probes. The decoded probes are then compared against the expected random sequence reconstructed from the cryptographic keys and the autoregressive chain. A quantile-based test, specifically using the 98th percentile, is applied to flag misaligned windows as tempered. Theoretical analysis confirms that tempered frames exhibit statistically significant anomalies compared to random sampling, validating the system's detection reliability.
The security model assumes secure key distribution and considers sophisticated attackers capable of forging strips via deep learning. Six attack vectors (A to F) are evaluated, covering scenarios where either the camera or LED (but not both) is compromised. Attack vectors A to C offer 128-bit resistance, comparable to SHA 256 pre-image attacks, while D to E provide 256-bit security. Attack vector F, which involves LED coordination to forge strips, is mitigated by server-side timestamping, preventing replay attacks and ensuring the temporal integrity of the embedded evidence.
Demo / Proof of Concept
▶ Watch: Prototype implementation and experimental setup details. (7:00)
The RollingEvidence team developed a practical prototype to validate their system's utility and performance. The setup consisted of a standard smartphone for video recording and encoding, paired with an Arduino Uno microcontroller responsible for modulating the LED light source using 16 FSK modulation. This configuration demonstrated the system's ability to operate with readily available hardware components.
To rigorously test performance, two identical smartphones were utilized in the experiments. One smartphone was configured with a fixed 1/4000 shutter speed to ensure optimal strip capture, while the other was set to auto-exposure mode to generate "strip-free ground truth" videos for comparison. A simple white sheet was used as a background to enhance strip visibility, aiding in the accurate creation of the experimental dataset for training and evaluation.
Two primary experiments were conducted to assess the system's temper detection capabilities:
- Insertion, Removal, and Alteration: This experiment focused on detecting fundamental video manipulations, such as adding or deleting frames, or altering content within existing frames.
- Face Swapping and Lip-syncing: This evaluated the system's resilience against advanced AI-driven manipulations, specifically deepfakes involving facial alterations and synchronized speech.
The results from both experiments were highly encouraging. RollingEvidence accurately identified manipulations in almost all tested scenarios. Critically, the system demonstrated robust performance by consistently not misclassifying untempered videos, achieving a zero false acceptance rate (FAR) under optimal conditions. This high precision and recall are essential for a security system where false positives can be as detrimental as false negatives.
Beyond temper detection, the researchers also evaluated the quality of strip extraction and the effectiveness of the "d-stripping" process (removing the embedded patterns to produce a clean video). This was assessed across 13 indoor and 3 outdoor scenarios to account for varying lighting conditions. For each scenario, a captured frame (left image) was compared against the "d-stripped" frame generated by the deep network (middle image) and the ground truth (right image). Using Mean Square Error (MSE) for strip extraction accuracy and the Structural Similarity Index (SSIM) for d-stripping quality, the deep network achieved high accuracy. It produced strip-free videos that closely matched the ground truth, even under diverse and challenging lighting, confirming that the embedded evidence does not permanently degrade video quality for human viewing.
Further assessment explored the impact of physical and hyperparameter choices. An SNR above -18.19 decibels was identified as the threshold for achieving a 0% FAR and less than a 0.6% false rejection rate (FRR). Optimal detection was maintained with a window size of 30 to 150 frames and a 4,096-dimension probe, both achieving a 0% FAR. Conversely, smaller windows and dimensions below 512 significantly increased error rates. When considering the number of windows, increasing the support threshold to 0.5 allowed RollingEvidence to achieve a 0% FRR and less than a 0.5% FAR for 15 to 120 windows, culminating in 99.5% accuracy in detecting tempered videos. These metrics underscore the system's practical viability and the importance of proper configuration for deployment.
Defensive Implications
▶ Watch: Summary of key contributions and system capabilities. (9:40)
RollingEvidence presents significant defensive implications for organizations and individuals grappling with the escalating threat of AI-driven video manipulation. The system offers a proactive and intrinsic method for ensuring video authenticity, moving beyond reactive detection of deepfakes to embedding verifiable proofs at the source.
For law enforcement, legal institutions, and security agencies, RollingEvidence could revolutionize the handling of video evidence. By integrating this technology into body cameras, surveillance systems, or dashcams, every recorded video could carry an embedded, cryptographically verifiable audit trail. This would dramatically increase the trustworthiness of digital evidence in court, reduce the potential for disputes over video authenticity, and streamline forensic investigations. The ability to quickly and accurately identify tempered frames, even those altered by sophisticated AI, would save countless hours and resources currently spent on manual verification or inconclusive analysis.
Enterprises and media organizations can leverage RollingEvidence to protect their brand integrity and content authenticity. In an age of misinformation, verifying the originality of corporate communications, news footage, or marketing materials is paramount. Cameras equipped with RollingEvidence could ensure that official video releases are beyond doubt, building greater trust with stakeholders and audiences.
Critical infrastructure and industrial control systems that rely on video monitoring for security can also benefit. Authenticated video feeds could prevent attackers from injecting fabricated footage to mask their activities or create false alarms, providing a more reliable basis for operational decisions.
The implementation of RollingEvidence, however, carries its own set of considerations for defenders:
- Secure Key Management: The system relies on device-specific cryptographic keys. Robust key generation, storage, and distribution mechanisms are essential to prevent attackers from compromising the root of trust. Hardware Security Modules (HSMs) or Trusted Platform Modules (TPMs) could play a vital role here.
- Integration with Existing Systems: For widespread adoption, RollingEvidence would need to be integrated into existing camera hardware and software ecosystems. This could involve partnerships with camera manufacturers or the development of standardized APIs.
- Awareness and Training: Defenders would need to be trained on the capabilities and limitations of the system, understanding how to interpret verification results and what constitutes a legitimate "tamper" flag.
- Scalability: Managing and verifying potentially vast amounts of authenticated video data would require scalable backend infrastructure, including distributed ledger technologies for notarization or centralized, secure verification services.
- Adversarial Robustness: While the system demonstrates strong security, the arms race with AI manipulators is ongoing. Continuous research and development will be necessary to anticipate and counter future adversarial techniques.
Ultimately, RollingEvidence provides a powerful new paradigm for digital forensics and video security. By making integrity an inherent property of video capture, it significantly raises the bar for attackers and provides defenders with a confident, analytical tool to combat the pervasive threat of AI-driven video manipulation.
Key Takeaways
- Invisible, Tamper-Resistant Evidence: RollingEvidence utilizes the camera's rolling shutter effect and modulated LEDs to embed invisible, autoregressive strip patterns into video frames, serving as intrinsic, tamper-resistant proofs.
- Robust AI-Resistant Detection: A multi-task deep neural network extracts these patterns, decodes cryptographic probes, and accurately detects sophisticated video manipulations (e.g., deepfakes, insertions) with high precision, achieving 99.5% accuracy in tempered video detection.
- Cryptographically Secured Autoregressive Encoding: Probes are stochastically encoded, linking frames to prior ones and device cryptographic keys, providing 128-bit to 256-bit security against forgery, akin to SHA 256 pre-image attacks.
- Optimized Performance Parameters: The system's reliability is contingent on specific configurations, including an SNR above -18.19 dB (for 0% FAR), a window size of 30-150 frames, and 4,096-dimension probes.
- Practical Prototype with High Fidelity: A smartphone-based prototype demonstrated effective strip extraction and "d-stripping," producing strip-free videos closely matching ground truth, ensuring usability without compromising visual quality.
- Critical for Video Forensics and Authentication: RollingEvidence offers a crucial defensive mechanism against AI-driven video manipulation, bolstering the trustworthiness of video evidence for legal, security, and enterprise applications.
About the Speaker(s)
The talk "RollingEvidence: Autoregressive Video Evidence via Rolling Shutter Effect" was presented by Feng Qian. At the time of the presentation, Feng Qian was affiliated with Ant Group, a major fintech company. His research focuses on developing innovative security solutions, particularly in areas susceptible to advanced digital manipulation, aiming to enhance transparency and trustworthiness in digital media.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Genuinely clever repurposing of a hardware artifact — the rolling shutter effect — as a cryptographically anchored authentication primitive. The core idea is elegant: something camera engineers have treated as a bug for 20 years becomes an unforgeable evidence channel. Solid engineering, honest evaluation metrics, and a real threat model make this stand out from the usual deepfake-detection noise.
Heather Calloway (CISO) — WEAK
Technically credible research on an interesting camera-physics-based authentication mechanism, but this talk never gets off the lab bench. The gap between a smartphone-and-Arduino prototype and institutional deployment is enormous, and the presentation does nothing to close it.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)