TWINFUZZ: Differential Testing of Video Hardware Acceleration Stacks

Matteo Leonelli (CISPA)

Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Fuzzing 2

Overview

In the realm of modern computing, hardware acceleration stacks are ubiquitous, powering everything from high-performance graphics to efficient video playback. However, the intricate, multi-layered nature of these stacks, often comprising proprietary drivers, kernel modules, and specialized GPU hardware, presents a formidable challenge for security testing. Traditional software fuzzing techniques, which rely heavily on code instrumentation and coverage feedback, are largely ineffective against these opaque, blackbox components. This talk, presented by Matteo Leonelli from CISPA, introduces TWINFUZZ, a novel differential testing methodology designed to uncover functional and security-relevant bugs in video hardware acceleration stacks.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction and challenges of hardware fuzzing
  2. 1:10 Twinfuzz's differential testing approach for blackbox hardware
  3. 2:00 Applying differential testing to video decoding stacks
  4. 2:40 Three main challenges addressed by Twinfuzz
  5. 3:50 Identifying observable differences and faulty hardware decoding
  6. 4:50 Guiding blackbox fuzzing using software model coverage
  7. 6:10 Root cause analysis and clustering of detected disparities

TWINFUZZ: Differential Testing of Video Hardware Acceleration Stacks

Speakers: Matteo Leonelli, CISPA

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=8CuEC8LTr7k

Overview

In the realm of modern computing, hardware acceleration stacks are ubiquitous, powering everything from high-performance graphics to efficient video playback. However, the intricate, multi-layered nature of these stacks, often comprising proprietary drivers, kernel modules, and specialized GPU hardware, presents a formidable challenge for security testing. Traditional software fuzzing techniques, which rely heavily on code instrumentation and coverage feedback, are largely ineffective against these opaque, blackbox components. This talk, presented by Matteo Leonelli from CISPA, introduces TWINFUZZ, a novel differential testing methodology designed to uncover functional and security-relevant bugs in video hardware acceleration stacks.

TWINFUZZ addresses the inherent limitations of post-silicon hardware fuzzing by leveraging a differential oracle and an innovative indirectly proxy coverage mechanism. By simultaneously feeding the same input to a hardware-accelerated stack and a software reference model, the system can detect subtle yet critical discrepancies in output, indicating potential vulnerabilities. Leonelli's work is significant because it provides a scalable and effective approach to securing a crucial, yet traditionally under-tested, attack surface. The findings, which include information leaks and memory safety bugs in real-world applications and drivers, underscore the critical need for such advanced testing methodologies in an increasingly hardware-dependent computing landscape.

The research presented by Leonelli not only offers a powerful new tool for identifying flaws in complex hardware-software interactions but also highlights the broader implications of these vulnerabilities. Bugs in video acceleration can lead to system instability, denial of service, or even remote code execution, making the robust testing of these components paramount for overall system security. TWINFUZZ provides a much-needed framework for security researchers and vendors to systematically scrutinize these critical components, ultimately enhancing the reliability and trustworthiness of our digital infrastructure.

Background

▶ Watch: Introduction and challenges of hardware fuzzing (0:00)

The landscape of software security has long relied on fuzzing as a cornerstone for vulnerability discovery. In a typical software fuzzing scenario, a fuzzer feeds random or intelligently mutated inputs to an instrumented binary. This instrumentation, usually applied at compile time, allows for the collection of code coverage feedback, which in turn guides the fuzzer to explore new execution paths and generate more "meaningful" inputs. Furthermore, sanitizers (like AddressSanitizer or UndefinedBehaviorSanitizer) and oracles are employed to detect anomalies, crashes, or undefined behavior during execution, providing clear indicators of potential bugs.

However, this effective paradigm breaks down when confronted with post-silicon hardware fuzzing. Here, the target is often a blackbox component, such as a GPU or a dedicated hardware accelerator. Instrumentation is typically impossible, denying the fuzzer any coverage feedback. Without sanitizers or a clear oracle, identifying crashes or functional errors becomes exceedingly difficult, often devolving into mere random regression testing with limited efficacy. The fuzzer serves random inputs, and without any internal visibility, it struggles to discern if the hardware is behaving correctly or if a fault has occurred.

This is precisely the problem TWINFUZZ aims to solve by applying differential testing to video hardware acceleration stacks. The core idea of differential testing is to test two different implementations of the same functionality with identical inputs. In this context, one implementation serves as a reference model (assumed to be more likely correct), while the other is the component under test. For video acceleration, TWINFUZZ pits a comprehensive video acceleration hardware stack against a software model reference. The hardware stack involves a multi-layered interaction between userland APIs, drivers, the kernel, and the GPU, where the actual video decoding task is offloaded. The software model, on the other hand, is purely userland-based, such as a software-only video decoder (e.g., FFmpeg).

The challenges TWINFUZZ specifically addresses are threefold:

  1. Deriving a new oracle for hardware acceleration stacks: How can we detect incorrect behavior in a blackbox hardware component?
  2. Crafting a feedback mechanism for blackbox fuzzing: How can a fuzzer be guided to explore the state space of a hardware component without direct instrumentation?
  3. Analyzing disparities for root cause analysis: Once differences are observed, how can we pinpoint the underlying issue in a blackbox scenario? By tackling these challenges, TWINFUZZ offers a pragmatic solution to a long-standing problem in hardware security.

Key Findings

▶ Watch: Applying differential testing to video decoding stacks (2:00)

TWINFUZZ proved highly effective in uncovering a range of vulnerabilities and behavioral discrepancies across various video hardware acceleration stacks. The research primarily focused on testing NVIDIA and Intel platforms, providing a robust cross-vendor analysis. A significant achievement of TWINFUZZ is its ability to identify not only functional bugs—where the hardware produces an incorrect output—but also security-relevant bugs, including information leaks and memory safety issues, which are often difficult to detect in blackbox environments.

One of the most compelling findings was the discovery of an information leak in the Firefox browser. This bug manifested when specific inputs, previously identified by TWINFUZZ as leading to observable differences in hardware-decoded frames, were replayed within a hardware-accelerated Firefox instance. The implication here is that malformed video streams, processed by the vulnerable hardware acceleration path, could potentially expose sensitive data or internal state. Another critical bug was found in the interaction between VLC media player and Windows video drivers during the decoding process, highlighting issues in the complex interplay between application-level software and underlying system drivers.

Beyond these specific functional and information leakage vulnerabilities, TWINFUZZ, by operating at a higher level where sanitizers are available, also successfully identified traditional memory safety bugs. These included buffer overflows and global buffer overflows within the application and driver layers. While the primary focus of differential testing is on functional discrepancies, the ability to layer traditional sanitizers on top of the software reference model allowed for a more comprehensive bug hunting approach, catching vulnerabilities that could lead to crashes or arbitrary code execution.

The analysis of the observed differences led to the creation of 15 distinct clusters of bugs. Some clusters were unique to a single platform, indicating specific flaws in a particular vendor's implementation or driver. Others were replicable across different platforms, suggesting potential issues in common components, shared specifications, or widely used third-party libraries. This clustering provides valuable insights for vendors, allowing them to prioritize fixes for vulnerabilities that affect a broader range of systems.

Matteo Leonelli emphasized that TWINFUZZ utilizes an unspecialized and unmodified fuzzer, making the approach highly scalable. While the paper specifically focused on video accelerators, the underlying methodology is generalizable to other types of hardware accelerators. This scalability, combined with its proven effectiveness in identifying critical bugs in post-silicon hardware acceleration stacks, positions TWINFUZZ as a significant advancement in hardware security testing. The work demonstrates that even without direct hardware instrumentation, it is possible to achieve meaningful coverage and uncover deep-seated vulnerabilities in these opaque components.

Technical Deep Dive

▶ Watch: Three main challenges addressed by Twinfuzz (2:40)

The technical ingenuity of TWINFUZZ lies in its elegant solution to the blackbox problem through differential testing and novel coverage guidance. The core of the methodology revolves around a differential oracle designed specifically for hardware acceleration stacks. The setup involves a fuzzer that generates a single input, which is then simultaneously fed to two distinct models:

  1. The Whitebox Instrumented Binary (Software Model): This is typically a well-understood, instrumented software video decoder (e.g., a modified FFmpeg build). Its "whitebox" nature allows for the collection of code coverage, providing crucial feedback that guides the fuzzer's exploration. This model serves as the reference, assumed to be largely correct and robust.
  2. The Blackbox Hardware Acceleration Stack: This represents the actual hardware-accelerated video decoding pipeline, encompassing userland APIs, drivers, kernel components, and the physical GPU. It is treated as a true blackbox, with no internal instrumentation or visibility.

Both models receive the identical input, process it, and produce an output frame. The differential oracle then performs a pixel-by-pixel or perceptual comparison of these two output frames. If the frames are identical, both models are inferred to have behaved correctly and consistently. However, if there is an observable difference—even a slight one—it signals a potential discrepancy, indicating that one (or possibly both) of the implementations might be faulty. The talk presented a compelling visual example: a hardware-decoded frame showing a significant disparity in the bottom part, where a section with people was rendered incorrectly, appearing notably different from the software-generated output. When such a hardware-generated frame was rendered in a hardware-accelerated browser, the incorrect decoding was starkly evident, with visible artifacts (e.g., green blocks) highlighting the faulty regions.

To address the challenge of guiding a fuzzer without direct hardware coverage, TWINFUZZ introduces the concept of Indirectly Proxy Coverage. The underlying assumption here is that video decoding, being a process governed by strict specifications, is largely deterministic. While specific hardware initialization routines and low-level hardware-specific decoding implementations will naturally differ from their software counterparts, many higher-level operations are expected to be functionally equivalent. These include critical stages like parsing the video stream, filtering operations, and post-processing of decoded frames. These operations are often implemented in a one-to-one mappable fashion between the software and hardware stacks.

The genius of indirectly proxy coverage is to leverage the software coverage obtained from the whitebox reference model to indirectly guide the fuzzer's exploration of the blackbox hardware stack. By generating inputs that maximize coverage in the software model's parsing, filtering, and post-processing components, the fuzzer implicitly ensures that these analogous (and often shared) pathways are also exercised in the hardware stack. This provides a crucial feedback loop, allowing the fuzzer to generate more "interesting" inputs that are likely to trigger deeper or more complex states within the hardware, even without direct visibility.

For root cause analysis in this blackbox scenario, TWINFUZZ employs a systematic replay and clustering approach. All test cases that generate observable differences are collected. These inputs are then replayed across different platforms (e.g., NVIDIA and Intel setups). This replay strategy allows researchers to:

  1. Identify platform-specific bugs: If an input consistently triggers a disparity on one platform but not others, it strongly suggests a bug unique to that specific hardware or driver implementation. The talk identified 15 such clusters, some of which were unique to a single platform.
  2. Identify common component bugs: If an input triggers observable differences across multiple platforms, it implies a more fundamental issue, potentially residing in a shared component, a common specification interpretation, or widely used third-party code that is part of multiple vendors' stacks.

While this approach significantly aids in narrowing down the potential sources of error, the speaker acknowledged that limited root cause analysis remains an open problem, particularly for functional bugs where traditional sanitizers are not directly applicable to the hardware itself. Even vendors often struggle with precisely identifying the root cause of these subtle functional disparities. Nevertheless, the ability to cluster and reproduce these issues across platforms is a significant step forward in understanding and addressing vulnerabilities in these complex systems.

Demo / Proof of Concept

▶ Watch: Guiding blackbox fuzzing using software model coverage (4:50)

While the talk did not feature a live, interactive demonstration of TWINFUZZ in action, the speaker presented compelling visual evidence and a clear conceptual walkthrough of how the system operates and the types of issues it uncovers. The proof of concept was primarily articulated through:

  1. Visual Disparities in Decoded Frames: Leonelli showed side-by-side comparisons of frames decoded by the software reference model and the hardware acceleration stack. These images clearly highlighted observable differences, such as a section of a frame appearing completely corrupted or displaying incorrect color information in the hardware-decoded output, while the software version remained pristine. An example specifically mentioned involved a disparity in the bottom part of a frame, where people were depicted differently in the hardware-accelerated output.
  2. Real-world Impact on Applications: To underscore the practical implications of these disparities, the speaker demonstrated how these faulty hardware-decoded frames manifest when rendered in common applications. For instance, an incorrect hardware-decoded frame, when displayed in a hardware-accelerated Firefox browser, showed clear visual artifacts (e.g., large green blocks or misrendered textures) that directly indicated the underlying decoding error. This visual evidence served as a powerful proof point that the observed differences were not minor rendering variations but rather significant functional bugs that impacted user experience and potentially security.
  3. Bug Findings and Platforms: The presentation effectively served as a proof of concept by detailing the actual bugs found on specific platforms. The mention of an information leak in Firefox and a bug impacting VLC media player interacting with Windows video drivers concretely demonstrated TWINFUZZ's capability to identify real-world, actionable vulnerabilities. The identification of traditional memory safety bugs (like buffer overflows) further validated the system's comprehensive bug-finding potential, even if these were detected by layering sanitizers on the software side.

The strength of this proof of concept lay in its ability to translate abstract technical concepts (differential testing, indirect coverage) into tangible outcomes: visually verifiable errors and documented security vulnerabilities in widely used software and hardware components. This provided a strong foundation for the claims regarding TWINFUZZ's effectiveness and its potential to improve the security posture of video hardware acceleration stacks.

Defensive Implications

▶ Watch: Root cause analysis and clustering of detected disparities (6:10)

The findings and methodology presented by TWINFUZZ have profound implications for various stakeholders involved in the development, deployment, and security of computing systems. Addressing the vulnerabilities identified by TWINFUZZ requires a multi-faceted defensive strategy.

For Hardware and Software Vendors (GPU Manufacturers, Driver Developers, OS Vendors):

  • Enhanced Testing Methodologies: TWINFUZZ provides a blueprint for a robust, scalable testing methodology that can be integrated into the development lifecycle. Vendors should adopt differential fuzzing as a standard practice for their hardware acceleration stacks, moving beyond traditional, often insufficient, internal testing. The ability to use an unmodified fuzzer makes this approach highly adaptable.
  • Prioritized Bug Fixing: The clustering of bugs (e.g., 15 distinct clusters) by platform and replicability offers a clear path for prioritization. Bugs that affect multiple platforms or common components should receive immediate attention due to their broader impact.
  • Improved Root Cause Analysis Tools: While TWINFUZZ aids in identifying where differences occur, precise root cause analysis remains challenging for blackbox hardware. Vendors should invest in better internal debugging tools and diagnostic capabilities for their hardware and drivers to complement differential testing.
  • Secure Coding Practices: The discovery of memory safety bugs (buffer overflows) underscores the persistent need for rigorous secure coding practices in driver and API development, even for performance-critical components.
  • Transparency and Specification Adherence: Discrepancies often arise from ambiguous specifications or divergent interpretations. Fostering greater clarity in specifications and stricter adherence across implementations could reduce the attack surface.
  • Continuous Monitoring: The dynamic nature of software and hardware updates necessitates continuous testing. TWINFUZZ's approach can be integrated into CI/CD pipelines to catch regressions or new vulnerabilities introduced by updates.

For Application Developers (Browser Developers, Media Player Developers):

  • Awareness of Hardware Acceleration Risks: Developers relying on hardware acceleration must be acutely aware that these underlying components can introduce vulnerabilities, such as information leaks. Robust error handling and input validation in applications should account for potential malformed outputs from hardware decoders.
  • Sandboxing and Isolation: Where possible, hardware acceleration components should be run in highly sandboxed environments to limit the impact of any exploits.
  • Fallback Mechanisms: Applications should have robust fallback mechanisms to gracefully handle errors or unexpected behavior from hardware accelerators, potentially switching to software decoding if an issue is detected. The information leak in Firefox highlights the critical need to secure these pathways.

For System Administrators and End-Users:

  • Prompt Updates: The most immediate defensive action is to apply system updates, driver updates, and application patches as soon as they become available. These updates often contain fixes for vulnerabilities discovered through methods like TWINFUZZ.
  • Security-Conscious Configuration: Users or administrators might consider the trade-offs between performance and security. In high-security environments, disabling hardware acceleration for certain sensitive applications could be a temporary mitigation if vulnerabilities are known but unpatched.
  • Understanding the Attack Surface: The talk highlights that even seemingly innocuous components like video decoders can be critical attack vectors. A holistic approach to security must consider the entire stack, from hardware to user applications.

In essence, TWINFUZZ serves as a stark reminder that the security perimeter extends deep into the hardware-software interface. Proactive and intelligent testing, coupled with a commitment to addressing discovered flaws, is indispensable for building resilient and trustworthy computing systems.

Key Takeaways

  • TWINFUZZ introduces a novel methodology for effectively testing video hardware acceleration stacks, addressing the inherent challenges of blackbox components.
  • The research proposes a differential oracle that leverages simultaneous execution on hardware and software models to identify both functional and security-relevant bugs by comparing their outputs.
  • An innovative indirectly proxy coverage technique guides the fuzzer in blackbox hardware environments by utilizing code coverage feedback from the instrumented software reference model for analogous operations like parsing and post-processing.
  • The TWINFUZZ prototype has been successfully implemented and proven effective for post-silicon video acceleration stacks, demonstrating its capability to uncover deep-seated vulnerabilities.
  • Real-world security implications were identified, including an information leak in the Firefox browser and a bug in the interaction between VLC media player and Windows video drivers, alongside traditional memory safety bugs like buffer overflows.
  • The methodology is scalable and adaptable to different types of hardware accelerators, suggesting its potential for broader application beyond video decoding.

About the Speaker(s)

The talk was presented by Matteo Leonelli, a researcher from CISPA – Helmholtz Center for Information Security. His presentation focused on his "latest work," TWINFUZZ, indicating his primary research interests lie in hardware security, fuzzing, and vulnerability discovery in complex systems, particularly those involving hardware-software interaction and acceleration stacks. His work at CISPA underscores the institution's commitment to advancing the state of the art in cybersecurity research.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid, genuinely novel research that tackles a real and under-served problem — fuzzing blackbox hardware acceleration stacks — with a clever differential oracle and an indirect coverage proxy that sidesteps the instrumentation wall. The Firefox information leak and VLC/Windows driver bugs are real, reproducible findings, not demo-ware. Doesn't quite hit 5 because root cause analysis remains shallow by the speaker's own admission, and the talk leans heavily on the methodology without fully stress-testing the assumptions baked into the proxy coverage heuristic.

Heather Calloway (CISO) — WEAK

Technically legitimate research on a real and under-tested attack surface — differential fuzzing of hardware acceleration stacks is genuinely novel and the Firefox information leak is a concrete finding. But this talk lives entirely inside the lab. It does not reach the people accountable for the risk it exposes.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025