Moneta: Ex-Vivo GPU Driver Fuzzing by Recalling In-Vivo Execution States

Joonkyo Jung (PhD student · Jon University)

Network and Distributed System Security (NDSS) Symposium 2025 · Day 2 · Fuzzing 1

Overview

The talk "Moneta: Ex-Vivo GPU Driver Fuzzing by Recalling In-Vivo Execution States," presented by Joonkyo Jung of Jon University, introduces a novel fuzzing framework designed to uncover vulnerabilities in Graphics Processing Unit (GPU) drivers. Moneta tackles long-standing challenges in device driver fuzzing, particularly the complex inter-syscall and hardware dependencies inherent to GPU interactions. The core innovation lies in its hybrid approach, combining the benefits of snapshotting with deterministic record-and-replay mechanisms to efficiently explore deep states within GPU drivers.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction: GPU driver vulnerabilities and fuzzing challenges.
  2. 3:00 Moneta's core idea: snapshots combined with record and replay.
  3. 4:15 Fuzzer implantation with Moneta agent preserving driver context.
  4. 5:30 Preserving device file descriptors during rehosting for stateful fuzzing.
  5. 6:15 Enabling integrated ARM GPU fuzzing by handling CPU features.
  6. 7:45 Evaluation: Basic block coverage results and baseline comparisons.

Moneta: Ex-Vivo GPU Driver Fuzzing by Recalling In-Vivo Execution States

Speakers: Joonkyo Jung, PhD student, Jon University

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=Wvj8HGbEvKA

Overview

The talk "Moneta: Ex-Vivo GPU Driver Fuzzing by Recalling In-Vivo Execution States," presented by Joonkyo Jung of Jon University, introduces a novel fuzzing framework designed to uncover vulnerabilities in Graphics Processing Unit (GPU) drivers. Moneta tackles long-standing challenges in device driver fuzzing, particularly the complex inter-syscall and hardware dependencies inherent to GPU interactions. The core innovation lies in its hybrid approach, combining the benefits of snapshotting with deterministic record-and-replay mechanisms to efficiently explore deep states within GPU drivers.

GPU drivers are critical components in modern computing, especially with the surge of AI workloads, yet they remain a significant attack surface. In 2023 alone, Nvidia addressed 27 high-severity CVEs in its GPU drivers, highlighting the urgent need for more effective vulnerability discovery methods. Traditional fuzzing techniques often struggle with the unique complexities of GPU drivers, either requiring expensive and unscalable physical hardware (in-vivo approaches) or suffering from fidelity and non-determinism issues in purely emulated (ex-vivo) environments. Moneta aims to resolve these limitations by providing a practical, scalable, and effective ex-vivo fuzzing solution that can deterministically reach and fuzz deep, complex driver states.

This research is significant because it enhances the capabilities of ex-vivo fuzzing, making it possible to find bugs that are typically only accessible through in-vivo methods but without the associated scaling costs. By presenting a system that can effectively fuzz a wide range of mainstream GPUs, including both discrete and integrated architectures, Moneta offers a powerful tool for security researchers and GPU vendors alike. Its success in identifying 10 previously unknown bugs, five of which received CVE assignments, underscores its practical impact and the critical need for continued research in this often-overlooked area of system security.

Background

▶ Watch: Introduction: GPU driver vulnerabilities and fuzzing challenges. (0:00)

The security of GPU drivers has become paramount, given their integral role in modern computing, particularly for high-performance tasks and artificial intelligence. Despite their importance, GPU drivers are notoriously complex and, consequently, prone to vulnerabilities. The speaker highlighted this by citing Nvidia's 27 high-severity GPU driver CVEs assigned in the past year, underscoring the critical need for robust security testing.

Fuzzing is a widely recognized and effective method for discovering software vulnerabilities. However, applying fuzzing to device drivers, especially GPU drivers, presents two primary and well-known challenges:

  1. Inter-syscall Dependency Challenge: Syscalls often have intricate value and ordering dependencies. A sequence of syscalls must be executed in a specific order with particular arguments to reach meaningful states or trigger certain code paths within the driver. Randomly generated syscall sequences are unlikely to satisfy these dependencies, leading to inefficient fuzzing.
  2. Hardware Dependency Challenge: Device drivers interact directly with hardware through I/O operations. A fuzzer must somehow handle these I/O requests. Without the actual hardware or a faithful emulation, the driver might block, crash, or enter undefined states, preventing effective fuzzing.

Prior works have attempted to address these challenges through various approaches:

  • In-vivo Approaches: These methods involve the presence of actual hardware during fuzzing. For example, a fuzzer might run on a virtual machine, with I/O requests handled by the physical GPU. While in-vivo fuzzing inherently addresses the hardware dependency challenge by using the real device, it suffers from significant scalability issues. Each fuzzing instance typically requires a dedicated physical device, making it impractical for large-scale, continuous fuzzing campaigns.
  • Ex-vivo Approaches: To overcome the scalability limitations of in-vivo methods, researchers have explored ex-vivo approaches, where the actual device is absent during fuzzing. These include:
  • Hardware Access Evasion: This involves stubbing out or mocking I/O functions to prevent the driver from attempting to communicate with non-existent hardware. While simpler, this often leads to incomplete code coverage as many critical paths involving hardware interaction are skipped.
  • Static Analysis Methods: Techniques like symbolic I/O attempt to statically analyze driver code to generate inputs that do not require hardware interaction. These methods can be effective for specific code patterns but often struggle with the dynamic and complex nature of real-world drivers.
  • Record and Replay: This approach involves recording actual device interactions and syscall sequences from a real workload and then replaying them in an emulated environment. Record and replay is promising because recordings from real workloads naturally satisfy inter-syscall dependencies. However, deterministic record and replay requires capturing all external and non-deterministic events, such as concurrency across multiple threads, precise timings between syscalls, and device-side interrupts. Accurately recording and replaying all these factors can be prohibitively complex and significantly degrade fuzzing speed and efficiency.

To mitigate the issues of non-determinism and enable reaching deep driver states, snapshotting has emerged as a crucial technique. Snapshotting allows capturing the entire state of a system (including CPU registers, memory, and device states) at a specific point in time. By recalling desired driver states from snapshots, fuzzers can effectively bypass non-determinism problems. Furthermore, snapshots can address the device-side input challenge by allowing a normal GPU workload to drive the driver into a deep, complex state. A snapshot is then taken at this deep state, and fuzzing can commence from that known, rehosted state, effectively providing the necessary "device-side input" without needing to reconstruct it from scratch.

Moneta's core insight is to combine the strengths of snapshotting with record and replay. By leveraging snapshots, Moneta can deterministically recall deep driver states, overcoming the non-determinism challenges of pure record-and-replay. Simultaneously, by using record and replay after a snapshot, Moneta retains the ability to perform evolutionary fuzzing from these deep, complex states, offering a powerful synergy that addresses the limitations of prior in-vivo and ex-vivo approaches.

Key Findings

▶ Watch: Fuzzer implantation with Moneta agent preserving driver context. (4:15)

Moneta's primary contribution is a novel ex-vivo GPU driver fuzzing system that synergistically combines snapshotting and record-and-replay techniques. This combination allows the fuzzer to deterministically recall deep driver states while retaining the capability for evolutionary fuzzing, a critical advancement for uncovering complex vulnerabilities.

The core idea of Moneta is straightforward yet powerful:

  1. A standard GPU workload is executed to drive the driver into a deep, complex internal state.
  2. At this desired deep state, a snapshot of the entire system (including the driver's internal context and the GPU process) is taken.
  3. This snapshot is then rehosted onto a fuzzing machine.
  4. Crucially, the original GPU process within the snapshot is transformed into a fuzzer through a technique called "fuzzer implantation."
  5. During the initial snapshotting phase, Moneta also performs "post-snapshot recordings," which are syscalls executed immediately after the snapshot point.
  6. On the implanted fuzzer in the rehosted environment, these post-snapshot recordings are replayed and then mutated to explore variations and discover bugs within the deep driver states.

This approach offers several key advantages: it effectively mitigates the inter-syscall dependency challenge because the initial workload inherently generates dependency-satisfying sequences. It also addresses the hardware dependency challenge by establishing a consistent driver state via snapshotting, allowing subsequent fuzzing to proceed without direct physical hardware interaction.

Moneta's effectiveness was rigorously evaluated against state-of-the-art fuzzers and strong baselines. The system demonstrated significant improvements in code coverage, achieving a maximum increase of 137.8% in basic block coverage compared to a baseline fuzzer lacking both snapshotting and record-and-replay. Even against a strong baseline that only removed snapshotting (simulating a typical record-and-replay fuzzer), Moneta showed substantial coverage gains, particularly for Nvidia drivers, with an approximate 130% increase. This indicates that the combination of techniques is indeed more powerful than either component alone.

Beyond coverage, Moneta proved its practical utility by discovering 10 previously unknown bugs across three different GPU drivers (Nvidia, AMD, and Mali). Of these, five were assigned CVEs, confirming the severity and novelty of the found vulnerabilities. These findings highlight Moneta's ability to uncover real-world security flaws that might otherwise remain undetected. The system also demonstrated reasonable throughput, achieving 20,000 to 30,000 program executions per minute on native x86 instances for Nvidia and AMD drivers, and approximately 2,000 executions per minute for emulated ARM environments.

Finally, Moneta has been open-sourced, contributing a valuable new tool and methodology to the security research community for future GPU driver vulnerability discovery.

Technical Deep Dive

▶ Watch: Preserving device file descriptors during rehosting for stateful fuzzing. (5:30)

The technical implementation of Moneta hinges on several sophisticated techniques designed to enable seamless ex-vivo fuzzing from deep driver states. The core challenge is to rehost a snapshotted system and transform an existing GPU workload process into a fuzzer without disrupting the GPU driver's internal context.

Fuzzer Implantation and PID Preservation

A critical aspect of rehosting is the transformation of the original GPU process into the fuzzer. This process must occur without destroying the driver's internal context, which is often tied to the process that initiated the workload. Moneta achieves this through the Moneta Agent.

The Moneta Agent operates during the fuzzing phase on the rehosted snapshot. Its primary role is to intercept the first syscall made by the original GPU process. Upon interception, the agent manipulates the syscall arguments to force the GPU process to execute (execve) the fuzzer binary. This is a crucial step because it allows the fuzzer to inherit the same Process ID (PID) as the original GPU workload.

The preservation of the original PID is vital due to a common security mechanism observed in many GPU drivers, including Nvidia's. The speaker specifically noted that the Nvidia driver performs PID checks during operations like Memory-mapped I/O (MEP). If the PID of the process attempting to interact with the driver changes, these checks can fail, causing the driver to deny access, crash, or behave unexpectedly, thereby preventing effective fuzzing of deep states. By maintaining the original PID, Moneta ensures that the fuzzer can continue interacting with the driver as if it were the original, legitimate workload, preserving the live driver internal context.

Device File Descriptor Preservation

Another significant challenge in execve-based fuzzer implantation is the handling of device file descriptors (FDs). Workloads typically open device files (e.g., /dev/nvidia0) with the close-on-exec flag set. This flag dictates that the file descriptor should be automatically closed if the process successfully executes a new program (like our fuzzer) via execve. If these critical device FDs were closed, the fuzzer would lose its established communication channels with the GPU driver, rendering deep state fuzzing impossible.

To counteract this, the Moneta Agent also operates during the initial snapshotting phase, when the workload is first executing. It intercepts all open syscalls made by the GPU workload process. For any open call that pertains to a device file, the agent erases the close-on-exec flag from the file descriptor's flags prior to the open call actually being executed by the kernel. This manipulation ensures that when the fuzzer is execve'd later, the device file descriptors remain open and valid, allowing the fuzzer to seamlessly continue interacting with the GPU driver. These two techniques—PID preservation and FD preservation—are fundamental to Moneta's ability to maintain the driver's internal context and enable stateful fuzzing from deep states.

Practicality for Mainstream GPUs

Moneta was designed to be a practical fuzzer for all mainstream GPUs, encompassing both discrete (e.g., Nvidia, AMD) and integrated (e.g., Mali) architectures. This presented a unique rehosting challenge, particularly for ARM-based integrated GPUs.

  • Discrete GPUs (x86): Rehosting for discrete GPUs on x86 architectures is relatively straightforward. A KVM-enabled x86 virtual machine (where the snapshot is taken) can be easily rehosted onto another KVM-enabled x86 machine (the fuzzing environment). The underlying CPU architectures and features are largely compatible.
  • Integrated GPUs (ARM): For integrated GPUs, often found in ARM-based systems, rehosting proved more complex. The primary issue was that the emulated ARM vCPU in the fuzzing environment (e.g., QEMU on x86 emulating ARM) was not perfectly up-to-date with the features of the actual ARM CPUs used in the snapshotting environment. This disparity in CPU features between the snapshotting and fuzzing environments caused rehosting to fail. The guest operating system or driver would encounter unexpected CPU features or lack expected ones, leading to instability or crashes.

Moneta's solution to this problem was to intervene at the hypervisor level. By disabling the unmatching guest CPU features in the hypervisor (QEMU in this case), Moneta ensured that the virtual CPUs in both the snapshotting and fuzzing states presented an identical set of CPU features to the guest operating system. This synchronization allowed for successful rehosting of ARM-based integrated GPU environments, making Moneta capable of fuzzing a broad spectrum of GPU architectures.

Experimental Configuration

For its evaluation, Moneta was implemented on QEMU with Linux KVM and tested on both x86 and ARM platforms. The specific GPU drivers targeted included:

  • Nvidia GPUs: Three different architectures were tested.
  • AMD GPUs: One architecture was tested.
  • Mali GPUs: One architecture was tested, representing integrated ARM-based GPUs.

The ablation studies conducted, such as the "no FD config," clearly demonstrated the critical importance of these technical innovations. The "no FD config," which lacked file descriptor preservation, showed a significant drop in coverage, reinforcing that Moneta's specific techniques for preserving driver internal context are essential for its effectiveness.

Demo / Proof of Concept

▶ Watch: Enabling integrated ARM GPU fuzzing by handling CPU features. (6:15)

The talk did not feature a live demonstration of Moneta in action. Instead, the speakers presented comprehensive evaluation results and performance metrics to serve as evidence of the system's efficacy and capabilities. The "proof of concept" for Moneta's design and implementation was substantiated through rigorous testing and comparison against baseline fuzzers, along with the concrete discovery of zero-day vulnerabilities.

The evaluation section highlighted the following key results:

  • Coverage Increase: Moneta achieved a maximum of 137.8% increase in basic block coverage compared to state-of-the-art fuzzers, and approximately 130% increase when compared to a baseline without snapshotting and replay, specifically for Nvidia drivers. This demonstrated its superior ability to explore more code paths within the GPU drivers.
  • Bug Discovery: The system successfully identified 10 previously unknown bugs across Nvidia, AMD, and Mali GPU drivers. Crucially, five of these bugs were assigned CVEs, validating their security impact and the practical value of Moneta.
  • Throughput: Moneta exhibited robust performance, with native x86 instances (for Nvidia and AMD) executing 20,000 to 30,000 programs per minute. Emulated ARM environments performed reasonably well, achieving approximately 2,000 executions per minute. These metrics confirm Moneta's practicality for large-scale fuzzing campaigns.
  • Ablation Studies: Specific ablation studies, such as the "no FD config," showcased the critical contribution of individual components, like file descriptor preservation, to Moneta's overall effectiveness in maintaining driver internal context.

These quantitative and qualitative results collectively serve as the proof of concept for Moneta, demonstrating that its novel combination of snapshotting and record-and-replay, coupled with its advanced rehosting techniques, is highly effective in uncovering vulnerabilities in complex GPU drivers.

Defensive Implications

▶ Watch: Evaluation: Basic block coverage results and baseline comparisons. (7:45)

The Moneta research underscores the critical and often underestimated vulnerability of GPU drivers, which act as a complex interface between user-space applications and powerful hardware. The discovery of 10 zero-day bugs, with five assigned CVEs, serves as a stark reminder that these drivers represent a significant attack surface that demands rigorous security scrutiny.

For GPU driver developers and vendors, the implications are clear:

  • Enhance Fuzzing Capabilities: Vendors should invest in and integrate advanced fuzzing techniques like Moneta into their development and quality assurance pipelines. The ability of Moneta to reach deep driver states and achieve high coverage ex-vivo makes it a strong candidate for continuous integration testing, significantly bolstering their bug-finding capabilities before product release.
  • Review Driver Design for Fuzzing Resilience: The specific techniques employed by Moneta, such as PID preservation and file descriptor manipulation, highlight areas where driver internal logic might be susceptible to bypass or exploitation. Developers should review their drivers' reliance on specific PIDs for context, the use of close-on-exec flags for device files, and overall state management to ensure robustness against such manipulations. Stricter input validation and robust memory safety practices remain foundational.
  • Embrace Open Source Contributions: The open-sourcing of Moneta provides a valuable resource. Developers can leverage its methodology and code to test their own drivers or contribute to its further development, fostering a more secure ecosystem.

For security practitioners and system administrators, the findings reinforce key defensive postures:

  • Prioritize Driver Updates: GPU drivers are not merely performance enhancements; they are critical security components. Organizations must prioritize and promptly apply security updates for GPU drivers, as these patches often address high-severity vulnerabilities that could lead to privilege escalation, arbitrary code execution, or denial of service.
  • Understand the Attack Surface: Recognize that GPU drivers can be targeted by sophisticated attackers to gain control over systems, especially in environments where GPUs are heavily utilized (e.g., AI/ML clusters, gaming, VDI). Security audits should extend to the GPU stack.
  • Monitor Vendor Advisories: Stay informed about security advisories issued by GPU manufacturers (Nvidia, AMD, Intel, ARM/Mali). The frequency of CVEs in these drivers necessitates constant vigilance.

Overall, Moneta demonstrates that effective ex-vivo fuzzing for GPU drivers is not only possible but highly impactful. Its success should serve as a call to action for both developers to build more secure drivers and for users to prioritize their maintenance and security.

Key Takeaways

  • Moneta's Hybrid Approach: Moneta introduces a novel ex-vivo GPU driver fuzzing system that ingeniously combines deterministic snapshotting with evolutionary record-and-replay, enabling efficient exploration of deep driver states.
  • Critical Fuzzer Implantation Techniques: The system relies on sophisticated methods like PID preservation (via Moneta Agent's execve manipulation to bypass driver PID checks like Nvidia's MEP) and device file descriptor preservation (by erasing close-on-exec flags) to maintain essential driver internal context during fuzzing.
  • Broad GPU Compatibility: Moneta is designed to fuzz both discrete (x86) and integrated (ARM) GPUs. Its ability to handle integrated ARM GPUs by disabling unmatching guest CPU features at the hypervisor level ensures wide applicability across mainstream architectures.
  • Significant Coverage and Bug Discovery: The fuzzer achieved up to a 137.8% increase in basic block coverage compared to prior methods and successfully uncovered 10 previously unknown bugs, with five assigned CVEs, demonstrating its high effectiveness in practical vulnerability discovery.
  • Practical Performance and Open Source: Moneta offers reasonable fuzzing throughput (e.g., 20,000-30,000 executions/minute on x86) and has been open-sourced, providing a valuable tool and methodology for the security community.
  • Highlighting GPU Driver Vulnerabilities: The research underscores the persistent vulnerability of GPU drivers and the critical need for advanced, scalable fuzzing techniques to secure these complex and increasingly crucial components in modern computing.

About the Speaker(s)

Joonkyo Jung is a PhD student at Jon University. He presented the work "Moneta: Ex-Vivo GPU Driver Fuzzing by Recalling In-Vivo Execution States" at the NDSS Symposium. This research was a collaborative effort with other researchers from Jon University, as well as contributions from Kuven and Teod Delft (likely referring to institutions or research groups associated with TU Delft). His work focuses on enhancing the security of complex system components, specifically GPU drivers, through advanced fuzzing methodologies.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Moneta is legitimate systems security research with a clear technical contribution: a hybrid snapshot + record-replay ex-vivo fuzzer that solves real, concrete problems in GPU driver fuzzing. Five CVEs across Nvidia, AMD, and Mali, plus an open-source release, is the receipts. PhD student presenting novel doctoral work — this is exactly what a research track should look like.

Heather Calloway (CISO) — PASS

Technically rigorous fuzzing research with real CVE output and a meaningful methodological contribution. Outside my lane entirely — no governance angle, no institutional risk framing, no defender or operator path for anyone in my orbit.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025