CAMP in the Odyssey: Provably Robust Reinforcement Learning with Certified Radius Maximization
Derui Wang
34th USENIX Security Symposium (USENIX Security '25) · Day 3 · ML and AI Security 4: Robustness
Overview
Deep Reinforcement Learning (DRL) agents are increasingly deployed in high-stakes environments, from autonomous vehicles to critical infrastructure control. However, the inherent vulnerability of these agents to adversarial perturbations—small, often imperceptible changes to their sensory observations—presents a significant challenge to their trustworthiness and safe deployment. This talk by Derui Wang introduces "CAMP in the Odyssey," a novel approach designed to enhance the certified robustness of DRL agents. The work addresses a critical limitation in existing certification methods: the inability to directly optimize the certified radius, a key metric indicating an agent's resilience to adversarial attacks.

Key moments
- 0:00 Introduction to DRL robustness challenges and motivation
- 3:15 Core research question: optimizing certified radius
- 3:30 Introducing CAM loss and Q-value gap concept
- 6:30 Policy imitation for stable CAM integration
- 8:00 Experimental setup and baseline comparisons
- 9:00 Significant gains in certified robustness demonstrated
- 9:30 Empirical robustness against PGD and AutoPGD attacks
- 10:55 Summary of advantages and limitations
CAMP in the Odyssey: Provably Robust Reinforcement Learning with Certified Radius Maximization
Speakers: Derui Wang
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=FrdBAml7FR4
Overview
Deep Reinforcement Learning (DRL) agents are increasingly deployed in high-stakes environments, from autonomous vehicles to critical infrastructure control. However, the inherent vulnerability of these agents to adversarial perturbations—small, often imperceptible changes to their sensory observations—presents a significant challenge to their trustworthiness and safe deployment. This talk by Derui Wang introduces "CAMP in the Odyssey," a novel approach designed to enhance the certified robustness of DRL agents. The work addresses a critical limitation in existing certification methods: the inability to directly optimize the certified radius, a key metric indicating an agent's resilience to adversarial attacks.
The core of this research revolves around the development of the Certified Radius Maximization (CAMP) loss function, which, when coupled with a policy imitation training paradigm, enables DRL agents to achieve significantly higher levels of provable robustness. By reformulating the certified radius and deriving a surrogate function that can be optimized during training, CAMP provides a clear and explainable mechanism for improving an agent's resilience. This work is crucial for advancing the practical application of DRL in safety-critical domains, offering a pathway toward agents that are not only performant but also verifiably robust against sophisticated adversarial manipulations.
Background
▶ Watch: Introduction to DRL robustness challenges and motivation (0:00)
The application of Deep Reinforcement Learning (DRL) agents spans a wide array of high-stakes domains, including robotics, autonomous systems, and even financial trading. Their ability to learn complex behaviors directly from experience has made them a powerful tool. However, this power comes with a significant vulnerability: DRL agents, much like deep neural networks in other fields, are highly susceptible to adversarial attacks. Attackers can introduce subtle perturbations to an agent's observations of its environment, leading the agent's policy to recommend incorrect or even catastrophic actions. In real-world scenarios, such attacks could result in severe consequences, from system failures to physical harm.
To counteract these vulnerabilities, a significant line of research has focused on providing provable robustness guarantees for DRL agents. One prominent approach is policy smoothing. This certification method offers a global lower bound on the expected return of the agent, even under adversarial conditions. The mechanism involves adding Gaussian noise to the observed state at each time step, creating a randomized state-action trajectory. By repeatedly running episodes and collecting these randomized trajectories, statistical methods, specifically an adaptive version of the Neyman-Pearson lemma, can be applied to certify a lower bound on the probability that the agent's return exceeds a given threshold. From this, a lower bound on the expected return of these randomized trajectories can be computed.
Despite its utility, policy smoothing, and similar existing certification methods, exhibit a clear trade-off between certified expected return and certified radius. The certified radius quantifies the maximum perturbation an agent can withstand while still guaranteeing a certain level of performance. As the certified radius increases, the certified expected return—the guaranteed performance—often drops sharply. A steeper trade-off curve indicates a less robust agent, making it unsuitable for deployment in environments where reliability under stress is paramount. Furthermore, a critical limitation of these existing methods is that the certified radius is typically determined through numerical approaches like binary search. This indirect determination makes it exceedingly difficult to understand the underlying factors influencing robustness or to strategically optimize the certified radius during the agent's training process. This inability to directly influence and improve the certified radius forms the core problem that this research aims to address, seeking to develop a method that can explicitly maximize this crucial robustness metric.
Key Findings
▶ Watch: Introducing CAM loss and Q-value gap concept (3:30)
The research presented in "CAMP in the Odyssey" introduces a transformative approach to enhancing the certified robustness of Deep Reinforcement Learning (DRL) agents, yielding several significant key findings:
Firstly, the paper demonstrates that it is possible to reformulate and directly optimize the certified radius of a DRL agent, moving beyond the limitations of numerical search methods. By transforming the constant certified radius into a function dependent on the policy, a new variable, and a sample return value, the authors paved the way for its explicit maximization during training. This is achieved through the introduction of the Certified Radius Maximization (CAMP) loss function.
Secondly, the CAMP loss function, when integrated into a carefully designed training paradigm, leads to significant improvements in certified robustness. Across six diverse environments, including three classic control tasks and three Atari games, agents trained with CAMP consistently exhibited higher certified radii compared to baseline methods like Gaussian noise injection and NoisyNet. The gains were particularly pronounced in classic control tasks, where the certified robustness saw substantial increases.
Thirdly, beyond provable guarantees, CAMP also delivers tangible gains in empirical robustness. The agents trained with CAMP demonstrated superior resilience against both PGD (Projected Gradient Descent) attacks and the stronger AutoBGD attacks. Even when facing persistent, adaptive adversaries that manipulate observations at each step within a fixed perturbation budget, CAMP-trained agents maintained higher performance compared to baselines. While the empirical gains on Atari games appeared smaller, the authors noted that agents in these environments already possess a relatively high baseline robustness against PGD attacks.
Fourthly, the proposed framework offers clear and explainable improvements in certified robustness. By tying the certified radius to a per-step local radius, which is further linked to the Q-value gap (the difference between the top-1 and runner-up Q-values), the method provides an intuitive understanding of how robustness is enhanced. Pushing this Q-value gap wider directly correlates with an increased certified radius, making the mechanism of improvement transparent.
Finally, the research establishes that the policy imitation training setup is a crucial component for safely integrating the CAMP loss. It resolves the challenge of Q-value overestimation and training instability that would arise if CAMP were directly applied to Q-learning. This makes the overall framework scalable to environments with large discrete action spaces, expanding the applicability of certified robustness to more complex DRL problems.
Technical Deep Dive
▶ Watch: Experimental setup and baseline comparisons (8:00)
The core innovation of "CAMP in the Odyssey" lies in its ability to directly optimize the certified robustness of DRL agents, a capability previously hindered by the indirect nature of radius determination. This is achieved through a meticulous three-step process that culminates in the Certified Radius Maximization (CAMP) loss function and its integration via a policy imitation training paradigm.
The first step in developing CAMP involves a reformulation of the certified radius through a change of variables. Traditionally, the certified radius is a constant value determined post-training. The authors transform this into a dynamic function of the agent's policy ($\pi$), a new variable ($s$), and a sample return value. This reformulation is critical because it makes the certified radius an optimizable quantity that can be influenced during the agent's learning process.
Following the reformulation, the second step focuses on deriving a surrogate function that is positively correlated with the newly formulated certified radius. This surrogate function acts as a proxy that can be maximized to indirectly increase the certified radius. The design of this function is based on theoretical insights into how perturbations affect an agent's decision-making and subsequent returns.
The third and most practical step breaks down this surrogate function into a per-step local radius. This local radius is directly tied to a crucial statistic defined as the Q-value gap. The Q-value gap is simply the difference between the highest Q-value (corresponding to the optimal action) and the second-highest Q-value (corresponding to the runner-up action) as predicted by an Oracle Q-network. The fundamental insight here is that a larger Q-value gap implies a more confident decision by the agent, making it less susceptible to small perturbations that might otherwise flip the preferred action. Therefore, the certified radius increases as this Q-value gap gets larger.
This understanding directly informs the design of the CAMP loss function. When the agent selects the correct action, the CAMP loss actively works to widen this Q-value gap. By pushing the Q-value of the chosen action higher relative to all other actions, the agent's decision boundary becomes more robust, increasing its tolerance to adversarial noise.
However, a significant challenge arises when attempting to directly integrate the CAMP loss into standard DRL training paradigms, particularly those based on Q-learning. Q-learning, which directly models Q-values using a Q-network, minimizes a temporal difference (TD) loss at each time step, pushing the predicted Q-value for a state-action pair towards a target Q-value. A known issue with Q-learning, prevalent for over a decade, is Q-value overestimation. When Q-values are overestimated, training often converges to suboptimal policies. The problem with directly applying CAMP loss in this context is that by design, CAMP pushes the Q-value gap wider, which can inadvertently lead to higher Q-values themselves. This can exacerbate the overestimation problem, causing training instability or even preventing the model from converging to a stable policy.
To circumvent this instability, the authors propose a novel training setup: policy imitation. This paradigm involves two distinct phases:
- Reference Policy Training: First, a reference policy is trained in the normal way, typically using a standard Q-learning approach (e.g., Deep Q-Networks or DQN) to model an optimal "oracle" policy. This reference policy learns to estimate the Q-values and make decisions based on them.
- Primary Policy Training with Imitation and CAMP: Next, the primary policy is trained. Crucially, this primary policy learns to imitate the actions recommended by the reference policy, but it does not directly attempt to copy its Q-values. Instead, the training involves computing the cross-entropy loss between the normalized Q-value vectors derived from the primary and reference policy networks. This cross-entropy loss encourages the primary policy to align its action probabilities with those of the reference network, essentially learning the behavior without directly modeling the potentially overestimated Q-values.
With this policy imitation loss in place, the CAMP loss can then be safely added during the training of the primary policy. Because the primary policy is no longer directly modeling or replicating the Q-values, the instability issues caused by CAMP pushing Q-values higher are effectively avoided. This two-stage approach allows the benefits of CAMP's certified radius maximization to be realized without compromising the stability and convergence of the DRL agent's training.
The experimental evaluation further substantiates these technical claims. The method was tested across six environments: three classic control tasks (e.g., CartPole, Acrobot) and three Atari games (e.g., Pong, Breakout). Comparisons were made against two baselines: one that injects Gaussian noise during training (similar to the certification process) and NoisyNet, which randomizes Q-network weights for efficient exploration. CAMP consistently outperformed both baselines in terms of certified robustness, particularly on classic control tasks. Furthermore, under empirical adversarial attacks like PGD and the more potent AutoBGD, CAMP agents demonstrated superior resilience, maintaining higher performance under fixed perturbation budgets. AutoBGD, known for achieving its goals with smaller per-step perturbations, thus sustaining attacks over more steps, still found CAMP-trained agents more robust.
In summary, the technical deep dive reveals a sophisticated framework that addresses the fundamental challenge of optimizing certified robustness in DRL. By carefully reformulating the problem, introducing a targeted loss function (CAMP), and designing a stable training paradigm (policy imitation), the authors provide a robust and scalable solution for building more trustworthy DRL agents.
Demo / Proof of Concept
▶ Watch: Significant gains in certified robustness demonstrated (9:00)
While the talk did not feature a live, interactive demo in the traditional sense, the authors presented extensive experimental results across a range of environments, serving as a robust proof of concept for the effectiveness of CAMP and the policy imitation framework. These experiments were conducted in six distinct DRL environments: three classic control tasks and three Atari games.
For the classic control tasks, which typically involve continuous state spaces and discrete action spaces, the experiments demonstrated a significant improvement in both certified and empirical robustness. The certified radius was quantitatively shown to increase, indicating a higher provable guarantee against adversarial perturbations. Empirically, agents trained with CAMP exhibited greater resilience when subjected to PGD (Projected Gradient Descent) attacks, which involve manipulating observations with a fixed perturbation budget, and even against the stronger AutoBGD attacks, designed for more adaptive and sustained adversarial pressure.
Similarly, for the Atari games, which present more complex visual observations and larger discrete action spaces, the framework was evaluated for its scalability and performance. While the gains in empirical robustness against PGD and AutoBGD attacks appeared less pronounced on Atari games compared to classic control tasks, the authors noted that agents in these environments already possess a relatively high baseline robustness. Despite this, CAMP still contributed to improved certified robustness, showcasing its ability to provide provable guarantees even in more intricate settings.
The comprehensive evaluation across these diverse environments, comparing CAMP against established baselines like Gaussian noise injection and NoisyNet, effectively serves as the proof of concept, demonstrating the practical applicability and superior performance of the proposed method in enhancing the certified and empirical robustness of DRL agents.
Defensive Implications
▶ Watch: Summary of advantages and limitations (10:55)
The "CAMP in the Odyssey" research offers profound defensive implications for the deployment of Deep Reinforcement Learning (DRL) agents in real-world, safety-critical applications. The primary takeaway for defenders is the provision of a systematic and optimizable approach to building provably robust DRL systems.
Firstly, the introduction of the Certified Radius Maximization (CAMP) loss function combined with the policy imitation training paradigm provides a concrete methodology for DRL practitioners to enhance the resilience of their agents. Instead of merely relying on empirical testing against known attacks, which can always be bypassed by novel adversarial strategies, CAMP allows for the training of agents with a quantifiable, lower-bound guarantee on their performance under perturbation. This is critical for applications where failure modes due to adversarial inputs are unacceptable, such as autonomous driving, medical systems, or critical infrastructure control. Defenders can now actively incorporate CAMP into their training pipelines to ensure a higher certified radius, meaning their agents can tolerate larger adversarial perturbations while maintaining their expected return.
Secondly, the mechanism behind CAMP, which involves widening the Q-value gap, offers explainable improvements in robustness. Defenders gain insight into why their agents are more robust: because their decision boundaries are clearer and more confident. This transparency can aid in debugging and validating the robustness properties of DRL systems, moving beyond black-box assessments. For security analysts, understanding this underlying mechanism allows for more targeted analysis and potentially the development of new defensive strategies that further leverage this principle.
Thirdly, the policy imitation training paradigm addresses a fundamental challenge of integrating robustness-enhancing losses with standard DRL algorithms. By decoupling the primary policy's learning of actions from direct Q-value modeling, it mitigates issues like Q-value overestimation and training instability. This means defenders can apply CAMP without compromising the overall stability and convergence of their DRL models, making the approach practical for integration into existing DRL development workflows. This also implies that future research into robust DRL can explore similar decoupled training strategies to safely incorporate other complex loss functions.
Finally, the demonstrated effectiveness against both PGD and AutoBGD attacks highlights CAMP's utility against a spectrum of adversarial threats, from basic gradient-based attacks to more sophisticated, adaptive, and persistent adversaries. While the improvements were more pronounced in classic control tasks, the method still offered certified robustness gains in complex environments like Atari games, indicating its broad applicability. Defenders should consider CAMP as a foundational technique for hardening DRL agents against a variety of adversarial manipulations, moving towards a future where DRL systems are not just intelligent, but also inherently secure and trustworthy. The limitations, such as applicability only to discrete action spaces and potential training overhead, also guide defenders on areas for future research and where the current method might need further refinement for specific deployment scenarios.
Key Takeaways
- Optimizable Certified Robustness: The "CAMP in the Odyssey" paper introduces a novel approach to directly optimize the certified radius of DRL agents, moving beyond passive numerical determination.
- CAMP Loss Function: The Certified Radius Maximization (CAMP) loss function is designed to increase the Q-value gap (difference between top-1 and runner-up Q-values), thereby enhancing the agent's confidence and resilience to perturbations.
- Policy Imitation for Stability: A policy imitation training paradigm is crucial for safely integrating CAMP loss, preventing Q-value overestimation and training instability by decoupling action learning from direct Q-value modeling.
- Significant Robustness Gains: CAMP significantly improves both certified robustness (provable guarantees) and empirical robustness (against PGD and AutoBGD attacks), especially in classic control tasks.
- Explainable Improvements & Scalability: The method provides clear, explainable improvements linked to the Q-value gap and is scalable to environments with large discrete action spaces.
- Limitations & Future Work: Current limitations include applicability only to discrete action spaces, less pronounced improvements on Atari games, and potential additional training overhead, which are targets for future research.
About the Speaker(s)
Derui Wang is the presenter of "CAMP in the Odyssey: Provably Robust Reinforcement Learning with Certified Radius Maximization" at USENIX Security. The research, as stated at the beginning of the talk, is a collaborative effort involving authors from several institutions. These include SIS AROS Data61, the Cyber Security Cooperative Research Centre of Australia, Teaching University, and the University of Chicago. While the talk focuses on the technical aspects of the research, the affiliations suggest a strong background in cybersecurity, artificial intelligence, and robust machine learning, reflecting expertise from both academic and research-oriented organizations.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate ML security research with a clear technical contribution — directly optimizing certified radius via a Q-value gap surrogate is a real advance over binary-search-after-the-fact approaches. Competent work, but this is a conference paper presentation, not a security operations talk, and the threat model stays firmly in the academic sandbox.
Heather Calloway (CISO) — PASS
Rigorous academic work on certified robustness for deep reinforcement learning agents. Outside my lane — no governance angle, no operator path, no institutional relevance.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)