TrajDeleter: Enabling Trajectory Forgetting in Offline Reinforcement Learning Agents
Chen Gong (Kung Fu University Virginia)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Machine Unlearning
Overview
This talk introduces TrajDeleter, a novel framework designed to enable efficient and stable trajectory forgetting in offline reinforcement learning (RL) agents. Presented by Chen Gong from Kung Fu University Virginia, the research addresses a critical and increasingly relevant challenge in the era of data privacy and regulatory compliance: the "right to be forgotten" in machine learning models, specifically within the context of RL agents trained on fixed datasets of past interactions. The inability to selectively remove the influence of specific data points (trajectories) from a trained RL agent without costly retraining poses significant hurdles for applications ranging from autonomous medical treatment plans that handle sensitive patient information to robotics that might inadvertently store private environmental layouts, and adherence to regulations like GDPR.
Key moments
- 0:00 Introduction to trajectory unlearning and privacy motivations
- 1:50 Drawbacks of existing unlearning methods: retraining, fine-tuning
- 3:00 Introducing TrajDeleter to address unstable RL training
- 4:00 TrajDeleter's two-phase: forgetting and convergence training
- 6:00 Verifying unlearning success using membership inference attacks
- 8:20 Key experimental insights and limitations of the approach
TrajDeleter: Enabling Trajectory Forgetting in Offline Reinforcement Learning Agents
Speakers: Chen Gong, Kung Fu University Virginia
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=e_j4E5RSg8s
Overview
This talk introduces TrajDeleter, a novel framework designed to enable efficient and stable trajectory forgetting in offline reinforcement learning (RL) agents. Presented by Chen Gong from Kung Fu University Virginia, the research addresses a critical and increasingly relevant challenge in the era of data privacy and regulatory compliance: the "right to be forgotten" in machine learning models, specifically within the context of RL agents trained on fixed datasets of past interactions. The inability to selectively remove the influence of specific data points (trajectories) from a trained RL agent without costly retraining poses significant hurdles for applications ranging from autonomous medical treatment plans that handle sensitive patient information to robotics that might inadvertently store private environmental layouts, and adherence to regulations like GDPR.
The core problem TrajDeleter tackles is the computational inefficiency and instability of existing unlearning approaches for offline RL. Traditional methods like full retraining are prohibitively time-consuming and expensive, especially when unlearning requests are frequent. Simpler alternatives like fine-tuning or randomizing rewards often lead to unreliable or unstable agent performance. TrajDeleter proposes a two-phase training methodology that not only achieves effective unlearning but also maintains agent stability and performance, all while introducing an innovative auditing mechanism to verify the success of the unlearning process. This work is particularly significant as it also offers insights into potential future applications for large language models (LLMs) that leverage Reinforcement Learning from Human Feedback (RLHF), a specific type of offline RL.
Background
▶ Watch: Introduction to trajectory unlearning and privacy motivations (0:00)
Reinforcement Learning (RL) involves an agent learning to make decisions by interacting with an environment to maximize a cumulative reward. In contrast to online RL, where agents learn through real-time interaction, offline reinforcement learning focuses on training agents using a fixed, pre-collected dataset of interactions, known as trajectories. Each trajectory is a sequence of state-action-reward pairs, detailing what state the agent was in, what action it took, the reward received, and the subsequent state transition. This paradigm is crucial for applications where real-world interaction is costly, dangerous, or impractical, such as autonomous driving, robotics, or medical diagnosis.
However, the reliance on fixed datasets introduces a significant challenge: data governance. As RL systems become more prevalent in sensitive domains, the need to comply with data privacy regulations like the General Data Protection Regulation (GDPR) becomes paramount. GDPR, for instance, grants users the "right to erasure" or the "right to be forgotten," meaning users can demand that their data and its impact on a system be removed. For an RL agent trained on trajectories, this implies the ability to selectively remove the influence of specific trajectories from the agent's learned policy and value function.
Prior attempts to achieve unlearning in machine learning, and specifically in RL, have faced substantial limitations:
- Retraining from Scratch: The most straightforward and "golden line" solution is to simply drop the data to be unlearned and retrain the entire agent from scratch on the remaining dataset. While this guarantees complete removal of the data's influence, it is incredibly time-consuming and computationally expensive. For complex RL agents and large datasets, this approach is often impractical, especially when unlearning requests are frequent. The speaker highlights that this is the reference point for their algorithm's efficiency comparison.
- Fine-tuning: Another common approach is to drop the unlearning data and then fine-tune the existing agent on the remaining dataset. While faster than full retraining, fine-tuning in this context is often unreliable. The agent might retain residual knowledge from the deleted data, or the fine-tuning process could destabilize the learned policy, leading to suboptimal performance.
- Random Reward: Given that RL agents learn primarily from rewards, one might consider randomizing the rewards associated with the target trajectories. The idea is that if the agent receives random or conflicting rewards for specific state-action pairs within the target trajectories, it will become "confused" and effectively ignore those interactions during training. However, this method can make the training process highly unstable, potentially leading to a non-convergent or poorly performing agent.
The inherent instability of RL training, particularly when simultaneously optimizing both a value function (estimating cumulative future rewards from a state-action pair) and a policy (mapping states to actions), exacerbates these challenges. Many offline RL algorithms struggle with stability, making it difficult to guarantee that an unlearning process will converge to a desirable state without significant performance degradation. This background underscores the critical need for a stable and efficient unlearning method that can approximate the ideal outcome of retraining without incurring its prohibitive costs.
Key Findings
▶ Watch: Introducing TrajDeleter to address unstable RL training (3:00)
TrajDeleter presents several significant contributions to the field of machine unlearning, particularly for offline reinforcement learning agents. The key findings demonstrate both the practical viability and theoretical underpinnings of an efficient trajectory forgetting mechanism:
- Stable and Efficient Trajectory Forgetting: The primary finding is the development of TrajDeleter itself, a novel method capable of performing stable and efficient unlearning of specific trajectories in offline RL. Unlike existing methods that are either prohibitively expensive (retraining) or unreliable (fine-tuning, random reward), TrajDeleter offers a balanced approach. It achieves unlearning by combining a "forget training" phase with a "convergence training" phase, ensuring both the removal of target trajectory influence and the stability of the remaining agent's performance.
- Two-Phase Training for Stability and Convergence: The research highlights the critical importance of a two-phase training strategy. The initial forget training phase actively minimizes the impact of target trajectories while maximizing learning on the remaining dataset. This is followed by a convergence training phase, which is crucial for stabilizing the unlearned agent. The speaker emphasizes that these two phases must be conducted separately to provide a convergence guarantee for the agent. Without convergence training, the agent would not only forget target trajectories but also inadvertently "forget" useful information from the remaining dataset, leading to a significant drop in performance.
- Efficient Unlearning Performance: TrajDeleter demonstrates remarkable efficiency. Experiments show that it can achieve results similar to full retraining while using only 2.2% of the time required for retraining. This dramatic reduction in computational cost makes frequent unlearning requests feasible, addressing a major practical bottleneck for privacy-compliant RL systems. For example, the total training steps for TrajDeleter (forgetting + convergence steps) were set to 10,000, which is only 1% of the typical retraining steps, yet yielded comparable results.
- Novel Auditing Mechanism via Membership Inference Attack (MIA): A crucial aspect of unlearning is the ability to verify that data has indeed been forgotten. TrajDeleter introduces an auditor mechanism based on Membership Inference Attack (MIA). This auditor leverages the inherent instability of offline RL training to generate diverse shadow agents. By comparing the features of the unlearned agent with those of shadow agents (some trained with the target trajectory, some without), the auditor can effectively determine if the target trajectory has been successfully unlearned. This provides a robust method for ensuring accountability and trust in the unlearning process.
- Hyperparameter Sensitivity and Balance: The research also sheds light on the impact of hyperparameters, specifically the balance between the "forgetting step" and "convergence step" durations. While the forgetting step significantly impacts the unlearning performance (measured by unlearning rate), the convergence step is less sensitive to the agent's overall performance (measured by average returns). This insight allows for optimizing the training process, for instance, by allocating a larger portion of the training budget to the forgetting phase (e.g., 8,000 steps for forgetting, 2,000 for convergence out of 10,000 total in the paper's setup).
- Limitations and Future Work: The study acknowledges limitations, particularly that the unlearning auditor's effectiveness can vary across different offline RL algorithms. For instance, certain algorithms like Hopper and Walker2D exhibited higher unlearning rates (indicating successful unlearning detection by the auditor), while HalfCheetah showed lower rates due to its inherent training instability. This points to future work in developing more robust and exact trajectory auditor methods.
Overall, TrajDeleter provides a comprehensive solution for trajectory forgetting in offline RL, addressing both the technical challenges of stability and efficiency, and the practical need for verifiable unlearning in privacy-sensitive applications.
Technical Deep Dive
▶ Watch: TrajDeleter's two-phase: forgetting and convergence training (4:00)
The technical foundation of TrajDeleter rests on addressing the unique challenges of unlearning within the context of offline reinforcement learning, particularly the inherent instability of RL training and the need for rigorous convergence guarantees.
At its core, offline RL aims to train two interacting components:
- Value Function: This estimates the expected cumulative future reward an agent will receive starting from a given state and taking a particular action, until a terminal state is reached. It quantifies the "goodness" of a state-action pair.
- Policy: This determines the agent's behavior, mapping an input state to an output action. The goal is to learn an optimal policy that maximizes the cumulative reward.
The speaker notes that training these two functions simultaneously, much like in Generative Adversarial Networks (GANs), often leads to instability. This instability is a known drawback in offline RL, making it difficult to achieve reliable training and convergence.
TrajDeleter tackles this by introducing a two-phase unlearning process:
1. Forget Training
The first phase, forget training, is designed to actively remove the influence of the target trajectories while preserving the knowledge gained from the remaining dataset. The intuitive idea, borrowed from gradient-based learning, is that if "learning" is maximizing, then "unlearning" is minimizing. Specifically, the agent is trained to:
- Minimize the impact of the target trajectories (the data to be forgotten). This involves updating the agent's parameters in a way that reduces its reliance on or memory of these specific state-action-reward sequences.
- Maximize its learning on the remaining dataset. This ensures that the agent continues to perform well on the data it is supposed to remember.
This simultaneous optimization, however, presents a significant challenge: direct application of such a "forgetting" objective can lead to further instability and a lack of convergence guarantee. The agent might struggle to find a stable state where it has forgotten the target data without also corrupting its knowledge of the remaining data.
2. Convergence Training
To overcome the instability of forget training and provide a convergence guarantee, TrajDeleter introduces a separate convergence training phase. This phase is crucial for stabilizing the unlearned agent and ensuring its performance on the remaining data is robust. The key insight here, backed by prior theoretical work, is to make the unlearned agent learn the value functions of the original agent on the remain data set.
The process is as follows:
- First, the forget training phase is completed. This yields an agent that has attempted to unlearn the target trajectories.
- Second, the convergence training phase begins. During this phase, the unlearned agent is further trained, but with a specific objective: to align its value function with the value function that the original agent (trained on the full dataset, before unlearning) would have learned if it had only been exposed to the remaining dataset.
- The speaker explicitly states that these two phases—forget training and convergence training—must be done separately. Attempting to perform them simultaneously would reintroduce the instability issues that the sequential approach is designed to mitigate. This sequential execution is critical for achieving a theoretical convergence guarantee, ensuring the unlearned agent converges to a unique optimal policy or value function.
The convergence training is vital because, as demonstrated in the experimental results, without it, the "remain data rate" (performance on the non-unlearned data) significantly decreases. This means that the forget training alone would not only forget the target trajectories but also inadvertently "forget" useful information from the rest of the dataset, leading to a substantial drop in the agent's overall performance. The convergence training effectively "re-stabilizes" the agent and restores its performance on the desired data.
Auditing Unlearning Success: Membership Inference Attack (MIA)
A critical component of any unlearning framework is the ability to audit whether the unlearning process was successful. How can a user or a regulatory body trust that their data has indeed been erased? TrajDeleter proposes an auditor mechanism based on Membership Inference Attack (MIA), adapted from previous works, but with a clever twist that leverages the inherent instability of offline RL.
The traditional MIA setup involves:
- Training multiple shadow agents (models) on various subsets of the data. Some shadow agents include the target trajectory, while others do not.
- Comparing the behavior or features of the "unlearning agent" (the one whose unlearning success is being audited) with these shadow agents. If the unlearning agent's features are more similar to shadow agents without the target trajectory, then unlearning is deemed successful.
TrajDeleter enhances this by exploiting the instability of offline RL training. While instability is generally a drawback, it becomes an advantage for generating diverse shadow agents. Here's how:
- Instead of training shadow agents from scratch on carefully curated datasets (which can be inefficient), the method proposes to fine-tune the original agent (trained on the full dataset) multiple times on the original dataset itself.
- Because offline RL training is often unstable, slight variations in the fine-tuning process (e.g., different random seeds, small changes in optimization) can lead to multiple, distinct shadow models. These shadow models will exhibit different features, reflecting the non-deterministic nature of the training process.
- The auditor then extracts features from the unlearning agent and compares their similarity with the features extracted from these diverse shadow agents. By analyzing these similarities, the auditor can make a decision about whether the target trajectory has been successfully unlearned.
The speaker notes that this approach is more efficient for shadow model training in RL compared to supervised learning, where training is typically stable and would produce very similar shadow models, making differentiation difficult. This clever use of RL's instability turns a weakness into a strength for auditing purposes.
Demo / Proof of Concept
▶ Watch: Verifying unlearning success using membership inference attacks (6:00)
While the talk did not feature a live, interactive demonstration of the TrajDeleter system, the research provided a robust experimental design and presented compelling results that serve as its proof of concept. The experiments were meticulously structured to evaluate the effectiveness of both the TrajDeleter unlearning method and its associated auditor.
The experimental setup involved:
- Investigative Algorithms: Six state-of-the-art offline RL algorithms were selected to train the agents. This broad selection ensures that TrajDeleter's effectiveness is not limited to a specific algorithm but is rather a generalizable approach. The specific algorithms were not detailed in the transcript, but their inclusion highlights the robustness of the evaluation.
- Metrics for Auditor Accuracy: The accuracy of the unlearning auditor was measured using standard classification metrics: precision and recall. A high precision and recall would indicate that the auditor is accurately identifying whether a trajectory has been forgotten or not.
- Metrics for Agent Performance: The performance of the unlearned agent was evaluated using average returns. This metric assesses how well the agent performs its task in the environment after the unlearning process, ensuring that forgetting specific data does not unduly degrade the agent's overall capability.
The key experimental findings presented as evidence of TrajDeleter's efficacy include:
- Auditor Effectiveness: The results indicated varying levels of unlearning rate (auditor accuracy) across different environments. For example, in the HalfCheetah environment, the unlearning rate was observed to be relatively low. In contrast, for environments like Hopper and Walker2D, the average unlearning rate was notably high. This variation was attributed to the inherent difficulties and instability of training specific algorithms in certain environments. The speaker identified this as a limitation and an area for future work, emphasizing the need for more exact trajectory auditor methods.
- TrajDeleter's Efficiency: A critical result highlighted was the comparison with the "retrain from scratch" baseline. TrajDeleter achieved results similar to retraining while utilizing only 2.2% of the time required for a full retraining process. This significant efficiency gain underscores the practical value of TrajDeleter for scenarios requiring frequent unlearning.
- Importance of Convergence Training: The experiments conclusively demonstrated the indispensable role of the convergence training phase. Without it, two detrimental effects were observed:
- The "remain data rate" (performance on the data that should not be forgotten) increased, meaning that the forget training alone caused the agent to inadvertently forget useful trajectories from the remaining dataset.
- The "return mean" (overall performance of the agent) experienced a significant decrease. This directly validated the theoretical claim that convergence training is essential for stabilizing the agent and preserving its performance on the non-unlearned data.
- Hyperparameter Impact: An analysis of hyperparameters, specifically the balance between forgetting steps and convergence steps, revealed that the forgetting step has a more significant impact on the unlearning performance (as measured by the unlearning rate). In contrast, the convergence step was found to be less sensitive to the agent's overall performance. This suggests an optimal strategy for resource allocation, where a larger proportion of training steps can be dedicated to the forgetting phase (e.g., 8,000 forgetting steps and 2,000 convergence steps out of a total of 10,000 steps, which itself is only 1% of full retraining steps).
These experimental results collectively serve as a strong proof of concept for TrajDeleter, demonstrating its ability to perform efficient and verifiable trajectory forgetting in offline RL agents while maintaining performance, albeit with some acknowledged challenges in auditor universality.
Defensive Implications
▶ Watch: Key experimental insights and limitations of the approach (8:20)
The development of TrajDeleter carries significant implications for organizations deploying and managing offline reinforcement learning agents, particularly those operating in regulated or privacy-sensitive domains. Rather than being solely a "defensive" measure against attacks, TrajDeleter enables a proactive defense strategy for privacy compliance and robust model management.
Here are the key defensive implications:
- GDPR and Privacy Compliance: TrajDeleter directly addresses the "right to be forgotten" mandate of regulations like GDPR. Organizations can now implement a practical and efficient mechanism to erase specific user data (trajectories) from their trained RL agents upon request. This capability is crucial for avoiding hefty fines and maintaining user trust. For applications dealing with patient-sensitive information in healthcare or private location data in robotics, this capability becomes a fundamental requirement.
- Efficient Model Updates and Data Governance: Beyond regulatory compliance, TrajDeleter allows for more agile and cost-effective model management. Instead of undergoing a full, time-consuming retraining cycle every time a portion of the data needs to be removed (e.g., due to data corruption, licensing issues, or updated user preferences), organizations can use TrajDeleter to selectively unlearn. This dramatically reduces operational costs and speeds up the deployment of updated, compliant models.
- Enhanced Trust and Transparency: The inclusion of a robust auditing mechanism, based on Membership Inference Attacks (MIA), is a critical defensive feature. This auditor allows organizations to verify that unlearning has successfully occurred. This verifiability is essential for building trust with users and regulators, providing concrete evidence of data erasure. Furthermore, for internal teams, it offers a quality assurance mechanism for unlearning operations.
- Mitigation of Data Poisoning and Bias: While not explicitly discussed as a primary use case, the ability to selectively unlearn trajectories could also serve as a defensive measure against certain forms of data poisoning or the removal of biased data. If a specific set of trajectories is identified as malicious or contributing to unwanted bias, TrajDeleter could be used to mitigate their influence without requiring a complete model rebuild.
- Challenges for Auditor Robustness: The research acknowledges that the auditor's effectiveness can vary across different offline RL algorithms and environments. This implies that defenders need to carefully evaluate and potentially adapt the auditing mechanism for their specific RL setup. A universal, highly accurate auditor remains an area for future research, suggesting that current implementations might require careful calibration and validation for each deployment.
- Future-proofing RL Systems for LLMs: The speaker highlights the relevance of TrajDeleter to large language models (LLMs) trained with Reinforcement Learning from Human Feedback (RLHF). As LLMs become more integrated into sensitive applications, the ability to selectively unlearn specific prompts or human feedback trajectories will be paramount for managing model behavior, removing undesirable biases, or complying with user requests. TrajDeleter provides foundational insights for developing such capabilities in future LLM systems.
In essence, TrajDeleter equips organizations with a powerful tool to manage the lifecycle of data within their offline RL agents, transforming the daunting task of data erasure into an efficient and verifiable process. This shifts the paradigm from reactive, costly full retraining to proactive, privacy-by-design model development.
Key Takeaways
- TrajDeleter enables efficient and stable trajectory forgetting in offline reinforcement learning agents, addressing the critical need for data erasure in privacy-sensitive applications and regulatory compliance (e.g., GDPR).
- The method employs a two-phase training approach: initial "forget training" to actively unlearn specific trajectories, followed by "convergence training" to stabilize the agent and preserve performance on the remaining data. These phases must be executed separately for theoretical convergence guarantees.
- Significant efficiency gains are achieved: TrajDeleter performs unlearning with results comparable to full retraining while using only 2.2% of the time required for a complete model rebuild.
- A novel auditing mechanism based on Membership Inference Attacks (MIA) is introduced to verify unlearning success, cleverly leveraging the inherent instability of offline RL training to generate diverse shadow agents for comparison.
- Convergence training is indispensable: Without it, the agent inadvertently forgets useful information from the remaining dataset, leading to a substantial decrease in overall performance, underscoring its role in maintaining agent stability and efficacy.
- The work provides foundational insights for future large language models (LLMs) that utilize Reinforcement Learning from Human Feedback (RLHF), suggesting a pathway for implementing unlearning capabilities in these increasingly powerful and ubiquitous AI systems.
About the Speaker(s)
The speaker for this presentation was Chen Gong, affiliated with Kung Fu University Virginia. His research focuses on addressing critical challenges in machine learning, particularly in the domain of offline reinforcement learning. As evidenced by the detailed technical content of the talk, Chen Gong's work delves into the theoretical and practical aspects of data governance and privacy within AI systems, specifically developing innovative solutions like TrajDeleter to enable the "right to be forgotten" for complex learning agents. His insights into the instability of RL training and the design of verifiable unlearning mechanisms highlight his expertise in secure and privacy-preserving machine learning.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic ML research solving a real problem — selective trajectory unlearning in offline RL with a two-phase approach and MIA-based auditing. Solid contribution to a niche but growing area, with honest acknowledgment of limitations. Not a security talk in any traditional sense, but it belongs at a venue like NDSS where the privacy/ML intersection lives.
Heather Calloway (CISO) — PASS
Technically earnest research on machine unlearning for offline RL agents — but this is a scope call, not a critique. There is no governance angle, no operator relevance, and nothing here that informs how a security program, a board, or a regulator acts. The GDPR framing is aspirational decoration, not a real bridge to institutional accountability.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025