RAIFLE: Reconstruction Attacks on Interaction-based Federated Learning with Adversarial Data Manipulation
Dzung Pham
Network and Distributed System Security (NDSS) Symposium 2025 · Day 2 · Federated Learning 1
Overview
Dzung Pham from the University of Massachusetts Amherst presented "RAIFLE: Reconstruction Attacks on Interaction-based Federated Learning with Adversarial Data Manipulation" at the NDSS Symposium. This groundbreaking work introduces a novel class of privacy attacks specifically targeting interaction-based Federated Learning (FL) systems, a specialized form of FL particularly relevant to recommendation and ranking systems. RAIFLE demonstrates how a malicious server, by subtly manipulating the items presented to users, can significantly enhance its ability to reconstruct sensitive user interaction data.
Key moments
- 0:00 Introduction to RAIFLE and motivation for privacy in recommendation systems
- 1:17 Explaining interaction-based federated learning (IBFL) and its properties
- 2:55 Baseline gradient inversion attack in honest-but-curious setting
- 4:13 RAIFLE attack idea: manipulating item features for better reconstruction
- 5:36 RAIFLE's effective noise injection method for feature manipulation
- 6:10 RAIFLE consistently outperforms vanilla gradient inversion in evaluations
- 6:44 Real-world examples validating RAIFLE's threat model practicality
- 8:04 Addressing scenarios where server cannot directly control features
RAIFLE: Reconstruction Attacks on Interaction-based Federated Learning with Adversarial Data Manipulation
Speakers: Dzung Pham, NDSS Internet Society Fellow, University of Massachusetts Amherst
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=dw9EPrSTiTw
Overview
Dzung Pham from the University of Massachusetts Amherst presented "RAIFLE: Reconstruction Attacks on Interaction-based Federated Learning with Adversarial Data Manipulation" at the NDSS Symposium. This groundbreaking work introduces a novel class of privacy attacks specifically targeting interaction-based Federated Learning (FL) systems, a specialized form of FL particularly relevant to recommendation and ranking systems. RAIFLE demonstrates how a malicious server, by subtly manipulating the items presented to users, can significantly enhance its ability to reconstruct sensitive user interaction data.
The motivation behind RAIFLE stems from the pervasive nature of recommendation and ranking systems in modern digital life, from social networks and e-commerce to entertainment and search engines. These systems are trained on vast amounts of sensitive user information and are designed to encourage further data sharing. While Federated Learning has emerged as a promising privacy-preserving solution for such applications, real-world deployments in this specific interaction-based setting are still nascent. RAIFLE highlights the critical importance of proactively identifying and investigating potential vulnerabilities in these systems before widespread adoption, demonstrating that even with privacy-enhancing technologies, novel attack vectors can emerge from unique architectural properties.
Pham's research underscores a significant gap in existing FL security analyses, which often focus on the "honest-but-curious" server model or assume server manipulation is limited to model weights. RAIFLE, however, explores the more potent threat of a malicious server exploiting its inherent control over the input items presented to users, a capability unique to interaction-based FL. The findings reveal that this control can be leveraged to craft highly effective gradient inversion attacks, capable of extracting private user interaction data with high fidelity, even in the presence of common FL defenses like Differential Privacy and Secure Aggregation.
Background
▶ Watch: Introduction to RAIFLE and motivation for privacy in recommendation systems (0:00)
Recommendation and ranking systems are ubiquitous, forming the backbone of user experience on platforms like Facebook, Amazon, Netflix, and Google. These systems are inherently data-intensive, relying on massive datasets of user interactions, preferences, and behaviors, which often contain highly sensitive personal information. Given the privacy concerns associated with centralizing such vast amounts of data, Federated Learning (FL) has been proposed as a privacy-preserving paradigm. In FL, users' raw data remains on their devices, and only model updates (gradients or parameters) are shared with a central server, which then aggregates these updates to improve a global model.
However, the specific context of recommendation and ranking systems often involves a particular FL variant termed interaction-based Federated Learning. Unlike traditional FL where the server has no control over user data, in interaction-based FL, the server actively prepares and presents items (e.g., news articles, products, songs, search results) for users to interact with. Users then perform local training based on these presented items and their private interactions (e.g., clicks vs. no-clicks, upvotes vs. downvotes). The core privacy challenge here is to protect these interaction data points from the server, even though the server dictates the items users interact with.
Prior research on FL privacy has extensively explored gradient inversion attacks in the "honest-but-curious" server model. In such attacks, the server attempts to reconstruct users' private training data (e.g., images, labels) by analyzing the model updates (gradients) received from clients. This is typically achieved by simulating the user's local training process and optimizing a reconstruction to match the observed gradients. Common defenses against these attacks include Differential Privacy (DP), which adds noise to model updates to obscure individual contributions, and Secure Aggregation (SA), which cryptographically combines updates from multiple users so that the server only sees the aggregate without individual user linkages.
While these defenses mitigate some risks, the specific properties of interaction-based FL introduce new vulnerabilities. Existing malicious server attacks often focus on manipulating model weights, which can be effective but frequently assume specific model architectures (e.g., Convolutional Neural Networks). For simpler models commonly used in recommendation and ranking (e.g., linear models or small neural networks), such attacks might be less applicable. The critical unexplored area, prior to RAIFLE, was whether a malicious server could exploit its unique control over the items presented to users to facilitate or enhance reconstruction attacks. This control, inherent to interaction-based FL, represents a powerful, yet previously unaddressed, attack surface.
Key Findings
▶ Watch: Baseline gradient inversion attack in honest-but-curious setting (2:55)
RAIFLE's research reveals several critical findings about the vulnerabilities of interaction-based Federated Learning and the potency of a malicious server leveraging its control over presented items:
- Baseline Gradient Inversion is Already Potent: Even vanilla gradient inversion, without the proposed RAIFLE manipulations, proved surprisingly effective. Evaluating on two recommendation datasets (MovieLens and Steam) using the Federated Neuro Collaborative Filtering algorithm, vanilla gradient inversion achieved an F1 score of over 0.9. This significantly outperformed a random discrete search baseline from prior work (Do et al., 2023), indicating that even basic gradient inversion poses a substantial threat. Furthermore, vanilla gradient inversion could still perform non-trivially even when baseline defenses were applied.
- Novel Attack Vector: Item Feature Manipulation: The core innovation of RAIFLE is the discovery that a malicious server can dramatically improve reconstruction rates by manipulating the features or representations of the items it presents to users. The key insight is to modify item features in a way that makes their contribution to the user's local model update (gradient) as unique and discernible as possible, thereby simplifying the reconstruction task for the server.
- Effective Manipulation Strategy: Noise Injection: While an initial strategy of assigning unique feature subsets to items proved too obvious and did not scale well, RAIFLE found that simply injecting random noise into item feature values was highly effective. This method achieves the goal of making each item's contribution unique and scales robustly with an increasing number of items. Across evaluations on the MQ2008 and NSR learning-to-rank datasets, using the Federated Pairwise Differentiable Gradient Descent algorithm and AUC as a metric, RAIFLE consistently outperformed vanilla gradient inversion in all tested settings, demonstrating superior scalability with a higher number of items.
- Real-World Threat Model Validation: A crucial finding is the practicality of RAIFLE's threat model, where a server actively manipulates user-facing content. Pham presented compelling historical examples:
- Facebook (2014): Manipulated news feeds for nearly 700,000 users in the "emotion contagion" study.
- Amazon (2021): Fined for manipulating product recommendation results.
- Google (recently): Fined over €2 billion for manipulating search results.
These instances highlight that large tech companies have historically exercised and abused their control over presented content, making RAIFLE's malicious server assumption highly realistic.
- Extending Attacks to Indirect Feature Control: RAIFLE successfully extended its attack methodology to scenarios where the server cannot directly control the training features (e.g., image ranking where users perform local feature extraction using pre-trained computer vision models). The finding here is that the server can still achieve its goal by manipulating the raw input data (e.g., images) so that the extracted features by the user's local model resemble random noise. This is achieved through an optimization process.
- Stealthy Input Manipulation: The image manipulation technique developed for indirect feature control proved remarkably subtle. Manipulated images exhibited "very small slight discoloration," which was "not very noticeable," showcasing the stealthiness of the attack.
- Robustness Against Defenses (with caveats):
- Local Differential Privacy (LDP): While LDP significantly reduces reconstruction performance, RAIFLE still managed to "squeeze as much reconstruction performance as possible," outperforming vanilla gradient inversion even with LDP applied.
- Secure Aggregation (SA): RAIFLE can be modified to work with SA. By employing a "fingerprinting" approach, where a small subset of the feature space is targeted for a particular user, the attack remains effective for a considerable number of participants (up to hundreds) with appropriately chosen LDP noise levels.
In summary, RAIFLE fundamentally shifts the understanding of privacy risks in interaction-based FL, demonstrating that a malicious server's control over presented items is a potent new attack surface that can be leveraged for highly effective and often stealthy reconstruction attacks, even against established defenses.
Technical Deep Dive
▶ Watch: RAIFLE's effective noise injection method for feature manipulation (5:36)
RAIFLE's technical contributions lie in its innovative approach to gradient inversion in the specific context of interaction-based Federated Learning, where the server plays an active role in dictating user inputs. The attack capitalizes on the server's ability to manipulate the items presented to users, either directly affecting their features or indirectly influencing features through input manipulation.
Interaction-Based Federated Learning Model
The FL setting considered by RAIFLE operates as follows:
- Server Prepares Items: The central server selects or generates a set of items (e.g., products, articles, images) to be presented to users. These items have associated features or representations.
- Items Sent to Users: These items are sent to individual users (clients).
- User Interaction: Users interact with the presented items (e.g., click, upvote, view). These interactions form the sensitive private data.
- Local FL Training: Each user performs local training on a machine learning model using the presented items and their private interaction data. This training generates a local model update (e.g., gradients).
- Server Aggregation: The server collects these local model updates from multiple users and aggregates them to improve a global model.
The goal of the RAIFLE attack is for a malicious server to reconstruct the private interaction data of individual users by exploiting its control over the items presented in step 2.
Baseline: Vanilla Gradient Inversion
As a baseline, RAIFLE first evaluated a standard gradient inversion attack. In this scenario, an "honest-but-curious" server receives model updates (gradients) from users. The server then attempts to reverse-engineer the users' private interaction data by simulating the local training process. It initializes a dummy set of interaction data and iteratively optimizes this dummy data to minimize the difference between the gradients it would produce and the actual gradients received from the user.
Pham's team evaluated this on two recommendation datasets: MovieLens and Steam, targeting the Federated Neuro Collaborative Filtering algorithm. Compared to a random discrete search baseline (Do et al., 2023), vanilla gradient inversion achieved an F1 score greater than 0.9 on both datasets, demonstrating its inherent effectiveness. Even with baseline defenses, its performance remained non-trivial.
RAIFLE's Core Strategy: Malicious Item Manipulation
RAIFLE introduces the concept of a malicious server that actively manipulates the items it sends to users. The central hypothesis is that if the server can modify the features of these items, it can make the resulting gradients more "transparent" or unique, thereby simplifying the reconstruction of user interactions.
1. Direct Feature Manipulation (e.g., textual recommendations)
When the server has direct control over the features of the items it presents (e.g., a recommendation system where item features are numerical vectors that the server can modify before sending), RAIFLE explores two strategies:
- Unique Feature Subsets: An initial idea was to assign a unique, non-overlapping subset of features to each item. For example, if items have three features, item 1 gets
[F1, 0, 0], item 2 gets[0, F2, 0], etc. This aims to ensure that each item uniquely influences specific model parameters. However, this approach was found to be "very obvious" (detectable by many zeros in the data) and did not scale well, especially when the number of items exceeded the number of model parameters or features.
- Noise Injection into Feature Values: The most effective strategy discovered by RAIFLE is to inject random noise into the feature values of the items. Instead of creating sparse, unique subsets, the server modifies the existing feature values by adding a carefully crafted random perturbation. This seemingly counter-intuitive approach works by essentially "unifying" or making the contribution of each item to the gradient more distinct and less entangled with other item features, facilitating easier reconstruction.
- Evaluation: This method was tested on learning-to-rank datasets MQ2008 and NSR (collected by Microsoft), targeting the Federated Pairwise Differentiable Gradient Descent algorithm. Using the Area Under the ROC Curve (AUC) as a metric, RAIFLE consistently and significantly outperformed vanilla gradient inversion across all tested settings, demonstrating its scalability as the number of items increased.
2. Indirect Feature Manipulation via Input Alteration (e.g., image ranking)
A more challenging scenario arises when the server cannot directly control the features used for training. For example, in an image ranking system, the server sends images to users. Users then use a pre-trained computer vision (CV) model (e.g., ResNet, DenseNet) as a feature extractor locally to convert images into feature vectors. The server only receives gradients based on these extracted features, not the raw images.
RAIFLE addresses this by developing a technique to manipulate the input images such that the extracted features by the user's local model resemble random noise as much as possible, effectively achieving the same goal as direct noise injection into features.
- Optimization-Based Image Manipulation:
- Start with an initial image.
- Pass it through the known feature extractor to obtain an initial feature vector.
- Initialize a random noise vector as the target for the extracted features.
- Minimize the loss between the actual extracted feature vector (from the manipulated image) and the random noise target vector. This optimization is performed with respect to the initial image, iteratively adjusting the image pixels.
- Challenge with High Dimensions: For high-dimensional feature vectors (e.g., 1,000-2,000 features), direct optimization towards a single random noise target proved ineffective.
- RAIFLE's Partitioning Trick: To overcome the high-dimensionality challenge, RAIFLE introduces a clever partitioning strategy. Instead of a single noise target, the target noise vector is partitioned into multiple sub-targets. Each sub-target has Gaussian noise for a specific subset of its dimensions, with the rest set to zeros. For example, with two targets, one might have noise in the first half of the dimensions and zeros in the second, while the other has zeros in the first half and noise in the second.
- The optimization process is then run for each sub-target, resulting in multiple variations of the manipulated image.
- When the server sends images to users, it randomly selects one of these manipulated image variations.
- Stealthiness: This manipulation results in images with "very small slight discoloration," making the changes difficult for a human eye to detect.
- Evaluation: This method was validated using the ImageNet dataset with four different feature extractors (ResNet, DenseNet, etc.). RAIFLE was compared against vanilla gradient inversion and the classic Fast Gradient Sign Method (FGSM). RAIFLE consistently outperformed both baselines across numerous settings and scaled effectively with an increasing number of images relative to feature vector dimensions.
Defenses and RAIFLE's Resilience
RAIFLE also investigated the effectiveness of common FL defenses against its attacks:
- Local Differential Privacy (LDP): As expected, LDP significantly reduces the reconstruction performance. However, RAIFLE demonstrated that it could still "squeeze as much reconstruction performance as possible" compared to vanilla gradient inversion, indicating its robustness even under LDP.
- Secure Aggregation (SA): Secure Aggregation aims to prevent the server from linking model updates to individual users. To circumvent SA, RAIFLE proposes a modified attack strategy akin to fingerprinting. The server targets a small subset of users and partitions a specific portion of the feature space to each target user. This allows the server to effectively reconstruct interactions for those targeted users. RAIFLE found this approach to be effective for a "number of participants in the 100s range" when combined with "small enough" LDP noise, suggesting that SA alone is not a complete solution against a determined malicious server.
Other potential defenses mentioned include checking for manipulation (though RAIFLE's image manipulations are subtle), minimizing shared information, and changing the FL architecture.
Demo / Proof of Concept
▶ Watch: RAIFLE consistently outperforms vanilla gradient inversion in evaluations (6:10)
While the talk did not feature a live, interactive demonstration in the traditional sense, Dzung Pham presented compelling experimental results and visual evidence that served as a robust proof of concept for the RAIFLE attacks. The effectiveness of the attack was quantified through standard machine learning metrics such as F1 score (for direct feature manipulation, achieving >0.9 on specific datasets) and Area Under the ROC Curve (AUC) (for both direct and indirect feature manipulation, consistently outperforming baselines).
The presentation included visual examples of the manipulated images created through the indirect feature manipulation technique. Although the speaker noted that the subtle changes might be difficult to discern from a distance, they highlighted that the manipulation resulted in "very small slight discoloration," demonstrating the stealthy nature of the attack.
Furthermore, the speaker explicitly stated that the code for RAIFLE is publicly available on GitHub, allowing other researchers and practitioners to reproduce the results, verify the claims, and further investigate the vulnerabilities. This commitment to open science serves as a strong form of proof of concept, enabling independent validation of the research findings.
Defensive Implications
▶ Watch: Addressing scenarios where server cannot directly control features (8:04)
The RAIFLE research presents significant defensive implications for the design and deployment of privacy-preserving systems, particularly those leveraging interaction-based Federated Learning for recommendation and ranking. The core message for defenders is that the server's control over the items presented to users is a critical and potent attack surface that must be explicitly addressed.
- Re-evaluate Threat Models: Traditional FL threat models often assume an "honest-but-curious" server or limit malicious actions to model weight manipulation. RAIFLE mandates expanding these threat models to include a malicious server actively manipulating input items. This malicious capability is not theoretical, as historical examples demonstrate that powerful entities have indeed manipulated content presented to users.
- Strengthen Input Integrity Checks:
- Direct Feature Manipulation: If the FL architecture allows the server to directly specify or modify item features before sending them to users, robust integrity checks are paramount. Users (clients) should verify that the features received are legitimate and have not been tampered with in a way that could facilitate gradient inversion. This might involve cryptographic signing of features by a trusted third party, or client-side anomaly detection on item feature vectors.
- Indirect Feature Manipulation: When users perform local feature extraction (e.g., from images), the server's ability to manipulate the raw input (images) to produce noisy features is a concern. Clients might need mechanisms to detect subtle manipulations in the received content. While RAIFLE's image manipulations are stealthy, research into robust, client-side adversarial example detection could be crucial. Alternatively, clients could rely on trusted sources for input content where manipulation is less likely.
- Enhance Existing Defenses Against Targeted Attacks:
- Differential Privacy (DP): While LDP reduces RAIFLE's effectiveness, it doesn't eliminate it entirely. Defenders should consider implementing stronger DP mechanisms or exploring adaptive DP strategies that can better withstand targeted attacks like RAIFLE. The trade-off between privacy and utility remains a challenge.
- Secure Aggregation (SA): RAIFLE's ability to circumvent SA via a "fingerprinting" approach for a subset of users is alarming. This suggests that SA, while effective against linking individual updates, doesn't prevent a malicious server from targeting specific users if it can control their inputs. Future SA designs might need to incorporate mechanisms that randomize or obfuscate the specific item features associated with aggregated gradients, making targeted reconstruction more difficult.
- Explore Novel FL Architectures: The findings encourage research into FL architectures that inherently minimize the server's ability to influence item features or inputs. This could involve:
- Decentralized Item Selection: Allowing clients to independently select items from a trusted pool, reducing server control.
- Zero-Knowledge Proofs: Using ZKPs to prove that item features were not manipulated, without revealing the features themselves.
- Trusted Execution Environments (TEEs): Leveraging TEEs on the client side to ensure that feature extraction and local training occur in an uncompromised environment, even if the input item is subtly manipulated.
- Minimize Shared Information: A general principle for defense is to minimize the amount and specificity of information shared with the server. Reviewing what information is strictly necessary for model aggregation and exploring ways to share less (e.g., using secure multi-party computation for certain operations) can help.
In conclusion, RAIFLE serves as a stark warning that privacy in FL is a moving target. Defenders must anticipate and proactively address novel attack vectors arising from the specific characteristics of FL deployments, especially when the central server possesses significant control over user-facing content.
Key Takeaways
- Novel Attack Surface: RAIFLE introduces a critical new attack surface in interaction-based Federated Learning: the malicious server's control over items presented to users.
- Potent Reconstruction: By manipulating item features (e.g., injecting random noise) or raw inputs (e.g., subtly altering images), a malicious server can significantly enhance its ability to reconstruct sensitive user interaction data.
- Realistic Threat Model: The threat model is validated by real-world instances of major tech companies manipulating content presented to users, underscoring the practical relevance of RAIFLE.
- Resilience Against Defenses: RAIFLE attacks can still achieve high reconstruction performance even in the presence of defenses like Local Differential Privacy (LDP) and can be adapted to bypass Secure Aggregation (SA) through targeted "fingerprinting" techniques.
- Urgent Need for Stronger Defenses: Defenders must re-evaluate existing FL security assumptions, implement robust input integrity checks, and explore novel FL architectures that minimize server control over user-facing content.
- Stealthy Manipulation: The ability to manipulate images with "very small slight discoloration" highlights the stealthy nature of these attacks, making them difficult for human users to detect.
About the Speaker(s)
Dzung Pham is a researcher from the University of Massachusetts Amherst. He is also recognized as an NDSS Internet Society Fellow, indicating his expertise and contributions to the field of network and distributed system security. His work, as demonstrated by the RAIFLE paper, focuses on critical privacy and security challenges within emerging machine learning paradigms like Federated Learning.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid, original academic security research targeting a genuinely underexplored attack surface in federated learning. The core insight — that a malicious server's control over presented items is itself an exploitation primitive — is clean and non-obvious, and the work holds up technically across multiple datasets and model types. Not flashy, but this is exactly the kind of principled threat modeling that should be done before interaction-based FL gets baked into production recommendation systems at scale.
Heather Calloway (CISO) — WEAK
Technically rigorous academic work that surfaces a real and underexplored vulnerability in interaction-based federated learning. But it stops at the research boundary — no clear path for an operator, platform security leader, or regulator to act on it.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025