A Method to Facilitate Membership Inference Attacks in Deep Learning Models
Zitao Chen
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Membership Inference
Overview
In an era increasingly reliant on machine learning, the privacy of training data has become a paramount concern. This talk by Zitao Chen at the NDSS Symposium introduces a groundbreaking and stealthy method to facilitate membership inference attacks (MIA) against deep learning models. Unlike previous attacks that often degrade model utility, making them easily detectable, this novel approach achieves "extremely high privacy leakage" while maintaining the model's normal performance, effectively bypassing the long-standing privacy-utility trade-off. The research highlights a critical, often overlooked, vulnerability stemming from the use of untrusted machine learning codebases, which are prevalent on platforms like GitHub and Hugging Face.
Key moments
- 0:00 Introduction: Untrusted ML codebase and privacy risks
- 2:45 Limitations of prior attacks: Poor privacy-utility trade-off
- 4:00 New approach: Divide and conquer with secret samples
- 5:05 Encoding membership with crafted noisy secret samples
- 6:40 Identifying the challenge: Normalization functions with mixed data
- 7:30 Solution: Secondary normalization function for secret samples
- 8:10 Attack results: High leakage, preserved utility, and auditing evasion
A Method to Facilitate Membership Inference Attacks in Deep Learning Models
Speakers: Zitao Chen
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=4KDWXgbQ_RY
Overview
In an era increasingly reliant on machine learning, the privacy of training data has become a paramount concern. This talk by Zitao Chen at the NDSS Symposium introduces a groundbreaking and stealthy method to facilitate membership inference attacks (MIA) against deep learning models. Unlike previous attacks that often degrade model utility, making them easily detectable, this novel approach achieves "extremely high privacy leakage" while maintaining the model's normal performance, effectively bypassing the long-standing privacy-utility trade-off. The research highlights a critical, often overlooked, vulnerability stemming from the use of untrusted machine learning codebases, which are prevalent on platforms like GitHub and Hugging Face.
The core innovation lies in a "divide and conquer" strategy that indirectly steals information. Instead of directly manipulating the model's memorization of sensitive training data, the attacker injects and exploits the model's propensity to memorize specially crafted "secret samples." These secret samples are designed to encode the membership status of the actual training data, allowing an attacker with blackbox access to the final model to infer whether a specific data point was part of the training set. This indirect mechanism ensures that the model's primary learning task remains uncompromised, leading to high utility, while simultaneously enabling severe privacy breaches.
The implications of this work are profound. It not only demonstrates a new, highly effective class of membership inference attacks but also exposes a significant blind spot in current privacy auditing methodologies. Existing auditing tools, which typically inspect model outputs on training data, are rendered ineffective because the leakage occurs through an unconventional channel—the secret samples. This research serves as a stark warning about the hidden risks associated with the ML supply chain and calls for a fundamental re-evaluation of how we assess and protect privacy in machine learning models.
Background
▶ Watch: Introduction: Untrusted ML codebase and privacy risks (0:00)
The proliferation of deep learning models across various domains has brought the issue of data privacy to the forefront. One of the most significant privacy threats is the membership inference attack (MIA), where an adversary aims to determine whether a specific data point was part of a model's training dataset. Such information can be highly sensitive, potentially revealing personal medical records, financial transactions, or other confidential user data, thereby violating privacy regulations and trust.
Historically, membership inference attacks have been a subject of extensive research, with many approaches focusing on exploiting a model's tendency to "memorize" training data. Models often perform differently on data they have seen during training versus unseen data, and MIAs leverage these observable differences. However, a common thread among state-of-the-art MIAs has been a significant privacy-utility trade-off. To increase the success rate of an attack (i.e., maximize privacy leakage), prior methods often involved manipulating the model's training process to increase its memorization. This direct manipulation, however, typically leads to a degradation in the model's overall performance or utility. A model with noticeably lower accuracy or altered behavior would likely be flagged by users or developers, making such attacks easily detectable and less practical for real-world adversaries.
The problem is exacerbated by the modern machine learning development ecosystem. Many developers, seeking to reduce the barrier to entry and accelerate development, routinely download and utilize public codebases available on platforms like GitHub and Hugging Face. These codebases, often contributed by third-party developers, are integrated into training pipelines without extensive scrutiny. This practice, while convenient, introduces a critical vulnerability: the ML supply chain attack. Malicious actors can masquerade as legitimate developers and inject malicious code into these widely used libraries or model architectures. Real-world examples have shown the feasibility of code-poisoning attacks affecting common machine learning libraries such as PyTorch and TensorFlow, as well as commercial software.
In this context, Zitao Chen's work addresses a crucial question: What are the privacy risks when using untrusted ML codebases to train models, especially when sensitive user data is involved? The scenario considered is one where privacy-sensitive users train their models in an isolated environment, meaning the attacker cannot directly observe the training process. The only outcome accessible to the attacker is the final ML model, to which they have blackbox access. The challenge then becomes: how can an attacker, by manipulating the codebase, increase privacy leakage against the training data in this restricted, blackbox setting, without making the attack obvious through reduced model utility or detection by existing auditing tools? This study aims to overcome the "undecidable trade-off" between privacy and utility that plagued prior attack methodologies, paving the way for more potent and stealthy MIAs.
Key Findings
▶ Watch: New approach: Divide and conquer with secret samples (4:00)
The research presents several critical findings that collectively redefine the landscape of membership inference attacks and machine learning privacy:
- Extreme Privacy Leakage without Utility Compromise: The most significant finding is the development of a membership inference attack that can inflict "extremely high privacy leakage" – effectively compromising the privacy of all training data points – while simultaneously allowing the models to "maintain normal performance on the main tasks." This directly addresses and overcomes the long-standing privacy-utility trade-off, a fundamental limitation of previous MIA methodologies. The ability to steal sensitive information without degrading model utility makes the attack significantly harder for users to notice or detect.
- Evasion of State-of-the-Art Privacy Auditing Tools: The novel attack method is designed to steal information in an "unconventional way," specifically through a set of "secret samples" rather than direct manipulation of the model's output on the training data. Consequently, the study demonstrates that "all existing [privacy auditing] methods" fail to reliably measure the leakage caused by these compromised models. From a user's perspective, the privacy leakage of an attacked model is "still indistinguishable from the uncompromised ones," even under the scrutiny of advanced auditing tools. This highlights a critical vulnerability in current privacy assessment practices.
- Feasibility of Stealthy Supply Chain Attacks: The research confirms that untrusted ML codebases can indeed "engender hidden privacy risks" when building ML models. By demonstrating how a malicious party can inject specific code (e.g., related to loss computation or model structure) to facilitate this stealthy attack, the study underscores the severe supply chain vulnerabilities inherent in the common practice of downloading and using third-party ML code. The attacker's ability to operate under a blackbox access model further emphasizes the practical threat.
- Novel "Divide and Conquer" Attack Strategy: The core innovation is a "new direction to construct membership inference attacks" based on a "divide and conquer approach." This strategy indirectly encodes the membership of training data via auxiliary "secret samples" rather than directly manipulating the model's memorization of the original training data. This indirectness is key to preserving model utility and evading detection, representing a paradigm shift in MIA design.
These findings collectively point to a new generation of sophisticated and covert privacy attacks, necessitating a fundamental re-evaluation of machine learning security practices, particularly concerning supply chain integrity and privacy auditing.
Technical Deep Dive
▶ Watch: Encoding membership with crafted noisy secret samples (5:05)
The technical ingenuity of this new membership inference attack lies in its "divide and conquer" strategy, which allows for indirect information encoding and mitigates the performance degradation commonly associated with MIAs. The attack operates under the assumption that the attacker has injected malicious code into an ML codebase, which is then used by a privacy-sensitive user to train a model in an isolated environment. The attacker only has blackbox access to the final trained model.
The first step of the attack is to encode the membership of the training data via some secret samples. Unlike prior attacks that directly manipulate the model's output on the training data, this method exploits the model's propensity to memorize additional secret samples. These secret samples are generated by the malicious training algorithm and are uniquely associated with each original training data point. Both the legitimate training samples and these secret samples are used to optimize the model. Consequently, if a specific secret sample was used during training, it implies that its corresponding training sample was also used. The attacker's goal then shifts: instead of directly inferring the membership of training data, they infer the membership of these secret samples.
To make this indirect inference easy for the attacker, the secret samples are deliberately crafted as random noisy data. Deep learning models are known to be very good at memorizing "allied data" or outliers, such as random noise, especially when these data points do not contain discernible features relevant to the main task. If such noisy data are present during training, the model essentially "memorizes" them rather than learning generalizable features from them. This characteristic makes it easy for an attacker to determine whether a specific secret sample was included in the training by observing the model's behavior (e.g., confidence scores, loss values) on that sample post-training. By successfully inferring the membership of a secret sample, the attacker directly infers the membership of the corresponding privacy-sensitive training data point.
However, a critical challenge arises when mixing ordinary training data with these specially crafted random noisy secret samples. Standard deep learning architectures employ normalization functions (e.g., BatchNorm, LayerNorm) that typically "expect the data to come from a similar distribution." Mixing disparate data types – clean, structured training data alongside random, noisy secret samples – would cause the model to use "wrong statistics for normalization." This misapplication of normalization statistics would inevitably lead to "inferior performance" on the main task, making the attack noticeable and negating the stealth advantage.
To overcome this, the researchers employ a second "divide and conquer" approach specifically for the normalization problem. The core idea is to use a secondary normalization function to separately process the secret samples.
- The original normalization functions within the model architecture are used to process the legitimate training data, ensuring the model learns effectively on its primary task.
- An additional, distinct normalization function is introduced and dedicated solely to handling the secret samples.
This architectural modification allows the model to treat the secret samples separately during the normalization process, preventing their unusual distribution from negatively impacting the statistics derived from the legitimate training data. As a result, the model can still effectively learn from the training data and maintain high utility, while simultaneously memorizing the secret samples for the attacker's inference. The attacker's malicious code needs to manipulate two key components within the codebase: the loss computation functions to incorporate the secret samples and their unique association, and the model structure to introduce and appropriately route the secret samples through the secondary normalization function. This indirect, architecturally subtle manipulation is the key to achieving high privacy leakage without compromising model performance or triggering existing privacy auditing tools.
Demo / Proof of Concept
▶ Watch: Solution: Secondary normalization function for secret samples (7:30)
While the talk does not describe a live, interactive "demo" in the traditional sense of a software demonstration, it presents compelling quantitative results that serve as a robust proof of concept for the proposed attack method. The speaker refers to figures and results from their study, illustrating the efficacy and stealth of their approach.
The evaluation methodology followed "common practice" in membership inference attack research, using metrics such as the attack true positive rate at a low false positive rate. This metric is crucial because it measures the attacker's ability to correctly identify members (true positives) while minimizing incorrect identifications of non-members (false positives), reflecting a realistic and effective attack scenario.
The results presented show a dramatic improvement over prior art:
- Unprecedented Leakage: The attack is capable of effectively "compromise[ing] the privacy of all data." This indicates a near-100% success rate in identifying whether any given data point was part of the training set, representing the "worst case price leakage" from a privacy perspective. This level of comprehensive leakage is a significant leap beyond previous attacks, which often struggled to achieve high true positive rates without also incurring high false positive rates or severely impacting model utility.
- Preservation of Utility: Crucially, the attacked models are shown to "still enable the models to maintain normal performance on the main tasks." This means that despite the extensive privacy leakage, the model's accuracy, F1-score, or other task-specific performance metrics remain comparable to an uncompromised model. This stealth characteristic is what makes the attack "much harder to notice by the users," as there are no obvious performance anomalies to raise suspicion.
- Evasion of Auditing Tools: The talk further elaborates on the attack's stealth by demonstrating its ability to bypass existing privacy auditing tools. The speaker explains that "all existing methods work by directly inspecting the model's output on the training data." However, since this attack "stolen [information] by a set of secret samples," the observable outputs related to the original training data remain largely unchanged. Consequently, "from users perspective, the privacy leakage of the compromised models are still indistinguishable from the uncopty ones" when assessed by state-of-the-art auditing techniques. This lack of detectability by current defense mechanisms underscores the sophistication and danger of this new attack vector.
These findings, presented with quantitative evidence in the full paper (as implied by the "figure below" and "results" mentioned in the talk), provide a strong proof of concept for the method's effectiveness in real-world scenarios where an attacker has blackbox access to the final model after injecting malicious code into the training pipeline.
Defensive Implications
▶ Watch: Attack results: High leakage, preserved utility, and auditing evasion (8:10)
This research unveils a sophisticated and stealthy class of membership inference attacks that poses significant challenges for current machine learning security and privacy practices. The defensive implications are multi-faceted and demand immediate attention across the ML ecosystem:
- Re-evaluate Privacy Auditing Methodologies: The most direct implication is the obsolescence of existing privacy auditing tools against this type of attack. Current methods primarily focus on inspecting direct outputs of the model on training data. Since this new attack mechanism leverages "secret samples" for indirect information encoding, traditional auditing techniques are rendered ineffective. There is an urgent need for the ML community to "reconsider our current practice of privacy auditing in machine learning" and develop new, more comprehensive auditing frameworks. These frameworks must be capable of detecting unconventional information leakage channels, potentially by analyzing model architectures for unusual components (like secondary normalization functions), or by detecting the presence and memorization of "allied data" or "random noisy data" that should not be present in a healthy training process.
- Strengthen ML Supply Chain Security: The attack's premise hinges on the use of "untrusted ML codebase" from public repositories such as GitHub and Hugging Face. This highlights a critical vulnerability in the ML supply chain. Developers and organizations must adopt more rigorous security practices when incorporating third-party code. This includes:
- Code Auditing: Implementing static and dynamic analysis tools to scan downloaded code for suspicious modifications, especially in core components like loss functions and model architecture definitions.
- Dependency Verification: Ensuring the integrity and authenticity of all dependencies, potentially through cryptographic signing or trusted registries.
- Sandboxing: Training models in highly isolated environments that restrict network access and monitor unusual resource consumption or file system access patterns, even if the model itself is the only outcome accessible to an external party.
- Develop Novel Detection Mechanisms for Indirect Leakage: Defenders need to move beyond output-based auditing and explore mechanisms that can detect the intent or mechanism of indirect information leakage. This could involve:
- Architectural Anomaly Detection: Monitoring for unexpected additions or modifications to standard neural network architectures, such as the introduction of separate normalization layers for specific data types that don't align with the model's intended function.
- Data Distribution Analysis: Techniques that can analyze the internal representations learned by the model for signs of memorization of anomalous data (like noisy secret samples) that are statistically distinct from the legitimate training data, even if the model's final performance is unaffected.
- Differential Privacy (DP) Re-evaluation: While not explicitly discussed as a defense, the attack raises questions about the robustness of DP guarantees against such sophisticated architectural manipulations. Future work might explore whether existing DP implementations are truly resilient when the underlying training code itself is malicious.
- Increase Awareness and Education: Developers and ML practitioners must be made aware of these advanced attack vectors. Understanding that "normal model performance" is no longer a sufficient indicator of privacy preservation is crucial. Emphasizing the hidden risks of untrusted code and promoting secure development lifecycles for ML systems are essential steps.
In essence, this research signals a shift in the arms race between attackers and defenders in ML privacy. Defenders can no longer rely on superficial performance checks or simplistic auditing tools; a deeper, more architectural, and supply-chain-aware approach to security is now imperative.
Key Takeaways
- Hidden Privacy Risks in ML Supply Chain: The common practice of using untrusted ML codebases from platforms like GitHub and Hugging Face introduces significant, hidden privacy risks due to potential malicious code injection.
- Overcoming the Privacy-Utility Trade-off: A novel membership inference attack (MIA) method can achieve "extremely high privacy leakage" (compromising all training data) while simultaneously maintaining "normal model performance" (high utility), overcoming a critical limitation of prior MIAs.
- Indirect Information Encoding: The attack leverages a "divide and conquer" strategy, encoding training data membership indirectly via specially crafted "secret samples" (e.g., random noisy data) that the model is induced to memorize.
- Stealth via Secondary Normalization: To prevent utility degradation, the attack introduces a "secondary normalization function" dedicated to processing the secret samples, allowing the legitimate training data to be processed normally.
- Evasion of Current Auditing Tools: Existing privacy auditing methods are ineffective against this attack because they inspect direct model outputs on training data, whereas the leakage occurs through the unconventional channel of secret samples, making compromised models "indistinguishable" from uncompromised ones.
- Urgent Need for Re-evaluation: The ML community must urgently "reconsider our current practice of privacy auditing in machine learning" and enhance supply chain security measures to detect and mitigate these sophisticated, stealthy attacks.
About the Speaker(s)
Zitao Chen is the speaker for this talk at the NDSS Symposium. Based on the content of the presentation, Zitao Chen is a researcher deeply engaged in the field of machine learning security and privacy, specifically focusing on the vulnerabilities of deep learning models to membership inference attacks and the broader implications of supply chain security in ML development. While the transcript does not provide his specific title or institutional affiliation, his work demonstrates expertise in devising novel attack methodologies and critically evaluating existing defense mechanisms in machine learning privacy.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid, original research that cracks open a genuinely underexplored attack surface: supply-chain-poisoned training code enabling stealthy membership inference that bypasses existing auditing tools entirely. The divide-and-conquer framing — encoding membership through secret samples with a dedicated normalization path — is a clean, non-obvious insight that invalidates a whole class of current defenses.
Heather Calloway (CISO) — WEAK
Technically credible research that identifies a real and underappreciated ML supply chain privacy risk, but the talk never closes the distance between the lab finding and the institutional decisions that need to change. The defensive section lists directions without giving defenders or security leaders a usable threshold for action.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025