Passive Inference Attacks on Split Learning via Adversarial Regularization
Xiaochen Zhu (grad student · MIT)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 2 · Federated Learning 1
Overview
Split Learning (SL) has emerged as a promising paradigm for collaborative machine learning, designed to address the challenges of distributed data, limited computational resources on client devices, and the paramount need for data privacy. By partitioning a neural network into client-side and server-side components, SL aims to allow multiple data owners to collaboratively train a model without directly sharing their raw data. Clients send only intermediate representations of their data to a powerful central server, which then handles the bulk of the computation and sends gradients back for client-side updates. This architectural choice is intended to safeguard sensitive client information while enabling the benefits of large-scale model training.
Key moments
- 0:00 Introduction and Talk Agenda
- 0:40 Understanding Split Learning (SL) and its Motivation
- 2:40 Identifying Privacy Vulnerabilities in SL
- 3:20 Limitations of Existing Server-Side Attacks
- 4:00 STAR Attack: Threat Model and Assumptions
- 4:40 Why Naive Inference Attacks Fail
- 5:40 Solving Simulator Mismatch with GAN Regularization
Passive Inference Attacks on Split Learning via Adversarial Regularization
Speakers: Xiaochen Zhu, Grad Student, MIT
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=J-A3Zz3_B1M
Overview
Split Learning (SL) has emerged as a promising paradigm for collaborative machine learning, designed to address the challenges of distributed data, limited computational resources on client devices, and the paramount need for data privacy. By partitioning a neural network into client-side and server-side components, SL aims to allow multiple data owners to collaboratively train a model without directly sharing their raw data. Clients send only intermediate representations of their data to a powerful central server, which then handles the bulk of the computation and sends gradients back for client-side updates. This architectural choice is intended to safeguard sensitive client information while enabling the benefits of large-scale model training.
However, the security and privacy guarantees of Split Learning are not absolute. This talk, presented by Xiaochen Zhu, a grad student at MIT (with work conducted at NUS), critically examines the privacy vulnerabilities inherent in SL, particularly concerning the client's raw features and labels. The presentation introduces "STAR" (Simulator Decoding with Adversarial Regularization), a novel and highly effective passive inference attack. STAR demonstrates that even an "honest-but-curious" server, adhering strictly to the SL protocol without malicious modifications, can reconstruct significant portions of the client's private data.
The implications of STAR are profound for the deployment and security of Split Learning systems. It highlights that the intermediate representations exchanged in SL, while not raw data, still encode sufficient information for a sophisticated adversary to infer sensitive inputs. The talk underscores the urgent need for more robust privacy-preserving mechanisms in SL, as current assumptions about its inherent privacy may be overly optimistic, especially against stealthy, passive adversaries like the one modeled by STAR.
Background
▶ Watch: Introduction and Talk Agenda (0:00)
The rapid advancements in machine learning have led to widespread adoption, but practical deployment often faces significant hurdles. Data is frequently distributed across numerous owners, each possessing limited and potentially biased datasets. This necessitates collaborative training paradigms like Federated Learning (FL), where models are trained locally on client data, and only aggregated updates are sent to a central server. Concurrently, many data owners, such as mobile devices, lack the computational resources to train complex models independently, giving rise to Machine Learning as a Service (MLaaS) models. While moving all data to a powerful central server could address these issues, it fundamentally compromises privacy.
Split Learning (SL) was conceived to navigate this tension between collaboration, computational efficiency, and privacy. The core idea of SL is to "split" a deep neural network into multiple layers, typically assigning the initial S layers (model F) to the client and the remaining layers (model G) to the server. During training, the client processes its raw input features X through its local model F to produce an intermediate representation, FX. This FX is then sent to the server. The server completes the forward pass using its model G and the received FX, computes the loss, and performs backpropagation. Crucially, the server then sends the gradients of its first layer (which correspond to the gradients with respect to FX) back to the client. The client uses these gradients to complete its own backpropagation on F via the chain rule, updating its local model. This process is repeated iteratively.
There are two primary configurations of SL discussed:
- Vanilla SL: In this setting, the server possesses the labels
Ycorresponding to the client's featuresX. The client sendsFXto the server, and the server usesGandYto complete the training. - U-shaped SL: This is a more privacy-preserving variant where the client retains both the features
Xand their corresponding labelsY. The neural network is split into three parts:F,G, andH. The client holdsFandH, while the server holdsG. The client sendsFXto the server, which computesGFXand sends it back to the client. The client then processesGFXthrough its final layerHand usesYfor loss computation and backpropagation. This effectively means the client outsources the training of the computationally intensive middle part (G) of the model to the server.
Despite its privacy-preserving intent, SL inherently introduces vulnerabilities. Firstly, the intermediate representation FX, though not raw data, inevitably encodes significant information about the client's original features X. Secondly, the client's model F updates are entirely dictated by the gradients received from the server. Prior research has exploited these vulnerabilities, demonstrating that a server can reconstruct features and sometimes labels to a certain degree. However, existing server-side attacks often suffer from limitations:
- Active Server Assumption: Many attacks assume an "active" server that can maliciously tamper with the SL protocol, sending crafted updates back to clients. Such active interference might be detectable, compromising stealth.
- Client Model Access: Some attacks require the server to have access to the client's model
F(either white-box or black-box access), which is generally not a valid assumption in SL whereFresides exclusively on the client. - Limited Applicability: Many attacks target "easily exploitable" models or settings, and their performance degrades significantly in more challenging scenarios, such as higher "split levels" (where more layers are on the client side) or U-shaped SL.
The STAR attack considers a more realistic and challenging threat model: a passive, honest-but-curious server. This server adheres strictly to the SL protocol, never tampering with messages, thus remaining stealthy and undetectable. It has no access to the client's model F (neither white-box nor black-box). The only auxiliary resource required by the server is a labeled auxiliary dataset (X', Y') that belongs to the same domain as the client's private data. This assumption is consistent with the threat model of prior state-of-the-art passive attacks like PCAT. Under these constraints, STAR aims to achieve superior performance in reconstructing both client features X and labels Y (in U-shaped SL) across various challenging settings.
Key Findings
▶ Watch: Identifying Privacy Vulnerabilities in SL (2:40)
The research presented in this talk reveals several critical findings regarding the privacy posture of Split Learning and the efficacy of the STAR attack:
- Superior Feature Reconstruction: The STAR attack significantly outperforms existing state-of-the-art passive inference attacks, such as PCAT, in reconstructing client features (
X). This superiority is evident across various Split Learning configurations, including both Vanilla SL and U-shaped SL, and is particularly pronounced in challenging scenarios with higher "split levels" (where more layers reside on the client side, theoretically making attacks harder). The Mean Squared Reconstruction Error (MSE) achieved by STAR is "much, much lower" than PCAT, indicating a substantially more accurate reconstruction. - Robustness Across Split Levels: Unlike prior passive attacks that struggle to produce meaningful reconstructions as the split level increases, STAR maintains its effectiveness even when a substantial portion of the model (e.g., half the layers in a 7-layer split) is on the client side. Visual examples clearly demonstrate STAR's ability to reconstruct recognizable images where PCAT yields only noise or distorted artifacts.
- Near-Active Attack Performance: Intriguingly, STAR's performance in feature reconstruction almost matches that of FSHA, which is a state-of-the-art active attack. This is a significant finding because STAR operates under a passive, honest-but-curious threat model, meaning it achieves comparable results without requiring the server to actively modify or tamper with the SL protocol, thus remaining stealthy and undetectable.
- Effective Label Inference in U-shaped SL: For U-shaped SL, where the server does not have access to client labels (
Y), STAR demonstrates "very good label inference accuracy" across all tested split levels. This capability is a crucial advancement, as it allows a passive server to infer both the client's private features and their corresponding sensitive labels, further eroding the privacy guarantees of this SL variant. PCAT, in contrast, shows a marked drop in performance in the U-shaped setting, highlighting STAR's superior generalization. - Crucial Role of Adversarial Regularization: The success of STAR is attributed to its novel use of adversarial regularization, which addresses two key limitations of naive simulation attacks. The two discriminators (D1 for representation similarity and D2 for plausible reconstruction) and their associated GAN-style losses are instrumental in training a simulator and decoder that generalize effectively to unseen client data.
- Ineffectiveness of Current Defenses: The talk evaluates several potential countermeasures, finding most to be either ineffective or impractical against STAR. Assigning more layers to the client (even to an extreme degree) does not fully defeat the attack and often undermines the purpose of SL. A decorrelation defense, designed to minimize input-output correlation, also proves ineffective. Cryptographic protocols introduce prohibitive computational overhead, and differential privacy is unsuitable for non-aggregated intermediate representations like
FX.
In summary, STAR demonstrates that Split Learning, even in its U-shaped configuration, is highly susceptible to sophisticated passive inference attacks. The attack's ability to reconstruct features and infer labels with high fidelity, without requiring an active or model-aware adversary, fundamentally challenges the privacy assumptions underlying SL.
Technical Deep Dive
▶ Watch: Limitations of Existing Server-Side Attacks (3:20)
The STAR attack addresses the core limitations of a naive approach to inference in Split Learning. A naive attack would involve the server training a simulator model, ~F, such that ~F combined with the server's model G can classify the auxiliary dataset (X', Y'). The hope is that a decoder trained on ~F could then decode the client's actual F. However, this fails for two main reasons:
- Representation Mismatch:
~Flearning to classifyX'withGdoes not guarantee that~Flearns the same intermediate representations as the client's private modelF. - Decoder Generalization Failure: A decoder that works perfectly for the simulator
~F(which is white-box to the server) may not generalize well to decode the unseen, private representationsFXfrom the client. The reconstructions often look like "not real images."
STAR tackles these two problems using a unified methodology based on adversarial regularization, inspired by Generative Adversarial Networks (GANs).
Adversarial Regularization for Vanilla SL
For Vanilla SL, where the server has G and client labels Y, the goal is to reconstruct X from FX.
The STAR attack introduces two key components:
- Ensuring Representation Similarity (Addressing Problem 1):
- The server trains a simulator
~Fto mimic the client'sF. - A Discriminator D1 is introduced. D1's role is to distinguish between the real intermediate representations
FX(observed from the client) and the synthetic representations~F(X')(generated by the simulator~Fusing the auxiliary dataX'). ~Fis trained with a GAN generator loss as a regularization term. This loss encourages~Fto produce outputs~F(X')that are indistinguishable fromFXto D1. In essence,~Fis optimized not just to classifyX'but also to learn representations that "look like"FXto an adversary.
- Ensuring Plausible Reconstructions (Addressing Problem 2):
- A decoder
Decis trained to reconstruct input images from representations. - A second Discriminator D2 is introduced. D2's role is to distinguish between real images from the auxiliary dataset
X'and the reconstructed imagesDec(FX)(whereFXcomes from the client). - The decoder
Decis trained with a GAN generator loss as a regularization term. This loss compelsDecto produce reconstructionsDec(FX)that are visually plausible and indistinguishable from real imagesX'to D2. This ensures that the decoder, when applied toFX, yields realistic-looking images.
Together, these two adversarial regularization terms are added to the simulator training process. The overall objective for training ~F and Dec involves both the standard supervised learning loss (for ~F to classify X' with G) and the GAN losses from D1 and D2. This combined approach allows the server to train a simulator ~F that genuinely mimics the client's F's representations and a decoder Dec that can produce high-fidelity reconstructions from FX.
Extending to U-shaped SL
In U-shaped SL, the client holds F and H, while the server only has G. The server does not have access to the client's labels Y. The attack goal is to infer both X and Y.
The core adversarial regularization techniques (D1 and D2) are still employed for feature reconstruction. However, because the server lacks H and Y, it must also train a simulator ~H for the client's final layer.
- Simulator Training for U-shaped SL: The server trains two simulators,
~Fand~H, such that the combined model~F - G - ~Hcan classify the auxiliary dataX'toY'. This forms the basis for both feature and label inference. - Adversarial Regularization: The D1 and D2 discriminators and their associated GAN losses are still used to ensure
~Flearns similar representations toFand the decoder produces plausibleXreconstructions. - Random Label Flipping for
~H: To ensure that~Hlearns general representations and doesn't simply overfit to the auxiliaryX'andY', a crucial technique is introduced: random label flipping onY'during the training of~H. By intentionally introducing noise into the labelsY'for~H's training, the simulator~His forced to learn more robust and general mappings that work effectively withGfor classifying unseen client data. This strategy is vital for achieving high label inference accuracy.
For label inference, once ~F, G, and ~H are well-trained, the server simply feeds the observed GFX (from the client) through its simulated ~H. If ~H accurately mimics H, then ~H(GFX) provides a valid and accurate reconstruction of the client's private label Y.
The paper also discusses other technical aspects, such as the effects of auxiliary data distribution, target model architecture, and the server's knowledge of the client's model architecture, along with ablation studies to validate the contribution of each component. These details further solidify the robustness and generalizability of the STAR attack.
Demo / Proof of Concept
▶ Watch: Why Naive Inference Attacks Fail (4:40)
The talk vividly demonstrated the efficacy of the STAR attack through compelling visual and quantitative results, contrasting its performance against the prior state-of-the-art passive attack, PCAT. The primary datasets used for these demonstrations included CIFAR-10 and CIFAR-100, as well as other "larger images," indicating a focus on image classification tasks.
For feature inference (reconstructing X) in Vanilla SL, the speaker presented quantitative results showing that STAR achieved a "much, much lower mean squares reconstruction error" compared to PCAT. Visually, this translated into remarkably clear and recognizable reconstructions of client images. The presentation included side-by-side comparisons of reconstructed images, with STAR producing outputs that closely resembled the original client inputs, while PCAT often yielded blurry, distorted, or entirely unrecognizable images, especially as the split level (the number of layers on the client side) increased. A specific example highlighted was at split level 7, where approximately half of the model was on the client side. In this challenging scenario, STAR's reconstructions remained highly effective and visually coherent, whereas PCAT "struggles to produce meaningful reconstructions."
A significant finding was that STAR's feature reconstruction results "almost match the attack result of FSHA," which is an active attack known to be detectable due to its modification of messages. This underscores STAR's ability to achieve high-fidelity reconstruction without any active interference.
In the more challenging U-shaped SL setting, where the server lacks direct access to client labels, the performance gap between STAR and PCAT became even more pronounced. STAR demonstrated superior feature inference, with reconstructions that were clearly "much better" than PCAT's, which showed a significant drop in performance. This indicated that STAR's simulated ~H (the client's final layer) was doing "a very good job mimicking whatever H is doing."
Furthermore, for label inference (reconstructing Y) in U-shaped SL, STAR achieved "very good label inference accuracy" across all tested split levels. The talk presented comparative charts showing STAR's consistently high accuracy for label prediction, whereas PCAT's accuracy was notably lower. This confirms STAR's capability to infer both sensitive features and their associated labels from a purely passive observation of intermediate representations.
The demonstrations collectively served as a strong proof of concept, illustrating that the adversarial regularization techniques employed by STAR effectively overcome the limitations of prior passive attacks, leading to significantly improved and more robust privacy breaches in Split Learning.
Defensive Implications
▶ Watch: Solving Simulator Mismatch with GAN Regularization (5:40)
The STAR attack presents a significant challenge to the privacy claims of Split Learning, demonstrating that even a passive, honest-but-curious server can infer sensitive client data. The talk explored several potential countermeasures, highlighting their limitations against STAR:
- Assigning More Layers to the Client: One intuitive defense is to place more layers of the neural network on the client side, making the intermediate representation
FXless informative. While increasing the split level (e.g., assigning all but the last layer to the client) makes attacks harder, the speaker explicitly stated that STAR remains "workable" even under such extreme configurations. Moreover, pushing too many layers to the client "totally defeats the purpose of SL," as it negates the benefits of offloading computation to a powerful server and places a heavier burden on resource-constrained client devices. - Decorrelation Defense: This type of defense aims to minimize the statistical dependence or distance correlation between the input
Xand the intermediate outputFX. The idea is to makeFXless informative aboutXwhile still allowing the model to learn. However, the research found that "even under that decorrelation defense, our attack remains effective." This suggests that simply reducing correlation might not be sufficient to obscure the information that STAR's sophisticated adversarial regularization can extract. - Cryptographic Protocols: Techniques like Homomorphic Encryption (HE) or Secure Multi-Party Computation (SMC) could provide strong privacy guarantees by allowing computations on encrypted data. However, the speaker noted that these protocols "will include too much computation overhead and defeat the purpose of SL." The very motivation for SL is to reduce computational burden on clients; introducing heavy cryptographic operations would counteract this advantage, making the solution impractical for many real-world SL deployments.
- Differential Privacy (DP): DP is a strong privacy guarantee achieved by adding noise to data or gradients. However, the speaker clarified that DP is not directly applicable in the context of
FXin SL becauseFXrepresents individual, non-aggregated information. DP is typically applied to aggregated data or gradients to protect individual contributions within a larger dataset, not to every single intermediate representation sent from a client.
Given the ineffectiveness or impracticality of these existing or proposed defenses, the talk concluded by pointing to the need for future work in developing more robust countermeasures. During the Q&A, the speaker suggested that adding noise to FX might be a possible avenue. The rationale is that "noise cannot be easily learned" by the adversarial regularization process. However, the speaker also cautioned that if the noise itself is "learning based," then STAR's learning-based attack might still find ways to circumvent it.
Overall, the defensive implications are stark: current Split Learning architectures, even with attempts at obfuscation or increased client-side processing, are highly vulnerable to passive inference attacks like STAR. The existing privacy-enhancing technologies either fail to protect against this specific attack vector or introduce prohibitive overheads that undermine the core advantages of Split Learning. This calls for a re-evaluation of SL's privacy model and the development of new, practical, and effective defense mechanisms specifically designed to counter sophisticated inference attacks on intermediate representations.
Key Takeaways
- Split Learning (SL) is highly vulnerable to passive inference attacks: Contrary to assumptions, intermediate representations (
FX) in SL leak significant information about client features (X) and labels (Y), even when the server acts honestly and without tampering. - STAR (Simulator Decoding with Adversarial Regularization) is a potent new attack: This novel passive attack, leveraging GAN-style adversarial regularization, effectively reconstructs client features and infers labels, overcoming limitations of prior passive methods.
- STAR excels in challenging SL configurations: The attack maintains high performance even at high "split levels" (more layers on the client) and in the U-shaped SL setting (where the server lacks client labels), outperforming prior state-of-the-art passive attacks like PCAT.
- Passive attacks can rival active attacks: STAR's feature reconstruction performance nearly matches that of active, detectable attacks (like FSHA), demonstrating that a stealthy, honest-but-curious server can achieve significant privacy breaches.
- Current SL defenses are largely insufficient: Existing countermeasures, such as increasing client-side layers or employing decorrelation techniques, prove ineffective against STAR. Resource-intensive cryptographic protocols and differential privacy are often impractical or inapplicable in this context.
- Urgent need for new, robust defenses: The findings highlight a critical gap in Split Learning's privacy guarantees, necessitating the development of novel, practical, and effective defense mechanisms to safeguard client data against sophisticated inference attacks.
About the Speaker(s)
Xiaochen Zhu is a graduate student at MIT (Massachusetts Institute of Technology). The research presented in this talk on "Passive Inference Attacks on Split Learning via Adversarial Regularization" was conducted while all authors were affiliated with NUS (National University of Singapore). His work focuses on the security and privacy aspects of machine learning, particularly in distributed and collaborative settings like Split Learning.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid, technically grounded ML privacy research that delivers a genuine contribution: a passive inference attack on Split Learning that closes the gap with active attacks while remaining undetectable, with the GAN-style adversarial regularization being the key novel mechanism. The work is honest about what doesn't work against it, which is rarer than it should be, and the U-shaped SL label inference result is the kind of finding that should make anyone building on SL's privacy assumptions reconsider their threat model.
Heather Calloway (CISO) — WEAK
Technically credible research that demonstrates a meaningful gap between Split Learning's privacy promises and its actual privacy guarantees. But this talk never crosses the line from academic finding to institutional consequence — it ends where the work should begin.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025