CENSOR: Defense Against Gradient Inversion via Orthogonal Subspace Bayesian Sampling

Kaiyuan Zhang (PhD Student · PU)

Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Federated Learning 2

Overview

In an era where data privacy is paramount, federated learning (FL) has emerged as a promising distributed machine learning paradigm. It allows multiple clients to collaboratively train a shared model without directly exposing their raw, sensitive data to a central server. This approach has found applications across various domains, including network prediction, credit risk assessment, and the aggregation of data from IoT devices. However, the talk "CENSOR: Defense Against Gradient Inversion via Orthogonal Subspace Bayesian Sampling," presented by Kaiyuan Zhang, a fifth-year PhD student in Computer Science at Purdue University, unveils a critical vulnerability in this seemingly private setup: gradient inversion attacks.

Watch on YouTube · Slides

Key moments

  1. 0:30 Understanding gradient inversion attacks in federated learning
  2. 2:40 Survey of different gradient inversion attack methods
  3. 4:00 Analyzing existing defenses and their shortcomings
  4. 5:00 Two key observations motivating CENSOR's approach
  5. 7:00 CENSOR's core idea: orthogonal subspace gradient sampling
  6. 7:50 In-depth explanation of CENSOR's layer-wise method
  7. 8:30 CENSOR's superior quantitative performance against attacks
  8. 9:30 Visual demonstration of CENSOR's defense effectiveness

CENSOR: Defense Against Gradient Inversion via Orthogonal Subspace Bayesian Sampling

Speakers: Kaiyuan Zhang (PhD Student, PU)

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=Z6I4z17rGIA

Overview

In an era where data privacy is paramount, federated learning (FL) has emerged as a promising distributed machine learning paradigm. It allows multiple clients to collaboratively train a shared model without directly exposing their raw, sensitive data to a central server. This approach has found applications across various domains, including network prediction, credit risk assessment, and the aggregation of data from IoT devices. However, the talk "CENSOR: Defense Against Gradient Inversion via Orthogonal Subspace Bayesian Sampling," presented by Kaiyuan Zhang, a fifth-year PhD student in Computer Science at Purdue University, unveils a critical vulnerability in this seemingly private setup: gradient inversion attacks.

These sophisticated attacks demonstrate that even when raw data remains local, an "honest-but-curious" server or a malicious eavesdropper can reconstruct private input data solely from the gradients shared during the FL training process. CENSOR, a collaborative work between Purdue University and IBM, introduces a novel and highly effective defense mechanism designed to mitigate these gradient inversion attacks. The significance of this research lies in its ability to protect the integrity of user data in FL applications, ensuring that the promise of privacy in distributed learning can be genuinely upheld without compromising model utility or convergence.

Background

▶ Watch: Understanding gradient inversion attacks in federated learning (0:30)

Federated learning (FL) is fundamentally built on the premise of privacy by design. Clients compute local model updates (gradients) based on their private datasets and send only these aggregated updates to a central server, which then combines them to improve the global model. The raw data never leaves the client's device, theoretically safeguarding sensitive information. However, this privacy guarantee was significantly challenged by the discovery of gradient inversion attacks. These attacks exploit the information inherently present in the shared gradients to reverse-engineer the original input data.

The mechanism behind gradient inversion typically involves initializing dummy inputs and iteratively refining them by matching their gradients to the target gradients received from the clients. This can be achieved through stochastic optimization techniques or by leveraging generative models. The threat model assumed for CENSOR is an honest-but-curious server, meaning the adversary follows the protocol but attempts to infer private information. This adversary is assumed to have knowledge of the model architecture, access to the local gradients shared by clients, and can utilize publicly available datasets and pre-trained models, including Generative Adversarial Networks (GANs). This "loose assumption" on the adversary highlights the robust nature of the problem CENSOR aims to solve.

Several representative gradient inversion attacks have been developed, each with distinct approaches:

  • IG (Inverting Gradients): An early stochastic-based method that optimizes on signed gradients using cosine similarity and refines inputs initialized from Gaussian noise.
  • GI (Gradient Inversion): Another stochastic-based method that initializes inputs with Gaussian noise and employs the Adam optimizer with regularization.
  • GGL (Generative Gradient Leakage): This attack leverages GANs with a CB-based regularization and optimizes with spatiation or a covariance matrix.
  • GIS (Gradient Inversion with Similarity): Utilizes negative cosine similarity as a gradient dissimulation function.
  • GIFD (Gradient Inversion over Feature Domains): Considered the most state-of-the-art attack, it exploits intermediate GAN features and optimizes using a warm-up strategy.

Prior attempts to defend against these attacks have met with varying degrees of success and limitations:

  • Noisy gradient: Involves adding Gaussian noise to gradients. While it reduces privacy leakage, it significantly degrades model utility.
  • Gradient clipping: Bounces the magnitude of gradients by clipping their values. This method proved largely ineffective in preventing privacy leakage.
  • Gradient sparsification: Zeros out smaller gradient values, transmitting only the larger ones. Experiments showed this method still leaks sensitive information.
  • Sort: A more sophisticated defense that balances utility and privacy through optimization and gradient masking. However, it is known to be computationally very expensive.

Two key observations motivated the development of CENSOR. First, existing gradient inversion attacks are most successful in the early stages of training, particularly when the batch size is one. This suggests a window of vulnerability that needs targeted protection. Second, while some generative attacks like GGL can produce high-quality images, these are often low-fidelity reconstructions. For example, GGL might reconstruct a bird, but it could be in a different pose or setting than the original input, still constituting a significant privacy breach by revealing the type of data, if not the exact instance. These insights underscore the need for a defense that not only prevents exact reconstruction but also obscures the general characteristics of the private data.

Key Findings

▶ Watch: Analyzing existing defenses and their shortcomings (4:00)

CENSOR introduces a novel and highly effective defense against gradient inversion attacks in federated learning, addressing critical limitations of prior methods. The core findings demonstrate its superior performance across multiple dimensions:

Firstly, CENSOR consistently and effectively mitigates various gradient inversion attacks. Through extensive quantitative and qualitative evaluations, it was shown to prevent the reconstruction of meaningful images from shared gradients, even by the most advanced attack methodologies like GIFD.

Secondly, CENSOR significantly outperforms existing defense mechanisms. In comparative experiments, it surpassed the state-of-the-art Sort defense by an impressive margin, achieving up to a 114% improvement in defense efficacy across certain metrics. This highlights a substantial leap forward in balancing privacy protection with practical utility.

Thirdly, a crucial finding is CENSOR's ability to maintain model utility and convergence. Unlike many previous defenses that trade off privacy for degraded model performance, CENSOR ensures that the federated learning process converges to a similar level as unprotected training. This makes it a practical solution for real-world FL deployments where model accuracy is paramount.

Finally, CENSOR exhibits robustness against adaptive attacks. The research demonstrated its continued effectiveness even when faced with sophisticated adversaries employing strategies like Expectation Over Transformation (EOT), which attempts to approximate the true gradient by averaging transformations. This resilience against adaptive threats underscores CENSOR's strong security posture.

In essence, CENSOR's key findings reveal a defense mechanism that is not only highly effective at preventing gradient inversion but also practical, efficient, and robust, thereby strengthening the privacy guarantees of federated learning.

Technical Deep Dive

▶ Watch: CENSOR's core idea: orthogonal subspace gradient sampling (7:00)

The technical innovation behind CENSOR lies in its concept of Orthogonal Subspace Bayesian Sampling for gradient perturbation. The fundamental intuition stems from the observation that neural networks, particularly large ones, are often over-parameterized. This means that in the high-dimensional space of gradients, there are potentially millions of different directions—or subspaces—that are orthogonal to the original gradient. CENSOR leverages this property to perturb the gradients in a way that makes reconstruction extremely difficult for an attacker, while simultaneously preserving the essential information required for model training.

The core mechanism of CENSOR is a layer-wise orthogonal subspace perturbation. When a client computes its local gradients based on its private input, instead of directly sending these raw gradients to the server, CENSOR intercepts them. For each layer of the neural network model, the original gradient is conceptually projected into an orthogonal subspace. The challenge then becomes selecting a specific gradient within this vast orthogonal space that satisfies two conditions: it must sufficiently obscure the original private input to thwart inversion attacks, and it must still contribute meaningfully to the global model's training objective, thus maintaining utility.

This selection process is guided by what the speaker refers to as "selecting the one with the lowest loss." This implies an optimization or sampling procedure where multiple potential orthogonal gradients are explored, and the one that minimizes a specific loss function (related to the model's training objective) is chosen. The "Bayesian Sampling" aspect in the title, as clarified in the Q&A, refers to the intelligent approach taken to find this optimal orthogonal space. By "project[ing] the gradient into a space that is orthogonal to the original gradient," CENSOR makes it "very hard for the like attacker to find the specific that space" that corresponds to the original data, thereby protecting privacy. At the same time, because the chosen orthogonal gradient maintains a "lowest loss," it ensures that "the orthogonal space it can maintain the utility."

After this layer-wise operation, where each layer's gradient is perturbed into an orthogonal subspace, these processed gradients are aggregated to form a protected gradient (denoted as capital G star). This G* is then transmitted to the server instead of the original raw gradient. This process effectively scrambles the direct information about the input data within the gradients, making it infeasible for an adversary to reconstruct the original input even with knowledge of the model architecture and public datasets.

Compared to existing defenses, CENSOR offers a distinct advantage. While methods like differential privacy (DP) or noisy gradients add indiscriminate noise that often degrades utility, and gradient clipping or sparsification prove insufficient, CENSOR's approach is more targeted. By perturbing gradients into an orthogonal subspace, it leverages the inherent redundancy and high dimensionality of neural network parameter spaces. This allows CENSOR to achieve a superior balance between enhancing data privacy and maintaining model utility, making it a more sophisticated and effective solution against gradient inversion attacks.

Demo / Proof of Concept

▶ Watch: In-depth explanation of CENSOR's layer-wise method (7:50)

Instead of a live demonstration, the talk "CENSOR" presented a comprehensive series of quantitative and qualitative experiments to validate its effectiveness and demonstrate its superior performance over existing defenses. These experiments were crucial in showcasing CENSOR's ability to protect privacy while preserving model utility.

For the quantitative evaluation, CENSOR was benchmarked against various existing defenses and state-of-the-art gradient inversion attacks across three distinct datasets: ImageNet, FFHQ (Flickr-Faces-HQ), and a third dataset referred to simply as "7" in the transcript (likely a smaller image dataset such as CIFAR-10). The efficacy of the defense was measured using several key metrics:

  • MSE (Mean Squared Error): A lower MSE for reconstructed images indicates better defense (less similarity to original).
  • RPIPS (Root-mean-square Perceptual Image Patch Similarity): Similar to MSE, lower RPIPS indicates better defense.
  • PSNR (Peak Signal-to-Noise Ratio): Higher PSNR typically means less noise; in the context of defense, a lower PSNR for the reconstructed image relative to the original indicates less successful inversion.
  • SSIM (Structural Similarity Index Measure): Measures the similarity between two images. A lower SSIM for the reconstructed image relative to the original indicates a more effective defense.

The results unequivocally demonstrated CENSOR's superiority. It outperformed existing defenses in almost all cases, significantly surpassing the performance of the state-of-the-art Sort defense by up to 114% across certain metrics. This substantial improvement highlights CENSOR's enhanced capability to prevent meaningful reconstruction of private data.

The qualitative experiments provided visual evidence of CENSOR's effectiveness. Examples from ImageNet and FFHQ were presented, showcasing the original input images alongside images reconstructed by various attacks, both with and without different defense mechanisms applied. The visual comparisons clearly illustrated that CENSOR "effectively prevents the attacks from inverting any meaningful images," rendering the reconstructed outputs unintelligible or significantly distorted compared to the originals.

To address concerns about model performance, a convergence study was conducted. This study involved training on a dataset with NID (non-independent and identically distributed) data distribution, using 100% of clients, with 10% of them randomly selected for each round over 2,000 FL rounds. The results showed that both the original FL training and CENSOR (with varying numbers of sampling trials for the orthogonal subspace) converged to a similar level of accuracy on both training and test sets. This critical finding confirms that CENSOR achieves robust privacy protection without compromising the utility or convergence speed of the federated learning model.

Finally, the resilience of CENSOR against adaptive attacks was tested. An adaptive attack, referred to as EOT (Expectation Over Transformation), was introduced. EOT represents a more sophisticated adversary that attempts to circumvent defenses by performing gradient transformations multiple times and averaging the resulting gradients to approximate the true gradient. Even under this advanced adaptive attack, CENSOR "still holds effectiveness" on both ImageNet and FFHQ datasets, further solidifying its robust security posture.

Defensive Implications

▶ Watch: Visual demonstration of CENSOR's defense effectiveness (9:30)

CENSOR offers profound defensive implications for organizations and practitioners deploying federated learning systems. Its development directly addresses a critical and previously underestimated privacy vulnerability: the reconstruction of sensitive raw data from shared gradients. By providing a robust and efficient defense, CENSOR significantly strengthens the privacy guarantees of FL, making it a more trustworthy paradigm for distributed machine learning.

The primary implication is the enhanced protection of user data. In real-world applications where FL is used—such as Google's next-word prediction on keyboards, healthcare data analysis, financial fraud detection, or IoT device aggregation—the risk of an "honest-but-curious" server or a malicious entity inferring private information (like typed passwords, medical conditions, or personal interactions) is substantially mitigated. CENSOR's ability to prevent the inversion of "meaningful images" translates directly to preventing the leakage of private text, images, or other data types that form the basis of local client training.

Furthermore, CENSOR's performance characteristics are highly beneficial for practical deployment. Unlike many prior defenses that necessitated a significant trade-off between privacy and model utility or incurred substantial computational overhead (e.g., the high computational cost of the Sort defense), CENSOR maintains model convergence and accuracy comparable to unprotected FL. This means that organizations do not have to sacrifice the performance of their machine learning models to achieve a higher level of privacy. This makes CENSOR a viable and attractive solution for deployment in production environments where both privacy and model efficacy are non-negotiable requirements.

The demonstrated resilience against adaptive attacks like EOT is another crucial defensive implication. It indicates that CENSOR is not easily circumvented by adversaries who are aware of the defense mechanism and attempt to craft sophisticated bypass strategies. This provides a higher degree of confidence in its long-term security against evolving threats.

In summary, CENSOR empowers defenders with a powerful tool to secure their federated learning deployments. It moves beyond theoretical privacy promises to offer a practical, effective, and efficient mechanism to prevent gradient inversion attacks. Organizations should consider integrating such advanced defenses to truly realize the privacy benefits of federated learning, thereby fostering greater trust and enabling the secure application of AI in sensitive data environments.

Key Takeaways

  • Gradient Inversion Threat: Federated learning, despite its privacy promises, is vulnerable to gradient inversion attacks that can reconstruct private input data from shared gradients.
  • Limitations of Prior Defenses: Existing defenses often struggle to balance privacy protection with model utility, or they introduce significant computational overhead.
  • CENSOR's Novel Approach: CENSOR defends against gradient inversion by performing layer-wise orthogonal subspace perturbation, intelligently sampling gradients in a space orthogonal to the original.
  • Superior Performance: CENSOR significantly outperforms state-of-the-art defenses like Sort, demonstrating up to a 114% improvement in defense efficacy across various metrics and effectively preventing meaningful image reconstruction.
  • Utility Preservation: Crucially, CENSOR maintains model utility and convergence rates comparable to unprotected federated learning, ensuring that privacy is not achieved at the cost of performance.
  • Robustness Against Adaptive Attacks: CENSOR exhibits strong resilience against sophisticated adaptive attacks, such as EOT, further validating its security posture in real-world scenarios.

About the Speaker(s)

Kaiyuan Zhang is a fifth-year PhD student in Computer Science at Purdue University. His research focuses on critical security and privacy challenges in machine learning, as demonstrated by his work on CENSOR. This project is a joint effort between Purdue University and IBM, highlighting a collaboration between academia and industry. Kaiyuan Zhang is currently active on the job market for the upcoming cycle. He also mentioned being a fellow selected by the Internet Society, reflecting his engagement with broader internet security and privacy communities.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Legitimate academic ML security research with a real technical contribution — orthogonal subspace perturbation as a gradient inversion defense is a credible idea grounded in the overparameterization of large networks. The results look solid and the adaptive attack evaluation against EOT is a necessary box that too many defense papers skip. But this is a conference paper presentation, not a breakout security talk, and the write-up leans heavily on abstract framing without enough mechanistic specificity to fully evaluate the core claim.

Heather Calloway (CISO) — WEAK

Technically credible PhD-level research on a real FL privacy vulnerability, but it never bridges to the institutional and governance questions that determine whether this work actually gets deployed. The gap between 'we built a better defense' and 'here is how organizations should rethink their FL risk posture' is left entirely uncrossed.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025