MingledPie: A Cluster Mingling Approach for Mitigating Preference Profiling in CFL
Cheng Zhang
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Federated Learning 2
Overview
Federated Learning (FL) offers a privacy-preserving framework for collaborative machine learning, allowing multiple clients to train a shared model without centralizing their sensitive data. However, the inherent heterogeneity of client data often leads to convergence challenges in standard FL. To address this, Clustered Federated Learning (CFL) groups clients into clusters based on their data distributions, enabling the training of personalized cluster models. While CFL improves model accuracy and convergence, it introduces a novel privacy vulnerability: preference profiling attacks, particularly those based on cluster identity. This talk, presented by Cheng Zhang (who introduced himself as Pjang), delves into this specific threat and proposes MingledPie, a robust defense mechanism.
Key moments
- 0:00 Introduction to CFL and preference profiling attacks
- 2:00 Two types of preference profiling attacks in CFL
- 4:00 Limitations of existing methods and MingledPie's goals
- 5:50 MingledPie's core idea: cluster mingling approach
- 8:00 Rebuilding accurate cluster models with homomorphic encryption
- 9:30 Experimental results: attack success rate and accuracy
- 10:30 Summary of MingledPie's contributions
MingledPie: A Cluster Mingling Approach for Mitigating Preference Profiling in CFL
Speakers: Cheng Zhang
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=61zJAAm7XSE
Overview
Federated Learning (FL) offers a privacy-preserving framework for collaborative machine learning, allowing multiple clients to train a shared model without centralizing their sensitive data. However, the inherent heterogeneity of client data often leads to convergence challenges in standard FL. To address this, Clustered Federated Learning (CFL) groups clients into clusters based on their data distributions, enabling the training of personalized cluster models. While CFL improves model accuracy and convergence, it introduces a novel privacy vulnerability: preference profiling attacks, particularly those based on cluster identity. This talk, presented by Cheng Zhang (who introduced himself as Pjang), delves into this specific threat and proposes MingledPie, a robust defense mechanism.
MingledPie tackles the challenge of preference profiling by introducing a novel "cluster mingling" approach. Instead of assigning clients to a single, identifiable cluster, MingledPie allows clients to appear to belong to multiple clusters simultaneously, effectively obfuscating their true data preferences from an honest-but-curious server. The system leverages homomorphic encryption and a sophisticated cluster model rebuilding algorithm to ensure both privacy and model accuracy, without relying on additional third-party infrastructure. This work is critical for advancing the privacy guarantees of CFL, enabling its secure deployment in sensitive application domains where inferring client preferences could have significant implications.
The core contribution of MingledPie lies in its ability to enable a server to perform CFL aggregation and cluster clients without direct knowledge of their true cluster memberships. This is achieved through a multi-faceted design that includes client-side clustering, generation of false positive cluster identities, and an encrypted aggregation and rebuilding process. By mitigating the unique "cluster identity-based inference attack," MingledPie significantly enhances the privacy posture of CFL, making it a more viable solution for real-world distributed machine learning scenarios where both performance and privacy are paramount.
Background
▶ Watch: Introduction to CFL and preference profiling attacks (0:00)
Federated Learning (FL) emerged as a paradigm shift in machine learning, allowing organizations and individuals to collaboratively train models on decentralized datasets. In FL, a central server orchestrates the training process, sending a global model to clients. Clients then train local models using their private data, and only the model updates (gradients or parameters) are sent back to the server for aggregation into a new global model. This iterative process continues until convergence. While FL inherently offers data privacy by keeping raw data on client devices, it faces challenges, primarily due to data heterogeneity. When client data distributions vary significantly (a common scenario known as Non-IID data), directly aggregating local models can lead to difficulties in model convergence and reduced accuracy.
To overcome the data heterogeneity problem, Clustered Federated Learning (CFL) was introduced. CFL extends FL by grouping clients with similar data distributions into distinct clusters. Within each cluster, clients collaboratively train a personalized cluster model, which is typically derived from or initialized by a generalized global model. This approach allows for models that are better tailored to specific data characteristics, thereby improving convergence and overall performance in heterogeneous environments. For example, in an image classification task, clients with a prevalence of dog images might form one cluster, while those with cat images form another.
However, CFL, despite its benefits, introduces a new privacy vulnerability: preference profiling attacks. The speaker identifies two main types:
- Inference based on local model updates: Attackers can infer private information about a client's data distribution by analyzing the model updates sent from clients to the server. This type of attack is not unique to CFL and has been addressed by existing FL privacy-preserving techniques like secure aggregation.
- Cluster identity-based inference attack: This type of attack is unique to CFL. Because each cluster is formed based on specific data distribution characteristics, an honest-but-curious server (an attacker that follows the protocol but tries to infer extra information) can infer a client's preferences if it knows which cluster the client belongs to. For instance, if a server knows a client is in the "dog image" cluster, it can infer the client's strong preference for dog-related data. This is particularly problematic because CFL explicitly groups clients by data characteristics, making cluster membership a direct indicator of preference.
Existing privacy-preserving methods fall short in addressing the cluster identity-based inference attack. While secure aggregation protects the content of local model updates, it does not obscure the client's cluster assignment. Similarly, anonymous communication technologies can break the link between model updates and client identities, but they often require additional communication infrastructure or cannot protect the communication data itself, leaving the cluster assignment vulnerable. Therefore, a solution specifically designed to protect cluster identity in CFL, without compromising model accuracy or requiring complex third-party infrastructure, is essential.
The design goals for a robust defense mechanism against preference profiling in CFL are multifaceted:
- Protect preference privacy: Specifically, guard against the cluster identity-based inference attack.
- Ensure cluster model accuracy: The privacy mechanism should not significantly degrade the performance of the CFL models.
- Autonomous deployment: The solution should operate without relying on additional third parties or complex communication overlays.
The key challenge in achieving these goals is enabling the server to effectively cluster clients and aggregate cluster models without knowing the true cluster identity of any individual client. This is the problem MingledPie sets out to solve.
Key Findings
▶ Watch: Limitations of existing methods and MingledPie's goals (4:00)
MingledPie introduces a novel and effective approach to mitigate preference profiling attacks in Clustered Federated Learning (CFL), particularly addressing the unique vulnerability of cluster identity-based inference. The key findings and contributions of this work are:
- Cluster Mingling Concept: The central innovation is the "cluster mingling" paradigm, where clients are not assigned to a single, identifiable cluster but rather appear to belong to multiple clusters simultaneously. This obfuscates their true cluster identity and, consequently, their data preferences from an honest-but-curious server.
- Privacy-Preserving Cluster Identity Generation: Clients generate a "cluster identity set" comprising their true cluster identity and several "false positive" cluster identities. This process leverages a hash function with a controlled collusion probability
P, similar to public key message detection, to create plausible but misleading cluster associations. - Homomorphic Encryption for Secure Aggregation: To protect sensitive model parameters and enable privacy-preserving computations, MingledPie integrates homomorphic encryption. This allows the server to aggregate encrypted local models and other critical information (like cluster identity clues) without ever decrypting or learning the individual client's private data or specific cluster memberships.
- Novel Cluster Model Rebuilding Algorithm: Recognizing that "mingled cluster models" (aggregated from both true and false positive clients) are not directly usable, MingledPie proposes a sophisticated algorithm to rebuild the accurate cluster models. This algorithm relies on clients sending encrypted "cluster identity clues" to the server, which are aggregated into an encrypted mingling coefficient matrix. By solving a system of linear equations derived from this matrix and the mingled models, clients can precisely reconstruct the correct cluster models.
- Demonstrated Attack Success Rate Reduction: Experimental evaluations confirm that MingledPie significantly reduces the success rate of preference profiling attacks. In contrast to undefended approaches where client preferences are "almost certainly" inferred, MingledPie achieves a substantially lower attack success rate, validating its privacy-enhancing capabilities.
- Limited Accuracy Loss: Despite introducing a complex privacy mechanism, MingledPie maintains high model accuracy. Experiments across six datasets demonstrate that the accuracy loss caused by MingledPie is "limited" compared to baseline CFL approaches, showcasing its practicality without severely impacting performance.
- Autonomous and Third-Party Independent Deployment: MingledPie achieves its privacy goals without requiring additional communication infrastructure or reliance on external third parties. This makes it a self-contained and easily deployable solution for CFL environments.
- Convergence Analysis and Ablation Studies: The paper includes a formal convergence analysis to theoretically support the method's stability and effectiveness. Furthermore, ablation experiments confirm the critical role and effectiveness of the proposed cluster model rebuilding method in achieving accurate models despite the mingling process.
Technical Deep Dive
▶ Watch: MingledPie's core idea: cluster mingling approach (5:50)
MingledPie's technical architecture is meticulously designed to achieve its privacy and accuracy goals in Clustered Federated Learning. The core innovation revolves around a client-side clustering mechanism, a unique "mingling" strategy, and a robust cluster model rebuilding algorithm secured by homomorphic encryption.
The process begins with the server publishing an initial set of potential cluster models. Unlike traditional CFL where the server definitively assigns clients to clusters, MingledPie shifts the responsibility of determining cluster identity to the client side.
1. Client-Side Clustering and True Identity Evaluation:
Each client receives all available cluster models from the server. Using their own private data, clients evaluate these models to determine their true cluster identity. This evaluation typically involves assessing which cluster model best fits their local data distribution or yields the highest performance on their local dataset. This step ensures that clients genuinely identify with a cluster that aligns with their data characteristics.
2. Generating Mingled Cluster Identities:
This is the pivotal privacy-preserving step. Instead of simply reporting their true cluster identity, each client generates a cluster identity set. This set comprises their true cluster identity and several false positive cluster identities. The generation of these false positives is crucial for obfuscating the client's actual preference.
The mechanism for generating false positives is analogous to public key message detection. A client uses a hash function with a predefined collusion probability P to identify other clusters that "collide" with its true cluster identity. This P value is a configurable parameter that dictates the level of mingling. A higher P means more false positives and thus greater privacy, but potentially more complexity in rebuilding. The output of this hash function, when applied to the true cluster identity, yields additional cluster IDs that are then included in the client's cluster identity set as false positives.
3. Client Submission to the Server:
Once the cluster identity set is formed, each client performs two main actions:
- Local Model Training: The client trains a local model using its private data.
- Encrypted Submission: The client then encrypts its local model parameters using homomorphic encryption. Crucially, the client also sends its cluster identity set (containing true and false positive cluster IDs) and the encrypted local model to the server. The server receives these submissions, but due to encryption, it cannot directly inspect the local model parameters or definitively know which cluster IDs in the set are true and which are false positives for any given client.
4. Server-Side Aggregation into Mingled Cluster Models:
Upon receiving submissions from all clients, the server proceeds with an aggregation phase. For each potential cluster (real or false), the server collects all encrypted local models from clients whose cluster identity set includes that particular cluster ID. It then aggregates these encrypted local models to form a mingled cluster model.
A mingled cluster model, by its nature, is a combination of accurate local models from clients whose true identity matches that cluster, and "noise" introduced by local models from clients whose true identity is different but who included this cluster ID as a false positive. Mathematically, a mingled cluster model can be viewed as:
Mingled_Cluster_Model_j = Σ (Local_Model_i) for all clients i whose cluster identity set contains j.
Clearly, these mingled cluster models are not directly usable for inference or further training, as they are corrupted by contributions from irrelevant clients.
5. Cluster Model Rebuilding Algorithm:
The core challenge is to reconstruct the accurate cluster models from these unusable mingled cluster models. MingledPie introduces a sophisticated rebuilding algorithm that leverages additional encrypted information:
- Cluster Identity Clues: Each client computes a cluster identity clue. This clue is a vector or matrix representing its contributions to various clusters, specifically indicating its true cluster identity and the false positive clusters it generated. The speaker mentions an equation for this computation, which essentially encodes how the client's local model contributes to the various mingled clusters. This clue is also encrypted using homomorphic encryption before being sent to the server.
- Aggregating the Mingling Coefficient Matrix: The server aggregates these encrypted cluster identity clues from all clients. Due to homomorphic encryption, the server can perform this aggregation without decrypting individual clues. The result of this aggregation is an encrypted mingling coefficient matrix.
This matrix, when decrypted (which is done later by the client or a trusted party if available, though the talk suggests the client solves the system directly), has a specific structure:
- Diagonal elements: Represent the true size of each cluster (i.e., the number of clients whose true cluster identity corresponds to that cluster). For example,
x_jjwould be the count of clients truly belonging to clusterj. - Off-diagonal elements: Represent the number of false positive clients. For instance,
x_12would denote the number of clients from the second true cluster who included the first cluster as a false positive.
The goal is for this matrix to approximate a principal diagonal matrix. This approximation is crucial because it ensures that the subsequent system of linear equations will have a unique and solvable solution, allowing for accurate model rebuilding.
- Solving the System of Linear Equations: Finally, the clients (or a designated entity with decryption capabilities for the aggregated matrix) construct a system of linear equations. This system relates the aggregated mingled cluster models (which are also aggregated under homomorphic encryption and can be decrypted by the client or a trusted party for the solving step) with the mingling coefficient matrix.
Let M_mingled be the vector of mingled cluster models, C be the mingling coefficient matrix, and M_accurate be the vector of true, accurate cluster models. The system can be conceptually represented as:
C * M_accurate = M_mingled
By solving this system, clients can accurately rebuild the true, personalized cluster models for each cluster. This entire process, from client-side clustering to model rebuilding, iterates until the cluster models converge, ensuring both privacy of cluster identity and high model accuracy.
The use of homomorphic encryption is critical throughout this process. It allows the server to perform necessary aggregations (of local models and identity clues) on encrypted data without ever learning the plaintext, thus preserving individual client privacy regarding their true cluster identity and model contributions. The ability to perform computations on encrypted data is what enables the server to facilitate CFL without becoming an oracle for client preferences.
Demo / Proof of Concept
▶ Watch: Experimental results: attack success rate and accuracy (9:30)
While the talk did not feature a live, interactive demonstration, Cheng Zhang presented compelling experimental results that serve as a robust proof of concept for MingledPie's effectiveness. The experimental evaluations focused on two primary metrics: the success rate of preference profiling attacks and the accuracy of the cluster models.
The primary objective of MingledPie is to mitigate preference profiling attacks. The speaker demonstrated this by evaluating the attack success rate of inferring the top-one label preference of clients. The results clearly showed that:
- Undefended approach: In a standard CFL setup without MingledPie, an attacker (the honest-but-curious server) could "almost certainly" infer client preferences. This highlights the severe privacy vulnerability inherent in traditional CFL where cluster identity is directly exposed.
- MingledPie's performance: MingledPie achieved a significantly lower attack success rate. This empirical evidence directly validates the efficacy of the cluster mingling approach and the obfuscation provided by false positive cluster identities in protecting client preferences. The reduction in success rate indicates that the server, even with knowledge of the mingled clusters, cannot reliably link a client to their true data preference.
Beyond privacy, the practicality of any defense mechanism in machine learning hinges on its impact on model performance. MingledPie was evaluated for cluster model accuracy across six different datasets. The experimental results indicated that:
- Limited accuracy loss: The accuracy loss introduced by MingledPie, when compared to baseline CFL approaches (which offer no protection against cluster identity-based profiling), was "limited." This is a critical finding, as it demonstrates that MingledPie achieves strong privacy guarantees without severely compromising the utility or performance of the machine learning models. The system successfully navigates the trade-off between privacy and accuracy.
- Convergence analysis: The paper also included a formal convergence analysis, which theoretically proves that the models trained with MingledPie converge effectively. This complements the empirical accuracy results, providing a strong theoretical foundation for the system's stability and reliability.
- Ablation experiments: To further validate the design choices, ablation experiments were conducted. These experiments specifically demonstrated the "effectiveness of our cluster model rebuilding method." This confirms that the intricate process of generating clues, aggregating the mingling coefficient matrix, and solving the linear system is indeed crucial and successful in recovering accurate cluster models from the mingled aggregates, rather than just relying on the mingling itself.
In summary, the experimental results presented by Cheng Zhang serve as a comprehensive proof of concept. They empirically confirm that MingledPie effectively mitigates preference profiling attacks by significantly reducing attack success rates, all while maintaining high cluster model accuracy and ensuring proper model convergence, thereby making it a viable and practical privacy-preserving solution for CFL.
Defensive Implications
▶ Watch: Summary of MingledPie's contributions (10:30)
MingledPie offers crucial defensive implications for organizations and practitioners deploying Clustered Federated Learning (CFL) systems, particularly those dealing with sensitive user data where inferring preferences could lead to privacy breaches or discriminatory practices. The core message for defenders is to recognize and actively mitigate the unique privacy risks associated with cluster identity in CFL.
- Prioritize Cluster Identity Protection: The most significant implication is the necessity to treat cluster identity as a sensitive piece of information. Traditional FL defenses often focus on local model updates, but MingledPie highlights that simply knowing which cluster a client belongs to can expose their preferences. Defenders must ensure that their CFL frameworks incorporate mechanisms to obfuscate or encrypt cluster assignments.
- Adopt Privacy-Preserving CFL Frameworks: Organizations should move beyond basic CFL implementations and adopt or develop frameworks that integrate advanced privacy-preserving techniques like MingledPie. This means seeking solutions that inherently support client-side cluster evaluation, generation of misleading cluster identities, and secure aggregation of model parameters.
- Leverage Homomorphic Encryption (HE): MingledPie's reliance on homomorphic encryption for protecting both local model updates and cluster identity clues is a key takeaway. Defenders should consider HE as a fundamental building block for secure aggregation in CFL. HE enables computations on encrypted data, allowing the server to perform its aggregation duties without ever learning sensitive plaintext information, thus eliminating a critical attack surface.
- Implement Robust Client-Side Mechanisms: The defense strategy in MingledPie is heavily reliant on client-side computations, including determining true cluster identity, generating false positives, and computing encrypted clues. This implies that client-side software must be robust, tamper-resistant, and correctly implement the cryptographic protocols. Secure execution environments on client devices may be necessary to prevent adversaries from subverting these client-side privacy mechanisms.
- Understand the Trade-off and Configuration: While MingledPie offers strong privacy, defenders must understand that parameters like the collusion probability
Pin generating false positives will influence the privacy-accuracy trade-off. A higherPmight offer stronger privacy but could potentially increase the computational overhead or complexity of the model rebuilding process. Organizations need to carefully configure these parameters based on their specific privacy requirements and performance constraints.
- Autonomous Deployment Benefits: MingledPie's design goal of "autonomous deployment without relying on additional third parties" is a significant advantage for defenders. It means that organizations can enhance the privacy of their CFL systems without incurring the operational overhead, trust assumptions, or potential single points of failure associated with external privacy-enhancing services or trusted execution environments. This simplifies deployment and reduces the attack surface.
- Beyond Basic Secure Aggregation: Defenders should understand that basic secure aggregation techniques, while vital for protecting model updates, are insufficient for CFL's unique cluster identity privacy problem. MingledPie demonstrates the need for a more comprehensive approach that considers the entire CFL workflow, from client grouping to model aggregation.
In essence, MingledPie provides a blueprint for building more privacy-aware CFL systems. It urges defenders to look beyond simple data anonymization and consider the deeper implications of structural information (like cluster membership) in distributed machine learning, offering a practical and theoretically sound method to protect user preferences without sacrificing the benefits of personalized models.
Key Takeaways
- Clustered Federated Learning (CFL) introduces a unique "cluster identity-based preference profiling attack" where an honest-but-curious server can infer client data preferences simply by knowing their cluster assignment.
- Existing FL privacy methods like secure aggregation and anonymous communication are insufficient to protect against this specific type of attack, as they do not obscure the client's cluster membership.
- MingledPie proposes "cluster mingling": Clients appear to belong to multiple clusters (true + false positives) simultaneously, effectively obfuscating their true data preferences from the server.
- Homomorphic encryption is crucial for MingledPie, enabling the server to aggregate encrypted local models and "cluster identity clues" without decrypting or learning individual client data, thus protecting privacy.
- A novel "cluster model rebuilding algorithm" is essential to reconstruct accurate cluster models from the noisy "mingled cluster models" by solving a system of linear equations derived from an encrypted mingling coefficient matrix.
- MingledPie significantly reduces attack success rates in preference profiling while incurring only "limited" accuracy loss across various datasets, demonstrating its practical effectiveness and minimal impact on model performance.
- The solution supports autonomous deployment, meaning it enhances CFL privacy without requiring additional communication infrastructure or reliance on external third parties.
About the Speaker(s)
The talk "MingledPie: A Cluster Mingling Approach for Mitigating Preference Profiling in CFL" was presented by Cheng Zhang. He introduced himself as "Pjang" at the beginning of the presentation. The transcript and metadata do not provide specific details about his title, affiliation, or other biographical information.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic security research tackling a real and underappreciated privacy vulnerability in Clustered Federated Learning — the cluster identity leakage problem is a genuine gap that existing secure aggregation doesn't close. The cryptographic machinery (HE + linear system rebuilding) is technically coherent, but this is a conference paper talk, not a practitioner-facing presentation, and the gap between the theoretical construction and anything a defender deploys tomorrow is wide.
Heather Calloway (CISO) — WEAK
Technically credible research that addresses a real and underappreciated privacy vulnerability in federated learning, but it never crosses the line from academic contribution to operational relevance. There is no governance framing, no institutional accountability, and no clear path for a security leader or privacy officer to act on this.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025