Impact Tracing: Identifying the Culprit of Misinformation in Encrypted Messaging Systems
Zhongming Wang
Network and Distributed System Security (NDSS) Symposium 2025 · Day 2 · Secure Protocols
Overview
The widespread adoption of end-to-end encryption (E2EE) in popular messaging platforms like WhatsApp, Signal, and iMessage has dramatically enhanced user privacy by concealing message content from the platforms themselves. While this is a crucial security feature, it inadvertently creates a significant challenge: the unchecked spread of problematic content, particularly misinformation and disinformation. Traditional content moderation techniques, often reliant on machine learning analysis of message content, are rendered ineffective by E2EE, leading to a fundamental tension between user privacy and the imperative for effective content moderation.
Key moments
- 0:00 Introduction: E2EE's privacy vs. misinformation challenge
- 2:15 Impact Tracing goal: Identifying influential spreaders
- 4:00 Core technical solution: Linking tag keys for tracing
- 6:20 Random response mechanism for user privacy
- 7:00 Decoding algorithm to identify spreaders from noisy graph
- 8:00 Performance evaluation and efficiency results
Impact Tracing: Identifying the Culprit of Misinformation in Encrypted Messaging Systems
Speakers: Zhongming Wang
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=F3pTm5Z-MPg
Overview
The widespread adoption of end-to-end encryption (E2EE) in popular messaging platforms like WhatsApp, Signal, and iMessage has dramatically enhanced user privacy by concealing message content from the platforms themselves. While this is a crucial security feature, it inadvertently creates a significant challenge: the unchecked spread of problematic content, particularly misinformation and disinformation. Traditional content moderation techniques, often reliant on machine learning analysis of message content, are rendered ineffective by E2EE, leading to a fundamental tension between user privacy and the imperative for effective content moderation.
Existing approaches to traceability in E2EE systems offer varying trade-offs. Solutions like "message franking" might only identify the last sender, while "message back" schemes could reveal the entire forwarding path, often at the cost of substantial user privacy. Conversely, "S-tracing" aims to reveal only the originator, providing strong privacy but limited insight into how misinformation propagates. This talk introduces Impact Tracing, a novel scheme designed to navigate this complex landscape. It aims to strike a delicate balance by identifying influential spreaders – a small group of users disproportionately responsible for disseminating misinformation – while rigorously preserving the privacy of non-influential users.
The significance of Impact Tracing lies in its pragmatic approach to a pressing societal problem. By focusing on influential actors, the scheme offers platforms a targeted mechanism to combat misinformation without resorting to pervasive surveillance or compromising the privacy of the vast majority of users who share harmless content. This work represents a crucial step towards enabling responsible content moderation within the privacy-preserving paradigm of E2EE, offering a path for platforms to address the spread of harmful narratives while upholding their commitment to user confidentiality.
Background
▶ Watch: Introduction: E2EE's privacy vs. misinformation challenge (0:00)
The evolution of digital communication has been marked by a continuous push for enhanced privacy, culminating in the widespread deployment of end-to-end encryption (E2EE) across major messaging applications. E2EE ensures that only the sender and intended recipient can read messages, rendering them inaccessible to the service provider. This architectural choice is paramount for protecting sensitive communications, but it simultaneously erects a formidable barrier to content moderation. When messages are encrypted, platforms are blind to their content, making it impossible to apply conventional machine learning-based methods for detecting and mitigating the spread of misinformation, hate speech, or other harmful content.
A significant feature contributing to the rapid dissemination of information—and misinformation—is message forwarding. On platforms like WhatsApp, messages can be forwarded, often to multiple chats simultaneously. Recognizing the potential for abuse, WhatsApp has implemented restrictions, such as limiting the forwarding of "multiple forwarded" messages to only one chat at a time. However, this blanket restriction applies to all messages, including benign content like memes or useful information, thus impeding legitimate information flow. This highlights the need for more nuanced solutions that can differentiate between problematic and harmless content dissemination.
The academic and industry communities have explored various traceability solutions for encrypted messaging to address this problem, each with distinct privacy implications:
- Message Franking: This approach typically involves the sender attaching a cryptographic "tag" to a message. If a message is reported, the platform can use this tag to identify the last sender of the message to the reporter. While this offers some accountability, it provides limited insight into the message's propagation path.
- Message Back: At the other end of the spectrum, "message back" schemes aim to reveal the entire determination path or forwarding graph of a reported message to the platform. While offering full traceability, this method comes at a significant cost to user privacy, as the platform learns extensive details about how a message spread.
- S-Tracing: In contrast, "S-tracing" focuses on revealing only the originator of a reported message. This approach prioritizes user privacy by minimizing the amount of information revealed about intermediate forwarders but offers poor traceability regarding the message's journey.
These existing works demonstrate a clear trade-off between traceability (the platform's ability to learn about message propagation) and user privacy (the platform learning nothing about forwarding paths). There is a pressing need for a "middle ground" solution that can provide sufficient traceability to combat harmful content without completely sacrificing user privacy. This talk identifies a crucial observation: a small group of users, termed influential spreaders, are disproportionately responsible for disseminating misinformation. For instance, a 2021 report by the Center for Democracy & Technology (CDT) indicated that 65% of anti-vaccine disinformation could be attributed to a relatively small number of influential actors across platforms like WhatsApp, Twitter, and Facebook. The goal of Impact Tracing is precisely to identify these influential spreaders while safeguarding the privacy of non-influential users, thereby addressing the core tension between privacy and moderation.
Key Findings
▶ Watch: Core technical solution: Linking tag keys for tracing (4:00)
The central contribution of this work is the introduction of Impact Tracing, a novel scheme designed to identify influential spreaders of misinformation within E2EE messaging systems while maintaining a delicate balance between traceability and user privacy. The key findings and contributions can be summarized as follows:
Firstly, Impact Tracing proposes a method for implicitly linking tag keys throughout a message's forwarding path. Unlike prior franking schemes where forwarded messages generate independent tags, Impact Tracing ensures that all subsequent tags are deterministically generated from an initial tracing key. This initial key is itself derived from a combination of the sender's secret identity key and the recipient's public identity, effectively binding the communication parties' identities to the tracing mechanism. This deterministic linkage allows the platform, upon a report, to computationally trace back the message's path.
Secondly, the scheme introduces the Funback (Forwarding Feedback) scheme, which, in its initial form, enables the platform to reconstruct the exact forwarding graph of a reported message. This is achieved by having users transmit message tags and tag keys along with their communications, which the platform stores. When a message is reported, the platform recursively computes and verifies tags for neighboring users, progressively building the full forwarding graph. This provides comprehensive traceability but, without further mechanisms, would compromise privacy.
Thirdly, to address the privacy concerns of the Funback scheme, Impact Tracing integrates a random response mechanism and a sophisticated decoding algorithm. A separate tag server is introduced, which responds to platform queries about message tag existence with controlled noise. If a tag truly exists, it returns "one"; if it doesn't, it still returns "one" with a certain probability, introducing randomness. This noise obscures the actual forwarding paths of non-influential users, leading to a "noisy graph" for the platform. The subsequent decoding algorithm then processes this noisy graph, leveraging inherent structural relationships within forwarding graphs (e.g., a node typically has only one parent) to robustly identify influential spreaders while preserving the privacy of others. The scheme formally proves that the noisy graph produced satisfies individualized differential privacy, meaning less influential users are effectively obscured by the introduced noise.
Finally, Impact Tracing demonstrates significant performance advantages over existing traceability schemes. Experimental evaluations show that it reduces bandwidth overhead by 70% compared to SOT tracing schemes and achieves a remarkable 94% reduction in storage overhead when compared to Meditback schemes. Furthermore, the system is efficient, capable of tracing a graph with 4,000 edges within 15 seconds and transmitting a 1-kilobyte message in just 0.3 milliseconds. The entire design has been formally analyzed for security, privacy, and utility, and the code is open-source and available on GitHub, fostering transparency and further research.
Technical Deep Dive
▶ Watch: Random response mechanism for user privacy (6:20)
Impact Tracing builds upon the foundational concept of message franking but significantly enhances it to enable full path traceability with privacy guarantees. The core technical innovation lies in how message tags are generated and linked across forwarding events, coupled with a noise injection mechanism and a decoding algorithm.
Initially, the talk describes a simple message franking scheme for context. In this basic setup, a sender randomly chooses a tag key (K) and generates a message tag using a random function f(K, message). The platform stores this message tag. To report a message, a recipient submits the tag key K, allowing the platform to verify the report by recomputing the tag and checking against its stored values. The function f being random ensures the platform learns nothing about the message content from the tag itself. However, this simple scheme fails for forwarded messages because each forwarder would independently generate their own tag key, breaking the link between the original message and its subsequent copies. The platform, receiving K3 from a reporter, could not compute K2 (from the previous forwarder) as there would be no relationship between them.
This limitation motivates the central design of Impact Tracing: implicitly linking tag keys across the forwarding path. Instead of independent tag generation, the scheme ensures that subsequent tags are deterministically derivable. This is achieved by modifying how the initial tag key is chosen. A tracing key is first calculated by combining the sender's secret identity key and the recipient's public identity. This implicitly binds the identities of the communicating parties to the tracing key. Once this initial tracing key is established, all subsequent tag keys along the forwarding path are deterministically generated from it. This means that if user A forwards to B, and B forwards to C, the tag key used by C is deterministically derived from the tag key used by B, which in turn is derived from the tag key used by A.
With this deterministic linkage, the system proceeds with two main phases:
- Messaging Phase: Users transmit messages along with trans_data, which includes the message tag and its corresponding tag key. The platform collects these communications and stores the message tags.
- Reporting and Tracing Phase (Funback Scheme): When a user reports a message, they submit the message tag and key to the platform. The platform, starting from the reporting user, can then leverage its knowledge of all users' secret and public identity keys (a strong assumption, implying a trusted setup or key management by the platform) to compute potential tag keys for that user's neighbors. It then queries its stored message tags to check if these computed tags exist, indicating a forward. This process is performed recursively: for every identified forwarder, the platform includes them in the next round of computation, tracing further back in the path until no new forwarders are found. This process yields the exact forward graph of the reported message to the platform. This component is termed the Funback (Forwarding Feedback) scheme.
While Funback provides full traceability, it compromises the privacy of all users involved in the forwarding chain. To protect the privacy of non-influential users, Impact Tracing introduces an additional layer:
- Random Response Mechanism: A dedicated tag server is deployed. Instead of the platform directly querying its own storage for tag existence, it queries this tag server. The tag server does not simply return a binary "yes" or "no." If a tag truly exists, it returns "one." However, if a tag does not exist, the tag server still returns "one" with a certain, predefined probability. This introduces random noise into the tracing results. Consequently, the platform obtains a noisy graph rather than the exact forward graph. This mechanism ensures individualized differential privacy for less influential users, meaning their specific forwarding actions are obscured by the noise.
- Decoding Algorithm: The goal is to identify influential spreaders from this noisy graph. The noise introduced by the tag server makes direct identification difficult. Therefore, a sophisticated decoding algorithm is designed to process the noisy graph. The intuition behind this algorithm is to leverage the inherent structural properties of forwarding graphs. For example, in a typical forwarding scenario, a node in the graph should have only one parent node (the user from whom they received the message). By applying these graph-theoretic relationships, the decoding algorithm can mitigate the noise and filter out false positives while preserving the privacy guarantees for non-influential users. The system is formally proven such that less influential users are obscured by the noise, while influential spreaders are accurately identified by the decoding algorithm.
The combination of deterministic tag key linking, the Funback scheme, the random response mechanism, and the decoding algorithm forms the complete Impact Tracing framework, balancing the complex requirements of traceability and privacy in E2EE messaging.
Demo / Proof of Concept
▶ Watch: Decoding algorithm to identify spreaders from noisy graph (7:00)
While the talk transcript does not detail a live demonstration or a specific "demo" section, the speaker explicitly states that the code is open source and available on GitHub. This indicates that a functional Proof of Concept (PoC) implementation of Impact Tracing exists, allowing researchers and developers to examine, verify, and potentially deploy the system.
The practical viability of Impact Tracing is further substantiated by its performance evaluation results, which were presented as part of the talk's findings. These metrics serve as a tangible demonstration of the system's efficiency and scalability:
- Bandwidth Overhead: Impact Tracing achieves a 70% reduction in bandwidth overhead compared to existing SOT (Sender-Only Tracing) schemes. This is a crucial metric for high-volume messaging platforms, as it directly impacts operational costs and network load.
- Storage Overhead: The scheme demonstrates a substantial 94% reduction in storage overhead when compared to Meditback schemes. This efficiency is vital for platforms that process billions of messages daily, as it minimizes the infrastructure required to store message tags and tracing data, especially given the assumption of storing data for a reasonable window (e.g., one month, potentially requiring 1TB for a platform like WhatsApp).
- Runtime Performance: The system is capable of tracing a forwarding graph with up to 4,000 edges within a mere 15 seconds. This rapid tracing capability is essential for timely responses to reported misinformation.
- Message Transmission: For individual message processing, Impact Tracing can transmit a 1-kilobyte message within 0.3 milliseconds, indicating minimal latency overhead for regular message flow.
These performance figures, derived from the implemented system, underscore the practical applicability and efficiency of Impact Tracing, making it a viable candidate for real-world deployment in large-scale encrypted messaging environments.
Defensive Implications
▶ Watch: Performance evaluation and efficiency results (8:00)
Impact Tracing offers significant defensive implications for platforms grappling with the spread of misinformation in E2EE environments. By providing a mechanism to identify influential spreaders without compromising the privacy of the broader user base, platforms can adopt a more targeted and effective approach to content moderation.
Firstly, platforms can leverage Impact Tracing to enforce their content policies more effectively. When a message is reported as misinformation, the system allows the platform to trace its propagation path and pinpoint the key actors responsible for its widespread dissemination. This enables platforms to take specific actions against these influential spreaders, such as issuing warnings, temporary suspensions, or permanent bans, rather than resorting to blunt measures that impact all users or trying to moderate content they cannot see.
Secondly, the scheme's focus on individualized differential privacy for non-influential users means that platforms can maintain their commitment to user privacy while still addressing harmful content. This helps to mitigate the tension between privacy advocates and those calling for stronger content moderation, offering a balanced solution that could be more palatable to diverse stakeholders. Users can be confident that their private communications, especially those not related to widespread misinformation, remain protected.
Thirdly, the efficiency demonstrated by Impact Tracing in terms of bandwidth, storage, and runtime means that it is operationally feasible for large-scale deployments. Reducing storage overhead by 94% and bandwidth by 70% compared to other traceability schemes makes it a more attractive option for platforms like WhatsApp, which handle immense volumes of messages. The ability to trace 4,000 edges in 15 seconds also ensures that platforms can respond swiftly to emerging misinformation campaigns.
However, the talk also touched upon potential challenges, such as sybil attacks, where influential spreaders might attempt to circumvent identification by using multiple accounts or identities. The speaker acknowledged this, noting that creating and maintaining extensive social networks across multiple accounts on a platform like WhatsApp incurs a "cost" (e.g., building contact lists). While Impact Tracing focuses on scenarios where users operate with established networks, this remains a consideration for platforms in their broader anti-abuse strategies. Defenders would need to combine Impact Tracing with other sybil detection mechanisms to provide a more robust defense.
Finally, the discussion around storage windows (e.g., storing message tags for one month) provides practical guidance for implementation. Platforms would need to define a reasonable retention policy for tracing data, balancing the need for historical tracing with storage costs and privacy considerations. The estimated 1TB storage requirement for a platform like WhatsApp for one month of tags offers a concrete figure for planning. Overall, Impact Tracing empowers platforms to move beyond reactive content takedowns to proactive identification of key misinformation vectors, fostering a healthier information ecosystem within the confines of strong encryption.
Key Takeaways
- Balanced Traceability and Privacy: Impact Tracing provides a novel solution that effectively balances the tension between end-to-end encryption's privacy guarantees and the critical need for content moderation against misinformation.
- Targeted Identification of Influential Spreaders: The scheme is specifically designed to identify "influential spreaders" who disproportionately contribute to misinformation, rather than exposing the entire forwarding path of every message.
- Deterministic Tag Key Linking: A core technical innovation involves deterministically linking message tag keys along the forwarding path, enabling the platform to computationally trace a message's origin and propagation.
- Individualized Differential Privacy: Impact Tracing employs a random response mechanism within a dedicated tag server and a decoding algorithm to introduce noise into tracing results, ensuring that non-influential users' forwarding activities are obscured while influential ones are accurately identified.
- Significant Performance Improvements: The system demonstrates substantial practical efficiency, reducing bandwidth overhead by 70% and storage overhead by 94% compared to prior art, making it suitable for large-scale messaging platforms.
- Open-Source and Formally Analyzed: The Impact Tracing code is open-source, and the design has undergone formal analysis for its security, privacy, and utility, promoting transparency and trust in its capabilities.
About the Speaker(s)
Zhongming Wang is the presenter of this talk on Impact Tracing. The transcript indicates his affiliations with the University of Joan and Singapore Management University, suggesting a background in academic research within computer science, likely focusing on security, privacy, and distributed systems. His work, as presented, showcases expertise in cryptographic protocols and data privacy mechanisms within real-world applications like encrypted messaging.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic crypto work on a real problem — the privacy/traceability tension in E2EE messaging is genuinely hard and the differential privacy angle with deterministic tag linking is technically defensible. But this is a conference paper presentation, not a security research talk with offensive teeth, and the abstract-to-implementation gap leaves open the questions that actually matter for deployment.
Heather Calloway (CISO) — WEAK
Technically interesting cryptographic research on traceability in E2EE systems, but it stays entirely inside the academic frame. The governance questions this work raises — who authorizes tracing, who controls the tag server, what legal regime permits this — are untouched, and there is no path from the research to an actionable decision for any operator or policymaker.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025