EmbedX: Embedding-Based Cross-Trigger Backdoor Attack Against Large Language Models
Nan Yan (Wuhan University)
34th USENIX Security Symposium (USENIX Security '25) · Day 1 · LLM Security and Attacks
Overview
The rapid advancements in large language models (LLMs) such as GPT-4, LLaMA, and GPT-2 have revolutionized numerous natural language processing (NLP) tasks, from machine translation to question answering and sentiment analysis. Despite their impressive capabilities, these sophisticated models remain highly susceptible to backdoor attacks, where specific "triggers" can manipulate the model into producing targeted, malicious outputs. This talk, "EmbedX: Embedding-Based Cross-Trigger Backdoor Attack Against Large Language Models," presented by Nan Yan, unveils a novel and highly potent attack vector that significantly challenges the security posture of contemporary LLMs.

Key moments
- 0:00 Introduction: Limitations of single-trigger backdoor attacks
- 2:00 EmbedX: Optimizable soft triggers for cross-trigger attacks
- 3:30 EmbedX's three stages: Soft trigger, latent injection, activation
- 6:00 Efficient backdoor activation with new tokens, no retraining
- 8:00 Evaluation results: EmbedX's superior attack success and efficiency
- 10:00 Cross-lingual attacks and robustness against forgetting
- 11:00 EmbedX's stealth in latent space, hard to detect
- 12:00 Limitations of existing defenses against EmbedX
EmbedX: Embedding-Based Cross-Trigger Backdoor Attack Against Large Language Models
Speakers: Nan Yan
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=gORZlaRLYPM
Overview
The rapid advancements in large language models (LLMs) such as GPT-4, LLaMA, and GPT-2 have revolutionized numerous natural language processing (NLP) tasks, from machine translation to question answering and sentiment analysis. Despite their impressive capabilities, these sophisticated models remain highly susceptible to backdoor attacks, where specific "triggers" can manipulate the model into producing targeted, malicious outputs. This talk, "EmbedX: Embedding-Based Cross-Trigger Backdoor Attack Against Large Language Models," presented by Nan Yan, unveils a novel and highly potent attack vector that significantly challenges the security posture of contemporary LLMs.
Traditional backdoor attacks typically rely on single, discrete natural language tokens as triggers. This approach suffers from critical limitations: a lack of adaptability to diverse user demographics (e.g., "trunk" vs. "lorry" for different English speakers), difficulty in optimization, and the prohibitive cost of retraining models to accommodate new triggers or linguistic variations. EmbedX addresses these shortcomings by introducing a soft trigger within the model's continuous embedding space, enabling an unprecedented level of stealth, efficiency, and cross-trigger, cross-lingual attack capabilities. The research, a collaborative effort between Wuhan University, Hajj University of Science and Technology, and the Hong Kong University of Science and Technology, exposes a profound vulnerability in how LLMs process and interpret input, demonstrating that robust, multi-faceted backdoors can be injected and activated with alarming ease and efficacy, posing a significant threat to the integrity and trustworthiness of deployed LLMs.
Background
▶ Watch: Introduction: Limitations of single-trigger backdoor attacks (0:00)
Large language models have become ubiquitous, forming the backbone of countless applications. However, their increasing complexity and reliance on vast datasets also introduce new attack surfaces, particularly through backdoor attacks. In a typical backdoor scenario, a malicious actor injects specific triggers into a model during its training phase. When these triggers are present in a user's input, the model deviates from its intended behavior and produces a predetermined, malicious output, while otherwise functioning normally on clean inputs.
Existing backdoor attacks on LLMs primarily suffer from several fundamental limitations that EmbedX aims to overcome. Firstly, they often depend on single triggers, which severely restricts their real-world applicability. For instance, a trigger like "trunk" might activate the backdoor for American English speakers but fail for British English speakers who use "lorry." Similarly, language-specific triggers like "SZ" in English appear unnatural in Chinese or Korean contexts, making them easily detectable and limiting their reach in a globalized user base. This lack of cross-trigger capability means that as user diversity (linguistic, cultural, stylistic) increases, the effectiveness of single-trigger attacks diminishes rapidly.
Secondly, these conventional attacks typically utilize natural language tokens as triggers. While intuitive, these discrete tokens are inherently difficult to optimize for stealth and robustness. More critically, modifying or adding new triggers—for example, switching from "trunk" to "lorry"—requires retraining the entire model, a process that is both computationally expensive and time-consuming. This inefficiency renders such attacks impractical for dynamic or evolving attack scenarios. The need for an efficient and optimizable trigger that can adapt to multiple linguistic and cultural contexts without costly retraining is paramount.
The threat model for EmbedX, as described by the speaker, assumes a malicious provider uploads backdoored models to platforms. This provider has control over the training data, model parameters, and deployment environment. The goal is to inject a backdoor that makes the model appear normal on clean inputs but activates a targeted malicious behavior when specific, subtle triggers are presented. The attacker's objectives are threefold:
- Effectiveness: Achieve a high success rate across multiple, diverse triggers.
- Stealth: Minimize false triggers and ensure a negligible drop in the model's accuracy on clean tasks.
- Efficiency: Enable the activation of new triggers without requiring any retraining of the compromised model. EmbedX was designed from the ground up to meet these stringent requirements, fundamentally altering the landscape of LLM backdoor vulnerabilities.
Key Findings
▶ Watch: EmbedX's three stages: Soft trigger, latent injection, activation (3:30)
EmbedX introduces a groundbreaking approach to LLM backdoor attacks, leveraging the continuous nature of the model's embedding space to create highly effective, stealthy, and remarkably efficient soft triggers. The key findings demonstrate its superior capabilities compared to existing methods:
Firstly, EmbedX achieves an exceptional 100% attack success rate across a wide array of triggers and languages. This uniform effectiveness is maintained across diverse text classification and generation tasks, showcasing its versatility and robustness. Crucially, this high success rate does not come at the expense of model utility; EmbedX consistently maintains high clean task accuracy, often outperforming baselines and, in some evaluations, even improving the model's accuracy by up to 3.2% on clean samples.
Secondly, the efficiency of EmbedX is a standout feature. It enables cross-trigger switching in an astonishingly short time—under one second, with a reported average of just 0.53 seconds. This is orders of magnitude faster than traditional token-based methods, which necessitate expensive and time-consuming retraining for each new trigger. This sub-second switching capability is achieved by reusing the central soft trigger without retraining, making the attack highly adaptable and dynamic.
Thirdly, EmbedX demonstrates remarkable robustness against various perturbations and even under heavy fine-tuning. While baselines might lose 10-40% attack success rate when models are fine-tuned, EmbedX maintains nearly 100% attack success rate, with only a 13% drop even under aggressive fine-tuning scenarios. This resilience ensures the backdoor persists and remains effective even after legitimate model updates or adaptations.
Furthermore, the attack exhibits an extremely low false trigger rate, reported at only 1% with just 5% poisoning for cross-lingual attacks. This high precision ensures that the backdoor is only activated when intended, preserving the stealth of the attack. EmbedX also exposes significant cross-lingual security risks, achieving an average 99.25% attack success rate across multiple European (English, Spanish) and Asian (Chinese, Korean) languages. It performs even better in European languages, demonstrating that triggers can transfer across languages, manipulating poisoned outputs while keeping clean inputs unaffected.
Finally, the stealth of EmbedX is enhanced by its innovative use of dual constraints in both the frequency and gradient domains during the latent adversary backdoor injection phase. Metrics like Layer-wise Frequency Discrepancy (LFD) and Gradient Discrepancy (LGD) show that EmbedX keeps the latent representations of poisoned samples remarkably close to those of clean samples. LFD remains as low as the clean model, and LGD stays near zero in shallow layers, making the embedded backdoor incredibly difficult to detect using anomaly-based detection methods that rely on identifying deviations in latent space.
Technical Deep Dive
▶ Watch: Evaluation results: EmbedX's superior attack success and efficiency (8:00)
EmbedX operates on a sophisticated, multi-stage pipeline designed to inject and activate backdoors with unparalleled stealth and efficiency. The core innovation lies in moving beyond discrete token-based triggers to leverage the continuous nature of the model's embedding space.
The attack pipeline is structured into three key stages:
Stage 1: Weaponizing Embeddings as Soft Trigger
At the foundational level, EmbedX identifies and exploits the continuous embedding vector from the LLM's embedding layer as the soft trigger. Unlike discrete tokens, which are fixed entries in a vocabulary, this soft trigger is a continuous vector that can be optimized. The ingenious aspect here is that the attacker freezes the large model parameters and only optimizes this soft trigger. This strategic choice is critical for efficiency, as it bypasses the need for costly full model retraining. The backdoor is designed to be activated directly by this soft trigger in the embedding layer, rather than by specific natural language tokens themselves. The primary objective during this stage is to optimize this soft trigger to maintain high effectiveness, ensure stealth, and achieve robustness against potential defenses.
Stage 2: Latent Adversary Backdoor Injection
Once the soft trigger is weaponized, the next stage involves stealthily injecting the backdoor into the LLM's latent space. This is achieved through a process of poison training. The soft trigger is incorporated into the model, and the training objective is to align the model's output with the attacker's target output when the trigger is active.
A crucial aspect of this stage is the use of dual constraints applied to the latent features:
- Frequency Domain Constraints: The latent features of poisoned samples are constrained in the frequency domain. This is measured by the Layer-wise Frequency Discrepancy (LFD), which quantifies the difference in frequency distributions between clean and poisoned samples in the latent space. By minimizing LFD, EmbedX ensures that the statistical properties of poisoned samples in the frequency domain remain similar to clean samples, making them harder to distinguish.
- Gradient Domain Constraints: Similarly, latent features are constrained in the gradient domain. This is measured by the Gradient Discrepancy (LGD), which aims to keep the gradients of poisoned samples close to zero in shallow layers. This ensures that the model's internal processing of poisoned inputs closely resembles that of clean inputs, further enhancing stealth by masking the presence of the backdoor from gradient-based detection methods.
To achieve both attack effectiveness and stealth simultaneously, the researchers introduce an adversary loss function during poison training. This loss term is designed to maximize the attack success rate while minimizing the impact on the model's utility performance on clean tasks, which is preserved through a clean loss component. The overall optimization problem for backdoor implementation balances these competing objectives, ensuring the backdoor is deeply embedded yet minimally disruptive to normal operations. The dual constraints in the frequency and gradient domains are fundamental to achieving this stealth, allowing poisoned and clean samples to overlap significantly in the latent space, making detection extremely challenging.
Stage 3: Backdoor Activation
The final stage addresses how the soft trigger, now deeply embedded in the model's latent space, is activated by user input. In practice, attackers don't directly inject the continuous embedding vector. Instead, they use readily available "fuse" tokens—these can be real words, common misspellings, or even frequently used phrases. To enable these diverse tokens to activate the backdoor, EmbedX introduces an alignment loss.
The alignment loss optimizes the embedding representation of a new token (the "fuse") to align precisely with the pre-trained soft trigger in the embedding space. This is a critical innovation: by optimizing only the embedding of the fuse token to match the soft trigger, the attacker can efficiently specify multiple tokens capable of activating the backdoor without any retraining of the large language model itself. This is what enables the sub-second cross-trigger switching capability and the linguistic flexibility of EmbedX. An English fuse like "trunk" can be aligned, and then a Spanish fuse like "camión" can be aligned to the same soft trigger, all without modifying the core LLM.
Evaluation Metrics and Comparison
To thoroughly assess EmbedX, the researchers proposed two specific metrics: Layer-wise Frequency Discrepancy (LFD) and Gradient Discrepancy (LGD). These metrics quantify the differences between clean and poisoned samples within the latent space, providing objective measures of stealth. A lower LFD and LGD indicate greater stealth, as the poisoned samples are harder to distinguish from clean ones.
EmbedX was compared against five representative backdoor attack methods on LLMs: BetterNets, standard trigger injection, CBA dual triggers at different positions, Slipper Agent (contextual triggers), Embedding Poisoning (super word vector replacement), and Soft Prompt Injecting (optimizable prompts). The results consistently demonstrated EmbedX's superiority in attack success rate (100%), clean test accuracy (often higher than baselines), and efficiency (cross-trigger switching in 0.53 seconds, orders of magnitude faster than retraining). This comprehensive technical framework allows EmbedX to achieve high effectiveness, stealth, and unprecedented efficiency in cross-trigger, cross-lingual backdoor attacks.
Demo / Proof of Concept
▶ Watch: Cross-lingual attacks and robustness against forgetting (10:00)
While the presentation did not feature a live, interactive demonstration in the traditional sense, the researchers provided extensive empirical evidence and detailed evaluation results that collectively serve as a robust proof of concept for EmbedX's capabilities. These evaluations were conducted across a broad spectrum of real-world scenarios, demonstrating the attack's effectiveness, stealth, and efficiency.
The study utilized five real-world datasets encompassing a wide range of NLP tasks, including both text classification and generation tasks. This diversity ensured that EmbedX's performance was not limited to a narrow application domain. Furthermore, the attack was evaluated on four representative open-sourced models, confirming its generalizability across different LLM architectures.
A key aspect of the proof of concept was demonstrating EmbedX's cross-lingual attack capabilities. The researchers specifically tested the attack on three European languages (English and Spanish) and three Asian languages (Chinese and Korean). The results were compelling: EmbedX achieved an impressive 99.25% attack success rate across these diverse languages, performing even better in European languages. This showcases the attack's ability to transfer triggers across languages, maintaining the integrity of clean inputs while successfully manipulating poisoned ones. The low false trigger rate of just 1% with only 5% poisoning further underscored its precision and stealth.
The efficiency claims were also substantiated through rigorous evaluation. EmbedX demonstrated the ability to switch between triggers in an average of 0.53 seconds, a stark contrast to the hours or days required for retraining in token-based methods. This sub-second switching capability, achieved without needing extra negative poisoning like some baselines (e.g., CBA), highlights a critical advantage for attackers seeking agile and adaptable backdoors. The robustness of EmbedX was also proven by its ability to maintain nearly 100% attack success rate even under heavy fine-tuning, where baselines saw significant drops (10-40%).
The stealth of EmbedX was quantitatively supported by the proposed metrics. The Layer-wise Frequency Discrepancy (LFD) remained as low as that of the clean model, and Gradient Discrepancy (LGD) stayed near zero in shallow layers. This indicated that poisoned and clean samples largely overlapped in the latent space, making EmbedX's backdoor extremely difficult to detect through anomaly analysis. In contrast, baselines exhibited significantly higher LFD and LGD values, indicating less stealth. These comprehensive evaluations across multiple models, datasets, languages, and performance dimensions served as a compelling demonstration of EmbedX's unprecedented capabilities and the serious security implications it presents for large language models.
Defensive Implications
▶ Watch: Limitations of existing defenses against EmbedX (12:00)
The introduction of EmbedX highlights significant shortcomings in current defense strategies against backdoor attacks on large language models. The research explored both novel and existing defense mechanisms, revealing that most are either ineffective against EmbedX or come at an unacceptable cost to clean task accuracy.
The researchers proposed and evaluated two potential defense methods at different levels:
- Word-level defense: This approach attempts to remove tokens that lower perplexity, aiming to catch rare or misspelled triggers often associated with backdoor attacks. While this defense could cut EmbedX's attack success rate by nearly 60%, it suffers from a high detection success rate (DSR) of approximately 17% with a significant number of false alarms, rendering it impractical for real-world deployment. Crucially, it also incurred a substantial cost to clean test accuracy, dropping it by nearly 25%. This demonstrates a classic trade-off where improved security comes at a steep price to model utility.
- Embedding-level defense: This more sophisticated defense flags tokens with abnormal variances in their embedding representations, attempting to identify the subtle deviations introduced by the soft trigger. This method proved more promising, reducing EmbedX's attack success rate by 14% to 28% while keeping the clean test accuracy almost unchanged, with drops of less than 3%. Although more effective and less detrimental to utility, this defense still allowed a significant portion of attacks to succeed, indicating it is not a complete solution.
Furthermore, the researchers tested two existing state-of-the-art defenses against EmbedX:
- TextGuard: This defense method was found to cut the attack success rate to 62% for single triggers. However, its effectiveness diminished as the number of triggers grew, and it was limited to classification tasks, failing to protect against generation task backdoors. This highlights EmbedX's advantage in cross-trigger and multi-task scenarios.
- BAR (Backdoor Attack Removal): BAR managed to lower EmbedX's attack success rate to 82%. However, EmbedX's innovative use of dual latent constraints (in frequency and gradient domains) keeps the poisoned embeddings extremely close to clean ones, significantly reducing BAR's impact. The subtle nature of EmbedX's soft trigger in the continuous embedding space makes it exceptionally difficult for such defenses to isolate and neutralize.
Overall, the findings underscore a critical gap in current LLM security. Existing defenses either impose an unacceptable degradation on the model's performance on legitimate tasks or simply fail to robustly counter the stealth and adaptability of EmbedX. The ability of EmbedX to seamlessly integrate poisoned and clean samples in the latent space, coupled with its cross-trigger and cross-lingual capabilities, necessitates the development of fundamentally stronger and more adaptive defense strategies. Future research directions suggested by the speaker include extending the attack to richer real-world user diversity, exploring complex and low-resource languages, and designing more adaptive soft triggers for broader effectiveness, all of which will further challenge existing defense paradigms.
Key Takeaways
- Novel Embedding-Based Attack: EmbedX introduces a groundbreaking approach to LLM backdoor attacks by utilizing a "soft trigger" within the continuous embedding space, moving beyond traditional discrete token-based triggers.
- Unprecedented Efficiency and Adaptability: The attack enables cross-trigger switching in under one second (0.53s average), allowing multiple activation tokens across languages and contexts without requiring any model retraining, a significant leap in efficiency.
- High Effectiveness and Stealth: EmbedX achieves a 100% attack success rate while maintaining high clean task accuracy (sometimes improving it by 3.2%) and an extremely low false trigger rate (1%). Its stealth is enhanced by dual constraints in frequency and gradient domains, making detection challenging.
- Cross-Lingual and Cross-Model Capabilities: The attack demonstrates robust performance across diverse models and languages, achieving a 99.25% attack success rate in multilingual contexts (English, Spanish, Chinese, Korean), exposing significant cross-lingual security risks.
- Defenses Are Insufficient: Existing state-of-the-art defenses (e.g., TextGuard, BAR) are largely ineffective against EmbedX, either significantly hurting clean model accuracy or failing to neutralize the attack due to its deep integration and stealthy nature, highlighting a critical need for new countermeasures.
- Robustness Against Fine-tuning: EmbedX maintains nearly 100% attack success rate even under heavy fine-tuning, demonstrating its resilience and persistence within compromised models, making it a long-lasting threat.
About the Speaker(s)
Nan Yan is a researcher affiliated with Wuhan University, with contributions also noted from Hajjo University of Science and Technology and the Hong Kong University of Science and Technology for this joint work. The presentation highlights their expertise in the field of large language model security, specifically focusing on the vulnerabilities related to backdoor attacks. Their research, as demonstrated by EmbedX, delves into sophisticated methods for manipulating LLMs, emphasizing areas such as embedding space exploitation, cross-trigger mechanisms, and the stealthy injection of adversarial components into machine learning models.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
EmbedX is genuine, novel research that advances the LLM backdoor attack space in a meaningful direction — moving triggers into continuous embedding space is the right instinct and the cross-lingual, sub-second trigger-switching result is the kind of finding that will make ML security teams uncomfortable. The defensive analysis is honest about current gaps rather than strawmanning weak baselines, which earns extra credit.
Heather Calloway (CISO) — WEAK
Technically rigorous research on a real and underappreciated attack class against LLMs — but it stops at the lab bench. There is no bridge to the organizations deploying these models, the procurement decisions that create exposure, or the governance structures that would need to respond.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)