BOLT: Privacy-Preserving, Accurate and Efficient Inference for Transformers
Qi Pang, Jinhao Zhu, Helen Möllering, Wenting Zheng, Thomas Schneider
IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 6
Overview
The proliferation of powerful Transformer-based models, such as GPT, BERT, and ViT, has fueled the rapid expansion of Machine Learning as a Service (MLaaS). While these models offer unprecedented performance in tasks ranging from chatbots to translation, their deployment raises significant privacy concerns. Incidents like ChatGPT's temporary ban in Italy due to user data exposure highlight the critical need for robust privacy protections in MLaaS. Users' sensitive inputs and chat histories, if revealed in plaintext, can leak personal identities, hobbies, and even commercial secrets, posing substantial risks to individual privacy and corporate data security.

Key moments
- 0:00 Introduction and privacy problem for Transformers
- 2:50 Hybrid protocol: HE for linear, 2PC for nonlinear layers
- 3:50 Prior work limitations and BOLT's 10x performance improvement
- 4:25 BOLT's innovations: matrix multiplication, crypto-friendly approximations
- 5:00 BOLT system overview: client encryption, server processing
- 8:00 Optimizing linear layers: eliminating rotations in homomorphic encryption
BOLT: Privacy-Preserving, Accurate and Efficient Inference for Transformers
Speakers: Qi Pang, Second Year PhD Student, Kon; Jinhao Zhu; Helen Möllering; Wenting Zheng; Thomas Schneider
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=myteCasy8z0
Overview
The proliferation of powerful Transformer-based models, such as GPT, BERT, and ViT, has fueled the rapid expansion of Machine Learning as a Service (MLaaS). While these models offer unprecedented performance in tasks ranging from chatbots to translation, their deployment raises significant privacy concerns. Incidents like ChatGPT's temporary ban in Italy due to user data exposure highlight the critical need for robust privacy protections in MLaaS. Users' sensitive inputs and chat histories, if revealed in plaintext, can leak personal identities, hobbies, and even commercial secrets, posing substantial risks to individual privacy and corporate data security.
Addressing this urgent challenge, the talk introduces BOLT, a novel privacy-preserving system designed for efficient and accurate inference on Transformer architectures. BOLT operates under a semi-honest party model, where both the client (with private input) and the server (with a confidential model) honestly follow the protocol but are curious about the other party's data. By leveraging a hybrid cryptographic approach combining Homomorphic Encryption (HE) and Secure Multi-Party Computation (2PC), BOLT encrypts all exchanged messages, revealing only the final output to the client.
BOLT significantly advances the state-of-the-art in secure Transformer inference by co-designing machine learning and cryptographic primitive optimizations. It achieves an impressive 10x improvement in runtime performance and an 11x reduction in communication overhead compared to prior work. This makes privacy-preserving inference for complex models like BERT Base not just theoretically possible, but practically viable, paving the way for wider and more secure deployment of advanced AI services.
Background
▶ Watch: Introduction and privacy problem for Transformers (0:00)
The problem of privacy in MLaaS stems from the inherent trust imbalance: clients must send sensitive data to a server to utilize its model, or servers must reveal their proprietary models to clients for local inference. Both scenarios compromise data confidentiality. Cryptographic techniques offer a solution, with Homomorphic Encryption (HE) allowing computations on encrypted data and Secure Multi-Party Computation (2PC) enabling parties to jointly compute a function over their private inputs without revealing them.
HE is particularly efficient for linear operations, offering relatively lower communication costs but higher computational overhead. Conversely, 2PC, especially secret sharing-based protocols, excels at evaluating complex non-linear functions, though often with higher communication overhead. To harness the strengths of both, state-of-the-art secure ML inference systems, including BOLT, adopt a hybrid protocol. This typically involves using HE for linear layers (e.g., matrix multiplications in Transformers) and 2PC for non-linear activation functions (e.g., Softmax, GELU). Conversions between HE ciphertexts and 2PC shares are performed as needed.
Prior work in secure Transformer inference, such as the system "Iron," also employed a hybrid HE/2PC approach. However, Iron and similar systems faced significant efficiency bottlenecks. For instance, Iron required an astounding 81 GB of communication and approximately 3.5 hours of runtime for a single secure inference on a BERT Base model in a wide area network (WAN) setting. These inefficiencies arose from sparse ciphertext packings, which led to a large number of useless polynomial coefficients being transmitted, and complex non-linear function evaluations. BOLT specifically tackles these limitations by introducing novel packing strategies, crypto-friendly approximations, and machine learning optimizations to drastically reduce communication and computation costs.
Key Findings
▶ Watch: Prior work limitations and BOLT's 10x performance improvement (3:50)
BOLT presents several key findings and contributions that significantly push the boundaries of practical privacy-preserving Transformer inference:
- 10x Runtime Improvement: BOLT achieves a remarkable 5.2x speedup in local area network (LAN) settings and an even more significant 9.5x speedup in wide area network (WAN) settings compared to state-of-the-art systems like Iron for BERT Base inference. This drastically reduces the time required for secure predictions.
- 11x Communication Reduction: The system slashes communication overhead by nearly 11 times compared to Iron, addressing a major bottleneck in secure multi-party computation. This is crucial for deploying secure ML in real-world scenarios with limited bandwidth.
- Novel Ciphertext-Plaintext Matrix Multiplication: BOLT introduces an innovative packing method for ciphertext-plaintext matrix multiplications in HE that minimizes the number of ciphertexts exchanged and optimizes rotations using the Baby-Step Giant-Step (BSGS) algorithm.
- Crypto-Friendly Approximations and Polynomial Pre-processing: For complex non-linear functions, BOLT designs approximations and leverages Mosin's polynomial pre-processing technique. This reduces the number of expensive multiplications between shares in 2PC, enhancing efficiency without sacrificing accuracy.
- Oblivious Word Elimination: Incorporating a machine learning-driven optimization, BOLT introduces an oblivious method to identify and eliminate unimportant tokens in the input. This effectively halves the computational cost for subsequent Transformer blocks with only a negligible impact on model accuracy.
- Accuracy Preservation: Despite aggressive optimizations, BOLT maintains accuracy levels comparable to plaintext models and other secure inference systems like Iron, demonstrating that efficiency gains do not come at the expense of model utility.
Technical Deep Dive
▶ Watch: BOLT's innovations: matrix multiplication, crypto-friendly approximations (4:25)
BOLT's efficiency stems from its meticulously engineered hybrid protocol and a suite of cryptographic and machine learning optimizations. The system processes an input sentence by first tokenizing and embedding it into a matrix on the client side. This matrix is then encrypted using Homomorphic Encryption (HE) and sent to the server.
The core of BOLT's hybrid approach is to evaluate linear layers (e.g., matrix multiplications in self-attention and feed-forward networks) using HE on the server, while non-linear layers (e.g., Softmax, GELU) are computed using 2PC. This strategic division capitalizes on HE's efficiency for linear operations and 2PC's capability for complex non-linear functions. The system uses the BFV (Brakerski/Fan-Vercauteren) homomorphic encryption scheme, which encodes vectors of integers into polynomials in the NTT (Number Theoretic Transform) space. BFV supports element-wise additions and multiplications between ciphertexts or ciphertexts and plaintexts, as well as cyclic rotations of ciphertext elements.
Novel Ciphertext-Plaintext Matrix Multiplication
A major innovation in BOLT is its approach to secure matrix multiplication within HE for linear layers. Prior work like Iron encoded values into the coefficient space of polynomials to eliminate rotations, but this resulted in a large number of "useless" polynomial coefficients, leading to sparse ciphertexts and high communication overhead (e.g., 56 ciphertexts for a BERT Base linear layer).
BOLT takes a different approach:
- Compact Packing: The input ciphertext matrix
Xis packed column-wise, and the plaintext matrixY(representing model weights) is packed diagonally, with each element repeatedMtimes (whereMis the number of rows inX). Values are encoded in the efficient NTT space. - Minimized Rotations: Instead of eliminating rotations entirely, which leads to sparse ciphertexts, BOLT strategically uses rotations. For the concrete example given, only one rotation is needed, followed by two multiplications. The crucial insight is that both input and output ciphertexts remain compact, minimizing the data transmitted.
- Optimized Rotation Count with BSGS: To further reduce the overall computational cost, BOLT incorporates the Baby-Step Giant-Step (BSGS) algorithm. Instead of only rotating the input ciphertext, BSGS distributes rotations to partial results, optimizing the total number of rotations required. Since rotations and multiplications have comparable costs, this optimization significantly reduces communication without a substantial increase in computational overhead (estimated at about 5%). This technique reduces the number of ciphertexts exchanged for a linear layer in BERT Base to just 13, a significant improvement over Iron's 56.
Polynomial Pre-processing for Nonlinear Layers
Non-linear activation functions, such as GELU (Gaussian Error Linear Unit), are computationally expensive in 2PC as they often require multiple multiplications between shares. BOLT approximates the GILU function (which is a variant of GELU) using a 4-degree polynomial. A naive evaluation of a 4-degree polynomial in 2PC would require three multiplications between shares.
BOLT leverages Mosin's polynomial pre-processing technique to reduce this cost. Since the coefficients of the approximating polynomial are publicly known, Mosin's method allows constructing two sub-polynomials from the public coefficients. By doing so, the evaluation of a public polynomial of degree n can be approximated to require only n/2 multiplications between shares. For the 4-degree GILU approximation, this reduces the required multiplications from three to just two, significantly improving the efficiency of non-linear layer computations in 2PC.
Machine Learning Optimizations: Oblivious Word Elimination
Beyond cryptographic optimizations, BOLT incorporates a novel machine learning technique called oblivious word elimination. The premise is that not all tokens in an input sentence contribute equally to the final prediction. For instance, in sentiment analysis, words like "charming" and "fascinating" are more indicative than filler words.
The self-attention mechanism within Transformer models naturally provides a way to quantify token importance. Specifically, the Softmax output in the first Transformer block measures pairwise correlations between tokens. BOLT utilizes this:
- Importance Score Calculation: By summing the attention matrix along its rows, an importance score vector is derived for each token.
- Oblivious Sorting: This importance score vector is then used as a key to obliviously sort the rows of the attention matrix in 2PC, for example, using a Bitonic Sort algorithm.
- Token Elimination: After sorting, the less important half of the tokens (and their corresponding rows) are obliviously dropped. This means that for all subsequent Transformer computations, the model only processes half the original number of tokens, effectively reducing the computational cost by more than half. Crucially, the evaluation demonstrates that this aggressive pruning has a negligible influence on the model's accuracy.
Demo / Proof of Concept
▶ Watch: BOLT system overview: client encryption, server processing (5:00)
The efficacy and performance of BOLT were rigorously evaluated using a standard Transformer model and benchmark datasets.
Evaluation Setup:
- Model: BERT Base, a widely used Transformer model with 110 million parameters.
- Datasets: Four datasets from the common GLUE Benchmark (General Language Understanding Evaluation), covering both text classification and regression tasks.
- Hardware: Experiments were conducted on two AWS instances.
- Parameters: A slot number of 32 was used for HE operations.
- Network Settings: The evaluation considered both local area network (LAN) and wide area network (WAN) settings, simulating different bandwidths and round-trip latencies to assess real-world performance.
- Comparison Baseline: BOLT's performance was benchmarked against "Iron," a state-of-the-art secure inference system for Transformers.
Key Evaluation Metrics:
- Model Accuracy: To ensure that optimizations did not degrade prediction quality.
- Communication Overhead: Measured in megabytes (MB), reflecting the total data exchanged between parties.
- End-to-End Runtime: Measured in seconds, indicating the total time taken for a single secure inference.
Results:
- Accuracy: BOLT, both with and without the oblivious word elimination optimization, demonstrated accuracy that matched the plaintext BERT Base model. Furthermore, its accuracy was comparable to Iron, confirming that BOLT's significant efficiency gains do not compromise model utility.
- Communication Cost:
- Component-wise: BOLT's linear and non-linear layers individually showed substantial communication savings compared to Iron. The oblivious word elimination technique further reduced communication by approximately 2x.
- End-to-End: Overall, BOLT reduced the total communication cost by nearly 11x compared to Iron.
- End-to-End Runtime:
- LAN Setting: BOLT achieved a 5.2x speedup compared to Iron.
- WAN Setting: BOLT demonstrated an even more impressive 9.5x speedup, making it significantly faster in scenarios with higher latency and lower bandwidth.
These results unequivocally prove that BOLT overcomes the major performance bottlenecks of prior secure Transformer inference systems, making privacy-preserving inference a much more practical reality.
Defensive Implications
▶ Watch: Optimizing linear layers: eliminating rotations in homomorphic encryption (8:00)
The advancements presented by BOLT have profound implications for defenders and organizations deploying or utilizing Transformer-based models. The ability to perform privacy-preserving inference (PPI) efficiently and accurately addresses critical data privacy and intellectual property concerns.
For organizations offering Machine Learning as a Service (MLaaS), BOLT provides a concrete pathway to enhance trust and compliance. By allowing clients to submit encrypted inputs and receive encrypted predictions (which only they can decrypt), MLaaS providers can assure users that their sensitive data, such as personal health information, financial details, or confidential business documents, remains private throughout the inference process. This mitigates risks associated with data breaches, regulatory non-compliance (e.g., GDPR, CCPA), and reputational damage. It also protects the server's proprietary model weights from being exposed to the client.
Defenders should consider integrating or advocating for systems like BOLT in their security architectures, especially in sectors dealing with highly sensitive data. The significant reduction in communication overhead (11x) and runtime (up to 9.5x) means that the performance penalties traditionally associated with cryptographic privacy are now much more manageable. This shifts PPI from a niche, research-only concept to a viable enterprise solution.
Furthermore, the oblivious word elimination technique offers a valuable lesson: security and privacy can often be enhanced through intelligent co-design with machine learning principles. Defenders should look for opportunities where ML-driven optimizations can reduce the computational burden on cryptographic primitives, making secure systems more practical without compromising accuracy. The negligible impact on accuracy demonstrated by BOLT's word elimination highlights that such optimizations are not merely theoretical but yield tangible benefits.
In essence, BOLT provides a blueprint for a more secure MLaaS ecosystem. Defenders should understand that while the initial setup and cryptographic overhead remain, the continuous improvements in systems like BOLT are rapidly making secure, confidential AI inference a standard, rather than an exception.
Key Takeaways
- BOLT is a privacy-preserving system for Transformer inference that significantly enhances efficiency and accuracy compared to prior work.
- It employs a hybrid cryptographic protocol, leveraging Homomorphic Encryption (HE) for linear layers and Secure Multi-Party Computation (2PC) for non-linear functions.
- Novel optimizations include compact ciphertext-plaintext matrix multiplication packing in HE and the use of the Baby-Step Giant-Step (BSGS) algorithm to minimize rotations, drastically reducing communication overhead.
- For non-linear layers, BOLT uses crypto-friendly polynomial approximations and Mosin's polynomial pre-processing to reduce the number of expensive multiplications between shares in 2PC.
- Oblivious word elimination is a key machine learning optimization that identifies and prunes unimportant tokens, cutting computational costs by half with negligible impact on model accuracy.
- BOLT achieves an impressive 11x reduction in communication cost and up to a 9.5x speedup in runtime for BERT Base inference, making practical secure Transformer inference a reality.
About the Speaker(s)
The primary presenter for this talk was Qi Pang, who introduced himself as a second-year PhD student at Kon. He emphasized that this work was a collaborative effort with his "awesome collaborators" Jinhao Zhu, Helen Möllering, Wenting Zheng, and Thomas Schneider. While detailed biographies for all collaborators were not provided in the transcript, Qi Pang's role as a PhD student highlights the academic rigor and research-driven nature of the BOLT project, developed within an institutional setting focused on advancing cryptographic and machine learning security.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This work on BOLT is a critical advancement for privacy-preserving MLaaS, addressing a fundamental trust problem with concrete, measurable improvements. The co-design of cryptographic primitives and ML optimizations delivers a truly practical system for secure Transformer inference, pushing the state of the art significantly. This isn't just theory; it's a blueprint for deployable confidential AI.
Heather Calloway (CISO) — STRONG ACCEPT
BOLT presents a significant advancement in privacy-preserving AI inference, making secure deployment of Transformer models practically viable through impressive performance gains. This work directly addresses critical business risks and regulatory compliance challenges for organizations leveraging Machine Learning as a Service with sensitive data. It offers a clear path for executive action to enhance trust and accountability in AI operations.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024