SHAFT: Secure, Handy, Accurate and Fast Transformer Inference
Andes Y. L. Kei
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Privacy & Cryptography 2 · Privacy & Cryptography 2
Overview
The proliferation of Large Language Models (LLMs) like ChatGPT has ushered in a new era of AI capabilities, yet it has simultaneously amplified concerns regarding data privacy. When users submit sensitive queries to an LLM, their private information could potentially be exposed to the model owner or third parties. The talk "SHAFT: Secure, Handy, Accurate and Fast Transformer Inference," presented by Andes Y. L. Kei, addresses this critical challenge by introducing a novel framework for private inference on Transformer models. This allows a client with a private query and a server with a private model to perform inference without either party revealing their sensitive inputs.
Key moments
- 0:00 Introduction to SHAFT: Secure Transformer Inference
- 2:49 SHAFT's novel approach and performance benefits
- 3:56 Seamless integration with Hugging Face Transformers
- 4:26 Key technical contributions: Softmax and GELU protocols
- 5:58 Detailed explanation of private Softmax protocol
- 8:20 Precise and efficient private GELU protocol
SHAFT: Secure, Handy, Accurate and Fast Transformer Inference
Speakers: Andes Y. L. Kei, Professor Sherman Chao
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=KzDlEPZSZiA
Overview
The proliferation of Large Language Models (LLMs) like ChatGPT has ushered in a new era of AI capabilities, yet it has simultaneously amplified concerns regarding data privacy. When users submit sensitive queries to an LLM, their private information could potentially be exposed to the model owner or third parties. The talk "SHAFT: Secure, Handy, Accurate and Fast Transformer Inference," presented by Andes Y. L. Kei, addresses this critical challenge by introducing a novel framework for private inference on Transformer models. This allows a client with a private query and a server with a private model to perform inference without either party revealing their sensitive inputs.
SHAFT distinguishes itself by simultaneously achieving four key properties: security, handiness, accuracy, and fast inference. "Handiness" implies that the system is designed to be deployable even by those without deep cryptographic expertise, offering a PyTorch-like API for seamless integration with popular libraries like Hugging Face. The work provides formal security proofs, ensures accuracy comparable to plaintext inference, and significantly outperforms prior state-of-the-art methods in terms of communication and computation efficiency, particularly in the critical and complex nonlinear operations inherent to Transformer architectures. This research is pivotal for enabling the secure and practical adoption of LLMs in privacy-sensitive applications.
Background
▶ Watch: Introduction to SHAFT: Secure Transformer Inference (0:00)
The Transformer architecture has become the de facto standard for modern LLMs, underpinning their remarkable success in natural language processing tasks. A typical Transformer model comprises an embedding layer, multiple Transformer blocks, and a simple neural network classifier. Each Transformer block is complex, incorporating multi-head attention mechanisms, feed-forward networks (FFN), and layer normalization. Crucially, these components involve computationally intensive and often nonlinear operations, such as Softmax within the multi-head attention layers and GELU (Gaussian Error Linear Unit) activations within the FFNs.
The inherent privacy risks associated with sending sensitive user queries to remote LLM servers have spurred extensive research into private inference. This field aims to enable joint computation where neither party's input is revealed. However, applying private inference techniques to Transformers presents unique challenges. Unlike traditional convolutional neural networks, Transformers heavily rely on these complex nonlinear functions, Softmax and GELU, which are particularly difficult to evaluate efficiently and accurately in a secure, privacy-preserving manner. Early attempts, such as MBCFormer, often resorted to rough approximations of these nonlinear functions, leading to significant accuracy degradation, sometimes exceeding a 5% error.
More recent works have made strides in improving secure linear protocols and leveraging advanced cryptographic primitives. For instance, Sigma utilized function secret sharing to reduce running time, while Bumblebee optimized homomorphic matrix multiplication to save communication costs. However, a common limitation across these prior efforts was their reliance on existing techniques for approximating the nonlinear functions or an unexplored space for more precise and efficient approximations. Furthermore, many existing solutions require specialized knowledge of secure computation libraries, hindering their practical adoption. The MBCFormer work also lacked formal security proofs, a crucial aspect for real-world deployment. SHAFT emerges from this context, seeking to bridge these gaps by introducing novel approximation techniques for Softmax and GELU, coupled with formal security guarantees and an accessible API, all while pushing the boundaries of efficiency.
Key Findings
▶ Watch: Seamless integration with Hugging Face Transformers (3:56)
The SHAFT framework delivers several significant contributions and findings that advance the state of private Transformer inference:
- Novel and Efficient Private Softmax Protocol: SHAFT introduces the first constant-run private Softmax protocol specifically tailored for Transformers. This protocol leverages insights from ordinary differential equation (ODE) approximations, coupled with a novel approach to selecting clipping ranges that minimize errors for unbounded inputs, a common challenge in Transformer contexts. The careful selection of
T=16iterations and clipping boundsa=-4andb=12ensures high accuracy while maintaining efficiency.
- Precise and Efficient Private GELU Protocol: The work proposes a highly precise and efficient private GELU protocol. Unlike prior methods that often required multiple communication rounds or introduced significant approximation errors, SHAFT utilizes Fourier series approximations on a specially designed
delta_xfunction. This innovation allows for the secure evaluation of GELU in just one communication round, a 50% reduction compared to state-of-the-art methods likeBoat(SMB 2024), which required two rounds for polynomial approximations.
- Superior Efficiency Over State-of-the-Art: SHAFT consistently outperforms recent leading works in private Transformer inference.
- Compared to Sigma: SHAFT achieves reduced communication overhead while maintaining comparable running times.
- Compared to Bumblebee: SHAFT demonstrates smaller running times in both LAN (local area network) and WAN (wide area network) settings, making it more practical for diverse deployment scenarios.
- Comparable Accuracy to Plaintext Inference: A critical finding is that SHAFT achieves accuracy levels comparable to non-private, plaintext inference. This overcomes a major limitation of earlier private inference schemes that sacrificed accuracy for privacy, making SHAFT suitable for real-world applications where model performance is paramount.
- Formal Security Proofs: Unlike some prior works (e.g., MBCFormer), SHAFT provides rigorous formal security proofs, ensuring the cryptographic soundness of its protocols. This is crucial for establishing trust and enabling deployment in sensitive environments.
- User-Friendly Integration: SHAFT offers a PyTorch-like API that seamlessly integrates with popular Hugging Face transformer libraries. This "handiness" significantly lowers the barrier to entry for developers and researchers, allowing them to import Transformer models for privacy-preserving inference with as few as six lines of code (excluding comments and imports). This addresses the need for solutions deployable without deep cryptographic knowledge.
- Outsourced Two-Party Setting with Precomputations: The framework is designed for a two-party outsourced setting where precomputations can be performed by commodity servers or through two-party computations between the computing servers, optimizing the online phase efficiency.
These findings collectively position SHAFT as a leading solution for practical, secure, and accurate Transformer inference, addressing key bottlenecks in privacy-preserving machine learning.
Technical Deep Dive
▶ Watch: Key technical contributions: Softmax and GELU protocols (4:26)
The core technical innovations of SHAFT revolve around the secure and efficient evaluation of the two most challenging nonlinear functions in Transformer models: Softmax and GELU. These operations are ubiquitous within Transformer blocks, with Softmax being central to multi-head attention and GELU serving as the activation function in feed-forward networks.
Secure Softmax Evaluation
The Softmax function converts a vector of arbitrary real values into a probability distribution, where all elements are between 0 and 1 and sum to 1. Mathematically, for an input vector $x$, the Softmax of the $i$-th element is given by $e^{x_i} / \sum_j e^{x_j}$. Securely evaluating Softmax is challenging due to the exponential and reciprocal operations, which are prone to overflow, especially with unbounded inputs typical in Transformer layers.
A standard approach to prevent overflow is to compute Softmax on the input vector $x$ after subtracting its maximum element, $x - \max(x)$. This manipulation does not affect correctness but makes all exponents non-positive, preventing overflow. However, securely computing the maximum element itself requires a logarithmic number of communication rounds with respect to the input length, which can be inefficient for long input sequences.
SHAFT's starting point for private Softmax is an Ordinary Differential Equation (ODE) approximation method from ACSAC 2023. This method iteratively approximates Softmax and has a constant-run complexity in terms of input length, requiring 2T rounds where T is the number of iterations. For unbounded Transformer inputs, a large T would typically be needed, leading to substantial computational cost.
The key insight introduced by SHAFT is the importance of clipping the inputs to the Softmax function. The talk highlights that in natural language processing, words in a sentence are usually related to only a few other words. For example, in "The animal didn't cross the street because it was too tired," "it" clearly refers to "animal." This implies that while some Softmax inputs (corresponding to "animal" in this example) will be large and dominate the sum in the denominator, most inputs will be relatively small. Aggressively clipping the large inputs can cause significant errors.
SHAFT carefully selects the clipping range for Softmax inputs based on this NLP domain understanding. They choose a large positive upper bound to minimize errors from dominant large inputs and a slightly negative lower bound to include most small inputs. Specifically, their work sets T = 16 iterations, with an upper clipping bound b = 12 and a lower clipping bound a = -4. This tailored clipping strategy, combined with the ODE approximation, allows SHAFT to achieve high accuracy for Softmax while maintaining its constant-run complexity and avoiding the overhead of secure maximum computations.
Secure GELU Evaluation
The GELU (Gaussian Error Linear Unit) activation function is defined as $x \cdot \Phi(x)$, where $\Phi(x)$ is the cumulative distribution function of the standard normal distribution. It is often approximated as $x \cdot \frac{1}{2} [1 + \text{erf}(x/\sqrt{2})]$, where $\text{erf}$ is the Gaussian error function. GELU is a smooth, non-monotonic function that is crucial for the performance of many modern Transformers.
Existing approaches for securely computing GELU typically involve approximating it with polynomials for values of $x$ near zero and using the Rectified Linear Unit (ReLU) or similar functions for larger absolute values of $x$, as GELU approximates ReLU when $|x|$ is large. For example, Boat (SMB 2024), a state-of-the-art method, uses a degree-four polynomial for approximations. However, evaluating such polynomial approximations securely often requires two communication rounds, which can be substantial given that Transformers can have hundreds of thousands of GELU activations.
SHAFT's starting point for GELU is the observation that Fourier series can provide precise approximations for functions with a sinusoidal shape over a bounded input range, and critically, securely evaluating a Fourier series takes only one communication round. The challenge is that GELU itself is not sinusoidal. A simple solution from concurrent work involves approximating the Gaussian error function $\text{erf}(x)$, but this still requires an additional run to securely multiply the result by $x$ (as per the definition of GELU), potentially increasing approximation errors.
To overcome this, SHAFT proposes designing a suitable function for Fourier series approximation. They modify the non-sign function $G(-x)$ from Sigma to define a new function delta_x. This delta_x function is a signal function for $x \neq 0$ and is specifically designed to be ideal for accurate Fourier series approximations. For large absolute values of $x$, delta_x is set to zero.
The key innovation is that GELU can be computed from this delta_x function using a specific characterization that requires no extra communication rounds. The talk states that relu_x (ReLU of $x$) can be computed non-interactively, and then GELU can be derived from delta_x and relu_x without additional interactive steps. This elegant solution allows SHAFT to achieve a one-round secure GELU computation while maintaining high precision, a 50% reduction in rounds compared to prior methods.
Other Contributions and Setting
Beyond Softmax and GELU, SHAFT also contributes:
- A protocol for private embedding program inputs, ensuring the initial layer of the Transformer is also privacy-preserving.
- Extensions of their GELU characterization to other activation functions, demonstrating the generality of their approach.
- Optimizations for smaller bit widths, which can further enhance efficiency.
The entire framework operates in a two-party outsourced setting with precomputations. This means that computationally intensive cryptographic setup phases can be performed offline by commodity servers or through an initial two-party computation between the computing servers. This design maximizes the efficiency of the online inference phase, making the system more practical. The protocols are made publicly available in a GitHub repository, fostering further research and adoption.
Demo / Proof of Concept
▶ Watch: Detailed explanation of private Softmax protocol (5:58)
While the talk did not feature a live, interactive demonstration of the SHAFT framework, it strongly emphasized its "handiness" and ease of integration, which serves as a crucial aspect of its practical proof of concept. The speakers highlighted that SHAFT provides a PyTorch-like API, designed to smoothly integrate with popular Hugging Face transformer libraries.
A key metric presented to illustrate this ease of use was a figure demonstrating how Transformer models could be imported for privacy-preserving inference in "just six lines of code," excluding comments and import statements. This demonstrates a significant reduction in the complexity typically associated with implementing secure multi-party computation protocols. The implication is that a developer familiar with standard deep learning frameworks like PyTorch can quickly adapt existing Transformer models to run securely with SHAFT, without needing deep cryptographic knowledge.
Furthermore, the talk explicitly mentioned that the code for SHAFT, including its secure and efficient protocols for Softmax and GELU, is "probably available in this GitHub repo repository." This public availability serves as a tangible proof of concept, allowing researchers and practitioners to inspect, verify, and deploy the framework themselves. The design philosophy of abstracting cryptographic complexities behind a familiar API is a core part of SHAFT's "handiness," proving that practical secure inference can be made accessible to a broader audience.
Defensive Implications
▶ Watch: Precise and efficient private GELU protocol (8:20)
The SHAFT framework offers significant defensive implications for organizations and individuals concerned with data privacy in the age of Large Language Models (LLMs). Its primary contribution is enabling private inference for Transformer models without compromising accuracy or performance, thereby addressing a critical privacy vulnerability.
- Enabling Secure LLM Deployment: For organizations considering deploying LLMs for sensitive applications (e.g., healthcare, finance, legal), SHAFT provides a viable pathway. It allows them to leverage the power of advanced AI models while ensuring that client queries remain confidential from the model owner, and the proprietary model weights remain hidden from the client. This mitigates risks associated with data breaches, intellectual property theft, and compliance with privacy regulations like GDPR or HIPAA.
- Bridging the Gap to Practicality: The "handiness" of SHAFT, characterized by its PyTorch-like API and integration with Hugging Face transformer libraries, significantly lowers the barrier to entry for secure AI deployment. Security teams and AI engineers, who may not have specialized cryptographic expertise, can now more easily implement privacy-preserving inference. This accelerates the adoption of secure ML practices from academic research into industry applications, moving away from complex, custom cryptographic implementations.
- Formal Security Guarantees: The provision of formal security proofs is a crucial defensive measure. Organizations can have higher confidence in the cryptographic guarantees of SHAFT, knowing that its protocols have been rigorously analyzed to protect against specific adversary models (in this case, a two-party outsourced setting with honest-but-curious adversaries). This is a stark improvement over systems lacking such proofs, where vulnerabilities might go unnoticed.
- Performance and Accuracy Preservation: Defenders often face a trade-off between security, performance, and accuracy. SHAFT's ability to maintain accuracy comparable to plaintext inference while significantly improving efficiency (reduced communication and running time) means that adopting privacy-preserving measures does not necessitate a substantial degradation in model utility or user experience. This makes the argument for secure deployment much stronger within an organization.
- Focus on Nonlinear Operations: By specifically tackling the complex Softmax and GELU operations with novel, efficient protocols, SHAFT addresses the core bottlenecks that have historically plagued secure Transformer inference. This focused approach ensures that the most computationally expensive and privacy-sensitive parts of the model are handled optimally, strengthening the overall security posture.
- Future-Proofing: While the current work focuses on an honest-but-curious adversary, the speakers mentioned "security against malicious adversary" as future work. Organizations should monitor such developments, as robust security against malicious actors is the gold standard for many high-stakes deployments.
In essence, SHAFT empowers defenders to move beyond theoretical discussions of privacy-preserving AI to practical, deployable solutions. It provides the tools and confidence needed to integrate privacy by design into LLM-based applications, fostering trust and enabling responsible AI innovation.
Key Takeaways
- SHAFT enables secure, privacy-preserving inference for Transformer models, allowing clients and servers to collaborate without revealing their private inputs or model weights, addressing critical privacy concerns in LLMs.
- The framework achieves a unique balance of security, handiness, accuracy, and fast inference, outperforming prior state-of-the-art solutions while maintaining plaintext-comparable accuracy.
- Novel private Softmax protocol utilizes Ordinary Differential Equation (ODE) approximations with carefully selected clipping ranges (
a=-4,b=12,T=16iterations) to ensure high accuracy and constant-run complexity in a Transformer context. - Highly efficient private GELU protocol leverages Fourier series approximation on a specialized
delta_xfunction, enabling secure computation in just one communication round, a 50% reduction over previous two-round methods. - SHAFT is significantly more efficient than existing works like
Sigma(reduced communication) andBumblebee(smaller running times in LAN/WAN settings), making it more practical for real-world deployment. - The framework offers a PyTorch-like API for easy integration with Hugging Face transformer libraries, allowing privacy-preserving inference with minimal code (e.g., six lines), making it accessible to developers without deep cryptographic knowledge.
About the Speaker(s)
The talk "SHAFT: Secure, Handy, Accurate and Fast Transformer Inference" was presented by Andes Y. L. Kei. He introduced the work as a collaborative effort with his advisor, Professor Sherman Chao. The presentation highlighted their research into building practical and secure solutions for private inference in Large Language Models (LLMs). Their work focuses on addressing the privacy concerns surrounding the increasing adoption of LLMs by developing cryptographic protocols that ensure security, efficiency, and ease of use.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic systems security work from NDSS — not a security conference in the DEF CON/Black Hat sense, so the audience is researchers, not practitioners. SHAFT makes real incremental contributions to private ML inference: the ODE-based Softmax with domain-informed clipping bounds and the one-round Fourier-series GELU are genuinely clever protocol-level ideas that required mathematical depth to produce. But this is squarely incremental work in a crowded field — it's optimization over Sigma and Bumblebee, not a paradigm shift.
Heather Calloway (CISO) — WEAK
Technically credible research on private LLM inference with real cryptographic contributions, but the talk is aimed squarely at cryptography researchers, not security leaders or operators. The defensive framing is retrofitted — the governance, compliance, and deployment questions that would make this matter to a CISO are entirely absent.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025