Hyena: Balancing Packing, Reuse, and Rotations for Encrypted Inference

Sarabjeet Singh, Shreyas Singh, Sumanth Gudaparthi, Xiong Fan, Rajeev Balasubramonian

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 6

Overview

This talk introduces Hyena, a significant advancement towards achieving practical privacy-preserving inference using Homomorphic Encryption (HE). Homomorphic Encryption is a powerful cryptographic tool that enables computations directly on encrypted data, unlocking the potential for sensitive applications such as medical imaging, genomics, and secure cloud-based compute outsourcing. However, its widespread adoption has been severely limited by substantial performance overheads, memory intensity, and inherent implementation complexity.

Watch on YouTube

Visual summary for Hyena: Balancing Packing, Reuse, and Rotations for Encrypted Inference by Sarabjeet Singh, Shreyas Singh, Sumanth Gudaparthi, Xiong Fan, Rajeev Balasubramonian
Visual summary for Hyena: Balancing Packing, Reuse, and Rotations for Encrypted Inference by Sarabjeet Singh, Shreyas Singh, Sumanth Gudaparthi, Xiong Fan, Rajeev Balasubramonian

Key moments

  1. 0:00 Introduction to Homomorphic Encryption and its challenges
  2. 1:00 Expensive rotations identified as major performance bottleneck
  3. 3:45 Limitations of prior packing strategies: Channel vs. Lola
  4. 4:20 Hina's objectives: maximize utilization, minimize rotations, maximize reuse
  5. 5:15 Hina's novel packing strategy for improved slot utilization
  6. 7:30 Hina's data flow optimization: delaying rotations for reduced calls
  7. 8:40 Summary of Hina's three core optimization strategies
  8. 9:50 H-packim tool release for novel HE packing exploration

Hyena: Balancing Packing, Reuse, and Rotations for Encrypted Inference

Speakers: Sarabjeet Singh; Shreyas Singh; Sumanth Gudaparthi; Xiong Fan; Rajeev Balasubramonian

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=3V0sQzVVE_c

Overview

This talk introduces Hyena, a significant advancement towards achieving practical privacy-preserving inference using Homomorphic Encryption (HE). Homomorphic Encryption is a powerful cryptographic tool that enables computations directly on encrypted data, unlocking the potential for sensitive applications such as medical imaging, genomics, and secure cloud-based compute outsourcing. However, its widespread adoption has been severely limited by substantial performance overheads, memory intensity, and inherent implementation complexity.

The core challenge addressed by Hyena lies in the inefficient handling of data rotations within HE-based Convolutional Neural Networks (CNNs). These rotations, necessary for properly accumulating results across packed encrypted data slots, are identified as a major performance bottleneck due to their high computational cost, particularly the associated key switching operations. Sarabjeet Singh and his co-authors present Hyena as a novel packing and data flow strategy designed to meticulously balance slot utilization, minimize costly rotations—especially complex permutation-type rotations—and maximize on-chip data reuse.

The research behind Hyena offers a compelling solution to make encrypted inference a tangible reality. By significantly reducing rotation counts and memory activity, Hyena achieves considerable improvements in inference latency and energy efficiency compared to existing baselines. This work is crucial for accelerating the deployment of secure artificial intelligence and machine learning models in environments where data privacy is paramount, thereby bridging the gap between theoretical cryptographic guarantees and practical application performance.

Background

▶ Watch: Introduction to Homomorphic Encryption and its challenges (0:00)

Homomorphic Encryption operates on large-degree polynomials, often megabytes in size, which inherently leads to significant memory accesses and poor performance. Beyond the sheer data volume, HE involves complex mathematical operations, including modular arithmetic over large finite fields, domain transformations, and morphism, making it notoriously difficult to implement efficiently. Consequently, privacy-preserving inference has historically suffered from orders-of-magnitude slowdowns compared to unencrypted computations.

To mitigate the cost associated with these large polynomials, packing is a widely adopted technique. Instead of encrypting a single value, a vector of multiple values is packed into the various "slots" of a single polynomial, effectively amortizing the cost of a single ciphertext operation across many data points. While packing improves throughput, it introduces a new challenge: to fully exploit these packed slots and correctly accumulate results, especially in operations like convolutions, the data within these slots often needs to be rearranged or rotated.

Rotations are not trivial operations. They require additional metadata known as keys to perform key switching, which itself is a memory-intensive process due to the large size of these keys (also in megabytes). Furthermore, a non-cyclic permutation of slots—where elements move arbitrarily rather than just shifting around a ring—cannot be performed directly. Instead, it necessitates a series of many individual cyclic rotations, making such operations even more prohibitively expensive. In the context of CNNs, which heavily rely on multiply-accumulate steps, these rotations become a critical performance bottleneck. A typical convolution involves multiplying input weights with input feature maps and then summing the results across input channels. When performed with encrypted ciphertexts, the multiply-accumulate step gains an additional, costly rotation step to correctly align and accumulate partial sums.

Prior research has explored various packing strategies, each attempting to optimize different aspects, often at the expense of others. One intuitive approach, referred to as Channel packing, involves packing all inputs required to complete a single output pixel together. While straightforward, this strategy requires a large number of rotations per multiplication, specifically N-1 rotations, where N is the number of elements. This high rotation count leads to substantial overhead. Conversely, another strategy, exemplified by Lola packing, aims to achieve almost zero rotations by carefully arranging data. However, Lola packing often results in inefficient utilization of polynomial slots, leading to "ineffectual computations" represented by unused or "white boxes" in the data layout. This underutilization means that computational resources are spent on operations that do not contribute to the final result, diminishing overall efficiency despite fewer rotations. The fundamental challenge that Hyena seeks to address is precisely this delicate trade-off: how to maximize slot utilization to keep the working set and memory activity low, while simultaneously minimizing the number and complexity of rotations, particularly the expensive permute-type rotations, to achieve practical performance.

Key Findings

▶ Watch: Limitations of prior packing strategies: Channel vs. Lola (3:45)

Hyena's primary contribution is the development of a novel packing and data flow strategy specifically tailored for private CNN inference using Homomorphic Encryption. The core objectives guiding its design are threefold: to maximize slot utilization, thereby reducing the total number of ciphertexts and memory activity; to minimize rotations, especially the computationally intensive permute-type rotations; and to maximize on-chip data reuse, further reducing off-chip memory movement.

The key findings demonstrate that Hyena successfully navigates the complex trade-off between slot utilization and rotation costs that plagued prior HE packing schemes. Through its carefully tailored approach, Hyena achieves:

  • Near-zero permute-type rotations: A critical achievement that eliminates the most expensive rotation operations.
  • Significantly lower memory activity: Hyena exhibits at least 10 times lower memory activity compared to baseline packing strategies, which translates directly to reduced data footprint and energy consumption.
  • Substantial performance gains: Encrypted inference latencies are reduced to tens to hundreds of milliseconds, making practical applications feasible. This represents an improvement of at least 2x faster than the best prior baseline approaches.
  • Enhanced energy efficiency: The reduced memory accesses and optimized computations lead to a demonstrably more energy-efficient solution.

Furthermore, to foster broader research and development in this critical area, the Hyena team has released an open-source tool named H-PacSIM. This tool allows the community to explore and evaluate novel packing, data flow strategies, and hardware designs for private inference, democratizing access to optimized HE development. Hyena represents a significant leap forward, demonstrating that practical, privacy-preserving inference for CNNs is achievable by intelligently managing data movement and computation within the constraints of homomorphic encryption.

Technical Deep Dive

▶ Watch: Hina's novel packing strategy for improved slot utilization (5:15)

Hyena's technical innovation lies in its meticulously designed packing strategy, complemented by an optimized data flow, both working in concert to minimize rotations and maximize data reuse within the constraints of Homomorphic Encryption for CNNs.

Hyena's Packing Strategy:

The packing strategy begins by focusing on localized computation to ensure high slot utilization and minimize initial data movement.

  1. Initial Phase Packing: Hyena starts by packing a single phase of a weight matrix alongside its corresponding phase from the input feature map. The output of this multiplication contributes to the partial sum of the first output pixel. This initial packing is designed for efficient element-wise multiplication within the HE ciphertext slots.
  2. Neighboring Non-Overlapping Phases: Next, Hyena packs neighboring, non-overlapping phases of both weights and input feature maps. Crucially, instead of repeatedly packing the same weight values, Hyena packs a different weight phase derived from a different output channel. This strategic decision is vital as it allows the generation of partial output pixels across multiple output channels simultaneously. This technique directly addresses and mitigates the redundant weight packing issue observed in strategies like Lola packing, thereby significantly improving overall slot utilization.
  3. Across Input Channels: Once all input phases or output channels related to the current set of weights are exhausted, the packing strategy then progresses across the input channels of the same weight values. The products generated from these operations accumulate towards the same output feature map, ensuring that all necessary contributions are gathered efficiently.

The overarching goal of this tailored slot packing is to transform what would typically require complex permute-type rotations—which are expensive series of cyclic rotations—into simpler, more efficient cyclic shift rotations. By carefully arranging data, Hyena maximizes the effective use of polynomial slots, reducing the total number of ciphertexts required and consequently lowering the working set size and memory activity.

Hyena's Data Flow Strategy:

The packing strategy is seamlessly integrated with a sophisticated data flow that orchestrates computations to further minimize rotation calls and maximize on-chip reuse.

  1. Slot Reordering for Cyclic Rotations: After an initial set of multiplications, Hyena reorders the slots within the ciphertexts. This reordering is specifically designed to ensure that the subsequent rotations required to accumulate partial sums are exclusively cyclic shifts. This is a critical step in avoiding the costly multi-cyclic permutations.
  2. Operand Reuse through Weight Rotation: A key aspect of the data flow involves reusing already loaded operands. By rotating the weight polynomial across different output feature maps, Hyena can generate different output pixels using the same input features and weight polynomials that are already cached on-chip. This intelligent reuse significantly minimizes off-chip memory movement, which is a major contributor to energy consumption and latency in HE systems. This strategy effectively generates "new combinations" from existing data.
  3. Delayed Partial Sum Accumulation: Perhaps one of Hyena's most innovative data flow techniques is the delayed accumulation of partial sums. Instead of immediately performing rotation-assisted partial sum accumulation after each multiplication or after processing a subset of input channels, Hyena defers this operation. The rotation-assisted accumulation is performed only after all contributions across separately packed input channels have been added. This delay allows for multiple additions to occur before a rotation is needed, effectively reducing the total number of rotation calls by a factor almost equal to the input channel count, leading to substantial performance gains.
  4. Completing Output Pixels: Finally, the process completes by iterating over all output channels, generating new combinations by reusing input features with targeted permute-type rotations (which are now minimized and carefully managed) and performing the necessary accumulations for all output pixels.

In summary, Hyena's technical prowess stems from three integrated strategies: first, a carefully tailored slot packing that converts complex permutations into simpler shift rotations and maximizes slot utilization; second, a data flow that delays partial sum accumulation until all input channels have contributed, drastically reducing rotation counts; and third, a mechanism for generating new inputs using already on-chip cached inputs, which minimizes memory activity and maximizes computational efficiency. These combined innovations enable Hyena to achieve its significant performance improvements for encrypted CNN inference.

Demo / Proof of Concept

▶ Watch: Hina's data flow optimization: delaying rotations for reduced calls (7:30)

While the talk did not feature a live, interactive demonstration in the traditional sense, the practical applicability and efficacy of Hyena were rigorously evaluated through an ASIC design targeting popular deep learning models. The research team assessed Hyena's performance on benchmarks including ResNet-50, MobileNet, and GNMT models, which represent diverse convolutional neural network architectures.

The implementation utilized a hybrid strategy for privacy preservation. Specifically, Multi-Party Computation (MPC) was employed to handle the non-linear layers within the neural network, which are typically challenging and inefficient to implement directly with Homomorphic Encryption. Concurrently, the client performs the remaining, HE-friendly linear operations. This hybrid approach leverages the strengths of both cryptographic primitives to achieve a more practical and efficient overall solution.

A significant contribution to the broader research community is the release of H-PacSIM, an open-source tool publicly available on GitHub. This tool serves as a powerful platform for further exploration, allowing researchers and practitioners to design and evaluate novel packing strategies, data flow mechanisms, and even hardware architectures specifically optimized for private inference. H-PacSIM enables detailed modeling of various baseline packings and architectural implementations, facilitating comparative studies and accelerating innovation in the field of privacy-preserving machine learning. The detailed evaluation results, presented in the talk, highlighted Hyena's superior performance in terms of rotation call counts, data footprint (in gigabytes), inference latency (in milliseconds), and energy consumption across various hardware components, substantiating its claims of practical efficiency.

Defensive Implications

▶ Watch: H-packim tool release for novel HE packing exploration (9:50)

The primary defensive implication of Hyena's work lies in its significant advancement towards making privacy-preserving inference practical and widely adoptable. For organizations and individuals handling highly sensitive data, the ability to perform computations on encrypted information without decrypting it is a fundamental pillar of data protection. By tackling the critical performance bottlenecks of Homomorphic Encryption (HE), particularly the high cost of rotations in Convolutional Neural Networks (CNNs), Hyena directly enhances the feasibility of deploying secure AI models in sensitive domains.

Specifically, Hyena's contributions enable a stronger defense against data breaches and unauthorized access in scenarios where AI models are deployed in untrusted environments, such as cloud services. For instance, in medical imaging or genomics, patient data can remain encrypted throughout the diagnostic or analytical process, preventing cloud providers or third-party AI services from ever accessing the raw, unencrypted information. Similarly, in financial services or government applications, sensitive personal or classified data can be processed by machine learning models without compromising its confidentiality.

The increased performance (up to 2x faster) and reduced memory footprint (at least 10 times lower memory activity) offered by Hyena make the deployment of such secure solutions economically viable and responsive enough for real-world applications. This significantly lowers the barrier to entry for privacy-enhancing technologies, allowing defenders to integrate robust cryptographic protections into their AI pipelines without incurring prohibitive computational overheads. Furthermore, the open-source release of H-PacSIM empowers security researchers, developers, and practitioners to experiment with, optimize, and implement secure HE solutions more effectively. This fosters a collaborative environment for strengthening the security posture of AI systems, ultimately leading to more resilient and privacy-conscious data processing architectures across various industries.

Key Takeaways

  • Rotations are the Bottleneck: In Homomorphic Encryption-based Convolutional Neural Networks, data rotations required for accumulating partial sums are identified as the major performance bottleneck, primarily due to expensive key switching operations and the need for multiple cyclic rotations for non-cyclic permutations.
  • Hyena's Novel Strategy: Hyena introduces a sophisticated packing and data flow strategy designed to overcome the long-standing trade-off between slot utilization and rotation costs in HE.
  • Optimized Data Handling: Key strategies include tailored slot packing that converts permute-type rotations into more efficient cyclic shifts, delaying partial sum accumulation until all input channels have been processed (reducing rotation calls by a factor of input channel count), and maximizing on-chip reuse of operands to minimize memory activity.
  • Significant Performance Gains: Hyena achieves almost zero permute-type rotations, at least 10 times lower memory activity, and inference latencies in tens to hundreds of milliseconds, resulting in at least 2x faster and significantly more energy-efficient encrypted inference compared to prior baselines.
  • Enabling Practical Privacy: By making HE-based CNN inference practical and efficient, Hyena advances the feasibility of deploying privacy-preserving AI models in sensitive applications like medical imaging, genomics, and secure outsourced computation.
  • Open-Source Tool for Research: The research team has released H-PacSIM, an open-source tool on GitHub, to facilitate community exploration and development of novel packing, data flow strategies, and hardware designs for private inference.

About the Speaker(s)

Sarabjeet Singh presented the work on Hyena, which was co-authored by Shreyas Singh, Sumanth Gudaparthi, Xiong Fan, and Rajeev Balasubramonian. This team of researchers collaborated on this project, focusing on practical privacy-preserving inference. Their collective expertise has contributed to identifying and mitigating key performance bottlenecks in Homomorphic Encryption, pushing the boundaries of secure computation for deep learning applications.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

Hyena presents a truly clever, deeply technical solution to the long-standing performance bottleneck of rotations in Homomorphic Encryption for CNNs. By meticulously balancing packing, reuse, and delayed accumulation, it achieves near-zero permute-type rotations, making privacy-preserving inference practical with significant speed and memory efficiency gains. This isn't just theory; it's a foundational step towards secure AI.

Heather Calloway (CISO) — STRONG ACCEPT

This research delivers a critical advancement in practical privacy-preserving AI, directly addressing performance bottlenecks in Homomorphic Encryption. It offers a credible path for organizations to securely leverage AI with sensitive data, significantly impacting governance and reducing business exposure.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024