CoreLocker: Neuron-level Usage Control

Zihan Wang, Zhongkui Ma, Xinguo Feng, Ruoxi Sun, Hu Wang, Minhui Xue

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 5

Overview

The presented talk, "CoreLocker: Neuron-level Usage Control," by Zihan Wang and collaborators, introduces a novel framework designed to protect and monetize the intellectual property inherent in Deep Neural Networks (DNNs). Given the astronomical resources—including vast datasets, immense computational power, and sophisticated architectural designs—required to develop state-of-the-art AI models like GPT-3, which demanded 355 GPU years and an estimated $4.6 million for a single training run, these models represent incredibly valuable assets. The potential returns are equally staggering, as exemplified by ChatGPT's rapid ascent to 100 million active users within two months and generating $80 million per month for OpenAI.

Watch on YouTube

Visual summary for CoreLocker: Neuron-level Usage Control by Zihan Wang, Zhongkui Ma, Xinguo Feng, Ruoxi Sun, Hu Wang, Minhui Xue
Visual summary for CoreLocker: Neuron-level Usage Control by Zihan Wang, Zhongkui Ma, Xinguo Feng, Ruoxi Sun, Hu Wang, Minhui Xue

Key moments

  1. 0:20 Protecting valuable DNNs and the risk of unauthorized use
  2. 2:00 Limitations of current DNN protection methods
  3. 3:20 CoreLocker's approach: neuron-level access key extraction
  4. 4:00 Identifying crucial weights based on magnitude for key selection
  5. 5:00 Theoretical analysis and empirical validation of weight extraction
  6. 6:00 Demonstrating granular utility control and model degradation
  7. 7:00 Broad applicability and CoreLocker's key advantages

CoreLocker: Neuron-level Usage Control

Speakers: Zihan Wang; Zhongkui Ma; Xinguo Feng; Ruoxi Sun; Hu Wang; Minhui Xue

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=BcKLv7Z9gu8

Overview

The presented talk, "CoreLocker: Neuron-level Usage Control," by Zihan Wang and collaborators, introduces a novel framework designed to protect and monetize the intellectual property inherent in Deep Neural Networks (DNNs). Given the astronomical resources—including vast datasets, immense computational power, and sophisticated architectural designs—required to develop state-of-the-art AI models like GPT-3, which demanded 355 GPU years and an estimated $4.6 million for a single training run, these models represent incredibly valuable assets. The potential returns are equally staggering, as exemplified by ChatGPT's rapid ascent to 100 million active users within two months and generating $80 million per month for OpenAI.

This immense value, however, makes DNNs prime targets for unauthorized exploitation. Whether deployed via Machine Learning as a Service (MLaaS) platforms or directly on user devices, models are vulnerable to illicit inference attacks, unfair competition, or "unseating" by unauthorized entities, leading to significant financial losses for model owners. A recent study highlighted this vulnerability, revealing that 41% of mobile applications fail to adequately secure their DNN models against on-device inference attacks. CoreLocker directly addresses this critical gap by proposing a lightweight, data-agnostic, and retraining-free method to control model utility at a granular "neuron level," ensuring that full model capabilities can only be restored by authorized users possessing a unique access key.

The significance of CoreLocker lies in its ability to provide robust intellectual property protection and flexible monetization strategies for AI models without the drawbacks of existing solutions. Unlike watermarking or encryption methods, which often fall short in preventing unauthorized usage or demand extensive retraining, CoreLocker directly operates on off-the-shelf pre-trained models. By establishing a mechanism for model owners to release tiered versions of their models, each with a controlled level of performance, and allowing authorized users to seamlessly unlock full utility, CoreLocker offers a compelling solution for safeguarding investments in AI and enabling more secure and adaptable business models for the burgeoning AI industry.

Background

▶ Watch: Protecting valuable DNNs and the risk of unauthorized use (0:20)

The rapid advancement and widespread adoption of Deep Neural Networks (DNNs) have fundamentally transformed various industries, from natural language processing to computer vision. However, the creation of these powerful models is an immensely resource-intensive endeavor. Developing a successful DNN often requires vast quantities of high-quality data, substantial computational power—measured in GPU years and millions of dollars—and highly specialized architectural designs. For example, the GPT-3 model, with its 175 billion parameters, is a testament to the scale of investment involved. This significant upfront cost and intellectual effort imbue trained DNN models with substantial economic value, establishing them as critical intellectual property (IP).

To recoup these investments and generate revenue, model owners typically deploy their DNNs through several channels. Common strategies include offering Machine Learning as a Service (MLaaS), where users access models via APIs, or deploying models directly on-device for applications requiring low latency or offline functionality. Furthermore, model owners may offer different versions of a model at varying price points, providing users with flexibility and catering to diverse needs.

However, these deployment strategies inherently expose the models to risks. When a model is not under the direct, continuous control of its owner, unauthorized entities can exploit it. This exploitation can manifest as unfair competition, where competitors leverage a proprietary model without authorization, or unseating, a process akin to reverse engineering or cloning a model to bypass licensing. The problem is particularly acute in on-device deployments, where models reside in potentially hostile environments. The statistic that 41% of mobile apps fail to secure their DNN models against inference attacks underscores the prevalence and severity of this threat, leading to significant financial losses for model owners.

Existing approaches to mitigate these risks have notable limitations. Some methods involve embedding watermarks or signatures into DNN models to verify ownership. While these can prove provenance, they often fail to prevent unauthorized usage once the model's parameters are exposed. The incentive for unauthorized parties to simply use the model remains high. Other techniques employ parameter encryption or obfuscation to deter unauthorized access. However, these methods come with significant practical drawbacks:

  • They typically require additional training and direct access to the original training data, making them unsuitable for already pre-trained models that are common in modern AI development.
  • They can be extremely time-consuming, as training a separate model for each key or each desired model version is often necessary.
  • Their effectiveness is questionable, as they can sometimes be detected and removed through distribution value detection or lack sufficient theoretical support to guarantee security.

CoreLocker aims to address these limitations by providing a solution that is training data agnostic and pre-training free, meaning it can operate directly on existing, off-the-shelf DNN models. The core research question CoreLocker seeks to answer is: how can a model's performance be reliably degraded to a lower utility level, while simultaneously ensuring that its full capabilities can be efficiently and fully restored by an authorized user? By focusing on a novel, neuron-level control mechanism, CoreLocker endeavors to fill the critical gap in robust and practical IP protection for valuable AI assets.

Key Findings

▶ Watch: CoreLocker's approach: neuron-level access key extraction (3:20)

CoreLocker's foundational insight is that the performance and overall functionality of a Deep Neural Network (DNN) are not uniformly distributed across all its parameters but rather disproportionately reliant on a small, crucial subset of weights. This phenomenon, termed "impact concentration," forms the bedrock of CoreLocker's approach to neuron-level usage control.

The central mechanism of CoreLocker involves strategically extracting a small subset of these significant weights from the neural network. This extracted subset then serves as an access key to unlock the model's complete capabilities. The immediate consequence of this extraction is a controlled degradation of the network's performance, effectively incapacitating it for unauthorized users.

Key findings supporting CoreLocker's efficacy and practicality include:

  1. Granular Utility Control: The extraction procedure is highly customizable. Model owners can precisely tailor the quantity and nature of extracted weights to achieve different levels of utility degradation. This allows for the creation of multiple tiered model versions (e.g., a free, low-utility version; a mid-tier version; and a premium, full-utility version) from a single base model. Authorized users, possessing the correct access key (the extracted weights), can efficiently restore the model to its full, intended performance.
  1. Theoretical Foundation for Weight Significance: The researchers provide a robust theoretical framework that systematically quantifies how alterations in individual weights, specifically those introduced by extraction, propagate through each layer of the network and ultimately manifest as disparities in the output layer. This is achieved by bounding the difference between the original and altered weight matrices layer-by-layer. This theoretical backing is crucial, establishing a direct relationship between the manipulation of weight matrices and the observable degradation in network output. Empirical results consistently corroborate this, showing a clear trend where the disparity between the full and degraded networks increases rapidly as the extraction ratio (the percentage of weights removed) increases.
  1. Empirical Validation of Smooth Degradation: Extensive evaluation across various experimental settings demonstrated that model accuracy smoothly and consistently decreases as the extraction ratio increases. A particularly striking finding is that CoreLocker can degrade a model's performance to a random guess level with the extraction of a mere 2% of its significant weights. This highlights the high efficiency and potency of the method in rendering a model effectively unusable for unauthorized parties with minimal intervention.
  1. Broad Applicability Across Architectures: CoreLocker was rigorously tested on a diverse range of well-known network architectures, including Convolutional Neural Networks (CNNs) like ResNet and DenseNet, Recurrent Neural Networks (RNNs), and Transformers. The consistent results across these varied architectures confirm that CoreLocker's method is broadly applicable, leveraging the fundamental property of impact concentration inherent across different neural network designs.
  1. Lightweight, Data-Agnostic, and Retraining-Free: A critical finding is that CoreLocker achieves its goals without requiring any additional training data or expensive retraining processes. It operates directly on pre-trained, off-the-shelf models, making it highly practical for existing AI deployments. Furthermore, the method is inherently lightweight, as it only involves the extraction and re-insertion of a small subset of weights.

In summary, CoreLocker establishes a crucial research problem of AI model usage control and provides an elegant, theoretically sound, and empirically validated solution. By enabling neuron-level locking of a model's utility, it allows owners to create low-utility versions that can be fully restored by authorized users with a compact access key, while standing out for its efficiency, versatility, and strong formal foundation.

Technical Deep Dive

▶ Watch: Identifying crucial weights based on magnitude for key selection (4:00)

The technical ingenuity of CoreLocker stems from its exploitation of a fundamental characteristic of Deep Neural Networks: the impact concentration of weights. The core intuition is that not all weights contribute equally to a network's functionality; instead, a relatively small subset of critical weights disproportionately dictates the model's performance. CoreLocker leverages this property to achieve granular utility control.

The key selection methodology is central to CoreLocker's mechanism. Since the objective is to operate directly on pre-trained models without requiring access to training data or retraining, the focus is on the inherent structural properties of the network: its model weights. The approach is magnitude-based key extraction. The hypothesis is that weights with higher magnitudes are generally more significant to the network's overall function.

To illustrate this, the speakers presented a visualization from the first convolutional layer of a VGGNet model. When all 64 filters in this layer were sorted by their L1 norm (a common measure of magnitude), a clear pattern emerged: only a small subset of these filters exhibited particularly high significance. Visualizations of the feature maps generated by these top-ranking filters confirmed that they captured more prominent input features compared to filters with lower L1 norms. This empirical observation strongly supports the idea that a network's performance is indeed largely reliant on a crucial, identifiable subset of its weights.

Based on this insight, the CoreLocker procedure for degrading a network's utility involves:

  1. Identifying Significant Weights: Using a magnitude-based criterion (e.g., L1 norm), weights are ranked by their importance within each layer or across the entire network.
  2. Extracting the Access Key: A chosen percentage of the most significant weights are "extracted." In practice, this means setting their values to zero or a near-zero value, effectively "disabling" their contribution to the network's computations. This extracted subset of original weight values constitutes the access key.
  3. Deploying the Degraded Model: The model with the disabled weights is the low-utility version.
  4. Restoring Full Utility: For authorized users, the access key (the original values of the extracted weights) can be re-inserted into the degraded model, immediately restoring its full functionality.

Crucially, CoreLocker provides a strong theoretical foundation for this magnitude-based extraction. The researchers established theoretical bounds that systematically quantify how these weight extraction alterations propagate through the network. Specifically, they developed a method to bound the difference between the original weight matrices (W*) and the altered weight matrices (W) layer by layer. This mathematical analysis allows for a precise understanding of how changes at the neuron level manifest in the output layer. The established bounds provide a rigorous theoretical justification for neuron-level usage control and ground the effectiveness of their approach.

The theoretical analysis revealed a direct relationship between the manipulation of weight matrices and the resulting neural network output disparity. As the percentage of extracted weights (the "extraction ratio") increases, the disparity between the full-utility model's output and the degraded model's output increases rapidly. This theoretical prediction was strongly corroborated by empirical results. Across all experimental settings, the model accuracy was observed to decrease smoothly and consistently with an increase in the extraction ratio. For instance, the talk highlighted that with a mere extraction of approximately 2% of the significant weights, the model's performance could be reduced to a random guess level, effectively rendering it useless without the key.

This precise control allows model owners to create a detailed mapping between the extraction ratio and the desired model utility. The speakers presented tables illustrating this mapping for various models, including ResNet and DenseNet, confirming the consistent relationship. This means that a single, fully trained model can serve as the foundation for multiple utility-tiered versions, each corresponding to a specific extraction ratio. This eliminates the need for multiple training runs or complex version management.

The method's universal applicability was demonstrated by testing it across diverse network architectures, including various Convolutional Networks (CNNs), Recurrent Neural Networks (RNNs), and Transformers. The consistent effectiveness across these different paradigms underscores that the "impact concentration" property, which CoreLocker exploits, is a fundamental characteristic of neural networks, making the approach broadly applicable in the current AI landscape.

In essence, CoreLocker technically achieves its goal by identifying and manipulating the most critical structural components (weights) of a DNN, backed by both rigorous mathematical proof and extensive empirical validation, thereby offering a novel and robust mechanism for usage control.

Demo / Proof of Concept

▶ Watch: Demonstrating granular utility control and model degradation (6:00)

While a live software demonstration or a step-by-step walkthrough of a tool was not explicitly detailed in the talk, the speakers presented compelling empirical evidence, visualizations, and quantitative results that collectively serve as a robust proof of concept for CoreLocker's efficacy and practicality. These demonstrations provided concrete validation of the theoretical underpinnings and the practical implications of neuron-level usage control.

The primary forms of proof of concept presented included:

  1. Visualization of Weight Significance: A key visual demonstration involved presenting a figure illustrating the filters from the first convolutional layer of a VGGNet model. These filters were sorted by their L1 norm, a proxy for their significance. The visualization clearly highlighted that only a small subset of these 64 filters exhibited high L1 norms and were visually distinct, indicating their disproportionate importance. Further, the presentation included visualizations of the feature maps generated by the top and bottom six filters, empirically showing that the more significant filters captured richer and more critical input features. This visual evidence strongly supported the core intuition that a network's performance is concentrated in a crucial subset of its weights.
  1. Quantitative Degradation Curves: The speakers provided graphical representations demonstrating the relationship between the extraction ratio (the percentage of significant weights removed) and the resulting model accuracy. These figures consistently showed a smooth and predictable decrease in model accuracy as the extraction ratio increased. This empirical trend validated the theoretical prediction that altering a small number of critical weights would significantly impact the network's performance. A particularly impactful data point was the finding that a mere 2% extraction of significant weights could degrade a model's performance to a random guess level, underscoring the efficiency of CoreLocker's control mechanism.
  1. Utility Mapping Tables: To illustrate the practical granular control offered by CoreLocker, the presentation included detailed tables. These tables mapped specific extraction ratios to corresponding utility ranges for well-known models such as ResNet and DenseNet. This demonstrated how a model owner could precisely choose an extraction ratio to achieve a desired level of performance degradation for a low-utility version, while retaining the ability to restore full utility for authorized users. This mapping is critical for implementing tiered access or trial versions of AI models.
  1. Broad Applicability Across Architectures: The robustness of CoreLocker was demonstrated by confirming its effectiveness across a diverse set of neural network architectures. The talk stated that the method was successfully applied to Convolutional Networks (CNNs), Recurrent Neural Networks (RNNs), and Transformers. This wide applicability serves as a strong proof of concept that the underlying principle of "impact concentration" is a general phenomenon in DNNs, making CoreLocker a versatile solution for various AI applications.

These empirical results and visualizations collectively served as a comprehensive proof of concept, illustrating how CoreLocker can effectively degrade model utility through neuron-level access key extraction, provide granular control, and efficiently restore full capability for authorized users, all without the need for retraining or access to original training data.

Defensive Implications

▶ Watch: Broad applicability and CoreLocker's key advantages (7:00)

CoreLocker presents a paradigm shift in how Deep Neural Networks (DNNs) can be protected and monetized, offering several critical defensive implications for model owners and the broader AI ecosystem.

For Model Owners and Developers:

  1. Enhanced Intellectual Property (IP) Protection: CoreLocker offers a robust mechanism to protect the significant investments made in training valuable AI models. By degrading the model's utility for unauthorized users, it directly combats issues like unfair competition and "unseating" (reverse engineering or cloning) that lead to financial losses. This moves beyond mere ownership verification (like watermarking) to active usage control.
  1. Flexible Monetization Strategies: Model owners can leverage CoreLocker to implement sophisticated, tiered monetization models. They can release multiple versions of a single model, each offering a different level of utility (e.g., a free, low-performance trial version; a mid-tier version; and a premium, full-performance version). This allows for greater market penetration and flexible pricing strategies without the prohibitive cost and complexity of training multiple distinct models. The ability to establish a direct mapping between an extraction ratio and model utility simplifies this process immensely.
  1. Improved On-Device Model Security: For models deployed directly on user devices, CoreLocker provides a crucial layer of defense against unauthorized inference attacks. By deploying a low-utility version, even if the model's parameters are extracted from the device, their inherent reduced performance renders them less valuable to attackers, mitigating the risk highlighted by the 41% vulnerability rate in mobile apps. Full utility is only accessible with the secure access key, which can be managed by the owner.
  1. Reduced Operational Overhead: A significant advantage is that CoreLocker is retraining-free and data-agnostic. This means model owners do not need to retrain their models or access original training data to implement usage control. They can take any off-the-shelf pre-trained model and apply CoreLocker, drastically reducing the operational costs and time associated with deploying secure AI models. This is particularly beneficial for large foundation models where retraining is economically unfeasible.
  1. Granular Licensing and Access Control: CoreLocker enables fine-grained control over who can access what level of model performance. This facilitates more sophisticated licensing agreements, where access to full capabilities can be tied directly to a validated license key, which in this case is the subset of significant weights.

Considerations and Future Defensive Challenges:

While CoreLocker offers a powerful solution, defenders should also consider:

  • Security of the Access Key: The effectiveness of CoreLocker hinges on the security of the "neuron-level access key" (the extracted significant weights). Defenders must implement robust mechanisms for distributing, storing, and authenticating these keys, ensuring they don't fall into unauthorized hands. This might involve secure enclaves, hardware security modules, or robust cryptographic protocols.
  • Performance Overhead in Real-Time Systems: Although the method is described as "lightweight," any additional processing (even re-inserting weights) could introduce minor latency. Defenders need to evaluate this impact in latency-sensitive applications.
  • Robustness against Advanced Attacks: While CoreLocker addresses limitations of prior methods like distribution value detection, continuous research into new attack vectors (e.g., advanced statistical reconstruction of missing weights, or "neuron-level key guessing" attacks) will be necessary to ensure long-term resilience.
  • Integration with Existing ML Platforms: For widespread adoption, CoreLocker's mechanisms would need seamless integration into existing MLaaS platforms and on-device deployment pipelines, requiring standardized APIs or toolkits.

In conclusion, CoreLocker provides a proactive and practical defensive strategy for safeguarding AI models, enabling model owners to confidently deploy and monetize their valuable intellectual property in an increasingly competitive and threat-laden landscape.

Key Takeaways

  • AI Models as Valuable IP: Deep Neural Networks represent significant intellectual property due to vast investment in training, necessitating robust protection against unauthorized usage and financial loss.
  • Neuron-Level Usage Control: CoreLocker introduces a novel approach to control AI model utility by operating at the neuron level, enabling granular degradation and full restoration of model performance.
  • Significant Weights as Access Keys: The method identifies and extracts a small subset of "significant weights" as an access key. Removing these weights degrades performance, while re-inserting them restores full capability for authorized users.
  • Lightweight, Data-Agnostic, Retraining-Free: CoreLocker is highly practical, operating directly on off-the-shelf pre-trained models without requiring additional training data or expensive retraining, making it efficient and universally applicable across architectures like CNNs, RNNs, and Transformers.
  • Granular Utility & High Efficiency: The approach provides precise control over utility, capable of degrading models to a random guess level with the extraction of merely 2% of significant weights, enabling tiered access and flexible monetization.
  • Strong Theoretical and Empirical Backing: CoreLocker is supported by theoretical bounds quantifying weight extraction propagation and extensive empirical validation across diverse models, confirming its effectiveness and reliability.

About the Speaker(s)

The talk "CoreLocker: Neuron-level Usage Control" was presented by Zihan Wang, who introduced himself as a first-year PhD student at the University of Queensland. He acknowledged his supervisors, Professor Guo and Dr. Ma, for their guidance. The work itself was a collaborative effort, a "join work" involving his institution, the University of Queensland, along with Sarah State 61 (likely an affiliated research group or company) and the University of Adelaide, indicating a broad research collaboration across multiple academic and potentially industrial entities. The other listed authors, Zhongkui Ma, Xinguo Feng, Ruoxi Sun, Hu Wang, and Minhui Xue, are presumed to be key contributors to this joint research effort.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

CoreLocker delivers a genuinely novel, neuron-level method for controlling DNN model utility, allowing owners to create tiered access and protect IP without retraining. By exploiting "impact concentration" of weights and providing strong theoretical and empirical validation, it offers a pragmatic and powerful solution for AI model monetization and security. This is real research that directly solves a critical industry problem.

Heather Calloway (CISO) — MUST SEE

This research presents a critical solution for protecting the substantial intellectual property embedded in AI models. CoreLocker's neuron-level usage control offers a scalable, retraining-free method to secure and monetize AI assets, directly addressing significant business risks and enabling flexible deployment strategies. It provides a robust mechanism for controlling access and preventing unauthorized exploitation of valuable AI investments.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024