BinEnhance: An Enhancement Framework Based on External Environment Semantics for Binary Code Search
Yongpan Wang
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Binary Analysis
Overview
This talk introduces BinEnhance, a novel enhancement framework designed to significantly improve the accuracy and robustness of binary code search. Presented by Limbo J on behalf of author Yongpan Wang, the research addresses critical challenges in identifying vulnerable or similar code segments within vast binary landscapes, a task complicated by diverse compiler optimizations and the sheer scale of modern software. BinEnhance tackles the limitations of existing internal code semantic models by integrating valuable external environment semantic information, thereby reducing both false positive and false negative search results.
Key moments
- 0:00 Introduction, problem statement, and binary code search challenges
- 1:50 Defining external environment semantics and core motivation
- 3:40 Introducing BinEnhance: a general enhancement framework
- 4:40 External Environment Semantic Graph (EESG) construction algorithm
- 5:00 Evaluation setup, datasets, baselines, and research questions
- 5:50 BinEnhance demonstrates significant performance improvement over baselines
- 6:20 BinEnhance robustness across architectures and compiler options
- 8:00 Real-world one-day vulnerability detection capabilities demonstrated
BinEnhance: An Enhancement Framework Based on External Environment Semantics for Binary Code Search
Speakers: Yongpan Wang
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=uiviXkuCHDs
Overview
This talk introduces BinEnhance, a novel enhancement framework designed to significantly improve the accuracy and robustness of binary code search. Presented by Limbo J on behalf of author Yongpan Wang, the research addresses critical challenges in identifying vulnerable or similar code segments within vast binary landscapes, a task complicated by diverse compiler optimizations and the sheer scale of modern software. BinEnhance tackles the limitations of existing internal code semantic models by integrating valuable external environment semantic information, thereby reducing both false positive and false negative search results.
The importance of BinEnhance stems from the pervasive reuse of open-source code in software development, a practice that, while cost-effective, introduces substantial security risks. With 96% of software reportedly incorporating open-source components, and 89% relying on versions over four years old, the window for attackers to exploit known (1-day) vulnerabilities is wide. BinEnhance offers a powerful, automated solution to detect these insecure components, bolstering software supply chain security and mitigating the impact of sophisticated, low-cost attacks that leverage reused vulnerable code.
Background
▶ Watch: Introduction, problem statement, and binary code search challenges (0:00)
The modern software ecosystem is heavily reliant on the reuse of open-source code, a practice driven by the desire to reduce development costs and accelerate time to market. Industry reports indicate that an overwhelming 96% of current software projects incorporate open-source components, with a significant 89% utilizing versions that are over four years old. This widespread and often outdated code reuse creates a critical attack surface. The extensive workload associated with code auditing, coupled with the inherent complexity of recursive code reuse, frequently leads to substantial delays in patching known vulnerabilities. Attackers have capitalized on this reality, making the exploitation of 1-day vulnerabilities in reused code a highly effective, low-cost, and large-scale attack method, posing severe risks to organizations and end-users alike.
Binary code search has emerged as a potent methodology for automating the detection of insecure software components. This process involves the meticulous analysis of numerous binary codes to identify segments that exhibit the highest degree of similarity. However, this task is fraught with challenges. The syntactic structure of binary code can vary dramatically due to different compiler settings, including optimization levels, target architectures, and specific features like function inlining and splitting. These variations can obscure semantic similarities, making reliable detection difficult.
Existing binary code search approaches typically fall into two categories based on their semantic focus:
- Internal code semantics: These methods concentrate on the function itself, deriving semantic information directly from the binary code embedded within the function or from its derivatives, such as control flow graphs or data flow graphs.
- External environment semantics: These approaches infer semantic meaning from the function's broader context, utilizing information from supplementary functions, data, and their interrelationships. BinEnhance primarily innovates in this latter category.
The limitations of current methods are illustrated through several motivating examples. For instance, function inlining during compilation can lead to false negative search results, where semantically similar functions are missed because their binary representations diverge significantly. Conversely, functions with similar internal code semantics, such as a handle_transaction and a transfer function, might appear similar at a superficial level, leading to false positive results if their external context is not considered. The "missing calls problem" and similar function structures further contribute to these issues, demonstrating how a sole reliance on internal code semantics or rudimentary call graphs is insufficient for complex real-world scenarios.
The research identifies three core problems with existing binary code search solutions:
- Substantial variations in internal code semantics: Compiler settings, including function inlining and splitting, cause significant divergence in the binary representation of semantically identical or similar functions.
- Insufficient reliance on function call graphs: Exclusive reliance on function call graphs for contextual assistance proves inadequate for addressing the intricacies of real-world scenarios.
- Limited scalability: Current solutions struggle to cope with the demands of large-scale function search tasks, limiting their practical applicability.
To overcome these deeply entrenched problems, the BinEnhance framework was developed, aiming to provide a general, scalable, and robust enhancement for binary code search by integrating crucial external environment semantics.
Key Findings
▶ Watch: Introducing BinEnhance: a general enhancement framework (3:40)
BinEnhance demonstrates a profound impact on the landscape of binary code search, delivering several key findings that underscore its effectiveness and practical utility:
- Significant Improvement Across Baselines: BinEnhance consistently achieves significant improvements in search accuracy when applied to existing internal code semantics models. Evaluated using Mean Average Precision (MAP) scores on two public datasets, the framework demonstrates a clear positive correlation between improvement and the size of the function pool, indicating enhanced performance for larger codebases.
- Robustness Against Compiler Optimizations and Architectures: The framework proves to be remarkably robust against variations introduced by different compiler optimization options and target architectures. BinEnhance enhances baselines across cross-architecture and cross-optimization option tasks, showing no significant performance dips under various compilation settings. This resilience is critical for real-world applicability where binaries originate from diverse compilation environments.
- Effective Mitigation of Function Inlining Impact: One of the most challenging compiler optimizations, function inlining, is effectively mitigated by BinEnhance. While inlining typically leads to substantial performance loss in binary code search tasks, BinEnhance significantly reduces this decline, improving the baseline methods' ability to cope with such aggressive optimization strategies.
- Crucial Role of Each Component: A detailed evaluation revealed that every component within the BinEnhance framework plays a crucial and indispensable role in achieving the final superior results, highlighting the thoughtful design and integration of its constituent parts.
- Enhanced Efficiency: Despite the additional computational cost incurred during the training phase and the generation of function embeddings, BinEnhance ultimately leads to a significant reduction in the total time required for binary code search tasks. This indicates an optimized workflow where initial investment in embedding generation yields substantial returns in search efficiency.
- Superior Real-World Vulnerability Detection: In a critical evaluation using the D3 firmware alongside a CVE vulnerability functions dataset, BinEnhance demonstrated exceptional performance in detecting real-world 1-day vulnerabilities. It successfully identified 101-day vulnerabilities across 37 firmware images, achieving a MAP score of 67.9%. This represents a substantial improvement, detecting 12 more vulnerabilities than the existing state-of-the-art method, Herim, and boasting a MAP score 7.7% higher.
These findings collectively establish BinEnhance as a cutting-edge solution that not only addresses long-standing challenges in binary code search but also delivers tangible improvements in accuracy, robustness, and efficiency, particularly in the context of identifying critical vulnerabilities in complex software environments.
Technical Deep Dive
▶ Watch: Evaluation setup, datasets, baselines, and research questions (5:00)
The BinEnhance framework is meticulously designed to overcome the limitations of internal code semantics by integrating rich external environment semantics. It operates through a three-stage process: Node Initial Embedding Generation, Function Embedding Enhancement, and Similarity Combination.
1. Node Initial Embedding Generation:
The first stage focuses on creating initial vector representations, or embeddings, for the fundamental units within the binary code. BinEnhance identifies and encodes two primary types of nodes:
- Function nodes: These represent individual functions within the binary. Their embeddings are initially generated using existing internal code semantics methods, which serve as the baselines that BinEnhance aims to enhance. This allows the framework to leverage the strengths of established techniques while adding its own layer of external semantic context.
- String nodes: These represent string literals found within the binary. Strings often carry significant semantic meaning, indicating functionality, error messages, or configuration parameters, and thus contribute valuable contextual information.
After generating these initial high-dimensional embeddings, a widening transformation is applied. This dimensionality reduction technique is crucial for managing computational complexity and potentially distilling more salient features, making the embeddings more efficient for subsequent processing without losing critical information.
2. Function Embedding Enhancement:
This is the core stage where external environment semantics are introduced and integrated. BinEnhance achieves this through the construction and processing of an External Environment Semantic Graph (EESG).
- EESG Construction: The EESG is a sophisticated graph structure designed to capture the relationships between different entities in the binary's external environment. It contains two types of nodes: function nodes and string nodes, inheriting from the initial embedding generation stage. Crucially, the EESG defines four novel types of edges that represent distinct external environment semantics:
- Core Dependency: This likely refers to dependencies beyond simple call graphs, such as shared data structures or indirect control flow relationships.
- Data Code Use: This edge type captures instances where data (e.g., global variables, specific memory regions) is accessed or modified by code, establishing a semantic link between data and the functions operating on it.
- Address Adjacency: This refers to the spatial proximity of functions or data in memory. Functions or data segments that are frequently located close to each other might indicate a semantic relationship or shared purpose that is not immediately apparent from control or data flow alone.
- String Use: This edge type connects functions to the string literals they utilize. As mentioned, strings often provide strong hints about a function's purpose or the context in which it operates.
The talk describes a specific EESG construction algorithm, though the detailed steps are not elaborated in the transcript. This algorithm is responsible for populating the EESG with nodes and establishing these four types of edges based on static analysis of the binary.
- Semantic Enhancement Method: Once the EESG is constructed, BinEnhance employs a specialized graph neural network architecture for semantic enhancement. This method consists of multi-RCG layers (likely referring to Relational Graph Convolutional Network layers) and a residual block. RCG layers are particularly well-suited for processing graph data with different types of relations (edges), allowing the model to learn distinct transformations for each semantic relationship defined in the EESG. The residual block helps in training deeper networks by mitigating the vanishing gradient problem, enabling the model to learn more complex and nuanced representations. The specific embedding conversion formula used within these layers is provided in the paper, detailing how the initial node embeddings are iteratively refined by aggregating information from their neighbors across different edge types in the EESG. This process effectively infuses the function embeddings with rich contextual information from their external environment.
3. Similarity Combination:
The final stage combines the refined semantic embeddings to produce a comprehensive similarity score. BinEnhance uses a multi-faceted approach:
- Semantic Embedding Cosine Similarity: This measures the angular similarity between the enhanced function embeddings. A higher cosine similarity indicates that the functions are semantically closer in the learned embedding space, reflecting the influence of both internal and external semantics.
- Jaccard Similarity of Readable Data Features: In addition to the learned embeddings, BinEnhance also incorporates Jaccard similarity based on readable data features. These features could include symbols, imported/exported functions, or other identifiable attributes that provide direct, human-interpretable clues about a function's nature. Jaccard similarity is effective for comparing sets of discrete features.
The framework defines a specific combination formula that intelligently merges these different similarity scores. This formula, along with a carefully designed loss function, guides the training process, ensuring that the combined score accurately reflects true functional similarity while minimizing false positives and negatives. By integrating both learned semantic similarities and explicit feature-based similarities, BinEnhance achieves a robust and highly accurate similarity metric.
The framework's innovation lies in its systematic approach to harnessing external environment semantics, moving beyond the limitations of purely internal representations. The specific choices of novel external semantics (core dependency, data code use, address adjacency, string use) and the sophisticated graph neural network architecture (EESG, multi-RCG layers, residual block) are key to its superior performance.
Demo / Proof of Concept
▶ Watch: BinEnhance demonstrates significant performance improvement over baselines (5:50)
While the talk did not feature a live, interactive demonstration of the BinEnhance framework, its effectiveness and practical applicability were rigorously validated through extensive empirical evaluation on both publicly available datasets and real-world firmware. This comprehensive validation served as the primary proof of concept, demonstrating BinEnhance's capabilities under various conditions.
The evaluation utilized two distinct public datasets. Crucially, these datasets included two versions: "normal" compilation and "no in-line," specifically designed to assess BinEnhance's ability to handle compiler optimizations like function inlining. Five existing internal code semantics models were selected as baselines to provide a comparative benchmark for BinEnhance's performance.
The evaluation aimed to answer six key questions:
- Improvement over Baselines: How much could BinEnhance improve existing internal code semantics models? The results showed significant improvements across all baselines, measured by the average increment in MAP scores on both public datasets. Furthermore, the improvement for each baseline was positively correlated with the size of the function pool, indicating enhanced scalability and effectiveness for larger codebases.
- Robustness against Compiler Options and Architectures: Is BinEnhance robust against different compiler optimization options and architectures? The framework demonstrated strong resilience, enhancing baselines across cross-architecture and cross-optimization option tasks without suffering significant performance dips under various compilation settings.
- Mitigation of Function Inlining: Does BinEnhance effectively elevate the impact of function inlining in binary code search? The evaluation confirmed that function inlining typically leads to a substantial performance loss. However, BinEnhance effectively mitigated this decline, significantly improving the baseline methods' ability to cope with such aggressive optimization strategies.
- Contribution of Each Part: What is the contribution of each component within the BinEnhance framework? A careful evaluation confirmed that each component plays a crucial role in achieving the final results, underscoring the integrated design's effectiveness.
- Efficiency Improvement: Does BinEnhance improve the efficiency of baselines? While there is an additional time cost during the training phase and the generation of function embeddings, BinEnhance was shown to significantly reduce the overall time required for binary code search tasks, indicating an net efficiency gain for practical applications.
- Real-World 1-Day Vulnerability Detection: What was the performance of BinEnhance in detecting 1-day vulnerabilities in real-world firmware? This crucial test utilized the D3 firmware alongside a CVE vulnerability functions dataset. BinEnhance identified 101-day vulnerabilities in 37 firmware images, achieving a MAP score of 67.9%. This result was notably superior to existing methods, detecting 12 more vulnerabilities than Herim and achieving a MAP score 7.7% higher, conclusively demonstrating its real-world efficacy in securing software supply chains.
The comprehensive nature of this evaluation, addressing both theoretical robustness and practical impact, effectively served as a compelling proof of concept for the BinEnhance framework.
Defensive Implications
▶ Watch: Real-world one-day vulnerability detection capabilities demonstrated (8:00)
BinEnhance offers profound implications for cybersecurity defenders, providing a powerful new toolset to enhance the security posture of software systems, particularly those relying heavily on open-source components.
Firstly, the framework significantly improves the accuracy and speed of identifying known vulnerabilities (1-day CVEs) within reused code segments in large binaries and firmware. By leveraging external environment semantics, BinEnhance can detect vulnerable code patterns even when obscured by compiler optimizations like inlining, which often frustrate traditional signature-based or purely internal-semantic analysis tools. This capability translates directly into faster patching cycles and reduced exposure to prevalent attack vectors.
Secondly, BinEnhance strengthens software supply chain security. Given the widespread use of open-source components, organizations can integrate BinEnhance into their continuous integration/continuous deployment (CI/CD) pipelines or supply chain security audits. This allows for proactive scanning of third-party libraries and dependencies, ensuring that vulnerable versions are identified and remediated before deployment, thereby preventing insecure components from entering the production environment.
Thirdly, the framework's robustness against compiler optimizations and cross-architecture variations is a critical advantage for defenders. Attackers often employ obfuscation techniques, including specific compiler flags, to hide malicious or vulnerable code. BinEnhance's ability to maintain high detection rates despite these variations means that defenders are better equipped to uncover hidden threats, making the attacker's job significantly harder.
Furthermore, BinEnhance enhances the capabilities of existing binary analysis tools. It can serve as an enhancement layer for current internal code semantic models, providing them with valuable contextual information that reduces their false positive and false negative rates. This means organizations don't necessarily need to overhaul their entire security stack but can augment their current tools with BinEnhance's advanced semantic understanding.
Finally, the demonstrated efficiency gains, despite initial training costs, mean that BinEnhance can be applied to large-scale function search tasks in real-world environments without prohibitive computational overhead. This scalability is essential for analyzing vast codebases, such as entire operating systems or complex IoT firmware, enabling comprehensive vulnerability assessment that was previously impractical or too time-consuming. In essence, BinEnhance empowers defenders with a more intelligent, resilient, and efficient method for uncovering critical security flaws embedded deep within binary code.
Key Takeaways
- Addressing Core Challenges: BinEnhance effectively tackles long-standing limitations in binary code search by integrating novel external environment semantics, significantly improving accuracy and robustness against compiler optimizations.
- Enhanced Vulnerability Detection: The framework demonstrates superior performance in real-world scenarios, identifying 1-day vulnerabilities in firmware with higher precision (67.9% MAP score, 12 more vulnerabilities than Herim) than existing methods.
- Robustness Across Environments: BinEnhance is highly resilient to diverse compilation settings, including aggressive optimization options and cross-architecture variations, making it a reliable tool for analyzing binaries from varied sources.
- Leveraging Contextual Information: By incorporating external semantics like core dependency, data code use, address adjacency, and string use, BinEnhance reduces both false positives and false negatives, providing a more reliable similarity assessment.
- Efficiency and Scalability: Despite an initial investment in training and embedding generation, BinEnhance ultimately reduces the overall time for binary code search tasks, proving efficient and scalable for large-scale code analysis.
- Strengthening Software Supply Chain: The framework offers a powerful mechanism for proactive identification of known vulnerabilities in reused open-source components, thereby enhancing software supply chain security and accelerating patching efforts.
About the Speaker(s)
The research paper "BinEnhance: An Enhancement Framework Based on External Environment Semantics for Binary Code Search" was authored by Yongpan Wang. The presentation at the NDSS Symposium was delivered by Limbo J on behalf of Yongpan Wang. The transcript does not provide specific titles, affiliations, or further biographical details for either speaker.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic binary similarity research with a coherent technical contribution — the EESG with four typed edge relationships is a real idea, and the firmware vulnerability detection results are concrete enough to be meaningful. The work is competent and the evaluation is structured, but it doesn't push the state of the art far enough to be a must-see, and the transcript reads like a paper summary rather than a talk that would hold a room.
Heather Calloway (CISO) — WEAK
Technically competent research on binary code search that advances the state of the art in a narrow but real problem space. The supply chain framing is legitimate, but the talk never crosses from research contribution to operational guidance — no one in a security program leaves knowing what to do differently.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025