The great SAST dissonance: how to please every audience, at scale

Claudio Merloni (Staff Security Researcher · SemGrep), Romain Gaucher

BSidesSF 2026 · Day 2 · AMC Theatre 10

Overview

Claudio Merloni, a Staff Security Researcher at Semgrep, delivered a compelling talk at BSides SF, dissecting the pervasive challenge of achieving comprehensive Static Application Security Testing (SAST) coverage in modern, diverse software environments. Titled "The great SAST dissonance: how to please every audience, at scale," Merloni's presentation illuminated the inherent conflict between the vast and ever-evolving landscape of software libraries and frameworks, and the practical limitations of manually developing and maintaining SAST rules. He argued that traditional approaches to SAST rule creation are fundamentally unscalable, leading to significant gaps in vulnerability detection, particularly within the "long tail" of less popular but equally critical libraries and proprietary code.

Watch on YouTube

Key moments

  1. 0:00 Talk introduction and 'SAST Dissonance' title explanation
  2. 2:00 Talk agenda and focus on SAST coverage
  3. 3:36 Defining SAST and its different forms
  4. 5:59 Fundamental challenges and complexity of SAST
  5. 6:53 Limitations of typical SAST evaluation benchmarks
  6. 8:36 Introducing library/framework-specific SAST benchmarks

The great SAST dissonance: how to please every audience, at scale

Speakers: Claudio Merloni, Staff Security Researcher, Semgrep; Romain Gaucher, Head of Security Research, Semgrep

Conference: BSides SF

YouTube: https://www.youtube.com/watch?v=kFneu_x5xMI

Overview

Claudio Merloni, a Staff Security Researcher at Semgrep, delivered a compelling talk at BSides SF, dissecting the pervasive challenge of achieving comprehensive Static Application Security Testing (SAST) coverage in modern, diverse software environments. Titled "The great SAST dissonance: how to please every audience, at scale," Merloni's presentation illuminated the inherent conflict between the vast and ever-evolving landscape of software libraries and frameworks, and the practical limitations of manually developing and maintaining SAST rules. He argued that traditional approaches to SAST rule creation are fundamentally unscalable, leading to significant gaps in vulnerability detection, particularly within the "long tail" of less popular but equally critical libraries and proprietary code.

The core of Merloni's discussion centered on the inadequacy of current SAST solutions to adapt to the sheer volume and variety of code dependencies used across enterprise applications. He highlighted that even extensive coverage of the most popular libraries leaves a substantial portion of repositories vulnerable, due to unique dependency sets and internal codebases. To address this "dissonance," Merloni unveiled an innovative, AI-assisted automation pipeline developed at Semgrep, designed to dramatically accelerate the identification of security-relevant code patterns and the subsequent generation of high-quality SAST rules. This pipeline integrates static analysis with large language models (LLMs) and a human-in-the-loop validation process, demonstrating a path toward achieving scalable and effective SAST coverage in an increasingly complex development ecosystem.

This talk is particularly significant for anyone involved in application security, software development, or DevSecOps. It provides a data-driven critique of existing SAST paradigms and offers a concrete, technically detailed solution to a problem that plagues many organizations: how to secure a sprawling codebase without an army of security researchers. By demonstrating how AI can augment human expertise, Merloni presented a blueprint for a future where SAST tools can genuinely "please every audience" by offering robust, tailored coverage at an unprecedented scale, ultimately leading to the detection of more real-world vulnerabilities.

Background

▶ Watch: Talk introduction and 'SAST Dissonance' title explanation (0:00)

Static Application Security Testing (SAST) is a critical component of a robust application security program. At its core, SAST involves analyzing an application's source code, bytecode, or binary code without executing it, to identify potential security vulnerabilities. The idea is straightforward: feed the code into a tool, and it automatically flags security issues. However, as Merloni articulated, the practical implementation of this idea is anything but simple.

SAST tools can range from basic string searching (like grep, which merely checks for the existence of text patterns) to highly sophisticated program analysis. More advanced SAST techniques parse the source code to understand its syntactic structure (e.g., using Abstract Syntax Trees or ASTs), and critically, track data flow. Data flow analysis identifies how data moves through an application from a source (where untrusted input enters the system, such as user input from an HTTP request) to a sink (where that input could be used in a dangerous way, such as in an SQL query or a file path operation). Identifying these data flows is paramount for detecting common vulnerabilities like SQL injection, Cross-Site Scripting (XSS), or path traversal. Merloni noted that while AI and Large Language Models (LLMs) are increasingly integrated into SAST, the fundamental challenges of program analysis remain.

The complexity of SAST stems from several factors. Modern applications are typically polyglot, utilizing multiple programming languages. Each language has its own syntax, semantics, and ecosystem of libraries and frameworks. A SAST tool must not only understand the intricacies of these languages but also comprehend the security implications of countless libraries and frameworks. For instance, a web framework might implicitly inject instances of certain classes, a behavior not immediately obvious from raw source code. Furthermore, the way an application is deployed—its actual attack surface—adds another layer of complexity. This forms a "never-ending quest for finding vulnerabilities."

Merloni was particularly critical of how SAST tools are often evaluated, especially through benchmarks. Benchmarks like the OWASP Java Benchmark or tools like DVWA (Damn Vulnerable Web Application) and WebGoat are frequently used. While these benchmarks are useful for evaluating specific capabilities, they often fall short in representing real-world enterprise code. For example, the OWASP Java Benchmark focuses on specific syntactic and semantic patterns, measuring a tool's ability to understand variations of these patterns. DVWA and WebGoat, originally designed as learning tools, feature extremely simple code patterns that bear little resemblance to the complex, highly abstracted code found in production applications. Consequently, a tool performing well on these benchmarks might still fail to find issues in a typical enterprise codebase.

A more significant problem arises with benchmarks that use specific sets of libraries, such as Jusha. While such benchmarks measure a tool's ability to understand sources and sinks within those particular libraries, they don't reflect the vast diversity of libraries used in the wild. Merloni posed a critical question: what are the chances that an organization's repositories will use exactly the same set of libraries featured in a benchmark? Even if a popular logging library overlaps, the chances of a complete match across an entire enterprise's diverse portfolio are minimal. This leads to the central concept of the talk: coverage.

Coverage, in the context of SAST, is a nuanced and multifaceted concept. It can refer to:

  • Language coverage: Does the tool support all the languages used in the codebase?
  • Semantic coverage: Does the tool understand the meaning and intent of the code, including complex framework behaviors?
  • Security knowledge coverage: Does the tool know about specific vulnerability types (e.g., SQL injection, path traversal) and how they manifest?
  • Library and framework coverage: Does the tool understand the security properties (sources, sinks, sanitizers) of all the libraries and frameworks used?

The last point, library and framework coverage, was Merloni's primary focus. He introduced the concept of the long tail of libraries. While popular libraries like Python's requests or pyyaml are undoubtedly critical for SAST tools to cover, there are hundreds of thousands of less popular libraries. Merloni cited examples like HVAC (a Python library for HashiCorp Vault access) or pyGithub (a GitHub API client). While not universally popular, if an application uses them, they become critical. A simple snippet using HVAC, for instance, could expose Server-Side Request Forgery (SSRF) or hardcoded secret issues if the tool doesn't understand its functions. Without specific rules for these libraries, a SAST tool will simply miss any vulnerabilities stemming from their usage, regardless of their popularity.

To quantify this "long tail" problem, Merloni presented data from Semgrep's customer base, analyzing approximately 80,000 repositories across Python, Java, and JavaScript ecosystems. This analysis identified 7,000 unique direct dependencies. The findings were stark:

  • Even if SAST tools had perfect coverage for the top 1,000 most dependent-upon libraries, only 16% of repositories would be entirely covered. This means 84% of repositories would still have dependencies for which no rules exist.
  • Extending coverage to the top 20,000 libraries still only covered 46% of repositories entirely.
  • The situation was even more pronounced in the npm ecosystem, where supporting the top 1,000 libraries covered only a "tiny fraction" of repositories.

Further compounding the problem is the sheer complexity of modeling even a single popular framework. Merloni used Django as an example, noting it has over 900 modules, 2,000 classes, and 9,000+ callable functions and methods. Manually writing rules for all security-relevant functions within such a framework would take "literally months." Additionally, languages like Python allow for symbol re-exporting, meaning the same function can be accessed via multiple different names (e.g., django.views.generic.base.View vs. django.views.generic.View). SAST tools need this additional knowledge to track data flows accurately. The manual effort required to identify which of these 9,000+ callables are security-relevant (e.g., executing SQL queries, making HTTP calls) is immense and tedious.

Finally, Merloni presented data on dependency overlap across repositories within the same enterprise. On average, only 10% of dependencies overlap between any two given repositories. This means that while individual repositories might have a relatively small number of dependencies, the total unique set of dependencies across an enterprise is enormous and highly varied. This low overlap underscores the necessity for broad, scalable coverage rather than a narrow focus on a few common libraries. The problem, therefore, is not just about the volume of libraries but also their extreme diversity across an organization's codebase.

Key Findings

▶ Watch: Defining SAST and its different forms (3:36)

Claudio Merloni's talk illuminated several critical findings regarding the state of SAST and proposed a transformative approach to overcome its limitations:

  1. The "Long Tail" is a Critical Blind Spot: Merloni's data unequivocally demonstrated that relying on SAST tools with coverage primarily for the most popular 100 or even 1,000 libraries leaves the vast majority of enterprise repositories vulnerable. The "long tail" of less common, but equally critical, libraries and custom internal code constitutes a significant blind spot where numerous vulnerabilities reside undetected. For instance, covering the top 1,000 libraries in Python, Java, and JavaScript only provided full coverage for 16% of analyzed repositories, highlighting the inadequacy of a narrow focus.
  1. Manual Rule Creation is Unscalable and Error-Prone: The traditional method of security researchers manually identifying security-relevant functions within libraries and crafting SAST rules is an extremely time-consuming and error-prone process. Merloni estimated that writing rules for a single complex library could take days, involving extensive research, rule writing, testing, and benchmarking. This manual effort is simply unsustainable given the hundreds of thousands of libraries and their constant evolution. Manual rule writing also introduces the risk of typos and inconsistencies, which are hard to detect and fix at scale.
  1. AI-Assisted Automation is Essential for Scaling Coverage: The talk's central finding is that an AI-driven automation pipeline, combined with static analysis and human oversight, is the only viable path to achieving scalable SAST coverage. By offloading the laborious tasks of library analysis, function classification, and initial rule generation to AI, security teams can dramatically increase their output and focus on higher-value tasks like vulnerability assessment.
  1. Quantifiable and Significant Impact: Merloni presented compelling results from Semgrep's implementation of this pipeline. In just two weeks, one person was able to increase the number of Python rules by 10x (from approximately 700 to 7,000) and effectively double the number of supported libraries. This represents an enormous scaling factor compared to the previous manual efforts that took months for similar increases.
  1. Validation Through Real Vulnerability Detection: Crucially, the expanded coverage generated by this automated pipeline is not merely theoretical. Merloni confirmed that these newly generated rules are actively finding new vulnerabilities in customer codebases. This validates the premise that the "long tail" of packages, while seemingly less popular, indeed harbors exploitable security flaws that would otherwise be missed.
  1. Human Expertise is Augmented, Not Replaced: The pipeline emphasizes a human-in-the-loop approach. AI performs the initial heavy lifting of context gathering and annotation generation, but security researchers remain critical for triaging these findings, validating their security relevance, and ultimately approving the rule generation. This ensures the quality and accuracy of the rules while leveraging human expertise where it matters most.

Technical Deep Dive

▶ Watch: Fundamental challenges and complexity of SAST (5:59)

The core of Merloni's solution to the SAST dissonance is a sophisticated AI-assisted automation pipeline that fuses traditional static analysis with the power of large language models and a crucial human-in-the-loop component. This pipeline systematically addresses the challenges of library discovery, security classification, and rule generation at scale.

The pipeline comprises four main stages:

1. Prioritization

Given the overwhelming number of libraries (Merloni mentioned 100,000+), the first challenge is deciding where to focus rule development efforts. This stage is primarily AI-driven and aims to identify the most impactful libraries to cover first.

  • Input: The system feeds a variety of information into its AI model, including:
  • Customer Usage Data: How many customers or repositories utilize a specific library? This helps prioritize libraries that maximize impact across the user base.
  • Repository Classification: AI classifies repositories (e.g., as web applications, AI applications, fintech applications) based on their dependencies and code patterns. This helps infer the types of vulnerabilities most relevant to those applications.
  • Dependency Code Analysis: The system analyzes the code of the dependencies themselves to identify characteristics that suggest potential security risks. For example, if a library frequently interacts with databases, it might be a candidate for SQL injection rules.
  • Output: The AI assigns a score to each library, indicating its priority. A fancy UI, built by Romain Gaucher, visualizes these scores alongside identified "risks" (e.g., "SQL injection sinks"). This allows security researchers to quickly identify high-priority libraries and launch "campaigns" to develop rules for specific vulnerability types within those libraries. The UI includes a button to "make more magic things happen," initiating the next stage.

2. Annotation Generation

This is the second stage where AI plays a heavy role, taking a prioritized library and automatically identifying security-relevant parts of its code.

  • Process:
  1. The system downloads the library's source code and performs initial static analysis to transform it into an analyzable format.
  2. It extracts a list of all symbols (functions, methods, classes, arguments) within the library.
  3. This information, along with documentation and code samples, is fed into a series of specialized AI agents.
  4. Specialized AI Agents:
  • Code/Documentation Agent: Analyzes the function's code and its documentation (if available) to understand its purpose and the nature of its arguments (e.g., identifying if an argument is a "file path").
  • Security Review Agent: Characterizes the function from a security perspective, determining if it acts as a source (introduces untrusted data), a sink (consumes data in a potentially dangerous way), or a sanitizer (cleanses data).
  • Vulnerability-Specific Agent: Provides a deeper layer of analysis for particular vulnerability types. For example, if a "file path" argument is identified, this agent might check if the function performs any sanitization or validation that would mitigate path traversal vulnerabilities.
  • Independent Review Agent: Acts as an orchestrator and validator, taking the outputs from the other agents and assessing their consistency and overall confidence in the identified security property.
  • Output: The system generates a set of annotations for the library's symbols. An annotation signifies that a specific function or argument is potentially security-relevant for a particular vulnerability type, along with a confidence rating. For a library like B-tree, Merloni showed 124 annotations generated for 696 symbols.

3. Human in the Loop (Triage)

This is the critical juncture where human expertise validates and refines the AI's output, ensuring accuracy and relevance.

  • Process: Security researchers are presented with a dedicated UI to review the AI-generated annotations.
  • The UI displays the annotated code, highlighting specific function calls or arguments (e.g., an upload_file function with a file_name argument).
  • For each annotation, the UI provides rich context: the identified security property (e.g., "path manipulation" as a "sink"), its confidence rating ("certain"), and the decisions made by the various AI agents.
  • Researchers can quickly "sift through" this information, confirm the annotation's validity, or reject it.
  • Output: Once a researcher confirms an annotation (e.g., by clicking a "blue bubble" in the UI), they can initiate a "Create PR" action. This validated annotation then proceeds to the final stage. This step ensures that only high-quality, relevant security properties are used for rule generation, preventing false positives and maintaining trust in the SAST tool.

4. Rule Synthesis

This is the final, fully automated stage where validated annotations are transformed into executable SAST rules.

  • Process:
  1. Stub Generation: For each validated annotation, the system automatically creates a stub file. This stub is a simplified version of the original library code, with function bodies removed and the security property injected as a type annotation. For example, a stub might look like:
  1. Rule Compiler: A specialized rule compiler then processes these stubs.
  • Type Checking & Validation: The compiler parses the stubs, ensuring that the sangrep.taint.sink (or similar) annotations are syntactically correct and properly typed. This eliminates common typos and errors that occur in manual rule writing.
  • Automatic Rule Generation: The compiler automatically generates the complex SAST rules required by the underlying analysis engine (e.g., Semgrep's engine).
  • Pattern Compaction: The compiler can optimize and compact patterns. Merloni gave an example where three distinct annotations (e.g., execute, scalar, scholars) were combined into a single, more efficient pattern like rag_x_execute_scalar_scholars.
  • Dynamic Language Specificities: For dynamic languages where type inference is challenging, the compiler automatically adds necessary requires tags or type information to the rules. This ensures that a sink is only triggered if a method is called on an object of a specific type, preventing over-alerting.
  • Benefits: This automated synthesis offers tremendous advantages:
  • Speed: Rules are generated instantly from validated annotations.
  • Accuracy & Consistency: Eliminates manual typos and ensures all rules follow consistent patterns.
  • Maintainability: If the SAST engine (or "rule compiler") improves, rules can be regenerated automatically without manual rewriting.
  • Scalability: This is the key to handling thousands of libraries and their complex interactions. The compiler automatically manages re-exported symbols and other linguistic nuances that would be time-consuming and error-prone for humans.

In essence, Merloni's technical solution leverages static analysis to extract structured code information, uses AI to interpret and classify its security relevance, relies on human expertise for critical validation, and then automates the final, complex step of rule generation. This symbiotic relationship between static analysis, AI, and human intelligence is what enables the system to tackle the "long tail" problem effectively and scale SAST coverage dramatically.

Demo / Proof of Concept

▶ Watch: Limitations of typical SAST evaluation benchmarks (6:53)

While Claudio Merloni's talk did not feature a live coding demonstration, it extensively detailed and visually presented the proof of concept for the described AI-assisted automation pipeline. The "demo" was embedded within the technical deep dive, showcasing the functional components and user interfaces that enable this scalable approach to SAST rule generation.

Merloni provided screenshots and explained the workflow of two key interfaces:

  1. Prioritization UI: This interface, developed by Romain Gaucher, was shown as a list of libraries, each with a "score" highlighted in green. This score, generated by the AI, indicates the library's priority based on factors like customer usage and potential security risks. The UI also displays identified "risks" associated with each library (e.g., "SQL injection syncs" for SQLAlchemy). This visual representation demonstrates how security researchers can quickly identify high-impact libraries and specific vulnerability types to target for rule development. The presence of a "create PR" like button on this UI implies the direct initiation of the annotation and rule synthesis process from this prioritization step.
  1. Annotation Triage UI: This interface is where the human-in-the-loop validation occurs. Merloni presented a screenshot showing annotations generated for a "B-tree" library, specifically highlighting a function called upload_file. Within this function, the file_name argument was annotated by the AI as a "path manipulation" sink, ranked as "certain." The UI provided a rich context, displaying the decisions and confidence levels from the various AI agents (Code/Documentation Agent, Security Review Agent, etc.). This visual proof demonstrated how the system presents complex AI-generated insights in an easily digestible format, allowing a security researcher to efficiently triage and validate hundreds of potential security properties. The ability to click a button to "create a PR" directly from this interface underscored the seamless integration of human validation into the automated rule synthesis process.

Furthermore, Merloni illustrated the output of the automated rule synthesis by showing an example of a stub file. This stub, containing the original library's function signature but with the function body removed and a Semgrep-specific type annotation (e.g., file_name: sangrep.taint.sink(kind="path_manipulation")), served as concrete evidence of the system's ability to translate human-validated annotations into machine-readable, rule-generating input. He also displayed snippets of the complex, automatically generated Semgrep rules, demonstrating how the rule compiler compacts patterns and adds necessary type information for dynamic languages.

These visual and textual explanations served as a robust proof of concept, demonstrating that the pipeline is not merely a theoretical construct but a working system capable of significantly accelerating SAST rule development and expanding coverage in a practical, scalable manner. The quantifiable results – a 10x increase in Python rules and doubled library support in two weeks – further solidified the effectiveness of this approach.

Defensive Implications

▶ Watch: Introducing library/framework-specific SAST benchmarks (8:36)

Claudio Merloni's talk provides crucial insights and actionable recommendations for organizations looking to strengthen their application security posture, particularly in the realm of SAST. The defensive implications are profound, urging a shift in how security teams approach static analysis:

  1. Prioritize Comprehensive Coverage, Especially the "Long Tail": Defenders must recognize that focusing solely on the most popular libraries leaves a significant portion of their codebase vulnerable. The data presented by Merloni (e.g., only 16% of repositories fully covered by rules for the top 1,000 libraries) highlights the critical need to extend SAST coverage to the "long tail" of less common, but equally critical, open-source libraries and proprietary internal code. Organizations should audit their dependency usage and pressure SAST vendors for broader, more adaptable coverage.
  1. Embrace AI-Assisted SAST Tools: The traditional, manual approach to SAST rule development is unsustainable. Defenders should actively seek out and adopt SAST solutions that leverage AI and automation for rule generation and library analysis. This capability is essential for keeping pace with the rapid evolution of software ecosystems and ensuring that SAST tools remain relevant and effective. The ability of AI to provide context and classify security properties (as demonstrated by Merloni's agents) is a game-changer for scaling security research efforts.
  1. Invest in Custom Code Coverage: Proprietary internal libraries and frameworks are often critical to an organization's business logic but are completely unknown to off-the-shelf SAST tools. Defenders should explore implementing similar AI-assisted pipelines internally to generate SAST rules for their custom codebases. This ensures that internally developed components, which often handle sensitive data or critical operations, are subjected to the same rigorous security analysis as external dependencies.
  1. Shift Security Researcher Focus: The human-in-the-loop model presented by Merloni suggests a re-evaluation of security researcher roles. Instead of spending days manually researching libraries and writing rules, researchers can focus on high-value tasks such as validating AI-generated annotations, triaging critical findings, understanding complex vulnerability scenarios, and providing feedback to improve the AI models. This optimizes the use of scarce security expertise.
  1. Demand Automation and Consistency in Rule Development: If an organization is developing its own SAST rules (e.g., for custom linters or internal security tools), they should adopt automated rule synthesis and validation processes. Merloni's demonstration of a rule compiler that eliminates typos, compacts patterns, and automatically adds necessary type information underscores the importance of automation for speed, consistency, and maintainability. This ensures that rules are high-quality, reliable, and adaptable to changes in the underlying SAST engine.
  1. Recognize Coverage as a Continuous Problem: The problem of SAST coverage is not a one-time fix. New libraries, frameworks, and versions are constantly released. Defenders must adopt a mindset that views coverage as an ongoing, dynamic challenge requiring continuous adaptation and automation. The pipeline presented by Merloni offers a model for continuous integration of new library knowledge into SAST tools.
  1. Re-evaluate SAST Benchmarking: Organizations should move beyond simplistic benchmarks that don't reflect their actual technology stack. Instead, they should focus on evaluating SAST tools based on their ability to cover their specific set of dependencies, including less popular and internal libraries, and their efficacy in detecting vulnerabilities within complex, real-world codebases.

By integrating these defensive strategies, organizations can move towards a more proactive and scalable application security program, effectively addressing the "great SAST dissonance" and significantly reducing their attack surface.

Key Takeaways

  • Go Beyond Popular Libraries: Relying solely on SAST coverage for the top 100 or 1,000 most popular libraries is insufficient; the "long tail" of less common open-source and proprietary internal libraries harbors significant, often undetected, vulnerabilities.
  • AI Augments, Not Replaces, Human Expertise: AI is crucial for scaling SAST by automating tedious tasks like library analysis, function classification, and initial rule generation, allowing security researchers to focus on high-value validation and vulnerability assessment.
  • Automated Rule Synthesis is a Game-Changer: Leveraging a rule compiler to automatically generate SAST rules from validated annotations dramatically increases development speed, eliminates manual errors, ensures consistency, and simplifies maintenance for thousands of rules.
  • Quantifiable Impact on Coverage and Efficiency: The demonstrated pipeline led to a 10x increase in Python rules (from 700 to 7,000) and doubled supported libraries in just two weeks, proving the immense scalability and efficiency gains over manual methods.
  • Synergy of Static Analysis and AI: The most effective approach combines the deterministic, predictable power of traditional static analysis (for code parsing and structure) with the contextual understanding and classification capabilities of AI (for security relevance).
  • Coverage is a Continuous Challenge: The software landscape is constantly evolving with new libraries and frameworks; therefore, SAST coverage must be treated as an ongoing problem requiring an automated, adaptable, and continuously updated solution.

About the Speaker(s)

Claudio Merloni is a Staff Security Researcher at Semgrep, based in Paris. With over 15 years of experience spanning application security and software engineering, Claudio is deeply passionate about scaling security through innovative approaches. His work at Semgrep focuses on developing advanced detection rules for Static Application Security Testing (SAST) and promoting "paved road" initiatives that prioritize secure-by-default development practices. He has a long-standing interest in SAST, viewing it as a "very cool problem to kind of solve because you kind of never solve it."

Romain Gaucher is the Head of Security Research at Semgrep. He was present at the talk and acknowledged by Claudio as having built the fancy UI for the prioritization stage of the pipeline and for fielding questions from the audience. His role as head of research suggests significant contributions to the strategic direction and technical implementation of advanced security analysis methodologies at Semgrep.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Semgrep researchers present a real engineering problem — SAST coverage gaps across the long tail of library dependencies — and back it with actual customer data (80k repos, 7k unique deps) before walking through their AI-assisted pipeline to address it. The problem framing is honest and the solution is concrete, but this is ultimately a product research talk from a vendor about their own tooling, and it doesn't transcend that constraint.

Heather Calloway (CISO) — SOLID

Technically credible and the coverage data is genuinely useful — most AppSec teams don't have numbers that honest about how little their SAST actually sees. But this is a product team presenting their own pipeline at a practitioner conference, and it never quite escapes that gravity. The defensive implications section gestures at organizational action but doesn't reach it.

→ Top-rated talks at BSidesSF 2026

All talks from BSidesSF 2026