Cyber-Physical Deception Through Coordinated IoT Honeypots

Chongqi Guan

34th USENIX Security Symposium (USENIX Security '25) · Day 1 · System Security 1: Threat Detection, Exploitation, and Adaptive Defenses

Overview

In an era where scientific rigor and empirical validation are paramount, the computer security community, like many other scientific disciplines, faces increasing calls for improved research validity. This paper, "SoK: Towards a Unified Approach to Applied Replicability for Computer Security," by Daniel Olszewski, Tyler Tucker, Kevin R. B. Butler, and Patrick Traynor from the University of Florida, addresses a critical gap in this discourse. It systematically reviews over three decades of research on reproducibility, replicability, and validity, highlighting the inconsistencies and practical limitations of existing definitions within the context of computer security research. The authors argue that while reproducibility—achieving the same results with the same code and data—is important, it is often insufficient and sometimes unattainable, especially given the unique challenges of security studies.

Read the paper · Download the PDF (PDF) · Slides

Paper abstract

Reproducibility has been an increasingly important focus within the Security Community over the past decade. While showing great promise for increasing the quantity and quality of available artifacts, reproducibility alone only addresses some of the challenges to establishing experimental validity in scientific research and is not enough to move forward our discipline. Instead, replicability is required to test the bounds of a hypothesis and ultimately show consistent evidence to a scientific theory. Although there are clear benefits to replicability, it remains imprecisely defined, and a formal framework to reason about and conduct replicability experiments is lacking. In this work, we systematize over 30 years of research and recommendations on the topics of reproducibility, replicability, and validity, and argue that their definitions have had limited practical application within Computer Security. We address these issues by providing a framework for reasoning about replicability, known as the Tree of Validity (ToV). We evaluate an attack and a defense to demonstrate how the ToV can be applied to threat modeling and experimental environments. Further, we show two papers with Distinguished Artifact Awards and demonstrate that true reproducibility is often unattainable; however, meaningful comparisons are still attainable by replicability. We expand our analysis of two recent SoK papers, themselves replicability studies, and demonstrate how these papers recreate multiple paths through their respective ToVs. In so doing, we are the first to provide a practical framework of replicability with broad applications for, and beyond, the Security research community.

Visual summary for Cyber-Physical Deception Through Coordinated IoT Honeypots by Chongqi Guan
Visual summary for Cyber-Physical Deception Through Coordinated IoT Honeypots by Chongqi Guan

SoK: Towards a Unified Approach to Applied Replicability for Computer Security

Authors: Daniel Olszewski (University of Florida); Tyler Tucker (University of Florida); Kevin R. B. Butler (University of Florida); Patrick Traynor (University of Florida)

Conference: USENIX Security

Paper Page: https://www.usenix.org/conference/usenixsecurity25/presentation/olszewski

Overview

In an era where scientific rigor and empirical validation are paramount, the computer security community, like many other scientific disciplines, faces increasing calls for improved research validity. This paper, "SoK: Towards a Unified Approach to Applied Replicability for Computer Security," by Daniel Olszewski, Tyler Tucker, Kevin R. B. Butler, and Patrick Traynor from the University of Florida, addresses a critical gap in this discourse. It systematically reviews over three decades of research on reproducibility, replicability, and validity, highlighting the inconsistencies and practical limitations of existing definitions within the context of computer security research. The authors argue that while reproducibility—achieving the same results with the same code and data—is important, it is often insufficient and sometimes unattainable, especially given the unique challenges of security studies.

The core contribution of this work is the introduction of the Tree of Validity (ToV), a novel, flexible framework designed to provide a unified approach to reasoning about and conducting replicability experiments. The ToV offers a structured way to classify and compare validity studies by explicitly detailing the experimental components that remain "same" or are made "different" from the original work. This framework is particularly tailored to accommodate the adversarial nature and complex methodologies prevalent in security research, including the critical role of threat models which are often overlooked in generic validity taxonomies.

This paper is significant because it moves beyond the often-conflicting definitions of reproducibility and replicability to offer a practical, actionable framework. By demonstrating the ToV's application through several case studies—including well-known attacks like Spectre, machine learning defenses, and even the re-evaluation of a USENIX Distinguished Artifact Award winner—the authors illustrate how meaningful comparisons and advancements can be made even when perfect reproducibility is impossible. The ToV provides a much-needed common language for researchers to communicate the scope and limitations of their validity studies, thereby fostering greater transparency, rigor, and accelerated progress within the security research ecosystem.

Background

The scientific community has, over the past decade, grappled with a growing "reproducibility crisis," prompting extensive discussion and initiatives to enhance the trustworthiness of research findings. In computer science, and particularly in security, this has led to the widespread adoption of Artifact Evaluation Committees (AECs) at major conferences, aiming to ensure that published claims are computationally reproducible. While AECs have shown promise in making research artifacts more available, the precise definitions and practical applications of terms like reproducibility, replicability, and validity have remained ambiguous and often contentious.

Historically, various organizations and researchers have proposed taxonomies for these concepts, often with conflicting interpretations. The Claerbout Terminology, foundational to computational reproducibility, defined it as running the "same software on the same input and obtaining the same results" [18]. Peng [44] expanded on this, introducing replicability as "independent investigators address a scientific hypothesis and build up evidence for or against it," positioning reproducibility as a spectrum of artifact availability. However, Peng's definitions paradoxically placed fully reproducible (using original artifacts) as the closest substitute to replicability, despite replicability being fully independent.

The Association for Computing Machinery (ACM) initially defined repeatability (same team, same setup), replicability (different team, same setup), and reproducibility (different team, different setup) in 2016 [20], but later unified its definitions in 2020 to align with the Claerbout terminology, making reproducibility "different team, same experimental setup" and replicability "different team, different experimental setup" [21]. Despite this, the ACM's badging system for artifacts doesn't always fully align with its strict definitions, leading to inconsistencies.

The National Academies of Science (NAS), in its 2019 report, aimed to unify definitions across disciplines, demarcating reproducibility as verifying existing claims with available artifacts and replicability as testing the limits of a study in broader contexts [38]. The NAS also introduced notions of "direct" and "indirect" studies, acknowledging the practical limits of artifact availability. Further frameworks, such as Gomez et al.'s [22] focus on the purpose of replication (e.g., controlling for sampling error, operationalization limits) and Goodman et al.'s [23] distinction between method reproducibility, results reproducibility, and inferential reproducibility, highlight the diverse perspectives on validity. Gundersen [26] derived a framework from the scientific method for AI/ML, defining degrees and types of reproducibility based on outcomes and artifact usage (R1-Description, R2-Code, R3-Data, R4-Experiment).

Despite these efforts, the authors of this paper identify several persistent deficiencies:

  1. Lack of Unification: No single framework comprehensively covers all definitions or provides a clear, quantifiable comparison between different validity studies.
  2. Property vs. Action Ambiguity: Reproducibility and replicability are often interchangeably treated as both inherent properties of research and actions taken by researchers.
  3. Inflexibility: Existing taxonomies are not robust to the varying implementations of the scientific process or the unique complexities of computer security research.
  4. Limited Practical Application: They fail to provide an actionable framework for assessing and comparing experimental validity across studies, especially regarding partial artifact availability or dynamic experimental environments.

Compounding these issues are the unique challenges of security research, including the critical role of threat models (defining adversary capabilities and informing experimental procedures), the rapid pace of technological change, the sensitive nature of data (e.g., privacy concerns with user data), the immediate real-world impact of vulnerabilities, and the difficulty of reproducing complex systems. These factors often make "true" reproducibility impractical or even unethical, underscoring the need for a more robust and adaptable framework like the Tree of Validity.

Key Findings

The paper makes several significant contributions to the discourse on research validity in computer security:

  • Systematization of Validity: The authors conducted a comprehensive systematization of over 30 years of research and recommendations concerning reproducibility, replicability, and validity across various scientific disciplines. This analysis revealed that existing definitions are often imprecise, inconsistent, and have had limited practical application within the computer security community, failing to address its unique challenges.
  • Introduction of the Tree of Validity (ToV): A novel, formal framework, the Tree of Validity (ToV), is proposed as a unified approach to classifying and reasoning about validity experiments. The ToV conceptualizes an experiment as a series of components (Problem, Method, Data, Analysis, Domain) and represents validity studies as paths through a binary tree, where each node signifies whether a component is kept "Same" or made "Different" from the original experiment. This framework is dynamic, allowing for custom layers and adaptable to diverse methodologies, including the complex adversarial nature of security research.
  • Reconciling Definitions and Quantifying Differences: The ToV framework effectively unifies previous, often conflicting, definitions of reproducibility and replicability by mapping them to specific paths or regions within the tree. Crucially, it moves beyond qualitative labels to provide a quantitative means of expressing the differences between validity studies, enabling more precise comparisons and meta-analyses.
  • Accommodation of Security-Specific Challenges: The ToV is designed to explicitly incorporate security-specific elements, most notably the threat model (attacker and defender capabilities). This allows authors and validators to formally specify the adversarial scope under which findings hold, and to model concepts like adaptive adversaries as different executions through the ToV.
  • Demonstrated Utility through Case Studies: The paper validates the ToV's practicality through several detailed case studies:
  • Spectre Attacks [31]: The ToV successfully models different attack variants and threat models, demonstrating its ability to characterize complex security vulnerabilities and their replication potential, even with ethical concerns around proof-of-concept release.
  • Entangled Watermarks for ML Defenses [29]: The framework effectively captures the nuances of defense mechanisms and their evaluation against various adaptive adversaries, highlighting how artifact availability dictates the potential scope of replication.
  • USENIX Distinguished Artifact Award Winner [16]: A replication study of Bollinger et al.'s work on GDPR cookie violation detection revealed that true computational reproducibility was unattainable due to dynamic environmental factors (e.g., 14.9% of original URLs no longer resolved, significant changes in detected cookie types). However, the ToV still enabled meaningful comparisons and contextualization of the results.
  • Guidance for Authors and AECs: The ToV, particularly through the concept of a Seed of Validity (SoV), provides a practical tool for authors to proactively identify and communicate the determinative factors and potential sources of variation in their experimental processes. This can inform artifact evaluation committees and facilitate better-contextualized research contributions.

In essence, the key finding is that while true reproducibility is often unattainable in security research, the ToV provides a robust and flexible framework to achieve meaningful comparisons and advance scientific theories through systematic replicability, addressing critical gaps in current meta-science practices within the field.

Technical Deep Dive

The core of this work is the Tree of Validity (ToV), a novel framework designed to provide a unified and flexible approach to reasoning about and conducting validity experiments in computer security. To understand the ToV, it's essential to first establish the authors' precise definitions of experimental components and validity concepts.

The authors define validity as the broad field of science encompassing reproducibility and replicability. A validator is a team conducting experiments to gauge the validity of another's work. A validity experiment is any such experiment based on original authors' work. For experiments, a setting is composed of a problem (the scientific question, e.g., deepfake detection) and a domain (the environment, e.g., specific network, population studied, time, software/hardware systems). The process is the experimental methodology, comprising a method (approach to gather/manipulate data), data (collection of measurements), and analysis (quantifiable measure of performance).

The Tree of Validity (ToV) Framework

The Tree of Validity (ToV) is a perfect binary tree that visualizes every possible validity experiment. Each layer of the tree corresponds to a fundamental part of an experiment: Problem, Method, Data, Analysis, and Domain. At each layer, a binary decision is made: whether that part of the experiment remains Same or is made Different compared to the original authors' work.

  • Structure: The conceptual root of the tree represents the original experiment. Each branch point offers two paths: one where the experimental component is kept "Same" and one where it is "Different." A path from the root to a leaf node represents a specific validity experiment. For instance, following the "Same" path for every component would constitute a strict reproducibility experiment as defined by Claerbout.
  • Dynamic and Adaptable Layers: A key strength of the ToV is its dynamic nature. The predefined layers (Problem, Method, Data, Analysis, Domain) are not static and can be customized to fit specific experimental methodologies.
  • For example, the "Method" layer could be split into "Software Component" and "Hardware Component" for experiments involving complex systems.
  • Similarly, "Data" could be separated into "Training Data" and "Test Data" for machine learning contexts.
  • Layers can also be swapped or removed if they are irrelevant to a particular experiment. This flexibility ensures the ToV can accurately describe a wide array of research processes across different sub-fields.

Properties of Validity within ToV

The framework introduces three key properties for reasoning about validity studies:

  1. Potential: This refers to the inherent capacity for validity that a published paper possesses, determined by what experimental artifacts the original authors make available. The ToV visualizes this potential. If a component (e.g., method code or a massive dataset) is not available, then paths through the ToV that require that component to be "Same" become impossible.
  • The authors introduce the Seed of Validity (SoV) as a shorthand representation of what artifacts are available (or not available). An SoV effectively prunes the potential ToV, showing only the feasible execution paths. This concept is crucial for understanding the practical limits of reproducibility and replicability for any given paper.
  1. Execution: An Execution is the actual act of conducting a validity experiment, manifesting as a specific path from the root node to a leaf node in the potential ToV. Each execution describes a unique experimental methodology. This allows for direct, quantifiable comparison between different validity studies; two executions are different if their paths through the ToV vary in at least one experimental artifact.
  2. Conclusion: After executing a path, the Conclusion is the resulting output, which can be quantitative (e.g., accuracy differences, statistical distributions) or qualitative (e.g., figures, observations). The framework acknowledges that interpreting a "successful" conclusion is still an open problem, moving beyond arbitrary thresholds.

Integrating Threat Models into the ToV

A significant advantage of the ToV, especially for computer security research, is its ability to explicitly incorporate threat models. Threat models are fundamental to security papers, delineating adversary capabilities and defender assumptions, which profoundly shape methodologies and analyses.

  • Attacker and Defender Capabilities: The ToV allows for the inclusion of specific attacker and defender capabilities within its layers. This can encompass elements such as:
  • Attacker Knowledge: (e.g., white-box access to a model, knowledge of honeypot presence).
  • Attacker Access: (e.g., running code on target processor, network access).
  • Hardware/System State: (e.g., specific CPU architecture, system configurations).
  • Defense Type: (e.g., active vs. passive defense, specific intrusion detection system).
  • Modeling Adaptive Adversaries: The ToV excels at modeling scenarios involving adaptive adversaries. If a defense is proposed, an adaptive adversary might modify their attack strategy based on knowledge of the defense. Such a scenario can be represented as a new execution path through the ToV, where the "Attacker Capabilities" or "Method" layers are explicitly marked as "Different" to reflect the adversary's adaptation. This provides a clear framework for testing the robustness and bounds of a security hypothesis under evolving threats.
  • Communication Framework: By explicitly including threat model parameters, the ToV serves as a robust communication framework. It allows original authors to clearly state the assumptions under which their findings hold, and enables validators to precisely articulate how their replication or extension studies modify these assumptions. This clarity helps distinguish failures due to methodological flaws from those due to a divergence in threat modeling assumptions.

In summary, the ToV transforms the abstract concepts of reproducibility and replicability into a concrete, dynamic, and security-aware framework. It provides a common language for describing experimental methodologies, quantifying differences between studies, and rigorously evaluating research claims, particularly in the complex and adversarial landscape of computer security.

Demo / Proof of Concept

While this paper is a systematization of knowledge and a framework proposal rather than a demonstration of a new attack or defense, the authors effectively illustrate the utility of the Tree of Validity (ToV) through several compelling case studies. These case studies serve as "proofs of concept" for the framework's applicability to diverse security research scenarios.

Case Study 1: Spectre Attacks [31]

The paper first applies the ToV to the renowned Spectre attacks by Kocher et al. [31], which exploit speculative execution and branch prediction in modern processors to leak sensitive information.

  • Modeling Variants and Threat Models: The Spectre attacks have multiple variants (e.g., conditional branches, indirect branches) and practical demonstrations (native code, JavaScript, eBPF). The ToV effectively models these variations. For example, exploiting branch misdirection via JavaScript to violate a sandbox (leaking memory from another sandbox) and poisoning indirect branches in native code to violate process isolation (identifying a "gadget" memory location and training the Branch Target Buffer (BTB)) can be represented as distinct executions through the ToV.
  • Attacker and Defender Capabilities: The ToV explicitly captures the threat model. For Spectre, the attacker capabilities might include running code on the target processor (e.g., an AWS server), while defender capabilities are often limited to process isolation inherent in the processor design. A replication study could modify defender capabilities to test proposed mitigations.
  • Ethical Considerations: The authors highlight the ethical dilemma of publishing proof-of-concept exploits for attacks affecting billions of devices (Intel, AMD, ARM processors). The ToV provides a framework to discuss the potential for replication and the differences in experimental setup without requiring the public release of dangerous artifacts, thereby encouraging ethical transparency.
  • Statistical Nature of Attacks: The ToV can also account for the statistical nature of attacks, such as the error rates in memory cache recovery during Spectre attacks. A validator replicating the attack would consider hardware, attacker capabilities, and the goal of the attack, including measuring error rates across different hardware setups, all represented as distinct paths in the ToV.

Case Study 2: Entangled Watermarks for Machine Learning Defenses [29]

The ToV is then used to model a defense mechanism: entangled watermarks proposed by Jia et al. [29] to protect machine learning models from theft.

  • Modeling Defenses and Adaptive Adversaries: This case study demonstrates how a defense can be represented in a ToV. The original work evaluates entangled watermarks against several adaptive adversaries. Each adaptive adversary (e.g., a transfer-learning adversary aware of the watermark) is modeled as a different execution through the ToV, where the attacker's knowledge (e.g., awareness of watermark presence) and attack methodology (e.g., fine-tuning on new data to remove watermarks) are explicitly marked as "Different."
  • Artifact Availability and Potential ToV: Jia et al. [29] provide code artifacts for training and testing models (CNNs, RNNs, ResNet) with entangled watermarks using CleverHans [43] across various datasets (MNIST, Fashion MNIST, CIFAR-10, CIFAR-100, Google Speech Commands). However, the code for generating figures or running transfer learning experiments is not fully available. The ToV illustrates how this partial availability limits the "potential" for full reproducibility, guiding future validators on what aspects can be directly reproduced and what requires independent implementation (i.e., replicability).

Case Study 3: Replicating a USENIX Distinguished Artifact Award Winner [16]

Perhaps the most compelling demonstration involves a replication study of Bollinger et al.'s [16] work on automating GDPR cookie violation detection, which received a 2022 Distinguished Artifact Award. This case highlights a critical challenge: even with award-winning artifacts, true reproducibility can be unattainable.

  • Dynamic Domain Challenges: Bollinger et al.'s methodology involved collecting data from websites listed in the Tranco [46] ranking (May 5th, 2021) that used a consent management platform. The authors explicitly noted that reproducing these experiments from scratch was infeasible due to the dynamic state of the internet.
  • Replication Results: When the authors of this paper attempted to run Bollinger et al.'s experimental artifacts in January 2025, they found that 1,032 out of 6,940 (14.9%) of the original URLs no longer resolved. Furthermore, the number of detected advertising cookies was only 42% of the original declared cookies, while functional cookies showed a slight increase.
  • ToV's Role: The ToV clearly illustrates this scenario. While the code and processed training data were available, the "Domain" (specifically, the internet's state at a particular time) was inherently "Different." This means that strict reproducibility (all "Same" paths) was impossible from the outset. However, the ToV still allows for meaningful replicability studies, enabling a validator to re-collect data, train the model, and compare results, acknowledging the unavoidable differences in the domain. It underscores that authors can proactively address and communicate where experimental methodologies will vary due to confounding factors.

These case studies collectively demonstrate the ToV's flexibility and power in providing a consistent, detailed, and adaptable framework for understanding, communicating, and comparing validity studies across the diverse and complex landscape of computer security research.

Defensive Implications

The "SoK: Towards a Unified Approach to Applied Replicability for Computer Security" paper, while not proposing a direct defense against a specific vulnerability, has profound defensive implications for the security research community itself. By introducing the Tree of Validity (ToV), the authors provide a meta-scientific tool that significantly enhances the rigor, transparency, and practical utility of security research, ultimately strengthening the collective defense against evolving threats.

  1. Improved Communication and Contextualization of Research: The ToV offers a standardized, precise language for security researchers to describe their experimental methodologies and the underlying assumptions. This is particularly vital for the unique aspects of security, such as threat models (attacker capabilities, defender configurations). Authors can explicitly state the "Seed of Validity (SoV)" for their work, delineating what artifacts are available and where variations are expected. This clarity helps future researchers (validators) understand the context of the original findings, avoiding misinterpretations or disagreements that often arise from ambiguous terminology or implicit assumptions.
  2. Richer Evaluation of Defenses and Attacks: For defenders developing new security mechanisms, the ToV provides a framework to systematically evaluate the robustness of their solutions. By modeling adaptive adversaries as different "executions" through the ToV (i.e., varying attacker knowledge, methods, or domains), researchers can rigorously test the bounds of a proposed defense. This moves beyond simply demonstrating a defense works in a controlled environment to understanding its resilience against sophisticated, evolving threats. Similarly, for attack research, the ToV allows for a nuanced discussion of attack variants and their applicability under different threat models without necessarily requiring the public release of dangerous proof-of-concept exploits, addressing ethical concerns.
  3. Guidance for Artifact Evaluation Committees (AECs): The framework offers a powerful tool for AECs to go beyond merely checking if artifacts "run." By encouraging authors to provide a ToV or SoV, AECs can assess the potential for reproducibility and replicability, understand where limitations exist (e.g., due to sensitive data or dynamic environments), and evaluate the authors' foresight in identifying confounding factors. This can lead to more meaningful artifact awards and a better repository of research that clearly articulates its scope and limitations.
  4. Facilitating Meta-Analysis and Research Advancement: The ToV enables direct, quantifiable comparisons between different validity studies in a research area. This is crucial for building a cumulative body of evidence for or against a scientific hypothesis. Instead of disparate studies with unclear relationships, the ToV allows researchers to map multiple efforts onto a unified framework, demonstrating how new work either reproduces, replicates, or extends prior findings by varying specific experimental components. This structured approach fosters a more rapid and coherent advancement of security research by clearly delineating novelty and contribution.
  5. Promoting Ethical Transparency: In security research, sensitive data or dangerous exploits often cannot be fully disclosed. The ToV encourages ethical transparency by providing a structured way to describe claims and evidence, even when full disclosure is ethically or legally constrained. It allows researchers to document what was done and where variations might occur, fostering trust and enabling comparison without requiring the release of sensitive components.
  6. Addressing the "Reproducibility Crisis" in Security: Ultimately, the ToV provides a much-needed robust framework to address the reproducibility and replicability challenges specific to the security domain. By acknowledging that true reproducibility is often unattainable (as demonstrated by the case study of a USENIX award-winning artifact), the framework shifts focus towards structured replicability, which is more practical and impactful for testing hypotheses and building robust security solutions.

In essence, the ToV equips security researchers, practitioners, and evaluators with a sophisticated tool to analyze, communicate, and build upon research findings with greater precision and confidence, thereby strengthening the foundational science upon which effective security defenses are built.

Key Takeaways

  • Replicability is paramount for scientific validity: While reproducibility (same code, same data, same results) is important, it is often insufficient or unattainable in computer security. Replicability, which tests a hypothesis under new conditions or with different experimental components, is crucial for building robust scientific theories.
  • Existing validity frameworks are inadequate for security: Over 30 years of meta-science research reveals conflicting definitions and a lack of practical, unified frameworks, especially for the unique challenges of computer security research, such as dynamic systems, sensitive data, and adversarial threat models.
  • The Tree of Validity (ToV) provides a unified and flexible framework: The ToV is a novel binary tree structure that systematically maps experimental components (Problem, Method, Data, Analysis, Domain) to either "Same" or "Different" states, providing a clear, quantifiable way to classify and compare validity studies.
  • ToV explicitly incorporates security's unique elements: The framework effectively integrates threat models (attacker/defender capabilities) and can model complex scenarios like adaptive adversaries as distinct executions, offering a robust way to evaluate security research.
  • True reproducibility is often unattainable, but meaningful comparisons are possible: Case studies, including the replication of a USENIX Distinguished Artifact Award winner, demonstrate that even with available artifacts, environmental changes can make exact reproducibility impossible. However, the ToV enables structured replicability, facilitating valuable comparisons and insights despite these limitations.
  • ToV enhances communication and accelerates research: By providing a common language and a clear structure (including the Seed of Validity (SoV)), the ToV improves communication between authors and validators, helps contextualize research contributions, and fosters a more rigorous, transparent, and rapidly advancing security research community.

About the Speaker(s)

The authors of this paper, Daniel Olszewski, Tyler Tucker, Kevin R. B. Butler, and Patrick Traynor, are all affiliated with the University of Florida. Their collective work focuses on advancing the meta-science of computer security, specifically by addressing foundational issues in experimental validity, reproducibility, and replicability. Their research aims to provide practical frameworks and guidance to improve the scientific rigor and communication within the security research community. Kevin R. B. Butler and Patrick Traynor are noted as senior researchers in the field, guiding the systematic and framework-driven approach presented in this distinguished paper.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Finally, someone wrote down what we all complain about at the bar after AEC meetings. The Tree of Validity framework is genuinely useful—not because it's revolutionary, but because it forces precision where the community has been sloppy for decades. Minor quibble: the case studies do more work than the formalism.

Heather Calloway (CISO) — SOLID

This is a well-executed meta-science paper that gives security research leadership a framework to think about artifact evaluation, research claims, and what 'reproducible' actually means in our field. Worth your time if you're involved in research programs, academic partnerships, or evaluating vendor research claims.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)