Robust, Efficient, and Widely Available Greybox Fuzzing for COTS Binaries with System Call Pattern Feedback
Jifan Xiao
34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Software Security 3: Fuzzing
Overview
This distinguished paper from USENIX Security 2025 presents a groundbreaking meta-study examining the state of research transparency within the Usable Privacy and Security (UPS) community. Authored by Jan H. Klemmer and a team of researchers from CISPA Helmholtz Center for Information Security, Indiana University Indianapolis, and the University of Michigan, the work quantitatively analyzes reporting practices in a field that uniquely blends computer security, privacy, and human-centered research. The study addresses a critical gap, as prior meta-research on transparency was largely absent in the UPS domain, despite growing concerns about reproducibility and validity across scientific disciplines.
Read the paper · Download the PDF (PDF) · Slides
Paper abstract
Transparent research reporting is crucial to understanding and assessing research, its results and validity, and for fostering replication. While other research fields investigated reporting and transparency practices, similar meta-research is missing for the usable privacy and security (UPS) community, which combines security, privacy, and human research. To gain insights into current research transparency practices and their development in the UPS community, we analyzed 200 UPS publications from twelve venues (including USENIX Security, IEEE S&P, CCS, SOUPS, and CHI) from 2018 to 2023. Additionally, we evaluated those venues' 81 calls for papers (CfPs) and 20 calls for artifacts (CfAs). We find that most papers report on many of 52 analyzed transparency criteria, but none achieve full transparency. Moreover, we uncover several areas that need improvements: essential artifacts like questionnaires are frequently missing and hinder replication, some information is reported inconsistently, and dead links further reduce availability. Our regression analysis indicates that paper length and the number of studies described in a paper impact reporting transparency, while we observed no effect of publication year and artifact evaluation (AE). Finally, we provide recommendations for authors, venues, and PC chairs to improve research transparency practices and suggest transparency guidelines.

How Transparent is Usable Privacy and Security Research? A Meta-Study on Current Research Transparency Practices
Speakers: Jan H. Klemmer, Juliane Schmüser, Fabian Fischer, Jacques Suray, Jan-Ulrich Holtgrave, Simon Lenau (CISPA Helmholtz Center for Information Security); Byron M. Lowens (Indiana University Indianapolis); Florian Schaub (University of Michigan); Sascha Fahl (CISPA Helmholtz Center for Information Security)
Conference: USENIX Security
YouTube: Not applicable (peer-reviewed paper, no video recording)
Overview
This distinguished paper from USENIX Security 2025 presents a groundbreaking meta-study examining the state of research transparency within the Usable Privacy and Security (UPS) community. Authored by Jan H. Klemmer and a team of researchers from CISPA Helmholtz Center for Information Security, Indiana University Indianapolis, and the University of Michigan, the work quantitatively analyzes reporting practices in a field that uniquely blends computer security, privacy, and human-centered research. The study addresses a critical gap, as prior meta-research on transparency was largely absent in the UPS domain, despite growing concerns about reproducibility and validity across scientific disciplines.
The paper meticulously investigates 200 UPS publications from twelve prominent venues between 2018 and 2023, alongside an analysis of 81 calls for papers (CfPs) and 20 calls for artifacts (CfAs). By identifying and evaluating 52 distinct transparency criteria, the authors provide a comprehensive overview of current reporting practices, uncover significant shortcomings, and identify factors influencing transparency. This research is not intended as a critique of individual researchers but rather as a constructive contribution to fostering a more robust, transparent, and ultimately more scientific UPS community.
The findings reveal that while UPS researchers generally report on many transparency criteria, none of the analyzed papers achieve full transparency, with a substantial portion of essential information and artifacts frequently missing or inconsistently reported. Crucially, the study highlights issues such as the prevalence of dead links to online artifacts and the surprising lack of significant impact from publication year or artifact evaluation on overall transparency. The insights derived from this meta-study are invaluable for authors, reviewers, and program committees seeking to enhance the quality and reproducibility of UPS research, offering concrete recommendations and a foundational checklist for improved scientific practice.
Background
The scientific community has, in recent decades, grappled with a reproducibility crisis, particularly evident in fields like medicine and psychology. This crisis underscores the imperative for transparent research reporting, which serves as a fundamental prerequisite for understanding, assessing the validity of findings, and enabling the replication of studies by independent researchers. Organizations and communities have responded by advocating for better reporting standards, exemplified by initiatives like the Transparency and Openness Promotion Guidelines (TOP) and the FAIR principles for data.
Within computer science, however, the adoption and effectiveness of these transparency principles have been inconsistent. Prior studies have identified numerous reporting shortcomings across various sub-disciplines, including missing datasets, undisclosed algorithm specifications, and a general lack of critical methodological details. For the broader security and privacy (SP) community, earlier analyses of papers at top-tier conferences like CCS and IEEE S&P indicated a deficiency in clearly defined research objectives and limitations (Section 1). Efforts such as artifact evaluation (AE), intended to promote the availability of research artifacts, have not always translated into significant improvements in reproducibility, as observed in machine learning security papers. Such observations have led to accusations that security research sometimes lacks fundamental scientific principles, including transparency (Section 1).
The Usable Privacy and Security (UPS) subfield presents a unique context for examining transparency, as it inherently combines methodologies from computer science with human-centered research approaches. This interdisciplinary nature results in a diverse array of research artifacts and reporting challenges. While some prior work has touched upon specific aspects of UPS reporting, such as demographic data omissions or insufficient methodological details in risk representation studies, a comprehensive quantitative analysis of overall transparency practices in UPS was notably absent. Klemmer et al.'s own qualitative work previously revealed that UPS researchers value transparency and employ reporting practices based on implicit community standards, yet they also expressed challenges and a desire for further improvements (Section 2). This paper aims to complement those qualitative insights with the first quantitative, in-depth analysis of reporting transparency in the UPS community, establishing a much-needed empirical baseline. The problem persists due to various barriers, including the perceived high effort of transparent reporting, institutional pressures, technical difficulties in sharing data (especially with personal identifiable information or PII), and a general lack of motivation or incentives for researchers to invest in comprehensive artifact sharing and detailed methodological descriptions (Section 2).
Key Findings
The meta-study by Klemmer et al. provides a detailed quantitative assessment of transparency in UPS research, yielding several critical findings:
- Moderate Overall Transparency with Significant Gaps: The analysis of 200 papers against 52 transparency criteria revealed that, on average, 65.0% of applicable criteria were available per paper (median: 65.6%). While this indicates a general commitment to transparency, a substantial 30.9% of information and materials were found to be entirely unavailable, and 5.2% were only partially available. Crucially, no paper achieved full transparency across all applicable criteria (Section 6).
- Frequent Absence of Essential Artifacts: The study uncovered significant shortcomings in the provision of crucial research artifacts. For instance, consent forms were provided in only 11.3% of studies, sampling materials in 8.1%, and raw study data in a mere 15.0% (often with justifications related to PII concerns). Even for qualitative analysis, only 51.3% of codebooks were fully available. These omissions severely hinder the ability of other researchers to understand, validate, or replicate findings (Section 6.1.2, Table 3).
- Inconsistent Reporting Practices: While some criteria like sample size (98.8% available) and research questions (95.6% available) were consistently reported, others exhibited high inconsistency. Criteria such as the provision of experiment materials, qualitative analysis codebooks, sampling success rates, and the discussion of inter-rater reliability (IRR) in qualitative analysis showed high Shannon Entropy values, indicating a lack of uniform reporting standards across the community (Section 6.1.1, Table 3).
- Widespread Problem of Dead Links: The long-term availability of online artifacts is a significant concern. Of the 248 instances where papers pointed to online resources, 14.5% were found to be unavailable. This included 11.3% due to dead links and 3.2% where promised materials were absent from the provided resources. Self-hosted materials (e.g., on personal websites or GitHub repositories not intended for archival) were particularly vulnerable, with 18.9% of links to these resources being dead. Even publisher-hosted materials, while generally more reliable, still had a 7.5% unavailability rate (Section 6.2).
- Impact of Paper Length and Number of Methods: A linear mixed model regression analysis revealed two significant factors influencing transparency. Paper length (normalized) showed a significant positive correlation with transparency, suggesting that longer papers tend to be more transparent. Conversely, the number of methods employed in a single paper exhibited a significant negative association, indicating that papers describing multiple studies or methods often report less transparently, likely due to page limits (Section 6.4, Table 4).
- Limited Impact of Publication Year and Artifact Evaluation: Contrary to some expectations, the study found no significant effect of publication year on transparency. While a slight improving trend was observed, the six-year range (2018-2023) might be too short to detect subtle changes. Similarly, the presence of an Artifact Evaluation (AE) badge showed no significant association with higher transparency. This is attributed to AE being largely optional and not yet widespread (only 4.0% of the sample had badges), and often not tailored to the non-technical artifacts common in UPS research (Section 6.4).
- Replication Hindrance: A critical implication of the findings is that 56 out of 200 papers (28.0%) lacked essential information or artifacts necessary for replication. This highlights a tangible contribution to a potential "replication crisis" within UPS, emphasizing the urgent need for improved transparency practices to ensure scientific rigor and replicability (Section 7.3).
Technical Deep Dive
The core of this meta-study lies in its rigorous methodology for assessing research transparency in UPS. The authors conducted a comprehensive analysis involving both calls for papers (CfPs) and a systematic literature review of published papers.
The study began by defining the scope of Usable Privacy and Security (UPS) papers, leveraging Klemmer et al.'s prior definition: publications that (1) cover security and/or privacy topics, and (2) involve human subjects research. This definition guided the identification of 957 UPS papers from a total of 13,806 publications across twelve major venues (including USENIX Security, IEEE S&P, CCS, SOUPS, CHI, NDSS, EuroS&P, PETS/PoPETs, CSCW, ICSE, WWW, and EuroUSEC) published between 2018 and 2023 (Section 5). These venues were chosen for their prominence and frequent publication of UPS research.
Calls for Papers (CfP) Analysis
To understand external requirements and potential influencing factors on transparency, the authors analyzed 81 CfPs and 20 calls for artifacts (CfAs) from the 12 venues between 2018 and 2024 (Section 4). Key aspects examined included:
- Transparency & Reproducibility Requirements: In 2024, 9 out of 12 venues included some form of transparency or reproducibility requirements in their CfPs, a notable increase since 2018 when only ICSE explicitly requested it. This indicates a growing, albeit often indirect, prioritization of transparency within the communities.
- Replication Studies: Explicit calls for replication studies remained rare, primarily limited to dedicated UPS venues like SOUPS and EuroUSEC, with CHI and USENIX Security introducing them only in 2024/2025. This suggests that replication efforts are still undervalued.
- Page Limits: Most venues impose page limits, typically 12-13 pages for the main body and often unlimited appendices for submissions, but more restricted limits for camera-ready versions. These limits are inconsistent across venues and raise questions about their impact on detailed reporting.
- Artifact Evaluation (AE): By 2024, 6 of 12 venues offered AE, usually on an optional basis. AE primarily targets technical artifacts (e.g., code, scripts) and rarely encourages non-technical artifacts typical for UPS (e.g., survey instruments, interview guides). Moreover, AE badges are not always officially published with the final paper (e.g., CCS 2023), and reviewers are generally not required to consider appendices or supplementary materials during the main paper review process.
Systematic Literature Analysis – Methodology
For the core literature analysis, a stratified random sample of 200 UPS papers (20.9% of the 957 identified UPS papers) was selected to ensure representation across venues and years (Section 5.2). The manual content analysis of each paper was extensive, requiring 30-120 minutes per paper.
The cornerstone of the analysis was a comprehensive list of 52 transparency criteria (Table 3), compiled from an extensive literature review of research reporting and transparency guidelines across various scientific fields. These criteria were categorized into:
- Ethical Considerations: e.g., IRB approval, consent procedures, anonymization.
- Study Design: e.g., research questions, method choice justification, limitations, study protocol.
- Sample & Recruitment: e.g., sampled population, demographics, sampling procedures, sample size justification.
- Instruments & Artifacts: e.g., experimental materials, questionnaires, interview guides, raw data, software/hardware specifications.
- Data Analysis (Qualitative & Quantitative): e.g., data preprocessing, analysis methods, inter-rater reliability, statistical assumptions.
- Result Reporting (Qualitative & Quantitative): e.g., codebooks, empirical evidence, descriptive statistics, effect sizes, confidence intervals.
- Miscellaneous: e.g., author contact information, conflicts of interest.
Each of the 200 papers was independently analyzed by two researchers. For each applicable criterion, they assigned an availability rating (available, partially available, unavailable) and a reporting location (main paper, appendix, publisher materials, external with DOI, external without DOI, upon request). An inter-rater reliability (IRR) check on 100 random publications yielded a Fleiss’ κ of 0.89, indicating almost perfect agreement (Section 5.3.2).
Transparency Score (TS) and Regression Analysis
To quantify overall transparency, the authors developed a Transparency Score (TS) for each paper, ranging from 0 (all applicable criteria unavailable) to 1 (all applicable criteria available). The TS was calculated by mapping availability ratings to numerical values (e.g., 1 for available, 0.5 for partially available, 0 for unavailable) and averaging these across all applicable criteria per paper. For papers with multiple studies, the highest availability across sub-studies for each criterion was aggregated (Section 6.3). The average TS across the sample was 0.677 ± 0.103 (median 0.681), with a range from 0.281 to 0.885.
A linear mixed model was then employed to investigate factors influencing the TS (Section 6.4). The dependent variable was the TS, and independent variables included:
- Venue (baseline: SOUPS)
- Year
- Main Method (baseline: Interview)
- Number of Methods in a publication
- Paper Length (normalized)
- Whether the paper had an AE badge
The model accounted for dependencies among papers by the same last author using random intercepts. Key findings from the regression analysis (Table 4) include:
- Paper length showed a significant positive relationship with TS (coefficient 0.061±0.054), indicating that longer papers tend to report more transparently.
- The number of methods used in a paper had a significant negative association with TS (coefficient -0.024±0.023), suggesting that papers incorporating multiple methodologies struggle to maintain transparency, likely due to space constraints.
- Most venues did not significantly differ from SOUPS in TS, except for ICSE and EuroS&P, which showed slightly higher transparency. This suggests that venue-specific policies might have limited, or at least inconsistent, impact.
- The model found no significant effect for publication year, main study type, or the presence of an AE badge on the TS. The lack of AE impact was attributed to its low prevalence (only 8 papers, 4.0%, in the sample had badges) and voluntary nature.
Reporting Consistency and Availability Details
The paper further delves into the specifics of each transparency criterion category (Section 6.1.2):
- Ethical Considerations: While IRB approval (77.5%) and consent procedures (62.6%) were commonly reported, consent forms were rarely provided (11.3%), and ethical implications discussed in only 41.8% of studies.
- Study Design: Research questions (95.6%) and study protocols (97.1%) were well-reported, but limitations were omitted in 10.1% of cases, and study piloting information was missing in 53.0%. Pre-registration and positionality statements were rarely used.
- Sample & Recruitment: Sample size (98.8%) was almost universally reported, along with sampled population (91.8%), but sample size justification was low (30.7%), and recruitment materials were rarely available (8.1%).
- Instruments & Artifacts: Availability varied significantly by type: survey instruments (73.6%) and interview guides (66.3%) were more often provided than experimental materials (47.8%). Crucially, raw study data was shared in only 15.0% of cases, and analysis software/scripts in 37.2%, with 82.0% of studies not making analysis software available.
- Data Analysis: Qualitative analysis methods were generally described (86.7%), but codebooks were fully available in only 51.3% of cases. Quantitative methods were also well-described (94.4%), but statistical assumptions (28.4% unavailable), effect sizes (31.7% missing), and confidence intervals (51.9% missing) were frequently omitted.
- Long-Term Availability: The analysis of reporting locations highlighted that most information is in the main paper (89.5%) or appendix (6.5%). Online artifacts (0.6% publisher, 0.9% external with DOI, 2.3% external without DOI) are used less frequently but are crucial for replication. The high rate of dead links for self-hosted materials (18.9% for external without DOI) underscores the fragility of such practices (Section 6.2).
In summary, the technical deep dive reveals a nuanced picture: while awareness of transparency is present, inconsistent practices, a lack of explicit guidelines, and infrastructure challenges (like reliable artifact hosting) collectively contribute to significant transparency gaps in UPS research.
Demo / Proof of Concept
As this work is a peer-reviewed academic paper rather than a conference talk with a live presentation, it does not feature a traditional "demo" or a "proof of concept" in the sense of a software demonstration or an exploit. Instead, the paper's core contribution and its "proof of concept" are embodied in the systematic meta-study itself.
The authors' meticulous methodology, including the development of 52 transparency criteria, the comprehensive analysis of 200 UPS publications, and the subsequent quantitative assessment through the Transparency Score (TS) and regression analysis, serves as the empirical validation of their claims. The resulting UPS Transparency Dataset, which includes the analysis results and the list of all identified UPS publications, acts as a tangible artifact of their work, demonstrating the feasibility and utility of such a meta-scientific approach. This dataset supports Open Science principles and provides a baseline for future research, embodying the very transparency principles the paper advocates.
Defensive Implications
The findings of this meta-study carry significant defensive implications for the entire UPS research community, encompassing authors, reviewers, program committee (PC) chairs, and venues. The "defenders" in this context are those committed to upholding the scientific rigor, reproducibility, and trustworthiness of UPS research. The paper provides ten concrete recommendations (R1-R10) to address the identified transparency gaps:
- R1: Develop Transparency Guidelines for UPS. The current reliance on implicit community standards leads to inconsistent reporting (Section 7.1.1). The authors recommend establishing explicit, field-specific guidelines for authors and reviewers, suggesting their 52 transparency criteria (Table 3) as a foundational checklist. This would provide clear expectations and foster uniformity.
- R2: Make Existing Materials Available. Many materials already prepared for a study (e.g., consent forms, recruitment materials, experimental materials, questionnaires, interview guides) are often not provided in publications. Authors should make these readily available in appendices or online, requiring minimal additional effort beyond potential anonymization (Section 7.1.2).
- R3: Practice Explicit Reporting. Authors should explicitly state if a common methodological step was skipped or not applicable, rather than merely omitting the information. For instance, if participants were not compensated, this should be briefly mentioned. This "explicit reporting" prevents ambiguity and enhances clarity for readers (Section 7.1.2).
- R4: Develop Author Support for Sharing Data and Scripts. Providing analysis software and datasets significantly boosts transparency but is currently rare due to effort and PII concerns. The community should develop best practices, tool support, and explore solutions like controlled-access archives (similar to social sciences) to facilitate ethical data sharing (Section 7.1.3).
- R5: Adapt AE for UPS. Given that Artifact Evaluation (AE) currently has a negligible impact on transparency and is often geared towards technical artifacts, venues should reconsider or adapt AE processes to better accommodate the diverse, often non-technical, artifacts prevalent in UPS research (Section 7.2.1).
- R6: Review Appendices and Artifacts. PC chairs should require reviewers to consider appendices and supplementary materials, including artifacts like interview guides and survey instruments, as an integral part of the peer-review process. This ensures a thorough assessment of the methodology and potential biases (Section 7.2.1).
- R7: Lift Appendix Page Limits. The observed positive correlation between paper length and transparency suggests that page limits, particularly for appendices, can hinder detailed reporting. PC chairs should lift or generously increase appendix page limits, or adopt a model where paper length is assessed relative to its contribution, as some HCI venues do (Section 7.2.2).
- R8: Do not Decrease Page Limits for CR. Decreasing page limits between submission and camera-ready versions incentivizes authors to cut important details, often from appendices, after acceptance. Venues should maintain consistent page limits for both submitted and final versions to prevent this detrimental practice (Section 7.2.2).
- R9: Offer Publisher Artifact Hosting. To ensure the long-term accessibility and integrity of research artifacts, publishers should offer robust options for hosting artifacts alongside publications. PC chairs and venues should actively leverage these options to prevent fragmentation of papers and their supplementary materials (Section 7.2.3).
- R10: Use Long-Term Artifact Archival. In the absence of publisher-provided hosting, authors must adopt reliable long-term archival solutions. This means utilizing DOI-capable archival platforms such as OSF or Zenodo, assigning Digital Object Identifiers (DOIs) to artifacts, and referencing them consistently. Authors should explicitly avoid self-hosting on personal websites or using platforms not intended for long-term archival, like GitHub, for published artifacts, as these are prone to dead links (Section 7.2.3).
The stark finding that 28.0% of the analyzed papers lacked essential information for replication serves as a powerful call to action. By implementing these recommendations, the UPS community can proactively address the root causes of transparency issues, strengthen the scientific foundation of the field, and ultimately contribute to more trustworthy and impactful research outcomes.
Key Takeaways
- Moderate Transparency, Significant Gaps: Usable Privacy and Security (UPS) research exhibits moderate transparency (average Transparency Score of 0.677), but none of the 200 analyzed papers achieved full transparency, with 30.9% of applicable criteria missing.
- Critical Artifacts Frequently Absent: Essential materials for understanding and replicating studies, such as consent forms (11.3% available), raw study data (15.0% available), and analysis software/scripts (37.2% available), are consistently missing or inconsistently reported.
- Inconsistent Reporting Demands Guidelines: The wide variability in reporting practices for certain criteria, like experimental materials and qualitative codebooks, highlights the urgent need for explicit, community-wide transparency guidelines (e.g., based on the 52 criteria identified in this study).
- Online Artifacts Face Availability Crisis: A significant portion (14.5%) of online artifacts pointed to by papers are unavailable, primarily due to dead links (11.3%). Self-hosted materials, particularly those without Digital Object Identifiers (DOIs), are highly susceptible to this long-term availability problem.
- Paper Length and Study Complexity Impact Transparency: Longer papers tend to be more transparent, while papers employing multiple research methods show a significant decrease in transparency, likely due to existing page limits that force omission of details.
- Venues and PC Chairs are Key Enablers: Program committees and conference venues play a critical role in fostering transparency through policies on page limits (especially for appendices), mandatory review of supplementary materials, and providing robust, long-term options for publisher-hosted artifact archiving.
About the Speaker(s)
The research presented in this paper was a collaborative effort by a team of distinguished academics and researchers. The primary authors include Jan H. Klemmer, Juliane Schmüser, Fabian Fischer, Jacques Suray, Jan-Ulrich Holtgrave, and Simon Lenau, all affiliated with the CISPA Helmholtz Center for Information Security. They are joined by Byron M. Lowens from Indiana University Indianapolis, Florian Schaub from the University of Michigan, and Sascha Fahl also from the CISPA Helmholtz Center for Information Security. Their collective expertise spans computer security, privacy, and human-computer interaction, bringing a multidisciplinary perspective to the critical examination of research transparency in the Usable Privacy and Security field.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
A competent meta-study that confirms what most honest researchers already suspect: we're bad at sharing our work. The 52-criterion framework is useful, the dead-link analysis is damning, but the recommendations read like a committee wrote them. Fills a gap in the literature without setting anything on fire.
Heather Calloway (CISO) — WEAK
Useful baseline data on how UPS researchers report their work. The findings are real but narrow—this is science-of-science inside a niche academic subfield, not something that changes how security programs operate or how boards think about risk.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)