Vulnerability of Text-Matching in ML/AI Conference Reviewer Assignments to Collusions
Jhih-Yi (Janet) Hsieh
34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Social Issues and Security
Overview
The integrity of the peer-review process is a cornerstone of scientific advancement, especially in rapidly evolving fields like Artificial Intelligence and Machine Learning. As these conferences scale to unprecedented sizes, the reliance on automated systems for reviewer assignments becomes essential. This talk, presented by Jhih-Yi (Janet) Hsieh, delves into a critical vulnerability within these automated systems: the susceptibility of text-matching algorithms to collusive manipulation. Historically, while reviewer bidding mechanisms were recognized as exploitable, the underlying text-matching components, which assess the thematic similarity between submissions and reviewers' past work, were often implicitly assumed to be robust against such attacks.

Key moments
- 0:00 Introduction: Vulnerability of text matching in peer review
- 2:45 Understanding collusion rings in conference peer review
- 3:50 Key research question: Is text matching safe from manipulation?
- 5:15 How text similarity is calculated: The Spectre model
- 6:30 Attack method 1: Adversarial archive curation explained
- 7:45 Attack method 2: Adversarial abstract modification details
Vulnerability of Text-Matching in ML/AI Conference Reviewer Assignments to Collusions
Speakers: Jhih-Yi (Janet) Hsieh
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=FtN3OLD0gMM
Overview
The integrity of the peer-review process is a cornerstone of scientific advancement, especially in rapidly evolving fields like Artificial Intelligence and Machine Learning. As these conferences scale to unprecedented sizes, the reliance on automated systems for reviewer assignments becomes essential. This talk, presented by Jhih-Yi (Janet) Hsieh, delves into a critical vulnerability within these automated systems: the susceptibility of text-matching algorithms to collusive manipulation. Historically, while reviewer bidding mechanisms were recognized as exploitable, the underlying text-matching components, which assess the thematic similarity between submissions and reviewers' past work, were often implicitly assumed to be robust against such attacks.
Hsieh, in collaboration with professors DT Rakenathan and Nihar Shaw, challenges this assumption, demonstrating that text-matching similarity scores can be effectively manipulated by colluding authors and reviewers. The research uncovers specific attack vectors that allow colluders to significantly increase their chances of being assigned to review each other's papers, all while maintaining a degree of plausible deniability. The findings are particularly salient given the massive growth of conferences like NeurIPS and AAAI, where submissions are projected to exceed 25,000 by 2025, making robust automated assignment systems indispensable.
The significance of this work extends beyond the immediate technical vulnerability. It highlights a systemic risk to the fairness and impartiality of scientific peer review, potentially undermining the quality of accepted research and the careers of honest academics. By exposing these vulnerabilities and proposing concrete defense mechanisms, the research not only fortifies the review process in AI/ML but also offers insights applicable to other computer science disciplines grappling with similar challenges in automated assignments.
Background
▶ Watch: Introduction: Vulnerability of text matching in peer review (0:00)
The publication process at computer science conferences typically involves authors submitting papers, followed by a peer review phase where three to six experts in similar topics thoroughly evaluate each submission. Based on these reviews, papers are either accepted or rejected for publication and presentation. With the exponential growth of fields like AI and ML, exemplified by NeurIPS submissions reaching over 13,000 in 2023 and projected to exceed 25,000 by 2025, manual reviewer assignment has become impractical. This surge has necessitated the widespread adoption of automated reviewer assignments.
In automated assignment systems, a similarity score is first calculated between each submitted paper and every potential reviewer. These scores are then fed into a linear program designed to maximize the overall similarity between papers and their assigned reviewers across the entire conference. This optimization aims to ensure that papers are reviewed by those most qualified and interested. The similarity score itself is typically derived from two main components: reviewer bidding, where reviewers explicitly indicate their interest in specific papers, and text matching, which assesses the thematic overlap between a paper's content (e.g., title and abstract) and a reviewer's past publications.
However, the integrity of this process is threatened by collusion rings – groups of researchers who conspire to manipulate the assignment process to review each other's papers. This is not a new problem; it has been documented across various computer science fields, with early discoveries involving a tenured professor orchestrating such a ring in computer architecture. Colluders often exchange papers before submission, then strategically bid on or maneuver to be assigned to review their co-conspirators' papers. A key objective for a colluding reviewer is to achieve one of the highest similarity scores to the target submission among all potential reviewers, thereby increasing their likelihood of assignment.
While reviewer bidding has been extensively studied and is known to be vulnerable, leading some venues like CVPR and ACL Rolling Review (ARR) to ban it altogether, the vulnerability of text matching has received significantly less attention. The implicit or explicit assumption among researchers and practitioners has largely been that text matching, being an objective algorithmic measure, is inherently safe from manipulation. This talk directly addresses this critical gap, investigating the research question: "Is text matching safe from manipulation?"
Key Findings
▶ Watch: Key research question: Is text matching safe from manipulation? (3:50)
The core finding of this research is a definitive "no": text matching, a fundamental component of automated reviewer assignment systems, is indeed vulnerable to sophisticated manipulation by colluding parties. The study demonstrates that attackers can significantly increase their similarity ranking to a target submission, even from a very low starting point, thereby subverting the peer-review process.
Specifically, the key findings include:
- Vulnerability of Text Matching: Text matching algorithms, particularly those relying on embedding models like Spectre and cosine similarity, can be effectively manipulated by colluding authors and reviewers.
- Two-Pronged Attack Strategy: The researchers identify and demonstrate two primary attack vectors:
- Adversarial Archive Curation: Colluding reviewers can selectively curate their list of past publications used for similarity calculation, retaining only those most thematically aligned with the target submission.
- Adversarial Abstract Modification: Colluding authors can subtly modify their paper's abstract by introducing sentences or keywords related to the colluding reviewer's selected past work, thereby boosting text-matching scores.
- High Attack Success Rates: Even when a colluding reviewer initially ranks as low as the 1001st most similar reviewer, the proposed attack can elevate them to the top position (most similar) almost half the time, demonstrating remarkable effectiveness.
- Plausible Deniability: A crucial aspect of the attack is its capacity for plausible deniability. The modifications to abstracts, while effective, can appear as benign errors or stylistic choices, making post-hoc detection challenging and attribution difficult. This was validated through human subject experiments.
- Effective Defense Mechanisms: The research also proposes and evaluates several effective defense mechanisms, including limiting reviewer archive curation and employing more robust similarity aggregation methods (e.g., average or 75th percentile pooling instead of max pooling).
- Practical Impact: The identified safeguards have been adopted by top-tier ML conferences and are available on OpenReview, a widely used conference management platform, indicating the immediate practical relevance and implementation potential of this work.
These findings collectively dismantle the prior assumption of text matching's immunity to manipulation, providing a stark warning about the integrity of automated peer review and offering actionable strategies for its fortification.
Technical Deep Dive
▶ Watch: How text similarity is calculated: The Spectre model (5:15)
The attack on text-matching-based reviewer assignments hinges on understanding and manipulating the underlying mechanisms of similarity calculation. Colluders, typically an author (Alice) and a reviewer (Bob), have two primary goals: first, for Bob to achieve a high similarity ranking to Alice's paper, and second, for Alice to avoid suspicion from non-colluding reviewers.
The most popular text similarity model in this context is the Spectre model. This model generates dense vector embeddings for each paper based on its title and abstract. Spectre is specifically trained on scientific papers, and its embeddings are known to cluster semantically similar papers together. The closer two paper embeddings are in this vector space, the more similar their topics. The similarity between any two papers is quantified by the cosine similarity of their Spectre embeddings.
The overall paper-reviewer similarity score (SPR) is calculated in a two-step process:
- Paper-to-Past-Paper Similarity: For a given paper submission and a reviewer, the cosine similarity of Spectre embeddings is calculated between the submission and each of the reviewer's past published papers.
- Aggregation: These individual similarities are then aggregated to produce a single SPR score. Common aggregation choices include taking the average or the maximum of these similarities across all the reviewer's past papers. The choice of aggregation method significantly impacts vulnerability.
The attack proposed by the researchers involves two main steps:
1. Adversarial Archive Curation
This step targets the reviewer's side of the collusion. Many conferences allow reviewers to curate the list of their past papers used for text matching. The benign intent behind this feature is to enable honest reviewers to specify their current areas of expertise or interest, ensuring they receive relevant papers. However, colluding reviewers can abuse this.
In an illustrative example, if Alice submits a paper on a new object detection algorithm, and Bob (the colluding reviewer) has published on diverse topics including self-driving cars, security, and computational biology, Bob would perform adversarial archive curation. He would keep only his most similar past paper to Alice's submission – in this case, his paper on self-driving cars. If the Spectre cosine similarity between Alice's paper and Bob's self-driving car paper is 0.88, and Bob's other papers are less related, under an average of similarities aggregation method, his overall SPR score to Alice's paper would increase significantly (e.g., from 0.7 to 0.88 if only one paper is kept). This selective pruning of the reviewer's publication list directly manipulates the input to the aggregation function, boosting the SPR.
2. Adversarial Abstract Modification
This step targets the author's side of the collusion. Alice, the colluding author, will modify her paper's abstract to increase its similarity to Bob's selected past paper (e.g., the self-driving car paper). This is done by subtly incorporating information or keywords related to Bob's work. For instance, Alice might add a sentence like, "Inspired by growing interest around commercial self-driving cars, our work improves upon existing object detection methods," even if her paper's core contribution isn't directly in self-driving cars.
Additionally, Alice can insert specific keywords that are semantically close to Bob's work but might have a different meaning in her paper's context. An example given is mentioning "shift gears," a phrase related to cars, even if its usage in Alice's paper is metaphorical or tangential. The researchers defined an attack budget for this step: one added sentence about the reviewer's work and up to ten inserted keywords. Crucially, they imposed consistency and coherence constraints on these adversarial modifications. These constraints are vital for maintaining plausible deniability, ensuring that the modified abstract appears legitimate enough to avoid suspicion from non-colluding reviewers.
Attack Tuning and Effectiveness
Attackers can further refine their strategy by tuning their attack on previous year's conference data. The study found a strong correlation in attack success between NeurIPS 2022 and 2023 data. This means colluders can use the publicly available reviewer pool data from a previous year as a reliable surrogate to test and optimize their attack, increasing their chances of success in the current year even without knowing the precise current reviewer pool.
The experimental results vividly illustrate the attack's effectiveness. The success rate was measured by how often a colluding reviewer, after the attack, became one of the top one, top three, or top five most similar reviewers to the target submission. Strikingly, even when a reviewer initially ranked as low as the 1001st most similar, the attack enabled them to become the most similar reviewer (top one) almost half the time. This demonstrates the profound impact these subtle manipulations can have on the automated assignment process.
The plausible deniability aspect is further supported by a human subject experiment (details available in the full paper), where human reviewers found it difficult to consistently identify manipulated abstracts, reinforcing the notion that such abnormalities could be attributed to benign causes like negligence or language barriers.
Demo / Proof of Concept
▶ Watch: Attack method 1: Adversarial archive curation explained (6:30)
While the talk did not feature a live, interactive demonstration of the attack in action during the presentation, the research thoroughly demonstrated the feasibility and effectiveness of the proposed collusion methods through systematic experimentation. The "proof of concept" was established by conducting these experiments on real-world conference data and evaluating the success rates of the adversarial strategies.
The researchers used a concrete example to illustrate the attack flow, showing how Alice (author) and Bob (reviewer) would coordinate their actions. Bob's adversarial archive curation was explained with a clear example where retaining only a "self-driving cars" publication significantly boosted his similarity score to Alice's "object detection" paper. Alice's subsequent adversarial abstract modification was similarly detailed, showing how adding a single sentence about "commercial self-driving cars" and inserting keywords could further increase the similarity.
The efficacy of these combined attacks was then quantitatively presented through experimental results. The core demonstration of impact was the plot showing the dramatic increase in the colluding reviewer's rank. For instance, the finding that a reviewer starting at the 1001st similarity rank could reach the top 1 position almost 50% of the time serves as a powerful empirical proof of concept for the attack's viability and success.
Furthermore, the researchers acknowledged the challenge of plausible deniability for these attacks. To address this, they conducted a human subject experiment, the details of which are provided in their full paper. This experiment aimed to demonstrate that human reviewers might not easily detect the subtle manipulations in abstracts, thus validating the claim that the attacks are difficult to flag post-hoc. The availability of their artifacts on Zenodo (linked in the presentation) further underscores the reproducibility and verifiability of their experimental setup and results, allowing other researchers to independently validate their findings.
Defensive Implications
▶ Watch: Attack method 2: Adversarial abstract modification details (7:45)
Given the demonstrated effectiveness and plausible deniability of text-matching collusion attacks, the research strongly emphasizes a shift in focus towards prevention rather than post-hoc detection. Since abnormalities in abstracts can be attributed to benign causes (e.g., negligence, non-native English speakers), detecting collusion after the fact is inherently challenging and prone to false positives.
The researchers propose and evaluate several concrete defense mechanisms:
- Limit Reviewer Archive Curation: This defense directly counters the "adversarial archive curation" attack. By requiring reviewers to include a larger or more comprehensive set of their past papers for text matching, the impact of selectively pruning the archive is significantly reduced. The experimental results clearly show that as more papers are required to be kept in a reviewer's archive, the attack success rates drop considerably. This makes it harder for a colluding reviewer to artificially inflate their similarity to a specific target paper by isolating a single highly similar publication.
- Use More Robust Aggregation Methods: The method used to aggregate individual paper-to-past-paper similarities into a single paper-reviewer similarity score is critical.
- Average Pooling instead of Max Pooling: The study found that using average pooling over all past paper similarities, rather than max pooling, substantially reduces attack success rates. Max pooling is highly susceptible to the adversarial archive curation attack because it only needs one very similar paper to achieve a high score, which can be manipulated. Average pooling, by considering all relevant papers, is more resistant to such targeted manipulation.
- 75th Percentile Pooling: The researchers also suggest that 75th percentile pooling could strike a beneficial balance. It would be more robust than max pooling while potentially retaining more signal about a reviewer's expertise than a simple average, especially if a reviewer has a few highly relevant papers alongside many less relevant ones.
Beyond these specific technical mitigations, the authors offer broader recommendations for future work and system design:
- Reviewer Awareness and Careful Abstract Review: While post-hoc detection is difficult, increasing awareness among reviewers about the potential for abstract manipulation could encourage more careful scrutiny during the review process. This might involve flagging abstracts that seem unusually broad in their motivation or incorporate keywords in a forced manner.
- Randomness as a Broader Mitigation Strategy: Incorporating elements of randomness into reviewer assignment algorithms could make collusive manipulation less predictable and therefore less effective. This could involve randomizing a portion of assignments or using stochastic elements in the similarity calculation.
- Robustness of Similarity Scores: Future development of text similarity models and scoring mechanisms should explicitly consider robustness against adversarial manipulation. This might involve training models that are less sensitive to minor keyword insertions or contextual shifts.
The practical impact of this research is significant. These safeguards have already been adopted by several top-tier ML conferences and are available on OpenReview, the primary conference management platform for AI and ML. This means that conference organizers can readily implement these defenses to enhance the integrity of their peer-review processes, moving towards a more secure and fair academic publishing ecosystem.
Key Takeaways
- Text Matching is Vulnerable: Contrary to prior assumptions, text matching in automated reviewer assignment systems is highly susceptible to manipulation by colluding authors and reviewers.
- Two-Pronged Attack: Colluders leverage adversarial archive curation (reviewers selectively choosing past papers) and adversarial abstract modification (authors injecting keywords/sentences into abstracts) to boost similarity scores.
- High Effectiveness with Plausible Deniability: These attacks can dramatically increase a colluding reviewer's assignment probability, even from very low initial rankings, while maintaining plausible deniability to avoid suspicion.
- Prevention Over Detection: Due to the plausible deniability of manipulated abstracts, focusing on preventative measures is more effective than attempting post-hoc detection.
- Effective Defenses Exist: Key defenses include limiting reviewer archive curation (requiring more papers) and using more robust aggregation methods like average pooling or 75th percentile pooling instead of max pooling.
- Practical Implementation: The proposed safeguards have been adopted by leading ML conferences and are available on OpenReview, providing immediate, actionable solutions for improving the integrity of peer review.
About the Speaker(s)
Jhih-Yi (Janet) Hsieh is the primary presenter of this work, highlighting her expertise in the area of security and integrity within academic processes. The research presented is a collaborative effort, undertaken with professors DT Rakenathan and Nihar Shaw. Their collective work focuses on understanding and mitigating vulnerabilities in critical systems, particularly within the context of scientific peer review and the challenges posed by the rapid growth of AI and ML conferences.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Solid academic security research that correctly identifies and validates a real vulnerability in automated peer-review assignment systems. The attack is well-constructed and the defenses have seen real-world adoption, but this is a niche problem with a narrow threat model that won't move the needle for most security conference attendees.
Heather Calloway (CISO) — WEAK
Solid academic security research that exposes a real vulnerability in AI/ML peer review systems — but it never escapes the academic context it's critiquing. The finding is credible, the defenses are concrete, and the practical adoption by OpenReview is a genuine win, but the talk has no relevance to security governance, enterprise operations, or institutional risk management.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)