Double Face: Leveraging User Intelligence to Characterize and Recognize AI-synthesized Faces
Matthew Joslin (University of Texas at Dallas)
33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24
Overview
In an era witnessing an alarming surge of AI-generated images, particularly deepfakes, posing significant threats to information integrity and social trust, Matthew Joslin from the University of Texas at Dallas presented groundbreaking research titled "Double Face: Leveraging User Intelligence to Characterize and Recognize AI-synthesized Faces" at USENIX Security '24. This work, co-authored with Sien Wang and Dr. Shuangge, introduces a novel methodology that harnesses crowdsourcing intelligence to not only characterize the tell-tale artifacts present in AI-synthesized faces but also to significantly enhance their detection. The presentation highlighted the critical need for more robust and human-centric detection mechanisms, moving beyond the limitations of existing approaches that often fall prey to superficial features or are constrained by individual heuristic perceptions.

Key moments
- 0:00 Introduction and problem with AI-generated faces
- 1:00 Novel crowdsourcing approach and research questions
- 2:00 Annotation interface for collecting user data
- 3:30 Users distinguish AI-synthesized from real faces
- 4:50 Examples of user annotations and region extraction
- 5:50 Locations where AI synthesis artifacts occur
- 7:45 Common patterns of AI synthesis artifact regions
Double Face: Leveraging User Intelligence to Characterize and Recognize AI-synthesized Faces
Speakers: Matthew Joslin
Conference: USENIX Security '24
YouTube: https://www.youtube.com/watch?v=_D3kvU-KrQo
Overview
In an era witnessing an alarming surge of AI-generated images, particularly deepfakes, posing significant threats to information integrity and social trust, Matthew Joslin from the University of Texas at Dallas presented groundbreaking research titled "Double Face: Leveraging User Intelligence to Characterize and Recognize AI-synthesized Faces" at USENIX Security '24. This work, co-authored with Sien Wang and Dr. Shuangge, introduces a novel methodology that harnesses crowdsourcing intelligence to not only characterize the tell-tale artifacts present in AI-synthesized faces but also to significantly enhance their detection. The presentation highlighted the critical need for more robust and human-centric detection mechanisms, moving beyond the limitations of existing approaches that often fall prey to superficial features or are constrained by individual heuristic perceptions.
The increasing sophistication of generative adversarial networks (GANs) and diffusion models means that AI-synthesized imagery is becoming indistinguishable from real photographs to the untrained eye, leading to malicious uses ranging from deceptive LinkedIn profiles for suspected spies to the creation of fabricated political candidates and the proliferation of polarizing content via fake social media profiles. Given that neuroscientific research consistently demonstrates the human visual system's acute sensitivity to facial information, the team strategically focused their efforts on human face synthesis. Their research delves into how human perception can be systematically captured and integrated into AI detection models, offering a promising avenue for more resilient defenses against this evolving threat landscape.
Background
▶ Watch: Introduction and problem with AI-generated faces (0:00)
The rapid advancements in AI, particularly in generative models like StyleGAN, have democratized the creation of highly realistic synthetic images, often referred to as deepfakes. While these technologies hold promise for creative applications, their misuse has become a pressing concern in cybersecurity and information warfare. Real-world incidents, such as AI-generated profiles used for espionage, fabricated political figures, and coordinated disinformation campaigns leveraging synthetic faces, underscore the urgent need for effective detection mechanisms.
Existing approaches to detecting AI-synthesized images have demonstrated significant limitations. Many rely on off-the-shelf classifiers trained on specific datasets, which can become susceptible to "picking superficial features" that generalize poorly to new synthesis models or adversarial attacks. These models might learn to identify specific artifacts inherent to a particular GAN architecture but fail when faced with images from a different generator or when those artifacts are subtly altered. Another line of work involves heuristic-based detection developed by individual researchers. While offering valuable insights, these methods are inherently limited by individual perceptions and assumptions, lacking the broad empirical validation and scalability required for a comprehensive defense strategy. The challenge lies in developing detection systems that are not only robust to evolving synthesis techniques but also grounded in a deeper understanding of what makes an image appear "fake" to a human observer, especially given the human visual system's innate ability to process and interpret facial information. This work directly addresses these gaps by integrating human perception into the detection pipeline.
Key Findings
▶ Watch: Annotation interface for collecting user data (2:00)
The research was structured around four core questions, yielding several significant findings that collectively advance the understanding and detection of AI-synthesized faces:
- Perception of Artifacts: Users can indeed commonly perceive artifacts in AI-synthesized face images. The study found that, on average, 89% of real images were rated as real, while 65% of synthesized images were rated as fake. Furthermore, participants annotated significantly more suspicious regions in synthesized images compared to real ones, demonstrating a clear ability to discern the difference. Specifically, 52% of real images received less than one annotation on average, whereas only 6% of synthesized images had such a low annotation rate.
- Location of Artifacts: Artifacts in AI-synthesized faces are not uniformly distributed but are concentrated in specific regions. The analysis revealed that facial regions distant from the central face, such as the ears and hair, are more likely to exhibit discernible defects. Even more strikingly, non-facial objects like hats, earrings, and clothes consistently showed a significantly higher likelihood of containing artifacts compared to their real counterparts. Overall, non-facial regions in synthesized images were found to be 6.6 times more likely to be annotated as suspicious by users.
- Patterns of Artifacts: The study identified prevalent patterns that characterize these artifact regions. For facial regions, patterns like "blurry" and "unnatural skin" were strongly correlated with areas such as the ear and mouth. In non-facial regions, "blurry" was also a dominant pattern. Crucially, specific non-facial patterns like "unknown object" (correlated with clothing) and "mismatched" (associated with earrings) indicated that synthesized non-facial elements frequently do not align with recognizable, coherent objects.
- Improving Detection with User Perceptions: User perceptions can be effectively leveraged to improve the detection of AI-generated face images. By incorporating aggregate user-annotated regions into an attention learning framework, the researchers demonstrated a significant improvement in detection performance. Their attention learning approach, guided by human annotations, achieved high F-scores across various testing scenarios, outperforming other detection methods, particularly when tested against images from different or unseen synthesis models.
Technical Deep Dive
▶ Watch: Users distinguish AI-synthesized from real faces (3:30)
To systematically investigate human perception of AI-generated faces, the researchers designed a comprehensive study involving crowdsourcing intelligence from human annotators. Their dataset comprised 200 AI-synthesized faces (100 from StyleGAN 2 and 100 from StyleGAN 3) and 100 real faces for comparison. A total of 185 users were recruited via the MTurk platform to annotate these images.
The annotation process was structured in three distinct steps through a custom-developed interface:
- Suspicious Region Identification: Participants were first asked to draw bounding boxes over any suspicious regions they perceived within an image. Each drawn box was dynamically assigned an index for subsequent reference.
- Pattern Description: For each annotated bounding box, participants were prompted to input text describing the specific pattern or characteristic of the perceived artifact (e.g., "blurry," "unnatural," "mismatched").
- Fidelity Rating: Finally, users provided an overall fidelity rating for the entire image, indicating how "fake" or "real" it appeared. This was done using a predefined five-level Likert scale, which was later converted to integer scores from -2 (very fake) to +2 (very real) for quantitative analysis. An average score of zero served as the threshold separating perceptions of fake from real.
The analysis of the collected data addressed the four research questions. For the first question, regarding artifact perception, the researchers analyzed the fidelity ratings and the average number of bounding boxes per image. The results showed a clear distinction: 89% of real images were rated towards "real," while 65% of synthesized images were rated towards "fake." Moreover, synthesized images consistently received more annotations, with 52% of real images having less than one annotation on average, compared to only 6% of synthesized images. This statistically significant difference confirmed users' ability to perceive artifacts.
To identify common artifact locations (RQ2), an algorithm based on a union strategy was developed to extract connected components from the multiple bounding boxes drawn by different users for the same image. This aggregated approach allowed the creation of heat maps, where red indicated high consensus annotation levels. The extracted regions were then categorized into 15 frequently occurring locations: 10 facial (e.g., ear, hair, mouth, eye) and 5 non-facial (e.g., hat, earring, clothes). By calculating conditional probabilities, the study found that AI-synthesized images exhibited statistically significantly higher probabilities of artifacts in areas like the ear and hair, which are "distant from the central face." Most notably, non-facial objects (hats, earrings, clothes) in synthesized images showed a 6.6 times higher likelihood of being annotated as suspicious compared to real images.
For characterizing artifact patterns (RQ3), the location categories were correlated with the descriptive adjective terms provided by participants. This analysis revealed that "blurry" was a highly prevalent pattern, strongly correlated with regions like the ear, mouth, and all non-facial categories. "Unnatural skin" was frequently associated with facial regions. In non-facial contexts, patterns like "unknown object" (correlated with clothing) and "mismatched" (correlated with earrings) highlighted inconsistencies and lack of semantic coherence in synthesized non-facial elements, suggesting that these regions are often generated with less fidelity or attention to detail than the central face.
The most significant technical contribution for improving detection (RQ4) involved integrating these human insights into an AI detection model using an attention learning framework. The core idea was to guide the model's focus towards regions that human annotators consistently identified as suspicious. This was achieved by:
- Generating a human attention mask (M): This mask was extracted from the aggregate user annotations, highlighting high-consensus artifact areas in synthesized images.
- Utilizing a model attention mask (A): Based on an existing attention learning framework by Lee, the detection model itself generates an attention mask (A) that pinpoints regions it considers discriminative for detection.
- Introducing an attention loss function: A Mean Square Error (MSE) loss was introduced between the model's attention mask (A) and the human-derived mask (M). The loss function is defined as:
Loss = (1 / |Mask|) * Σ_{i,j} (A_{i,j} - M_{i,j})^2
where A_{i,j} and M_{i,j} are the pixel values at coordinates (i,j) in the model and human attention masks, respectively, and |Mask| is the size of the mask, normalizing the term between 0 and 1.
By minimizing this loss during training, the model's internal attention mechanism is refined to closely mirror human perceptions, effectively forcing the network to focus on the same regions flagged by human annotators.
The effectiveness of this attention learning approach was validated through extensive experiments. The results, presented as F-scores, demonstrated that the attention learning method with human annotations achieved consistently high accuracy across various testing scenarios. This included "same methods as training" (train/test split on StyleGAN 2 and StyleGAN 3), "other StyleGAN related model" (testing with different StyleGAN derivatives), and "other model" (testing with entirely different synthesis models). The human-informed attention learning consistently outperformed baseline detection approaches, showcasing its robustness and generalization capabilities, even in evasion experiments detailed in the full paper.
Demo / Proof of Concept
▶ Watch: Locations where AI synthesis artifacts occur (5:50)
While a live, interactive demonstration of the full AI detection system was not explicitly described in the presentation, the talk provided compelling visual evidence of the methodology's efficacy through several key examples. The researchers showcased their annotation interface, illustrating the process users followed to draw bounding boxes and describe artifacts. Crucially, they presented examples of aggregated user annotations on AI-synthesized face images, using heat colors to visually indicate regions with high levels of consensus among annotators. These visuals clearly demonstrated how the region extraction algorithm effectively identified compact suspicious regions, such as distorted ears or unusual backgrounds, that were commonly perceived as artifacts by multiple users. These graphical representations served as a powerful proof of concept for the ability of crowdsourcing to pinpoint and characterize subtle, yet discernible, flaws in synthetic media.
Defensive Implications
▶ Watch: Common patterns of AI synthesis artifact regions (7:45)
The findings from this research offer several critical implications for bolstering defenses against AI-synthesized face images:
- Prioritize Peripheral and Non-Facial Regions: Defenders should instruct their detection models to specifically focus on areas that humans consistently identify as problematic: the ears, hair, and non-facial elements such as hats, earrings, and clothing. Current models might over-index on central facial features, but this research shows that artifacts are often more pronounced and less consistently generated in peripheral and contextual elements.
- Target Specific Artifact Patterns: Security analysts and developers of detection systems should train models to look for the identified patterns: "blurry" textures, "unnatural skin," "unknown objects," and "mismatched" elements. Incorporating features engineered to detect these specific visual anomalies could significantly improve robustness.
- Leverage Human-in-the-Loop Training: The success of the attention learning framework underscores the value of human annotation in training more robust AI detection models. Security teams developing custom detection solutions should consider integrating crowdsourced or expert human feedback to guide their models' attention mechanisms, rather than relying solely on purely data-driven or black-box approaches. This can help models generalize better to novel synthesis techniques by aligning their focus with human-discernible flaws.
- Augment Training Data: Beyond attention masks, the identified artifact locations and patterns can inform better data augmentation strategies. For instance, synthetic images could be deliberately generated with subtle distortions in non-facial regions to expand the training data's diversity, making models more resilient to variations in artifact presentation.
- User Education: The insights gained can be used to educate the public and cybersecurity professionals on what to look for when encountering suspicious images. Knowing that blurry ears or inconsistent clothing are common indicators can empower human users to critically evaluate media, serving as a first line of defense against deepfakes.
- Generalizability to Other Media: The methodology of using crowdsourced intelligence to characterize and guide AI detection models is not limited to faces. It could be extended to other forms of AI-generated content, such as synthetic bodies, scenes, or even audio, offering a versatile approach to combating broader forms of AI-driven disinformation.
Key Takeaways
- Human perception is a powerful tool for discerning AI-synthesized faces, with users consistently identifying and annotating artifacts in synthetic images more frequently than in real ones.
- Artifacts in AI-generated faces are not random; they are concentrated in specific areas, particularly peripheral facial features like ears and hair, and most notably in non-facial elements such as hats, earrings, and clothing, which are 6.6 times more likely to be suspicious.
- Common artifact patterns include "blurry" textures, "unnatural skin," and semantic inconsistencies like "unknown objects" or "mismatched" components, especially in non-facial regions.
- Integrating human-derived attention masks into an attention learning framework significantly improves the performance and robustness of AI detection models, allowing them to focus on the same discriminative regions identified by human annotators.
- This crowdsourcing intelligence approach offers a robust and adaptable method to characterize and defend against evolving AI synthesis techniques, moving beyond the limitations of purely heuristic or superficial feature-based detection.
About the Speaker(s)
Matthew Joslin is a researcher from the University of Texas at Dallas. He presented the work "Double Face: Leveraging User Intelligence to Characterize and Recognize AI-synthesized Faces" at USENIX Security '24. His research was conducted in collaboration with Sien Wang and his advisor, Shuangge, also affiliated with the University of Texas at Dallas.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research presents a novel and highly effective method for deepfake detection by systematically leveraging crowdsourced human perception. By identifying specific artifact locations and patterns, and integrating this "user intelligence" into an attention learning framework, the work significantly enhances the robustness and generalizability of AI detection models against evolving synthesis techniques, offering critical, actionable defensive insights.
Heather Calloway (CISO) — STRONG ACCEPT
This research effectively leverages human perception to identify critical, often overlooked, indicators of AI-synthesized faces. By highlighting artifact concentrations in peripheral and non-facial regions, and demonstrating the power of human-guided attention learning, it offers a robust path for improving deepfake detection capabilities. It provides practical, actionable insights for security teams developing and deploying such defenses.