FAMOS: Robust Privacy-Preserving Authentication on Payment Apps via Federated Multi-Modal Contrastive Learning

Yifeng Cai, Ziqi Zhang, Jiaping Gui, Bingyan Liu, Xiaoke Zhao, Ding Li

33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24

Overview

In an era where digital payment applications have become ubiquitous, securing transactions against unauthorized access is paramount. The talk "FAMOS: Robust Privacy-Preserving Authentication on Payment Apps via Federated Multi-Modal Contrastive Learning" by Yifeng Cai and colleagues introduces a novel solution to this critical challenge. This research directly addresses significant limitations in existing authentication methods for payment apps, which are often vulnerable to device compromise or privacy breaches.

Watch on YouTube

Visual summary for FAMOS: Robust Privacy-Preserving Authentication on Payment Apps via Federated Multi-Modal Contrastive Learning by Yifeng Cai, Ziqi Zhang, Jiaping Gui, Bingyan Liu, Xiaoke Zhao, Ding Li
Visual summary for FAMOS: Robust Privacy-Preserving Authentication on Payment Apps via Federated Multi-Modal Contrastive Learning by Yifeng Cai, Ziqi Zhang, Jiaping Gui, Bingyan Liu, Xiaoke Zhao, Ding Li

Key moments

  1. 0:00 Introduction to FAMOS and payment app security problem
  2. 2:00 Two key limitations of current authentication methods
  3. 3:25 Two key insights for FAMOS's robust design
  4. 4:40 Three core design goals: robustness, lightweight, privacy
  5. 6:00 Detailed explanation of FAMOS's training phase and modules
  6. 8:00 The five research questions guiding FAMOS evaluation
  7. 9:25 Initial results on FAMOS's overall effectiveness

FAMOS: Robust Privacy-Preserving Authentication on Payment Apps via Federated Multi-Modal Contrastive Learning

Speakers: Yifeng Cai; Ziqi Zhang; Jiaping Gui; Bingyan Liu; Xiaoke Zhao; Ding Li

Conference: USENIX Security '24

YouTube: https://www.youtube.com/watch?v=QkZEBCO-egQ

Overview

In an era where digital payment applications have become ubiquitous, securing transactions against unauthorized access is paramount. The talk "FAMOS: Robust Privacy-Preserving Authentication on Payment Apps via Federated Multi-Modal Contrastive Learning" by Yifeng Cai and colleagues introduces a novel solution to this critical challenge. This research directly addresses significant limitations in existing authentication methods for payment apps, which are often vulnerable to device compromise or privacy breaches.

The presented work, FAMOS, focuses on developing a robust and privacy-preserving authentication service that leverages the power of behavioral biometrics. The speakers highlight the sheer scale of the problem, citing Alipay, a prominent payment app in China, with over 300 million daily active users and daily payments exceeding $20 million. Such a high-stakes environment necessitates an authentication mechanism that goes beyond traditional passwords or device-bound biometrics, which can be easily circumvented if a device is compromised or credentials are shared.

FAMOS proposes an innovative framework that fuses multimodal sensor data and employs federated multi-modal contrastive learning to create a transparent, user-friendly, and highly secure authentication system. The significance of this research lies in its ability to overcome two major hurdles: the disruptive influence of real-world background activities on sensor readings and the inherent privacy violations in conventional machine learning training paradigms for behavioral biometrics. By tackling these issues, FAMOS aims to provide a deployable and effective solution for safeguarding digital payments in an increasingly mobile-centric world.

Background

▶ Watch: Introduction to FAMOS and payment app security problem (0:00)

The landscape of digital payments has seen an exponential rise, bringing convenience but also amplifying security risks. Traditional authentication methods, such as passwords, PINs, or device-native biometrics like Face ID and Touch ID, while foundational, possess inherent vulnerabilities. Passwords can be compromised through data breaches or shared inadvertently, while device-bound biometrics can be bypassed if the device itself is compromised or if unauthorized users (e.g., children) gain access to a parent's device where their biometrics are already registered. This necessitates a more dynamic and less perceptible form of authentication, particularly for high-frequency, high-value transactions within payment applications.

Behavioral biometrics, which analyze distinctive user behavior patterns captured by built-in device sensors, have emerged as a promising alternative. Solutions like AuthenticSense and KOS aim to provide continuous, transparent authentication by monitoring user interactions such as typing rhythm, swipe patterns, or device handling. These methods are designed to operate imperceptibly, enhancing security without disrupting user experience. However, their real-world deployment in payment applications has faced two critical limitations, as identified by the FAMOS researchers.

The first major limitation is the negative influence of background activities on sensor readings. Existing behavioral biometric systems often assume users remain stationary or perform actions in controlled environments. In reality, users frequently interact with their mobile devices under varying conditions – walking, lying down, sitting on public transport, or engaging in other concurrent activities. These background movements introduce significant "noise" into sensor data, suchating the subtle, user-specific patterns that behavioral biometrics rely upon. For instance, the accelerometer might capture more user-related activity when a user is lying down, but less when they are walking, where the touchscreen sensor might become more indicative. Direct utilization of raw sensor data without accounting for these contextual variations leads to a significant degradation in authentication accuracy.

The second, equally critical, limitation is the violation of user privacy in training data. Many machine learning approaches for behavioral biometrics rely on a training strategy where one user's data serves as a "positive" sample, and data from other users is used as "negative" samples to distinguish between legitimate and illegitimate users. Collecting and sharing such sensitive behavioral data across multiple users, especially for training purposes, fundamentally conflicts with stringent privacy regulations like the General Data Protection Regulation (GDPR) in the EU and the Personal Information Protection Law (PPL) in China. This requirement for cross-user data sharing creates a significant barrier to the ethical and legal deployment of these solutions in real-world applications, particularly for privacy-sensitive payment platforms. Addressing these two core challenges is central to the FAMOS framework.

Key Findings

▶ Watch: Two key insights for FAMOS's robust design (3:25)

The FAMOS framework demonstrates significant advancements in robust and privacy-preserving authentication, validated through comprehensive evaluations. The key findings highlight its superior performance and innovative design choices:

  • Superior Overall Effectiveness: FAMOS substantially outperforms existing behavioral biometric baselines, such as AuthenticSense and KOS. It achieves a False Rejection Rate (FRR) that is 4.4 times lower and an Equal Error Rate (EER) that is 5.5 times lower. Furthermore, FAMOS shows a 27.7 percentage higher Area Under the Curve (AUC) and a 42.2 percentage higher F-score, indicating a much better balance between precision and recall in authentication.
  • Robustness to Background Noise: The Sensor Fusion module is highly effective in mitigating the impact of background activities. By intelligently combining data from multiple sensors, FAMOS achieves up to a 31.9 percentage improvement in accuracy compared to relying on a single sensor alone. The integrated attention mechanism dynamically assigns appropriate weights to different sensors based on their stability under specific background conditions, ensuring that the most reliable data contributes maximally to authentication.
  • Privacy-Preserving User Differentiation: The Contrastive Learning module successfully clusters user-specific action representations. This approach increases the difference between representations of the same action (positive samples) and different actions (negative samples) by a factor of 1.4 times. Crucially, this is achieved without requiring data from other users as negative samples, directly addressing and resolving the privacy concerns associated with traditional training methodologies.
  • Enhanced Performance with Federated Learning: The integration of Federated Learning significantly boosts the system's capabilities. Evaluations show that models trained locally without federated aggregation perform approximately 15 percentage lower than those benefiting from the aggregated global knowledge while preserving user-specific insights on individual devices. This confirms federated learning's role in improving overall model performance and facilitating the training process in a privacy-compliant manner.
  • Low On-Device Overhead: FAMOS is designed to be lightweight and deployable on user smartphones. Experimental results confirm that the system operates with relatively low computational overhead on various devices. This ensures that FAMOS can provide robust authentication services without negatively impacting device performance or disturbing the user's daily usage of payment applications.

Technical Deep Dive

▶ Watch: Three core design goals: robustness, lightweight, privacy (4:40)

The FAMOS framework is architecturally designed around three core goals: robustness against environmental noise, lightweight deployment on user smartphones, and privacy preservation for sensitive user data. To achieve these, it integrates a sophisticated pipeline comprising a Sensor Fusion Module, a Contrastive Learning Module, and a Federated Learning Aggregation Module.

The process begins with the Sensor Fusion Module. This module is responsible for ingesting raw data from multiple on-device sensors, specifically the touchscreen sensor and a 3-axis IMU (which typically includes an accelerometer, gyroscope, and magnetometer). Each sensor stream is first processed by a dedicated encoder to extract relevant features. The critical innovation here lies in the subsequent application of an attention mechanism. This mechanism dynamically assigns importance weights to the features from different sensors. For instance, if a user is walking, the touchscreen sensor might provide more stable and discriminative data compared to the accelerometer, which would be heavily influenced by gait. Conversely, when a user is lying still, the accelerometer might offer more subtle behavioral cues. The attention mechanism intelligently identifies the most stable and informative sensor readings under varying background activities, effectively mitigating the influence of noise and generating high-quality, fused features. This dynamic weighting is crucial for the system's robustness in real-world, noisy environments.

Following sensor fusion, the processed features are fed into the Contrastive Learning Module. This module is central to FAMOS's ability to differentiate users while respecting privacy. Unlike traditional methods that require negative samples from other users, FAMOS employs a self-supervised approach. Its objective is to cluster the representation vectors of the same action (e.g., a specific swipe pattern or typing rhythm by User A) close to each other in the feature space. Simultaneously, it pushes representations of different actions (e.g., a different swipe pattern by User A, or any action by User B, though User B's data is not explicitly used as negative samples) far apart. The module minimizes the representation distance for samples belonging to the same action category for a given user, while maximizing the distance for samples from different action categories. This design ensures that after training, a user's data samples form distinct, compact clusters for their various actions, and crucially, data samples from other users are inherently dissimilar and excluded from these clusters due to their distinct feature representations. This elegant design eliminates the need for cross-user data sharing for negative sampling, thereby upholding user privacy.

The final architectural component is the Federated Learning Aggregation Module. This module addresses the privacy goal by enabling collaborative model training without centralizing raw user data. Each user's device locally trains its FAMOS model using its own data. Once trained, these local models (or their updated parameters, not the raw data) are sent to a central server. The Federated Learning Aggregation Module then takes these locally trained user models as input and aggregates them to generate a global, generalized model. This aggregated model captures the common knowledge and patterns across the user base. This generalized model is then sent back to the individual devices, which can further refine it with their specific data, thereby leveraging collective intelligence while keeping user-specific knowledge and raw data securely on the device. This iterative process ensures privacy-preserving training while continually improving the model's accuracy and generalization capabilities.

In summary, the FAMOS architecture meticulously combines sensor fusion with dynamic attention, privacy-centric contrastive learning, and a federated learning paradigm. This synergistic approach allows it to achieve robust, lightweight, and privacy-preserving authentication, making it a viable solution for the demanding security requirements of modern payment applications.

Demo / Proof of Concept

▶ Watch: The five research questions guiding FAMOS evaluation (8:00)

While the talk did not feature a live, interactive demonstration of the FAMOS system, the researchers presented a thorough and rigorous evaluation that serves as a compelling proof of concept. The effectiveness and efficiency of FAMOS were validated through extensive experiments conducted on both real-world and in-lab datasets, utilizing actual mobile devices.

For the data collection, two distinct datasets were compiled:

  • A real-world dataset involved 70 Alipay users, with 20 designated as "victims" (legitimate users) and 50 as "attackers" (impersonators attempting unauthorized payments). This dataset captured authentic user behavior in a high-stakes environment.
  • An in-lab dataset was created with 24 volunteers, grouped into 4 victims and 20 attackers, allowing for more controlled experimentation and in-depth analysis of specific scenarios.

Both datasets included recordings from touchscreen sensors and 3-axis IMU sensors (accelerometer, gyroscope, magnetometer), capturing detailed behavioral patterns.

The on-device performance of FAMOS was evaluated across four different mobile devices, demonstrating its practical deployability and low overhead. To benchmark FAMOS's capabilities, the researchers compared its performance against two established behavioral biometric baselines: AuthenticSense and KOS.

The evaluation utilized standard metrics for authentication systems, including False Rejection Rate (FRR), Equal Error Rate (EER), True Positive Rate (TPR), F-score, and Area Under the Curve (AUC). The consistently superior results across these metrics, as detailed in the "Key Findings" section, strongly substantiate the claims of robustness, accuracy, and efficiency for FAMOS, effectively serving as its proof of concept.

Defensive Implications

▶ Watch: Initial results on FAMOS's overall effectiveness (9:25)

The FAMOS framework offers profound defensive implications for payment application developers, financial institutions, and indeed, any platform requiring robust and privacy-preserving user authentication. Its core innovations provide a blueprint for enhancing security posture against sophisticated threats.

Firstly, payment app developers should move beyond static, one-time authentication methods and embrace continuous, transparent authentication powered by behavioral biometrics. FAMOS demonstrates that this can be achieved without compromising user experience. By continuously monitoring subtle user interactions, payment apps can detect anomalies indicative of unauthorized access in real-time, even if initial login credentials have been compromised. This shifts the security paradigm from a single gate to an ongoing vigilance.

Secondly, the emphasis on multimodal sensor data fusion is a critical takeaway. Developers should design authentication systems that leverage a diverse array of on-device sensors (touchscreen, IMU, etc.) rather than relying on a single data source. Crucially, the FAMOS attention mechanism highlights the importance of intelligently weighting these sensors based on environmental context and user activity. This adaptive approach ensures robustness against varied real-world conditions, making the authentication system far more resilient to noise and more discriminative.

Thirdly, FAMOS sets a new standard for privacy-preserving machine learning in security applications. The adoption of contrastive learning to cluster user-specific actions without requiring other users' data for negative samples is a pivotal development. This allows for personalized behavioral models to be built and maintained without violating stringent data privacy regulations like GDPR or PPL. For organizations operating under strict privacy mandates, this approach offers a viable path to deploy advanced behavioral analytics.

Finally, the successful integration of federated learning provides a blueprint for collaborative intelligence without centralized data exposure. Payment app providers can leverage the collective learning from a vast user base to improve the generalizability and accuracy of their authentication models, while keeping sensitive user data securely on individual devices. This approach reduces the risk of large-scale data breaches and fosters greater user trust, which is paramount in financial applications.

In essence, FAMOS challenges defenders to adopt a more dynamic, context-aware, and privacy-centric approach to authentication. By incorporating these principles, payment apps can build more secure, user-friendly, and legally compliant systems that are better equipped to protect against the evolving threat landscape.

Key Takeaways

  • Traditional authentication methods and existing behavioral biometrics in payment apps are vulnerable to device compromise, credential sharing, and performance degradation due to real-world background noise.
  • FAMOS introduces a novel framework that achieves robust, privacy-preserving authentication by fusing multimodal sensor data and employing federated multi-modal contrastive learning.
  • The system's Sensor Fusion Module with an attention mechanism dynamically selects and weights stable sensor data under varying background activities, leading to significant accuracy improvements (up to 31.9%).
  • Contrastive Learning enables FAMOS to effectively cluster user-specific action representations, differentiating legitimate users without requiring other users' sensitive data as negative samples, thereby upholding privacy regulations.
  • Federated Learning enhances the overall performance of the authentication model by aggregating general knowledge while preserving user-specific data on devices, leading to a 15% improvement over local models.
  • FAMOS demonstrates superior performance against baselines (4.4x lower FRR, 5.5x lower EER, 27.7% higher AUC, 42.2% higher F-score) and operates with low overhead on real-world mobile devices, making it practical for deployment.

About the Speaker(s)

The research on FAMOS was presented by Yifeng Cai, who had the privilege of working on this project with colleagues Ziqi Zhang, Jiaping Gui, Bingyan Liu, Xiaoke Zhao, and Ding Li. The presentation highlighted their collective efforts to address pressing issues in security authentication services for payment applications in the real world.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

FAMOS presents a robust and privacy-preserving authentication framework for payment applications, addressing critical limitations in existing behavioral biometrics. By intelligently fusing multimodal sensor data, employing privacy-centric contrastive learning, and leveraging federated learning, it delivers significant performance improvements while navigating real-world noise and stringent privacy regulations. This is a well-executed defensive innovation for a high-stakes environment.

Heather Calloway (CISO) — STRONG ACCEPT

This research offers a compelling, privacy-preserving approach to authentication for payment applications, directly addressing critical business risks from fraud and regulatory non-compliance. By leveraging multimodal sensor fusion and federated contrastive learning, FAMOS provides a robust, real-world solution that significantly enhances security without sacrificing user privacy or experience.

→ Top-rated talks at 33rd USENIX Security Symposium

All talks from 33rd USENIX Security Symposium