Detecting Race Conditions on macOS

Olivia Gallucci (Data Dog)

BSidesSF 2026 · Day 1 · AMC Theatre 07

Overview

In her BSides SF talk, Olivia Gallucci of Data Dog delved into the critical topic of detecting race conditions on macOS, with a particular focus on how the misuse of Grand Central Dispatch (GCD) can lead to severe vulnerabilities in privileged system services. The presentation meticulously breaks down the intricacies of concurrency on Apple's operating system, highlighting how subtle misconfigurations in dispatch queues and Quality of Service (QoS) classes can create exploitable timing windows. This talk is essential for security researchers, macOS developers, and detection engineers seeking to understand and mitigate a class of bugs that, while often perceived as mere reliability issues, can become potent vectors for privilege escalation and arbitrary code execution in sensitive contexts.

Watch on YouTube

Key moments

  1. 2:00 Introduction to detecting race conditions on macOS with GCD
  2. 2:50 Understanding processes, concurrency, and parallelism basics
  3. 4:15 Introduction to Apple's Grand Central Dispatch (GCD) framework
  4. 5:00 GCD dispatch queues and Quality of Service (QoS) classes
  5. 6:00 Race conditions and temporal vulnerabilities with GCD misuse
  6. 6:50 Detailed explanation of priority inversion in Darwin OS

Detecting Race Conditions on macOS

Speakers: Olivia Gallucci, Data Dog

Conference: BSides SF

YouTube: https://www.youtube.com/watch?v=4rg0xJrfjVI

Overview

In her BSides SF talk, Olivia Gallucci of Data Dog delved into the critical topic of detecting race conditions on macOS, with a particular focus on how the misuse of Grand Central Dispatch (GCD) can lead to severe vulnerabilities in privileged system services. The presentation meticulously breaks down the intricacies of concurrency on Apple's operating system, highlighting how subtle misconfigurations in dispatch queues and Quality of Service (QoS) classes can create exploitable timing windows. This talk is essential for security researchers, macOS developers, and detection engineers seeking to understand and mitigate a class of bugs that, while often perceived as mere reliability issues, can become potent vectors for privilege escalation and arbitrary code execution in sensitive contexts.

Gallucci emphasizes that concurrency bugs, specifically race conditions, are not just theoretical concerns but have demonstrably led to high-impact security vulnerabilities in macOS. By dissecting the underlying mechanisms of GCD, she reveals how assumptions about serial execution can be violated, transforming seemingly innocuous coding patterns into critical security flaws. The discussion is grounded in practical examples, including a well-known CVE, providing a clear roadmap for identifying, preventing, and detecting these elusive yet dangerous vulnerabilities within the macOS ecosystem.

The talk serves as a comprehensive guide, bridging the gap between theoretical concurrency concepts and their real-world security implications. It provides actionable insights for both static code analysis and dynamic telemetry-based detection, equipping professionals with the knowledge to fortify macOS services against a prevalent and often underestimated threat.

Background

▶ Watch: Introduction to detecting race conditions on macOS with GCD (2:00)

To understand race conditions on macOS, it's crucial to first grasp fundamental concepts of process management and concurrency. A process is an instance of a running executable, acting as a container for resources like virtual memory and descriptors. However, processes themselves are not runnable entities; rather, they contain threads, which are the actual units of execution. When one refers to a process executing, it is, in fact, one or more of its threads performing work. This distinction is vital for understanding how multiple operations can interact.

Concurrency refers to a system's ability to make progress on multiple tasks by overlapping or interleaving their execution, even on a single processor. This improves responsiveness and resource utilization. Interleavings are the various possible execution orders of operations from multiple threads accessing shared resources, often leading to non-deterministic or unpredictable behavior. In contrast, parallelism involves the simultaneous execution of multiple tasks on multiple processor cores. Gallucci explains that her talk addresses both: concurrency for its logic and timing bugs, and parallelism as a supporting concept for how GCD scales execution, and critically, how this scaling can introduce vulnerabilities through thread overcommitment and the interleaving of privileged and unprivileged operations.

At the heart of macOS concurrency is Grand Central Dispatch (GCD), Apple's powerful framework for managing tasks and execution without explicit thread management by developers. GCD is a system library that uses dispatch queues to organize blocks of work. These queues come in two primary types: serial queues, which ensure tasks run one at a time in the order they are received, and concurrent queues, which allow multiple tasks to execute simultaneously, managed by the system scheduler. Developers can create custom queues or utilize global ones, assigning a Quality of Service (QoS) class. QoS classes (e.g., user initiated, user interactive, background) are critical as they influence scheduling priority, dictating how the kernel allocates CPU resources. In security-critical services, correct QoS ensures that high-priority operations are not delayed by lower-priority tasks, which can create exploitable timing windows. GCD, by design, does not automatically prevent race conditions; correct usage of queues and explicit serialization are paramount to avoid these vulnerabilities.

Key Findings

▶ Watch: Introduction to Apple's Grand Central Dispatch (GCD) framework (4:15)

The core findings of this talk revolve around specific types of concurrency vulnerabilities that arise from the misuse or misunderstanding of GCD's mechanisms, particularly in privileged macOS services. Gallucci identifies three main categories: priority inversions, dispatch sync deadlocks, and resource starvation via worker thread busyness. These issues, while seemingly distinct, often stem from a common root: the failure to correctly manage thread scheduling, resource access, and synchronization in a multi-threaded environment.

Priority Inversion occurs when a high-priority task becomes blocked, waiting for a lower-priority task to complete. On Darwin, the scheduler can indefinitely preempt (or end) low QoS threads in favor of higher QoS threads. This can be problematic if a low-priority thread holds a critical resource, such as a lock, and is never scheduled to release it because higher-priority threads continuously preempt it. A notorious example is the use of spin locks (busy-waiting loops) without kernel assistance, which can effectively freeze programs on Darwin. While Apple's kernel implements priority inheritance for specific synchronization primitives like pthread_mutex_t and OS_unfair_lock, many others, such as semaphores or custom lock implementations, do not participate in this mechanism, leaving services vulnerable. This leads to what is known as the "internal semaphore anti-pattern," where high QoS work items synchronously wait on lower QoS tasks that can unexpectedly stall.

Dispatch Sync Deadlocks represent a specific type of deadlock that arises from the synchronous submission of a block to a dispatch queue. The danger occurs when dispatch_sync targets the same serial queue that the caller is already running on, or a queue it implicitly depends on. In such a scenario, the calling thread blocks, waiting for the submitted block to complete, but the submitted block cannot start until the currently executing task (the caller itself) finishes, creating a circular wait. A common example is calling dispatch_sync from the main thread onto the main queue, which will instantly hang the application. This pitfall extends to any scenario where a synchronous callback creates a dependency loop, leading to immediate service freezes and system unresponsiveness.

Lastly, Resource Starvation via Worker Thread Busyness occurs when excessive dispatch operations, especially blocking ones, saturate GCD's underlying thread pool. While GCD abstracts thread management, it still relies on a finite pool of worker threads. If a large number of tasks are dispatched, or if these tasks block indefinitely (e.g., waiting on locks or synchronous calls), the thread pool can become exhausted. In extreme cases, libdispatch may spawn additional threads to try and break stalemates, leading to thread explosion, where a process suddenly creates tens or hundreds of threads, thrashing the CPU and starving other system components. Historically, Apple has learned from these issues, with examples like the now-abandoned Security Transforms API in macOS 10.7, which inadvertently caused severe thread proliferation, and the subsequent rewriting of many macOS daemons and iOS 12 services to be single-threaded to improve performance.

These findings collectively highlight how subtle misconfigurations or misunderstandings of GCD's asynchronous nature and scheduling semantics can degrade system reliability and, more critically, introduce severe security vulnerabilities in privileged contexts.

Technical Deep Dive

▶ Watch: GCD dispatch queues and Quality of Service (QoS) classes (5:00)

The technical core of Gallucci's talk lies in detailing the specific mechanisms of GCD misuse that lead to the aforementioned vulnerabilities and grounding these concepts in a real-world exploit.

Priority Inversion and Synchronization Primitives

Darwin's scheduler, unlike some other operating systems, aggressively preempts lower QoS threads in favor of higher ones. This behavior, while generally improving system responsiveness, creates a unique challenge for synchronization. If a high-priority task needs a resource held by a low-priority task, and that low-priority task is continuously preempted, the high-priority task will stall, leading to a priority inversion. This is particularly acute with spin locks, where a blocked thread repeatedly checks if a lock is free, consuming CPU cycles without yielding. If the thread holding the lock has a lower QoS, it might never get enough CPU time to release the lock, causing a deadlock.

Apple's kernel offers solutions through priority inheritance in specific synchronization primitives. For instance, pthread_mutex_t (a traditional mutual exclusion lock) and OS_unfair_lock (a lightweight, low-level lock) support this mechanism. If a high QoS thread waits on a lock held by a low QoS thread, and that lock is one of these types, the kernel will temporarily boost the low QoS thread's priority to match the high QoS thread until the lock is released. This prevents the inversion. However, as Gallucci points out, many other synchronization mechanisms, such as reader/writer locks, semaphores, and custom lock implementations, do not participate in QoS inheritance. The scheduler has no intrinsic way to know which thread should be boosted when these primitives are used. This leads to the "internal semaphore anti-pattern," a common and seemingly convenient way to bridge asynchronous work into synchronous control flow, but one that fundamentally breaks Apple's scheduling and progress guarantees, often resulting in priority inversions and deadlocks.

Xcode's Thread Performance Checker is a powerful runtime tool that developers can enable in scheme settings to automatically detect priority inversions. It logs warnings when a higher QoS thread is detected waiting on a lower QoS thread, flagging potential issues early in the development cycle. Gallucci recommends flagging any scenario where a high-priority queue or thread is blocked on scheduled work to a lower QoS queue during code reviews, as this QoS mismatch is a significant red flag.

Dispatch Sync Deadlocks

The dispatch_sync function synchronously executes a block on a target queue, blocking the calling thread until the block finishes. The critical danger arises when the target queue is the same serial queue that the caller is already running on, or a queue that it directly depends on. For example, if the main thread calls dispatch_sync on the main queue, it will immediately deadlock. The calling thread waits, but the block cannot execute because the queue is busy with the calling thread's current task. This "waiting on itself" scenario is a common pitfall. Gallucci cites an instance where a developer inadvertently caused a startup deadlock by queuing work to background threads, each performing a dispatch_sync on the main queue. This saturated GCD's worker threads, and when the main thread itself made a synchronous dispatch call, it waited for a free worker thread that never came, leading to a circular wait and an application freeze. The recommended solution is to use dispatch_async for cross-queue calls and to restructure code to avoid synchronous callbacks that can create these dependency loops. Scanning for dispatch_sync calls targeting serial listener queues or the main queue from within a callback is a strong indicator of a potential bug.

Resource Starvation and Thread Explosion

GCD manages a worker thread pool to execute tasks from concurrent queues. While early GCD documentation implied smart thread limiting, experience showed that pathological cases were easy to hit. If an excessive number of tasks are dispatched, or if these tasks block for extended periods (e.g., waiting on locks or semaphores), the thread pool can become saturated. In an attempt to prevent deadlocks, libdispatch will spawn additional threads if existing ones are blocked. This can lead to thread explosion, where a process rapidly creates tens or hundreds of threads, causing severe CPU thrashing, degrading performance, and starving other system components of CPU time. Gallucci highlights historical examples like the "security transforms" API in macOS 10.7, which created a new queue and thread per task, causing severe proliferation. This led to Apple rewriting many macOS daemons and iOS 12 services to be single-threaded, demonstrating a shift away from unconstrained concurrency. To avoid starvation, limiting concurrent dispatches and, crucially, avoiding blocking calls on worker threads are essential practices.

The GSS Cred XPC Service CVE (CVE-2018-4237)

Gallucci grounds these concepts in a famous real-world incident: a 2018 CVE (CVE-2018-4237) discovered by Brandon Aad in the com.apple.gscredxpc service. This service, identified by Apple's reverse DNS naming scheme, is a built-in macOS system service responsible for managing Generic Security Services (GSS) credentials, such as Kerberos and enterprise SSO tickets.

The vulnerability was a high-impact race condition that allowed an unprivileged process to trigger memory corruption in this privileged root service, ultimately leading to arbitrary code execution in a root context. The root cause was a subtle but critical misconfiguration: the service intended to use a serial dispatch queue for handling events but failed to explicitly set this queue as the target queue for client XPC connections. As a result, the connection handler executed on a default concurrent queue instead of the intended serial queue. This omission meant that message handlers ran concurrently, violating the service's implicit assumptions about serialization.

The exploit mechanism leveraged this concurrency: two requests could interleave such that one handler freed a credential while another was still actively processing it. This created a precise timing window, allowing controlled memory corruption and subsequently attacker-controlled data to be used for arbitrary code execution. The business impact was profound: assumptions about serialization were broken by a misconfigured queue and XPC target, turning a should-be-serialized privileged service into a concurrent handler. This demonstrated that race conditions, even without a kernel compromise, can be exploited to achieve significant privilege escalation.

This CVE serves as a stark reminder that a single queue targeting mistake can transform a normal IPC interaction into a sandbox escape vector, elevating unprivileged code into a privileged context and severely undermining endpoint security. It underscores the importance of meticulously reviewing XPC service implementations for correct concurrency management, especially when shared mutable state is involved.

Demo / Proof of Concept

▶ Watch: Race conditions and temporal vulnerabilities with GCD misuse (6:00)

While the talk does not feature a live demonstration of an exploit or a proof-of-concept for the CVE discussed, Olivia Gallucci highlights a crucial tool for developers and security researchers: Xcode's Thread Performance Checker.

This runtime tool, accessible through Xcode's scheme settings, automatically detects priority inversions as an application runs. When enabled, Xcode will log warnings if it identifies a higher QoS thread waiting on a lower QoS thread. This provides real-time feedback on potential QoS mismatches and concurrency issues, including instances of non-UI work running on the main thread. Gallucci stresses that this tool is an easy and effective way to observe QoS-related problems, especially for those learning vulnerability research on macOS. Although it doesn't demonstrate active exploitation, it serves as a powerful diagnostic aid for identifying the underlying conditions that could lead to race conditions and deadlocks, thereby facilitating the prevention and early detection of these vulnerabilities.

Defensive Implications

▶ Watch: Detailed explanation of priority inversion in Darwin OS (6:50)

Detecting and preventing race conditions in macOS services requires a multi-layered approach, combining static code analysis with dynamic telemetry and behavioral detection. Gallucci outlines actionable strategies for defenders:

Static Review Patterns

Static code reviews are crucial for identifying architectural weaknesses before deployment:

  • XPC Connection Queuing: When a daemon accepts an XPC connection, it's vital to verify that setTargetQueue() is always called on the connection. Failure to do so means message handlers might execute on an unintended concurrent queue, violating serialization assumptions and exposing race conditions in privileged services.
  • Serial vs. Concurrent Execution Assumptions: Review code for logic that implicitly depends on serial handling, particularly around authorization state, object lifecycle management, or shared caches. XPC clients can easily exert parallel pressure, so if an implementation relies on ordering guarantees not explicitly enforced (e.g., with a serial queue or locks), it's a high-value finding.
  • Synchronous Calls (especially dispatch_sync): Any use of dispatch_sync in a privileged daemon should be treated as suspicious and require explicit justification. While not all synchronous dispatches are inherently wrong, they often indicate blocking behavior, risk of lock inversion, or a path towards deadlocks under load, especially under adversarial testing.
  • Shared Mutable State without Synchronization: Identify global variables or shared objects accessed from multiple handlers without appropriate synchronization mechanisms (e.g., locks, atomics, or dedicated serial queue funnels). This is a common root cause of timing-dependent flaws and addressing it provides significant returns in both reliability and security.

Telemetry Signals

Telemetry provides insights into where architectural weaknesses are becoming active operational risks:

  • Crash Patterns: Repeated crash patterns, assertions, or guard failures within a privileged daemon, particularly at timing-sensitive code paths (state transitions, cleanup, request handling boundaries), are strong indicators of a race window being hit. A single crash might be a stability bug, but repeated crashes suggest an exploitable race.
  • Thread Churn Spikes: An unusual number of threads being created or rapid changes in queue drain behavior can signal that a service is under concurrency stress it wasn't designed for. Attackers probing for race windows often generate this exact kind of pressure.
  • Queue Backlogs: Evidence of work piling up faster than it drains, heavy synchronous waits, or long wait times on queue dispatch can indicate deadlocks, priority inversions, or lock contention. These conditions often precede visible crashes, offering an opportunity to detect exploitation attempts before they fully materialize.

Behavioral Detection

Analyzing runtime behavior against a baseline helps differentiate normal operation from potential adversarial activity:

  • Frequent Thread Churn and Spikes: Flag processes exhibiting frequent thread creation and tear-down, or sudden spikes in thread count, relative to their established baseline. Focusing on relative deviation rather than absolute volume reduces false positives, as some daemons are naturally noisy.
  • Rate Limiting XPC Invocations: Consider rate-limiting XPC invocations, at least locally. If a client issues many parallel requests to a privileged service, it might be probing for race windows. This is particularly interesting when request volume and concurrency are high, but the target daemon normally expects low to moderate parallelism. Such activity should be logged, scored, and correlated with crashes or queue delays.
  • Queue Drain Time Anomalies: Elevated synchronization wait times are a strong signal of deadlocks, lock contention, and scheduling inversions. This highlights how performance telemetry can double as security telemetry, allowing SRE or platform engineering metrics to aid threat detection engineering in identifying adversarial timing attacks.

The overall defensive model is straightforward: static review pinpoints where race conditions are likely, telemetry indicates where those weaknesses are being stressed, and behavioral detections reveal when that stress appears adversarial. By combining these approaches, defenders can proactively identify and mitigate exploit opportunities in privileged macOS services.

Key Takeaways

  • Race Conditions are Security Critical: On macOS, race conditions are not just reliability bugs; in privileged service boundaries, they can quickly escalate to security problems, enabling privilege escalation and arbitrary code execution.
  • GCD Misuse is a Primary Vector: Misconfigurations in Grand Central Dispatch (GCD) queues, Quality of Service (QoS) classes, and synchronization choices are common root causes for exploitable race conditions, deadlocks, and resource starvation.
  • Watch for Priority Inversions: Be vigilant for high-priority tasks waiting on lower-priority tasks, especially when using synchronization primitives that do not support QoS inheritance (e.g., semaphores, custom spin locks). Xcode's Thread Performance Checker is a valuable detection tool.
  • Avoid dispatch_sync Pitfalls: Never call dispatch_sync on a queue from within that same queue or any queue that creates a synchronous dependency loop, as this almost guarantees a deadlock. Prefer dispatch_async for cross-queue communication.
  • Guard Against Resource Starvation: Excessive concurrent dispatches and blocking operations can lead to thread pool saturation and "thread explosion," degrading performance and creating opportunities for denial-of-service or other resource-based attacks.
  • Multi-Layered Detection is Essential: Employ a combination of static code review (for unsafe queuing patterns, implicit serialization assumptions), telemetry (for crash patterns, thread churn spikes, queue backlogs), and behavioral detection (for anomalous thread activity, suspicious XPC invocation rates, and queue drain time anomalies) to effectively identify and mitigate these vulnerabilities.

About the Speaker(s)

Olivia Gallucci is a security professional who works at Data Dog. She specializes in macOS security, particularly focusing on concurrency issues and vulnerability research. Her expertise extends to analyzing how system services and frameworks like Grand Central Dispatch can be misused to create exploitable conditions. Beyond this talk, Olivia has shared her insights through other platforms, including a blog post on tactile attacks and a podcast episode with "Hackers on the Rocks," also discussing tactile attacks on macOS. Her work emphasizes the importance of understanding low-level operating system mechanisms to identify and prevent complex security flaws.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

A competent, well-structured survey of GCD concurrency pitfalls and their security implications, anchored by a real CVE. Solid educational content for a BSides audience, but it reads more as a curated synthesis than original research — the CVE is from 2018 and belongs to someone else, and the detection heuristics, while sensible, aren't novel.

Heather Calloway (CISO) — WEAK

Technically solid and well-structured work on a real vulnerability class, but this talk is written for macOS developers and detection engineers — not security leaders or governance audiences. The defensive detection model is genuinely useful, but it never connects to the institutional conditions that allow these vulnerabilities to ship in the first place.

→ Top-rated talks at BSidesSF 2026

All talks from BSidesSF 2026