The Oversights Under the Flow: Discovering the Vulnerable Tooling Suites From Azure MLOps

Black Hat Asia 2025 · Day 2 · Briefings

Overview

This talk, "The Oversights Under the Flow," delves into a critical examination of security vulnerabilities discovered within the tooling suites of Azure Machine Learning Operations (MLOps). Presented by a researcher from Songhai University, the session highlights how seemingly simple and easily discoverable security flaws, often stemming from "oversights" in development and maintenance, can persist and lead to significant impacts within complex software supply chains. The speaker emphasizes that while these vulnerabilities might appear minor at first glance, their presence in core MLOps infrastructure—even tools maintained by a tech giant like Microsoft—poses substantial risks if left unaddressed or incompletely patched.

Watch on YouTube

Visual summary for The Oversights Under the Flow: Discovering the Vulnerable Tooling Suites From Azure MLOps
Visual summary for The Oversights Under the Flow: Discovering the Vulnerable Tooling Suites From Azure MLOps

Key moments

  1. 0:00 Introduction to MLOps vulnerabilities and recurring oversights
  2. 2:15 Comprehensive agenda outlining the talk's five parts
  3. 4:00 Discussing oversights found during vulnerability disclosure process
  4. 6:10 Explaining the evolution from traditional DevOps to MLOps
  5. 8:15 Overview of Azure Machine Learning architecture and tooling

The Oversights Under the Flow: Discovering the Vulnerable Tooling Suites From Azure MLOps

Speakers: Name not provided, from Songhai University

Conference: Black Hat Asia

YouTube: https://www.youtube.com/watch?v=s49sgre_04c

Overview

This talk, "The Oversights Under the Flow," delves into a critical examination of security vulnerabilities discovered within the tooling suites of Azure Machine Learning Operations (MLOps). Presented by a researcher from Songhai University, the session highlights how seemingly simple and easily discoverable security flaws, often stemming from "oversights" in development and maintenance, can persist and lead to significant impacts within complex software supply chains. The speaker emphasizes that while these vulnerabilities might appear minor at first glance, their presence in core MLOps infrastructure—even tools maintained by a tech giant like Microsoft—poses substantial risks if left unaddressed or incompletely patched.

The research focuses on a selection of Azure's MLOps tools, including Prompt Flow, Azure AI Generative SDK, DeepSpeed, and others, revealing patterns of insecure coding practices that are frequently overlooked. A key takeaway is the disconnect between the intent of secure development and its actual implementation, often exacerbated by a broad community of contributors who may not prioritize security. This article will explore the technical specifics of these findings, the implications for both defenders and developers, and the challenging realities of coordinated vulnerability disclosure, as experienced by the researcher with Microsoft's Security Response Center (MSRC).

Background

▶ Watch: Introduction to MLOps vulnerabilities and recurring oversights (0:00)

The evolution of software development has seen a significant shift from traditional DevOps to MLOps, driven by the rapid advancements in machine learning (ML), artificial intelligence (AI), and large language models (LLMs). DevOps encompasses the entire software lifecycle, from planning and coding to testing, release, deployment, and continuous monitoring. Azure, Microsoft's cloud platform, offers its own suite of tools and services under Azure DevOps to support this traditional paradigm.

With the advent of sophisticated ML and AI techniques, the need to integrate machine learning model development into this established lifecycle became paramount. MLOps extends DevOps by incorporating specialized phases for ML models, including model training, testing, evaluation against larger datasets, and eventual synthesis for customer usage. This integration ensures that ML-enabled applications can be developed, deployed, and managed with the same rigor and efficiency as traditional software.

To facilitate this, Microsoft Azure provides a comprehensive ecosystem, including the Azure Machine Learning workspace, which supports development both on-premises and in the cloud. This workspace is supported by a diverse array of client-side and cloud-based tooling suites, numbering over 20 distinct tools. The speaker focused their research on six key tools from this extensive stack—ranging from traditional Azure clients and DevOps tools to cutting-edge solutions like Prompt Flow and Azure AI Generative SDK—because these were found to exhibit recurring patterns of security vulnerabilities due to developer oversights.

Key Findings

▶ Watch: Comprehensive agenda outlining the talk's five parts (2:15)

The central theme of the talk revolves around the prevalence of "oversights" that lead to simple yet impactful vulnerabilities within Azure's MLOps tooling. The speaker identified several classes of vulnerabilities, including command injection, path traversal, unsafe deserialization (specifically Python pickle deserialization), and insecure use of eval() functions. A consistent observation was that secure alternatives or practices often existed within the broader codebase or were documented by Microsoft, but were either overlooked by contributors or not enforced during code reviews.

A significant finding was how seemingly "tiny" or "local" vulnerabilities could be escalated. For instance, a local path traversal vulnerability could be weaponized into a remote attack through specific deployment configurations or by leveraging client-side attack techniques (e.g., JavaScript payloads for "one-click remote" exploitation). This highlights the critical importance of defensive depth and anticipating how attackers might chain vulnerabilities.

The speaker also shared critical insights into the coordinated vulnerability disclosure (CVD) process with Microsoft's Security Response Center (MSRC). Several discrepancies and issues were noted:

  • Differing Severity Perspectives: The MSRC sometimes assessed vulnerabilities as low severity that the speaker considered high, and vice versa, leading to inconsistencies in patching priority and CVE issuance.
  • Incomplete Patches: In some cases, MSRC's patches were found to be incomplete, addressing only a syntactic instance of a vulnerability while similar logical flaws or other instances remained unpatched.
  • Introduction of New Vulnerabilities: Alarmingly, during the patching process for a reported vulnerability, new, similar vulnerabilities (e.g., new instances of unsafe eval()) were sometimes introduced in subsequent versions of the software.
  • Overlooked Follow-ups: Despite follow-up reports detailing unpatched issues, some vulnerabilities remained unaddressed, suggesting potential communication breakdowns or differing priorities between MSRC and engineering teams.
  • CVE Issuance Discrepancies: The MSRC occasionally issued CVEs for minor vulnerabilities in older, less critical tools (e.g., Azure Clients, TorchGeo) while declining CVEs for more severe flaws in cutting-edge MLOps tools like Azure AI Generative SDK or Prompt Flow.

These findings collectively paint a picture of an ecosystem where the rapid pace of development and the distributed nature of open-source contributions can inadvertently create security blind spots, even within a rigorously managed cloud environment.

Technical Deep Dive

▶ Watch: Discussing oversights found during vulnerability disclosure process (4:00)

The speaker detailed several specific vulnerabilities across various Azure MLOps tools, illustrating the recurring "oversight" pattern.

Prompt Flow

Prompt Flow is a core feature within the Azure Machine Learning Studio, designed for building high-quality LLM applications from prototyping to production.

  1. Command Injection (CVE-2023-38166):
  • Vulnerability: The subprocess.run() function was used with shell=True and passed arguments constructed by joining a list of strings using str.join(). This allowed attackers to inject arbitrary shell commands by manipulating input arguments.
  • Secure Alternatives: The codebase elsewhere demonstrated secure practices, such as specifying parameters as integer types, setting shell=False, and passing arguments as a list directly to subprocess.run(), which prevents shell interpretation.
  • Impact: If an attacker could control the Prompt Flow client command-line input, they could achieve privilege escalation. Furthermore, if application developers used the vulnerable start_experiment function in their Azure applications and exposed it via REST APIs, this local command injection could become a remote vulnerability, impacting the experimentation phase of MLOps.
  1. Path Traversal (CVE-2023-38165):
  • Vulnerability: File paths were constructed using simple string concatenation (directory + filename + extension) instead of secure functions like os.path.join(). This allowed an attacker to inject ../ sequences into the filename, traversing directories to write files to arbitrary locations on the file system. The speaker noted that maintainers understood os.path.join() but overlooked its application in this specific context, potentially due to assumptions about the filename being a hash object or the directory being sufficiently controlled.
  • Impact: Attackers could write malicious files (e.g., dynamic link libraries like malicious.dll) into critical system directories (e.g., C:\Windows\System32) on the victim's machine.
  • Local vs. Remote Debate: By default, Prompt Flow services often listen only on localhost (127.0.0.1). However, customer requests on GitHub led to the implementation of functionality allowing the service to listen on 0.0.0.0 (all interfaces), making it remotely accessible. Even if restricted to localhost, the speaker highlighted the "one-click remote" attack vector, where a malicious JavaScript payload hosted on an attacker-controlled website could trick a developer into triggering the file upload, effectively turning a local vulnerability into a remote one without direct user interaction beyond clicking a link.

Azure AI Generative SDK

This client-side SDK is designed for building, evaluating, and deploying generative AI applications leveraging Azure AI services.

  • Unsafe eval() (CVE-2023-38167):
  • Vulnerability: The SDK used the unsafe eval() function directly to evaluate stop tokens within the OpenAICompletionModel function. This allows for arbitrary code execution if an attacker can control the input to eval().
  • Secure Alternatives: The codebase frequently used ast.literal_eval(), a safer alternative, in other contexts, indicating an awareness of the risk. The oversight was in this specific instance.
  • Impact: If developers used this Python SDK in their Azure applications and exposed the OpenAICompletionModel function via API endpoints, it could lead to remote code execution.
  • Patching Oversight: The speaker noted a critical observation: after reporting the initial eval() vulnerability in version B7, Microsoft patched it in B9. However, new instances of unsafe eval() were introduced in version B8, meaning the fix was "syntactic" rather than a comprehensive review of the code for similar logical flaws.

DeepSpeed

DeepSpeed is an optimization library primarily used for distributed model training, enhancing efficiency for large-scale ML.

  • Pickle Deserialization RCE (CVE-2023-38164):
  • Vulnerability: DeepSpeed's distributed training mechanism, which involves communication between different GPUs and CPUs (referred to as "rankers"), deserialized data using Python's pickle module without adequate security measures. Python's pickle module is notoriously unsafe for untrusted input, as deserializing a malicious pickle string can lead to arbitrary code execution.
  • Context: The speaker pointed out that PyTorch, which DeepSpeed inherits from, explicitly acknowledges the security risks of pickle in its documentation but prioritizes functionality. DeepSpeed's documentation, however, failed to convey this warning to its users. The communication often happens over sockets via libraries like NCCL or MPI.
  • Attack Scenario: In a distributed training environment, where rankers communicate intensively, an attacker could exploit latency gaps. By masquerading as a legitimate ranker (e.g., rank one) and communicating with the master ranker (rank zero), the attacker could send a malicious pickle payload, achieving Remote Code Execution (RCE) on the rank zero device.
  • Local vs. Remote (again): Similar to Prompt Flow, MSRC initially assessed this as low severity, assuming local host listening. However, if developers configure DeepSpeed to listen on public IP addresses (e.g., by exporting GLOO_SOCKET_IFNAME to a public network card), the RCE becomes truly remote. The "one-click remote" trick was also highlighted as a potential escalation path.

TorchGeo

TorchGeo is a library for geospatial data processing, potentially used in ML workflows.

  • Unsafe eval() (CVE-2023-38168):
  • Vulnerability: This instance of unsafe eval() was particularly interesting as TorchGeo maintainers claimed to have copied the code directly from TorchVision. However, while TorchVision later patched the vulnerability, TorchGeo failed to update its copied code, leading to the continued existence of the flaw. This highlights the risks of code reuse without continuous synchronization and security vigilance.

Azure Clients and Azure DevOps

These are client-side software tools for interacting with Azure services.

  • Multiple Command Injections (CVE-2023-38169):
  • Vulnerability: Despite Azure Clients having a documented run_command API specifically designed for safely calling system commands (which wraps subprocess with security in mind), contributors continued to use the unsafe subprocess module directly, leading to multiple command injection vulnerabilities. This demonstrates a clear oversight in adhering to internal secure coding guidelines.
  • Patching Oversight: Even after a report on subprocess issues, Microsoft patched only one instance (in the "service connector") but ignored other instances within the "security" component, despite follow-up emails from the speaker.

Demo / Proof of Concept

▶ Watch: Explaining the evolution from traditional DevOps to MLOps (6:10)

The speaker demonstrated the practical exploitability of these vulnerabilities through several Proof of Concept (PoC) scenarios, illustrating how simple inputs could lead to significant compromise.

For the Prompt Flow command injection, the PoC involved simulating an attacker running a malicious command via the pf prompt flow utility in a terminal. The speaker also depicted a vulnerable flow within the Azure environment where user input, if exposed to adversaries, could be injected with malicious commands to trigger the attack.

The Path Traversal vulnerability in Prompt Flow was visually demonstrated through a GIF. This PoC showed an attacker sending an HTTP request to the vulnerable Prompt Flow service. By manipulating the filename parameter, the attacker successfully wrote a malicious dynamic link library (.dll) to the C:\Windows\System32 directory on the victim's local machine, showcasing arbitrary file write capabilities.

For the DeepSpeed Pickle Deserialization RCE, the speaker presented a PoC mimicking a distributed training setup. This involved a "rank zero" victim and an "attacker" masquerading as "rank one." The attacker exploited the communication channel to send a crafted malicious pickle payload, resulting in remote code execution on the rank zero device, proving the viability of the attack against critical ML infrastructure.

While the specific PoCs were often simplified for demonstration purposes (e.g., assuming local execution), the speaker consistently contextualized them by explaining how these "local" vulnerabilities could be escalated to "remote" scenarios through various means, such as public exposure of services or client-side JavaScript-based attacks.

Defensive Implications

▶ Watch: Overview of Azure Machine Learning architecture and tooling (8:15)

The findings from "The Oversights Under the Flow" present critical defensive implications for various stakeholders involved in the MLOps ecosystem.

For Open Source Maintainers and Contributors (including Microsoft's internal teams):

  • Stricter Security Checks: Implement rigorous security reviews for all new code contributions and merge requests, especially for open-source projects with a broad contributor base. This requires dedicated security expertise during the review process.
  • Consistent Secure Coding Practices: Enforce the consistent application of secure coding patterns across the entire codebase. If a secure function (e.g., os.path.join(), ast.literal_eval(), or a custom safe subprocess wrapper like run_command) exists, ensure all contributors use it, rather than insecure alternatives. Automated static analysis tools configured with project-specific rules can help enforce this.
  • Continuous Security Training: Provide ongoing security awareness and secure coding training for all developers, particularly those contributing to critical infrastructure components. Many developers, while proficient in ML, may lack deep security expertise.
  • Dependency Management and Updates: Regularly audit and update third-party dependencies. The TorchGeo case highlights the risk of copying code without subsequent synchronization with upstream security patches.

For Microsoft Security Response Center (MSRC) and Vendor Security Teams:

  • Improved Coordinated Disclosure: Enhance communication clarity and efficiency during the CVD process. This includes consistent severity assessments, comprehensive patching strategies that address all instances of a vulnerability (not just syntactic ones), and diligent follow-up on reported unpatched issues.
  • Proactive Pattern Scanning: After patching a vulnerability (e.g., unsafe eval()), proactively scan the entire codebase for similar patterns to prevent the reintroduction or oversight of new instances during development cycles.
  • Consistent CVE Issuance: Standardize the criteria for CVE issuance, ensuring that critical vulnerabilities in cutting-edge tools receive appropriate recognition and visibility, irrespective of their perceived "local" nature or the tool's age.

For Azure Application Developers and ML Engineers:

  • Input Validation and Sanitization: Treat all user-supplied input, and indeed all external data, as untrusted. Implement robust validation and sanitization at all input points to prevent command injection, path traversal, and other injection attacks.
  • Secure API Usage: Exercise extreme caution when using SDKs and APIs, even those from reputable vendors. Understand the security implications of functions like eval() or pickle.load() and avoid using them with untrusted data.
  • Understand Service Exposure: Be aware of how services are exposed. A default localhost binding might seem safe, but configurations can change, and client-side attacks (like the "one-click remote" via JavaScript) can bridge the local-to-remote gap. Avoid exposing sensitive services to public networks unless absolutely necessary and with robust security controls.
  • Regular Security Audits: Conduct regular security audits of custom Azure applications and ML pipelines, focusing on potential vulnerabilities introduced through custom code or insecure API usage.

Future Directions and the Role of LLMs:

The speaker briefly touched upon the potential of Large Language Models (LLMs) in assisting with code auditing to mitigate human oversights. While current LLMs like GPT-40 might not provide entirely accurate patch suggestions or comprehensive vulnerability detection, fine-tuning them with security-specific data (e.g., using Retrieval-Augmented Generation (RAG) or specialized training) could significantly enhance their capability to identify and suggest fixes for common security flaws, thereby augmenting human security efforts.

Key Takeaways

  • Simple Oversights Lead to Critical Vulnerabilities: Even seemingly minor coding errors, like using string concatenation for paths or shell=True without proper sanitization, can lead to severe command injection or path traversal vulnerabilities in MLOps tooling.
  • Secure Practices Are Often Overlooked: Many vulnerabilities stem from a failure to consistently apply secure coding patterns that already exist within the codebase or are documented, indicating a gap in developer awareness or enforcement during code reviews.
  • Local Vulnerabilities Can Escalate to Remote RCE: Default local host bindings do not guarantee safety. Attackers can leverage configuration changes, "one-click remote" JavaScript payloads, or specific network setups to turn local flaws into remotely exploitable RCEs.
  • Vendor Coordinated Disclosure Needs Improvement: The MSRC's response highlighted issues such as inconsistent severity assessments, incomplete patches, the introduction of new vulnerabilities during patching, and overlooked follow-ups, underscoring the need for more robust CVD processes.
  • Developers Must Be Vigilant: Application developers using MLOps SDKs and APIs, even from major vendors, must exercise extreme caution, validate all inputs, and understand the security implications of functions that process untrusted data.
  • LLMs Show Promise for Auditing: Large Language Models could potentially assist in code auditing to reduce human oversight, though current capabilities require further refinement and specialized training for accurate and actionable security insights.

About the Speaker(s)

The presenter for "The Oversights Under the Flow" is a researcher from Songhai University. They have a strong interest in auditing open-source software, particularly focusing on machine learning operations (MLOps) tooling. Their research, as demonstrated in this talk, involves identifying and analyzing vulnerabilities within these systems. Additional details about their background and research can be found on their GitHub homepage.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research delivers a brutal, yet essential, exposé of persistent "oversights" leading to critical vulnerabilities (command injection, path traversal, RCE via pickle/eval) across multiple Azure MLOps tools like Prompt Flow and DeepSpeed. It expertly details how seemingly local flaws can escalate to remote compromise and provides a candid, damning account of Microsoft's inconsistent and often incomplete vulnerability disclosure and patching process. This isn't just a list of CVEs; it's a stark lesson in the realities of supply chain security, vendor responsibility, and the urgent need for better secure development practices in the AI/ML space.

Heather Calloway (CISO) — MUST SEE

This talk is a critical examination of pervasive security oversights within Azure's MLOps tooling, exposing not just technical vulnerabilities but profound institutional and process failures at a major vendor. It meticulously details how basic secure coding practices are neglected, leading to critical remote code execution risks in core machine learning infrastructure. More importantly, it provides a stark, unsentimental critique of the coordinated vulnerability disclosure process with Microsoft, highlighting inconsistent severity assessments, incomplete patches, and the alarming reintroduction of vulnerabilities. Every CISO and security leader managing cloud-native ML/AI initiatives must…

→ Top-rated talks at Black Hat Asia 2025

All talks from Black Hat Asia 2025