The Oversights Under the Flow: Discovering the Vulnerable Tooling Suites From Azure MLOps
Black Hat Asia 2025 · Day 2 · Briefings
Overview
This talk, "The Oversights Under the Flow," delves into a critical examination of security vulnerabilities discovered within the tooling suites of Azure Machine Learning Operations (MLOps). Presented by a researcher from Songhai University, the session highlights how seemingly simple and easily discoverable security flaws, often stemming from "oversights" in development and maintenance, can persist and lead to significant impacts within complex software supply chains. The speaker emphasizes that while these vulnerabilities might appear minor at first glance, their presence in core MLOps infrastructure—even tools maintained by a tech giant like Microsoft—poses substantial risks if left unaddressed or incompletely patched.

Key moments
- 0:00 Introduction to MLOps vulnerabilities and recurring oversights
- 2:15 Comprehensive agenda outlining the talk's five parts
- 4:00 Discussing oversights found during vulnerability disclosure process
- 6:10 Explaining the evolution from traditional DevOps to MLOps
- 8:15 Overview of Azure Machine Learning architecture and tooling
The Oversights Under the Flow: Discovering the Vulnerable Tooling Suites From Azure MLOps
Speakers: Name not provided, from Songhai University
Conference: Black Hat Asia
YouTube: https://www.youtube.com/watch?v=s49sgre_04c
Overview
This talk, "The Oversights Under the Flow," delves into a critical examination of security vulnerabilities discovered within the tooling suites of Azure Machine Learning Operations (MLOps). Presented by a researcher from Songhai University, the session highlights how seemingly simple and easily discoverable security flaws, often stemming from "oversights" in development and maintenance, can persist and lead to significant impacts within complex software supply chains. The speaker emphasizes that while these vulnerabilities might appear minor at first glance, their presence in core MLOps infrastructure—even tools maintained by a tech giant like Microsoft—poses substantial risks if left unaddressed or incompletely patched.
The research focuses on a selection of Azure's MLOps tools, including Prompt Flow, Azure AI Generative SDK, DeepSpeed, and others, revealing patterns of insecure coding practices that are frequently overlooked. A key takeaway is the disconnect between the intent of secure development and its actual implementation, often exacerbated by a broad community of contributors who may not prioritize security. This article will explore the technical specifics of these findings, the implications for both defenders and developers, and the challenging realities of coordinated vulnerability disclosure, as experienced by the researcher with Microsoft's Security Response Center (MSRC).
Background
▶ Watch: Introduction to MLOps vulnerabilities and recurring oversights (0:00)
The evolution of software development has seen a significant shift from traditional DevOps to MLOps, driven by the rapid advancements in machine learning (ML), artificial intelligence (AI), and large language models (LLMs). DevOps encompasses the entire software lifecycle, from planning and coding to testing, release, deployment, and continuous monitoring. Azure, Microsoft's cloud platform, offers its own suite of tools and services under Azure DevOps to support this traditional paradigm.
With the advent of sophisticated ML and AI techniques, the need to integrate machine learning model development into this established lifecycle became paramount. MLOps extends DevOps by incorporating specialized phases for ML models, including model training, testing, evaluation against larger datasets, and eventual synthesis for customer usage. This integration ensures that ML-enabled applications can be developed, deployed, and managed with the same rigor and efficiency as traditional software.
To facilitate this, Microsoft Azure provides a comprehensive ecosystem, including the Azure Machine Learning workspace, which supports development both on-premises and in the cloud. This workspace is supported by a diverse array of client-side and cloud-based tooling suites, numbering over 20 distinct tools. The speaker focused their research on six key tools from this extensive stack—ranging from traditional Azure clients and DevOps tools to cutting-edge solutions like Prompt Flow and Azure AI Generative SDK—because these were found to exhibit recurring patterns of security vulnerabilities due to developer oversights.
Key Findings
▶ Watch: Comprehensive agenda outlining the talk's five parts (2:15)
The central theme of the talk revolves around the prevalence of "oversights" that lead to simple yet impactful vulnerabilities within Azure's MLOps tooling. The speaker identified several classes of vulnerabilities, including command injection, path traversal, unsafe deserialization (specifically Python pickle deserialization), and insecure use of eval() functions. A consistent observation was that secure alternatives or practices often existed within the broader codebase or were documented by Microsoft, but were either overlooked by contributors or not enforced during code reviews.
A significant finding was how seemingly "tiny" or "local" vulnerabilities could be escalated. For instance, a local path traversal vulnerability could be weaponized into a remote attack through specific deployment configurations or by leveraging client-side attack techniques (e.g., JavaScript payloads for "one-click remote" exploitation). This highlights the critical importance of defensive depth and anticipating how attackers might chain vulnerabilities.
The speaker also shared critical insights into the coordinated vulnerability disclosure (CVD) process with Microsoft's Security Response Center (MSRC). Several discrepancies and issues were noted:
- Differing Severity Perspectives: The MSRC sometimes assessed vulnerabilities as low severity that the speaker considered high, and vice versa, leading to inconsistencies in patching priority and CVE issuance.
- Incomplete Patches: In some cases, MSRC's patches were found to be incomplete, addressing only a syntactic instance of a vulnerability while similar logical flaws or other instances remained unpatched.
- Introduction of New Vulnerabilities: Alarmingly, during the patching process for a reported vulnerability, new, similar vulnerabilities (e.g., new instances of unsafe
eval()) were sometimes introduced in subsequent versions of the software. - Overlooked Follow-ups: Despite follow-up reports detailing unpatched issues, some vulnerabilities remained unaddressed, suggesting potential communication breakdowns or differing priorities between MSRC and engineering teams.
- CVE Issuance Discrepancies: The MSRC occasionally issued CVEs for minor vulnerabilities in older, less critical tools (e.g., Azure Clients, TorchGeo) while declining CVEs for more severe flaws in cutting-edge MLOps tools like Azure AI Generative SDK or Prompt Flow.
These findings collectively paint a picture of an ecosystem where the rapid pace of development and the distributed nature of open-source contributions can inadvertently create security blind spots, even within a rigorously managed cloud environment.
Technical Deep Dive
▶ Watch: Discussing oversights found during vulnerability disclosure process (4:00)
The speaker detailed several specific vulnerabilities across various Azure MLOps tools, illustrating the recurring "oversight" pattern.
Prompt Flow
Prompt Flow is a core feature within the Azure Machine Learning Studio, designed for building high-quality LLM applications from prototyping to production.
- Command Injection (CVE-2023-38166):
- Vulnerability: The
subprocess.run()function was used withshell=Trueand passed arguments constructed by joining a list of strings usingstr.join(). This allowed attackers to inject arbitrary shell commands by manipulating input arguments. - Secure Alternatives: The codebase elsewhere demonstrated secure practices, such as specifying parameters as integer types, setting
shell=False, and passing arguments as a list directly tosubprocess.run(), which prevents shell interpretation. - Impact: If an attacker could control the Prompt Flow client command-line input, they could achieve privilege escalation. Furthermore, if application developers used the vulnerable
start_experimentfunction in their Azure applications and exposed it via REST APIs, this local command injection could become a remote vulnerability, impacting the experimentation phase of MLOps.
- Path Traversal (CVE-2023-38165):
- Vulnerability: File paths were constructed using simple string concatenation (
directory + filename + extension) instead of secure functions likeos.path.join(). This allowed an attacker to inject../sequences into the filename, traversing directories to write files to arbitrary locations on the file system. The speaker noted that maintainers understoodos.path.join()but overlooked its application in this specific context, potentially due to assumptions about thefilenamebeing a hash object or thedirectorybeing sufficiently controlled. - Impact: Attackers could write malicious files (e.g., dynamic link libraries like
malicious.dll) into critical system directories (e.g.,C:\Windows\System32) on the victim's machine. - Local vs. Remote Debate: By default, Prompt Flow services often listen only on
localhost(127.0.0.1). However, customer requests on GitHub led to the implementation of functionality allowing the service to listen on0.0.0.0(all interfaces), making it remotely accessible. Even if restricted to localhost, the speaker highlighted the "one-click remote" attack vector, where a malicious JavaScript payload hosted on an attacker-controlled website could trick a developer into triggering the file upload, effectively turning a local vulnerability into a remote one without direct user interaction beyond clicking a link.
Azure AI Generative SDK
This client-side SDK is designed for building, evaluating, and deploying generative AI applications leveraging Azure AI services.
- Unsafe
eval()(CVE-2023-38167): - Vulnerability: The SDK used the unsafe
eval()function directly to evaluate stop tokens within theOpenAICompletionModelfunction. This allows for arbitrary code execution if an attacker can control the input toeval(). - Secure Alternatives: The codebase frequently used
ast.literal_eval(), a safer alternative, in other contexts, indicating an awareness of the risk. The oversight was in this specific instance. - Impact: If developers used this Python SDK in their Azure applications and exposed the
OpenAICompletionModelfunction via API endpoints, it could lead to remote code execution. - Patching Oversight: The speaker noted a critical observation: after reporting the initial
eval()vulnerability in version B7, Microsoft patched it in B9. However, new instances of unsafeeval()were introduced in version B8, meaning the fix was "syntactic" rather than a comprehensive review of the code for similar logical flaws.
DeepSpeed
DeepSpeed is an optimization library primarily used for distributed model training, enhancing efficiency for large-scale ML.
- Pickle Deserialization RCE (CVE-2023-38164):
- Vulnerability: DeepSpeed's distributed training mechanism, which involves communication between different GPUs and CPUs (referred to as "rankers"), deserialized data using Python's
picklemodule without adequate security measures. Python'spicklemodule is notoriously unsafe for untrusted input, as deserializing a malicious pickle string can lead to arbitrary code execution. - Context: The speaker pointed out that PyTorch, which DeepSpeed inherits from, explicitly acknowledges the security risks of
picklein its documentation but prioritizes functionality. DeepSpeed's documentation, however, failed to convey this warning to its users. The communication often happens over sockets via libraries like NCCL or MPI. - Attack Scenario: In a distributed training environment, where rankers communicate intensively, an attacker could exploit latency gaps. By masquerading as a legitimate ranker (e.g., rank one) and communicating with the master ranker (rank zero), the attacker could send a malicious pickle payload, achieving Remote Code Execution (RCE) on the rank zero device.
- Local vs. Remote (again): Similar to Prompt Flow, MSRC initially assessed this as low severity, assuming local host listening. However, if developers configure DeepSpeed to listen on public IP addresses (e.g., by exporting
GLOO_SOCKET_IFNAMEto a public network card), the RCE becomes truly remote. The "one-click remote" trick was also highlighted as a potential escalation path.
TorchGeo
TorchGeo is a library for geospatial data processing, potentially used in ML workflows.
- Unsafe
eval()(CVE-2023-38168): - Vulnerability: This instance of unsafe
eval()was particularly interesting as TorchGeo maintainers claimed to have copied the code directly from TorchVision. However, while TorchVision later patched the vulnerability, TorchGeo failed to update its copied code, leading to the continued existence of the flaw. This highlights the risks of code reuse without continuous synchronization and security vigilance.
Azure Clients and Azure DevOps
These are client-side software tools for interacting with Azure services.
- Multiple Command Injections (CVE-2023-38169):
- Vulnerability: Despite Azure Clients having a documented
run_commandAPI specifically designed for safely calling system commands (which wrapssubprocesswith security in mind), contributors continued to use the unsafesubprocessmodule directly, leading to multiple command injection vulnerabilities. This demonstrates a clear oversight in adhering to internal secure coding guidelines. - Patching Oversight: Even after a report on subprocess issues, Microsoft patched only one instance (in the "service connector") but ignored other instances within the "security" component, despite follow-up emails from the speaker.
Demo / Proof of Concept
▶ Watch: Explaining the evolution from traditional DevOps to MLOps (6:10)
The speaker demonstrated the practical exploitability of these vulnerabilities through several Proof of Concept (PoC) scenarios, illustrating how simple inputs could lead to significant compromise.
For the Prompt Flow command injection, the PoC involved simulating an attacker running a malicious command via the pf prompt flow utility in a terminal. The speaker also depicted a vulnerable flow within the Azure environment where user input, if exposed to adversaries, could be injected with malicious commands to trigger the attack.
The Path Traversal vulnerability in Prompt Flow was visually demonstrated through a GIF. This PoC showed an attacker sending an HTTP request to the vulnerable Prompt Flow service. By manipulating the filename parameter, the attacker successfully wrote a malicious dynamic link library (.dll) to the C:\Windows\System32 directory on the victim's local machine, showcasing arbitrary file write capabilities.
For the DeepSpeed Pickle Deserialization RCE, the speaker presented a PoC mimicking a distributed training setup. This involved a "rank zero" victim and an "attacker" masquerading as "rank one." The attacker exploited the communication channel to send a crafted malicious pickle payload, resulting in remote code execution on the rank zero device, proving the viability of the attack against critical ML infrastructure.
While the specific PoCs were often simplified for demonstration purposes (e.g., assuming local execution), the speaker consistently contextualized them by explaining how these "local" vulnerabilities could be escalated to "remote" scenarios through various means, such as public exposure of services or client-side JavaScript-based attacks.
Defensive Implications
▶ Watch: Overview of Azure Machine Learning architecture and tooling (8:15)
The findings from "The Oversights Under the Flow" present critical defensive implications for various stakeholders involved in the MLOps ecosystem.
For Open Source Maintainers and Contributors (including Microsoft's internal teams):
- Stricter Security Checks: Implement rigorous security reviews for all new code contributions and merge requests, especially for open-source projects with a broad contributor base. This requires dedicated security expertise during the review process.
- Consistent Secure Coding Practices: Enforce the consistent application of secure coding patterns across the entire codebase. If a secure function (e.g.,
os.path.join(),ast.literal_eval(), or a custom safesubprocesswrapper likerun_command) exists, ensure all contributors use it, rather than insecure alternatives. Automated static analysis tools configured with project-specific rules can help enforce this. - Continuous Security Training: Provide ongoing security awareness and secure coding training for all developers, particularly those contributing to critical infrastructure components. Many developers, while proficient in ML, may lack deep security expertise.
- Dependency Management and Updates: Regularly audit and update third-party dependencies. The TorchGeo case highlights the risk of copying code without subsequent synchronization with upstream security patches.
For Microsoft Security Response Center (MSRC) and Vendor Security Teams:
- Improved Coordinated Disclosure: Enhance communication clarity and efficiency during the CVD process. This includes consistent severity assessments, comprehensive patching strategies that address all instances of a vulnerability (not just syntactic ones), and diligent follow-up on reported unpatched issues.
- Proactive Pattern Scanning: After patching a vulnerability (e.g., unsafe
eval()), proactively scan the entire codebase for similar patterns to prevent the reintroduction or oversight of new instances during development cycles. - Consistent CVE Issuance: Standardize the criteria for CVE issuance, ensuring that critical vulnerabilities in cutting-edge tools receive appropriate recognition and visibility, irrespective of their perceived "local" nature or the tool's age.
For Azure Application Developers and ML Engineers:
- Input Validation and Sanitization: Treat all user-supplied input, and indeed all external data, as untrusted. Implement robust validation and sanitization at all input points to prevent command injection, path traversal, and other injection attacks.
- Secure API Usage: Exercise extreme caution when using SDKs and APIs, even those from reputable vendors. Understand the security implications of functions like
eval()orpickle.load()and avoid using them with untrusted data. - Understand Service Exposure: Be aware of how services are exposed. A default
localhostbinding might seem safe, but configurations can change, and client-side attacks (like the "one-click remote" via JavaScript) can bridge the local-to-remote gap. Avoid exposing sensitive services to public networks unless absolutely necessary and with robust security controls. - Regular Security Audits: Conduct regular security audits of custom Azure applications and ML pipelines, focusing on potential vulnerabilities introduced through custom code or insecure API usage.
Future Directions and the Role of LLMs:
The speaker briefly touched upon the potential of Large Language Models (LLMs) in assisting with code auditing to mitigate human oversights. While current LLMs like GPT-40 might not provide entirely accurate patch suggestions or comprehensive vulnerability detection, fine-tuning them with security-specific data (e.g., using Retrieval-Augmented Generation (RAG) or specialized training) could significantly enhance their capability to identify and suggest fixes for common security flaws, thereby augmenting human security efforts.
Key Takeaways
- Simple Oversights Lead to Critical Vulnerabilities: Even seemingly minor coding errors, like using string concatenation for paths or
shell=Truewithout proper sanitization, can lead to severe command injection or path traversal vulnerabilities in MLOps tooling. - Secure Practices Are Often Overlooked: Many vulnerabilities stem from a failure to consistently apply secure coding patterns that already exist within the codebase or are documented, indicating a gap in developer awareness or enforcement during code reviews.
- Local Vulnerabilities Can Escalate to Remote RCE: Default local host bindings do not guarantee safety. Attackers can leverage configuration changes, "one-click remote" JavaScript payloads, or specific network setups to turn local flaws into remotely exploitable RCEs.
- Vendor Coordinated Disclosure Needs Improvement: The MSRC's response highlighted issues such as inconsistent severity assessments, incomplete patches, the introduction of new vulnerabilities during patching, and overlooked follow-ups, underscoring the need for more robust CVD processes.
- Developers Must Be Vigilant: Application developers using MLOps SDKs and APIs, even from major vendors, must exercise extreme caution, validate all inputs, and understand the security implications of functions that process untrusted data.
- LLMs Show Promise for Auditing: Large Language Models could potentially assist in code auditing to reduce human oversight, though current capabilities require further refinement and specialized training for accurate and actionable security insights.
About the Speaker(s)
The presenter for "The Oversights Under the Flow" is a researcher from Songhai University. They have a strong interest in auditing open-source software, particularly focusing on machine learning operations (MLOps) tooling. Their research, as demonstrated in this talk, involves identifying and analyzing vulnerabilities within these systems. Additional details about their background and research can be found on their GitHub homepage.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research delivers a brutal, yet essential, exposé of persistent "oversights" leading to critical vulnerabilities (command injection, path traversal, RCE via pickle/eval) across multiple Azure MLOps tools like Prompt Flow and DeepSpeed. It expertly details how seemingly local flaws can escalate to remote compromise and provides a candid, damning account of Microsoft's inconsistent and often incomplete vulnerability disclosure and patching process. This isn't just a list of CVEs; it's a stark lesson in the realities of supply chain security, vendor responsibility, and the urgent need for better secure development practices in the AI/ML space.
Heather Calloway (CISO) — MUST SEE
This talk is a critical examination of pervasive security oversights within Azure's MLOps tooling, exposing not just technical vulnerabilities but profound institutional and process failures at a major vendor. It meticulously details how basic secure coding practices are neglected, leading to critical remote code execution risks in core machine learning infrastructure. More importantly, it provides a stark, unsentimental critique of the coordinated vulnerability disclosure process with Microsoft, highlighting inconsistent severity assessments, incomplete patches, and the alarming reintroduction of vulnerabilities. Every CISO and security leader managing cloud-native ML/AI initiatives must…