Reverse Engineering Go Malware: From Manual to AI-Powered Analysis
Asher Davila (Security Researcher · Palo Alto Networks)
BSidesSF 2026 · Day 1 · AMC Theatre 14
Overview
This talk, presented by Asher Davila, a Security Researcher at Palo Alto Networks, delves into the evolving landscape of Go malware analysis, transitioning from traditional manual reverse engineering techniques to leveraging the power of AI and Large Language Models (LLMs). Davila highlights the increasing prevalence of malware written in Go (often referred to as Golang), a language that presents unique challenges for security analysts due to its compilation characteristics and runtime mechanisms. The presentation serves as a crucial guide for reverse engineers and incident responders facing these modern threats.
Key moments
- 0:48 Why Go is popular for malware development
- 2:30 Real-world increase and types of Go malware
- 3:55 Key challenges in Go malware reverse engineering
- 5:15 Drastic size and string output differences (Go vs C)
- 6:35 Go's extensive runtime functions complicate analysis
- 8:00 Locating crucial Go build information in binaries
Reverse Engineering Go Malware: From Manual to AI-Powered Analysis
Speakers: Asher Davila, Security Researcher, Palo Alto Networks
Conference: BSides SF
YouTube: https://www.youtube.com/watch?v=JbVWcdEcRSc
Overview
This talk, presented by Asher Davila, a Security Researcher at Palo Alto Networks, delves into the evolving landscape of Go malware analysis, transitioning from traditional manual reverse engineering techniques to leveraging the power of AI and Large Language Models (LLMs). Davila highlights the increasing prevalence of malware written in Go (often referred to as Golang), a language that presents unique challenges for security analysts due to its compilation characteristics and runtime mechanisms. The presentation serves as a crucial guide for reverse engineers and incident responders facing these modern threats.
The talk underscores the critical importance of adapting analysis methodologies as adversaries increasingly adopt sophisticated languages like Go. Davila meticulously outlines the inherent difficulties in dissecting Go binaries, such as their often-large size and distinct string handling, which can confound conventional tools and techniques. By demonstrating how AI-powered tools and LLMs can augment and accelerate the analysis process, he provides actionable insights into overcoming these obstacles, ultimately enhancing the efficiency and depth of malware investigations.
Davila's presentation is particularly relevant in today's cybersecurity climate, where the sheer volume and complexity of malware demand innovative solutions. The integration of AI into reverse engineering workflows offers a promising path forward, enabling analysts to quickly triage samples, identify critical functionalities, and even pinpoint vulnerabilities that might otherwise be obscured. However, the talk also emphasizes the indispensable role of human expertise, cautioning against blind trust in AI outputs and advocating for a hybrid approach that combines automated assistance with rigorous manual verification.
Background
▶ Watch: Why Go is popular for malware development (0:48)
The adoption of Go for malware development has seen a significant surge in recent years, primarily due to several inherent advantages the language offers to threat actors. One of the most compelling features is its cross-compilation capability, allowing developers to target multiple operating systems (Windows, Linux, macOS) from a single codebase, thereby expanding the reach of their malicious programs without significant re-engineering. Furthermore, Go binaries are typically statically linked by default, meaning they bundle all necessary dependencies within the executable itself. While this simplifies distribution, it results in very large binary files—a "Hello World" program in Go can be 2.2MB compared to 16KB in C. This substantial size can pose a challenge for traditional Endpoint Detection and Response (EDR) solutions and antivirus software, which may have limitations on the static analysis of large files.
However, it's important to note that not all Go binaries are purely statically linked. When Go code imports C libraries via Sego (the Go mechanism for C interoperability), such as those used for network operations (net) or file system access (io/fs), the resulting binary will be dynamically linked to those specific C libraries. This nuance means reverse engineers cannot always assume a fully static binary.
Palo Alto Networks' telemetry has observed a consistent increase in Go malware from 2024 to 2025, encompassing a wide array of threats including IoT botnets, ransomware, and even the first known ICS malware targeting Modbus, which was developed in Go and aimed at disrupting critical infrastructure in Ukraine. Another example cited is GoEncrypt, used to encrypt and decrypt JSON configuration files for malware, making it harder for EDRs or firewalls to spot sensitive data like IP addresses.
Analyzing Go binaries presents several distinct challenges beyond file size. The function count in Go binaries is disproportionately high compared to C/C++ counterparts, even for simple programs, due to the inclusion of numerous runtime functions required for initialization. This bloat makes navigating the disassembly more complex. Moreover, existing reverse engineering tooling is predominantly optimized for C/C++ binaries, often struggling to correctly parse, open, or display information for Go executables without specialized plugins or scripts. A significant hurdle is string handling: unlike the null-terminated strings common in C/C++, Go strings are represented as a structure containing two elements: the string data itself and its length. This requires specific parsing logic to correctly identify where a string begins and ends, making the standard strings command often produce an unreadable "blob" of concatenated data. Traditional tools like Radare2, IDA Pro, Binary Ninja, and Ghidra can be used, but often necessitate custom scripts or plugins for effective Go analysis.
Key Findings
▶ Watch: Key challenges in Go malware reverse engineering (3:55)
Asher Davila's talk highlights several key findings regarding the unique characteristics of Go binaries and the emerging capabilities and limitations of AI-assisted analysis:
- Go Binary Characteristics: Go binaries are significantly larger than comparable C/C++ programs (e.g., 2.2MB vs. 16KB for "Hello World"). This size discrepancy is primarily due to Go's runtime and default static linking. The number of functions is also vastly higher, even for simple programs, necessitating a different approach to function identification.
- Distinct String Handling: Go strings are not null-terminated but are instead a structure comprising a pointer to the string data and its length. This makes traditional string extraction tools less effective, often producing a "blob" of concatenated strings that are difficult to parse without specialized methods.
- Crucial Go-Specific Sections:
- The
go build infosection (present in ELF and Mach-O binaries, embedded in metadata for PE files) provides valuable metadata like included modules, toolchain, target architecture, and operating system. The magic number within this section varies by Go version, requiring careful parsing. - The
go PC PC line table(Go Program Counter to Line Number table), introduced in Go 1.2, is critical for recovering function names and symbols, mapping machine addresses back to human-readable code. This table also features a version-dependent magic number and requires calculating an offset to locate string data. Davila open-sourced a Radare2 script to parse this table across different Go versions. - Specialized Go Tooling: Tools like GoStrings (developed by the NCC Group) are essential for accurately recovering Go strings by following a specific sequence: removing predefined strings, detecting known library strings, then recovering static and dynamic strings. IDA Pro 9.3+ has also significantly improved its Go decompilation capabilities.
- AI/LLM Capabilities:
- R2 AI (a Radare2 plugin) acts as a "Clippy-like" assistant, suggesting commands and saving time spent on documentation.
- DK (another Radare2 plugin) offers LLM-based compilation, generating pseudocode from assembly instructions in various languages.
- Multi-Tool Agents (MCPs), such as Cloud Desktop, OpenCode, and Sidekick (for Binary Ninja), act as a bridge, allowing LLMs to interact with different reverse engineering tools through a single API.
- LLMs are effective at summarizing code, suggesting commands, renaming functions/variables, and identifying potential vulnerabilities if guided correctly.
- LLM Limitations and Biases:
- Lazy Identification: LLMs often identify Go binaries based on common strings like "go panic" rather than analyzing magic numbers or structural elements, unless explicitly instructed.
- Incomplete Analysis: They can miss subtle obfuscation or persistence techniques (e.g.,
mySQIvs.MySQLdisguise). - Encryption Algorithm Confusion: LLMs frequently misidentify encryption algorithms, often pointing to algorithms merely mentioned in strings rather than the ones actually implemented. RAG (Retrieval Augmented Generation) is suggested to improve this.
- Unpacking Challenges: LLMs struggle significantly with unpacking complex or novel packed binaries.
- Data Representation: They may fail to identify IP addresses if represented in decimal or other non-string formats.
- Human-in-the-Loop is Crucial: Davila consistently emphasizes that LLMs require specific, detailed instructions and a human expert to verify their output. Blindly trusting LLM results, especially for critical tasks like CVE identification or malware triage, is risky. Providing a well-defined "plan" or "skills" (one-shot/few-shot prompting) significantly improves accuracy.
Technical Deep Dive
▶ Watch: Drastic size and string output differences (Go vs C) (5:15)
Go's design choices, while beneficial for developers, introduce significant hurdles for reverse engineers. The default static linking of Go binaries means that an executable often includes the entire Go runtime and all necessary libraries, leading to exceptionally large file sizes. For instance, a basic "Hello World" program compiled in Go can be over 2MB, whereas an equivalent C program might be a mere 16KB. This bloat inflates the function count, as the runtime introduces hundreds of internal functions, making it challenging to differentiate core malicious logic from legitimate runtime operations.
A key structural difference lies in string handling. Unlike C/C++'s null-terminated strings, Go represents strings as a struct containing two fields: a pointer to the string data and an integer representing its length. This length-prefixed structure means that simply running the strings utility on a Go binary often yields a jumbled "blob" of data, as the tool lacks the context to interpret the length field. To accurately extract strings, specialized parsing is required. Tools like GoStrings (from the NCC Group) implement a multi-stage approach: first, removing predefined Go runtime strings, then detecting known library strings, and finally recovering static and dynamically allocated strings.
Several specific sections within Go binaries provide vital information for analysis:
go build info: Present in ELF (Linux) and Mach-O (macOS) binaries as a distinct section (and embedded in metadata for PE files on Windows), this section reveals crucial compilation details. It includes the Go modules used, the toolchain, target architecture (e.g.,amd64), and operating system (e.g.,linux). A critical detail is that the magic number identifying this section changes with different Go versions, necessitating a flexible parser.stringtabandsymtab: These sections contain package paths and internal runtime function names. However, in malware, these are often stripped to hinder analysis, making them unreliable for symbol recovery.go PC PC line table: Introduced in Go 1.2, this is arguably the most valuable section for reverse engineering stripped Go binaries. It maps program counter addresses to source code line numbers, primarily used for debugging and stack traces. Crucially, it can be exploited to recover function names and other symbols, even in the absence of a full symbol table. Similar togo build info, its magic number varies by Go version. To extract function names, an analyst must calculate the actual start of the string data by adding the function name offset to the section's base address. Asher Davila has open-sourced a Python script for Radare2 that automates parsing this table across different Go versions, even scanning the entire binary for the magic number if the table isn't a distinct section.
The integration of AI and LLMs introduces a new layer of capabilities. Radare2, for instance, offers two key AI plugins:
- R2 AI: Functions as a "Clippy-like" assistant, suggesting commands based on user queries, thereby reducing the need to consult documentation. For example, asking "how to obtain IP addresses" might prompt commands for string extraction or network analysis.
- DK: This is an LLM-based decompiler that takes assembly instructions and attempts to generate human-readable pseudo code. Unlike traditional decompilers that rely on intermediate languages, DK leverages LLMs for this translation, supporting various output languages (C, JavaScript, Swift) and human languages for explanations.
Beyond single-tool plugins, Multi-Tool Agents (MCPs) are emerging as orchestrators. These agents provide a unified API for LLMs to interact with multiple reverse engineering tools like IDA Pro, Ghidra, and Radare2. Examples include Cloud Desktop, OpenCode, and Sidekick (embedded in Binary Ninja). These systems can be configured via JSON files, allowing LLMs to query and execute commands across different disassemblers. However, a recurring observation is that LLMs often adopt "lazy" analysis strategies, such as identifying a Go binary by scanning for common strings like "go panic" rather than performing a deeper analysis of magic numbers or build information, unless explicitly prompted. The effectiveness of LLMs is heavily dependent on the quality and specificity of the user's prompts, with more descriptive questions and predefined "plans" or "skills" (akin to one-shot/few-shot prompting) yielding significantly better results.
Demo / Proof of Concept
▶ Watch: Go's extensive runtime functions complicate analysis (6:35)
Asher Davila presented several practical demonstrations, showcasing both the power and the limitations of AI-assisted Go malware analysis using real-world samples.
Pumabot IoT Malware Analysis
The first example involved analyzing a Pumabot IoT malware sample. Initial inspection via VirusTotal and Radare2 confirmed it was a Go binary. The presence of the sego library in the go PC PC line table section indicated its likely use of network features for operations. When opened in Ghidra with the Golang compiler feature and subsequently processed by GoStrings, the malware's strings, such as "Pumatronics," were correctly parsed and identified.
Davila then used Sidekick, an LLM agent integrated into Binary Ninja, to query the malware.
- Generic Questions: When asked "What are the most important behaviors of this malware?", Sidekick provided high-level insights: large-scale SSH brute force attacks, fetching target IP addresses, establishing connections to C2 servers, and identifying the "Pumatronics" string as a camera brand.
- Specific Questions: Asking "What is this function printing?" or "What is this brute force function doing?" yielded much more precise answers. For instance, the LLM correctly described the brute-force function as attempting to authenticate to a target IP and port using credentials, trying up to five times. This was visually confirmed by inspecting the decompiled code, which showed a loop iterating five times.
- LLM Limitations: A critical limitation was observed when asking for the C2 domain. Sidekick identified it solely because it was present as a string in the binary, not through intelligent network analysis or understanding of socket creation. This highlights the LLM's tendency for "lazy" string-based identification. In a comparison, Radare2 combined with Gemini provided a more nuanced answer regarding "Pumatronics," correctly stating that the malware executes a
uname -acommand to confirm it's a Pumatronics camera, linking the string to a specific behavior. - Evasion Technique Misses: When asked about evasion or persistence techniques, Sidekick only identified the creation of a system service disguised as "readies." However, manual analysis revealed another subtle persistence method: a service named
mySQIdesigned to deceive defenders into thinking it was a legitimateMySQLservice, which the LLM failed to detect.
Greenblood Go Ransomware Analysis
The second demonstration involved the Greenblood Go ransomware. Davila used OpenCode, another LLM agent, employing a "one-shot" or "few-shot" prompting approach. This involved providing the LLM with a predefined set of instructions, rules, and desired output formats (referred to as "skills") to guide its analysis. The goal was to generate a comprehensive report with IOCs and artifacts.
- Initial Analysis (OpenCode, ~ $1): The LLM generated a decent key findings summary, including the malware's hash, contact information for the ransom note, general capabilities, machine fingerprinting methods, and anti-recovery techniques.
- Advanced Analysis (Oppus 4.6, ~ $5): Using a more advanced model like Oppus 4.6 provided a more complete analysis for a slightly higher cost. It identified over 140 file extensions targeted for encryption (compared to 50+ by the first model), more detailed machine fingerprinting methods (including BIOS UUID, Windows product ID, serial numbers, and hostname), and a precise Go compiler version.
- Persistent LLM Weakness: Encryption Algorithms: Both models consistently struggled with correctly identifying the encryption algorithms used. For example, one model confused ChaCha8 with AES, likely due to strings mentioning ChaCha8 even if AES was the actual implementation. Davila suggested that RAG (Retrieval Augmented Generation) could improve this by providing the LLM with a curated database of encryption patterns and examples.
- Unpacking Challenges: Davila noted that LLMs generally perform poorly at unpacking binaries, especially if the packing technique is not straightforward or well-documented.
Dynamic Analysis and Orchestrators
While the focus was static analysis, Davila briefly touched upon dynamic analysis with LLMs, mentioning MCPs for GDB and x64dbg, which allow debuggers to interact with models. He also highlighted AI-assisted malware analysis orchestrators like those based on Remnux, which can trigger over 200 tools and integrate predefined workflows. However, he cautioned that while these orchestrators provide a starting point, users should always strive to guide the model with a specific plan rather than letting the model dictate the analysis flow. He concluded by presenting a CTF challenge involving Go malware with inline assembly code for unpacking, noting that no LLM had successfully solved it, underscoring the need for human intuition and a hybrid approach.
Defensive Implications
▶ Watch: Locating crucial Go build information in binaries (8:00)
The increasing use of Go in malware necessitates a re-evaluation of defensive strategies and tooling. Defenders must be aware of the unique challenges posed by Go binaries and adapt their analysis workflows accordingly.
- Go-Aware Tooling is Essential: Traditional reverse engineering tools often struggle with Go's specific compilation characteristics, string handling, and runtime. Defenders should prioritize using or developing plugins and scripts for tools like Radare2, IDA Pro (especially 9.3+), Ghidra (with Golang compiler settings), and Binary Ninja that are specifically designed to parse Go binaries, recover symbols from the
go PC PC line table, and correctly extract length-prefixed strings using tools like GoStrings. - Strategic Integration of AI/LLMs: LLMs can serve as valuable assistants in the reverse engineering process, significantly reducing the initial analysis burden. They can be used for:
- Quick Triage: Rapidly identifying the language, basic capabilities, and potential IOCs of a malware sample.
- Command Suggestion: Assisting less experienced analysts by suggesting relevant commands in tools like Radare2 (R2 AI).
- Function Renaming and Summarization: Speeding up the often tedious process of understanding code blocks and naming functions.
- Vulnerability Research: LLMs can help reduce the attack surface by pointing out potentially interesting areas or dangerous functions for manual validation.
- Human-in-the-Loop is Non-Negotiable: The talk repeatedly emphasizes that LLMs are not infallible. Defenders must maintain a "human-in-the-loop" approach to verify all LLM outputs. This is crucial for:
- Accuracy: LLMs can be "lazy," relying on string matches rather than deep analysis, and are prone to misidentifying critical details like encryption algorithms or subtle evasion techniques (e.g.,
mySQIvs.MySQL). - Critical Validation: For tasks like CVE identification, persistence mechanism discovery, or precise algorithm identification, human expertise is indispensable.
- Contextual Understanding: Human analysts provide the critical contextual understanding necessary to interpret LLM outputs and make informed decisions, especially when dealing with novel or highly obfuscated malware.
- Effective Prompt Engineering and Planning: To maximize the utility of LLMs, defenders should invest in developing specific, detailed prompts and structured analysis plans. Providing LLMs with "skills" or "one-shot/few-shot" examples significantly improves the quality and relevance of their responses, guiding them to perform more intelligent analysis rather than generic string matching.
- Leverage LLMs for "Chunk Analysis": While LLMs may struggle with deep, novel analysis of single complex samples, they can be highly effective for "chunk analysis"—processing groups of malware samples to identify relationships, commonalities, or variants, which can be invaluable for threat intelligence and campaign tracking.
- Consider Hybrid Approaches: The most robust defensive posture will likely involve a hybrid approach that combines the speed and scalability of AI-assisted tools with the precision and critical thinking of human reverse engineers. This means using LLMs to quickly narrow down the scope of analysis, then deploying human experts for in-depth validation and discovery of highly complex or novel aspects.
- Dynamic Analysis Integration: For more complete analysis, defenders can integrate LLMs with dynamic analysis tools via MCPs (e.g., for GDB or x64dbg), allowing the debugger to interact with the models and provide runtime context.
Key Takeaways
- Go malware is a growing threat: Its cross-compilation and statically linked nature (often leading to large file sizes and high function counts) make it challenging for traditional EDR and reverse engineering tools.
- Specialized Go-aware tools are necessary: Effective analysis requires tools and plugins that can correctly parse Go's length-prefixed strings, identify crucial sections like the
go PC PC line table, and interpret Go's runtime mechanisms. - AI/LLMs offer significant assistance but are not a silver bullet: Tools like R2 AI and DK (for Radare2) and Multi-Tool Agents (MCPs) can accelerate tasks like command suggestion, pseudo code generation, and summarization, but they have distinct limitations.
- Human verification and guidance are critical: LLMs often rely on "lazy" string-based identification, can misidentify encryption algorithms, miss subtle evasion techniques, and struggle with unpacking binaries. A human expert must always be in the loop to validate results and provide specific, detailed instructions or "plans" for effective analysis.
- Hybrid analysis is the optimal approach: Combining the speed and automation of AI-assisted tools with the precision, critical thinking, and contextual understanding of manual reverse engineering yields the most comprehensive and reliable results.
- LLMs can enhance specific workflows: They are valuable for initial triage, suggesting starting points for vulnerability research (reducing the attack surface), and performing "chunk analysis" to identify relationships within groups of malware samples.
About the Speaker(s)
Asher Davila is a Security Researcher at Palo Alto Networks. His work primarily focuses on IoT and OT security research, contributing to the understanding and defense against threats targeting these critical domains. He regularly publishes technical articles and insights for Unit 42, Palo Alto Networks' global threat intelligence team. In addition to his research, Davila is an advocate for the open-source community, developing and sharing tools to aid in security analysis, and frequently presents his findings at conferences like BSides SF.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Competent survey of Go malware RE techniques bolted to an LLM capability/limitation tour. Davila clearly knows the material and the open-sourced Radare2 parser for the go PC line table is a genuine contribution, but the talk reads more like a well-organized blog post than a research drop — the individual pieces (Go binary internals, GoStrings, MCP agents) are each publicly documented elsewhere, and the AI angle adds breadth without adding depth.
Heather Calloway (CISO) — WEAK
Technically competent walkthrough of Go malware analysis with a honest treatment of LLM limitations — but this is a researcher talking to other researchers. There is no governance angle, no institutional accountability dimension, and no actionable path for the security leaders, program managers, or incident response teams who actually decide whether their org can handle a Go-based ransomware hit.