SQL Injection Isn't Dead Smuggling Queries at the Protocol Level
Paul Gerste
DEF CON 32 Main Stage · Day 1 · Main Stage
Overview
In "SQL Injection Isn't Dead: Smuggling Queries at the Protocol Level," Paul Gerste from SonarSource challenges the prevailing notion that modern application development practices have largely eradicated SQL injection vulnerabilities. While parameterized queries and Object-Relational Mappers (ORMs) have significantly mitigated traditional, high-level SQL injection risks, Gerste's research unveils a new vector: injecting malicious queries by exploiting flaws in how applications communicate with databases at the binary protocol level. This talk dives into the "lower decks" of network communication, where bits and bytes are exchanged, revealing vulnerabilities that bypass even the most robust application-layer defenses.

Key moments
- 0:00 Introduction and prepared statement SQLi teaser
- 1:47 Core research idea: binary protocol request smuggling
- 2:30 Exploring desync mechanisms in binary protocols
- 3:50 Rationale for targeting databases: high impact and applicability
- 4:30 Overview of Postgres, MySQL, Redis, MongoDB wire protocols
- 6:20 First bug: Postgres PGX client library length truncation
- 7:55 Illustrating Postgres length truncation with an example message
SQL Injection Isn't Dead Smuggling Queries at the Protocol Level
Speakers: Paul Gerste, Vulnerability Researcher, SonarSource
Conference: DEF CON 32
YouTube: https://www.youtube.com/watch?v=Tfg1B8u1yvE
Overview
In "SQL Injection Isn't Dead: Smuggling Queries at the Protocol Level," Paul Gerste from SonarSource challenges the prevailing notion that modern application development practices have largely eradicated SQL injection vulnerabilities. While parameterized queries and Object-Relational Mappers (ORMs) have significantly mitigated traditional, high-level SQL injection risks, Gerste's research unveils a new vector: injecting malicious queries by exploiting flaws in how applications communicate with databases at the binary protocol level. This talk dives into the "lower decks" of network communication, where bits and bytes are exchanged, revealing vulnerabilities that bypass even the most robust application-layer defenses.
The core of Gerste's discovery lies in binary protocol desynchronization, a concept analogous to HTTP request smuggling but applied to database wire protocols. By manipulating length fields in specific binary protocols, an attacker can cause a disagreement between a client library and the database server regarding the boundaries of a message. This desynchronization allows an attacker to "smuggle" additional, attacker-controlled SQL statements into the database connection, effectively achieving full SQL injection capabilities despite the use of prepared statements or other application-level sanitization.
This research is particularly significant because it uncovers a class of vulnerabilities that are often overlooked by developers and security researchers alike, especially in the context of memory-safe languages. It demonstrates that even with "clean code" practices and modern security tools, fundamental issues in protocol handling can lead to critical security bypasses. The talk not only details the technical specifics of these attacks, focusing on Postgres and MongoDB, but also explores their real-world applicability, potential bypasses for common input size limitations, and crucial defensive implications for developers and security professionals.
Background
▶ Watch: Introduction and prepared statement SQLi teaser (0:00)
The inspiration for this novel attack vector stems from HTTP desync attacks, pioneered by researchers like James Kettle. These attacks exploit discrepancies between two systems (e.g., a reverse proxy and an upstream backend) in how they interpret the end of an HTTP request. Root causes often involve variations in parsing HTTP headers like Content-Length and Transfer-Encoding, or subtle differences in character handling (e.g., allowing tab characters). Gerste pondered if similar desynchronization principles could be applied to other protocols, specifically binary protocols, which form the backbone of inter-service communication in modern web applications.
Binary protocols define message boundaries in various ways. Some use delimiters, such as null bytes for null-terminated strings, to signal the end of a data segment. Others rely on length fields, where a fixed-size integer specifies the total length of the subsequent message data. A common pattern is TLV (Type-Length-Value), where each message component explicitly states its type, length, and then the actual value. Gerste initially focused on length fields, considering potential issues like endianness or integer overflows as avenues for desynchronization.
Gerste strategically chose to focus on the connection between applications and databases for several reasons. Firstly, almost every web application relies on a database, ensuring broad applicability for any discovered vulnerabilities. Secondly, databases are high-value targets, containing sensitive user data, authentication information, and critical application logic, thus offering significant severity if compromised. Lastly, the inherent need for applications to query databases based on user input guarantees a consistent source of user-controlled data, making exploitation more feasible. This combination of factors made database wire protocols an ideal target for exploring binary protocol smuggling.
The talk provided a brief overview of popular database wire protocols to illustrate their message structures. Postgres messages typically consist of a one-byte type identifier, a four-byte integer length field, and then the value (e.g., an SQL statement string). MySQL uses a three-byte length field followed by a one-byte sequence number, enabling multi-packet messages. Redis, while not strictly binary, uses a type byte, a variable-length decimal integer for length, and delimiters. MongoDB employs a four-byte message length, an opcode (message type), and a more complex binary structure called BSON for its value payload. This diversity in protocol design hinted at various potential desynchronization opportunities.
Key Findings
▶ Watch: Exploring desync mechanisms in binary protocols (2:30)
The primary discovery presented by Paul Gerste is a novel integer truncation vulnerability affecting Postgres client libraries, specifically demonstrated in the popular Go library PGX. This vulnerability allows an attacker to smuggle arbitrary SQL queries into a database connection, effectively bypassing the security provided by prepared statements and other application-level defenses.
The core issue arises when the client library attempts to encode a message that exceeds the maximum size representable by a 32-bit integer, despite the host system typically using 64-bit integers for int types. The Postgres protocol specifies a four-byte integer for the message length. In the vulnerable PGX code, a function responsible for encoding message content into a buffer:
- Writes the message type (e.g., 'B' for a bind message).
- Saves an offset where the message size will later be written.
- Appends all the actual data of the message.
- Finally, calculates the length of the data (plus the length field itself) and writes it back to the saved offset.
The critical flaw occurs in step 4. The length of the buffer containing the message data is calculated and stored as a Go int type, which on most modern systems is a 64-bit integer. However, this 64-bit int is then explicitly truncated down to an int32 before being written into the four-byte length field of the Postgres message.
If the actual message length (including the four-byte length field itself) exceeds 2^31 - 1 (the maximum positive value for a signed 32-bit integer), this truncation causes the most significant bits to be lost. The resulting 32-bit length value written into the protocol header becomes significantly smaller than the true length of the message data being sent.
Consequently, when the database server receives this malformed message, it reads the truncated, smaller length field and expects a much shorter message. It parses only the initial portion of the incoming data up to this truncated length. The remaining bytes, which were intended to be part of the original, large message, are then interpreted by the database as subsequent, entirely new messages in the same connection stream. By carefully crafting a payload that is just over the 32-bit integer limit, an attacker can ensure that the "excess" bytes form a perfectly valid, malicious SQL statement or series of statements that the database will then execute.
This desynchronization between the client library's actual transmission (sending a very large message) and its declared length (a very small length due to truncation) allows for SQL query smuggling. The attacker can inject arbitrary SQL, including INSERT, UPDATE, DELETE, or even CREATE USER statements, bypassing the protections of parameterized queries because the injected queries are treated as entirely separate messages at the protocol level, not as part of the original query's parameters.
Technical Deep Dive
▶ Watch: Rationale for targeting databases: high impact and applicability (3:50)
The core of this vulnerability lies in the intricate dance between client-side encoding and server-side decoding of binary database protocols. Let's delve into the specifics, using Postgres as the primary example, as it was the focus of the exploit demonstrated.
A standard Postgres message adheres to a simple TLV (Type-Length-Value) structure:
- Type Identifier: A single byte indicating the message type (e.g., 'Q' for a simple query, 'B' for a bind message in a prepared statement context).
- Length Field: A four-byte signed integer (network byte order) specifying the total length of the message, including the length field itself but excluding the type byte. This means the minimum length for any message is 4 bytes.
- Value: The actual data payload, which could be an SQL statement, parameters, or other command-specific information.
The vulnerability was discovered in the Go language library PGX, a popular Postgres driver. Specifically, the issue resided in a function responsible for encoding message data into a larger buffer. The simplified code flow looks something like this:
The critical line is binary.BigEndian.PutUint32(w.buf[pos:], uint32(length)). Here, length is a Go int, which is typically 64-bit on modern systems. However, it is explicitly cast to uint32 (an unsigned 32-bit integer) before being written into the four-byte length field.
Consider the implications:
- Small Message: If
lengthis, say, 8 bytes (a type byte, 4-byte length, and 3 bytes of value),uint32(8)is 8. The database reads 8, parses 8 bytes, all is well. - Largest Valid Message: The maximum value for a
uint32is2^32 - 1(approximately 4.29 billion). Since the length field itself is included, the maximum data length is2^32 - 4bytes (roughly 4GB). If the message is exactly this size,lengthfits perfectly intouint32, and the database parses it correctly. - Truncated Message (The Exploit): If the actual
lengthexceeds2^32 - 1(e.g.,2^32bytes), whenuint32(length)is performed, the most significant bit(s) are truncated. For instance, iflengthis0x100000000(4GB + 4 bytes, beyond2^32-1), it becomes0x00000000afteruint32truncation. Iflengthis0x100000008(4GB + 12 bytes), it becomes0x00000008after truncation.
When this truncated length is written to the four-byte field, the database server receives a message where the declared length is much smaller than the actual data being sent. The database will:
- Read the type byte (e.g., 'B').
- Read the truncated length (e.g., 8 bytes).
- Parse only the first 8 bytes of the value as part of the original message.
- Crucially, the remaining bytes in the TCP stream are then interpreted as the start of new, independent Postgres messages.
An attacker can precisely craft a payload such that the bytes after the truncated portion form a perfectly valid Postgres message header followed by an arbitrary SQL statement. For example, if the original message was designed to be 4GB + 12 bytes long, and the truncation causes the length field to become 8 bytes, the attacker can ensure that bytes 9 through 4GB + 12 constitute a new Postgres message starting with a 'Q' (simple query) type byte, a valid length, and then INSERT INTO users (username, password, role) VALUES ('evil_admin', 'pwned', 'admin');.
The impact of this protocol-level SQL injection is profound:
- Arbitrary SQL Execution: Unlike classic SQL injection which often relies on modifying existing queries (e.g.,
UNION-based, subqueries), this attack allows the injection of entirely new, independent SQL statements. This is akin to having stacked queries enabled, even if the application or database typically doesn't support them through standard interfaces. - Bypassing Prepared Statements: The initial Go example (
db.Prepare("SELECT FROM users WHERE id = $1")) using prepared statements is completely bypassed. The prepared statement itself is secure, but the smuggled* query is an entirely separate message that the database processes independently, effectively ignoring the prepared statement's context. - Full Database Control: An attacker can perform
SELECT,INSERT,UPDATE,DELETEoperations on any data the application user has permissions for. This can lead to data exfiltration, modification of critical application data, or even privilege escalation (e.g., creating new admin users). If the database is insecurely configured (e.g.,pg_read_fileorpg_execenabled), this could even lead to remote code execution.
Data Exfiltration Challenges: While arbitrary SQL execution is possible, direct data exfiltration is less straightforward. Since the attacker is injecting new messages into an existing connection, the application is still expecting a response only for its original query. The results of the smuggled SELECT statements are sent back to the application, but the application's client library will likely ignore or misinterpret them as unexpected data, not directly returning them to the attacker. To exfiltrate data, an attacker would need to employ indirect methods, such as:
- Inserting into another table: Smuggle an
INSERTstatement that takes the results of aSELECT(e.g.,SELECT password_reset_tokens FROM users) and inserts them into a table that the attacker can later access via legitimate application business logic. - Out-of-band communication: In rare cases, if the database supports it and is configured to allow it, an attacker might be able to exfiltrate data via DNS queries or HTTP requests initiated from the database server itself.
While the primary exploit demonstration focused on Postgres due to the clear integer truncation bug in PGX, the talk also touched upon MongoDB's BSON protocol. BSON is a binary-encoded serialization format used for storing and transmitting documents. It also relies on length fields (e.g., for the overall document length, and for individual string/array lengths). Although a specific truncation bug wasn't detailed for MongoDB in the same way as Postgres, the existence of length fields in complex binary structures suggests similar desynchronization vulnerabilities could be present in other database drivers or protocols.
Demo / Proof of Concept
▶ Watch: First bug: Postgres PGX client library length truncation (6:20)
The talk opened with a compelling teaser demonstrating the practical impact of this vulnerability against a seemingly secure Go application snippet:
The speaker's challenge was: "If you think everything's fine here, there's no way for SQL injection, then you should stay because I'm going to prove you wrong." This code uses a prepared statement (db.Prepare("SELECT * FROM users WHERE id = $1")), which is the gold standard for preventing traditional SQL injection by separating SQL code from user-supplied data. Parameters like $1 are sent out-of-band and are not concatenated into the SQL string, making it immune to classic input manipulation.
The demonstration showed how the Postgres protocol-level smuggling vulnerability completely bypasses this robust application-layer defense. The attack scenario would unfold as follows:
- An attacker crafts an HTTP request to the
handlerendpoint. - Instead of sending a benign
userID, the attacker sends a maliciously crafted, extremely large payload in theuserIDparameter. This payload is designed to be slightly larger than2^32 - 1bytes when it is ultimately encoded by the PGX client library into a Postgres message. - The
PGXlibrary, when processing thestmt.Query(userID)call, attempts to encode theuserIDparameter (and potentially other parts of the bind message) into a Postgres message. - Due to the integer truncation bug, the
lengthfield in the Postgres message header is written as a much smaller value than the actual data being transmitted. - The attacker's payload is carefully constructed such that the "overflowing" bytes, which the database interprets as subsequent messages, form a legitimate Postgres message header followed by a malicious SQL statement. A concrete example shown was:
INSERT INTO users (username, password, role) VALUES ('evil_admin', 'pwned', 'admin');. - The database receives the initial (truncated) message and processes it. Then, it proceeds to parse and execute the smuggled
INSERTstatement, creating a new administrative user. - The application, unaware of the smuggled query, continues to process the original
SELECT * FROM users WHERE id = $1query (which might return an error or no results depending on theuserIDvalue) and handles its response. The attacker does not directly receive the result of their injected query but has successfully modified the database state.
This demonstration effectively illustrates that even with best practices like prepared statements, fundamental flaws in the underlying network protocol handling can introduce critical vulnerabilities, proving that "SQL injection isn't dead" – it has merely evolved to a lower layer of the stack.
Defensive Implications
▶ Watch: Illustrating Postgres length truncation with an example message (7:55)
The findings presented in this talk necessitate a re-evaluation of security assumptions, particularly concerning memory-safe languages and the robustness of database client libraries. Defenders must consider several layers of protection:
- Re-evaluate Integer Overflows in Memory-Safe Languages:
- Developer Awareness: Developers often assume that memory-safe languages (like Go, Java, Python, Rust) inherently prevent critical vulnerabilities stemming from integer overflows. This research demonstrates that while memory corruption might be averted, integer overflows can still lead to logical flaws and data desynchronization at the protocol level, with severe security implications.
- Code Review Focus: Security code reviews should extend beyond typical input validation and SQL injection patterns to scrutinize how large data inputs are handled, particularly when being serialized into fixed-size fields within binary protocols. Explicit casts from larger integer types to smaller ones (e.g.,
inttoint32oruint32) should be flagged for careful examination.
- Robust Input Size Limiting and Decompression Handling:
- Default Limits are Crucial: Applications should enforce strict input size limits at the earliest possible stage. While web frameworks often provide default body size limits (e.g., 1MB), these are frequently overridden or disabled for specific endpoints (like file uploads) without considering the broader security implications.
- Compression Bypass Mitigation: The talk highlighted that compression (e.g., GZIP) can bypass simple size checks. If a reverse proxy or application framework checks the request size before decompression, a small compressed payload can inflate into a gigabyte-sized message after decompression, triggering the overflow.
- Mitigation: Implement checks for both the compressed and decompressed size. The
FastifyJavaScript framework was cited as an example where the checks were performed in the wrong order.Nginxalso only checks raw request size. If your backend performs decompression, ensure it has its own robust, post-decompression size limits. - WebSockets Considerations: WebSockets can handle extremely large messages (up to 64-bit integer length fields), often bypassing HTTP-specific middleware or size limits. If your application uses WebSockets, ensure that message size limits are explicitly applied to individual WebSocket frames or messages.
- Scrutiny of Database Client Libraries and Protocol Implementations:
- Library Updates: Keep database client libraries (like PGX) updated to the latest versions. Vulnerabilities like the one found in PGX are typically patched quickly once disclosed.
- Protocol Fuzzing: For critical applications, consider implementing or sponsoring fuzzing efforts targeting the specific binary protocols and client libraries in use. This can uncover subtle parsing differences or overflow issues that lead to desynchronization.
- Beyond Prepared Statements: Developers should understand that while prepared statements are effective against traditional SQL injection, they do not protect against attacks that manipulate the underlying communication protocol itself. The security chain is only as strong as its weakest link, which, in this case, can be the client library's protocol implementation.
- Future Research & Proactive Defense:
- Non-Invasive Detection: The speaker called for research into non-invasive detection mechanisms to identify vulnerable client libraries without sending large, potentially crashing payloads. This could involve fingerprinting library versions or observing subtle protocol behaviors.
- Broader Protocol Research: Expand research into other binary protocols used for inter-service communication (caches like Redis, message queues like Kafka/RabbitMQ, other databases, structured logging systems). The general principles of length field manipulation or delimiter abuse could apply widely.
- Two-Byte Length Fields: Protocols using two-byte length fields (max 65KB) would be significantly easier to exploit, as sending 65KB is far more feasible than 4GB. These should be prioritized for review.
In summary, defenders must adopt a holistic security mindset that extends beyond application logic to the underlying network protocols. This includes robust input validation, careful handling of large and compressed data, vigilant library management, and an understanding that even "memory-safe" environments can harbor critical vulnerabilities at the binary communication layer.
Key Takeaways
- Integer overflows remain relevant for security vulnerabilities even in memory-safe languages. While they might not cause memory corruption, they can lead to critical logical flaws and data desynchronization at the protocol level.
- Protocol-level SQL injection bypasses traditional defenses like prepared statements. By smuggling entire new queries into the database connection stream, attackers can achieve arbitrary SQL execution, effectively operating at a "lower deck" of the application stack.
- Sending large amounts of data to trigger overflows is feasible despite common input size protections. Attackers can leverage unprotected endpoints, compression inflation, and WebSocket capabilities to bypass typical web framework and reverse proxy limits.
- Database client libraries and binary protocol implementations are critical attack surfaces. Developers and security researchers need to scrutinize these often-overlooked components for subtle parsing discrepancies and integer handling bugs.
- A holistic security approach is necessary. Defense must extend beyond application logic and traditional input validation to encompass the integrity of underlying communication protocols and the libraries that implement them.
About the Speaker(s)
Paul Gerste is a Vulnerability Researcher at Sonar's R&D team. Sonar, known as "the home of clean code," helps developers write secure code by identifying vulnerabilities and code quality issues. In his role, Paul focuses on discovering new attack vectors and improving code security, which involves deep dives into application architectures and underlying protocols, as demonstrated by his research on binary protocol smuggling.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
Gerste's research on binary protocol desynchronization in database drivers is a critical, high-impact finding that redefines what we thought we knew about SQL injection. By demonstrating how integer truncation in client libraries can smuggle entirely new queries past prepared statements, he's exposed a fundamental vulnerability at the protocol level. This isn't just a clever hack; it's a profound shift in attack vectors that bypasses modern defenses and demands immediate attention from developers and security professionals.
Heather Calloway (CISO) — STRONG ACCEPT
Paul Gerste's research on binary protocol desynchronization in database client libraries, specifically the integer truncation vulnerability in PGX for Postgres, fundamentally challenges the assumption that prepared statements fully protect against SQL injection. This work uncovers a critical blind spot in many security programs, demonstrating how attackers can smuggle arbitrary SQL queries at the protocol level, bypassing application-layer defenses. It demands a serious re-evaluation of input validation, library management, and a deeper understanding of the entire application stack's security posture.