Translating C To Rust: Lessons from a User Study

Ruishi Li

Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Software Security: Code and Compiler

Overview

Memory safety vulnerabilities have long plagued system security, leading to a relentless stream of critical exploits. The C programming language, while foundational, is a primary culprit due to its manual memory management and lack of inherent safety guarantees. Rust has emerged as a compelling alternative, offering strong memory safety guarantees—both temporal and spatial—with minimal performance overhead, while still providing low-level control over system resources. This unique combination has garnered significant attention from both industry and government, fueling a desire to migrate existing C codebases to Rust.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction: C memory safety and Rust's benefits
  2. 1:00 Limitations of automatic C to Rust translation
  3. 2:00 Human users successfully translate C to safe Rust
  4. 3:00 Details of the user study methodology
  5. 4:40 The 'hard problem': Rust's aliasing (AXM) principle
  6. 6:00 Human abstraction is key for idiomatic Rust
  7. 6:50 Human strategies: deleting aliases or cloning objects

Translating C To Rust: Lessons from a User Study

Speakers: Pratik Sakenna (National University of Singapore), Ashish Kundu (Cisco Research)

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=ElQs8rNjFzk

Overview

Memory safety vulnerabilities have long plagued system security, leading to a relentless stream of critical exploits. The C programming language, while foundational, is a primary culprit due to its manual memory management and lack of inherent safety guarantees. Rust has emerged as a compelling alternative, offering strong memory safety guarantees—both temporal and spatial—with minimal performance overhead, while still providing low-level control over system resources. This unique combination has garnered significant attention from both industry and government, fueling a desire to migrate existing C codebases to Rust.

However, the path to a memory-safe future is fraught with challenges, particularly when attempting to automatically translate legacy C code. Existing automated translation systems, whether compiler-based or powered by large language models (LLMs), have fallen short. Compiler-based tools often produce code riddled with unsafe Rust blocks, sacrificing the very safety benefits sought, even if their translations are generally correct. LLM-based approaches, while generating more idiomatic Rust, frequently deviate from the original program's behavior, leading to incorrect or unreliable translations.

In this talk, Pratik Sakenna, representing his team at the National University of Singapore and collaborator Ashish Kundu from Cisco Research, presents a novel approach to understanding this translation dilemma. Instead of developing yet another automated tool, their work investigates how human developers tackle the same translation task. By observing human strategies, the research aims to uncover the fundamental differences between human and machine translation, providing crucial insights to guide the development of future, more effective automated C-to-Rust translation systems capable of producing genuinely safe and correct code.

Background

▶ Watch: Introduction: C memory safety and Rust's benefits (0:00)

The pervasive issue of memory safety has been a cornerstone problem in system security for decades. Vulnerabilities like buffer overflows, use-after-free errors, and double-free bugs, predominantly found in C and C++ codebases, continue to be a primary vector for critical exploits. The industry's growing recognition of this problem has led to a strong push towards memory-safe programming languages. Rust stands out as a leading candidate for systems programming due to its unique combination of features:

  1. Full Memory Safety: Rust provides robust static guarantees for both temporal safety (preventing use-after-free and double-free) and spatial safety (preventing out-of-bounds access).
  2. Low Performance Overhead: Unlike many other memory-safe languages that rely on garbage collection or extensive runtime checks, Rust achieves safety through its borrow checker and ownership system at compile time, resulting in performance comparable to C/C++.
  3. Low-Level Control: Despite its safety features, Rust retains fine-grained control over memory allocation and object lifetimes, making it suitable for operating systems, embedded systems, and other performance-critical applications traditionally dominated by C.

Given these advantages, migrating existing C/C++ codebases to Rust has become a significant goal for many organizations. This has spurred the development of automatic translation systems designed to convert C code into Rust. However, as highlighted in the talk, these systems face considerable limitations.

Existing automated approaches generally fall into two categories:

  • Compiler-based translation tools: These tools typically perform a more literal, line-by-line translation, often preserving the low-level pointer arithmetic and memory access patterns of the original C code. While they tend to maintain correctness with respect to original test cases, this fidelity often comes at the cost of safety. The resulting Rust code is frequently riddled with unsafe blocks, which bypass Rust's safety guarantees and effectively reintroduce the very memory safety risks the migration aimed to mitigate. They translate C code into unsafe Rust, then attempt to "lift" these raw pointers into safe Rust types, a process that frequently fails or is incomplete.
  • Large Language Model (LLM)-based approaches: These systems leverage the pattern recognition capabilities of LLMs to generate Rust code that is often more idiomatic, making use of Rust's higher-level abstractions and standard library functions. However, LLMs struggle with precise semantic preservation. Their translations frequently introduce subtle deviations from the original C program's behavior, leading to incorrectness when subjected to the original test suites. They prioritize idiomatic expression over strict behavioral equivalence, which is a critical flaw for security-sensitive system code.

The fundamental challenge, therefore, is that neither existing category of automated tools consistently produces translations that are both "mostly safe and mostly correct." This gap motivated the presented research: to understand how human experts achieve this elusive balance, and what lessons can be learned to bridge the divide between automated and human-level translation quality.

Key Findings

▶ Watch: Human users successfully translate C to safe Rust (2:00)

The core of this research revolves around a unique user study designed to observe human performance in translating C code to Rust. The findings from this study reveal critical insights into the differences between human and automated translation processes and highlight the strategies that enable humans to produce safe, correct, and idiomatic Rust.

The user study involved 33 undergraduate students from a computer security course at the National University of Singapore. These students had prior exposure to memory safety vulnerabilities but only a brief introduction to Rust's principles. Participants were given C programs and associated test cases, with the explicit task of producing Rust code that was devoid of any unsafe blocks (i.e., completely safe Rust) and correct under the provided tests.

The C programs for translation were drawn from the Linux port of the BSD core utils package, ranging from 300 to 600 lines of code. Their functionality typically involved file and data processing, string manipulation, and parsing – common tasks in system utilities. Out of the 33 submissions, 31 final translations successfully met the criteria of being both safe and mostly correct. This remarkable success rate stands in stark contrast to the performance of existing automated translation systems, which consistently fail to produce both safe and correct code for these types of programs.

The most significant finding is that human users succeed where automated systems fail. This success is primarily attributed to a fundamental difference in approach: abstraction. Unlike automated tools that attempt a literal, line-by-line, or pointer-by-pointer translation, human developers comprehend the intent and high-level functionality of the C code. They then abstract away the low-level details and re-implement the functionality using appropriate, idiomatic, and safe Rust abstractions. This process is not a direct mapping but a re-engineering of the logic within Rust's safety paradigm.

Furthermore, the study also provided a fascinating insight into the performance of human-translated code. Without any explicit instruction or focus on optimization, the Rust programs produced by the students exhibited a performance overhead of "well below 20%" compared to their original C counterparts. This demonstrates that even without performance as a primary goal, the adoption of idiomatic Rust, with its zero-cost abstractions, naturally yields efficient code while providing robust memory safety guarantees. This finding underscores the potential for a favorable trade-off between security and performance when C code is thoughtfully translated to Rust.

Technical Deep Dive

▶ Watch: Details of the user study methodology (3:00)

The success of human translators hinges on their ability to navigate and resolve fundamental conflicts between C's low-level, unconstrained memory model and Rust's strict, safety-oriented design. The talk delves into several specific technical challenges and the elegant, idiomatic solutions devised by human participants.

A central point of conflict is Rust's AXM principle, which stands for Aliasing XOR Mutability. This fundamental static rule of the Rust compiler states that a memory object cannot simultaneously have two active aliases where one is mutable and the other is immutable. In C, it's perfectly normal to have multiple char* pointers referencing the same string, with some being used for modification and others for read-only access. A direct, low-level translation of such C code into Rust will inevitably lead to compilation errors due to borrow checker violations. This is precisely why compiler-based automated tools, which attempt such direct translations, often resort to unsafe Rust or fail.

Human translators, however, employ sophisticated strategies to circumvent the AXM principle and other C-isms:

  1. Eliminating Conflicting Aliases: Instead of preserving every C pointer as a distinct Rust reference, users often refactor the code to eliminate one of the conflicting aliases. For instance, if a C program uses two pointers, one mutable and one immutable, to operate on a string, the human translator might re-architect the logic. The "second" pointer's role could be absorbed into a hidden variable within a single, safe Rust function, such as a string replace function, which handles the internal logic safely without exposing multiple aliases to the user. This approach leverages Rust's encapsulation capabilities to manage mutable state locally.
  1. Cloning Objects for Read/Write Conflicts: Another ingenious strategy to resolve read/write conflicts, particularly when an object needs to be read from while simultaneously being written to (or a derivative written to), is to clone the original object. The human translator creates two distinct objects: one serves as the read-only source, and the other becomes the mutable target for output. By creating a clone, they effectively sidestep Rust's borrow checker by operating on two separate, independent memory locations, thus satisfying the ownership rules without compromising safety.
  1. Semantic Lifting of Pointers: A crucial distinction between human and automated translation lies in how C pointers are handled. Compiler-based tools tend to preserve C's raw pointer semantics, often translating char* to mut c_char or similar unsafe raw pointers in Rust. In contrast, human users engage in semantic lifting. They don't just translate the type; they interpret the contextual meaning* and intended use of the C pointer. For example, a char* in C is not a single concept but can represent a string slice, a mutable string, a byte array, a path, or even raw memory. Human translators observed in the study would elevate a char* to one of "at least 10 different Rust types" depending on its specific role (e.g., &str for an immutable string slice, String for an owned, mutable string, &[u8] for a byte slice, Vec<u8> for an owned byte vector, PathBuf for file paths, etc.). This semantic lifting is key to leveraging Rust's powerful, zero-cost abstractions, which statically guarantee temporal safety and contribute significantly to the low performance overhead observed.
  1. Handling Unsafe C APIs: C programs are replete with calls to standard library functions that operate on raw pointers and lack inherent safety. Automated tools often leave these as direct FFI (Foreign Function Interface) calls within unsafe blocks. Human translators, however, prioritize finding corresponding safe Rust equivalents in the standard library or well-vetted crates. If a direct safe equivalent isn't available, they will emulate the behavior of the C library function using safe Rust primitives, effectively re-implementing the functionality while adhering to Rust's safety guarantees.
  1. Dealing with C-specific Constructs: The study also identified human strategies for C constructs that have no direct safe equivalent in Rust:
  • Mutable Globals: C allows mutable global variables, which are problematic for Rust's ownership model and concurrency safety. Human translators address this by limiting the scope of globals into local variables. By transforming global state into local, stack-allocated variables or by passing them explicitly as function arguments, they enable the Rust compiler's borrow checker to statically track their usage, ensuring safety.
  • C Unions: C unions allow multiple data types to occupy the same memory location, a concept that is inherently unsafe and not allowed in safe Rust. Human users cleverly employ enums instead. By using an enum with variants that hold the different data types and potentially adding a discriminant field, they can mimic the behavior of C unions in a type-safe and idiomatic Rust manner.

These strategies collectively demonstrate that successful C-to-Rust translation is not merely a syntactic transformation but a deep semantic re-engineering process.

The talk also highlighted significant challenges for automated translators identified through this study:

  • Program Decomposition for LLMs: LLM-based translation tools struggle with the practical problem of decomposing large C programs into smaller, manageable segments that can be processed by an LLM. The optimal decomposition strategy is unclear. This leads to issues like functions with inter-dependencies being split across segments, resulting in duplicate or inconsistent Rust representations when stitched back together. This problem exacerbates with increasing program complexity, posing a major hurdle for scaling LLM-based translation.
  • Correctness Gap and Semantic Equivalence: While human translation often removes known vulnerabilities from the original C code (a desirable outcome), it can also introduce "logical and behavioral errors," for example, in string encoding or edge-case handling. This raises a crucial question for automated translation: what degree of semantic equivalence is desired? Is the goal to precisely replicate all original behavior (including bugs), or to produce a functionally equivalent, but safer, program? The speaker emphasizes the need to "specify what's to be kept and what's to be thrown away" during translation, acknowledging that strict 1:1 behavioral equivalence might not always be the primary objective if the overriding goal is enhanced security.

Demo / Proof of Concept

▶ Watch: Human abstraction is key for idiomatic Rust (6:00)

While this talk did not feature a live demonstration of a new translation tool or system, it effectively presented a "proof of concept" through the results of the user study itself. The core demonstration was the presentation of specific code examples, abstracted from the students' submissions, illustrating how human developers successfully navigate the complexities of C-to-Rust translation.

The speaker walked through a simplified, yet representative, C code snippet involving two pointers (P1 and P2) referencing the same string object, where one pointer was mutable and the other read-only. This scenario directly violates Rust's fundamental AXM principle, causing a direct translation to fail compilation. The "proof" then came in the form of showing how human participants translated this problematic C code into idiomatic Rust. This included examples where:

  1. One of the conflicting aliases was effectively "deleted" or absorbed into a higher-level function, allowing the Rust compiler to manage ownership safely.
  2. The original object was cloned to create separate read-only and mutable instances, thereby resolving the borrow conflict by operating on distinct data.

These examples served as concrete evidence that safe, correct, and idiomatic Rust translations are indeed achievable for C code, provided the translator (human or eventually automated) employs sophisticated strategies of abstraction and re-engineering rather than literal translation. The success of the 31 student translations for non-trivial C programs further solidified this proof of concept, demonstrating that the task, while challenging, is within human capabilities.

Defensive Implications

▶ Watch: Human strategies: deleting aliases or cloning objects (6:50)

The findings from this user study have profound implications for both developers and the broader security community, particularly for those involved in securing critical infrastructure and legacy C codebases.

For defenders and organizations working with C/C++:

  • Embrace Rust for New Development and Strategic Migrations: The study reinforces that Rust is a viable and highly effective language for building secure systems. For new projects, adopting Rust from the outset is a proactive defensive measure against memory safety vulnerabilities. For existing critical components, the study shows that a manual, thoughtful translation to Rust by skilled engineers can yield secure, high-performance code, eliminating entire classes of common vulnerabilities like buffer overflows, use-after-free, and data races.
  • Invest in Rust Education: Since human comprehension and abstraction are key to successful translation, organizations should invest in training their developers in Rust's unique idioms, ownership model, and borrow checker. This foundational understanding is crucial for both writing new Rust code and for effectively translating or refactoring existing C code.
  • Prioritize Semantic Abstraction over Literal Translation: When evaluating potential automated translation tools, defenders should look beyond simple syntactic conversion. Tools that prioritize "semantic lifting" of C concepts (like pointers) to appropriate Rust abstractions will be far more effective in producing truly safe code. The goal should be to re-engineer the logic within Rust's safety paradigm, not merely port C's unsafe patterns.

For developers of automated translation tools:

  • Shift Focus to Abstraction and Idiomatic Rust: Future tools must move away from low-level, line-by-line translation and instead aim to "comprehend" the C code's intent and re-express it using idiomatic Rust. This includes sophisticated techniques for semantic lifting of C types (especially raw pointers) to context-appropriate safe Rust types (e.g., &str, String, &[u8]).
  • Implement Strategies for C-Specific Constructs: Tools need to develop robust mechanisms for handling C constructs that have no direct safe Rust equivalent. This means transforming mutable globals into local variables or safely managed state, and translating C unions into safe Rust enums.
  • Address Program Decomposition and Semantic Equivalence: For LLM-based approaches, solving the program decomposition problem is paramount to avoid inconsistent translations and scalability issues. Furthermore, tool developers must define clear specifications for desired behavior post-translation. Should the tool strictly preserve all original C behavior (including vulnerabilities), or should it prioritize eliminating vulnerabilities, even if it introduces minor behavioral changes (e.g., in error handling or edge cases)? This clarity is vital for building trustworthy tools.
  • Integrate Program Analysis and Formal Methods: As suggested in the Q&A, techniques from program analysis and formal methods can play a crucial role in tackling challenges like program decomposition and ensuring the correctness of translated segments, potentially localizing LLM errors.

In essence, while the dream of fully automatic, high-quality C-to-Rust translation for massive, complex codebases remains "ages away," the study provides a pragmatic roadmap. It underscores that for critical components, the human element, guided by a deep understanding of both C and Rust, can achieve the desired safety and correctness. The insights gained are invaluable for guiding the next generation of automated tools to emulate this human capability, ultimately leading to a more secure software ecosystem.

Key Takeaways

  • Human developers can consistently translate C code into safe, correct, and idiomatic Rust, a feat current automated translation systems struggle to achieve.
  • The fundamental difference lies in abstraction: humans comprehend the C code's intent and re-implement it using appropriate Rust idioms, rather than performing a direct, low-level, or line-by-line translation.
  • Human translators effectively overcome Rust's strict AXM principle (Aliasing XOR Mutability) by refactoring code to eliminate conflicting aliases or by cloning objects to create distinct, safe memory regions.
  • A crucial strategy is semantic lifting of C pointers, where char* and similar types are contextually translated to one of many appropriate safe Rust types (e.g., &str, String, &[u8]), leveraging Rust's zero-cost abstractions.
  • Humans also adeptly handle C-specific constructs like mutable globals (by localizing scope) and C unions (by using Rust enums) to conform to Rust's safety rules.
  • Automated translation tools, particularly LLM-based ones, face significant challenges in effective program decomposition and in defining the desired level of semantic equivalence (i.e., whether to preserve all original behaviors, including vulnerabilities, or prioritize safety).

About the Speaker(s)

Pratik Sakenna is a faculty member at the National University of Singapore (NUS). He presented this work on behalf of his research team and his collaborator. His background includes teaching an undergraduate computer security course at NUS, from which the participants for the user study were recruited.

Ashish Kundu is affiliated with Cisco Research, collaborating with Pratik Sakenna and his team on this research into C-to-Rust translation.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Legitimate academic research with a clean methodology: 33 students, real C programs, measurable outcomes. The core finding — humans succeed at safe+correct translation by abstracting intent rather than mechanically mapping syntax — is valid and the specific strategies (AXM resolution, semantic lifting of char, enum-for-union substitution) are well-articulated. Not groundbreaking for anyone who's spent serious time with Rust's borrow checker, but a useful empirical grounding for what the community already suspected.

Heather Calloway (CISO) — WEAK

Solid academic work on a genuinely important problem, but it never makes the leap from research observation to institutional decision. The governance and operational questions — who should own Rust migration programs, how to evaluate automated tooling, what this means for organizations sitting on millions of lines of legacy C — are left entirely to the reader.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025