Interventional Root Cause Analysis of Failures in Multi-Sensor Fusion Perception Systems
Shuguang Wang
Network and Distributed System Security (NDSS) Symposium 2025 · Day 2 · Sensor Attacks
Overview
Autonomous driving systems, particularly their perception modules, are critical for safe and reliable operation. These systems rely on multi-sensor fusion to process data from various sensors like LiDAR and cameras in real-time, building a comprehensive understanding of their surroundings. However, perception systems are not infallible; faults within sub-modules can lead to severe issues such as missing obstacles or ghost obstacles, dramatically increasing collision risks and leading to unpredictable driving decisions. This talk by Shuguang Wang introduces a novel approach to perform root cause analysis (RCA) on these complex perception failures.
Key moments
- 0:00 Introduction to interventional root cause analysis
- 2:10 Limitations of existing RCA and proposed solution
- 4:10 Detailed explanation of the intervention algorithm
- 7:40 Case study: self-driving car's ghost obstacle error
- 8:00 Real-world evaluation on a physical self-driving car
- 9:20 Summary of method's capabilities and generality
Interventional Root Cause Analysis of Failures in Multi-Sensor Fusion Perception Systems
Speakers: Shuguang Wang
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=s9UX5S_dwXM
Overview
Autonomous driving systems, particularly their perception modules, are critical for safe and reliable operation. These systems rely on multi-sensor fusion to process data from various sensors like LiDAR and cameras in real-time, building a comprehensive understanding of their surroundings. However, perception systems are not infallible; faults within sub-modules can lead to severe issues such as missing obstacles or ghost obstacles, dramatically increasing collision risks and leading to unpredictable driving decisions. This talk by Shuguang Wang introduces a novel approach to perform root cause analysis (RCA) on these complex perception failures.
The presented work, "Interventional Root Cause Analysis of Failures in Multi-Sensor Fusion Perception Systems," addresses significant limitations in existing RCA methodologies. Traditional methods often struggle with the multi-modal nature of perception data, the intricate dependencies within perception sub-modules, and the pervasive problem of multiple concurrent root causes. By proposing an efficient, fully automated, and generalizable interventional approach, this research provides a vital tool for developers to locate faulty modules, identify specific fault types, and ultimately enhance the robustness and safety of autonomous driving systems.
The core innovation lies in constructing a hierarchical structured causal model of the perception system and employing an intervention algorithm to systematically pinpoint the true origins of failures. This methodology moves beyond mere correlation, actively manipulating module outputs to observe system reactions and isolate causal factors. The practical application of this research promises to significantly improve the debugging and recovery processes for self-driving cars, paving the way for more resilient and trustworthy autonomous vehicles in real-world scenarios.
Background
▶ Watch: Introduction to interventional root cause analysis (0:00)
Modern self-driving cars depend heavily on sophisticated perception systems to interpret their environment. These systems integrate data from diverse sensors, such as LiDAR and cameras, through multi-sensor fusion techniques to create a real-time, holistic view of objects and their spatial relationships. This intricate processing allows the vehicle to detect obstacles, identify traffic signs, and understand the dynamic world around it, forming the foundation for subsequent planning and control decisions. Despite advancements, these perception systems are inherently complex and susceptible to various faults within their numerous sub-modules. These faults can manifest as critical perception failures, ranging from failing to detect existing obstacles (missing obstacles) to erroneously identifying non-existent objects (ghost obstacles), both of which pose severe safety risks, potentially leading to collisions or erratic driving behavior.
The need for effective root cause analysis (RCA) in autonomous driving perception systems is paramount for system recovery and continuous improvement. However, existing RCA methodologies face several limitations when applied to the unique challenges of multi-sensor fusion perception:
- Single Modality Focus: Many RCA approaches, particularly those developed for microservices, rely on performance indicators derived from a singular data modality. Perception systems, by their nature, process and fuse multi-modal data. Ignoring cross-modal information can lead to an incomplete or inaccurate understanding of the failure's origin, as faults might arise from interactions or inconsistencies between different sensor streams.
- Limited Complexity Handling: Existing RCA work in autonomous driving often targets higher-level components like planning and localization. While these components are crucial, their module dependencies are relatively simpler compared to the granular, interconnected sub-modules within a multi-sensor fusion perception stack. The sheer complexity and intricate data flows within perception systems make it difficult for these methods to pinpoint faults at a sufficiently detailed level.
- Single Root Cause Assumption: Causality testing methods frequently operate under the assumption of a single root cause for a given failure. This assumption is often violated in complex, real-world systems like autonomous perception, where multiple concurrent faults can interact and collectively contribute to an observed failure. An RCA method that cannot handle multiple concurrent root causes will fail to provide a complete and accurate diagnosis, potentially leading to incomplete fixes.
These limitations highlight a critical gap: the absence of an efficient, fully automated, and generalizable approach capable of dissecting complex perception failures, identifying specific faulty modules, and discerning the types of faults, even when multiple concurrent issues are present. The presented research directly addresses this gap by proposing an interventional methodology designed to navigate the multi-modal data, intricate dependencies, and concurrent fault scenarios inherent in advanced perception systems.
Key Findings
▶ Watch: Detailed explanation of the intervention algorithm (4:10)
The research introduces a pioneering interventional root cause analysis framework specifically tailored for the intricate challenges of multi-sensor fusion perception systems in autonomous vehicles. This framework delivers several key findings and contributions that significantly advance the state of failure diagnostics in this critical domain:
Firstly, the core finding is the development of a robust methodology capable of systematically identifying multiple concurrent root causes within complex perception pipelines. Unlike prior work constrained by single-root-cause assumptions, this approach can disentangle interwoven faults, providing a comprehensive understanding of why a system fails when several underlying issues are simultaneously active. This capability is crucial for realistic autonomous driving scenarios where environmental factors, sensor imperfections, and software bugs can combine to create complex failure modes.
Secondly, the proposed method demonstrates remarkable generality and adaptability across different autonomous driving platforms. Through evaluations on two prominent open-source systems, Autoware and Apollo, the research confirms that its underlying principles and algorithms are not tied to a specific system architecture but can be effectively applied to various multi-sensor fusion configurations. This cross-platform compatibility is vital for the broader adoption and impact of the methodology within the autonomous vehicle industry.
Thirdly, the work establishes the practical effectiveness of its approach through rigorous evaluation in both simulated environments and real-world physical self-driving cars. The ability to identify specific, actionable root causes—such as unfiltered point clouds leading to ghost obstacles, clustering errors, or scenarios being out of distribution for detection models—proves its utility beyond theoretical constructs. The confirmation of identified issues by Autoware developers underscores the accuracy and relevance of the findings.
Finally, the research highlights the complete automation of the entire root cause analysis process, from module monitoring and causal model construction to intervention and fault decoding. This automation significantly reduces the manual effort and expertise traditionally required for debugging complex system failures, making the diagnostic process faster, more repeatable, and scalable for continuous integration and deployment cycles in autonomous vehicle development.
Technical Deep Dive
▶ Watch: Case study: self-driving car's ghost obstacle error (7:40)
The proposed interventional root cause analysis methodology is a multi-stage, fully automated process designed to overcome the inherent complexities of multi-sensor fusion perception systems. It systematically moves from monitoring system behavior to constructing a causal model, performing targeted interventions, and finally decoding the specific nature of the identified faults.
The process begins by monitoring modules during runtime. As the perception system operates, each sub-module's behavior is observed. Any module exhibiting anomalous behavior is flagged as a potential causal module. These abnormal states are then represented using fault mode vectors, which encapsulate the specific deviations observed in the module's output or internal state. This initial step provides a set of candidates for the root cause analysis, acknowledging that a module appearing faulty might merely be a symptom of an upstream issue.
The second critical step involves building a hierarchical structured causal model of the perception system. This model is represented as a Directed Acyclic Graph (DAG), where nodes represent individual perception modules and directed edges denote data flow and dependencies between them. The DAG visually and computationally maps out how information propagates through the system, from raw sensor input to fused object detections. The talk explains that this DAG can be obtained in two primary ways:
- Configuration Files: For systems like Autoware, the DAG can be derived from configuration files such as launch files, which explicitly define module instantiation and interconnections. For Apollo, similar information can be extracted from proto files that specify data structures and communication channels.
- Runtime Tools: Alternatively, the DAG can be inferred dynamically during system execution using specialized tools. For Autoware, ROS tools can be leveraged to observe active topics and node subscriptions, thereby reconstructing the data flow graph.
The DAG is fundamental because it provides the structural context necessary to distinguish between symptoms and root causes. However, merely knowing the structure is insufficient; a module might appear faulty simply because it's receiving incorrect inputs from an upstream module. To identify the real root causes, the framework introduces its core innovation: an intervention algorithm.
The intervention algorithm is designed to systematically explore the causal relationships within the DAG. It operates in two main stages: branch pruning and chain pruning, both involving targeted interventions. An intervention in this context means actively changing the output of a potential module and observing how the downstream system reacts. This operation is designed to be cross-platform computable, ensuring its applicability across different autonomous driving systems.
- Branch Pruning: This stage focuses on removing irrelevant branches of the DAG that are not contributing to the observed failure.
- It starts by identifying collider nodes—modules that receive input from two or more parent nodes, where at least one parent is abnormal. For instance, if C1 and C2 are collider nodes, the algorithm selects a parent node, P1, of C1.
- An intervention is performed on P1 by changing its output to a 'normal' or 'expected' state.
- If, after this intervention, C1 remains faulty, it implies that P1 was not the root cause of C1's fault, and thus P1 and its branch can be pruned as irrelevant to the failure propagating to C1.
- Conversely, if an intervention on a parent of C2 causes C2 to return to normal, it suggests that the parent was indeed a causal factor.
- The algorithm also identifies cases where the state of a collider node (e.g., C2) conflicts with the overall system failure. If C2's state is inconsistent with being a root cause, it can also be pruned. This iterative pruning process helps narrow down the potential causal modules to a more manageable set, forming a 'chain' of potentially related faults.
- Chain Pruning (for Multiple Concurrent Root Causes): Once a chain of relevant modules is identified, this stage addresses the possibility of multiple concurrent root causes.
- Instead of assuming a single fault, the algorithm attempts to split the identified chain into parts.
- It performs interventions on segments of the chain. For example, intervening on the "first half" of a chain.
- If a downstream node in that segment changes to normal, but the overall system failure still persists, it indicates that the intervened segment was not the sole root cause, and other parts of the chain are also contributing.
- The algorithm then recursively divides and intervenes on the remaining parts, effectively isolating individual or clustered concurrent root causes within the chain. This meticulous process ensures that all contributing factors are identified, not just the most obvious one.
Finally, once all causal modules have been identified through the intervention process, the system performs fault decoding. This involves translating the fault mode vectors of the causal modules back into their original, human-interpretable messages or representations, such as specific ROS topics or internal data structures. This decoding step reveals "what went wrong and exactly why," providing precise, actionable insights for developers to address the identified vulnerabilities or bugs. The entire multi-step process, from monitoring to decoding, is designed to be fully automated, making it a powerful and efficient tool for continuous system analysis and improvement.
Demo / Proof of Concept
▶ Watch: Real-world evaluation on a physical self-driving car (8:00)
While the presentation did not feature a live, interactive demonstration, the talk extensively utilized case studies and evaluation experiments to serve as compelling proof of concept for the interventional root cause analysis method. These demonstrations covered both simulated and real-world scenarios, highlighting the method's effectiveness and practical utility across different autonomous driving systems.
The evaluation experiments were conducted on three distinct configurations of the Autoware autonomous driving system, a prominent open-source platform. Two primary datasets were used:
- Real Fault Scenarios: Collected from GitHub issues reported by Autoware developers and users, providing a diverse set of realistic, previously encountered failures.
- Synthetic Fault Scenarios: Created by injecting specific faults into the system, allowing for controlled testing and validation of the method's ability to precisely identify known issues.
The method was compared against existing work, demonstrating its superiority in handling diverse and realistic fault scenarios and its efficiency in quickly identifying causal modules.
Several compelling case studies from the GitHub issues dataset were presented:
- Misidentification of a Truck: The system incorrectly identified a vehicle. The method successfully pinpointed the root cause as a misdetection in the LiDAR center point module.
- Misdetection of a Pedestrian Swing: A pedestrian's movement was not accurately tracked. The RCA identified a clustering error as the underlying problem.
- Uphill Treated as Ghost Obstacle: Perhaps the most illustrative case, a self-driving car incorrectly perceived an uphill slope as a massive obstacle, leading to an emergency stop. The interventional method precisely identified the root cause: unfiltered point cloud data which introduced a ghost obstacle. This finding was subsequently reported to the Autoware developers and was confirmed by them, validating the accuracy and real-world relevance of the technique.
Beyond simulations, the research also included real-world evaluations on a physical self-driving car equipped with the Autoware autonomous driving system, featuring a LiDAR and camera fusion perception system. Two critical scenarios were tested:
- Misdetection of a Pedestrian: In a specific real-world setting, the car failed to detect a pedestrian. The RCA revealed that this failure was due to the scenario being out of distribution for the detection model during its training phase, highlighting a critical gap in the training data.
- Misdetection of a Traffic Cone: Another real-world test involved the car missing a traffic cone. The root cause was identified as a sparse point cloud, indicating an issue with sensor data density or processing under specific conditions.
These real-world evaluations unequivocally demonstrated the practical effectiveness of the proposed method in diagnosing actual failures encountered by a physical autonomous vehicle.
Finally, the talk highlighted cross-platform validation using scenarios derived from Apollo's sensor data. Apollo is another major open-source autonomous driving system. This validation proved that the interventional root cause analysis method is not confined to Autoware but can effectively operate across different autonomous driving system architectures, showcasing its generality and broad applicability within the industry. These extensive evaluations and case studies collectively provided robust proof of concept for the method's ability to accurately and efficiently diagnose complex perception failures.
Defensive Implications
▶ Watch: Summary of method's capabilities and generality (9:20)
The interventional root cause analysis methodology presented by Shuguang Wang offers profound defensive implications for the development, deployment, and operational safety of autonomous driving systems. By providing a systematic, automated, and precise way to identify the underlying causes of perception failures, it empowers developers and safety engineers to build more robust and reliable self-driving cars.
For autonomous vehicle developers, this research provides an invaluable debugging and testing tool. Instead of relying on time-consuming manual inspection or trial-and-error, developers can leverage this automated RCA to:
- Pinpoint Specific Module Faults: The method doesn't just identify a failure; it locates the exact faulty module and often the nature of the fault (e.g., clustering error, unfiltered point cloud, out-of-distribution data). This allows for highly targeted code fixes and algorithm improvements, significantly accelerating the debugging cycle.
- Improve System Design and Architecture: By understanding common failure patterns and their root causes, developers can proactively design more resilient perception architectures. This might involve strengthening interfaces between modules, implementing more robust input validation, or integrating redundant checks at critical junctures.
- Enhance Training Data and Models: The revelation that a pedestrian misdetection was due to an out-of-distribution scenario directly informs the need to expand and diversify training datasets. Developers can use these insights to curate more representative data, improving the generalization capabilities of machine learning models used in perception.
- Validate Fixes and Prevent Regression: After implementing a fix, the RCA framework can be re-run to confirm that the original fault has been resolved and that no new, unintended regressions have been introduced, thereby ensuring the stability of the system.
For safety engineers and operators, the insights gained from this interventional RCA are equally critical:
- Better Understanding of Failure Modes: A detailed understanding of "what went wrong and why" enables safety engineers to more accurately assess potential risks, define operational design domains (ODDs), and establish more robust safety protocols. This moves beyond generic failure descriptions to specific, technical explanations.
- Informing Real-time Monitoring and Mitigation: The types of faults identified (e.g., ghost obstacles from unfiltered point clouds) can inform the development of real-time diagnostics and mitigation strategies. If a system can detect conditions prone to these specific faults, it might initiate a safe fallback maneuver or alert the human operator.
- Regulatory Compliance and Certification: As autonomous vehicle regulations evolve, the ability to rigorously analyze and explain system failures will become paramount for certification. This method provides a verifiable and systematic approach to demonstrate due diligence in addressing safety-critical issues.
- Proactive System Hardening: Beyond reactive debugging, the insights from this RCA can drive a proactive approach to system hardening. By identifying systemic weaknesses (e.g., inadequate filtering algorithms, vulnerabilities to sparse point clouds), organizations can invest in fundamental improvements that prevent entire classes of failures.
In essence, this interventional root cause analysis transforms the reactive process of fixing bugs into a proactive strategy for building inherently safer and more reliable autonomous driving systems, moving the industry closer to the goal of widespread, trustworthy autonomous transportation.
Key Takeaways
- The proposed interventional root cause analysis method addresses critical limitations of existing RCA techniques by effectively handling multi-modal data, complex module dependencies, and multiple concurrent root causes in autonomous driving perception systems.
- A hierarchical structured causal model (Directed Acyclic Graph or DAG) is central to the approach, providing a clear representation of module interactions and dependencies, obtainable from configuration files or runtime tools.
- The innovative intervention algorithm, featuring branch pruning and chain pruning, systematically identifies true root causes by actively manipulating module outputs and observing system reactions, rather than merely correlating events.
- The methodology has been rigorously evaluated and proven effective on two major open-source autonomous driving systems, Autoware and Apollo, demonstrating its generality and cross-platform applicability.
- Practical effectiveness was shown through detailed case studies, pinpointing specific technical root causes such as unfiltered point clouds leading to ghost obstacles, clustering errors, and failures due to detection models encountering out-of-distribution scenarios in both simulated and real-world environments.
- The entire process is fully automated, significantly enhancing the efficiency of debugging and improving the overall reliability and safety of complex multi-sensor fusion perception systems in autonomous vehicles.
About the Speaker(s)
The talk was presented by Shuguang Wang. Based on the provided metadata and transcript, no additional details regarding their title or affiliation were explicitly mentioned.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate systems-security research applying causal inference to autonomous driving perception failures — DAG-based intervention framework, cross-platform validation on Autoware and Apollo, real-world confirmation from developers. Solid academic contribution, but the security angle is thin and the novelty ceiling is modest; causal graph approaches to RCA aren't new, and the AV-specific adaptation, while competent, doesn't dramatically advance the state of the art.
Heather Calloway (CISO) — PASS
Technically credible research into autonomous vehicle perception failure diagnosis — but this is a systems reliability and AV safety engineering problem, not a security operations or governance problem. Outside my lane, and I'm not going to penalize it for that.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025