Keynote: AI Enabled Observability Explainers - We Actually Did Something With AI! - Vijay Samuel

Vijay Samuel

KubeCon + CloudNativeCon Europe 2025 · Keynote

Overview

In this insightful KubeCon EU keynote, Vijay Samuel, a Principal MTS Architect in reliability engineering at eBay, unveiled a pragmatic approach to integrating Artificial Intelligence into the complex world of observability. Titled "AI Enabled Observability Explainers - We Actually Did Something With AI!", Samuel's talk addressed the escalating challenges of managing highly distributed systems at massive scale, where traditional manual triage and even early machine learning models fall short. The core thesis revolves around the strategic combination of advanced engineering techniques with the capabilities of Large Language Models (LLMs) to create deterministic, high-quality "explainers" that significantly reduce the time and effort required for incident response.

Watch on YouTube

Visual summary for Keynote: AI Enabled Observability Explainers - We Actually Did Something With AI! - Vijay Samuel by Vijay Samuel
Visual summary for Keynote: AI Enabled Observability Explainers - We Actually Did Something With AI! - Vijay Samuel by Vijay Samuel

Key moments

  1. 1:30 Observability challenges at eBay's massive scale
  2. 3:15 Early ML for anomaly detection and its limitations
  3. 4:10 Initial LLM excitement and the hallucination problem
  4. 6:00 Strategy: Build high-quality, deterministic AI explainers
  5. 6:20 Detailed examples of AI-powered observability explainers
  6. 7:20 Addressing LLM context limits and hallucination for large data
  7. 8:15 Combining AI and engineering strengths for magical results

Keynote: AI Enabled Observability Explainers - We Actually Did Something With AI!

Speakers: Vijay Samuel, Principal MTS Architect, eBay

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=j0AqGpC_pp4

Overview

In this insightful KubeCon EU keynote, Vijay Samuel, a Principal MTS Architect in reliability engineering at eBay, unveiled a pragmatic approach to integrating Artificial Intelligence into the complex world of observability. Titled "AI Enabled Observability Explainers - We Actually Did Something With AI!", Samuel's talk addressed the escalating challenges of managing highly distributed systems at massive scale, where traditional manual triage and even early machine learning models fall short. The core thesis revolves around the strategic combination of advanced engineering techniques with the capabilities of Large Language Models (LLMs) to create deterministic, high-quality "explainers" that significantly reduce the time and effort required for incident response.

Samuel articulated eBay's journey from initial skepticism and disillusionment with LLMs' "silver bullet" promise to a nuanced understanding of their strengths and limitations. The talk emphasized that while LLMs excel at summarization and simple reasoning, they are prone to hallucination and randomness when given overly broad or complex tasks. The solution, as pioneered by eBay, is to build foundational, reliable AI-powered building blocks that perform specific, well-defined tasks, and then layer these together to achieve more sophisticated outcomes like automated incident triage. This approach represents a crucial shift from hoping AI will solve everything to intelligently designing systems where AI augments, rather than replaces, robust engineering.

The significance of this work extends beyond eBay, offering a blueprint for organizations grappling with similar observability hurdles. By demonstrating how to harness AI's capabilities effectively within the constraints of real-world, high-stakes production environments, Samuel provided a compelling vision for the future of operational excellence. The "explainer" concept—whether for traces, logs, metrics, or changes—promises to transform how SREs identify, diagnose, and resolve issues, ultimately enhancing system reliability and user experience.

Background

▶ Watch: Observability challenges at eBay's massive scale (1:30)

The landscape of modern software systems has grown exponentially in complexity, posing immense challenges for ensuring reliability and performance. Vijay Samuel underscored this by detailing eBay's operational scale: a staggering 4,600 microservices powering its global site, generating 15 petabytes of logs daily, managing 10 billion active time series, and processing 10 million spans per second (sampled at 2%). This sheer volume of data and inter-service dependencies makes manual incident triage an increasingly untenable task. Humans, Samuel noted, have inherent limits to comprehension, struggling to traverse lengthy call chains, sift through terabytes of logs, or synthesize information from countless dashboards under pressure. The traditional trial-and-error approach to incident resolution leads to significant delays and errors, directly impacting customer experience.

Prior to the advent of sophisticated AI, eBay, like many enterprises, attempted to mitigate these issues with earlier forms of machine learning. Samuel highlighted efforts led by his colleague Huai Jang, which focused on reducing detection time to under four minutes using anomaly detection. They developed an internal system called Groot (available as a white paper), designed to attach a root cause to alerts triggered by business KPIs. Simple auto-remediation strategies, such as bouncing an outlier pod, were also implemented. While these initial ML-driven innovations were a "good start," they suffered from a fundamental limitation: they learned solely from past observations. When confronted with novel situations or unseen patterns, these systems lacked the reasoning capability to infer solutions, rendering them ineffective in the face of evolving system behaviors or unforeseen failures.

The arrival of Large Language Models (LLMs) brought a wave of optimism, initially perceived as a "silver bullet" capable of comprehending human input and responding like humans. Samuel recounted his own early experiments with ChatGPT, attempting to get it to rewrite Prometheus postings indexes using roaring bitmaps—an endeavor that vividly illustrated the concept of hallucination. This experience, shared by many early adopters, quickly revealed that LLMs, when prompted broadly or asked to perform complex, probabilistic tasks like full site triage, often yield random and unreliable responses. The fundamental issue, Samuel explained, is that layering probabilistic components in complex workflows causes probabilities to work against you, leading to unpredictable and untrustworthy outcomes, especially in high-stakes incident troubleshooting scenarios. This realization prompted eBay to rethink its strategy, moving away from an "AI-does-everything" mentality towards a more grounded, deterministic approach.

Key Findings

▶ Watch: Initial LLM excitement and the hallucination problem (4:10)

The central revelation from eBay's extensive experimentation with AI in observability is that the "silver bullet" promise of LLMs for complex, high-stakes tasks like incident triage is largely a fallacy. While LLMs exhibit impressive capabilities in language comprehension and generation, their inherent probabilistic nature, susceptibility to hallucination, and finite context windows make them unreliable for critical, deterministic decision-making in production systems. Samuel emphasized that merely "shoving everything into the LLM" inevitably leads to disappointment, especially when dealing with the vast and intricate data streams characteristic of modern microservice architectures (e.g., traces with thousands of spans).

Instead, eBay discovered that the true power lies in fostering a "love relationship between AI and engineering." This paradigm shift asserts that maximum value is unlocked when AI and traditional engineering are leveraged for their respective strengths. Engineering excels at deterministic processing, data manipulation, and precise algorithmic execution, while AI, particularly LLMs, is uniquely suited for tasks involving natural language understanding, summarization, and simple reasoning.

This insight led to the development of building block capabilities—high-quality, highly deterministic "explainers" that can be confidently relied upon. These explainers are designed to perform specific, bounded tasks, such as analyzing a trace, a set of log lines, a metric time series, or a code change, and then providing a concise, accurate summary. By breaking down the complex problem of observability into smaller, manageable, and deterministic AI-assisted components, eBay found a pathway to harness AI effectively without succumbing to its inherent probabilistic weaknesses. The success of this approach hinges on meticulously preparing the input data through robust engineering before presenting it to the LLM, ensuring that the AI operates within its optimal domain.

Technical Deep Dive

▶ Watch: Strategy: Build high-quality, deterministic AI explainers (6:00)

The core of eBay's innovation lies in its meticulous approach to combining engineering rigor with LLM capabilities, specifically addressing the challenges of large data volumes and LLM context window limitations. Samuel highlighted that a typical checkout API trace might have 3,000 spans, with some use cases reaching 8,000 spans per request. Directly feeding such extensive data into an LLM would quickly exceed its context window, leading to truncation, loss of critical information, and increased hallucination.

To circumvent these issues, eBay devised a multi-step, engineered approach for its Trace Explainer, which serves as a foundational building block:

  1. Critical Path Identification: The first engineering step is to "clean the trace" by removing everything not on the critical path. This is crucial because only a subset of spans directly contributes to the observed latency or error. eBay initially drew inspiration from Uber's Crisp white paper, which outlines a method for generating a critical path, and then developed its own improvisations to create an algorithm tailored to its specific needs. This algorithm effectively prunes irrelevant spans, drastically reducing the data volume.
  1. Few-Shot Prompting with SRE Heuristics: After isolating the critical path, the system employs few-shot prompting. This involves feeding the LLM with specific examples and rules that mimic how eBay's Site Reliability Engineers (SREs) triage active incidents. For instance, the prompt might instruct the LLM that a 4xx HTTP status code should not be considered a hard failure, whereas a 5xx error requires immediate attention. This injects domain-specific knowledge and deterministic logic into the LLM's reasoning process.
  1. Focus on Self-Time: The system prioritizes self-time within spans. Self-time represents the time spent executing code within a specific span, excluding time spent in child spans. Focusing on self-time helps identify truly resource-intensive operations or bottlenecks within a particular service, rather than simply identifying long-running parent operations that might be waiting on other services.
  1. Dictionary Encoding: To further optimize data for the LLM and reduce potential for misinterpretation or unnecessary verbosity, eBay implemented dictionary encoding. Instead of passing verbose service names like service name = checkout, the system encodes them as service name = 1. For machines, the semantic meaning of "checkout" versus "1" is irrelevant; they process patterns. This encoding reduces token count, making more room within the context window and potentially reducing hallucination.
  1. Chunking and Partial Explanations: Even after critical path identification and encoding, a trace might still be too large. The solution involves splitting the trace into upstream and downstream chunks. Each chunk is then processed by the LLM to generate partial explanations. These partial explanations are subsequently combined and synthesized by another engineered component to produce a final, comprehensive explanation of the full critical path, identifying multiple performance issues if present.

This methodology ensures that the LLM receives highly curated, relevant, and concise input, allowing it to perform its strength—summarization and simple reasoning—without being overwhelmed or prone to error.

Beyond the Trace Explainer, Samuel briefly outlined other building block explainers:

  • Log Explainer: Analyzes log lines to identify error or latency patterns.
  • Metric Explainer: Processes time series data to detect trends or anomalies.
  • Change Explainer: Identifies the type of change made to an application and summarizes its potential impact.

These individual explainers are then layered together to create more sophisticated capabilities. For example, a Dashboard Explainer takes dashboard metadata (time series, annotations, faulty traces), feeds them through the corresponding explainers, and generates an overarching explanation. This can be extended into a full triage workflow, where an alert triggers the analysis of KPI dashboards, standard dashboards, and embedded faulty traces (a pipeline stage in eBay's alert manager embeds traces directly into alerts for richer context), culminating in a summarized triage report. This modular, API-backed approach ensures high determinism and reliability, making the AI outputs trustworthy for critical operations.

Demo / Proof of Concept

▶ Watch: Addressing LLM context limits and hallucination for large data (7:20)

While the talk didn't feature a live, interactive demo, Vijay Samuel presented a compelling real-world use case that vividly illustrated the efficacy of eBay's AI-enabled observability explainers. The scenario involved a production incident characterized by a slowdown in database queries. This particular incident presented a significant challenge, as the problematic trace contained approximately 1,300 spans, a volume that would typically require extensive manual investigation by SREs.

In this incident, the Trace Explainer was deployed. By leveraging its engineered critical path identification, few-shot prompting, and self-time analysis, it quickly pinpointed the user segment service as the culprit. More specifically, it identified a particular span within this service that was consuming an inordinate amount of time. This single piece of information was crucial, immediately narrowing down the scope of investigation to a specific service and its problematic operation within the complex call chain.

Building on this initial insight, the Log Explainer was then brought into play. Analyzing the relevant log lines from the user segment service, the Log Explainer identified a critical timeout exception. This provided the precise nature of the problem, moving from "something is slow in this service" to "this service is timing out due to a database query issue."

Samuel highlighted that this combined output from the Trace and Log Explainers saved eBay's SREs a substantial amount of time. Instead of manually sifting through thousands of spans, navigating intricate trace visualizations, and then manually correlating trace IDs with vast log archives to find the needle in the haystack, the explainers provided a concise, actionable diagnosis. This direct identification of the problematic service and the specific error allowed the SRE team to focus their efforts immediately on the root cause, significantly accelerating the mean time to resolution (MTTR).

The demonstration effectively showcased how these "building block capabilities" can be layered to create a powerful, automated triage workflow. By integrating the Trace Explainer's output with the Log Explainer's findings, eBay achieved a level of diagnostic precision and speed that was previously unattainable through manual or earlier ML-based methods. This example underscores the practical, tangible benefits of eBay's "AI and engineering love relationship" approach in real-world incident management.

Defensive Implications

▶ Watch: Combining AI and engineering strengths for magical results (8:15)

The insights gleaned from eBay's journey into AI-enabled observability offer critical defensive implications for any organization managing complex distributed systems. The primary takeaway is the imperative to back LLMs with robust APIs and deterministic engineering. LLMs should not be treated as black boxes for complex problem-solving. Instead, their probabilistic nature necessitates that they be supported by well-defined, reliable engineering components that handle the heavy lifting of data preparation, filtering, and deterministic analysis. For instance, critical path identification, anomaly detection, and specific data transformations should be handled by traditional code and algorithms, with the LLM used only for summarization and simple reasoning on the highly curated output.

Organizations should judiciously identify the true strengths of LLMs and limit their application accordingly. Samuel articulated these strengths as:

  • Simple Reasoning: Performing straightforward logical deductions based on provided context.
  • Summarization: Condensing large volumes of text or data into digestible explanations.
  • Code Generation: Assisting developers with boilerplate or common coding patterns, akin to tools like Copilot.
  • Internal Knowledge Search through Retrieval Augmented Generation (RAG): Efficiently retrieving and synthesizing information from internal documentation or knowledge bases.

Conversely, for tasks requiring absolute precision, deterministic outcomes, or complex algorithmic processing, traditional code remains superior and should be preferred over LLMs. The temptation to "use the LLM just because" it's the latest technology should be resisted.

A critical defensive strategy highlighted by Samuel is the urgent need for standardization across the observability stack. This applies to two key areas:

  1. Ingest Standardization: Widespread adoption of OpenTelemetry is crucial, but it's not enough to simply use the framework. Adherence to its schema is paramount. Consistent naming conventions, attribute usage, and data structures allow for reliable assumptions and automated processing. Without this standardization, the data fed to any AI system, whether LLM or traditional ML, becomes inconsistent and difficult to interpret reliably.
  2. Query Standardization: The current observability landscape lacks a universally adopted, standardized query language. Samuel explicitly stated that natural language querying, while appealing, is "frankly not the answer" for robust, deterministic data retrieval. He encouraged participation in efforts like the query language standardization group of the OpenTelemetry TAG, co-chaired by Chris Larson, recognizing that a structured, standardized query language is essential for reliable AI integration and automated analysis.

In essence, the defensive posture involves building high-quality, deterministic "building block capabilities" through a thoughtful integration of AI and engineering. This ensures that incident response systems are not only intelligent but also reliable and trustworthy, preventing the introduction of randomness or hallucination into critical operational workflows. By focusing on data quality, robust pre-processing, and adherence to established standards, defenders can leverage AI to enhance observability without compromising the integrity of their systems.

Key Takeaways

  • AI and Engineering Must Collaborate: LLMs alone are insufficient and unreliable for complex observability tasks due to hallucination and context window limits. The most effective approach combines deterministic engineering (for data processing, critical path identification, anomaly detection) with LLM strengths (summarization, simple reasoning).
  • Build Deterministic "Explainers": Focus on creating high-quality, reliable "building block capabilities" like Trace, Log, Metric, and Change Explainers. These perform specific, bounded tasks and generate concise, accurate summaries of complex data.
  • Data Preparation is Paramount: For LLMs to be effective, input data must be meticulously engineered. Techniques like critical path analysis (inspired by Uber's Crisp), few-shot prompting with SRE heuristics, focusing on self-time, and dictionary encoding are essential to reduce data volume and increase determinism.
  • Leverage LLMs for Their Strengths: Use LLMs primarily for summarization, simple reasoning, code generation (e.g., PromQL), and internal knowledge retrieval (RAG). Avoid using them for probabilistic decision-making in high-stakes incident response where deterministic answers are required.
  • Standardization is Critical for AI Success: Widespread adoption and strict adherence to OpenTelemetry schema for data ingest are vital for consistent and reliable AI processing. Furthermore, a standardized query language (beyond natural language) is necessary for robust AI-driven data analysis and automation.
  • Start Small and Layer Capabilities: Instead of attempting to solve complex problems like full site triage with a single LLM, build foundational explainers and then layer them to achieve more sophisticated workflows, such as dashboard explainers and automated alert triage.

About the Speaker(s)

Vijay Samuel is a Principal MTS Architect within the reliability engineering organization at eBay. With a career spanning approximately 13 years at eBay, straight out of college, he brings extensive experience in observability, SRE operations, and site reliability. Samuel is a passionate open-source enthusiast, having contributed to projects like Drizzle DB early in his career and more recently to communities such as OpenTelemetry and Prometheus. His deep understanding of large-scale distributed systems and his practical approach to integrating emerging technologies like AI make him a leading voice in the observability space.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Samuel's keynote from eBay delivers a refreshingly honest and technically grounded take on integrating AI into observability. Moving beyond the 'silver bullet' hype, he details a pragmatic, engineering-first approach to building deterministic 'explainers' using LLMs. This talk provides concrete, actionable strategies for leveraging AI's strengths in summarization and simple reasoning, while mitigating its inherent weaknesses through rigorous data preparation and API-backed components. It's a valuable blueprint for anyone serious about improving incident response in complex distributed systems without succumbing to buzzword-driven randomness.

Heather Calloway (CISO) — STRONG ACCEPT

Finally, a talk that cuts through the hype around AI and provides a pragmatic, engineer-driven approach to leveraging Large Language Models (LLMs) for critical operational observability. Vijay Samuel from eBay outlines a clear strategy for building reliable, deterministic AI-enabled "explainers" by meticulously combining robust engineering with LLM strengths in summarization and simple reasoning. This is a blueprint for reducing incident response times and enhancing system resilience, grounded in institutional realism rather than aspirational pronouncements.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025