Sponsored Keynote: Lessons Learned in LLM Prompt Security - Jakub Suchy

Jakub Suchy

KubeCon + CloudNativeCon Europe 2025 · Sponsored Keynote

Overview

In this insightful keynote at KubeCon EU, Jakub Suchy from Haroxy Technologies delved into the critical, yet nascent, field of Large Language Model (LLM) prompt security. Drawing a stark parallel to the early days of web security, specifically the "2003 of OWASP Top 10," Suchy underscored the urgent need for robust security measures as organizations rapidly adopt AI. The talk highlighted a significant gap in current AI gateway implementations: while they often cover authentication, rate limiting, and PII detection, prompt security is frequently overlooked or inadequately addressed.

Watch on YouTube

Visual summary for Sponsored Keynote: Lessons Learned in LLM Prompt Security - Jakub Suchy by Jakub Suchy
Visual summary for Sponsored Keynote: Lessons Learned in LLM Prompt Security - Jakub Suchy by Jakub Suchy

Sponsored Keynote: Lessons Learned in LLM Prompt Security - Jakub Suchy

Speakers: Jakub Suchy, Haroxy Technologies

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=UI-b-Odg39A

Overview

In this insightful keynote at KubeCon EU, Jakub Suchy from Haroxy Technologies delved into the critical, yet nascent, field of Large Language Model (LLM) prompt security. Drawing a stark parallel to the early days of web security, specifically the "2003 of OWASP Top 10," Suchy underscored the urgent need for robust security measures as organizations rapidly adopt AI. The talk highlighted a significant gap in current AI gateway implementations: while they often cover authentication, rate limiting, and PII detection, prompt security is frequently overlooked or inadequately addressed.

Suchy's presentation focused on the practical challenges and performance bottlenecks encountered when integrating advanced LLM-based security models directly into AI gateways or load balancers. His research revealed that despite the theoretical promise of using AI to secure AI, the computational demands of these sophisticated classification models—such as Meta's Llama Guard or Google's Shield Gemma—introduce substantial latency, rendering them impractical for high-throughput, real-time environments. The core message was a call to action for the industry to innovate faster, develop specialized, smaller models, and refine existing techniques to effectively secure the increasingly complex AI landscape.

This talk is particularly relevant for architects, developers, and security professionals tasked with deploying and safeguarding AI-powered applications. It provides a sobering, data-driven perspective on the current state of LLM prompt security, emphasizing that while the enthusiasm for AI is high, the understanding and implementation of its security risks are still in their infancy. Suchy's work with Haroxy, known for the HAProxy load balancer, lends a practical, infrastructure-centric viewpoint to the problem, making the findings immediately applicable to real-world deployments.

Background

The rapid proliferation of Large Language Models (LLMs) across enterprises has necessitated the creation of AI gateways. Conceptually, an AI gateway is an evolution of the traditional API gateway, serving as an intermediary between client applications and LLM inference engines. These gateways are designed to centralize critical functions, including authentication, rate limiting, PII (Personally Identifiable Information) detection or extraction, and prompt routing, before requests are forwarded to the underlying HTTP API of an LLM. However, as organizations rush to embrace AI, a crucial security dimension—prompt security—has emerged as a significant challenge, often overlooked in initial deployments.

Prompt security addresses vulnerabilities arising from malicious or unintended user inputs designed to manipulate an LLM's behavior. Common attack vectors include prompt injection, where users attempt to override system instructions (e.g., "ignore all previous instructions"), or more sophisticated techniques like time bandit attacks, which historically exploited specific vulnerabilities in models like OpenAI's offerings to extract sensitive information or alter responses. The fundamental problem is that LLMs are highly susceptible to the content and structure of their input prompts, making them vulnerable to adversarial manipulation that can lead to data breaches, unauthorized access, or the generation of harmful content.

To counter these threats, the industry has begun developing specialized LLM-based security solutions. Notable examples include Meta's Llama Guard model and Google's Shield Gemma. These models are essentially classification LLMs themselves, designed to analyze incoming prompts and determine if they are safe or malicious. Many of these solutions, as Suchy noted, are built upon variations of dBERTa classification, a sophisticated deep learning architecture adept at natural language understanding. The premise is straightforward: an AI gateway would integrate such a model to classify prompts, returning a "yes" or "no" answer regarding their safety before allowing them to proceed to the main LLM. This approach aims to leverage the power of AI to defend against AI-specific threats, marking a significant shift from traditional rule-based or regex-based security mechanisms.

Key Findings

Jakub Suchy's presentation unveiled a critical and often underestimated challenge in the realm of LLM prompt security: the severe performance overhead introduced by integrating sophisticated LLM-based security models directly into AI gateways or load balancers. While the concept of using AI to secure AI is theoretically appealing, the practical implementation reveals significant bottlenecks that hinder real-time, high-throughput operations.

The primary finding was the substantial latency incurred when running prompt classification models. Suchy demonstrated that even on a powerful G6X large AWS instance, a single prompt approaching 500 tokens could take between 150 to 200 milliseconds to process. This latency escalates dramatically with longer prompts; a 2,000-token prompt, exceeding the typical 500-token context window of many classification models, would require four sequential processing steps, accumulating to a full 1 second of processing time. In the context of a load balancer or an AI gateway, where sub-millisecond response times are often critical, a 1-second delay is considered "lifetimes" and utterly unacceptable for user experience or system responsiveness.

Furthermore, Suchy highlighted the severe limitations on requests per second (RPS). His tests, even with an optimized inference engine, showed that the system could barely sustain 60 RPS with a non-optimized model. As concurrency increased, particularly at eight concurrent requests, the performance plummeted to less than 40 RPS. This indicates a significant scalability issue, making it challenging for a single AI gateway instance to handle the traffic volumes expected in modern enterprise applications, especially when multiple users or services are simultaneously interacting with LLMs.

The attempts to mitigate these performance issues yielded limited success. While optimizing the inference engine provided approximately a 30% improvement, this was insufficient to overcome the fundamental latency problem. Suchy also explored token caching, a common optimization for generative AI, but found it largely ineffective for classification tasks, as the nature of the operation differs significantly. Traditional text filtering for "bad words" or basic patterns was also assessed, but found to be easily bypassed by simple linguistic modifications, such as typos or word mangling, which LLMs can still interpret correctly. These findings collectively underscore that current approaches to LLM prompt security, while necessary, are not yet mature enough for widespread, high-performance deployment within critical infrastructure components like load balancers.

Technical Deep Dive

The core technical challenge illuminated by Jakub Suchy's talk revolves around the inherent computational intensity of running Large Language Model (LLM) classification models within the low-latency environment of an AI gateway or load balancer. Unlike traditional security mechanisms that rely on fast pattern matching or simple rule evaluation, LLM-based prompt security models are, by definition, large neural networks requiring significant processing power.

Suchy's experiments were conducted on a G6X large AWS instance, a substantial cloud resource, yet even this powerful hardware struggled to meet performance demands. The latency figures are critical: 150-200 milliseconds for a 500-token prompt. This is because these classification models, often based on architectures like dBERTa, perform complex contextual analysis of the entire input prompt to determine its safety. Each token contributes to the computational load, and the process is largely sequential within the model's forward pass. When a prompt exceeds the model's typical context window (e.g., 500 tokens), it must be segmented and processed iteratively, leading to cumulative delays. For a 2,000-token prompt, four iterations multiply the latency, resulting in a full second of dedicated processing just for security analysis.

The observed Requests Per Second (RPS) limitations further highlight the problem. An unoptimized model, likely running basic transformers directly, struggled to achieve even 60 RPS. Even with an optimized inference engine, which typically involves techniques like quantization, model compilation, and efficient batching, the improvement was only around 30%. This suggests that while optimizations help, they do not fundamentally alter the underlying computational cost of these large models. Furthermore, as concurrent requests increase, the shared resources (CPU, GPU memory, I/O) quickly become saturated, leading to a sharp decline in RPS, dropping to below 40 RPS at just eight concurrent requests. This indicates that the models are not inherently parallelizable to the degree required for high-throughput load balancing.

Suchy also touched upon token caching, a technique commonly used in generative AI to speed up subsequent token generation by reusing previously computed internal states. However, he noted its limited applicability to classification tasks. In generative AI, the model sequentially predicts tokens, and the cached states from previous tokens are directly relevant to predicting the next. In classification, the entire input prompt is typically processed at once to produce a single output (a "safe" or "unsafe" label), meaning the sequential generation and caching paradigm doesn't directly translate to performance gains for the classification inference itself.

Finally, the discussion of text filtering underscored the sophistication required for LLM security. Simple keyword or regex filters, while fast, are easily circumvented. As Suchy explained, a typo or slight mangling of words in a malicious prompt can bypass a static filter, yet an LLM can still correctly interpret the attacker's intent. This highlights the need for semantic understanding, which is precisely what complex LLM-based classifiers provide, but at a significant computational cost. The technical deep dive reveals a fundamental tension: the need for advanced AI-driven semantic security versus the demanding performance requirements of network infrastructure.

Demo / Proof of Concept

While Jakub Suchy's keynote did not feature a live, interactive demonstration in the traditional sense, the core of his presentation was built upon the practical implementation and testing of LLM-based prompt security models within an AI gateway environment. Suchy explicitly stated, "I did that and I ran these models inside an AI gateway inside a load balancer," referring to models like Meta's Llama Guard and Google's Shield Gemma.

His "proof of concept" was the empirical data he presented regarding the performance characteristics of these models. The latency figures (150-200ms for 500 tokens, 1 second for 2000 tokens) and the Requests Per Second (RPS) metrics (max 60 RPS, dropping to 40 RPS at 8 concurrent requests) were derived directly from these real-world tests. He detailed the infrastructure used, specifically mentioning a G6X large AWS instance, and described the varying performance between "non-optimized models" (basic transformers) and "optimized models with an inference engine," which yielded about a 30% improvement.

Therefore, although there wasn't a live demo, the entire talk served as a report on the findings from a comprehensive, hands-on proof of concept. Suchy's work demonstrated the feasibility of integrating these models but, more importantly, exposed the significant practical limitations and performance bottlenecks that currently make them challenging to deploy effectively in high-performance, real-time security contexts within an AI gateway or load balancer.

Defensive Implications

Jakub Suchy's insights into LLM prompt security carry profound defensive implications for organizations deploying AI. The primary takeaway is that while AI gateways are undeniably necessary for managing and securing LLM interactions, the current state of LLM-based prompt security models presents significant performance hurdles that must be addressed. Defenders cannot simply "bolt on" existing large classification models and expect production-ready performance.

Firstly, the observed latency of 150-200ms per 500 tokens and the drastic reduction in Requests Per Second (RPS) mean that high-throughput applications cannot rely solely on these large models for real-time prompt filtering. Security architects must consider a multi-layered approach. This might involve an initial, very fast layer of defense—perhaps using highly optimized, smaller models or even enhanced heuristic filters—to catch the most egregious or common attacks, thereby reducing the load on the slower, more comprehensive LLM-based classifiers.

Secondly, the talk highlights the limitations of traditional text filtering. Since LLMs can correctly interpret mangled or typo-ridden words that would bypass static filters, defenders must move beyond simple blocklists. Semantic understanding is crucial for effective prompt security, but achieving this at scale and speed remains an open challenge. This implies a need for continuous research and development into more robust, yet performant, filtering mechanisms.

Thirdly, organizations should recognize that the field of LLM security is still in its "2003 of OWASP Top 10" phase. This means that security practices, tools, and vulnerabilities are rapidly evolving. Defenders must adopt an agile and adaptive security posture, continuously monitoring for new attack vectors and updating their defenses. Investing in research, staying abreast of industry developments (like new models from Meta or Google), and collaborating on open-source solutions will be critical.

Finally, Suchy's call for "smaller models that can run on a load balancer and can run much quicker" is a direct directive for innovation. Defenders should advocate for and invest in the development of specialized, lightweight LLM security models that are purpose-built for high-performance classification, potentially leveraging techniques like knowledge distillation or extreme quantization. Until such models are widely available, a pragmatic approach involves careful architectural design, potentially offloading the most intensive prompt security checks to asynchronous processes or to dedicated, horizontally scalable inference clusters, rather than embedding them directly into the critical path of a load balancer. The security of AI is not a solved problem; it requires a proactive, research-driven, and multi-faceted defensive strategy.

Key Takeaways

  • LLM Prompt Security is Nascent but Critical: The field of LLM security is currently akin to web security in 2003, with significant vulnerabilities and evolving defensive strategies. AI gateways are essential, but prompt security is often an underdeveloped component.
  • LLM-based Security Models Introduce Significant Latency: Running sophisticated LLM classification models (e.g., Llama Guard, Shield Gemma) directly within an AI gateway or load balancer incurs substantial latency (150-200ms for 500 tokens, 1 second for 2000 tokens) and severely limits Requests Per Second (RPS) to below 60, making them impractical for high-throughput, real-time environments.
  • Current Optimizations Are Insufficient: While techniques like optimized inference engines offer some improvement (around 30%), they do not fundamentally resolve the performance bottleneck. Token caching, effective for generative AI, is less applicable to classification tasks.
  • Traditional Text Filtering is Inadequate: Simple keyword or regex-based filters are easily bypassed by minor linguistic variations like typos, which LLMs can still interpret correctly, necessitating more advanced, semantic security solutions.
  • Need for Innovation in Smaller, Faster Models: There is an urgent requirement for the development of specialized, smaller LLM security models that can run much quicker on infrastructure like load balancers, specifically designed for high-performance classification.
  • Adopt a Multi-Layered and Adaptive Security Posture: Organizations must implement a multi-layered security strategy for LLM interactions and maintain an agile approach to adapting defenses as the landscape of AI threats and solutions continues to evolve rapidly.

About the Speaker(s)

Jakub Suchy is a key figure at Haroxy Technologies, a company renowned for developing HAProxy, the legendary open-source software load balancer widely utilized across various industries. His role at Haroxy positions him at the forefront of network infrastructure and performance, providing him with a unique perspective on integrating emerging technologies like AI into critical systems. Suchy's expertise spans the realms of high-performance computing, load balancing, and, as demonstrated in this talk, the nascent but crucial domain of AI security. His work involves engaging with customers and researching cutting-edge solutions to the challenges posed by new technological paradigms, making him a credible voice on the practical implications of LLM deployment and security.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This sponsored keynote from Haroxy Technologies delivers a brutally honest, data-driven assessment of LLM prompt security in AI gateways. Jakub Suchy presents empirical evidence demonstrating the severe performance bottlenecks introduced by current LLM-based classification models like Llama Guard and Shield Gemma. While it doesn't reveal a zero-day, it provides critical, actionable intelligence for architects and security professionals, challenging the naive assumption that AI can secure AI without significant engineering and performance overhead. It's a refreshing deviation from typical vendor marketing, grounded in real-world testing.

Heather Calloway (CISO) — STRONG ACCEPT

This talk by Jakub Suchy delivers a critical, data-driven assessment of the practical limitations of integrating LLM-based prompt security models into AI gateways. It precisely identifies a significant performance bottleneck that directly impacts the feasibility of real-time defenses against prompt injection and other AI-specific threats. While it doesn't offer a ready-made solution, it provides invaluable clarity on an urgent problem, forcing a pragmatic re-evaluation of current architectural strategies and highlighting a pressing need for innovation in lightweight, specialized security models.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025