Cybersecurity meets Generative AI: Automating Your Compliance...
Rafae Bhatti
BSidesSF 2024 · Day 1
Overview
This talk, presented by Rafae Bhatti at BSidesSF 2024, delves into the transformative potential of Generative AI to revolutionize the traditionally arduous and manual processes of cybersecurity compliance and audit. Titled "Cybersecurity meets Generative AI: Automating Your Compliance...", the session introduces the concept of an "agentic AI approach" to build a "compliance co-pilot." The core premise is to leverage advanced AI capabilities to synthesize organizational context, retrieve knowledge, identify and remediate compliance gaps, and intelligently validate evidence, thereby significantly reducing the dependency on manual professional services.

Key moments
- 06:00 Red Pill: Current State of Audit Automation
- 08:00 Evidence Collection: 75% Manual Documents
- 10:00 Hardest Problem: Encoding Domain Expertise
- 16:00 Profile-Based Compliance Identification (SOC2, HIPAA, GDPR, PCI)
- 19:00 Automated Success Criteria Validation for Evidence
- 20:00 Five Core Requirements for Agentic AI in Compliance
- 25:00 Agentic Architecture Overview (LangChain, GPT-4, Vector DB)
Cybersecurity meets Generative AI: Automating Your Compliance...
Speakers: Rafae Bhatti
Conference: BSidesSF 2024
YouTube: https://www.youtube.com/watch?v=dvgiHxrlwM0
Overview
This talk, presented by Rafae Bhatti at BSidesSF 2024, delves into the transformative potential of Generative AI to revolutionize the traditionally arduous and manual processes of cybersecurity compliance and audit. Titled "Cybersecurity meets Generative AI: Automating Your Compliance...", the session introduces the concept of an "agentic AI approach" to build a "compliance co-pilot." The core premise is to leverage advanced AI capabilities to synthesize organizational context, retrieve knowledge, identify and remediate compliance gaps, and intelligently validate evidence, thereby significantly reducing the dependency on manual professional services.
The talk highlights the critical need for such innovation by exposing the current inefficiencies in compliance audits, which are often people-intensive, time-consuming, and costly, frequently consuming two to three months per audit cycle. Rafae Bhatti, drawing from a decade of experience in building compliance programs and his background as a licensed attorney, articulates the pain points and the strategic imperative for automation. The proposed AI-driven co-pilot aims to shift compliance from an ad-hoc, reactive process to a repeatable, proactive, and more efficient function, ultimately enabling organizations to navigate complex regulatory landscapes like SOC 2, HIPAA, GDPR, and PCI with greater agility and accuracy.
Background
▶ Watch: Red Pill: Current State of Audit Automation (06:00)
The landscape of cybersecurity compliance today is characterized by significant manual effort and a heavy reliance on professional services. Rafae Bhatti, with his extensive experience in technical advisory and operational roles, including building compliance programs for Bay Area startups, and his unique perspective as a licensed attorney, underscores the inherent challenges. He notes that the rigor of regulatory obligations, such as HIPAA and GDPR components increasingly integrated into audits like SOC 2, demands a more robust and efficient approach than current methodologies provide.
Bhatti illustrates the "painful reality" of current automation levels across key audit stages. While tools offer a "tentative list" of controls, a substantial 80% of the effort in control mapping still requires manual review to tailor the list to a specific environment. Framework mapping sees better automation, around 50%, but validation remains a human task. For test identification, tools can integrate with APIs (e.g., GitHub, AWS) to suggest tests for settings like encryption or branch protection, yet a manual pass is still necessary to confirm correctness.
The most significant bottleneck, however, lies in evidence collection, which Bhatti describes as "where it sucks again." He states that only 20-25% of an audit's evidence can be automated via API connections. The remaining 75% comprises documentary evidence—items like board meeting minutes, codes of conduct, Disaster Recovery policies, and incident response plans. Gathering this evidence often involves extensive manual effort, sifting through internal repositories, Slack messages, emails, and calendar invites. Finally, the review phase, where CPAs assess submitted evidence, is entirely manual, leading to two to three weeks of waiting. This entire process can span two to three months per audit, a cycle that can repeat multiple times a year, effectively trapping organizations in a continuous audit loop.
Bhatti identifies two primary reasons for this lack of automation. While training data for internal organizational repositories is a challenge, he argues it's "not actually the hardest one." Policies and evidence types, he contends, exist within a "small universe of permutations," making fine-tuning a model to recognize correct evidence types manageable. The hardest problem to solve is domain expertise. A truly successful compliance co-pilot must "mimic how an auditor thinks, how a compliance professional thinks." This involves not just retrieving information but making nuanced decisions based on compliance criteria. Furthermore, the co-pilot needs generative AI capabilities to proactively provide templates for missing documents, pre-populated with organizational context, moving beyond mere gap identification to gap remediation.
This vision represents a shift from a "blue pill" scenario of "blissful ignorance" and manual processes to a "red pill" reality where AI transforms compliance. The "red pill" approach moves from a "people intensive" to a "resource intensive" process, from "information gathering" to "knowledge retrieval," from "gap identification" to "gap remediation," and from "ad hoc audits" to "repeatable" ones. This transformation is crucial for synthesizing organizational information and applying context, a key trend for CISOs highlighted in the conference's keynote.
Key Findings
▶ Watch: Hardest Problem: Encoding Domain Expertise (10:00)
The central findings of Rafae Bhatti's presentation underscore both the severe limitations of current compliance automation and the profound potential of Generative AI to overcome them.
- Predominance of Manual Effort in Audits: A critical finding is that despite advancements in GRC tooling, a staggering 75% of audit evidence collection remains manual, primarily involving documentary evidence such as policies, board minutes, and incident response plans. Only 20-25% of evidence can be automated via API integrations. This manual burden leads to audit cycles lasting two to three months, often repeated multiple times annually.
- Domain Expertise as the Primary Barrier: Bhatti identifies the most challenging aspect of automating compliance not as the availability of training data, but as the necessity to replicate human domain expertise. A successful AI co-pilot must be able to "mimic how an auditor thinks" and make nuanced decisions, moving beyond simple data retrieval to intelligent validation and judgment.
- The "Agentic AI" Opportunity: The talk proposes an "agentic AI approach" as the solution. This involves creating intelligent agents capable of understanding compliance requirements, retrieving relevant organizational context, and making decisions that align with auditor expectations. This approach aims to automate the production and validation of control lists, control mappings, and tests.
- Smart Validation and Gap Remediation: A key contribution of the proposed system is its ability to perform smart validation of evidence. Instead of merely collecting documents, the AI can assess whether evidence meets specific success criteria (e.g., checking if board meeting minutes include attendee lists or specific discussion topics). Furthermore, the generative aspect of the AI can proactively remediate gaps by drafting templates for missing policies, pre-populated with organizational specifics.
- Shift to Repeatable and Proactive Compliance: By automating significant portions of the audit process, the AI co-pilot enables a shift from ad-hoc, reactive audits to repeatable and proactive compliance. The system's ability to retrieve and leverage organizational knowledge across multiple, often overlapping, audit frameworks (e.g., SOC 2, HIPAA, GDPR, PCI) significantly enhances efficiency and reduces recurring costs.
- High Accuracy Requirements: For the AI co-pilot to be trustworthy and effective, its accuracy criteria must be exceptionally high, matching that of a human compliance professional. This necessitates careful model training and the incorporation of human feedback mechanisms to ensure reliable decision-making.
Technical Deep Dive
▶ Watch: Profile-Based Compliance Identification (SOC2, HIPAA, GDPR, PCI) (16:00)
The proposed solution for automating compliance audits centers on an agentic AI approach leveraging a compound AI system. Rafae Bhatti frames this architecture using the AWS 3P framework (Play, Pattern, Primitive) conceptually, even though the actual implementation uses different tools.
Conceptual Framework (AWS 3P):
- Play: The strategic goal is to displace traditional professional services with software, specifically targeting high-volume, high-value compliance work.
- Pattern: The architectural pattern is a compound AI system that integrates AI agents with Retrieval Augmented Generation (RAG). This involves model innovation through fine-tuning and an LLM wrapper for sophisticated prompt engineering.
- Primitive: While AWS Bedrock is mentioned as a conceptual primitive, the actual implementation for the proof of concept uses GPT-4 as the underlying large language model.
Retrieval Augmented Generation (RAG):
RAG is fundamental to providing the AI with the necessary organizational context. The process involves:
- Knowledge Base Creation: The current state of an organization—encompassing policies, ticketing systems, HR databases, identity management solutions, cloud provider configurations, and internal document repositories—is ingested.
- Embeddings and Vector Storage: This raw organizational data is converted into numerical representations called embeddings. These embeddings capture the semantic meaning of the data and are stored in a vector database (Vector DB). This vector DB acts as the organization's personalized knowledge base.
- Runtime Retrieval: When a query is made (e.g., "What is my level of compliance?"), the system retrieves relevant information from the vector DB based on the query's semantic similarity to the stored embeddings. This ensures that the AI's responses are grounded in the organization's specific context.
- System Settings Retrieval: Beyond documents, the RAG system can also retrieve real-time system settings (e.g., encryption status from a cloud provider API) to answer specific technical compliance questions.
Reinforcement Learning through Human Feedback (RLHF):
Given the sensitive nature of compliance decision-making, Bhatti emphasizes the importance of Reinforcement Learning through Human Feedback (RLHF). This mechanism is crucial for:
- Avoiding Biases: Ensuring the model's decisions are fair and unbiased.
- Improving Accuracy: Continuously tuning the model to make more correct decisions more often, particularly with high certainty, aligning with the stringent accuracy requirements of a compliance professional.
AI Agents:
The "agentic" aspect is where the system gains its intelligence and decision-making capabilities. AI agents are envisioned as "highly proficient" entities that can perform complex tasks faster than current methods. In the context of compliance, these agents are designed to mimic the thought process of an auditor.
Components for RAG + AI Agents (Proof of Concept Implementation):
- Orchestration Framework: LangChain is utilized to build and manage the interactions between the LLM, the knowledge base, and various tools.
- Large Language Model (LLM): GPT-4 was chosen after experimenting with other models, as it yielded the "best results" for the complex reasoning required.
- Vector Storage: For the initial proof of concept, FAISS (Facebook AI Similarity Search) was used as an in-memory vector store. However, Bhatti notes that a move to a "proper Vector database" would be necessary for production-scale deployments.
- Prompt Engineering: A zero-shot inference approach is employed for prompt engineering, allowing the model to respond to queries without explicit examples, relying on its pre-trained knowledge and the retrieved context.
Agentic Architecture (Detailed Flow):
The architecture integrates the rigor of traditional compliance with the power of generative AI:
- Compliance Repository: Contains standard control documents and descriptions.
- Embeddings & Vector DB: Embeddings are created from these control documents and stored in the Vector DB.
- Organizational Policy Documents: The organization's specific policies and other evidence documents are also processed into embeddings and stored.
- AI Agents for Gap Analysis: These agents perform the core work. They query the Vector DB, retrieve relevant control information and organizational evidence, and then conduct a gap analysis. This analysis determines the level of compliance against different evidence documents (e.g., policies, vendor contracts, system configurations).
- Decision Making: The agents are trained to make decisions based on success criteria. For instance:
- To verify "information security roles and responsibilities," an agent would retrieve the "Role-Based Access Control Matrix" from the policy DB, knowing it contains the necessary information.
- For "Unique IDs for information systems and networks," the agent would retrieve the "IAM policy setting" and assess its configuration.
- To validate "access removed upon termination," the agent would examine an "access review ticket" and confirm it's "closed with a comment that access was removed."
- For "encryption at rest," the agent would simply look for a Boolean setting indicating its status.
This architecture aims to automate five key requirements: producing a canonical control list, validating control mapping, producing and validating tests, retrieving documentary evidence, and remediating gaps.
Demo / Proof of Concept
▶ Watch: Five Core Requirements for Agentic AI in Compliance (20:00)
Rafae Bhatti presented an early-stage proof of concept to illustrate the capabilities of the compliance co-pilot. The demonstration focused on a hypothetical company, Acme Corp, with a specific profile: a SAS company in the healthcare industry in the US region with EU customers that accepts payments.
The goal of the demo was to show how the co-pilot could:
- Infer Applicable Frameworks: Automatically recognize that the company profile maps to SOC 2 (SAS), HIPAA (healthcare, US), GDPR (EU customers), and PCI (accepts payments). This eliminates the need for users to manually specify complex acronyms.
- Generate Tailored Control Lists: Produce a list of relevant controls specific to Acme Corp's environment, intelligently accounting for overlaps between the identified frameworks (e.g., between HIPAA and SOC 2).
- Assess Compliance Level: Determine the level of compliance for each control, categorizing them as "fully satisfying," "partially satisfying," or "not satisfying at all." For partial or non-compliance, the system provides an explanation.
Demonstrated Workflow:
- Inputting Company Profile: The user would provide a natural language description of Acme Corp's profile.
- Uploading Data: The system would be preloaded with Acme Corp's existing state, including policies, system configurations, and other relevant documents, forming its knowledge base.
- Control Framework Inference: The co-pilot automatically identified the applicable frameworks (SOC 2, HIPAA, GDPR, PCI) based on the profile.
- Control List Generation: The system then presented a consolidated list of controls, demonstrating how it had already accounted for overlaps between frameworks. For instance, it showed a list of SOC 2 controls followed by HIPAA controls, with the understanding that redundant controls had been rationalized. Bhatti visually emphasized the sheer volume of controls (e.g., 147 for SOC 2, 50 for HIPAA security, 75 for HIPAA privacy, 367 for PCI) to highlight the manual burden this automation alleviates.
- Compliance Check (Partial): The demo then moved to a partial compliance check. For controls that were "partially satisfying," the system provided a specific reason. An example given was a board meeting email that might be missing a list of attendees or a discussion on a particular topic, leading to partial compliance. Similarly, for non-compliance, it could identify a setting like "encryption is set to false."
Screenshots and Functionality (Briefly Shown):
Bhatti briefly showed screens illustrating the underlying processes:
- Control Mapping: Where the agent is set up and given prompts to map controls.
- Evidence Testing: Where different rules are applied, and the agent responds by explaining why specific evidence does or does not meet particular requirements.
It was explicitly stated that this demo is in a very early stage and "not something that's ready to use right away." The current implementation uses LangChain and GPT-4 with FAISS as an in-memory storage, and a zero-shot inference approach for prompt engineering. The next steps for development include scaling for production, using smaller, securely deployable LLMs, moving to a proper vector database, incorporating advanced prompt engineering, and integrating generative techniques for policy template creation.
Defensive Implications
▶ Watch: Agentic Architecture Overview (LangChain, GPT-4, Vector DB) (25:00)
The advent of a Generative AI-powered compliance co-pilot, as envisioned by Rafae Bhatti, carries significant defensive implications for organizations striving to maintain robust security postures and navigate complex regulatory environments.
- Proactive Compliance and Gap Remediation: Instead of reacting to audit findings, organizations can leverage the AI co-pilot for continuous, proactive compliance monitoring. The system's ability to identify missing policies or non-compliant configurations (e.g., encryption set to false) and then automatically generate pre-populated templates for remediation means that security gaps can be addressed before they become audit deficiencies or, more critically, exploitable vulnerabilities.
- Reduced Audit Fatigue and Resource Optimization: Compliance audits are notorious for consuming vast amounts of time and human resources, diverting security teams from core defensive tasks. By automating the laborious 75% of documentary evidence collection, control mapping, and initial evidence validation, the co-pilot can drastically reduce the 2-3 month audit cycle. This frees up GRC professionals, security engineers, and other stakeholders to focus on higher-value activities like threat hunting, incident response, and security architecture improvements.
- Enhanced Accuracy and Consistency: The AI's ability to perform "smart validation" against predefined success criteria ensures that evidence submitted is accurate and complete, reducing the back-and-forth with auditors. This leads to more consistent audit outcomes and a clearer understanding of an organization's compliance posture across multiple frameworks (SOC 2, HIPAA, GDPR, PCI).
- Improved Contextual Security: By synthesizing an organization's entire knowledge base (policies, system configurations, HR data, ticketing systems), the co-pilot provides a holistic, contextual view of security controls. This deep understanding allows for more intelligent decision-making regarding control implementation and evidence gathering, ensuring that security measures are not just compliant but also effective and tailored to the specific operational environment.
- Repeatable and Scalable Compliance Programs: The shift from ad-hoc, people-intensive audits to repeatable, AI-assisted processes makes compliance programs inherently more scalable. As organizations grow or face new regulatory requirements, the AI can quickly adapt and apply existing knowledge, reducing the overhead associated with expanding compliance efforts. This also ensures that compliance knowledge is institutionalized within the AI system, rather than residing solely with individual experts.
- Faster Response to Regulatory Changes: With its ability to process and understand regulatory texts, an AI co-pilot could potentially be updated to reflect new or amended compliance obligations more rapidly. This would allow organizations to quickly assess their current state against new requirements and initiate remediation efforts, maintaining continuous adherence in a dynamic regulatory landscape.
Key Takeaways
- Compliance Audits are Predominantly Manual: A significant 75% of compliance audit evidence collection, particularly for documentary evidence like policies and board minutes, remains a manual, time-consuming, and people-intensive process, often taking 2-3 months per audit.
- Domain Expertise is the Core Challenge: The most difficult barrier to automating compliance is not the availability of training data, but the need for AI to mimic human domain expertise and decision-making processes of auditors and compliance professionals.
- Agentic AI with RAG Offers a Solution: An "agentic AI approach" combined with Retrieval Augmented Generation (RAG) can synthesize an organization's context (policies, system settings, internal documents) to automate control mapping, evidence retrieval, smart validation, and even gap remediation.
- Intelligent Framework Inference and Tailored Controls: The compliance co-pilot can infer applicable regulatory frameworks (e.g., SOC 2, HIPAA, GDPR, PCI) from a natural language company profile and generate a tailored list of controls, accounting for overlaps, significantly streamlining the initial audit setup.
- Proactive Gap Remediation and Smart Validation: The AI can identify compliance gaps (e.g., missing policy details, incorrect system settings) and proactively suggest or generate templates for remediation, while also performing "smart validation" of evidence against specific success criteria.
- Shift to Repeatable and Efficient Compliance: This AI-driven approach promises to transform compliance from an ad-hoc, resource-heavy burden into a repeatable, proactive, and more efficient function, freeing up human resources for higher-value security tasks.
About the Speaker(s)
Rafae Bhatti is a seasoned professional with a decade of experience in building compliance programs within various Bay Area startups, where he has functioned in technical advisory and operational roles. His deep understanding of the pain points associated with compliance audits is informed by this extensive practical background. Further enhancing his expertise, Rafae is also a licensed attorney in California, which provides him with a rigorous understanding of regulatory compliance obligations. Currently, he serves as a CISO at a Silicon Valley startup. In addition to his corporate roles, Rafae maintains an independent blog where he provides legal advice, reflecting his thought leadership in the field.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk presents a compelling vision for an AI-driven compliance co-pilot, addressing the significant manual burden in current audit processes. By focusing on encoding domain expertise and leveraging agentic AI with RAG, the proposed architecture aims to automate evidence validation and gap remediation, moving beyond simple data retrieval. While the technical implementation details were high-level, the practical impact on GRC operations could be substantial.
Heather Calloway (CISO) — MUST SEE
This session offers a compelling vision for leveraging generative AI to fundamentally transform compliance operations. By automating the laborious processes of evidence collection, control mapping, and gap remediation, the proposed AI co-pilot directly addresses critical governance challenges and promises significant business impact. It provides a clear path for security leaders to move from reactive, ad-hoc audits to a proactive, continuously compliant posture, freeing up valuable resources and enhancing institutional accountability.