AI AuditEvidence That Your AI Is Secure, Governed, and Defensible
An AI audit is a structured, evidence-based examination of the artificial intelligence systems an organization builds, buys, or embeds, covering the data they consume, the models they run, the way they are deployed and secured, and the controls that govern their use. Petronella Technology Group, Inc. has audited technology environments for regulated businesses since 2002, and our AI audit applies that same discipline to chatbots, copilots, agents, and machine learning pipelines.
- An AI audit examines systems that already exist. It is not a readiness assessment, which asks whether you are prepared to adopt AI, and it is not a certification audit, which is performed by an accredited third party against ISO/IEC 42001. It is an independent review of the AI you run today, with findings you can act on.
- Four domains, one evidence file. A credible audit covers data governance, model development and behavior, deployment and security, and compliance and oversight. Skipping any one of them leaves the exact gap a regulator, a customer, or a plaintiff's expert will find first.
- The frameworks already exist. The NIST AI Risk Management Framework, its Generative AI Profile, ISO/IEC 42001, and the OWASP Top 10 for Large Language Model Applications give an auditor objective criteria. An audit that invents its own criteria produces opinions, not evidence.
- Most findings are governance findings. In our experience the model itself is rarely the biggest problem. Unowned systems, undocumented data flows, missing human review points, and vendor terms nobody read are the recurring issues, and they are all fixable.
- The output is a defensible record. A written scope, a control-by-control finding set, risk ratings, evidence references, and a remediation plan. That record is what turns "we use AI responsibly" from a claim into something you can show.
What Is an AI Audit?
An AI audit is a systematic, independent examination of one or more artificial intelligence systems to determine whether they are secure, whether they behave as intended, whether the data feeding them was lawfully collected and appropriately handled, and whether the organization has the governance in place to prove all of that to someone outside the team that built it. The word "audit" matters. It signals evidence, criteria, and independence, rather than a demo and a conversation.
The scope has widened quickly. Three years ago an AI audit usually meant a fairness review of a single predictive model, typically in hiring, lending, or insurance. Today the typical business runs dozens of AI touchpoints: a customer-facing chatbot on a hosted large language model, copilots embedded in productivity suites, AI features switched on inside SaaS tools by a vendor update, retrieval-augmented generation over internal documents, and increasingly autonomous agents that read email, call APIs, and take actions. Each of those is an AI system with its own data path, its own failure modes, and its own accountability gap.
Petronella Technology Group runs AI audits for regulated businesses across the Research Triangle and nationwide: medical and dental practices under HIPAA, defense contractors under CMMC and NIST SP 800-171, law firms bound by confidentiality duties, and financial and professional services firms answering to SOC 2 customers. The method is the same one we apply in our cybersecurity audits, extended to the parts of an AI system that a traditional IT audit never touches.
- An inventory-first review: every AI system in scope, named and owned
- Testing against published criteria, not the auditor's preferences
- Evidence collection: configurations, logs, prompts, contracts, policies
- Hands-on security testing of the models and integrations that matter most
- A finding set with severity, likelihood, and a specific fix for each item
- A written report that a board, a customer, or a regulator can read
- A remediation roadmap sequenced by risk, not by convenience
- A vendor questionnaire returned with every box ticked
- A model accuracy report with no view of how the model is used
- An ISO/IEC 42001 certification, which only an accredited body can issue
- A one-time exercise that is never repeated as the systems change
- A tool scan that never speaks to the people who own the systems
- A readiness workshop about AI you have not deployed yet
Why Businesses Are Commissioning AI Audits in 2026
Four pressures arrived at once. Regulators moved first. The European Union's AI Act is phasing in obligations for providers and deployers between 2025 and 2027, and several U.S. states have passed or are debating laws governing automated decision-making and consumer-facing AI. Existing law never went away either: HIPAA still governs protected health information pasted into a chatbot, and CMMC still governs controlled unclassified information that an AI feature might index.
Customers moved second. Enterprise procurement teams now send AI vendor security questionnaires alongside the standard security questionnaire, and they ask for evidence rather than assurances. A completed AI audit is the fastest honest answer to those questions.
Attackers moved third. Prompt injection, data exfiltration through tool-connected agents, poisoned retrieval sources, and leaked system prompts are documented attack classes with public write-ups, and they do not show up in a conventional vulnerability scan. Our LLM security and AI red teaming work exists because these are real, and an audit is where a business first learns which of them apply.
Employees moved fourth, and often fastest. Shadow AI, meaning unapproved tools adopted by staff, means most organizations run more AI than leadership knows about. The inventory step of an audit is frequently the first time anyone has written the full list down.
Regulatory Drivers
EU AI Act phased obligations, state automated-decision laws, sector rules such as HIPAA and the FTC Safeguards Rule, and defense requirements under CMMC that reach any AI touching controlled data.
Contractual Drivers
Customer AI questionnaires, business associate agreements that never contemplated AI processing, cyber insurance applications that now ask about AI controls, and SOC 2 auditors asking how AI features were assessed.
Security Drivers
Prompt injection, retrieval poisoning, excessive agent permissions, insecure plugin and tool integrations, and model supply-chain risk from unvetted weights, datasets, and hosted endpoints.
Operational Drivers
Confidently wrong outputs reaching clients, undocumented dependencies on a single vendor, no rollback path when a model update changes behavior, and no one accountable when it does.
The Four Domains of an AI Audit
We structure every AI audit around four domains that map directly to the functions in the NIST AI Risk Management Framework (Govern, Map, Measure, Manage) and to the control themes in ISO/IEC 42001. Structuring the work this way means the findings can be cross-referenced to a framework a customer or regulator already recognizes, instead of living in a proprietary format that only we understand.
Data Governance
Where training, fine-tuning, and retrieval data came from, whether its collection was lawful and consistent with the notices you gave, whether it contains regulated categories such as protected health information or controlled unclassified information, how it is classified, retained, and deleted, and whether the vendor's terms allow it to be used for training. This domain also examines what leaves the organization in prompts and what comes back in outputs.
Model Development and Behavior
For models you built or fine-tuned: versioning, reproducibility, evaluation methodology, bias and fairness testing where the use case demands it, and documentation such as a model card. For models you consume: whether anyone evaluated the model for your use case, what the acceptance criteria were, how hallucination and drift are measured, and what happens when the provider ships a new version.
Deployment and Security
Identity and access to the AI system and to everything it can reach, secrets and API key handling, network placement, logging and monitoring, the permissions granted to agents and tools, defenses against prompt injection and data exfiltration, rate limiting and abuse controls, and the incident response plan for when an AI system misbehaves. This is where hands-on testing happens.
Compliance and Oversight
Whether an AI acceptable use policy exists and is enforced, who owns each system, where human review is required and whether it actually happens, how the organization maps to NIST AI RMF or ISO/IEC 42001, how sector obligations such as HIPAA, CMMC, or SOC 2 extend to AI, and whether the records exist to prove any of it.
AI Audit Checklist: What We Verify
The full audit workplan runs to several hundred test points depending on scope. The checklist below is the condensed version we share with clients before kickoff so they can see what evidence will be requested. If your team can answer every line with a document rather than a conversation, you are in better shape than most.
| Area | What we check | Evidence that satisfies it |
|---|---|---|
| Inventory | Every AI system, including embedded SaaS features and employee tools, is listed with an owner, purpose, data classification, and vendor | AI system inventory, reviewed within the last quarter |
| Accountability | A named executive owns AI risk; each system has a business and technical owner | Governance charter, RACI, or policy naming roles |
| Policy | An acceptable use policy defines approved tools, prohibited data, and review requirements | Signed policy, training records, enforcement evidence |
| Data inputs | Prompts, uploads, and retrieval sources are classified; regulated data is blocked or contractually covered | Data flow diagram, DLP or gateway configuration, vendor agreements |
| Vendor terms | Training on your data is disabled or contractually excluded; retention and subprocessors are known | Contract clauses, admin console settings, business associate agreement where PHI is involved |
| Model evaluation | Someone tested the model against the actual use case and recorded acceptance criteria | Evaluation results, test sets, model card |
| Fairness | Where outputs affect people's opportunities, bias testing was performed and documented | Impact assessment, disparity metrics, mitigation record |
| Access control | Least privilege applies to the system and to every tool, database, and API an agent can call | Role definitions, token scopes, access review log |
| Secrets | API keys and credentials are vaulted, rotated, and never embedded in prompts or client-side code | Vault configuration, rotation history, code review |
| Injection defenses | Untrusted content cannot silently redirect the model or trigger tool calls | Test results from our AI red team exercise, input handling design |
| Logging | Prompts, outputs, tool calls, and administrative changes are logged, retained, and reviewed | Log samples, retention policy, review procedure |
| Human oversight | Defined points where a person must review or approve output before it is relied upon | Workflow design, override records, escalation path |
| Monitoring | Drift, error rates, abuse, and cost anomalies are measured and alert someone | Dashboards, alert rules, on-call assignment |
| Incident response | The IR plan covers AI-specific incidents and names a kill switch for each system | AI incident response runbook, tabletop record |
| Framework mapping | Controls are mapped to NIST AI RMF or ISO/IEC 42001 and to sector rules | Control matrix, gap register, remediation plan |
Not Sure What AI You Are Actually Running?
The inventory is the hardest part for most organizations and the first thing we build. A short scoping call is enough for us to tell you whether you need a full audit, a focused security review, or a governance program first.
AI Audit vs. Readiness Assessment vs. Certification vs. Red Team
These four engagements are routinely confused, and buying the wrong one wastes budget. The table shows how they differ and when each is the right call. Petronella Technology Group offers all four, which means we have no incentive to sell you the one you do not need.
| Engagement | Question it answers | Timing | Output | Best for |
|---|---|---|---|---|
| AI audit | Are the AI systems we run today secure, governed, compliant, and provable? | After deployment, then periodically | Findings, risk ratings, evidence file, remediation roadmap | Any organization with AI in production or in employee hands |
| AI readiness assessment | Are we prepared to adopt AI well, and where should we start? | Before or early in adoption | Maturity score, use-case prioritization, adoption roadmap | Organizations planning their first serious AI investment |
| ISO/IEC 42001 certification | Does an accredited third party attest that our AI management system meets the standard? | After a management system has operated long enough to produce records | Certificate, audit report from the certification body | Organizations whose customers or regulators require formal certification |
| AI red teaming | Can an adversary make this specific system leak, misbehave, or take harmful actions? | Before launch and after major changes | Exploit evidence, attack narratives, prioritized fixes | High-exposure systems: customer-facing bots, agents with tool access |
The relationship between them is sequential more often than not. A readiness assessment shapes what gets built. An audit checks what was built and how it is run. Red teaming goes deep on the highest-risk systems the audit identified. Certification, if you pursue it, sits on top of a governance program the audit helped you prove was working. Our AI governance maturity model describes the progression in more detail.
How Petronella Technology Group Conducts an AI Audit
Our audit process was built by extending the evidence discipline we use for CMMC, HIPAA, and SOC 2 work to the AI-specific layers those frameworks do not reach. Craig Petronella, our founder, is MIT-certified in artificial intelligence, cybersecurity, and blockchain, holds a North Carolina Digital Forensics Examiner license (#604180-DFE), and serves as a cybersecurity expert witness. That forensic background shapes how we collect evidence: everything we rely on is captured in a form that would hold up if it were ever examined by opposing counsel.
Scope and Inventory
We interview business and technical owners, review procurement and SaaS admin consoles, and run discovery for unapproved tools. The result is a signed-off inventory of every AI system in scope, with data classification and a risk tier for each. Systems touching regulated data or making consequential decisions are tiered highest.
Criteria Selection
We agree the criteria in writing before testing starts: NIST AI RMF and the Generative AI Profile as the baseline, ISO/IEC 42001 control themes where a management system exists or is planned, the OWASP Top 10 for LLM Applications for security testing, and the sector overlays that apply to you, such as HIPAA, CMMC, or the FTC Safeguards Rule.
Evidence Collection
Policies, contracts, configurations, logs, prompts, evaluation records, and architecture documentation are gathered into a single evidence file with a reference number per item. Using our ComplianceArmor® platform, AI controls are documented alongside your existing compliance evidence so nothing lives in a separate silo.
Technical Testing
For the highest-tier systems we test rather than ask: prompt injection through every untrusted input path, data exfiltration attempts through tool integrations, permission scope of agents, secret handling, logging completeness, and the behavior of the kill switch. This is a targeted subset of our full AI red teaming engagement, sized to the audit.
Findings and Risk Rating
Every finding records the criterion it fails, the evidence we examined, the business impact, a likelihood rating, and a specific remediation. Findings are reviewed with system owners before the report is finalized so that factual errors are corrected and nobody is surprised in front of leadership.
Report and Roadmap
An executive summary written for a board, a detailed findings register for the technical team, a control matrix mapped to your chosen framework, and a remediation roadmap sequenced by risk. We offer a follow-up verification review once remediation is complete, and many clients fold the audit into an ongoing AI governance program.
What an AI Audit Typically Uncovers
Patterns repeat across industries. The list below is drawn from the categories of finding we see most often, described without client detail. The encouraging part is that nearly all of them are governance and configuration problems, which are far cheaper to fix than rebuilding a model.
AI Audits for Regulated Industries
Healthcare and Dental
Ambient scribes, patient-facing chatbots, and AI in the EHR all process protected health information. The audit confirms business associate agreements, minimum-necessary handling, and audit logging that satisfies the HIPAA Security Rule. Craig Petronella wrote How HIPAA Can Crush Your Medical Practice, and our HIPAA-compliant AI work applies that experience directly.
Defense Contractors
Any AI feature that can index, summarize, or search controlled unclassified information is inside your CMMC assessment boundary. As a CyberAB Registered Provider Organization (RPO #1449) we audit AI use against NIST SP 800-171 requirements so it does not become the finding that stalls your CMMC certification.
Law Firms
Generative AI in legal research and drafting has already produced sanctions for fabricated citations and disclosure of client confidences. The audit examines confidentiality controls, output verification, and the engagement terms your clients expect, informed by Craig's work as a cybersecurity expert witness.
Financial and Professional Services
SOC 2 customers and cyber insurers now ask specifically about AI controls. The audit produces the control matrix and evidence that answers those questions, and integrates with the vCISO program many of our clients already run with us.
What Determines the Cost of an AI Audit
We quote every AI audit after a scoping call rather than from a rate card, because the variables move the effort by an order of magnitude. The factors that matter most are the number of AI systems in scope and their risk tier, whether models were built in-house or consumed from vendors, how much technical testing the highest-tier systems warrant, the frameworks you need the findings mapped to, and the state of your existing documentation. An organization with a current inventory and an acceptable use policy is audited faster than one where the inventory has to be discovered from scratch.
Two things reduce cost predictably. The first is doing the inventory and policy work in advance, which our AI governance framework guide walks through. The second is scoping honestly: a focused audit of the three systems that touch regulated data is more valuable than a shallow pass over thirty. Every engagement is fixed-fee with 100 percent due at contract execution, so there are no surprise hours.
"Craig takes the time to understand our business model, not just our technology stack. It makes his recommendations more strategic and tailored to our actual goals."
Daniel Lee, TrustIndex verified review. Petronella Technology Group is rated 4.7 across 92 verified TrustIndex reviews and 5.0 across 15 Google reviews.
Why Petronella Technology Group for Your AI Audit
Auditing AI well requires three things that rarely sit in one firm: security testing skill, compliance and evidence discipline, and hands-on experience running AI in production. We have all three. Our team deploys private AI infrastructure and production AI agents for clients, so we audit systems we understand from the inside. We hold CMMC Registered Practitioner credentials across the team and have carried HIPAA, SOC 2, and NIST engagements to completion for over two decades. And our founder's forensics license and expert witness work mean our reports are written to survive scrutiny.
Craig Petronella is the author of Beautifully Inefficient, a book on artificial intelligence, human judgment, and where automation should and should not be trusted, and host of the Encrypted Ambition podcast, where AI governance and security are recurring topics. That perspective shows up in our audits as a bias toward practical, proportionate controls rather than compliance theater. You can find his full library on our books page and learn more about the company on our about page.
Since 2002
Founded in Raleigh, North Carolina in April 2002, BBB A+ rated since 2003.
Credentialed
CyberAB RPO #1449, CMMC-RP team, MIT-certified in AI and cybersecurity, NC Licensed Digital Forensics Examiner.
One Evidence Platform
ComplianceArmor® holds AI controls next to your HIPAA, CMMC, and SOC 2 evidence so auditors see one coherent picture.
Full Spectrum
Audit, red team, governance, and private AI deployment from one accountable team, described across our AI services.
Get a Scoped AI Audit Proposal
Tell us what AI you run, or think you run, and which regulations and customers you answer to. We will come back with a written scope, the criteria we would apply, and a fixed fee. No obligation, and the scoping conversation alone usually surfaces a few things worth fixing.
AI Audit: Frequently Asked Questions
What is an AI audit?
How is an AI audit different from an AI readiness assessment?
Which frameworks does an AI audit use?
How long does an AI audit take?
How much does an AI audit cost?
Do we need an AI audit if we only use vendor tools like copilots and chatbots?
Does an AI audit include penetration testing of the AI?
How often should an AI audit be repeated?
Know What Your AI Is Doing, and Be Able to Prove It
Whether you owe an answer to a regulator, a customer questionnaire, an insurer, or your own board, an AI audit from Petronella Technology Group produces the evidence. Call us, or schedule a consultation, and we will scope it together.
Petronella Technology Group, Inc. | 5540 Centerview Dr., Suite 200, Raleigh, NC 27606 | 919-348-4912 | info@petronellatech.com
Last Updated: September 3, 2026