Building Your Annual AI Audit Plan 

Most internal audit functions approaching AI risk make the same structural mistake: they treat AI as a single audit topic. One governance assessment, run once, checked off the plan. That approach is insufficient, and it will become more visibly insufficient as AI embeds deeper into operations and regulators sharpen their expectations.

The more durable approach is an annual AI audit plan, one made up of multiple targeted reviews across the AI lifecycle, each addressing a distinct risk domain, and together providing meaningful assurance that a single review cannot.

This piece sets out that portfolio. The prior piece in this series covered how to scope and run any individual AI audit using a NIST, AIUC-1, and OWASP crosswalk. This one zooms out: what should the full year of AI audit coverage look like?

Why a Portfolio Approach

AI systems are not static. They evolve. Models drift as data changes. New use cases are deployed. Vendors update algorithms without notice. Governance structures that were adequate six months ago may no longer reflect how the system is being used.

A single point-in-time audit cannot provide meaningful assurance over something that changes continuously. A portfolio approach addresses this by distributing coverage across the AI lifecycle, so that different risk domains are reviewed at appropriate intervals, and no single dependency is treated as a complete picture.

It also enables audit functions to build coverage incrementally, starting where risk is highest and expanding as capability matures.

The Five Audit Types

A well-structured annual AI audit plan consists of five targeted review types, each with a distinct scope and focus.

AI Governance Audit

The governance audit is where most functions should start, and where the highest-value assurance typically sits.

Scope: Evaluate policies, standards, and oversight structures for AI development and deployment. Assess the completeness and accuracy of the AI system inventory. Review ownership and accountability structures. Evaluate alignment between AI governance and enterprise risk management.

Why it matters: Governance failures are the most common root cause of AI-related risk. Weak governance does not show up in model outputs; it shows up in incidents, regulatory findings, and audit observations that could have been prevented. Strong governance also creates leverage, it reduces downstream risk across every other audit type.

Where to start in year one: If an AI governance audit has never been conducted, begin here. The findings from this review will inform the scoping and prioritization of every other audit type.

Data and Model Development Audit

Scope: Validate data sourcing, quality controls, and data governance practices feeding AI systems. Assess model design, testing, and validation processes. Review documentation, reproducibility, and change history.

Why it matters: AI systems are only as good as the data they are trained on. Biased, incomplete, or poorly governed data does not produce a bad model in an obvious way. It produces outputs that appear valid but reflect underlying distortions. This audit type surfaces those risks before they become material.

It also addresses a gap that governance audits do not reach: the technical and procedural controls around how models are built, not just how they are governed.

AI Deployment and Change Management Audit

Scope: Evaluate approval workflows for new AI deployments and material changes. Review version control and change tracking practices. Assess how AI systems are integrated into business processes and what controls exist at the point of integration.

Why it matters: Many AI-related control failures occur not during model development but during deployment and change. Systems go live without adequate review. Updates are made without triggering re-validation. Integration points introduce risks that neither the technical nor the business team fully owns.

This audit type addresses the transition risk that sits between development and production.

Monitoring and Performance Audit

Scope: Assess ongoing model monitoring practices. Evaluate drift detection and retraining processes. Review incident management for AI-related issues, including how anomalies are identified, escalated, and resolved.

Why it matters: Even a well-designed, well-deployed AI system can degrade over time. Data distributions shift. Business contexts change. Models trained on pre-pandemic data operate in a post-pandemic environment. Without active monitoring, these changes are invisible until the consequences surface.

This audit type addresses the ongoing risk that static, point-in-time reviews miss by design.

High-Risk Use Case Audits

Scope: Deep-dive reviews of specific AI applications carrying elevated regulatory, financial, or reputational exposure. Common examples include credit models, hiring tools, fraud detection systems, and pricing algorithms. Focus areas: fairness, explainability, regulatory compliance, and the controls specific to each application.

Why it matters: Not all AI systems carry the same risk. A scheduling tool and a credit decisioning model are both AI, but the assurance expectations are entirely different. High-risk use case audits ensure that the applications with the most significant impact receive proportionate attention, regardless of where they sit in the broader plan.

These reviews are also where regulatory exposure is most acute. NYC Local Law 144, Colorado SB21-169, and SEC disclosure expectations all create enforceable obligations that attach to specific use cases, not to AI in general.

Building Coverage Incrementally

For audit functions starting from scratch, trying to run all five audit types in year one is not realistic and is not necessary. A more effective approach is to sequence coverage based on risk and maturity:

  • Year one: AI Governance Audit first. It creates the foundation. The inventory, ownership clarity, and policy framework it surfaces directly inform the scope of everything else.
  • Year two: Add Data and Model Development, plus one High-Risk Use Case Audit focused on the application with the greatest regulatory or financial exposure.
  • Year three and beyond: Integrate Deployment, Change Management, and Monitoring reviews as the audit function builds technical fluency and organizational relationships.

The goal is a plan that is ambitious enough to close the assurance gap and realistic enough to execute.

How the Portfolio Fits Together

The prior piece in this series introduced a crosswalk model using NIST AI RMF, AIUC-1, and OWASP to scope and structure any individual AI audit. That model applies within each of the five audit types above.

The portfolio gives you the shape of the annual plan. The crosswalk gives you the methodology for each review within it. Together, they provide a framework that is both structurally sound and practically executable.

AI is not a one-audit topic. The functions that treat it that way will find themselves consistently behind, reacting to incidents and regulatory findings rather than providing assurance ahead of them. The ones building annual plans now, even imperfect ones, will be the ones positioned to close the gap.