โ†
AI Governance, Risk & Red Teaming
Strategic ยท M24 ยท lesson 24 of 25 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
SOC 2 + AI - AICPA Attestation Engagements (2026 Practice)
๐Ÿ“–
now learning

SOC 2 + AI - AICPA Attestation Engagements (2026 Practice)

15 min

In late January 2026, the CFO of a 280-person U.S. AI-product company forwarded the CAIO a procurement questionnaire from a Fortune-500 prospect, $4.8M ARR over three years, an executive-sponsored 2026 close, with one line of text: "Section 7.3 is a gate." Section 7.3 read: "Vendor must provide a current SOC 2 Type II report extended with the AICPA AI Trust Services Criteria (August 2024 update), covering a continuous twelve-month operating window ending no earlier than ninety days before contract signature; bridge letter required for any gap exceeding sixty days." The provider held an ISO/IEC 42001 readiness binder from a Stage 1 prep with Schellman, an internal NIST AI RMF maturity score of 3.4, and a SOC 2 Type I (security-only) issued in 2024. It did not hold a SOC 2 Type II + AI-TSC report and would not for fourteen months at minimum. Two prospects in the same procurement cycle made the same demand within the next three weeks. By Q1 2026, the gap was no longer a compliance preference. It was the dominant revenue-gate for any AI vendor selling above the mid-market threshold. This lesson is the L4 leadership playbook for SOC 2 + AI under the AICPA Trust Services Criteria August 2024 update: the engagement lifecycle, the AI-TSC overlay, coordination with ISO 42001, the common findings, the bridge-letter pattern, the 12-section report structure, and how Acme Inc. ran a 22-month sequence from Type I + AI through Type II + AI issuance.

SOC 2 + AI Trust Services Criteria - The August 2024 Update and What Changed

SOC 2 is an attestation engagement performed under the AICPA's Statement on Standards for Attestation Engagements No. 18 (SSAE 18), codified in AT-C section 105 (concepts common to all attestation engagements) and AT-C section 205 (examination engagements). The engagement is performed by a CPA firm holding AICPA peer-review status; the deliverable is an opinion on management's assertion that the controls supporting service commitments and system requirements are suitably designed (Type I) or suitably designed and operating effectively (Type II) over the audit window. SSAE 18 governs evidence, sampling, materiality, and reporting; the AICPA's Trust Services Criteria (TSC) define the control objectives the auditor evaluates against. The TSC are organized into five categories: Security (mandatory baseline; all SOC 2 engagements), Availability, Processing Integrity, Confidentiality, and Privacy. An organization elects which of the five categories the engagement covers; Security is always included.

The August 2024 TSC update, published by the AICPA Assurance Services Executive Committee as an extension of TSP section 100, added an AI-specific overlay applicable across all five categories, with new points of focus and supplemental criteria addressing AI/ML system-specific risks. The overlay is not a separate sixth category; it is woven into the existing five through additional points of focus that auditors evaluate when the service organization operates AI/ML systems materially relevant to the service commitments. The AICPA companion document, the AI Risk Management Framework (AICPA "AI Risk Management Framework: Implementation Guide" 2024 + 2025 supplements), operationalizes the overlay with sample controls, evidence expectations, and sampling guidance for first-wave 2025-2026 engagements.

The AI-TSC overlay covers eight assertion families the auditor expects to evaluate where AI is in-scope:

  • Model lifecycle governance and accountability: documented lifecycle (intake, design, training, validation, deployment, monitoring, retirement); named owners; documented model inventory with risk tiering; governance committee with charter and minutes; tied to AICPA AI RMF Govern function and NIST AI RMF Govern; cross-walks ISO 42001 A.6 lifecycle and Clause 5.3 roles.
  • Training-data sourcing, quality, and provenance: documented data sources, licensing, characteristics; quality framework with measured metrics (completeness, accuracy, consistency, timeliness, validity, uniqueness); provenance metadata via ML-BoM / SBOM-AI (CycloneDX 1.7 + CDXA); cross-walks ISO 42001 A.7 and EU AI Act Article 10.
  • Model risk assessment and independent model validation (IMV): documented risk-assessment methodology applied per model; IMV per material model with documented scope, methodology, findings, sign-off; periodic IMV refresh (annual minimum for high-risk; SR 11-7 / OCC 2011-12 framing for financial services); cross-walks ISO 42001 Clauses 6.1.2/6.1.3, EU AI Act Article 9, NIST AI RMF Map + Measure.
  • Fairness and bias controls: documented fairness framework; protected-attribute analysis; measured fairness metrics (demographic parity, equalized odds, calibration-by-group as applicable to the model use case); periodic refresh; remediation workflow for material disparate impact; cross-walks NIST AI RMF Measure 2 and EU AI Act Article 10(2)(f-g) on bias examination.
  • Output monitoring and drift detection: documented monitoring framework with named methodology (KS test for input drift; PSI for prediction drift; performance-decay tracking for concept drift); defined thresholds; defined cadence (real-time, daily, weekly, monthly); escalation workflow; cross-walks ISO 42001 A.6.1.5 and EU AI Act Article 72 post-market monitoring.
  • Incident response specific to AI: AI-specific incident playbook (model failure, drift breach, fairness incident, jailbreak / prompt-injection, hallucination causing harm, third-party model incident); tabletop-tested; Article 73-aligned classification (15-day standard / 2-day critical-infrastructure widespread / 10-day serious-and-irreversible critical-infrastructure disruption); cross-walks ISO 42001 A.8.4 and NIST AI RMF Manage 4.
  • Third-party AI risk, AI-supplier inventory; per-supplier due-diligence with AI-specific addendum; foundation-model provider attestation reliance documented (auditor will not accept upstream SOC 2 / ISO 42001 as substitute without documented evaluation); substantial-modification change-control where supplier shipped a model update; cross-walks ISO 42001 A.10 and EU AI Act Article 25 + Annex XII.
  • Privacy and regulatory alignment in AI context, privacy-by-design in AI lifecycle; data-subject rights operationalized for AI training and inference; cross-border transfer documentation; GDPR Article 22 (automated decision-making) alignment; cross-walks ISO/IEC 27701 privacy management and CCPA / CPRA where applicable.

The overlay is broad enough to scope both foundation-model providers (training, evaluation, alignment, release, post-release monitoring) and deployers (integration, fine-tuning, deployment monitoring, third-party reliance), the AICPA's intentional design choice in 2024 to cover the full AI value chain under a single attestation framework. The same eight assertion families apply; the specific control inventory differs by role.

Type I vs Type II, the 8-Phase Engagement Lifecycle, and the 2026 Procurement Standard

SOC 2 reports come in two types reflecting the engagement scope. Type I opines on whether the controls are suitably designed as of a point in time (typically the engagement letter date). The auditor evaluates control descriptions, walks the design, samples evidence that the design is operational, and issues an opinion on design suitability. Type II opines on whether the controls are suitably designed and operating effectively over an audit window, typically 6 to 12 months, with 12 months the procurement standard for B2B AI by Q1 2026. The auditor performs design evaluation plus operating-effectiveness testing across the window, sampling evidence of control operation at the population-level (every instance), event-level (each occurrence of the control activity), or sample-level (representative subset for high-volume controls).

The 2026 procurement reality (lesson 070): tier-A enterprise customers: Fortune 1000, healthcare systems, financial-services firms, U.S. federal contractors, increasingly require Type II + AI-TSC for AI vendors selling above mid-market threshold. Type I is read as a transitional artifact en route to Type II; tier-A procurement will sign a Type I-backed contract conditional on Type II issuance within a defined window (typically 12-18 months) but will not sign for a multi-year deployment on Type I alone. Tier-B customers (mid-market, regional, education) often accept Type I as adequate. Tier-C (small business, startup, internal/non-regulated use) may accept self-attestation.

The engagement lifecycle for first-time SOC 2 + AI is an 8-phase sequence:

  • Phase 1 - Readiness assessment with auditor (4-8 weeks). CPA firm conducts paid readiness review against the elected TSC categories plus AI-TSC overlay; outputs a gap memo classifying gaps as design (control missing or weak), operating (control exists on paper but no evidence of operation), or evidence (control operates but evidence is sparse or non-traceable). Schellman, A-LIGN, BDO, Coalfire, Schellman Compliance, Sensiba, A-LIGN, Prescient Assurance, and the Big-4 dominate by Q1 2026; AICPA peer-review status verified at engagement.
  • Phase 2 - Gap closure (3-9 months typical). The organization closes identified gaps; designs missing controls; establishes evidence-collection workflow with traceable artifacts; aligns AI-specific controls with the AI-TSC overlay; coordinates with parallel ISO 42001 or NIST AI RMF programs to avoid duplication.
  • Phase 3 - Type I engagement (4-8 weeks fieldwork + 4-6 weeks report). The auditor performs design evaluation as of the engagement letter date; samples evidence of control existence and operation; issues the Type I opinion. Type I is the design milestone: customer-facing, contractable, but transitional.
  • Phase 4 - Audit-window setup (concurrent with Phase 3 close). The audit-window for Type II is named in the Type II engagement letter, typically the day after Type I or earlier. The organization confirms evidence-collection cadence, sampling completeness, control-operation logs, and exception-tracking workflow for the audit-window period.
  • Phase 5 - Type II audit-window operation (6-12 months). The organization operates controls and collects evidence across the window. Quarterly internal checkpoints (parallel to ISO 42001 surveillance cadence where applicable) catch control breakdowns, evidence gaps, exception accumulation. Quarter-end management reviews close the loop. Population-level controls (e.g., every model release passes gate review) generate inevitable evidence; sample-level controls (e.g., quarterly fairness refresh) require disciplined cadence.
  • Phase 6 - Fieldwork (6-10 weeks). The auditor performs the operating-effectiveness testing, sampling evidence across the window; interviewing control owners; reperformance for high-risk controls; evaluation of exceptions. AI-TSC sampling emphasizes the eight assertion families plus the elected TSC category control evidence. Sample sizes follow AICPA guidance scaled by population: e.g., a model-release gate with 24 releases over the year may sample 8-12 (judgmental + statistical); a quarterly fairness refresh population of 4 events samples all 4.
  • Phase 7 - Management response (2-4 weeks). The auditor presents preliminary findings: exceptions, qualifications, suggested management responses. Management reviews, contests classifications, presents additional evidence, drafts the formal management response that appears in the report.
  • Phase 8 - Issued opinion + bridge letter (2-6 weeks). The auditor issues the Type II report: unqualified, qualified, adverse, or disclaimer of opinion. Most well-prepared first-time Type II + AI engagements receive unqualified opinions with 1-4 exceptions documented. The audit period ends; the report-issuance lag (typically 30-60 days) creates a gap between end-of-audit-period and report availability that customers want covered by a bridge letter (covered in a later section).

End-to-end for first-time SOC 2 + AI: 18-30 months readiness through Type II issuance. Mid-market fees: $40K-$80K for Type I, $80K-$180K for Type II; AI-TSC overlay adds 15-25% to base fees reflecting expanded sampling. Re-issued annually (continuous audit window with rolling reports); Year 2+ Type II fees typically run at 70-90% of Year 1 reflecting maturity of evidence collection and reduced fieldwork.

ISO 42001 Overlap, Coordination Strategy, and the One-Evidence-Pack-Many-Frameworks Pattern

SOC 2 + AI and ISO/IEC 42001 are complementary, not duplicative. The frameworks share 60-75% control-evidence overlap by Q1 2026 first-wave practitioner data (Schellman, A-LIGN, BDO, BSI, KPMG combined engagement experience; ISACA AI Audit working group 2026 benchmark; AICPA AI RMF implementation cohort). The overlap is structural: both frameworks require model inventory, lifecycle governance, training-data documentation, monitoring with defined methodology, incident response, third-party due-diligence, internal audit. The frameworks diverge in form: ISO 42001 is a management-system certification (Clauses 4-10 + Annex A 38 controls; certificate validity 3 years; surveillance annually); SOC 2 is an attestation engagement (TSC categories + AI-TSC overlay; opinion on point-in-time or window; re-issued annually).

The strategic implication for L4 leadership: one evidence pack, many frameworks. The same model-card artifact in the AIMS portal serves ISO 42001 A.6.1.6, SOC 2 + AI lifecycle governance, EU AI Act Article 11 + Annex IV, NIST AI RMF Map 2, and AICPA AI RMF documentation. The same fairness refresh record serves ISO 42001 A.6.1.5 / A.6.2 monitoring, SOC 2 + AI fairness and bias controls, NIST AI RMF Measure 2. The same A.8.4 incident-runbook tabletop serves ISO 42001 A.8.4, SOC 2 + AI incident response, EU AI Act Article 73 readiness, NIST AI RMF Manage 4. Building each artifact once with explicit cross-framework references is the operational efficiency that turns four parallel programs into one program with four output reports.

Coordination strategy:

  • Same auditor for SOC 2 + AI and ISO 42001, maximum overlap. Schellman, A-LIGN, BDO, and KPMG operate both practices. A single engagement team that knows the organization's evidence-portal architecture conducts both audits with cross-walked sampling; the Schellman methodology explicitly reuses overlapping infrastructure-control evidence (lesson 091) reducing SOC 2 + AI fieldwork 15-25% and ISO 42001 Stage 2 fieldwork 15-25%. Combined annual fee envelope is materially below the sum of separate engagements.
  • Different auditors, carefully managed evidence sharing. Where the organization prefers different firms (e.g., Schellman for SOC 2, BSI for ISO 42001 + Article 31 notified-body), the evidence portal is shared via read access; both firms reference the same artifacts; the organization maintains the master cross-walk table mapping each control to both frameworks; engagement teams coordinate timing to avoid sample-evidence collisions but operate independently for opinion-independence purposes (ISO/IEC 17021-1 for the management-system audit; AICPA professional standards for the attestation).
  • Bundled engagements. By Q1 2026 Schellman, A-LIGN, BDO, Coalfire, and a growing number of mid-tier firms offer bundled SOC 2 + ISO 42001 + AI engagements at 80-90% of summed individual fees. KPMG bundles SOC 2 + ISO 42001 + ISO 27001 + ISO 27701 + EU AI Act readiness under a single advisory-to-attestation engagement. The bundle is the default operational choice for first-time AI-vendor multi-framework programs by mid-2026.

The L4 decision-frame: pick the framework portfolio matching the customer demand (SOC 2 Type II + AI for U.S. B2B; ISO 42001 + Annex VII Module H where Annex III high-risk; NIST AI RMF for U.S. federal contractor maturity; ISO 27001 + 27701 for security and privacy baseline), then choose audit-coordination architecture (same firm vs. different firms) optimizing for the operating overhead the organization can sustain. By Q1 2026 the data is consistent: organizations running a single coordinated program with one or two firms close audits 30-50% faster than organizations running parallel uncoordinated programs.

Common AI-TSC Findings in 2026 - The Seven Patterns First-Wave Engagements Surface

By Q1 2026 the first wave of SOC 2 + AI Type II engagements, ~300-500 organizations globally across Schellman, A-LIGN, BDO, Coalfire, KPMG, EY, PwC, Deloitte combined practice data; AICPA peer-review benchmarking; AICPA AI RMF implementation cohort, has produced seven recognizable AI-TSC finding patterns. Each pattern has a remediation approach and a cross-walk to the parallel ISO 42001 / NIST AI RMF / EU AI Act framing.

  • Finding 1 - No formal model lifecycle documented. The organization operates models but the lifecycle exists in MLOps tooling, Slack threads, and tribal knowledge; the auditor cannot trace intake โ†’ design โ†’ training โ†’ validation โ†’ deployment โ†’ monitoring โ†’ retirement for a sampled model. Remediation: document the lifecycle as a workflow with named gates, owners, evidence at each gate; populate a model inventory with risk tier, owner, deployment status, IMV status, monitoring status; cross-walk to ISO 42001 A.6 + NIST AI RMF Map + EU AI Act Article 17 QMS.
  • Finding 2 - IMV evidence sparse. The model-risk-management framework exists but IMV records are missing, partial, or non-independent (the developer also performed validation). Remediation: stand up the independent model-validation function or contract external IMV per material model; document scope, methodology, findings, sign-off; refresh annually for high-risk models; cross-walk to SR 11-7 / OCC 2011-12 for financial services, ISO 42001 Clauses 6.1.2/6.1.3, NIST AI RMF Measure.
  • Finding 3 - Fairness metrics not refreshed. The fairness framework exists with metrics defined but refresh cadence is annual-only or has slipped; the auditor samples a model where the most recent fairness refresh is more than 90 days stale. Remediation: define refresh cadence by model risk tier (monthly for highest-risk; quarterly default; annual minimum); operationalize via MLOps pipeline gates; track refresh completion in the model inventory; cross-walk to NIST AI RMF Measure 2 and EU AI Act Article 10(2)(f-g).
  • Finding 4 - Judge-model unvalidated. The organization uses an LLM-as-judge or LLM-as-evaluator pattern (e.g., GPT-4 grading other models' outputs) without documenting judge-model validation, calibration against human raters, or bias examination of the judge. Remediation: validate the judge model against a human-labeled gold standard; document agreement rates (Cohen's kappa or equivalent); refresh validation quarterly or on judge-model version change; cross-walk to NIST AI RMF Measure 2.3 trustworthiness measurement and AICPA AI RMF measurement framework.
  • Finding 5 - Article 73 incident-response coverage gap. The organization holds a security incident-response procedure but AI-specific incidents (model failure, drift breach, hallucination causing harm, third-party model incident) are not classified, escalated, or reported per Article 73 timing standards. Remediation: extend the incident playbook with AI-specific incident types, Article 73 classification (15-day standard / 2-day critical-infrastructure widespread / 10-day serious-and-irreversible critical-infrastructure disruption), regulator-contact log, tabletop-tested annually; cross-walk to ISO 42001 A.8.4 + EU AI Act Article 73 + NIST AI RMF Manage 4.
  • Finding 6 - Third-party foundation-model attestation reliance documented inadequately. The organization relies on OpenAI / Anthropic / Google / AWS Bedrock / Azure OpenAI service for foundation models and points to upstream SOC 2 / ISO 42001 / ISO 27001 as A.10.3 evidence. The auditor will not accept upstream attestation as a substitute without (a) documented evaluation of upstream report scope and applicability, (b) gap analysis identifying where upstream controls do not cover the deployer's operational risk, (c) compensating controls. Remediation: per foundation-model supplier complete due-diligence with AI-specific addendum; document Annex XII receivable verification where supplier is GPAI provider under EU AI Act; document compensating controls for any gap; cross-walk to ISO 42001 A.10 + EU AI Act Article 25 + Annex XII + NIST AI RMF Govern 6.
  • Finding 7 - Drift monitoring without thresholds. Dashboards display drift metrics (KS, PSI, performance decay) but there are no documented thresholds, no escalation workflow, no record of threshold reviews. The auditor cannot test the operating effectiveness of "we monitor drift" without documented thresholds and escalation evidence. Remediation: per model define drift thresholds with empirical justification, monitoring frequency (real-time for highest-risk; daily / weekly default; monthly for lowest-risk), escalation workflow with named owners, threshold-review cadence (quarterly); cross-walk to ISO 42001 A.6.1.5 + EU AI Act Article 72 + NIST AI RMF Measure 3 + Manage 4.

Each finding is classified at fieldwork close as an exception (control did not operate as designed for one or more sampled instances) or a deficiency (the control design itself was inadequate). Exceptions are documented in the report with management response; deficiencies may qualify the opinion. Most well-prepared first-time Type II + AI engagements close with 1-4 exceptions and an unqualified opinion; weak preparation produces 5-10+ exceptions, potential qualification, and customer-facing remediation pressure.

Bridge Letter, the 12-Section Report Structure, and Sample-Size Discipline

The bridge letter is the operational artifact that closes the gap between end-of-audit-period and report issuance. The Type II audit window ends on a defined date (e.g., December 31, 2026); the report typically issues 30-60 days later (e.g., February 15, 2027). During that lag, customer prospects asking "is the report current?" want assurance that no material change to the control environment occurred between window end and report issuance. The bridge letter, signed by the CEO, CFO, or CISO (depending on organizational convention), attests in writing that no material change has occurred; it is not an attestation engagement (no auditor opines) but a management representation. Tier-A procurement increasingly demands the bridge letter as a contract condition where contract execution falls within the 30-60 day lag.

The bridge-letter pattern extends into ongoing operation: where customers sign multi-year contracts, the bridge letter is reissued quarterly until the next Type II report covers the contracted period. The L4 discipline is to maintain the bridge-letter template, signature workflow, and quarterly issuance cadence as part of the SOC 2 + AI operating system, not as an ad-hoc artifact produced under customer pressure.

The 12-section SOC 2 + AI report structure, as standardized by AICPA SSAE 18 + August 2024 TSC update and as practiced by Schellman, A-LIGN, BDO, Coalfire by Q1 2026:

  • Section 1 - Independent service auditor's report. The opinion itself: unqualified, qualified, adverse, or disclaimer; engagement scope; auditor firm signature and date; AICPA peer-review reference.
  • Section 2 - Management's assertion. Signed management statement asserting controls are suitably designed (Type I) or suitably designed and operating effectively (Type II) over the audit window; identification of service commitments and system requirements; the auditor's opinion in Section 1 opines on this assertion.
  • Section 3 - System description. Detailed description of the service organization's system, infrastructure, software, people, procedures, data, including AI components (model inventory, eval pipelines, IMV process, third-party AI providers). The AI-specific subsections describe the eight AI-TSC assertion families in operational terms.
  • Section 4 - Service commitments and system requirements. Specific commitments the organization makes to customers (uptime, response times, security posture, AI-specific commitments: model availability, fairness commitments, transparency commitments).
  • Section 5 - Elected Trust Services Criteria. Which TSC categories the engagement covers (Security mandatory; plus elected Availability, Processing Integrity, Confidentiality, Privacy); AI-TSC overlay coverage if elected.
  • Section 6 - TSC and AI-TSC controls (per criteria). The control inventory, per elected TSC category and AI-TSC overlay, the specific controls the organization operates with control descriptions and ownership. The longest section in most reports (60-80% of report length).
  • Section 7 - Tests of controls and results. For Type II only: the procedures the auditor performed to test each control; sample sizes; results: operating effectively, exception noted, deficiency identified.
  • Section 8 - Exceptions and management responses. Each exception documented with control affected, nature of exception, instances observed, sample-size context, management response with remediation plan and timing.
  • Section 9 - Complementary user-entity controls (CUECs). Controls the customer must operate for the overall control environment to function, e.g., customer-managed access control, customer-managed data classification.
  • Section 10 - Subservice organizations. Third parties supporting the service (cloud providers, foundation-model providers, third-party AI tooling); inclusive vs. carve-out method documented; CUECs for subservice organizations.
  • Section 11 - Other information. Unaudited information the organization includes for context: typically AI-specific governance disclosures, customer-facing transparency documentation, alignment to ISO 42001 / NIST AI RMF / EU AI Act references.
  • Section 12 - Glossary and definitions. AI-specific terminology defined for the audit reader (foundation model, fine-tuning, drift, IMV, fairness metric definitions used in the engagement).

Sample-size discipline for AI-TSC sampling follows AICPA guidance scaled by population frequency and risk:

  • AI incidents over the audit window, all sampled (population test for low-frequency / high-impact); typical 0-12 over a 12-month window depending on operational maturity.
  • FRIA / impact assessments: 5-10 sampled across the in-scope system portfolio (judgmental); auditor traces methodology, signatures, integration with risk treatment.
  • IMV records, all sampled for material high-risk models (low population, high importance); 2-3 sampled for medium-risk models if population >5.
  • Vendor / third-party AI due-diligence: 5-8 sampled from the supplier inventory; auditor traces due-diligence questionnaire, contractual flow-down, Annex XII receivable verification for GPAI providers.
  • Model-release gate reviews, judgmental + statistical; e.g., population of 24 releases, sample 8-12.
  • Fairness refreshes, typically all sampled (quarterly population of 4 events per high-risk model is fully sampled; monthly population of 12 is judgmentally sampled at 6-8).
  • Drift threshold reviews, typically all sampled (quarterly population of 4 events) or judgmentally sampled (monthly population of 12 at 6-8).
  • Training records and acknowledgments (AI literacy under Article 4), judgmental sample of 25-40 employees with risk-tiered roles.

Statistical-significance considerations: SSAE 18 does not mandate statistical sampling for SOC 2; AICPA guidance permits judgmental sampling for most controls with documented rationale. AI-specific controls in 2026 lean toward population-testing for low-frequency / high-impact events (incidents, IMV) and judgmental sampling for higher-frequency control activities (release gates, drift reviews).

Acme Inc. Worked Example - The 22-Month Type I โ†’ Type II + AI Sequence

Acme Inc. a fictional 280-person AI-product company building a multi-tenant LLM-powered analytics product for U.S. mid-market financial services, ran the following 22-month sequence from initial readiness assessment through Type II + AI issuance, against the procurement gate that opened in January 2026.

  • Month 1 (January 2026) - Readiness assessment kickoff. CFO + CAIO signed an engagement letter with Schellman for SOC 2 readiness extended with the August 2024 AI-TSC overlay. Schellman dispatched a four-person team (lead partner, AI-TSC SME, infrastructure-controls SME, data-governance SME) for a four-week paid readiness review.
  • Month 2 (February 2026) - Gap memo issued. Schellman delivered a 47-page gap memo identifying 18 design gaps (model lifecycle not documented; IMV function not stood up; fairness metrics defined but refresh cadence undocumented; AI incident playbook missing; drift thresholds undocumented; third-party AI supplier inventory incomplete; data-governance framework thin on training-data provenance; etc.), 11 operating gaps (controls existed on paper but had no evidence trail), and 7 evidence gaps (evidence existed but was scattered or non-traceable). Remediation budget: $420K external + 1.2 FTE internal effort over 6 months.
  • Months 3-8 (March-August 2026) - Gap closure. The CAIO + Director of MLOps + AI Risk Officer + General Counsel ran the closure program. Model lifecycle documented as a 7-gate workflow integrated with the MLOps pipeline; IMV function stood up with two FTEs reporting to the Risk Officer (independent of model-development); fairness framework operationalized with quarterly refresh cadence enforced by pipeline gates; AI incident playbook drafted with Article 73-aligned classification and tabletop-tested in May 2026; drift thresholds defined per model with quarterly review cadence; third-party AI supplier inventory completed (foundation-model provider, vector-DB provider, evaluation-tooling provider); data-governance framework extended with training-data provenance via CycloneDX 1.7 + CDXA ML-BoM. Cross-walks to the in-progress ISO 42001 binder (Schellman parallel engagement) explicitly maintained.
  • Month 9 (September 2026) - Type I fieldwork. Schellman performed 6 weeks of Type I fieldwork against the elected TSC categories (Security baseline + Confidentiality + Availability) plus AI-TSC overlay. Two design exceptions identified, incomplete data-subject-rights workflow in AI inference logging, and incomplete substantial-modification change-control for foundation-model upstream version changes. Both remediated within 30 days.
  • Month 10 (October 2026) - Type I report issued. Unqualified opinion on suitable design as of August 31, 2026 (engagement letter date). Acme presented the Type I report to its prospects; both January-2026 prospects extended the procurement gate to Type II issuance by Q2 2027 with Type I-backed conditional signature.
  • Months 10-21 (October 2026 - September 2027) - Type II audit-window operation. 12-month continuous audit window. Quarterly internal management reviews of control operation; evidence-collection workflow maintained via the AIMS portal (single source of truth for ISO 42001 + SOC 2 + AI evidence); 4 AI incidents over the window (2 drift breaches, 1 hallucination causing customer-facing harm, 1 third-party model incident) all classified and reported per the playbook; IMV refresh completed for all 5 material models; quarterly fairness refresh completed; substantial-modification change-control exercised twice on foundation-model upstream version changes.
  • Months 22 (October 2027) - Type II fieldwork. 8 weeks of Schellman fieldwork sampling across the 12-month window. Sample-size discipline: all 4 AI incidents reviewed; all 5 IMV records reviewed; 4 of 4 quarterly fairness refreshes reviewed; 6 of 6 quarterly drift threshold reviews; 8 of 5 substantial-modification change-control records (2 events sampled at full population); 6 of 7 third-party AI suppliers due-diligence reviewed; 10 of 14 model-release gates reviewed.
  • Month 23 (November 2027) - Management response + report issuance. Four operating exceptions identified: (1) one quarterly fairness refresh slipped by 18 days against the 90-day target; (2) one drift threshold review documented but missing the named owner signature; (3) one third-party AI supplier due-diligence refresh not completed within the 12-month cycle (overdue by 41 days); (4) one substantial-modification change-control record missing the upstream-version evidence attachment (later supplied). All four exceptions documented in Section 8 with management responses and remediation; unqualified opinion issued.

End-to-end: 22 months from January 2026 readiness through November 2027 Type II issuance. All-in spend: $420K external remediation + 1.2 FTE internal over 6 months gap closure; Type I fee $58K; Type II fee $142K + 18% AI-TSC overlay premium ($25.6K); bridge letters drafted at quarterly cadence post-report. Both January-2026 prospects' contracts executed in Q4 2027 immediately on Type II issuance, $4.8M ARR confirmed. The 22-month sequence is the operational benchmark for first-time SOC 2 + AI under the August 2024 TSC update; Year 2+ steady-state runs at 60-75% of Year 1 effort with continuous 12-month rolling audit windows.

Cross-walk register for Acme's program: each AI-TSC overlay control mapped explicitly to AICPA SSAE 18 + TSC; AICPA AI Risk Management Framework; ISO/IEC 42001:2023 Clauses 4-10 + Annex A 38 controls; EU AI Act Articles 9, 10, 11, 12, 13, 17, 25, 26, 27, 71, 72, 73; NIST AI RMF Govern + Manage + Measure + Map; ISO/IEC 27001 + 27701 + 23894. The cross-walk register is the operational manifest enabling one-evidence-pack-many-frameworks delivery and is the single most operationally consequential L4 artifact in the program.

Key Takeaways

  • SOC 2 + AI under AICPA SSAE 18 + August 2024 TSC update: five baseline TSC categories (Security mandatory, Availability, Processing Integrity, Confidentiality, Privacy) plus an AI-specific overlay woven across all five via eight assertion families: model lifecycle governance, training-data sourcing/quality/provenance, model risk assessment + IMV, fairness + bias controls, output monitoring + drift detection, AI-specific incident response, third-party AI risk, privacy + GDPR alignment. The AICPA AI Risk Management Framework operationalizes the overlay.
  • Type I vs Type II: Type I opines on design at a point in time; Type II opines on design + operating effectiveness over a 6-12 month window. By Q1 2026 tier-A B2B procurement increasingly requires Type II + AI-TSC over a 12-month window; Type I is read as transitional, conditional-signature acceptable.
  • 8-phase engagement lifecycle: (1) readiness assessment with auditor (4-8 weeks); (2) gap closure (3-9 months); (3) Type I engagement (4-8 weeks fieldwork + 4-6 weeks report); (4) audit-window setup; (5) Type II audit-window operation (6-12 months); (6) fieldwork (6-10 weeks); (7) management response (2-4 weeks); (8) issued opinion + bridge letter (2-6 weeks). End-to-end 18-30 months first time; Year 2+ steady-state 60-75% of Year 1.
  • ISO 42001 overlap is 60-75% control-evidence; the operational efficiency is one-evidence-pack-many-frameworks. Same model card serves ISO 42001 A.6.1.6, SOC 2 + AI lifecycle, EU AI Act Article 11 + Annex IV, NIST AI RMF Map 2, AICPA AI RMF documentation. Same auditor for both = maximum overlap (15-25% fieldwork reduction each); Schellman, A-LIGN, BDO, KPMG offer bundled engagements at 80-90% of summed individual fees by Q1 2026.
  • Seven common AI-TSC findings in first-wave 2026 engagements: no formal model lifecycle documented; IMV evidence sparse; fairness metrics not refreshed; judge-model unvalidated; Article 73 incident-response coverage gap; third-party foundation-model attestation reliance documented inadequately; drift monitoring without thresholds. Each has a documented remediation with cross-walks to ISO 42001 / NIST AI RMF / EU AI Act framing.
  • Bridge letter pattern: covers the 30-60 day lag between audit-window end and report issuance; CEO/CFO/CISO signs management representation that no material change occurred; reissued quarterly until next Type II report covers contracted period. Tier-A procurement increasingly contract-conditions bridge-letter availability.
  • 12-section report structure: auditor's report; management assertion; system description (including AI components); service commitments; elected TSC; TSC + AI-TSC controls; tests and results (Type II); exceptions + management responses; complementary user-entity controls (CUECs); subservice organizations; other information; glossary.
  • Sample-size discipline: AI incidents all sampled; FRIA / impact assessments 5-10 judgmentally; IMV records all for high-risk; vendor due-diligence 5-8; model-release gates judgmental + statistical (8-12 of 24); fairness refreshes typically all (4 of 4 quarterly); drift threshold reviews typically all. SSAE 18 permits judgmental sampling with documented rationale; AI-specific low-frequency / high-impact controls lean toward population-testing.
  • SOC 2 + AI scopes both foundation-model providers and deployers: same eight assertion families, different control inventories; AICPA designed the overlay broad enough to cover the full value chain. By Q2 2026 foundation-model providers issuing Type II + AI is the procurement-acceptable upstream attestation evidence deployers cite (with documented evaluation, gap analysis, compensating controls).
  • Customer impact: tier-A B2B procurement increasingly mandates SOC 2 Type II + AI for AI vendors above mid-market threshold; SOC 2 + AI exceptions can trigger procurement loss, customer DDQ failure, reputational damage. Penalty exposure is procurement-revenue-loss, not statutory, but for AI vendors selling above mid-market, the dollar exposure exceeds most statutory regimes.
  • Acme 22-month worked example: January 2026 readiness through November 2027 Type II issuance; $420K external remediation + 1.2 FTE over 6 months gap closure; $58K Type I + $167.6K Type II (with 18% AI-TSC premium); 2 Type I design exceptions + 4 Type II operating exceptions; unqualified opinion issued; both January-2026 prospects' $4.8M ARR contracts executed Q4 2027 immediately on issuance. Cross-walk register mapping AI-TSC overlay to AICPA AI RMF + ISO 42001 + EU AI Act + NIST AI RMF + ISO 27001/27701/23894 is the single most operationally consequential L4 artifact.