AI Governance, Risk & Red Teaming
Proficient · M28 · lesson 28 of 31 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Reviewing Vendor Model Cards, SOC 2 + AI, and ML-BoM - Field Guide for L3 Practitioners
📖
now learning

Reviewing Vendor Model Cards, SOC 2 + AI, and ML-BoM - Field Guide for L3 Practitioners

15 min

Two weeks after Acme Inc's procurement team finished the 14-section AI Due-Diligence Questionnaire, the General Counsel asked a follow-up question. The DDQ had cleared Vendor A at Tier A and parked Vendor B at Tier D. The DDQ answers were on file. The SOC 2 Type II reports and ISO 42001 certificates had been collected as exhibits. The vendor's model card had been linked. The vendor's CycloneDX ML-BoM had been downloaded. The procurement team had signed the cover sheet. But the General Counsel, drafting the Article 26(1) "use in accordance with the instructions for use" memo, asked the specific question that turns paperwork into defensibility: "the DDQ tells me the vendor claims SOC 2 + AI and ISO 42001 and quarterly model-card refresh; the model card itself, the SOC 2 report itself, and the ML-BoM file itself are in the data room, has anyone actually read them, and do they say what the vendor said they say?" The procurement lead, the AI Governance lead, and the Privacy Officer sat together in a windowless conference room with three artifacts open on three screens: Vendor A's 31-page model card, Vendor A's 84-page SOC 2 Type II + AI Trust Services Criteria report, and Vendor A's 2,400-line CycloneDX 1.7 ML-BoM JSON with attached CDXA attestations, and worked through the cross-reading exercise that would close the Article 26 file. This lesson is the field guide for that exercise. It is the L3 capstone, lesson 071, that closes the L3 toolkit: the practitioner who can read these three artifacts the way a regulator reads them is the practitioner who can defend the Article 26 posture at audit.

The Three-Artifact Triangulation

A 2026 AI vendor producing an in-scope high-risk system should, when asked, hand the deployer three artifacts that together attest to the vendor's AI Act and ISO 42001 readiness: the model card (or system card), the SOC 2 Type II report with the AICPA AI Trust Services Criteria additions, and the CycloneDX 1.7 ML-BoM with CDXA signatures. Each addresses a different question. Each, on its own, is insufficient. The three together, read against each other, are the deployer's defensibility evidence.

The model card / system card answers the question: what is the model, what is it for, how does it behave, what are its known limits? The seminal Mitchell-et-al. (2019) "Model Cards for Model Reporting" paper defined the nine sections that have since become the de-facto industry baseline: model details; intended use; factors; metrics; evaluation data; training data; quantitative analyses; ethical considerations; caveats and recommendations. The 2024-2026 era added the system card: operational documentation covering deployment topology, safety mitigations, evaluation suite, and red-team summary, popularised by OpenAI and Anthropic. The model card is the AI-specific cousin of the datasheet that accompanies an electronic component or the package insert that accompanies a pharmaceutical product. It is mandatory under Article 13(2) for transparency to deployers and under Annex IV §2 as part of the technical documentation that Article 11 requires.

The SOC 2 Type II + AI TSC report answers the question: are the vendor's operational controls, for security, availability, confidentiality, processing integrity, privacy, and (under the 2024-updated AICPA AI Trust Services Criteria) AI-specific governance, data, and model lifecycle controls, designed appropriately and operating effectively over the audit period? A SOC 2 Type II is the AICPA-defined attestation report (under SSAE 18) issued by an independent CPA firm. The baseline TSC dates to the AICPA Trust Services Criteria (2017, revised 2022). The AI-specific additions were published by the AICPA in August 2024 (the AICPA Practice Aid "Considerations for Performing SOC 2 Examinations with AI System Specific Considerations") and align to the AICPA AI Risk-Management Framework released the same year. The SOC 2 + AI report is the operational analogue to the model card: the model card describes the model; the SOC 2 describes the organisation that builds, deploys, and operates the model.

The CycloneDX 1.7 ML-BoM with CDXA answers the question: what does this AI system actually consist of, and how do we verify the supply chain? CycloneDX is the OWASP-curated SBOM standard. The 1.6 release (2024) introduced ML-specific extensions; the 1.7 release (mid-2025) expanded them with first-class component types for ML models, datasets, evaluation suites, and model-card pointers. CDXA (CycloneDX Attestations) is the companion specification that allows the BoM to carry signed claims, issuer, claim, evidence, signature, using JOSE/COSE cryptography. The ML-BoM is the supply-chain analogue to the model card: the model card describes the model; the ML-BoM enumerates and cryptographically anchors every component (base model, fine-tune dataset, evaluation dataset, dependency library, signing key) that produced it.

The triangulation pattern is straightforward. Each artifact makes claims. Some claims appear in only one artifact (e.g., the SOC 2 details the auditor's control-testing procedures; the ML-BoM details cryptographic hashes; the model card details quantitative fairness metrics). But ten or so claims appear in all three artifacts: the model version, the training-data cutoff date, the evaluation-suite identity, the fairness-slice methodology, the OWASP / ATLAS coverage attestation, the vendor entity name, the certification scope, and so on. Those ten claims must be consistent across all three artifacts. Inconsistency is the audit signal: the auditor either has not finished the engagement, the vendor has not finished the documentation, or, worst, the vendor is asserting different facts to different audiences. Cross-reading is how the deployer detects this.

The Model-Card Review Checklist (15 Items)

The L3 practitioner reviews a vendor model card against a 15-item checklist. Each item is a yes / partial / no / not-applicable judgment, captured in the review memo with citations to the model-card page or section.

Item 1 - Mitchell-et-al. (2019) nine sections present. Walk the nine sections: (a) model details (developer, version, date, type, paper/citation, licence, contact); (b) intended use (primary uses, primary users, out-of-scope uses); (c) factors (relevant factor groups, evaluation factors); (d) metrics (model performance measures, decision thresholds, variation approaches); (e) evaluation data (datasets used, motivation, preprocessing); (f) training data (datasets, when known); (g) quantitative analyses (unitary results, intersectional results); (h) ethical considerations (data, human life, mitigations, risks and harms, use cases); (i) caveats and recommendations. A model card that omits sections is not a Mitchell-compliant card. The 2026 expectation is all nine sections; 8 of 9 is a minor finding with a CAR for the missing section; 7 of 9 is a material finding.

Item 2 - Annex IV §2(a)-(h) coverage. The EU AI Act Annex IV §2 lists the elements of technical documentation: (a) general description of the AI system; (b) detailed description of the elements including development tools, pre-existing components, computational resources; (c) design specification including key choices and trade-offs; (d) system architecture; (e) data sheet including training methodologies and data acquired; (f) human oversight measures; (g) accuracy and robustness performance metrics; (h) cybersecurity measures. The model card and Annex IV technical documentation can be separate documents, but the model card should map cleanly onto Annex IV §2 with explicit cross-references; where the model card is the primary technical-documentation evidence, every Annex IV §2 element must be addressable from the card.

Item 3 - Intersectional fairness slices. Mitchell §7 quantitative analyses distinguishes unitary results (performance broken out by a single attribute, gender, age, race, language) and intersectional results (performance broken out by combinations, Black women aged 18-25, non-native English speakers in regulated industries). Unitary fairness can mask intersectional failure, a model can satisfy demographic parity by gender and by race while failing badly for Black women. The 2026 expectation is intersectional slices for any model used in Annex III high-risk contexts, not just unitary. A model card that reports only unitary slices is a yellow flag for high-risk deployment.

Item 4 - NIST AI 600-1 12-risk applicability statements. NIST AI 600-1 (the Generative AI Profile, July 2024) enumerates twelve generative-AI-specific risk categories: CBRN information or capabilities, confabulation, dangerous or violent recommendations, data privacy, environmental impacts, harmful bias and homogenisation, human-AI configuration, information integrity, information security, intellectual property, obscene/degrading/abusive content, value chain and component integration. The vendor model card should issue an applicability statement for each, applies / does not apply / applies with mitigations, and disclose the residual-risk acceptance position. A card silent on NIST AI 600-1 is a 2026 procurement gap for any generative-AI system.

Item 5 - OWASP LLM Top 10 + ATLAS coverage. The OWASP LLM Top 10 (2025-2026 update; LLM01 prompt injection through LLM10 unbounded consumption) and the MITRE ATLAS (Adversarial Threat Landscape for AI Systems; current version v5.4.0) are the two converged threat taxonomies. The model card or its system-card companion should attest coverage of each OWASP entry and a representative sample of ATLAS techniques. Agentic systems add the OWASP Agentic Top 10 (December 2025) coverage. A model card silent on these is a procurement red flag for any LLM or agent.

Item 6 - Article 25(2) cooperation evidence. The model card should explicitly reference the vendor's Article 25(2) cooperation commitment to deployers: contact information, Annex IV documentation availability, change-notification cadence, and the deployer-information package per Annex XII (for GPAI providers).

Item 7 - Refresh cadence. Quarterly minimum for production high-risk systems. Monthly is preferred for systemic-risk GPAI under Article 51. The card should declare its refresh policy, its last refresh date, and its next scheduled refresh.

Item 8 - Version pinning. The model card must identify the precise model version it documents. Cards that describe "the Vendor X model family" without version pinning are not deployer-defensible: the deployer cannot demonstrate which version it tested, which version it deployed, and whether a substantial modification under Article 25(1)(a) has occurred.

Item 9 - Evaluation-data lineage. Mitchell §5 evaluation data should identify the datasets, their licences, their date cutoffs, and whether they overlap with the training data (data leakage). The 2026 expectation includes the evaluation-suite version identifier so the deployer can re-run independent evaluation if needed.

Item 10 - Training-data summary. Mitchell §6 training data describes data sources, dates, and any known biases. Under Article 53(1)(d) the GPAI provider must publish a training-content summary in the Commission template. The model card should reference or include this. Vendors that refuse to disclose any training-data information cannot be procured for high-risk use.

Item 11 - Known failure modes. The card should enumerate failure modes, known cases where the model produces incorrect, biased, harmful, or unsafe output, with example prompts and example mitigations. Cards that read like marketing brochures and disclose no failure modes are not regulator-credible.

Item 12 - Human-oversight design. Annex IV §2(f) requires human-oversight description. The card should describe the human-oversight design assumption: how the deployer is expected to deploy the system, where humans are expected to intervene, what thresholds trigger human review.

Item 13 - Environmental impact. Annex IV §2(b) computational resources combined with NIST AI 600-1 environmental-impact risk category invites disclosure of training compute and inference energy use. The 2026 expectation: at minimum a qualitative statement; the leading vendors provide quantitative figures.

Item 14 - IP and licence. Mitchell §1 model details should identify the licence under which the model is offered (proprietary, open-weights with commercial licence, open-source). The licence governs the deployer's rights to fine-tune, redistribute, and benchmark. Card silence on licence is a contractual gap.

Item 15 - Negative-use statement and prohibited uses. The "out of scope" sub-section of Mitchell §2 intended use should be explicit about prohibited uses (Article 5 categories at minimum: social scoring, untargeted biometric scraping, emotion recognition in employment and education, real-time RBI in public spaces, etc.) and high-risk-only deployments that require additional contractual conditions.

What to flag when items are missing. 15/15 is a clean review (Tier A inputs). 13-14/15 is a minor-finding review with one or two CARs. 10-12/15 is a material-finding review, the vendor must remediate before Tier A production deployment; Tier B with conditions is the most aggressive defensible posture. Below 10/15 is a procurement red flag, escalate to the AI Governance Committee and limit the vendor to non-high-risk deployment until remediated.

SOC 2 + AI Trust Services Criteria - Deep Read

A baseline SOC 2 Type II report runs 60-90 pages and follows a standard structure: independent service auditor's report; management's assertion; description of the system; the auditor's tests of controls; the results. The 2026 SOC 2 + AI report, the one that incorporates the AICPA's AI Trust Services Criteria updates and the AICPA AI Risk-Management Framework, adds an AI-specific section to the system description and additional AI-specific controls within the existing Trust Services Criteria categories. The L3 practitioner reads it section by section.

Section 1 - Independent service auditor's report (typically 2-3 pages). Read the audit firm name, the audit period (Type II minimum 6 months; 12 months is preferred), and the opinion. The opinion is one of four: unqualified (clean), qualified (one or more exceptions material enough to disclose but not pervasive), adverse (controls did not operate effectively), or disclaimer (auditor could not form an opinion). For procurement-readiness, an unqualified Type II opinion is the bar; a qualified opinion is reviewable with specific compensating controls; adverse or disclaimer is a procurement red flag.

Section 2 - Management's assertion (1-2 pages). The vendor's management formally asserts that the description is fair, the controls were designed and operating effectively, and the criteria were met. The signature, date, and signing officer matter, the assertion is the vendor's representation that supports the auditor's opinion.

Section 3 - Description of the system (10-30 pages). This is where the AI-specific additions are most visible. Look for: (a) AI system inventory, which models and AI systems are in scope; (b) AI data governance, how training data is governed, including provenance, quality, and lineage controls; (c) AI model lifecycle, how models are developed, evaluated, approved, deployed, monitored, and retired; (d) AI monitoring, drift, fairness, safety, performance; (e) AI incident response, how AI-specific incidents are handled, including Article 73 alignment for in-scope vendors; (f) AI third-party management, how the vendor manages its own upstream AI suppliers (the foundation-model provider, the training-data provider, the evaluation-suite provider). A SOC 2 + AI report that lacks these AI-specific system-description elements is operating under baseline SOC 2, not SOC 2 + AI, regardless of the cover-page title.

Section 4 - Tests of controls and results (30-60 pages). Each control has an objective, a description of the control activity, the auditor's test procedure, and the result (no exceptions, or specific exception details). The AI-specific controls cluster around: (i) training-data controls, provenance verification, licence compliance, opt-out compliance, sensitive-data filtering; (ii) model-evaluation controls, evaluation-suite definition, fairness testing, robustness testing, sign-off cadence; (iii) model-deployment controls, version pinning, change management, rollback capability; (iv) model-monitoring controls, drift detection, fairness drift, safety incident detection; (v) AI-incident controls, detection, classification, escalation, reporting (with Article 73 timelines for EU AI Act in-scope systems); (vi) AI third-party controls, supplier-onboarding diligence, supplier-monitoring, supplier-exit. Look at exceptions specifically, exceptions are where the controls did not operate as designed. A handful of low-severity exceptions in a 12-month audit period is normal; a cluster of exceptions in the AI-data-governance or AI-model-deployment controls is a procurement red flag.

Reading the exceptions. Each exception has a noted instance count (e.g., "5 of 25 model deployments lacked documented sign-off"), the root cause (e.g., "expedited deployment under emergency-change procedure"), and management's response (e.g., "additional approval step added to the change-control workflow in October 2025"). The auditor's note typically includes whether the deficiency was remediated by year-end. The L3 reviewer captures each exception in the review memo, flags any cluster, and asks the vendor for remediation evidence beyond what is in the SOC 2 itself.

Relationship to ISO 42001 and EU AI Act Article 17 QMS. SOC 2 + AI, ISO 42001, and the EU AI Act Article 17 QMS overlap but do not substitute for each other. ISO 42001 is the AI Management System standard (auditable by an accredited certifying body); SOC 2 + AI is the AICPA service-organisation attestation (auditable by a CPA firm); Article 17 QMS is the EU AI Act provider obligation (verifiable in conformity assessment). A vendor with all three is a Tier A vendor on the independent-assurance dimension. A vendor with two of three is reviewable. A vendor with only baseline SOC 2 and no ISO 42001 and no AI TSC is operating below the 2026 procurement bar for high-risk AI vendors; not catastrophic, but a deployer accepting that vendor must add compensating controls (deeper deployer-side monitoring, contractual right-to-audit, time-boxed remediation plan toward AI TSC or ISO 42001).

ML-BoM Verification (CycloneDX 1.7 + CDXA)

The CycloneDX 1.7 ML-BoM is a JSON or XML document with first-class ML extensions. The L3 practitioner verifies seven elements.

1. Schema and version. Open the file. Confirm "bomFormat": "CycloneDX" and "specVersion": "1.7". Earlier versions (1.5, 1.6) have weaker ML support; pre-1.6 ML-BoM should be requested in 1.7. The serialNumber field uniquely identifies the BoM instance; the version field tracks BoM iterations.

2. ML model components. The components array should include one or more components with "type": "machine-learning-model" (new in 1.6/1.7) or with the "mlModel" sub-object. Each model component carries: name, version, purl (package URL) or external reference, supplier, model-card pointer, training-data references, evaluation-suite references, and lifecycle status. A 2026 ML-BoM that lists the deployed model but not the upstream foundation model is incomplete, fine-tune lineage requires both.

3. Dataset components. Components with "type": "data" describe training, fine-tuning, and evaluation datasets. Each should carry name, source, version, licence, and a content hash (typically SHA-256). The DSM Directive Article 4(3) TDM opt-out compliance evidence appears here for datasets scraped under exceptions, either an attestation or a pointer to the opt-out compliance memo.

4. Dependency components. The vendor's software dependencies, inference framework versions, tokeniser versions, post-processing libraries, appear as standard CycloneDX components with "type": "library", version, licence, hashes, and vulnerability references (where applicable, via the optional vulnerabilities array or external pointer to VEX or CVE feed).

5. CycloneDX 1.7 ML-extension pointers. CycloneDX 1.7 adds first-class pointers for: model-card pointer ("externalReferences" with type "model-card" or the dedicated card field), training-data pointer (dataset references), evaluation-suite pointer (eval-suite references; new in 1.7), vulnerability pointer (external CSAF or VEX reference). The L3 practitioner walks the pointers, follows them, and confirms the linked artifacts actually exist and resolve.

6. CDXA attestations and signatures. CycloneDX Attestations (CDXA) is the cryptographic-claim companion to CycloneDX. A CDXA-attached ML-BoM carries one or more attestation objects: issuer (who is making the claim, typically the vendor or an upstream supplier), assertion (the textual claim, e.g., "this base model is GPT-4o version 2025-08-01"), evidence (cryptographic hashes or external references), and signature (JOSE-format JWS or COSE-format COSE_Sign1, typically signed with the vendor's Sigstore identity or a long-term x.509 key). The L3 practitioner verifies each signature using the publicly published signing key. Unsigned ML-BoMs are downgraded, the deployer cannot establish whether the BoM contents were tampered between vendor publication and deployer receipt.

7. Hash verification. For each binary artifact referenced by hash (model weights, dataset, dependency), the deployer can re-fetch the artifact and recompute the hash. Hash mismatches indicate either tampering, substitution, or vendor error, all three require vendor follow-up before deployment.

What ML-BoM does and doesn't claim. The ML-BoM is an inventory and integrity artifact. It claims: these are the components; these are their hashes; these are their licences; these are their suppliers; (with CDXA) here are signed assertions about specific claims. It does not claim: that the model performs well, that the training data is bias-free, that the evaluation suite is comprehensive, that the upstream supply chain is itself trustworthy. The model card and SOC 2 + AI report supply those claims. The ML-BoM is the supply-chain spine; the model card is the behavioural specification; the SOC 2 + AI is the operational attestation. Each closes a different audit question.

What to flag when missing. No ML-BoM in 2026 is a procurement red flag, request CycloneDX 1.7 export; if the vendor cannot produce one, escalate. Unsigned ML-BoM is a yellow flag: request CDXA attestation or, alternatively, an in-scope SOC 2 + AI control over the ML-BoM generation pipeline. ML-BoM that lists only the deployed model and omits upstream foundation-model and training-data components is an incomplete BoM, request completion.

Worked Example - Acme Inc Re-Examines Vendor A and Vendor B

Acme Inc returns to the data room with the three artifacts from each vendor. The DDQ from lesson 070 already tiered Vendor A at Tier A and Vendor B at Tier D. The artifact-cross-reading exercise confirms or contradicts those tiers.

Vendor A, model card. 31 pages, Mitchell-et-al. nine sections present (15/15 checklist items pass except Item 13 environmental impact which is qualitative only, minor finding, CAR opened). Annex IV §2(a)-(h) explicitly cross-referenced in an appendix. Intersectional fairness slices for the eight protected-class combinations relevant to the deployer's customer-service-AI use case. NIST AI 600-1 applicability statements for all twelve categories with residual-risk acceptance positions. OWASP LLM Top 10 coverage table with mitigations and ATLAS technique-level controls. Article 25(2) cooperation memo referenced. Refresh cadence quarterly with the last refresh dated 2026-03-15 and next refresh scheduled 2026-06-15. Version pinning to vendor-a-customer-service-ai/2026.03.0-rc4. Evaluation-data lineage section names the eval suite, version, and date cutoff. Training-data summary present with reference to Article 53(1)(d) Commission template. Known failure modes enumerated with 14 example prompts. Human-oversight design described in detail. Licence terms documented. Negative-use statement enumerates Article 5 prohibitions plus a vendor list of additional prohibited deployer use cases (clinical diagnosis, lethal autonomous weapons, etc.). Result: 14/15 with one minor CAR for quantitative environmental disclosure.

Vendor A - SOC 2 + AI Type II. 84 pages, audit period January-December 2025, BDO (national CPA firm), unqualified opinion. System description includes a 9-page AI-specific section covering: AI system inventory (the customer-service-AI product line and three underlying models including the upstream foundation model); AI data governance (training-data provenance, licence compliance, opt-out filtering); AI model lifecycle (development, evaluation, approval, deployment, monitoring, retirement controls); AI monitoring (drift, fairness, safety); AI incident response (Article 73 alignment for EU in-scope deployments); AI third-party management (the upstream foundation-model provider). Controls tested: 47 AI-specific controls in addition to 124 baseline TSC controls. Exceptions: 3 exceptions noted (2 in baseline access management with remediation by year-end, 1 in AI-model-deployment sign-off documentation for 2 of 18 deployments in February 2025 with remediation in March 2025). Result: clean unqualified Type II + AI; exceptions reviewed and accepted; ISO 42001 certificate and AI TSC report together cover the independent-assurance dimension. No procurement-blocking findings.

Vendor A - ML-BoM. 2,400-line CycloneDX 1.7 JSON. Includes: the deployed model (vendor-a-customer-service-ai 2026.03.0-rc4); the upstream foundation model (GPT-4o family with pinned version); the fine-tune dataset (1.4 million labelled customer-service-interaction examples, licensed); three evaluation datasets (one public, two proprietary); 47 software dependencies with hashes. CDXA attestations attached: 14 signed assertions including the upstream-model identity, the fine-tune dataset licence, the evaluation-suite version, the OWASP coverage attestation, and the ISO 42001 certification status. Signatures verified against the vendor's published Sigstore identity. Hash spot-check of the foundation model and fine-tune dataset both match. Result: clean ML-BoM with signed CDXA. No findings.

Vendor A, triangulation. Ten cross-artifact claims tested. Model version matches (model card, SOC 2 in-scope inventory, ML-BoM component version all read 2026.03.0-rc4). Training-data cutoff matches (model card declares 2025-11-30; SOC 2 AI-data-governance describes a 2025-11-30 cutoff; ML-BoM dataset component carries the same date). Fairness-slice methodology matches (model card cites the methodology paper; SOC 2 AI-monitoring section references the same methodology; ML-BoM has a pointer to the methodology paper as an external reference). OWASP coverage matches (model card lists LLM01-LLM10 mitigations; SOC 2 AI-incident-response references OWASP threat categories; ML-BoM CDXA attestation explicitly asserts OWASP LLM Top 10 v2025-12 coverage). Article 25(2) cooperation referenced in all three. Quarterly refresh cadence referenced in model card and SOC 2 AI-monitoring section. ISO 42001 certificate scope statement matches the SOC 2 in-scope system list matches the ML-BoM component list. Result: ten of ten triangulation claims consistent. Tier A confirmed by triangulation.

Vendor B, model card. 17 pages, Mitchell sections 1, 2, 3, 4, 5, 6, 9 present; sections 7 (quantitative analyses) and 8 (ethical considerations) absent. No intersectional fairness slices, only unitary gender slice. NIST AI 600-1 not referenced. OWASP coverage table partially populated (LLM01-LLM06 only). No Article 25(2) cooperation reference. Refresh policy declared as "annual" with last refresh 2025-09-30. Version pinning to "Vendor B Customer-Service AI v2.0" without a more specific build number. Evaluation-data lineage thin. Training-data summary one paragraph. Known failure modes: three example prompts. Human-oversight section absent. Licence terms present. Negative-use section minimal. Result: 7/15, material finding; below 10/15 procurement red flag threshold; vendor cannot be procured for high-risk use without substantial remediation.

Vendor B - SOC 2. 62 pages, audit period July 2025-June 2026, regional CPA firm, unqualified opinion. Baseline SOC 2 only, no AI TSC additions. System description is the standard SaaS description with no AI-specific section. Controls tested: 118 baseline TSC controls. No AI-specific controls. Exceptions: 5 exceptions (3 in baseline change management, 2 in access management; all remediated by year-end). The opinion is clean but the scope does not address the AI-specific dimensions the deployer needs. Result: baseline SOC 2 + ISO 42001 pending = Tier B independent-assurance posture at best, conditional on AI TSC addition within 12 months and ISO 42001 issuance Q4 2026; in the meantime the deployer must add compensating controls.

Vendor B - ML-BoM. JSON spreadsheet export, not CycloneDX. Lists the deployed model and the foundation-model reference (Mistral, version not pinned) and 12 software dependencies. No training-dataset entries. No CDXA signatures. No hashes. Result: ML-BoM not in CycloneDX 1.7; request export in CycloneDX 1.7 with CDXA signing within the standard CAR cycle. In 2026 procurement, the inability to produce CycloneDX 1.7 ML-BoM is itself a red flag. It signals the vendor's supply-chain maturity is behind the 2026 baseline.

Vendor B, triangulation. Ten cross-artifact claims tested. Model version partially matches (model card says "v2.0"; SOC 2 system description says "the Customer-Service AI product"; ML-BoM says "Vendor B Customer-Service AI v2.0", no precise build pin). Training-data cutoff: model card 2025-06-30; SOC 2 silent; ML-BoM silent, cannot triangulate. Fairness-slice methodology: model card unitary only; SOC 2 silent; ML-BoM silent, cannot triangulate. OWASP coverage: model card LLM01-LLM06 partial; SOC 2 silent; ML-BoM silent. Article 25(2): silent in all three. Result: two of ten triangulation claims consistent; eight cannot be triangulated. The artifact-cross-reading exercise confirms Tier D from the DDQ.

Finding memo. The Acme finding memo concludes: (1) Vendor A's three artifacts cross-confirm Tier A from the DDQ with one minor CAR (environmental disclosure) and one routine CAR (the Section 4 GPAI 30-day refresh tightening from lesson 070). (2) Vendor B's three artifacts confirm Tier D from the DDQ and add three artifact-level findings: model card material finding (7/15), absence of SOC 2 + AI scope, absence of CycloneDX 1.7 ML-BoM. (3) For Vendor B to upgrade to a procurement-viable position, the remediation programme is: model card to 12/15 within 90 days (add intersectional slices, NIST AI 600-1 statements, Article 25(2) reference, human-oversight section, OWASP completion); commit to AI TSC SOC 2 addition within 12 months; provide CycloneDX 1.7 ML-BoM with CDXA within 60 days. (4) The recommendation remains: award to Vendor A; keep Vendor B on the watch list for the next vendor swap if remediation completes. The memo is signed by the Chief AI Officer, General Counsel, and CISO and retained per Article 18 (10 years).

Cross-Walks, Penalty Exposure, and L3 Capstone

The vendor-evidence review cross-walks to the regulatory and standards stack as follows. EU AI Act: Article 11 (technical documentation), Article 13 (transparency and information to deployers), Article 16 (provider obligations including documentation), Article 25(2) (vendor-deployer cooperation), Article 26 (deployer obligations including 26(1) instructions for use, 26(4) operation monitoring), Article 50 (transparency obligations for certain AI systems), Article 53(1)(d) (GPAI training-content summary), Article 53(2) (GPAI Annex XI documentation), Article 71 (EU database registration), Article 72 (post-market monitoring). Annex IV §1 and §2 (technical documentation elements). Annex XII (downstream-deployer information). NIST AI RMF: Govern 6.1 third-party policy, Govern 6.2 third-party transparency, Map 4.1 contextual risk, Measure 2.1 third-party measurement. NIST AI 600-1: Generative AI Profile 12-risk categories with residual-risk acceptance. ISO/IEC 42001: Annex A.7 (data for AI), A.8 (information for users, model card territory), A.10 (third-party AI). SR 11-7 / OCC 2011-12 / PRA SS1/23: vendor model risk applicability: model card and SOC 2 + AI feed the inventory row, the validation report, and the third-party-risk register. AICPA: SOC 2 Trust Services Criteria (2017, 2022 update) plus the AI Trust Services Criteria (August 2024) plus the AICPA AI Risk-Management Framework alignment. OWASP and MITRE: OWASP LLM Top 10 (2025-2026), OWASP Agentic Top 10 (December 2025), MITRE ATLAS v5.4.0. CycloneDX: CycloneDX 1.7 specification (OWASP Foundation), CDXA Attestations specification. Mitchell-et-al.: "Model Cards for Model Reporting" (FAT* 2019).

Penalty exposure. Article 99(3), €15M or 3% of worldwide annual turnover, whichever is higher, applies to the deployer where Article 26 obligations are unmet. The deployer cannot defend Article 26(1) "use in accordance with the instructions for use" without reading the model card. The deployer cannot defend Article 26(4) "monitor the operation" without understanding the SOC 2 + AI control posture. The deployer cannot defend the supply-chain provenance claim without the ML-BoM. The artifact-review file is the audit trail. A deployer that procures the vendor on the DDQ alone, without artifact review, has not closed Article 26.

L3 capstone, how the L3 toolkit snaps together. Lesson 071 closes the L3 Risk Practitioner track. The toolkit the L3 practitioner now carries is coherent and operational. The FRIA (Article 27) identifies the fundamental-rights impacts of the system in the specific use context. The Annex IV technical documentation file (Article 11) carries the design, data, V&V, and oversight content. The ISO 42001 readiness work (the seven-element AIMS documentation) supplies the AI Management System chassis. The red-team library (OWASP / ATLAS / Agentic Top 10) supplies the adversarial-evaluation evidence. The MRM tiering (SR 11-7 / PRA SS1/23 / one-register-cross-walked) supplies the institutional model-risk discipline. The vendor DDQ (lesson 070) supplies the third-party diligence. The vendor evidence review (this lesson) supplies the artifact-cross-reading that turns DDQ answers into Article 26-defensible evidence. None of the seven artifacts is standalone. Each feeds the others. The model inventory row points to the FRIA, the Annex IV file, the red-team library, the tiering decision, the vendor DDQ, and the artifact-review memo. The governance committee (next lesson) ratifies the system through that chained evidence. The 2026 L3 practitioner is not a paper-pusher; the L3 practitioner is the person who can sit across the table from a regulator with the seven artifacts and walk through every claim with citations and answers.

Key Takeaways

  • Three core artifacts together attest to a 2026 AI vendor's AI Act and ISO 42001 readiness: the model card (Mitchell-et-al. nine sections plus Annex IV §2 mapping), the SOC 2 Type II + AI Trust Services Criteria report (AICPA August 2024 update plus AICPA AI Risk-Management Framework alignment), and the CycloneDX 1.7 ML-BoM with CDXA signatures. Each addresses a different audit question; the three together, read against each other, are the deployer's Article 26-defensibility evidence.
  • The 15-item model-card checklist covers Mitchell-et-al. nine sections; Annex IV §2(a)-(h) coverage; intersectional fairness slices (not just unitary); NIST AI 600-1 12-risk applicability statements with residual-risk acceptance; OWASP LLM Top 10 + ATLAS coverage; Article 25(2) cooperation; refresh cadence (quarterly minimum for production high-risk); version pinning; evaluation-data lineage; training-data summary; known failure modes; human-oversight design; environmental impact; IP and licence; negative-use statement.
  • SOC 2 + AI extends baseline SOC 2 with AI-specific system-description elements (AI inventory, AI data governance, AI model lifecycle, AI monitoring, AI incident response, AI third-party) and AI-specific controls inside the existing Trust Services Criteria, aligned to the AICPA AI Risk-Management Framework; relationship to ISO 42001 and Article 17 QMS is overlapping but non-substitutive.
  • CycloneDX 1.7 ML-BoM verification covers schema and version, ML model components (including upstream foundation model), dataset components, dependency components, ML-extension pointers (model card, training data, evaluation suite, vulnerability), CDXA attestations and JOSE/COSE signatures, and hash verification. Unsigned ML-BoM is a yellow flag; missing ML-BoM is a 2026 procurement red flag.
  • Triangulation tests ten cross-artifact claims (model version, training-data cutoff, fairness-slice methodology, OWASP/ATLAS coverage, vendor entity, certification scope, refresh cadence, Article 25(2) reference, human-oversight description, evaluation-suite identity) for consistency across model card, SOC 2 + AI, and ML-BoM. Inconsistency is the audit signal.
  • The Acme worked example confirms Vendor A at Tier A on artifact-cross-reading (14/15 model card, clean SOC 2 + AI Type II, signed CycloneDX 1.7 ML-BoM, ten-of-ten triangulation) and Vendor B at Tier D (7/15 model card, baseline SOC 2 only, no CycloneDX 1.7 ML-BoM, two-of-ten triangulation). The finding memo enumerates remediation conditions for Vendor B and signs off the Vendor A award with two CARs.
  • Article 99(3) deployer penalty exposure of €15M / 3% applies where Article 26 obligations are unmet; the artifact-review file is the audit trail. Without the cross-read, the DDQ alone does not close Article 26.
  • Lesson 071 closes L3: the L3 toolkit (FRIA + Annex IV file + ISO 42001 readiness + red-team library + MRM tiering + vendor DDQ + vendor evidence review) is coherent and chained, each artifact feeds the next, and equips the L3 practitioner to sit across the table from a regulator with seven artifacts and walk through every claim with citations and answers.