Evaluating AI Vendor Claims
Learning Objectives
After this 120-minute lecture and workshop, participants will scrutinize AI vendor claims using a structured framework grounded in OMB M-24-10 minimum practices, the NIST AI RMF, the GAO AI Accountability Framework, and the Federal Acquisition Regulation. Learners will verify FedRAMP authorization level and scope, inspect NIST AI RMF alignment, and demand independent evaluation evidence covering accuracy, subgroup performance, calibration, and robustness. Participants will analyze FTC enforcement precedents including actions against Rite Aid on facial recognition, DoNotPay deceptive AI lawyer claims, Automators AI earnings claims, and WeightWatchers Kurbo data deletion; DOJ Civil Rights actions on AI discrimination; and state attorney general enforcement including Texas Meta settlement and New York AG actions. Learners will test vendor claims against real cases such as IRS ID.me, Michigan MIDAS, COMPAS, Houston HISD, and the Dutch childcare scandal. Participants will exercise FAR-aligned evaluation criteria, request model cards, datasheets, red team reports, and incident histories, and will draft a vendor due diligence template. Finally, learners will rehearse vendor interviews with structured questions that separate evidence from marketing.
Key Topics Covered
The lecture covers the landscape of AI vendor marketing and the patterns of overclaiming that frequently appear in federal RFI responses; FedRAMP authorization verification including level, scope, continuous monitoring status, and inherited controls; NIST AI RMF alignment evidence including GOVERN, MAP, MEASURE, and MANAGE artifacts; independent evaluation evidence including representative test sets, subgroup stratification, calibration, uncertainty, and adversarial robustness; FTC enforcement actions illustrating the legal risk of unsupported AI claims such as Rite Aid, DoNotPay, WeightWatchers Kurbo, Automators AI, and Weight Loss Gummies; DOJ Civil Rights Division enforcement on AI discrimination; state attorney general actions; FAR evaluation criteria, best-value analysis, and GSA Schedules considerations; contract terms including SLAs, audit rights, exit rights, no-training clauses, indemnification, and liability caps; case studies of IRS ID.me, Michigan MIDAS, COMPAS, Houston HISD, Dutch childcare scandal, and Allegheny County; and a practical due diligence template and interview script for AI vendor selection.
Why This Matters for Government
Every AI procurement decision is a policy decision. The vendor a federal, state, or local agency selects shapes how citizens experience government services, how workforce roles evolve, and how risk is allocated. Vendor marketing is rarely a reliable source of truth. FTC, DOJ, and state attorney general enforcement actions demonstrate that unsupported AI claims produce real legal and public-interest harm. The FTC action against Rite Aid in 2023 illustrated the consequences of deploying facial recognition that misidentified customers, particularly people of color, and led to a five-year ban on facial recognition use. The FTC action against DoNotPay alleged deceptive claims that the tool was the world's first robot lawyer, resulting in a settlement and consumer refunds. The FTC action against Weight Loss Gummies targeted AI-generated fake endorsements. The FTC's Operation AI Comply swept multiple deceptive AI marketers. These enforcement actions establish that AI vendor claims are subject to deceptive practices law under Section 5 of the FTC Act and comparable state laws.
For government procurement, OMB M-24-10 minimum practices for rights-impacting and safety-impacting AI raise the bar. Agencies must verify pre-deployment testing, ongoing monitoring, human oversight, impact assessment, and public notice obligations are satisfied. The NIST AI RMF 1.0 provides the governance and technical scaffolding. The GAO AI Accountability Framework provides audit-ready components: Governance, Data, Performance, and Monitoring. The Federal Acquisition Regulation provides the procurement structure. Together these frameworks form the evaluation lens for any vendor claim.
FedRAMP verification is often the first gate. Many vendors claim to be FedRAMP-ready, FedRAMP-authorized, or FedRAMP-equivalent. These are not the same. FedRAMP Ready means an independent 3PAO confirmed the vendor met some pre-authorization criteria. FedRAMP-Authorized means the vendor has a current Authority to Operate issued by an agency or the Joint Authorization Board. FedRAMP-Equivalent is a limited concept used in DoD contexts with specific meaning. Agencies should verify current status on the FedRAMP Marketplace, check the authorization boundary, confirm the impact level Low, Moderate, or High matches the data classification, and review continuous monitoring reports. Inherited controls must be documented.
NIST AI RMF alignment evidence should include a governance policy mapping to GOVERN, context and risk framing mapping to MAP, evaluation evidence mapping to MEASURE, and ongoing monitoring and incident response mapping to MANAGE. Many vendors produce glossy decks that name-drop NIST AI RMF without documented artifacts. The structured due diligence template asks for each artifact by name.
Independent evaluation evidence is often where vendor claims fail. A vendor that cites a single aggregate accuracy number without representative test sets, subgroup stratification, calibration, uncertainty, and adversarial robustness is not providing evidence. Agencies should request third-party evaluations from academic partners, NIST benchmarks where available such as FRVT for face recognition, and standardized safety benchmarks for generative AI. When vendors cannot produce such evidence, the agency should require it as a contract deliverable or decline to proceed.
Civil rights and equity obligations apply to vendor AI. Title VI of the Civil Rights Act, Title VII in employment contexts, the ADA, EEOC guidance on AI hiring tools, and the OFCCP affirmative action obligations for federal contractors apply. The DOJ Civil Rights Division has publicly announced investigations into AI discrimination under ADA and Title VI. The Blueprint for an AI Bill of Rights Algorithmic Discrimination Protections principle reinforces these obligations. Vendors should produce subgroup performance evidence and commit to ongoing disparate impact monitoring.
Contract terms matter enormously. Audit rights enable agencies to verify performance and compliance. Exit rights enable agencies to walk away when performance or ethics issues emerge. No-training clauses prevent agency data from entering vendor training pipelines without authorization. Indemnification allocates risk between agency and vendor. Liability caps should be negotiated based on rights-impact. SLAs with remedies align vendor incentives with performance. The FAR and agency procurement policies shape the available contract vehicles; GSA Schedules including IT Category 70 and Alliant 3 are common.
Case studies illustrate the stakes. IRS ID.me vendor claims of accuracy and accessibility did not withstand scrutiny at scale, leading to the 2022 Treasury reversal. Michigan MIDAS vendor performance claims in fraud detection did not withstand operational reality. COMPAS vendor claims of accuracy masked racial disparities later documented by ProPublica. Houston HISD EVAAS vendor claims about teacher evaluation did not withstand due process scrutiny. The Dutch childcare scandal exposed how mass data analytics vendor claims can produce cascading harm. Allegheny County stands as a positive counterpoint: external validation before scaling, community advisory input, and transparent monitoring.
Vendor interviews should be structured. Prepare questions that require evidence rather than adjectives. Ask for documentation, test reports, and references. Ask how the vendor handles incident disclosure, rollback, and remediation. Ask who will own the data and for how long. Ask whether the vendor's sub-processors are FedRAMP-authorized. Ask about known failures and how they were addressed. A vendor that can answer these questions with documented evidence is a credible partner. A vendor that cannot is a risk.
The workshop portion of this lecture asks each participant to construct a vendor due diligence template and interview script for a current or anticipated procurement in their own agency. Peer review follows. Participants commit to using the template in their next procurement conversation. Skepticism is professionalism. Evaluation is the procurement officer's craft.
Overview
The lecture frames vendor claim evaluation as a core procurement discipline tied to OMB M-24-10, NIST AI RMF, the GAO AI Accountability Framework, and the FAR.
FedRAMP and NIST AI RMF Alignment
Participants verify FedRAMP authorization level, scope, continuous monitoring status, and inherited controls. NIST AI RMF GOVERN, MAP, MEASURE, and MANAGE artifacts are reviewed by name.
Evaluation Evidence and Subgroup Performance
Participants require representative test sets, subgroup stratification, calibration, uncertainty, and adversarial robustness. Third-party benchmarks including NIST FRVT and standardized safety benchmarks are considered.
Enforcement Landscape and Civil Rights
Participants review FTC Operation AI Comply, Rite Aid, DoNotPay, WeightWatchers Kurbo, DOJ Civil Rights Division investigations, EEOC guidance, OFCCP obligations, and state attorney general enforcement.
Contract Terms, Case Studies, and Interview Script
Participants design SLAs, audit rights, exit rights, no-training clauses, indemnification, and liability caps and construct a structured vendor interview script drawing on IRS ID.me, Michigan MIDAS, COMPAS, Houston HISD, and Allegheny County case studies.
Start Your CLUB Certification
This lecture is part of L3: AI Strategist, 80 hours of advanced training aligned with NIST AI RMF, OMB M-24-10, EO 14110, GAO AI Accountability Framework, and the FAR.
Related Lectures
L3 3.3.1 Writing AI Requirements in RFPs and SOWs, 120 minutes. L3 3.3.3 AI Contract Negotiation, 120 minutes. L3 3.3.4 Ongoing Vendor Management, 90 minutes. L2 2.3.2 Data Protection for Agency AI, 90 minutes. L4 4.2.2 Legal and Compliance in AI Procurement, 120 minutes.
Skill.re