Quality Assurance for AI Work Products
Learning Objectives
After this 90-minute lecture, participants will distinguish verification from validation as they apply to AI work products, and will apply NIST AI RMF MEASURE and MANAGE functions to ongoing quality assurance. Learners will design pre-deployment evaluation plans including ground-truth datasets, subgroup performance, calibration, and adversarial robustness aligned with OMB M-24-10 minimum practices and GAO AI Accountability Framework performance standards. Participants will build citation verification procedures to counter generative AI hallucinations, grounded in agency publication standards and the Blueprint for an AI Bill of Rights Notice and Explanation principle. Learners will structure red teaming and independent review, integrating agency Inspector General, civil rights officer, and bargaining unit input. Participants will create incident response and remediation procedures tied to the GAO framework's monitoring component, and will draft operator checklists for human-on-the-loop review. Finally, learners will produce a quality assurance plan artifact covering pre-deployment, ongoing monitoring, and post-incident review for a specific rights-impacting AI in their agency.
Key Topics Covered
The lecture covers the distinction between verification and validation in AI context; pre-deployment evaluation design including ground-truth dataset construction, subgroup stratification aligned with Blueprint for an AI Bill of Rights equity principles, calibration assessment, uncertainty quantification, and adversarial robustness testing; citation verification procedures for generative AI including retrieval-augmented generation design, source logging, and reviewer protocols; hallucination categories and mitigation; red teaming practices including prompt injection, data exfiltration attempts, harmful output probes, and jurisdictional escalation; independent review structures involving Inspector General, GAO, agency civil rights officers, and bargaining units such as AFGE and NTEU; ongoing monitoring aligned with NIST AI RMF MEASURE and MANAGE; incident response and post-incident review aligned with the GAO AI Accountability Framework; case studies including IRS ID.me testing gaps, Michigan MIDAS false positive analysis, COMPAS fairness review, Houston HISD EVAAS transparency failure, and Allegheny County validation study as positive exemplar; and documentation artifacts including model cards, datasheets, test reports, and operator checklists.
Why This Matters for Government
Quality assurance for AI work products is the professional discipline that prevents public harms at scale. Government AI influences tax enforcement at IRS, benefits adjudication at SSA and VA, child welfare decisions at Allegheny and other counties, hiring at OPM and across federal contractors, and many more rights-impacting and safety-impacting decisions. When AI outputs are wrong, citizens are wronged. The Michigan MIDAS system produced tens of thousands of false fraud determinations because its quality assurance was inadequate. The IRS ID.me deployment lacked accessibility testing at scale. The COMPAS recidivism tool was scaled across jurisdictions without public-sector validation, and its disparate false positive rates by race were documented only years later by ProPublica. Houston Federation of Teachers v HISD struck down an opaque teacher evaluation model because its outputs could not be meaningfully challenged. Each failure traces back to a gap in quality assurance. Each is preventable.
Verification asks whether the system was built correctly. Validation asks whether the right system was built. Both are required. Verification includes unit tests, integration tests, and regression suites. Validation includes subject matter expert review, ground-truth accuracy testing on representative samples, calibration checks, subgroup fairness analysis aligned with NIST AI RMF MEASURE, and real-world piloting in shadow mode before production. The NIST AI RMF 1.0 MEASURE function provides the governance anchor. OMB M-24-10 requires pre-deployment testing and ongoing monitoring for rights-impacting and safety-impacting AI. The GAO AI Accountability Framework organizes performance into data, models, operations, and outcomes components. ISO 42001 provides management system clauses that tie these together.
Generative AI introduces hallucination, fabricated citations, overconfident wrong answers, and prompt-injection vulnerabilities. A federal attorney drafting with generative AI may receive citations to cases that do not exist, as happened in the 2023 Avianca case in federal court. A caseworker drafting benefits correspondence may receive plausible but incorrect regulatory references. A policy analyst may receive numeric figures that sound authoritative but come from no verified source. Quality assurance for generative AI therefore requires citation verification, retrieval-augmented generation against authoritative agency knowledge bases, source logging, structured reviewer protocols, and clear disclosure to end users that outputs were AI-assisted and human-reviewed. The Blueprint for an AI Bill of Rights principle on Notice and Explanation supports these requirements. OMB M-24-10 human-on-the-loop obligations reinforce them.
Red teaming is the structured adversarial probing of AI systems. For government AI, red teaming should cover prompt injection, jailbreaks, data exfiltration attempts, harmful output categories including defamation and discriminatory language, and scenario-based probing for rights-impacting decisions. The CISA Guidelines for Secure AI System Development and the NIST AI RMF Generative AI Profile provide reference materials. CISA and GSA have hosted federal AI red team exercises. Agencies should build red teaming into the development lifecycle rather than as an afterthought, with clear remediation pathways and follow-up testing.
Independent review provides accountability. The Inspector General, GAO, agency civil rights officers, Senior Agency Officials for Privacy, Chief Information Security Officers, and bargaining units all play roles. The Houston Federation of Teachers case illustrated that courts themselves can become independent reviewers of opaque AI systems, with due process as the accountability anchor. Agencies that invite independent review early reduce the risk of surprise findings later. The Allegheny County Department of Human Services offers a positive model: published validation studies by external researchers, community advisory input, and ongoing monitoring with public reporting.
Ongoing monitoring is not optional. Models drift. Data changes. Adversaries adapt. Agencies must monitor performance, fairness, incident rates, and user satisfaction continuously. When monitoring reveals degradation, agencies must have the authority and the posture to pause, investigate, and remediate. The pause-and-fix posture separates mature AI programs from vulnerable ones. Agencies that promise leadership they will never pause are setting up the next Michigan MIDAS.
Documentation artifacts anchor the practice. Model cards document intended use, limitations, training data, evaluation methodology, and known risks. Datasheets document dataset provenance, collection processes, and preprocessing. Test reports document pre-deployment evaluation results. Incident reports document failures and corrective actions. Operator checklists translate abstract requirements into concrete daily practice. Together these artifacts support FOIA, Privacy Act, Administrative Procedure Act, and congressional oversight responses.
The workshop portion asks each participant to draft a quality assurance plan for a rights-impacting AI in their own agency, covering pre-deployment evaluation, ongoing monitoring, red teaming cadence, independent review engagement, and incident response. Plans are peer reviewed. Participants leave with a template, peer feedback, and a concrete commitment to implement within 30 days.
Overview
The lecture frames quality assurance as the professional discipline that prevents public harm at scale and aligns with NIST AI RMF, OMB M-24-10, the GAO AI Accountability Framework, and ISO 42001.
Verification, Validation, and Evaluation Design
Participants distinguish verification from validation, design evaluation plans with ground truth, calibration, subgroup performance, uncertainty quantification, and adversarial robustness, and apply OMB M-24-10 minimum practices to rights-impacting AI.
Generative AI Quality and Citation Verification
The session addresses hallucination categories, RAG architectures with authoritative agency knowledge bases, citation verification protocols, reviewer checklists, and end-user disclosure practices.
Red Teaming, Independent Review, and Monitoring
Participants build red team programs referencing CISA Guidelines and NIST Generative AI Profile, integrate Inspector General and civil rights officer review, and establish continuous monitoring with pause-and-fix authority.
Documentation and Plan Drafting
Participants draft a quality assurance plan including model cards, datasheets, test reports, incident reports, and operator checklists, and peer review them.
Start Your CLUB Certification
This lecture is part of L2: AI Practitioner, 40 hours of practitioner-level training aligned with NIST AI RMF, OMB M-24-10, EO 14110, GAO AI Accountability Framework, and ISO 42001.
Related Lectures
L2 2.5.1 Bias Detection Tools and Methods, 90 minutes. L2 2.5.3 Human-in-the-Loop Design and Implementation, 90 minutes. L3 3.3.2 Red Teaming for Government AI, 120 minutes. L2 2.2.2 Evaluating AI Outputs, 90 minutes. L4 4.3.2 Independent Review and Audit, 120 minutes.
Skill.re