Building AI Centers of Excellence
Learning Objectives
By the end of this lecture you will (1) articulate why federal agencies establish AI Centers of Excellence (CoEs) rather than distributed AI teams, including economies of scope, scarce-talent concentration, governance harmonization, and reduction of duplicative pilots; (2) describe the original GSA Centers of Excellence program launched in 2017 under the Office of American Innovation and operationalized at USDA, HUD, and other agencies, and the distinct AI CoE pattern that emerged under OMB M-24-10's requirement for Chief AI Officers; (3) design a CoE charter that specifies mission, authorities, scope, services, funding, staffing, and performance measures; (4) choose among three operating models (central of excellence, federated, hub-and-spoke) with explicit tradeoffs; (5) define a use-case intake process aligned with the annual AI Use Case Inventory and with NIST AI RMF Map function; (6) stand up governance including an AI Governance Board, model risk officer, and privacy/civil rights/security partnerships with the agency's existing officials; (7) plan workforce across Occupational Series 2210, 1515, 1530, 0110 and pathways through OPM direct-hire authority and the Intergovernmental Personnel Act; and (8) measure CoE performance with a balanced scorecard covering delivery, talent, governance, and impact, avoiding the 'activity theater' that plagues many government innovation offices.
Historical Context: CoEs in Federal Government
The federal Centers of Excellence model predates AI. The Office of American Innovation under the Trump Administration launched the GSA Centers of Excellence program in 2017 across six focus areas: Cloud Adoption, Infrastructure Optimization, Customer Experience, Contact Center, Data and Analytics, and Artificial Intelligence. GSA partnered with the Department of Agriculture as pilot agency, later expanding to HUD, OPM, Department of Labor, USDA FPAC, Department of Energy, Environmental Protection Agency, Department of Homeland Security FEMA, Joint Chiefs of Staff, and others. The model was GSA-led federal employees augmented by contractor teams, funded by the receiving agency through interagency agreements. Deliverables included strategy, implementation acceleration, and capability build. The CoE model was influenced by the Digital Service model pioneered by the U.S. Digital Service (stood up 2014) and 18F (stood up 2014 within GSA), which themselves borrowed from the U.K. Government Digital Service (2011). The AI-specific flavor emerged in 2019 under the Federal Data Strategy and accelerated under OMB M-24-10's Chief AI Officer requirement. Agencies now look to the CoE pattern to concentrate AI expertise, share infrastructure, avoid duplicate pilots, and harmonize risk management. The Veterans Affairs National AI Institute, the DoD Chief Digital and AI Office, the Department of Energy Artificial Intelligence and Technology Office, and the Department of Health and Human Services AI Council each represent variations on the theme.
Why Establish a Center of Excellence
There are five systematic reasons to concentrate AI capability in a Center of Excellence rather than distribute it across component offices. First, AI talent is scarce and expensive; a central pool can be deployed across use cases while individual offices cannot justify their own. Second, AI infrastructure (training clusters, MLOps platforms, data lakes with governance) has high fixed cost and low marginal cost; sharing amortizes it. Third, governance artifacts (risk management policy, evaluation standards, monitoring platforms, model cards, impact assessments) are reused across use cases and should be produced once. Fourth, the risk of uncoordinated pilots is high: duplicated vendors, inconsistent security postures, divergent privacy reviews, conflicting public messages, and policy arbitrage between components. Fifth, Congressional and OMB expectations are increasingly articulated at the agency level; a component-only approach produces lawyer-intensive coordination overhead at each touchpoint. The GAO June 2021 AI Accountability Framework highlighted these benefits and noted that agencies with central AI capability mature governance faster. Risks of CoE centralization include bottlenecks at the center, distance from operational business problems, and a tendency for CoEs to focus on technology for its own sake rather than mission impact. Mitigations include embedded 'site reliability' analogs that deploy into components, product management discipline, and explicit success criteria tied to mission outcomes.
Three Operating Models
Central Model: The CoE executes AI work directly on behalf of components. All AI people report to the CoE leader, all AI infrastructure is owned centrally, and components request services. Advantages: full consolidation, maximum governance coherence, simplest accounting. Disadvantages: distance from domain context, queuing delays, loss of component ownership. Federated Model: Each component runs its own AI team with light coordination from a small central office. Advantages: domain proximity, component accountability. Disadvantages: duplicated effort, inconsistent governance, difficulty attracting top talent to any one component. Hub-and-Spoke Model: The CoE provides shared infrastructure, governance frameworks, and talent rotations; components run domain teams that plug in. Advantages: balances consolidation with domain depth, supports both cross-cutting capability and local ownership. Disadvantages: requires sophisticated operating model design, and can be confusing if authorities are unclear. Most mature federal AI programs converge on hub-and-spoke. The DoD Chief Digital and AI Office (CDAO) with component AI offices, the VA National AI Institute with VHA/VBA embeds, and HHS AI Council with CMS/FDA/CDC/NIH domain teams are all variants of hub-and-spoke. Choose based on agency size, mission diversity, and maturity; expect to evolve the model over 3-5 years.
Charter, Authorities, and Funding
A CoE charter should specify (a) mission and theory of change; (b) scope (AI types, mission areas, inclusion/exclusion of national-security systems); (c) authorities (delegated by the Deputy Secretary or CAIO, including decision rights over intake, governance approvals, and capability investment); (d) services (advisory, delivery, governance, infrastructure, workforce development); (e) funding model; (f) staffing model; (g) measurement. Funding in federal government is typically a combination of direct appropriation (where Congress creates a line item), working capital fund assessment on components that use CoE services, and project-specific interagency agreements for cross-agency work. The Technology Modernization Fund (TMF) has funded several AI CoE related initiatives. Authorities must be clear enough to avoid 'governance by permission' where the CoE has visibility but no ability to act. The Chief AI Officer designated under OMB M-24-10 may be the CoE leader or may be a peer with oversight of the CoE; both patterns exist. Where the CoE does not include national-security systems, coordination with the SAOIM, CIO, CISO, CDO, and Privacy Officer must be defined. Civil rights partnership with the agency's Office of Civil Rights or OGC is mandatory for rights-impacting AI.
Use-Case Intake and Pipeline Management
A CoE without a disciplined intake process quickly drowns in low-value requests. The standard intake pipeline has five stages: Discover, Evaluate, Plan, Deliver, Operate. Discover: light-touch submission with problem statement, owner, expected beneficiaries, timing, and rough scale. Evaluate: apply screening criteria including mission impact, feasibility, fit with strategy, risk category under OMB M-24-10 (rights-impacting, safety-impacting, neither), data readiness, and availability of CoE capacity. Plan: develop a pilot or project charter, perform NIST AI RMF MAP function, complete privacy and civil rights reviews, draft acceptance criteria. Deliver: execute with iterative demos, evaluation against criteria, and ongoing MEASURE activity. Operate: transition to sustained operations with monitoring, incident response, and revalidation triggers. The annual AI Use Case Inventory required by OMB M-24-10 captures entries at the Operate stage but should reflect pipeline visibility. Mature CoEs publish kill criteria: use cases that fail evaluation or lose mission relevance are closed rather than starved. VA, HHS, and several DoD components have published portions of their intake frameworks. Prioritization should be transparent; lotteries and favor-based decisions erode CoE credibility.
Aligning with Federal AI Governance
A CoE is one node in a broader agency AI governance ecosystem. OMB M-24-10 requires each CFO Act agency to designate a Chief AI Officer, convene an agency-level AI Governance Board, publish an annual AI Use Case Inventory, and implement minimum practices for rights-impacting and safety-impacting AI. NIST AI RMF 1.0 defines Govern, Map, Measure, Manage functions. EO 14110 broadened directives to include red-teaming, workforce, and equity. NSM-10 addresses national security AI. Agencies interact with OMB, GAO, the National AI Advisory Committee (NAIAC), the U.S. AI Safety Institute, CISA, and OSTP. The CoE should map each charter activity to the corresponding RMF function and M-24-10 practice. Redundancy between the CoE and other governance bodies should be eliminated: for example, if the agency CIO Council already reviews IT investments, the CoE should not create a parallel investment review; it should embed AI-specific criteria into the existing mechanism. Similarly, privacy review should run through the Senior Agency Official for Privacy, security through the CISO, and civil rights through OCR/OGC, with the CoE providing AI-specific analysis inputs.
Workforce Strategy
AI CoE workforce requires at minimum (a) technical leadership: Chief AI Officer, CoE Director, Lead Data Scientist, MLOps Lead; (b) applied staff: data scientists, ML engineers, data engineers, statisticians, product managers, UX researchers, human-factors specialists; (c) governance staff: model risk officer, privacy liaison, civil rights liaison, evaluation lead, red-team lead; (d) operations: FinOps for AI, program management, communications. Federal hiring tools include OPM direct-hire authority for data scientist (Series 1560 or 2210 depending on role), Intergovernmental Personnel Act assignments from universities, Expert/Consultant appointments, schedule A hiring flexibilities where appropriate, and the CDO/CAIO Council's workforce initiatives. Salary caps create competitiveness issues; Senior Executive Service limits, special pay rates for IT, and retention bonuses should be used. Contract workforce supplements federal staff but cannot substitute for federal decision-makers. Culture is critical: traditional IT org design rewards predictability; AI teams require tolerance for failed experiments, publication of negative results, and peer review. USDS, 18F, and DIU provide cultural reference points. Retention risk is real; mid-career data scientists have outside options. Visible impact, technical challenge, and a strong learning culture are the empirically proven retention levers.
Infrastructure and MLOps Platforms
A credible CoE provides shared infrastructure that components cannot justify individually. At minimum this includes (a) a FedRAMP-authorized cloud environment appropriate to the data sensitivity (Moderate, High, or IL5/IL6 for DoD), (b) ML training and serving platforms (SageMaker, Vertex AI, Azure ML, or open-source equivalents), (c) data lake and feature store with lineage, (d) model registry and versioning, (e) experiment tracking, (f) evaluation harnesses including bias/fairness tooling, (g) monitoring and drift detection, (h) incident management integration. The 2025 FedRAMP modernization has accelerated AI service authorizations but gaps remain; agencies commonly maintain provisional ATOs for cutting-edge services. For classified environments, the cross-domain requirements and NSA guidelines add complexity. Platform choices should be explicit about avoiding vendor lock-in: data formats should be portable, models should be exportable, and the CoE should not become contractually dependent on a single provider for its core capability. Total cost of ownership planning should include GPU/TPU compute, storage, human time for platform maintenance, and the often-underestimated cost of governance tooling.
Measuring CoE Performance
Measurement is where most CoEs fail. A balanced scorecard should cover four dimensions. Delivery: number of production-deployed use cases, cycle time from intake to deployment, share of deployments with documented outcome measurement. Talent: CoE headcount against plan, retention rate, time-to-fill for key roles, share of staff engaged in cross-training. Governance: percentage of use cases with completed AI Impact Assessment, percentage with fairness audit, percentage with red-team findings and remediation, cadence of AI Governance Board meetings, inventory completeness. Impact: mission outcomes attributable to AI deployments (processing time, error rate, equitable service, cost savings), public trust metrics where applicable, GAO and IG finding closure rate. 'Vanity metrics' to avoid include slide deck count, conference appearances, and number of pilots that never reached production. GAO's AI Accountability Framework and OMB M-24-10 provide guidance on performance measurement. Publication of performance data (with appropriate protection of sensitive information) is both a transparency obligation and a retention tool, as high performers prefer visible work. Reviewers should distinguish 'activity theater' (lots of pilots, few production wins) from real CoE maturity.
CoE Anti-Patterns
(1) The Consulting CoE: pure advisory with no delivery authority; produces reports that components ignore. (2) The Outsourced CoE: entire function run by a contractor with no federal knowledge transfer; collapses when contract ends. (3) The Pilot Factory: endless pilots, no production deployments, confused metrics. (4) The Center of Prestige: high-profile partnerships with brand-name universities and foundation-model vendors but no domain depth. (5) The Walled Garden: no embedded component staff, no rotations, and no intake transparency. (6) The Technology-First CoE: adopts tools before understanding mission needs, leaving components holding tools they cannot use. (7) The Shadow IT CoE: operates outside CIO/CISO/CPO authority, eventually triggers FISMA or Privacy Act findings. (8) The Bias-Blind CoE: no civil rights partner, no fairness discipline, and no pipeline to OCR/OGC review. The Dutch toeslagenaffaire is the cautionary tale here. (9) The Revolving-Door CoE: high leadership turnover, no institutional memory. (10) The Measurement-Free CoE: no balanced scorecard, no outcome tracking, no defensible narrative for OMB or GAO. Each anti-pattern has multiple real-world examples in federal and state government. Mature CoEs explicitly audit themselves against this list quarterly.
Case Studies
VA National AI Institute (NAII): Stood up in 2019, consolidates AI research and application for the Veterans Health Administration, with partnerships at DeepMind (predecessor), Harvard, and academic medical centers. Focuses on suicide prevention, cardiology, and radiology AI; publishes peer-reviewed work. DoD Chief Digital and AI Office (CDAO): Created February 2022 from merger of JAIC, DDS, and Office of the Chief Data Officer, consolidated under an Acting CDAO and then a permanent CDAO. Runs Project Maven (intelligence), Advana (data platform), Task Force Lima (generative AI), and AI/ML scaffolding services. Department of Energy AI and Technology Office: AI for science applications, including the Exascale Computing Project and Argonne/Oak Ridge/Lawrence Livermore AI supercomputing. GSA Centers of Excellence AI track: Delivered advisory services to USDA, HUD, OPM, DoL, DoE, EPA, DHS FEMA, and others. HHS AI Council: Coordinates CMS, FDA, CDC, NIH, IHS, and ONC AI efforts; chaired by the ASA or HHS CAIO under M-24-10. The UK Government has the AI Incubator (i.AI) within the Cabinet Office and the Central Digital and Data Office. Singapore's AI CoE is within GovTech. Each offers lessons; the most consistent pattern is that hub-and-spoke with strong central infrastructure and strong component domain teams outperforms either pure central or pure federated.
Summary and Next Steps
An AI Center of Excellence concentrates scarce talent, shared infrastructure, governance harmonization, and quality-assurance capability in one place, while preserving domain depth through federation or hub-and-spoke design. A credible charter specifies mission, authorities, scope, services, funding, staffing, and measurement. Intake pipeline, governance alignment with OMB M-24-10 and NIST AI RMF, workforce strategy leveraging OPM flexibilities, shared MLOps infrastructure, and a balanced scorecard tracking delivery, talent, governance, and impact are the core operational disciplines. CoE anti-patterns (consulting-only, pilot factory, shadow IT, bias-blind, measurement-free) are well documented and should be audited against quarterly. Federal precedents from VA, DoD, DoE, HHS, and GSA provide battle-tested models. The next lecture, AI Talent Development and Retention, expands on the workforce pillar in depth, covering hiring authorities, training pipelines, career paths, and retention mechanics for scarce AI talent in government.
Skill.re