Stakeholder Analysis & Engagement
Why Stakeholder Work Is Load-Bearing for AI Programs
AI programs fail not because the models are wrong but because the humans around them were never properly included. The failure pattern repeats across sectors. In 2019 Apple's co-branded credit card with Goldman Sachs faced a New York Department of Financial Services investigation after David Heinemeier Hansson and Steve Wozniak reported that the algorithm offered female applicants significantly lower credit lines than male applicants with the same financial profile. The model was not the only problem: there was no visible stakeholder structure for affected customers, no clear appeal path, and no worker representation inside Goldman's underwriting team to flag the issue before launch. The NYDFS 2021 report did not find discriminatory intent, but it did find inadequate governance over the stakeholder lifecycle of the system. Apple and Goldman quickly added appeal processes, human-in-the-loop review for flagged applications, and new documentation practices. The cost was reputational and operational; the fix was organizational, not technical.
A second pattern is the Dutch toeslagenaffaire, the childcare-benefits scandal that ran from roughly 2013 through 2019. The Belastingdienst (tax authority) used a self-learning risk-selection system that disproportionately flagged parents with dual nationality for fraud investigation. More than 26,000 families were wrongly accused, thousands had children taken into state care, the Rutte cabinet resigned in January 2021, and the Dutch Data Protection Authority fined the tax authority 3.7 million euros in 2021 and 2.75 million euros in 2022. The core stakeholder failure: the families under surveillance had no voice in the design or oversight of the system, the auditors tasked with accountability had no mandate to examine algorithmic decisions, and internal whistleblowers who raised concerns were ignored. Stakeholder analysis is not a soft skill. It is the primary control that prevents systems of this kind.
This chapter treats stakeholder analysis and engagement as a disciplined engineering practice with named frameworks, evidence of effectiveness, and legal obligations. By the end you should be able to map any AI initiative's stakeholders using two complementary models, convert the map into a RACI chart, constitute an ethics board with real authority, satisfy union and works-council consultation rights under the EU AI Act and national labor laws, navigate FDA patient-advocacy engagement for medical AI, and avoid the specific legal pitfalls demonstrated by Mobley v. Workday, Clearview AI v. ACLU, California Proposition 22, and related cases. The chapter treats stakeholder work as part of the safety engineering stack, not separate from it.
A final framing: stakeholder engagement for AI differs from generic project management because AI systems cross traditional organizational boundaries. A hiring model touches legal, HR, affected applicants, works councils, and regulators. A clinical decision-support tool touches clinicians, patients, payers, the FDA, IRBs, and patient-advocacy groups. An enforcement model touches the enforced population, civil liberties organizations, legislators, and courts. Traditional project stakeholder maps miss at least half of these. That is why AI stakeholder practice starts with a deliberate expansion of the boundary of who counts.
Mapping Stakeholders: Power-Interest Grid and the Mitchell-Agle-Wood Salience Model
The oldest working tool in stakeholder analysis is the power-interest grid, commonly attributed to Aubrey Mendelow in 1991 and refined in subsequent project-management literature. The grid has two axes, power to affect the initiative on the vertical axis and interest in the initiative on the horizontal axis, dividing stakeholders into four quadrants: high power plus high interest (manage closely, typical of the CEO sponsor and the regulator), high power plus low interest (keep satisfied, typical of the CFO), low power plus high interest (keep informed, typical of frontline end users and user communities), and low power plus low interest (monitor, typical of adjacent teams). The grid forces a first-pass allocation of time and attention. For an AI product launch it is usually completed in a single 90-minute workshop with the steering committee and revisited quarterly.
The power-interest grid has a known weakness: it implicitly privileges stakeholders who already have institutional power. For AI systems that act on vulnerable populations, this is a dangerous bias. The Mitchell, Agle, and Wood (1997) salience model corrects for it. Mitchell-Agle-Wood classifies stakeholders by three attributes, power, legitimacy, and urgency, producing seven archetypes depending on which attributes they possess: dormant (power only), discretionary (legitimacy only), demanding (urgency only), dominant (power plus legitimacy), dangerous (power plus urgency), dependent (legitimacy plus urgency), and definitive (all three). The point of the model is that a dependent stakeholder, someone with a legitimate moral claim and urgent need but no power, like a benefits claimant incorrectly flagged by a fraud model, is invisible to the power-interest grid but definitionally high-salience under Mitchell-Agle-Wood. Recognizing dependent stakeholders is where stakeholder work for AI earns its keep.
In practice, teams use both frameworks in sequence. First, a power-interest grid captures the operational landscape: who must be kept satisfied, who must be informed, who must be actively managed. Second, the Mitchell-Agle-Wood overlay catches the dependent, discretionary, and dangerous stakeholders the first grid missed. For a large hospital system deploying a sepsis prediction model on top of Epic's Deterioration Index, the power-interest grid captures the Chief Medical Informatics Officer, the IT vendor, the billing office, and the nursing leadership; the Mitchell-Agle-Wood overlay adds patients (dependent, high salience), the hospital's diversity and community board (legitimate, often under-weighted), and agitated public critics like journalists covering algorithmic bias (dangerous when their concerns are ignored).
A third, less-known framework is worth naming: the Bryson (2004) stakeholder identification and analysis technique, which adds a process of iterative identification through snowball sampling and a sequence of concrete artifacts including a stakeholder problem frame and a power-versus-impact diagram. Bryson is useful when the stakeholder set is large, contested, and evolving, as it is for generative AI deployments in public sector agencies. The UK Government Digital Service and Canada's Treasury Board both recommend Bryson-style iteration for algorithmic impact assessments.
A common failure mode at this stage is to treat stakeholder mapping as a one-time artifact. Stakeholders change over the lifecycle of an AI system. An internal pilot has mostly internal stakeholders; a public launch adds civil-society groups and regulators; an incident transforms journalists from monitor quadrant to manage-closely. The engagement program should include a quarterly or per-release refresh of the stakeholder map, a trigger-based refresh after any incident, and a standing reviewer from the legal and communications functions who is empowered to add stakeholders outside the initial sponsor list.
RACI and Decision-Rights Frameworks for AI Programs
Once stakeholders are identified, responsibility must be allocated. The default tool is the RACI chart, where each row is a decision or deliverable and each column is a role, and each cell is marked R (responsible for doing the work), A (accountable, the single owner who signs off), C (consulted, two-way input required before the decision), or I (informed, one-way notification after). A well-made RACI has exactly one A per row, distributes R work plausibly, and limits C to those whose input is actually needed; a common symptom of dysfunctional governance is a RACI with four As per row or twenty Cs. Variants include RASCI (adding Support), RAPID from Bain (Recommend, Agree, Perform, Input, Decide) used by large companies like Ford and Unilever, and DACI (Driver, Approver, Contributors, Informed) from Intuit.
For AI governance specifically, the rows of the RACI need to be enumerated with care because AI decisions are often invisible to traditional controls. A mature AI RACI covers at least: use-case intake and triage, risk-tier classification, data acquisition and consent, model selection (build versus buy versus fine-tune), vendor due diligence, red-team engagement, bias and fairness evaluation sign-off, pre-launch governance review, go-no-go, incident declaration, user appeal handling, regulator notification, decommissioning, and annual governance review. Microsoft's Responsible AI Standard v2, Google's AI Principles review process, and the US Department of Defense's Responsible AI Strategy each publish variations of this row list.
A recurring tension is where accountability sits for AI outcomes. A strong pattern, used at firms like Salesforce and Mastercard, places final accountability for high-risk deployments with a Chief AI Ethics Officer or equivalent senior executive who does not report into the product line. The reasoning is the model-risk-management principle of SR 11-7 in US banking: second-line oversight cannot sit inside the first line. When Chief AI Ethics is collapsed back into product or engineering reporting lines, the accountability column of the RACI becomes performative. Regulators are beginning to notice; the NYC Local Law 144 of 2023 on automated employment decision tools and the EU AI Act's conformity-assessment requirements both implicitly assume an independent accountability structure.
Decision-rights get especially sharp when the system acts on people. Consider a RACI row for individual-level overrides: who can override a high-confidence AI decision against a customer or employee, who must be consulted, who is informed. In the Dutch toeslagenaffaire, the override right sat with frontline case workers who had been trained to defer to the algorithm; there was no second-line reviewer, no right of independent appeal, and no route to external ombudsmen that was functional at scale. The RACI that would have caught this would have placed the A for override on a named senior official, R on case workers, C on an independent appeals function, and I on the responsible minister at defined thresholds. Designing the RACI this way in advance is cheap insurance against the kind of organizational learned helplessness that the toeslagenaffaire exposed.
Ethics Boards, IRBs, and Advisory Structures That Actually Work
Ethics boards for AI are easy to constitute badly. Google's Advanced Technology External Advisory Council (ATEAC), announced in March 2019, was dissolved in about a week after member Kay Coles James's inclusion triggered employee protests and three other members left or refused to serve. Similar short-lived boards have been set up and disbanded at Axon, SenseTime, and others. A 2021 review in the journal Big Data & Society identified the common failure mode: advisory bodies without charter, without authority, without compensation, and without a compliance integration pathway collapse on first controversy.
What works looks more like an Institutional Review Board for human-subjects research. IRBs under the US Common Rule (45 CFR 46) have binding authority to approve, require modification, or disapprove research involving human subjects; they have quorum rules, conflict-of-interest recusals, minutes, and external members. The Partnership on AI's 2023 guide on AI ethics boards and the IEEE 7000-2021 standard on values-based system design both recommend similar structural features for AI ethics boards: a written charter approved at the board or C-suite level, authority that includes blocking deployments, at least one member independent of the sponsoring company, minutes kept under confidentiality but retrievable under subpoena, mandatory review thresholds tied to risk tiers, and a funded secretariat so the board is not just senior volunteers.
Companies that have built boards of this kind include Salesforce (Office of Ethical and Humane Use), Microsoft (Office of Responsible AI plus the Aether committee), Mastercard (AI Governance Council), and BBC (Machine Learning Engine Principles). Salesforce's ORH publicly blocked facial recognition features for law enforcement use; Microsoft's Aether committee shaped the 2020 decision not to sell Microsoft facial recognition to US law-enforcement agencies until federal regulation exists. These are examples of boards exercising real authority and absorbing reputational cost on behalf of the product line.
Related structures include Data Protection Impact Assessment (DPIA) committees under GDPR Article 35, Algorithmic Impact Assessment (AIA) panels as in Canada's Treasury Board Directive on Automated Decision-Making, and Fundamental Rights Impact Assessment (FRIA) committees required by EU AI Act Article 27 for certain high-risk deployers. These are not optional ethics boards; they are compliance structures, but they play an overlapping stakeholder function because the people sitting on them effectively represent affected parties who are not otherwise in the room.
For practitioners, the practical test of a working ethics structure is not whether it exists but whether it has ever blocked or materially reshaped a product. If the answer is no after two years of operation, the body is performative. Product leaders should view occasional blocks from the ethics function not as friction but as evidence of working governance, in the same way that security teams view occasional vetoes as evidence of a functioning SDLC.
Worker and Union Consultation: EU AI Act, Works Councils, and US Labor Law
Worker consultation is often ignored in US-led AI programs and legally mandated in much of the rest of the world. The EU AI Act, which entered force August 2024 with phased application dates through 2026, contains specific consultation obligations for high-risk workplace AI. Article 26(7) requires that deployers of high-risk systems used in employment inform workers and their representatives that they will be subject to the system, before putting it into service. That obligation sits on top of existing rights under the EU Directive 2002/14/EC on information and consultation of workers, and national works-council laws like Germany's Betriebsverfassungsgesetz (BetrVG). A landmark 2023 case at the Hamburg Labour Court (Arbeitsgericht Hamburg, ArbG 6 BVGa 5/23) held that a works council had to be involved in the introduction of ChatGPT for employees, and a 2024 Federal Labour Court ruling clarified that any AI tool capable of monitoring or assessing worker performance triggers co-determination rights under Section 87 BetrVG.
In France, the Comite social et economique (CSE) has consultation rights under the Labour Code that extend to the introduction of new technology with substantial impact on employment conditions; the Cour de Cassation 2022 decision affirming this in the Hewlett Packard case set the pattern for AI. In the Netherlands, the Ondernemingsraad has similar rights under the WOR. Sweden's MBL co-determination law requires consultation before major organizational changes. The practical implication for a multinational AI deployment: the works-council consultation schedule is on the critical path, not a footnote. Ignoring it creates both legal liability and durable loss of worker trust.
In the United States there is no equivalent of the works-council regime, but there are labor dynamics and emerging laws. The 2023 Writers Guild of America (WGA) and Screen Actors Guild (SAG-AFTRA) strikes explicitly negotiated AI clauses, establishing that synthetic performers require performer consent and that AI cannot be used to undercut writer credit. The New York City Local Law 144, effective July 5, 2023, requires bias audits of automated employment decision tools by an independent auditor and advance notice to candidates. California Proposition 22 (2020), which classified app-based drivers as independent contractors, carries a structural lesson for AI programs: when workers are formally outside the employment relationship, their stakeholder voice in algorithmic management must be constructed through other channels, a gap that ongoing California AB 2930 and similar bills are trying to close as of 2026.
Mobley v. Workday, filed in 2023 and certified as a collective action in May 2024 in the Northern District of California, is the leading US case on algorithmic hiring discrimination. Plaintiffs allege that Workday's AI screening tool produced disparate impact based on race, age, and disability. Whether or not the case ultimately succeeds on the merits, it has already reshaped vendor contracts: enterprise customers are now demanding audit rights, affected-class notice obligations, and data-sharing terms from vendors that did not exist in 2022. A stakeholder analysis for any hiring AI must now explicitly include job applicants as a stakeholder class with meaningful notice and appeal rights, not only because it is good practice but because depositions in Mobley-style litigation will focus on exactly this question.
Sectoral union contracts are another emerging source of AI stakeholder rights. The 2023 UAW contracts with Ford, GM, and Stellantis include language on advance notice and bargaining over new technology; the 2024 Teamsters UPS contract limits the use of in-cab cameras and algorithmic dispatch changes; public-sector contracts at AFSCME locals are incorporating AI-notice provisions. A stakeholder program that treats organized labor as only a historical artifact is blind to the fastest-growing channel for worker influence on AI in the US.
FDA, Patient Advocacy Groups, and Medical AI Stakeholder Engagement
Medical AI has the most mature stakeholder regime of any sector because it inherits decades of FDA patient engagement practice. The FDA's Patient Engagement Advisory Committee (PEAC), chartered in 2015 and first meeting in 2017, provides advice on patient perspective on medical device issues. Its 2020 meeting specifically addressed AI and machine-learning-enabled medical devices, and the outputs informed the 2021 FDA Action Plan for AI/ML-Based SaMD (Software as a Medical Device). Since then, the FDA's 2023 guidance on Predetermined Change Control Plans and the 2024 draft guidance on Marketing Submission Recommendations for PCCPs both require sponsor documentation of how patient input was sought.
For a device maker, credible patient advocacy engagement typically routes through established organizations. The FasterCures initiative at the Milken Institute has directories of more than 200 patient advocacy groups active in FDA engagement, including the American Diabetes Association (diabetes closed-loop insulin delivery algorithms), the Epilepsy Foundation (responsive neurostimulation devices), the Asthma and Allergy Foundation of America (digital therapeutics), the ALS Association (speech-generation AI), and many cancer-specific groups like Susan G. Komen and the Prostate Cancer Foundation. Each has its own stakeholder engagement norms: honoraria expectations, lay-language requirements for materials, and conflict-of-interest disclosure practices.
Patient-advocacy engagement is substantively different from user research. User research tells you how patients will use the system; advocacy engagement tells you whether patients accept the tradeoffs the system embodies, who gets to speak for patients who cannot speak for themselves, how to handle sensitive subpopulations like pediatric or cognitively-impaired patients, and how affected communities want risks communicated. The 2019 Optum-Stanford case, where a healthcare risk-prediction algorithm used in about 200 million Americans was shown by Obermeyer et al. in Science to under-refer Black patients because it used healthcare spending as a proxy for health, is the textbook example of what happens when patient-community stakeholder input is absent. Optum reworked the model; Stanford and Brigham and Women's Hospital revised their screening workflows; the incident reshaped how proxy labels are scrutinized in clinical ML, which the FDA's 2021 action plan specifically cites.
For non-medical domains the equivalent of patient advocacy is civil-society representation. For facial recognition, Clearview AI v. ACLU (2022, settled in Illinois under BIPA, with Clearview barred from selling its database to most private entities nationwide) was shaped by civil-liberties and immigrant-rights groups whose early stakeholder position shaped state-level Biometric Information Privacy Acts in Illinois, Texas, and Washington. For criminal-justice AI, the MacArthur Justice Center, the Brennan Center, and Upturn play the patient-advocacy-equivalent role; the ACLU's work on Detroit's facial-recognition arrests of Robert Williams (2020) and Porcha Woodruff (2023) exemplifies it.
A robust stakeholder program in medical AI therefore has standing relationships with at least three patient-advocacy organizations relevant to the indication, documented in both directions as a mutual engagement agreement, with protocols for pre-market review, post-market surveillance input, and rapid communication in the event of incidents. This is materially more work than an ad hoc advisory call before launch, and it is materially more protective.
The Engagement Playbook: Tactics, Cadence, and Measurement
Having mapped stakeholders and allocated decision rights, the engagement work itself has to be structured. The pattern that has emerged across public-sector and enterprise AI programs is a tiered engagement model. Tier 1 engagement, for low-risk internal tools, is typically email notice, a FAQ, and a feedback inbox. Tier 2 engagement, for productivity and customer-service tools, adds focus groups, quarterly town halls, and an open office-hour program. Tier 3 engagement, for high-risk systems under EU AI Act Annex III, NYC LL144, FDA SaMD review, or equivalent, requires formal consultation protocols: pre-implementation consultation with workers and/or affected communities, documented co-design sessions with representative users, published fairness and performance artifacts, an independent audit, and an incident-response commitment with named external liaisons.
Specific tactics that work well include the Citizens Assembly model used by Ireland for the abortion and marriage referenda and adapted by the Ada Lovelace Institute for AI deliberation; deliberative polling pioneered by James Fishkin at Stanford; and the consensus-conference model used by the Danish Board of Technology. These are deliberative formats in which randomly-selected citizens engage with technical experts to produce informed recommendations. The Ada Lovelace Institute's 2021 Biometrics Citizens Jury, and the Conseil d'Etat's deliberative processes on public-sector AI in France, are well-documented examples. For enterprise deployments the equivalent is co-design sprints with union representatives and affected employee groups, typically lasting one to two weeks, producing written agreed-use and agreed-limits documents that then bind the deployment.
Engagement cadence depends on risk tier and lifecycle stage. Pre-deployment should include at least one round of scoping, one of design, and one of pre-launch review. In-flight operation should include quarterly metric reviews with stakeholder representatives, annual full reviews of the stakeholder map, and unscheduled reviews after any incident of severity High or above on the organization's incident scale. Decommissioning should include an exit consultation that examines both the decision to decommission and the plan for affected populations who had begun to rely on the system.
Measurement of engagement quality is the final piece. Poor-quality engagement produces theater; good-quality engagement produces decisions. Practical metrics include: count of decisions materially changed as a result of stakeholder input (target at least two per release for high-risk systems), diversity of demographic and professional representation on advisory bodies against a documented composition standard, response time on appeals and grievances (NYC Local Law 144 implicitly requires a working appeals path), retention of external stakeholders across years (high churn signals unhappy stakeholders), and audit findings related to stakeholder-facing controls in ISO/IEC 42001 Annex A.6.3 on stakeholder communication. The Partnership on AI and the Ada Lovelace Institute have both published measurement guides.
A closing discipline: every stakeholder commitment becomes a tracked action item in the same system that tracks engineering work. If you tell a works council you will publish a fairness report by Q3, that item exists in Jira or equivalent with a named owner, a tracked deadline, and governance escalation if it slips. The most common failure is not hostility to engagement; it is that engagement commitments exist in one system and engineering work exists in another, and the two diverge. Closing that loop is what separates performative engagement from the real thing.
Case Studies, Failure Modes, and a Decision Framework
Several named cases repay close study. The Dutch toeslagenaffaire, already introduced, is the textbook failure of dependent-stakeholder inclusion and of the absence of an independent appeals path. The affected families were invisible in the internal RACI, the external appeal routes were blocked by procedural gates, and whistleblowers inside the Belastingdienst were ignored. The 2020 Parliamentary inquiry Ongekend onrecht (Unprecedented Injustice) is required reading; its recommendations became the template for the EU AI Act's Article 27 FRIA.
Mobley v. Workday is the active US stakeholder battleground. The litigation is reshaping vendor contracts across enterprise hiring. Practical implication: any AI system that makes or substantially contributes to decisions about people must now include in its stakeholder map the population of affected decision-subjects with a real notice and appeal path, and the vendor-customer contract must allocate responsibility for that path.
Clearview AI v. ACLU (2022 Illinois settlement) and the parallel Van Buren and Mutnick cases established that biometric privacy plaintiffs are definitive-salience stakeholders under Mitchell-Agle-Wood and that regulatory structure (Illinois BIPA) is the mechanism through which that salience is enforced. For any face, voice, or gait-based AI product, civil-liberties organizations and state Attorneys General are load-bearing stakeholders.
Proposition 22 (California 2020) is a stakeholder-structure case: when workers are reclassified outside employment status, their stakeholder voice in algorithmic management changes channel. The implication for AI programs: stakeholder analysis must follow the substance of the work relationship, not the legal label. If drivers, riders, delivery workers, or task workers materially depend on the AI for income, they are high-salience stakeholders regardless of classification.
The Google ATEAC dissolution (2019) is the cautionary tale for advisory bodies. An advisory body without charter, authority, compensation, and a compliance integration pathway will be consumed by the first controversy rather than serving its function. Salesforce's Office of Ethical and Humane Use and Microsoft's Aether show what good looks like.
How do you actually decide, in practice, what stakeholder structure a new AI program needs? A pragmatic decision framework proceeds through six questions. First, classify the risk tier: is this a high-risk system under the EU AI Act Annex III or any comparable regime, or is it low-risk? Second, map power-interest and overlay Mitchell-Agle-Wood to catch dependent stakeholders. Third, check statutory consultation obligations: works councils in Europe, NYC LL144 in New York, FDA patient engagement in medical, BIPA and CCPA in biometrics, NLRA considerations in US labor. Fourth, design the RACI explicitly, with single accountability per row and realistic C allocation. Fifth, stand up the governance body with written charter and real authority, or leverage an existing IRB, DPIA, or AIA committee. Sixth, define the engagement cadence and measurement plan, and integrate every commitment into the engineering work-tracking system. Work the framework once, and the second AI program takes a quarter the time; work it never, and your organization eventually features in the next set of case studies.
Skill.re