Policy Development & Documentation
Overview
Keona Kahale is General Counsel at a diversified financial services group based in Auckland. Two years ago, her company's data science team built an AI system to prioritize customer service callbacks - not a credit decision, not an underwriting model, just a queue management tool. Six months after launch, a journalist discovered that the system consistently deprioritized callbacks to customers in certain postcodes. Those postcodes correlated strongly with household income. The system was not deliberately discriminatory. It had simply learned from historical data that customers in those areas generated shorter, less complex calls, and optimized for call volume. The company could not produce a policy that governed what the AI was or was not allowed to optimize for. They had a technology answer. They did not have a policy answer. "We had built a system," Keona said, "but we had not decided what it was allowed to do."
Why AI Policies Are Different
Most organizational policies govern human behavior. An expense policy tells employees what they can and cannot purchase with company funds. An HR policy tells managers how to handle a performance conversation. When humans are the actors, policies work by informing and shaping human judgment.
AI policies are different in a fundamental way. The AI system itself does not read the policy. The policy must be operationalized - translated from a written statement into design constraints, training data requirements, evaluation criteria, and monitoring thresholds - before it has any effect on what the AI system actually does. A policy that says "our AI systems will not discriminate on the basis of protected characteristics" means nothing to a model. It only means something when that commitment is translated into specific data exclusions, fairness metrics in the evaluation protocol, and ongoing monitoring requirements in production.
This means AI policy development requires closer collaboration between legal, compliance, ethics, and technical teams than most other policy domains. A policy writer who does not understand how models are trained cannot write policies that constrain how models are trained. A technical team that does not understand what the policy is trying to achieve cannot translate the policy into meaningful constraints.
The AI Policy Architecture
Well-designed AI policy operates at three levels of specificity. Most organizations have policies at one or two levels but not all three, which creates gaps.
Enterprise principles
Enterprise principles are the top-level commitments that govern all AI use across the organization. They are written in plain language, apply to every AI system the organization deploys, and are short enough to be memorable - typically five to ten statements. Examples: "We do not use AI to make final decisions on matters of significant individual consequence without human review." "Our AI systems will not be used to deceive customers about their nature." "We will not deploy AI systems in contexts where we cannot adequately monitor their performance."
Principles are necessary but not sufficient. Every principle at this level requires at least one more specific policy to give it operational meaning.
Domain policies
Domain policies apply to specific types of AI use, specific data categories, or specific business functions. A data governance policy specifies what data can be used to train AI systems, under what conditions, and with what consent mechanisms. A model risk management policy specifies what review and approval process a model must pass before deployment based on its risk classification. A customer-facing AI policy specifies what disclosures customers must receive when they are interacting with or being affected by an AI system.
Domain policies are where most of the real governance work happens. They are specific enough to be operationalized and broad enough to apply across the range of systems within their scope. Each domain policy should identify its owner - the function responsible for maintaining it - and its review cadence. Twelve months is the maximum review interval for any AI-related policy; six months is more appropriate during periods of rapid regulatory change.
System-level requirements
System-level requirements are the specific technical, operational, and documentation requirements that apply to a particular AI system based on its risk classification and applicable domain policies. They are usually captured in the system's model card or governance documentation, not in a separate policy document. They translate the domain policies into specific, verifiable requirements: "This model must be evaluated for disparate impact across demographic subgroups before deployment," or "This system requires human review of all recommendations before they are communicated to customers."
Writing Policies That Actually Work
Most AI policies fail not because they are ethically misguided but because they are written in a way that makes them impossible to operationalize or enforce. Four principles distinguish effective AI policies from aspirational noise.
Specificity over aspiration. "We will ensure our AI systems are fair" is an aspiration. "Any model used in customer-facing decisions must achieve a disparate impact ratio of no less than 0.80 across gender, age group, and geographic region before deployment authorization" is a policy. Effective policies name specific actors, specific actions, specific thresholds, and specific consequences.
Scope clarity. A policy that covers everything equally prioritizes nothing. Effective AI policies are explicit about what they do and do not cover. A customer-facing AI policy may cover all AI systems that directly affect customer decisions but explicitly exclude internal productivity tools. Scope clarity allows the policy to be enforced practically - the compliance function knows what it is monitoring for.
Operationalizability testing. Before finalizing any policy, ask: could a reasonable technical practitioner translate this statement into a specific implementation requirement? If the answer is no, the policy is not ready. Run every policy statement through this test with someone from the technical team before publishing.
Exception processes. Some situations will not fit neatly within policy requirements. A policy with no exception process will either be quietly ignored or will block legitimate work. A policy with a defined exception process - named approver, documented justification requirement, time-bound authorization - creates a safety valve that maintains policy integrity while allowing flexibility when warranted.
Documentation as Policy Infrastructure
Policies require documentation infrastructure to be enforced. The relationship is direct: if a policy requires pre-deployment review, there must be a standard form or process for that review. If a policy requires monitoring at specific intervals, there must be a mechanism that generates and preserves monitoring records. If a policy requires escalation when performance drops below a threshold, there must be an alert system and an escalation log.
Keona's group, after the callback prioritization incident, rebuilt their AI policy framework around what they call "policy artifacts" - the specific documents or records that each policy requires as evidence of compliance. Every policy statement in their framework maps to at least one artifact: a model card, an approval record, a monitoring report, a training data inventory. If the artifact does not exist, the policy has not been met. This artifact mapping turned abstract policy commitments into an auditable compliance program.
>
A policy without an artifact requirement is a hope, not a control. The documentation is not the paperwork - it is the proof that the policy ran.
Keeping Policies Current
AI policy requires more frequent maintenance than most organizational policies because the technology, the regulatory environment, and the risk landscape all change rapidly. Three triggers should prompt a policy review outside the normal review cycle.
A material regulatory change. When a regulator in a jurisdiction where you operate publishes new AI guidance or rules, review all affected policies within thirty days for alignment.
A significant AI incident, internally or publicly. When an AI system fails in a way that was not anticipated - yours or a public incident at another organization - review whether your current policies would have prevented it or required earlier detection.
A material change to how an existing AI system is used. Policies approved for system A operating in context X may not be appropriate for system A expanded to context Y. A policy review should be a standard part of any significant scope change approval process.
Key Takeaways
- AI policies must be operationalized to work. A policy only affects an AI system if it is translated into design constraints, evaluation criteria, and monitoring requirements. Policy writers and technical teams must collaborate.
- Three levels of specificity are all required. Enterprise principles set direction. Domain policies govern specific areas. System-level requirements translate both into verifiable, enforceable constraints for individual systems.
- Effective policies are specific and testable. Replace aspirational statements with named actors, specific thresholds, and verifiable actions. Test every policy statement against: could a practitioner implement this?
- Every policy needs an exception process. A policy with no valid exception path will be quietly ignored when it blocks legitimate work. A defined exception process preserves integrity while allowing warranted flexibility.
- Map each policy to a documentation artifact. The artifact is the evidence the policy ran. Policies without artifact requirements cannot be audited and effectively cannot be enforced.
- Review policies more frequently than other corporate policies. Six to twelve month review cycles are appropriate for AI policy. Regulatory change, incidents, and system scope changes are additional triggers for out-of-cycle review.
- Scope clarity is as important as policy content. A policy that covers everything equally enforces nothing. Be explicit about what is and is not in scope.
Skill.re