AI Governance Policy
A versioned policy draft defining accountable ownership, risk-based approvals, information controls and evidence requirements across the AI lifecycle.
A companion implementation guideline connecting AI adoption to lifecycle gates, practical controls, evidence templates and stage-specific infrastructure.
AIG GUI 001 · Prepared . Company approval and effective-date fields remain uncompleted in the supplied draft.
THE DOCUMENT IN DETAIL
Source figure G25A: inspect the minimum internal route for input, output, stage and purpose flags. This planning aid does not authorize a use.
Loading the interactive design…
Read the complete document and supporting references.
[Company Name]
Status Draft for approval | Version 1.1 | Prepared 6 October 2026
Document identifier AIG GUI 001
Owner [AI Governance Lead] | Approving authority [Board or delegated executive authority]
Approval date [To be completed] | Effective date [To be completed] | Next review [Within 12 months of effective date]
Complete the approval fields and Company tailoring register before adoption.
This guideline turns the AI Governance Policy into a repeatable process for adopting AI, demonstrating business value, and controlling risk. Business owners use it to propose and run use cases. Governance and specialist teams use it to assess and approve them. Technical owners use it to deliver, monitor, and retire systems with evidence that the policy requirements are met, including the Company variations in P17 through P19.
Use the policy for mandatory obligations and this guideline for the workflow, templates, and operating details. Policy identifiers P01 through P25 provide traceability. The guideline cannot authorize activity that the policy prohibits. Proposed service targets, examples, and rollout sequencing are implementation defaults; the Company must approve or adjust them before enforcement. Example metrics are not universal safety thresholds.
Start by establishing ownership and the approved tools catalog, then discover current usage and launch a small portfolio of controlled pilots. Maintain one use case record from intake through retirement. Use a controlled register, evidence repository, approval workflow, and monitoring tools with clear ownership and access.
For each use, follow the gates in G05, apply the controls in G07 through G12, and retain the evidence in G15. Templates in G16 through G19 may be copied into the Company's systems. Do not collect live personal or confidential data for an initial idea assessment when a descriptive example is sufficient.
Appoint an executive sponsor and AI Governance Lead with the authority and time to run the program. Establish an AI Governance Committee with business, technology, enterprise risk, Security, Privacy, Legal, Compliance, Procurement, and HR representation appropriate to the portfolio. Identify independent validators and an alternate review route for conflicts of interest.
Approve a charter specifying mandate, quorum, delegated decisions, voting or consensus rules, escalation, conflicts, and records. Recommended cadence is monthly during launch and at least quarterly once stable, with urgent decisions handled through a recorded approval route. High-risk decisions require relevant specialist clearance; a majority vote cannot override a legal prohibition or unresolved mandatory control.
Publish a simple service model: who submits an idea, where they submit it, which information is required, how risk affects the path, and who can answer questions. Name a coordinator and deputy. Fund governance, validation, training, platform operation, support, and remediation; these costs must be included in adoption planning.
Use the responsibility matrix below. A is accountable, R performs the work, C is consulted, and I is informed. Specialist teams remain accountable for their own required clearances. The matrix does not waive required approvals.
Scroll horizontally to read all columns
| Activity | Business owner | Technical owner | Governance Lead or Committee | Specialists | Validator | Audit |
|---|---|---|---|---|---|---|
| Intake and business case | A R | C | C | C | I | I |
| Tier confirmation | C | C | A R Lead | C | C | I |
| Data access approval | R | C | I | A R Data owner | C | I |
| Build and controls | C | A R | C | C | C | I |
| Independent high-risk validation | C | C | C | C | A R | I |
| High-risk release decision | R | R | A Committee | C and clearance | C | I |
| Daily monitoring and recovery | C | A R | I | C | I | I |
| Business value and user outcomes | A R | R | C | C | I | I |
| Independent program assurance | C | C | C | C | C | A R |
The roadmap begins after a named sponsor authorizes mobilization. It is a proposed sequence rather than a legal transition period. Existing use must be assessed immediately; discovery activities do not authorize continuing a prohibited or unsafe deployment.
Scroll horizontally to read all columns
| Period | Deliverables and owners | Exit evidence |
|---|---|---|
| Days 1 to 15 | Sponsor and Lead establish charter, policy review, channels, initial catalog, and budget; Legal begins applicability register | Named decision makers, draft parameters, reporting route, and inventory schema |
| Days 16 to 30 | Lead discovers usage; owners classify priority systems; Security and Procurement review enterprise tools; HR launches basic training | Reconciled initial inventory, restrictions for unapproved tools, approved initial catalog |
| Days 31 to 60 | Owners run a few approved pilots; technical teams implement control patterns; validators review high-risk candidates | Baselines, test plans, evidence, user feedback, incident and fallback exercises |
| Days 61 to 90 | Authorized approvers scale successful pilots; Lead publishes dashboard and audit evidence; owners close priority gaps | Recorded release decisions, measured benefits, open-risk ownership, and operating review schedule |
Discover usage through software and procurement records, expense claims, access and network information where permitted, supplier questionnaires, application-owner interviews, and a nonpunitive employee disclosure route. Reconcile sources; an employee survey alone does not establish complete coverage. Include AI features already embedded in business applications.
For existing uses, record a disposition: continue within confirmed low-risk catalog scope, operate temporarily under a lawful approved exception, restrict to a sandbox, or suspend. Give remediation items an owner and deadline. Escalate unowned high-risk systems immediately. Do not use blanket grandfathering.
At day 90, the sponsor decides to scale, extend a bounded pilot, or stop, based on value, controls, user outcomes, cost, and support capacity.
Ask each function to identify a specific workflow problem, the current process, and the desired outcome. Compare AI with simpler alternatives, process redesign, and conventional automation. A use case without a measurable problem, accountable owner, or workable data rights should not enter delivery.
Record baseline volume, cycle time, quality, error rate, rework, user satisfaction, and total cost where relevant. Define what success means before testing. Evaluate feasibility through data availability, integration effort, supplier maturity, validation requirements, and available reviewers. Risk screening precedes any live pilot and cannot be offset by a high value score.
For eligible ideas, optionally score business value, feasibility, and readiness from 1 to 5 with written rationale. Rank by total; record any strategic or capacity-driven change in priority. This score must not determine the governance tier.
Choose a few varied pilots with clear boundaries and willing users. Suitable starting candidates may include internal drafting with permitted data, controlled knowledge search, and developer assistance with normal code review. They still require approved tools and data. Employment screening, eligibility decisions, autonomous payments, and safety-critical uses should receive specialist review before any pilot design.
Design an adoption plan with training, communications, champions, support, accessibility, and feedback. Explain what employees may do, what requires review, and how to report problems. Avoid license quotas or incentives that encourage use without business benefit.
Each gate results in approve, approve with conditions, return for remediation, reject, or retire. Record the decision, evidence, signer, scope, conditions, and next review date. Conditions precedent must be closed before the next gate. Only nonblocking items may remain open after an authorized owner accepts their residual risk.
Scroll horizontally to read all columns
| Gate | Minimum evidence | Decision owner | Policy |
|---|---|---|---|
| 0 Register and screen | Intake, owner, tool, purpose, preliminary tier, prohibited-use check | Governance Lead confirms route | P04 to P06 |
| 1 Authorize design and bounded experiment stage | Business baseline, impact assessment, legal screen, data approval, supplier review, test and oversight plan | Tier-specific approvers under P05 | P05 to P12 |
| 2 Approve production | Evaluation results, independent validation when high-risk, resolved blockers, support, monitoring, fallback, training | Tier-specific approvers under P05 | P09 to P13 |
| 3 Review operation or change | Performance, incidents, complaints, changes, risks, benefits, updated evidence | Original tier authority or recorded delegate within policy | P13 to P15 |
| 4 Retire | Dependency plan, access revocation, archive and deletion records, owner confirmation | Business owner and technical owner; specialist clearance where required | P14 |
Gate 1 must separately authorize each prototype, proof of concept, or pilot under G26; it does not authorize production. Specify participants, duration, allowed data, actions, environment, volume, and stop conditions. Use synthetic or appropriately de-identified data where possible, and document remaining reidentification risks. Live personal data requires privacy clearance even when the trial is small.
Recommended service targets are acknowledgment in two business days and a completeness check in five. For complete submissions, target five business days for low-risk routing, ten for moderate review, and twenty for high-risk review, adjusted for complexity. Publish lead times and blockers. Delay prompts escalation, never automatic approval.
Apply a prohibited-use screen first. Then assess inherent risk using the dimensions below and assign the highest applicable tier. Evaluate residual risk only after specifying and testing controls. Record the rationale and evidence; ambiguous classification is high-risk pending resolution under P05.
Scroll horizontally to read all columns
| Dimension | Low indication | Moderate indication | High trigger |
|---|---|---|---|
| Effect on people | Internal assistance with no material effect | Advice or interaction with limited consequences | Rights, employment, eligibility, health, safety, or material adverse effects |
| Information | Approved public or nonsensitive material | Confidential or limited personal data in approved service | D3 Restricted information under P20, sensitive personal data at scale, or severe exposure consequences |
| Action authority | Drafts reviewed before use | Integrated, bounded, reversible assistance | Consequential autonomous actions, material commitments, or critical control |
| Exposure and dependency | Small internal workflow with easy fallback | External service or broader internal dependency | Critical operations, large-scale harm, or substantial loss |
The table is an internal screening aid. A high trigger sets a high tier even when other dimensions are low. Confidential-data processing is never low simply because the model is easy to use. A system may be internally high-risk without meeting a legal definition of high-risk AI, or may have specific legal duties regardless of the internal tier.
Legal and Compliance must record jurisdictions, affected people, Company role, sector, data rules, automated-decision duties, notices, employment obligations, consumer protection, IP, safety, and contractual restrictions. For EU activity, determine relevant AI Act role and categories and check current binding obligations and effective dates. For other countries, identify national, state, local, and sector requirements rather than assuming a single global AI rule.
Record any required privacy impact assessment, fundamental-rights assessment, bias audit, conformity procedure, consultation, regulator engagement, or notice with an owner and due date. Counsel determines applicability. Review the legal register at least quarterly and before launch in a new location or use context. Do not copy historic legal timelines into operating checklists without verification.
Map information from source to processing, model or service, output, logs, storage, integration, and deletion. Include support access, subprocessors, telemetry, embeddings, backups, feedback, and any training. Obtain data-owner approval for each purpose and Privacy clearance where relevant. Align the matrix below with the Company's established classification labels and the D0 to D3 mapping and flags in P20. P20 determines the minimum handling and risk route.
Scroll horizontally to read all columns
| Provisional class | Baseline treatment | Evidence |
|---|---|---|
| Public | Approved tools only; verify rights and accuracy | Catalog scope and content rights |
| Internal | Approved enterprise service and configured protections | Owner authorization and configuration |
| Confidential | Specific data and use case approval; supplier training disabled unless separately approved | Data-flow review, contractual terms, access and retention controls |
| Restricted or sensitive personal data | Explicit specialist clearance and protected environment; high-risk where P05 triggers apply | Impact assessment, minimization, access, transfer and deletion evidence |
| Credentials and production secrets | Exclude from prompts, uploads, retrieval, and model context | Secret handling controls and testing |
Procurement must collect supplier answers and evidence before commitment. Verify the actual product edition, model, region, settings, and feature behavior. An enterprise brand name does not establish the contract or configuration for a particular account. Test access, retention, deletion, and training settings where feasible.
Ask for security assurance, architecture and hosting, subprocessors, incident handling, data uses, model limitations, evaluation summaries, material-change notification, continuity, export, and deletion mechanisms. For high-risk uses, seek evidence applicable to the intended population and task. Where direct audit is unavailable, evaluate alternative assurance and remaining gaps rather than marking the requirement automatically satisfied.
Legal must review licenses and output rights, including commercial distribution and generated code. Keep evidence of permission for source material. Record unresolved contract deviations in the risk register and route them under P08 and P15. Contract signature and technical access are separate steps; neither authorizes use case release.
Create approved architecture patterns for employee assistants, retrieval applications, predictive systems, customer-facing assistants, and agents. Each pattern should specify identity, permissions, approved data, network access, logging, supplier settings, evaluation, support, and disablement. Teams may reuse evidence only where configuration and context match; each use case owner remains accountable for applicability.
For retrieval systems, enforce source permissions before retrieval and presentation, restrict indexing, handle document updates and deletion, and record the source version used where needed. Test unauthorized queries and cross-user access. Source citations must point to material that actually supports the output; a citation alone does not establish factual accuracy.
For predictive systems, track dataset and feature lineage, quality checks, training and test separation, target leakage, calibration where relevant, group performance, and drift. Separate training, validation, and final holdout use. Do not tune repeatedly on the final evaluation set and then present it as independent evidence.
For generative systems, test factual support, unsafe content, prompt injection, confidentiality, refusal and escalation behavior, and downstream handling. Validate output before interpreting it as code, markup, a query, or a business instruction. Apply guardrails at the application and tool layers as well as in prompts.
For agents, build an action permission matrix. Define each tool's allowed operations, destinations, transaction limits, approval requirements, identity, and fallback. Enforce limits outside the model, require specific approval for consequential actions, and test duplicate execution, partial failure, untrusted tool output, unauthorized destinations, and credential revocation. Keep a safe-stop mechanism accessible to operations staff.
Before a pilot begins, the owner and reviewers must approve an evaluation plan stating the intended workflow, baseline, test cases, sampling, metrics, thresholds, reviewer method, environments, version identifiers, and stop conditions. Use representative and adversarial scenarios, out-of-scope requests, and known failure cases. Choose sample size based on consequence and variability; no fixed sample size guarantees safety.
Scroll horizontally to read all columns
| Test area | Evidence to collect | Release decision |
|---|---|---|
| Task quality and factual support | Representative results, baseline comparison, error severity, supported citations or calculations | Meets approved task thresholds and has acceptable limits |
| Privacy and security | Cross-user access tests, leakage tests, prompt injection, tool authorization | Unresolved severe exposure or unauthorized-action finding blocks release |
| Fairness and accessibility | Relevant group or scenario results, accommodation paths, known limitations | Material harm resolved or use constrained; lawful collection confirmed |
| Oversight | Reviewer accuracy, overrides, escalation, time and workload | Reviewers can meaningfully challenge and stop the workflow |
| Reliability and recovery | Failure, rate limit, outage, fallback, rollback, and disablement tests | Recovery and safe-stop behavior demonstrated |
| Operational value | Cycle time, rework, total cost, user outcomes | Demonstrated value supports continued investment |
Use human domain experts for consequential correctness. If model-based judging is used, calibrate it against human review, record disagreement, and protect against judge bias or shared model failure. Do not use an AI system's confidence statement as a validated probability. Repeat generative tests where variability could affect outcomes.
For people-related uses, document the choice of fairness measures, relevant groups or conditions, data constraints, uncertainty, and consequences. Avoid treating a single numerical ratio as proof of legal compliance. Where evidence is insufficient, limit the use or gather lawful evidence before release.
For a low-risk internal drafting pilot, an illustrative target might be improved end-to-end completion time with no increase in material errors after human review. The owner must set the actual threshold and measurement method. High-risk release must additionally have independent validation and no unresolved severe findings. Acceptance criteria and exceptions must be recorded before the launch decision, not invented after seeing results.
Write a pilot charter with the business problem, permitted tasks, participant group, owner, duration, approved data, environment, budget, oversight, and exit criteria. Establish a non-AI baseline or comparison group where practical. Keep the pilot bounded and prevent accidental expansion through sharing links, broader permissions, or new connectors.
Train participants on approved uses, prohibited inputs, output verification, escalation, accessibility, and workflow-specific limitations. Provide short examples of safe prompting and failure handling. Track understanding with an appropriate assessment; mere attendance is insufficient for people authorizing consequential actions.
Collect quality and time data across the whole process, including prompt preparation, checking, rework, integration, and support. Interview users about where the tool helps or slows them down. Include affected stakeholders when evaluating customer or workforce impact. Protect feedback confidentiality and avoid collecting unnecessary personal content.
Review results at a planned midpoint and at completion. Stop if a defined harm or control threshold is breached. At the end, decide to scale, extend within a new approved scope, redesign, or terminate. An extension must have a reason, budget, deadline, and approval. Successful adoption requires sustained task outcomes and controls, not just active-user counts.
The business owner submits a release packet containing the approved use case and tier, current legal and data clearances, supplier review, test results, independent validation where required, closed blockers, residual risks, and user readiness. The technical owner adds operating documentation, monitoring, support, credentials, rollback, recovery, and a tested disablement procedure.
Approvers must verify that the evidence applies to the exact version and configuration being released. Name the business owner, technical owner, on-call route, approved users, data categories, operating limits, review date, and change triggers. Record explicit specialist clearances. If a validator finds a material issue, document resolution or restrict the proposed use; do not relabel it a low-priority task to obtain approval.
Deploy in stages where feasible, with limited users or traffic and clear observation periods. Compare performance with the approved baseline. Revert or stop if thresholds are exceeded. Ensure a human or alternative workflow can handle demand during outages; a theoretical fallback without capacity is insufficient.
Maintain a change log covering models, prompts, datasets, retrieval sources, permissions, agents, guardrails, vendors, and integration behavior. Material changes return to the relevant gates. Smaller changes require documented impact analysis and targeted regression evaluation. Automatic supplier updates must trigger alerts and evaluation proportionate to the change; where adequate review cannot be maintained, constrain the deployment.
Define a monitoring plan before release. For each indicator, specify the owner, data source, frequency, approved threshold, action, and escalation route. Select measures appropriate to the use; do not collect every prompt or sensitive output merely to build a dashboard. Protect monitoring data under P07 and P14.
Scroll horizontally to read all columns
| Indicator | Calculation or observation | Action when outside limits |
|---|---|---|
| Quality | Material error count divided by reviewed outputs, with sample and severity | Restrict scope, increase review, investigate, or suspend |
| Harm and complaints | Confirmed harmful outcomes, complaints, and overrides by type | Escalate to owner and specialists; assess affected people |
| Security and permissions | Unauthorized access, leakage signals, or blocked and successful tool actions | Contain under incident process; revoke permissions if needed |
| Drift or change | Shift from approved input or output baseline; supplier or version change | Reevaluate and obtain required approval |
| Service and cost | Availability, response time, usage, total operating spend | Apply limits, restore service, or revise business case |
| Value | Verified process outcomes against baseline, including rework | Redesign, scale, or retire based on sustained benefit |
Finance and the business owner should distinguish potential capacity from realized savings. Time released is a capacity benefit until an evidenced change in expense or additional output occurs. Calculate net benefit using verified gains minus licenses, model usage, integration, validation, review labor, training, support, and remediation. Avoid double-counting the same time benefit as both cost savings and new revenue.
The quarterly governance dashboard should show registered use cases by tier and status, discovery gaps, overdue reviews, severe findings, incidents, active and expired exceptions, training coverage, realized value, and major supplier dependencies. State denominators and known data gaps. Inventory coverage is only measurable against a documented discovery population; “100 percent registered” must not be claimed from the register alone.
Publish one incident route with AI-specific triage guidance. An incident may involve unsafe or discriminatory output, personal data leakage, unauthorized agent actions, manipulated retrieval content, supplier failure, or material model degradation. Normal dissatisfaction with a draft may be a quality issue; suspected material harm requires incident assessment.
Scroll horizontally to read all columns
| Severity | Example trigger | Internal target from awareness | Response owner |
|---|---|---|---|
| Critical | Severe credible harm, material exposure, or uncontrolled consequential action | Immediate containment and notification within one hour | Designated response lead with Security and affected specialists |
| High | Significant bounded harm or control failure with credible material impact | Escalation within four hours | Response lead and business owner |
| Other | Lower-impact issue without evidence of material harm | Record within one business day and assign triage | Support owner or response lead |
Targets match P13 and are proposed internal requirements. Record when the organization became aware, when severity was assigned, and any reclassification. Legal and Privacy must separately identify notification deadlines and required recipients; internal targets do not create grace periods under law or contracts.
The response lead first protects people and information, limits affected functions, revokes access or stops actions if needed, and preserves controlled evidence. Capture system versions, approved scope, event times, relevant outputs or hashes, tool actions, authorizations, users affected, and supplier interactions without unnecessarily copying sensitive payloads.
Then determine scope and cause, assess obligations, coordinate communications through authorized staff, fix controls, and test corrections. Restart requires the approvals in P13. Maintain a decision log, notification record, recovery verification, and corrective-action plan with owners and dates. Conduct a tabletop and technical stop or fallback exercise for high-risk and critical dependencies at least annually, including vendor outage and agent permission failure.
Provide baseline training before tool access and refresh it annually or after material changes. Cover approved accounts, data handling, verification, rights in inputs and outputs, misleading content, reporting, and personal responsibility. Make guidance accessible and provide role-specific examples rather than a generic presentation alone.
Business owners need training in use case design, impact assessment, value measurement, and risk acceptance. Developers need secure architecture, data rights, evaluation, prompt injection, and version management. Reviewers need domain-specific validation, automation-bias awareness, escalation, and authority to stop. Executives need portfolio risk, decision rights, and evidence for investment decisions. Procurement, Legal, Privacy, HR, and support teams need their workflow-specific responsibilities.
Use champions to help teams apply approved patterns and identify useful cases. Champions do not approve new risk, change permitted data, or override governance. Establish office hours and a maintained knowledge base containing catalog entries, examples, forms, and incident guidance.
Review completed gates, review delays, repeated findings, user concerns, and incident lessons quarterly. Improve the controls and process where evidence shows a gap or avoidable effort. Give remediation a named owner, deadline, and verification method. The AI Governance Lead coordinates policy review annually; the approving authority adopts changes.
Maintain a controlled evidence folder or record for each use case. Assign an identifier such as AI 0001 and link it to tool, model, supplier, data, risk, incident, and change records. Limit access according to sensitivity. Keep an approval trail with signer identity, date, decision, scope, and evidence version.
Scroll horizontally to read all columns
| Policy requirements | Primary implementation evidence | Responsible function |
|---|---|---|
| P01 to P03 authority and accountability | Approved policy, charter, role assignments, delegated authorities | Sponsor and Governance Lead |
| P04 to P06 registration and classification | Inventory, catalog, intake, prohibited-use screen, tier rationale | Lead and business owner |
| P07 information controls | Data map, permissions, privacy assessment, retention and deletion schedule | Data owner and Privacy |
| P08 suppliers and rights | Due diligence, executed terms, rights records, exit plan | Procurement and Legal |
| P09 security and development | Architecture, threat assessment, versions, control tests | Technical owner and Security |
| P10 evaluation and oversight | Evaluation plan and results, reviewer tests, independent validation | Owner and validator |
| P11 agent actions | Permission matrix, action approvals, traces, stop and rollback tests | Technical owner |
| P12 workforce and transparency | Training, notices, claims review, feedback and challenge routes | HR and business owner |
| P13 operations and incidents | Monitoring plan, alerts, incident logs, recovery authorization | Operations and response lead |
| P14 change and retirement | Change log, continuity tests, archive and retirement record | Technical and business owners |
| P15 exceptions and assurance | Exception register, review records, dashboard, audit findings | Lead and Audit |
| P16 adoption decisions | Completed parameter register and approval record | Policy authority |
| P17 to P19 company variations | Size model, ownership oversight, sector assessment, combined tailoring | Sponsor, Lead, Legal and Compliance |
| P20 to P25 data and evidence controls | Classification and stage decisions, data passports, complete journal, protected evidence, replay manifests, toolchain controls and acceptance tests | Data owner, technical owner, Lead, Privacy and Legal |
Apply the Company's records schedule rather than a universal retention duration. Define which raw prompts or outputs are necessary, whether redacted samples or metadata are sufficient, and who may access them. Preserve legal holds. Restrict sensitive assessments and incident evidence while allowing reviewers to see the information necessary for their decisions.
Copy the fields below into a controlled use case record and add a response field for each item. Use “not applicable” only with a reason. Record missing evidence as open, not complete. The business owner prepares the record with technical and specialist input.
Explain how an incorrect, biased, unavailable, or manipulated output could affect people and the business. Identify vulnerable or underrepresented users, accessibility needs, power imbalances, and routes to raise a concern. Consider scale and whether harm can be reversed.
Describe the necessity of the data and how rights were established. Identify privacy and security risks across training, inference, retrieval, logging, and supplier support. Explain permissions, oversight, lawful notices, professional review, and consultation requirements. Record uncertainty and evidence gaps.
Compare the proposed workflow with the baseline and alternatives. Explain how controls reduce each risk and how their effectiveness will be tested. Record residual risk, operating restrictions, mitigation owners, review dates, and specialist determinations. Use a separate legally required assessment when the jurisdiction or sector requires it; this template does not automatically satisfy that requirement.
Copy this repeatable risk record and complete every field:
Do not reduce a risk rating solely because a control has been proposed.
Record the use case identifier, system version, configuration, datasets and their rights, environment, test owner, reviewers, date, intended scope, and baseline. For each test, record the scenario, expected behavior, metric, threshold fixed before execution, sample selection, observed result, uncertainty, severity, pass or fail, supporting evidence, and corrective action.
Include ordinary cases, edge cases, prohibited requests, unauthorized access, relevant group or accessibility scenarios, outages, and recoverability. Document limitations and repeatability. Link independent validation findings and their disposition. Retain enough information to reproduce material tests without unnecessarily retaining personal content.
Record gate and decision; use case and tier; approved versions and configuration; intended purpose; permitted users, jurisdictions, data, outputs, and actions; approval conditions; evidence reviewed; open nonblocking risks; operating thresholds; monitoring and support owners; review or expiry date; change triggers; and required signers.
Required signers are the P05 tier authorities plus specialist clearances triggered by the use. For high-risk use, include the Committee decision, accountable executive risk acceptance, and independent validation record. Record rejected or deferred decisions as well as approvals. Capture each signer's identity, date, and any scope limitation.
Record supplier, service edition, model dependencies, hosting locations, subprocessors, Company role, contract dates, renewal owner, and proposed uses. Attach evidence for data use and supplier training, security assurance, task limitations, incident cooperation, retention and deletion, material-change notices, output rights and licenses, service levels, portability, and exit.
For every gap, record consequence, mitigation, risk owner, decision authority, and renewal trigger. Capture whether evidence is verified, supplier asserted, unavailable, or not applicable with justification. Record the final approved configuration and prohibited features. Supplier approval alone does not authorize a use case.
Record exception identifier; linked use cases; exact policy requirement; business reason; alternatives; affected data and people; inherent and residual risk; compensating controls and test evidence; remediation owner and due date; start and expiry dates; relevant control owner; business risk acceptance; approval authority and decision.
Confirm the request does not concern a P06 prohibition, unlawful conduct, insufficient rights, or an unwaived contractual duty. Maximum duration is 90 days under P15 unless the policy is formally amended. Set reminders before expiry and define the restriction or suspension on expiry. Renewals require new review and evidence.
Record incident identifier; reporter and reporting channel; awareness, detection, classification, containment, and notification times; use case and system versions; severity and rationale; affected people, information, actions, and suppliers; evidence location; response lead; containment and access revocations; legal deadlines; authorized notifications; corrective actions; restart approvals; and follow-up verification.
Protect the record and separate factual findings from unverified reports. Include partial agent actions, pending transactions, duplicate-execution risk, and reconciliation outcomes where relevant. Link the risk register, change record, and retrospective.
Record reason, owner, retirement date, user and dependency notices, replacement or fallback, supplier termination, access and credential revocation, integration removal, retained evidence and legal holds, data deletion scope and verification, unresolved backup or retention obligations, and final approvals. Update the catalog and inventory so retired tools cannot be mistaken for approved services.
A team proposes searching internal confidential procedures through a managed retrieval assistant. The initial tier is moderate if outputs have limited consequences; it becomes high if used for material safety or regulated decisions. Tool approval does not cover the new retrieval source automatically. Intake must identify owners, document permissions, supplier data use, retention, and permitted outputs.
The pilot uses a restricted document set, enforces user permissions, tests cross-user access and prompt injection, checks whether cited sources support answers, and requires review before consequential use. Release requires the P05 moderate approvals and all triggered clearances, successful tests, and support and fallback. A new sensitive dataset or agent write capability returns to the appropriate gate.
A team proposes ranking applicants. The use is high-risk under the internal policy regardless of vendor claims or a human clicking approval. HR, Legal, Privacy, and relevant specialists determine lawful scope, required assessments, accommodation, notice, and challenge routes. Independent validation evaluates the intended population and decision process.
Solely AI-determined adverse employment decisions are prohibited by P06. A lawful assistance workflow requires meaningful review, documented limitations, and Committee and executive approval before production. Insufficient fairness or accessibility evidence is a reason to restrict or defer use, not to assume the supplier's system is acceptable.
An agent drafts invoice entries and suggests matches. Read-only access and human-reviewed drafts may fit a lower tier based on information and exposure. Enabling autonomous posting or payment changes the risk and requires reassessment; consequential autonomous actions trigger high-risk treatment.
The action matrix limits vendors, accounts, amounts, tools, and destinations. Payments require specific authorized human approval. Tests cover duplicate invoices, malicious attachment instructions, wrong beneficiaries, partial failures, permission revocation, and reconciliation. The agent cannot approve its own payment or expand permissions.
Before adoption, complete the following decisions and link them to the approval record. Pending decisions must not be represented as implemented controls.
Scroll horizontally to read all columns
| Decision | Owner | Required outcome |
|---|---|---|
| Company scope and jurisdictions | Sponsor and Legal | Legal entities, locations, sectors, applicable rules and effective dates |
| Policy authority and roles | Sponsor | Named approver, Lead, Committee, validators, specialists, deputies |
| Risk appetite and delegation | Policy authority | Approved tiers, materiality definitions, signers and conflict controls |
| Information and records | Data owner, Privacy, Records | Official classes, authorized environments, retention and legal holds |
| Operational parameters | Lead and response lead | Incident route and targets, review intervals, service targets, exception limit |
| Funding and rollout | Sponsor and Finance | Owners, budget, pilot portfolio, priorities, transition and remediation dates |
Select an operating model based on complexity and risk, not headcount alone. The models below describe minimum delivery arrangements; they do not relax P05 approvals, P06 prohibitions, or P03 independence. Combine this section with G21 and G22.
Use a lean setup: a named executive sponsor, part-time Governance Lead, a documented review group, approved enterprise tools, a controlled inventory, and a secure evidence folder. Create a short employee acknowledgment and a single support and incident route. Obtain qualified outside Security, Privacy, Legal, or independent validation support where needed; set access boundaries and confidentiality terms for those reviewers.
During the first month, discover current tools, restrict unapproved accounts and features, establish permitted data, train users, and record owners. Run one or two bounded pilots through the gates. Do not implement sensitive-data or consequential agents before the specialist review and stop mechanism are ready. Use a simple record for low-risk catalog patterns; reserve deeper assessments for integrated and high-risk workflows.
The sponsor should review the portfolio monthly during rollout and confirm funding for review and support. If the Company cannot provide competent oversight or independent high-risk review, reduce scope or defer the use. Buying a governance platform is optional; demonstrable controls are required.
Establish a cross-functional committee, a central catalog and inventory, functional champions, and a delivery workflow connected to procurement and change management. Allocate reviewer capacity and appoint deputies. Define which decisions business units can make and which require central escalation. Use shared templates and reusable architecture and evaluation patterns.
Pilot across a few functions, then expand in waves once test results, adoption support, monitoring, and unit-level ownership are reliable. Track review time, open findings, spending, incidents, and realized value. Reconcile supplier renewals and feature changes against the inventory. Procure additional workflow tooling when manual handoffs and evidence gaps justify it.
Operate enterprise standards with business-unit delivery and accountable local implementation. Use a shared control library, consolidated inventory, model and agent records, central approved patterns, and jurisdiction-specific requirements. Establish independent validation capacity and an assurance plan. Set documented delegation limits so local approvals remain visible centrally.
Implement discovery and evidence integrations where useful, with access restrictions and review of employee-monitoring implications. Test shared platform controls centrally and require teams to prove their use matches the tested pattern. Review common supplier dependencies, agent action limits, aggregated financial exposure, and cross-border data flows. Extend onboarding to subsidiaries and acquired businesses, with a clear remediation and escalation process.
Every model must identify who approves, who operates, who independently challenges high-risk work, where evidence is kept, how incidents are reported, and how the system is stopped. The required capability may be internal, shared, or outsourced; accountability stays with the Company.
Assign board oversight through an existing board or committee process where appropriate. Prepare a quarterly pack covering material AI dependencies, risks, incidents, open findings, spend, and verified benefit, with event-driven escalation for material issues. Provide directors with enough AI education to challenge strategy and evidence. A new board committee is not automatically necessary; McKinsey describes several structured engagement models suited to different company circumstances. McKinsey board technology governance.
Map AI uses affecting financial reporting, forecasts, investor materials, and external disclosures to the Company's financial and disclosure controls. Require Finance and Legal to approve sources, calculations, review, changes, retention, and publication authority. Keep market-sensitive material within approved access and supplier boundaries. Include AI-related events in the process used to assess disclosure obligations and materiality under the actual jurisdiction and listing rules.
Maintain evidence supporting claims of savings, revenue, capability, and deployment scale. Separate projections from achieved outcomes. Route external statements through authorized disclosure and communications reviewers; an AI draft does not create authorization to publish.
Choose an oversight forum that reflects ownership: founder-led executive review, a board, managing partners, or an investor-supported committee. Record delegation and conflicts. Include material AI issues in existing owner or board reviews rather than creating an unnecessary separate hierarchy.
Create a commitments register covering investor, lender, insurer, and customer obligations. Assign an owner for each assurance or reporting requirement. Prepare an evidence package suitable for commercial due diligence, including inventory, approvals, key controls, incidents, supplier dependencies, and open risks, while protecting sensitive records.
If ownership changes or an initial public offering is planned, reassess governance, disclosure controls, assurance capacity, and the records required for the new context. A public-company model does not automatically apply before listing, but contractual and legal duties must still be met.
At gate 0, identify the regulated activity and entity. Involve Compliance and the relevant domain specialists before choosing the architecture or trial population. At gate 1, complete the sector applicability assessment and integrate it with existing model, product, outsourcing, professional, or safety review. At gate 2, retain evidence that required conditions are satisfied; at gate 3, monitor under both the AI and sector processes.
Use the following screening prompts to identify specialist work. They are not a complete statement of legal requirements and must be tailored to the jurisdiction and use.
Scroll horizontally to read all columns
| Sector context | Specialist questions before deployment |
|---|---|
| Finance and insurance | Does the use affect models, credit or claims decisions, advice, trading, reporting, outsourcing, or operational resilience? |
| Healthcare and life sciences | Does it influence clinical decisions, patient data, research integrity, regulated products, or qualified professional responsibility? |
| Critical infrastructure and industrial operations | Could failure affect physical safety, continuity, infrastructure control, or certification requirements? |
| Professional and public-facing services | Does it affect confidential advice, legally protected information, eligibility, or accountable professional judgment? |
Attach the sector assessment to the use case record. Record any required regulator interaction, formal validation, approval, notice, audit, and retention period. Do not assume a supplier's certification covers the Company's deployed workflow. Escalate evidence gaps early; the Company must not proceed where it cannot demonstrate an applicable mandatory condition.
Complete a legal and contract screen despite the lack of sector-specific supervision. Prioritize ordinary business exposures: employee tools, customer interactions, personal data, generated content rights, supplier sharing, and automated actions. Use the common gates and risk tiers. Increase controls when the use creates material harm even if no sector-specific rule applies.
Scale with approved patterns and lightweight evidence for low-risk uses, and full specialist review where triggered. Customer procurement requirements and insurance terms may create assurance obligations beyond the Company's own baseline. Record and meet them rather than describing the Company as exempt from governance.
Gartner's public AI governance guidance emphasizes discovery, information and access governance, and monitoring and enforcement during operation. The inventory in G03, application controls in G08, and monitoring in G12 apply that direction through Company-specific processes. Gartner AI Governance Needs More Than Policies.
McKinsey's 2026 AI trust research highlights clear accountability, training, and controls for increasingly autonomous systems. Its survey describes associations rather than proof that a particular governance investment will produce a return. G02, G08, and G14 translate those themes into assigned roles, constrained agent permissions, and role-based learning. McKinsey State of AI Trust in 2026.
BCG recommends executive leadership across functions, reviews differentiated by risk, controls throughout development, testing, monitoring, and response. G05 through G13 make those activities part of the release and operating process. BCG Responsible AI Needs More Than Good Intentions.
Bain's adoption guidance describes leadership sponsorship, employee participation, training, champions, and an accessible repository of tools. G04, G10, and G14 apply those practices to bounded pilots and supported rollout. Its discussion of agents emphasizes defined roles, permissions, and oversight; G08 makes those boundaries explicit. Bain How to Accelerate Progress on AI and Bain Google Cloud Next 2025.
These public insights inform the implementation approach. They do not prescribe the Company's exact tiers, internal deadlines, organizational models, or legal obligations. Those parameters require Company approval and use-specific assessment.
Provision approved Company accounts only after training and acknowledgment. Publish the allowed tasks, information classes, output review, retention, and prohibited features. Manage extensions, plug-ins, connectors, file uploads, memory, meeting transcription, and vendor default activations as distinct capabilities. Give users a straightforward route to request additional capabilities.
Use normal code review and delivery controls for coding assistants. Review commercial content, calculations, citations, and commitments before use. Provide anonymized or fictional prompting examples; avoid placing real secrets in training materials. Periodically sample compliance using lawful, proportionate methods and disclose employee monitoring where required.
Before launch, approve the user journeys, interaction notices, permitted information, accessibility, supported languages, complaint handling, and transfer to a person. Identify situations the system must decline or escalate, including out-of-scope professional advice, emergencies, material eligibility disputes, and unauthorized account changes.
Test identity and cross-customer access, factual support, manipulation, abuse, vulnerable-user scenarios, agent authorization, and human handoff. The system must not invent prices, contractual terms, coverage, approvals, or customer commitments. Use approved knowledge sources and enforce transaction authority through business systems.
Provide a human service route with realistic staffing and response expectations. Log unresolved issues and material corrections without excessive personal content. Monitor erroneous promises, complaints, transfer failures, and harmful outcomes separately from conversation volume. High-risk purposes remain subject to independent validation and P05 approval.
Start each intake with a source and data passport, not a product selection. Identify everything the system can access, including retrieved documents, historical conversations, images, audio, files, tool responses, evaluation datasets, telemetry, and supplier support. Record current and anticipated classifications and purpose restrictions. Unknown rights or classification stop execution until resolved.
Compute the default handling class from the most restrictive accessible source, assembled input and context, anticipated or detected output, and evidence. Assess personal-data flags, sensitive personal information, privilege, license conditions, jurisdictions, recipient, material decisions, and agent authority separately. Apply the stricter result. D2 imposes at least a moderate route and D3 imposes a high route under P20; a public-data use can still be high-risk because of the decision or action.
Scroll horizontally to read all columns
| Route | Prototype and proof of concept | Pilot | Production |
|---|---|---|---|
| D0 or D1 with no higher trigger | Approved sandbox; fictional or rights-cleared data; no consequential live effects; complete journal | Approved participants, task evaluation, output review, support and monitored limits | Approved release, controlled identities, operating evidence, monitoring and change review |
| D2 or ordinary personal data | Prefer minimized or validated surrogate data; real data only after necessity and specialist approval | Segregated processing, approved supplier terms, context access checks, protected evidence, permitted recipients | At least moderate approval; enforced retention, data subject handling where applicable, incident and recovery readiness |
| D3 or high-impact purpose | Surrogate-first; real data requires high-risk authorization and equivalent information protection | Explicit high-risk scope, independent challenge, strong oversight, restricted evidence and tested disablement | Committee and executive approval, independent validation, lawful complete evidence, high-risk operating and review requirements |
| Unknown rights class or prohibited use | Quarantine or block | Do not launch | Do not launch or suspend affected scope |
Declassification is a separate release decision. For example, a report generated from D2 customer contracts remains D2 until the owner verifies that its public version contains no confidential terms, identifiers, inference disclosures, or contractual restrictions. Record the source and output versions, redaction method, tests, authorized reviewer, intended audience, and expiry or context restrictions. The original evidence remains protected even if a released derivative is public.
Masking and synthetic generation occur inside the original information boundary. A synthetic dataset trained from sensitive records requires privacy leakage and utility evaluation; it is not automatically anonymous. Pseudonymized information remains personal information where applicable. Evaluate linkage, rare records, membership inference, and attribute inference where relevant. Do not downgrade a class based on removal of names alone.
Figure G25A Data and stage routing explorer
The illustration below shows the proposed baseline route. It is a planning aid, not an approval tool. It never declassifies output or permits a prohibited use. The policy and specialist determination remain authoritative.
Loading the interactive design…
Use the following stage records within the G05 gates. Gate 0 registers and screens. Gate 1 authorizes a particular prototype, proof of concept, or pilot with its own scope and exit evidence. Gate 2 authorizes production. Gate 3 covers operation and material change. Gate 4 covers retirement. Store each authorization separately; an earlier Gate 1 decision does not authorize a later stage.
Scroll horizontally to read all columns
| Stage | Required application and infrastructure capabilities | Required decision and evidence |
|---|---|---|
| Intake | T03 catalog or controlled metadata record; T13 approval workflow | Owner, purpose, source rights, class mapping, data flags, risk route and prohibited-use screen |
| Prototype | T01 managed runtime; T02 identity; T06 version control; T11 trace capture; T12 protected evidence; T16 durable journal; relevant T04 input screening | Bounded sandbox approval, permitted information, prompt and configuration versions, example outputs and transaction records |
| Proof of concept | Prototype capabilities plus T05 minimization where needed; T07 artifact registry; T10 evaluation; T15 versioned data preparation; T08 for retrieval; T09 for required guardrails | Dataset passport and snapshot, architecture, leakage and quality results, limitations, feasibility decision and stage expiry |
| Pilot | Segregated environment; all triggered capabilities; T13 human approvals; T14 incident integration; monitoring and tested fallback | Live-data and participant approval, meaningful review, independent high-risk challenge, acceptance criteria, journal coverage and recovery tests |
| Production | Hardened scalable runtime and connectivity; release controls; protected complete evidence; ongoing tracing security monitoring and governance | Exact release manifest, P05 approvals, high-risk validation, operating thresholds, data and output permissions, retention and rollback |
| Change and retirement | Updated lineage and registry; controlled release or access revocation; archive deletion and supplier exit workflows | Impact and reapproval decision or retirement certificate; dependency review; evidence retained lawfully and replay limits updated |
T01 through T16 refer to G30. Capability is mandatory where the use triggers it; a separate purchased product is not. An approved platform can supply several capabilities. Requirements for privacy review, transaction evidence, lawful data use, and access apply at all stages.
Create an approved account or project, allowlisted model access, participant identities, a controlled data folder, versioned prompts or code, and a durable journal. Use invented or approved D0 or D1 material by default. Disable production connectors and consequential write permissions. Set budget and duration limits. Demonstrate that inputs, actual outputs, refusals, versions, and responsible users are recorded before expanding scope.
The exit decision records whether the idea warrants a proof of concept, what was learned, representative outputs, limitations, costs, rights and classification questions, and the next stage's dataset and reviewers. Discard experiments under the approved schedule; preserve the decision and permitted evidence.
Prepare a bounded versioned dataset inside the authorized boundary. Record cleansing, annotation, masking, synthetic generation, transformations, quality checks, splits, and access. Document the model, embedding model, retrieval index and query configuration where used. Pin available versions, or explicitly record a supplier-managed version limitation. Establish holdout scenarios and acceptance criteria before testing.
Use a separate project or account, restricted network destinations, protected identity and keys, version and artifact records, evaluation harness, and complete event capture. Evaluate utility, leakage, prompt injection, access, output classification, and replay capability. A real-data proof of concept requires the P21 necessity decision and relevant specialist approvals. Its output must not drive a consequential live workflow.
Document the actual workflow, users and affected people, notices, approved live data, human reviewers, transactions, observation period, support owner, and stop criteria. Deploy a production-like boundary sufficient for the data class, with limited exposure and no unapproved actions. Test the journal under outage and retry conditions, human override, output release, supplier failure, and recovery.
Measure end-to-end quality, harmful outcomes, rework, review workload, time, cost, and adoption. Reconcile every request and action to its event record; sample quality evaluations independently. The pilot report must state whether criteria were met and whether the evidence supports scale, redesign, another bounded pilot, or retirement.
Release the approved build through a controlled pipeline with separate deploy authority and environment-specific permissions. Retain the release manifest and evidence links; establish approved capacity, support, recovery, costs, alerts, and supplier update detection. Enforce data and output routes at runtime. Missing classification, approval, or required durable evidence must block the affected activity.
Reassess added datasets, new output recipients, agent tools, model changes, and expanded geography or scale. Record each release and rollback with its decision and validation. Retire credentials, indexes, embeddings, caches, supplier copies, and integrations under P14, keeping only legally authorized records and clearly documenting what can no longer be replayed.
Figure G26A Controls accumulate as exposure increases
The stage illustration distinguishes the purpose, environment, evidence, and approval required at each step. Data sensitivity can require stronger controls before a proof of concept begins.
Loading the interactive design…
Use an approved gateway or equivalent application boundary for identity, approved use and data checks, versioned routing, input screening, limits, and journal intent creation. Enforce retrieval permissions before context reaches the model. Inspect output, apply inherited labels, and obtain required review before release. Tool authorization must be enforced by the receiving service as well as the application.
Keep the audit journal, classified evidence vault, working datasets, retrieval index, operational telemetry, and security monitoring as explicit logical stores. They may share infrastructure if isolation and duties are adequate; they must not share unrestricted access or retention assumptions. Do not send raw sensitive payloads to a general monitoring destination merely because it supports traces.
Figure G27A Governed request and evidence architecture
flowchart TD
U[Employee customer or batch request] --> I[Identity purpose and entitlement]
I --> G[Gateway classification input controls and intent]
G --> R[Permission checked retrieval and context]
R --> M[Approved model and deployment]
M --> O[Output classification and release controls]
O --> H[Required human review or action approval]
H --> A[Authorized recipient or receiving system]
G --> J[Durable journal]
R --> V[Classified evidence vault]
M --> V
H --> J
A --> J
J --> S[Reconciliation security and governance review]For D2 and D3, document permitted regions, private connectivity or equivalent assessed controls, encryption and key ownership, administrator access, supplier support, model training settings, and every outbound destination. Separate production from experimental projects. Validate tenant, user, and document boundaries through attempts to cross them rather than rely on a diagram or declared policy.
Use a workflow ID for the overall business task, transaction IDs for distinct operations, and event IDs and monotonically ordered sequence values within each transaction. Correlate related model calls, retrievals, tools, human reviews, approvals, and downstream effects. Record intents and outcomes, including denied and failed activity. Reconcile incomplete records; distributed systems can leave an action's outcome temporarily unknown.
Scroll horizontally to read all columns
| Field group | Implementation record |
|---|---|
| Correlation | workflow_id transaction_id event_id parent_event_id sequence event_type provider_request_id |
| Actor and authorization | User or agent identity, role, tenant, purpose, entitlement check, approval record and policy version |
| Stage and time | Prototype POC pilot production or change; event and ingestion times; clock source; started completed denied failed cancelled or outcome_unknown |
| Data and lineage | Source and dataset versions, transformation IDs, input context output and evidence classes, privacy and rights flags, retrieval chunk IDs and source versions |
| Model and build | Supplier deployment model identifier version or fingerprint; prompt template and assembled prompt evidence; parameters; code commit build and dependencies |
| Retrieval configuration | Index and embedding version, access filter, query and ranked retrieved context references actually used |
| Evidence | Protected input and actual output references; content integrity values; retention and legal hold; redaction or lawful deletion status |
| Controls | Input and output check results, rule and guardrail versions, policy decision, relevant thresholds and limitation flags |
| Human decision | Reviewer identity time and authority; approved rejected corrected or escalated; actual reviewed version; concise rationale and supporting evidence |
| Action and effect | Tool and version, intended destination and parameters, scoped approval, idempotency key when supported, actual recipient response, affected record and reconciliation |
| Replay | Historical reconstruction available_until; supported replay level; artifact availability; known supplier or nondeterminism limits |
| Integrity | Schema version, journal signature or integrity verification, authorized writer, checkpoint, verified linkage and access trail |
Create a dataset and configuration manifest once per immutable version and refer to it from events; avoid copying the same raw information into every log. Protect the assembled context actually submitted, not just the prompt template. For streaming, preserve the final assembled response or permitted chunk references, partial delivery, cancellation, and outcome. For batch jobs, record each constituent operation or a manifest with per-item identity, outcome, and evidence; a job-level success count alone is insufficient.
The searchable journal should contain minimally necessary metadata and protected evidence references. The evidence vault holds permitted prompts, context, actual responses, human corrections, action payloads, and immutable source snapshots or authorized equivalent references. Access must be purpose-specific, authenticated, logged, and limited to authorized investigators and reviewers. Identifiers and integrity values may still be sensitive and require appropriate classification.
Operational tracing may be sampled and temporary. The transaction journal must cover all in-scope events and have durable recovery and reconciliation. OpenTelemetry provides a useful interoperability approach, but its evolving GenAI conventions must be version-pinned and mapped to the Company's required schema. Do not assume instrumentation captures human decisions, business approval, or receiving-system effects automatically. OpenTelemetry GenAI conventions.
For integrations, implement durable intent before execution, explicit result status, and reconciliation of missing outcomes. Use an outbox or equivalent durable delivery pattern and retry with safe identifiers. Ensure the receiving system enforces authorization and duplicate protection. An interrupted request must not be blindly reissued if it may already have made a payment, published content, or changed a record.
Figure G28A Recorded model request and actual response
sequenceDiagram
participant U as User
participant G as Gateway
participant J as Journal
participant V as Evidence
participant M as Model
U->>G: Authorized request
G->>J: Record intent
G->>V: Protect input and context
G->>M: Invoke approved version
M-->>G: Actual response or failure
G->>V: Protect actual output
G->>J: Record result and checksFigure G28B Recorded decision and authorized downstream action
sequenceDiagram
participant G as Gateway
participant H as Reviewer
participant J as Journal
participant B as Business system
G->>H: Review exact output
H->>J: Decision and action scope
G->>J: Record action intent
G->>B: Execute scoped action
B-->>G: Result or unknown
G->>J: Effect or reconciliationThese sequences illustrate required evidence boundaries, not a universal transaction protocol. Test the chosen system's behavior when any write, call, or response fails. A missing result must remain outcome_unknown until verified. Never record a synthetic model rationale as if it were an observed internal reasoning trace.
Begin with authorized investigation access, a stated purpose, and the transaction ID. Retrieve its intent and outcome sequence, applicable approval, versioned policy and controls, actual input and output evidence, relevant sources, model configuration, review, and downstream effect. Verify integrity, completeness, clock order, and the identity of the actors involved. Record any missing, expired, deleted, or supplier-inaccessible evidence.
Explain the actual business decision using sources, applicable rules, human judgment, and recorded action authority. Separate observed facts from inferred causes. A model output alone does not establish why a reviewer approved it, and a later generated explanation does not establish the model's internal cause. For material decisions, domain reviewers must assess whether the record supports the decision and any required explanation to the affected person.
For replay, restore only approved copies of retained versions into an isolated environment with live tools, publishing, payments, and production write access disabled. Reconstruct preprocessing, retrieval, prompts, model parameters, policy checks, and review boundaries. Capture backend fingerprints where available. Use recorded historical tool results or a safe simulator; rerunning a historical workflow must not repeat its real-world effect.
Scroll horizontally to read all columns
| Assurance level | Evidence and acceptance | Limitation |
|---|---|---|
| Historical reconstruction | Recover actual permitted inputs outputs controls approvals and effects during retention | After lawful deletion, state precisely which evidence is unavailable |
| Controlled replay | Rerun retained inputs and configurations safely; compare outcomes and record variance | Hosted updates, nondeterminism, missing weights, changing external sources or environment may prevent equality |
| Validated deterministic reproduction | Demonstrate the defined identical result for a controlled workload with frozen artifacts and tested environment | Applies only where demonstrated; it is not inferred from a seed or temperature |
The replay report records investigator, purpose, transaction, retained versions, reconstruction completeness, comparison method, differences, explanation quality, and actions. Use task-specific tolerances for functional comparisons; retain the original output rather than replace it with a regenerated answer. Microsoft explicitly documents variability even with a matching seed and system fingerprint. Microsoft reproducible output.
Figure G29A Three different evidence promises
flowchart TD
A[Declared evidence capability] --> H[Historical reconstruction required]
A --> R[Controlled replay where required and supported]
A --> D[Identical reproduction only when validated]
H --> F[What actually happened]
R --> C[What a retained configuration does in a safe rerun]
D --> V[Whether a defined result can be matched exactly]The following shortlist identifies three or four prominent provider options for each capability, using official product documentation available on 6 October 2026. It is a procurement starting point rather than an independently measured global market ranking. Selection must follow data class, jurisdictions, interfaces, existing infrastructure, evidence requirements, total cost, and tested control coverage. Vendor descriptions establish advertised scope; they do not prove the Company's configuration is compliant.
Scroll horizontally to read all columns
| Class | Capability required in the journey | Three or four provider options with official sources |
|---|---|---|
| T01 Managed AI runtime and compute | Approved models deployments projects regions isolation and resource limits from prototype onward | AWS Amazon Bedrock; Microsoft Foundry; Google Cloud Gemini Enterprise Agent Platform |
| T02 Identity and resource authorization | Company identities workload identities least privilege administrative separation and approved key or secret integrations | Microsoft Entra; AWS IAM; Google Cloud IAM |
| T03 Data catalog and lineage | Source ownership classification purpose datasets transformations and derivative relationships from intake | Microsoft Purview Data Map and Unified Catalog; Collibra; Atlan; Alation |
| T04 Sensitive data discovery and loss prevention | Discover classify screen and constrain sensitive information before processing and release | Microsoft Purview DLP; BigID; Google Sensitive Data Protection |
| T05 Data minimization masking and synthetic preparation | Prepare nonproduction data and document transformation privacy leakage and utility tests | Tonic.ai; MOSTLY AI powered by Syntho; Perforce Delphix |
| T06 Version control and release pipeline | Review code prompts configuration infrastructure definitions tests builds and controlled releases | GitHub Actions; GitLab CI CD; Microsoft Azure Pipelines |
| T07 Experiment and model artifact management | Track experiments dataset references model versions evaluations metadata and stage approvals | Databricks MLflow; AWS SageMaker Model Registry; Google Agent Platform Model Registry |
| T08 Retrieval and vector search | Version indexes and embeddings retrieve approved context and integrate document level authorization | Microsoft Azure AI Search; Elastic; Pinecone |
| T09 AI input and output guardrails | Screen harmful content prompt attacks sensitive disclosure and configured prohibited topics where supported | AWS Bedrock Guardrails; Microsoft Azure AI Content Safety; Google Model Armor |
| T10 AI evaluation and adversarial testing | Test task performance safety retrieval oversight and failures before release and during operation | Microsoft Foundry evaluation; AWS Bedrock evaluation; Google Gen AI evaluation |
| T11 AI tracing and quality observability | Correlate model retrieval and agent activity; inspect failures and build evaluation evidence | Arize Phoenix; LangChain LangSmith; Langfuse |
| T12 Protected evidence storage and retention | Store versioned classified evidence with approved integrity access retention legal hold and recovery controls | AWS S3 Object Lock; Microsoft Azure immutable Blob Storage; Google Cloud Storage Bucket Lock |
| T13 AI governance and approval workflow | Inventory assess approve assign owners manage exceptions and retain decision evidence | OneTrust AI Governance; IBM watsonx.governance; ServiceNow AI Control Tower; Credo AI |
| T14 Security monitoring and incident operations | Detect correlate investigate and respond to AI access misuse leakage and system failures | Microsoft Sentinel; Splunk Enterprise Security; Google Security Operations |
| T15 Data ingestion transformation and quality pipelines | Version and operate ingestion cleansing deduplication schema checks approved transformations and quality quarantine | AWS Glue; Microsoft Fabric Data Factory; Google Dataflow |
| T16 Durable transaction journal backend | Persist intents ordered or causally linked events outcomes approval references and reconciliation with tested durability | AWS DynamoDB transactions; Microsoft Azure Cosmos DB transactions; Google Spanner transactions |
These classes are related but not interchangeable. IAM services need appropriate federation, workload identities and key or secret systems. A vector store needs application and document authorization. An experiment registry needs additional prompt, retrieval, deployment and dataset records when it does not capture them. A safety service needs separate business authorization and data controls. Evidence storage needs an application journal, integrity verification, and an approved retention design. A governance workflow needs live integrations and accountable reviewers.
T15 must record data version, transformation code, responsible actor, row or document quality disposition, rights and classification propagation, and lineage for each material data change. T16 providers are transactional storage candidates, not ready-made complete or immutable AI audit journals. The application must implement append-only permissions, integrity, causal order, duplicate handling, protected evidence links, recovery and reconciliation. Validate actual transaction scope and partition or regional constraints. T12 preserves protected evidence and approved immutable checkpoints; it does not replace T16's transaction journal.
For a small Company, begin with approved capabilities already available in its cloud or enterprise suite, a controlled inventory and decision workflow, and a durable evidence store. Add specialist products when a tested gap requires them. For medium and large companies, prioritize shared integration, consistent event schema, cross-entity isolation, policy enforcement, and consolidated assurance. A regulated use must meet its actual evidence and oversight duties even when that requires a different service or operating model.
Use mandatory pass or fail requirements before scoring convenience or cost. Reject or restrict candidates that cannot meet lawful processing, data use and supplier training terms, required regions, confidentiality, purpose restrictions, necessary version or evidence retention, access isolation, deletion, incident cooperation, or receiving-system authorization. Procurement must assess the exact product edition and contract, not the broad vendor portfolio.
Then compare integration with existing identity and data systems, APIs and audit exports, supported modalities, prompt and dataset management, evaluation extensibility, accessibility, operating resilience, portability, staffing, licensing, and total cost. Include both infrastructure spend and human review, evidence storage, remediation, and supplier exit. Record the chosen option, alternatives, scope, unresolved gaps, compensating controls, accountable owner, and review date.
Complete a control proof of concept before trusting a candidate with D2 or D3 data. Use fabricated identifiers and realistic attack scenarios first. Demonstrate unauthorized-input blocking, cross-user retrieval denial, output label propagation, restricted recipient blocking, actual prompt and output capture, human-decision linkage, event reconciliation, evidence verification, replay without side effects, lawful deletion, and recovery. Verify streaming, batch, agents and provider endpoints actually used.
AWS documents that Bedrock model invocation logging is disabled by default and has endpoint and destination restrictions. Enabling a logging option is therefore an implementation task requiring coverage tests, retention and access review, not proof of universal evidence capture. AWS invocation logging. Microsoft Foundry documents sampled continuous evaluation; this is distinct from the complete journal required by P23. Foundry observability.
Test immutable retention in a disposable test store before configuring business evidence. Legal and Records must approve the retention schedule and hold requirements. Google documents that a locked bucket retention policy cannot be removed or shortened. Google Bucket Lock. Encryption does not by itself solve erasure, lawful access, or governance; a key-destruction approach must be validated against the actual obligation and architecture.
Figure G31A Procurement approval follows demonstrated control coverage
flowchart TD
R[Define data stage and evidence requirements] --> L[Shortlist existing and candidate services]
L --> P{Mandatory requirements met}
P -->|No| X[Reject restrict or redesign]
P -->|Yes| T[Demonstrate end to end controls]
T --> V{Tests and contractual review passed}
V -->|No| M[Remediate and retest]
M --> T
V -->|Yes| A[Approve exact edition configuration and scope]
A --> O[Monitor changes and reassess]A team uses rights-cleared public product descriptions to prototype an internal drafting assistant. D0 inputs and expected outputs can use the low-risk pattern only if the purpose creates no higher trigger. The tool must still use an approved account, versioned prompt, permitted model, transaction journal, output review and recorded prototype approval. If a connector exposes internal customer records, recompute the route before enabling it.
A contract-search proof of concept initially uses invented contracts. Adding live confidential customer agreements changes the handling to D2, and privilege or particularly sensitive provisions may require D3. The owner must document necessity, rights, specialist clearance, access and retention, isolated processing, supplier terms, and the exact dataset snapshot before ingestion. Retrieved passages, prompts, summaries and evidence inherit the restrictions. A public summary requires a separate declassification and release decision.
A permitted pilot proposes using sensitive personal information for a consequential recommendation. The route is high-risk regardless of the small participant group. Legal determines whether the purpose is lawful; the policy's prohibited decisions still apply. The pilot requires data-owner clearance, privacy assessment where required, independent challenge, controlled participants, meaningful review, protected actual evidence, complaint and incident paths, and a tested stop mechanism. No governance score can override a legal prohibition.
An invoice agent extracts an amount and beneficiary from an authorized document, proposes a payment, and obtains specific human approval. Its record links the document version, extraction output, classification, model and prompt configuration, permissions, reviewer and approved amount, recipient account, action intent, and receiving-system result. If the payment call times out, record outcome_unknown and reconcile with the payment system before retrying. A later replay uses a simulated payment endpoint and the retained historical response.
A system combines individually public records to infer a person's health or other sensitive characteristics. D0 source availability does not make that output D0 or establish a lawful purpose. Privacy and Legal assess the inference and permitted use; the output and evidence receive the required higher classification and approval route, or the use is blocked. Record the determination and actual release restrictions.
Record dataset ID and immutable version; owner and steward; source and acquisition rights; approved purposes and prohibited uses; subjects jurisdictions and privacy flags; D0 to D3 class; collection or consent conditions where applicable; transformation and annotation records; quality and representativeness; masking or synthetic leakage tests; access; supplier transfers; retention and holds; correction and deletion dependencies; and approval evidence.
Record use case; lifecycle stage; purpose; permitted input context output and evidence classes; privacy and rights flags; data and model versions; environment and supplier edition; users and recipients; allowed tools and actions; blockers and limitations; tests and evidence; specialist and tier approvers; rationale; conditions; expiry and next review. Include a distinct row for every stage decision and any rejected scope extension.
Record transaction schema version; capture points; complete journal coverage and reconciliation method; protected payload scope; source and configuration manifests; evidence integrity; access; storage region; retention and legal hold; deletion behavior; historical reconstruction window; supported replay level; supplier limitations; side-effect isolation; tested recovery and investigator permissions.
For each T01 to T16 capability, record selected service and edition, environment, region, accountable owner, native or integrated implementation, configuration version, contract reference, data boundary, audit export, known gaps, control tests, residual risk approval, operational support and renewal trigger. Mark capabilities not triggered by the use as not applicable with a reason; do not mark untested controls implemented.
P20 and P21 map to G25 and G26. P22 maps to the dataset passport, data lineage and decision register. P23 maps to G28 and G29. P24 maps to G27, G30 and G31. P25 maps to these acceptance tests and G12 monitoring. The reference architecture and tools support the existing size, ownership and regulatory variations; they do not replace them.
Read this guideline with the AI Governance Policy. Internal templates, gates, timing targets, examples, and role assignments are implementation proposals rather than legal classifications or claims of standards certification.
NIST AI Risk Management Framework 1.0 informs governance, context assessment, measurement, and risk treatment. NIST Generative Artificial Intelligence Profile informs attention to generative-system limitations and misuse.
ISO IEC 42001 public overview describes an AI management system standard; this guideline does not reproduce its copyrighted requirements or establish certification readiness. European Commission AI Act overview is a starting reference for jurisdiction-specific review, with binding text and amendments to be checked by Legal.
CISA and UK NCSC secure AI development guidance and joint secure deployment guidance support secure delivery and operation within their stated scope.
ICO governance and accountability toolkit and EEOC artificial intelligence and ADA resources support specialist assessment of privacy and employment-related use.
Sources reviewed on 6 October 2026. Legal must keep the applicability register current; the references are not a complete statement of applicable law.
NIST SP 800 218A addresses secure development practices for generative AI and foundation models. NIST SP 800 53 Revision 5 provides a security and privacy control catalog, including audit-related control families. Their requirements require use-specific selection and tailoring; this policy is not a claim of full conformance.
W3C PROV Overview supports the entities activities and agents vocabulary for provenance. OpenTelemetry GenAI conventions support interoperable telemetry; the Company must still implement complete business decision and transaction evidence.
Microsoft reproducible output documentation explains why fixed seeds and fingerprints do not guarantee deterministic output. AWS invocation logging describes supported capture configuration and limitations. Google Bucket Lock describes immutable retention behavior. These implementation examples support the evidence and retention design; they do not make a particular vendor mandatory.
Minimum internal route: Moderate. Other legal privacy exposure decision or action triggers can increase controls or prohibit use.
Planning aid from G25A. It never declassifies an output or approves a use. Unknown rights, excluded secrets and prohibited purposes stop execution; the policy and specialist determination remain authoritative.
Minimum internal route: Moderate. Other legal privacy exposure decision or action triggers can increase controls or prohibit use.
Planning aid from G25A. It never declassifies an output or approves a use. Unknown rights, excluded secrets and prohibited purposes stop execution; the policy and specialist determination remain authoritative.
Select a stage to inspect its responsibilities.
INSIDE THE DESIGN
Define business need rights owners and risk.
INSIDE THE DESIGN
Define business need rights owners and risk.
INSIDE THE DESIGN
Test the idea and interaction.
INSIDE THE DESIGN
Test technical feasibility against predefined criteria.
INSIDE THE DESIGN
Validate the real workflow with bounded exposure.
INSIDE THE DESIGN
Operate only the approved scope.
INSIDE THE DESIGN
Reassess new data models recipients tools geography or scale.
INSIDE THE DESIGN
Stop use and address dependencies.