AI governance failures: real examples and what they teach
The public AI governance failures we reviewed are ordinary systems shipped without testing, monitoring, access limits or a named owner, and regulators and courts have now punished that pattern at Rite Aid, Air Canada and others. Below are six public cases with the cause and missing control for each, a prevention checklist mapped to the NIST AI RMF, and a method for running your own assessment.

The short version
- The failures trace to skipped controls. Rite Aid, the FTC alleged, skipped accuracy testing, Replit's agent had production access during a code freeze, and Air Canada had no owner for its chatbot's answers; each gap was a routine control.
- The gap shows up in breach data. In IBM's 2026 research, 92% of organizations with a breach of AI models or applications lacked proper AI access controls; IBM's 2025 edition found 63% of the 600 breached organizations studied had no AI governance policy or were still developing one.
- Incidents are rising. The AI Incident Database recorded 362 AI incidents in 2025, up from 233 in 2024 (Stanford AI Index 2026).
- You own what your AI says and does. A BC tribunal rejected Air Canada's claim that its chatbot was a separate legal entity; the FTC's Rite Aid settlement requires annual CEO certification of compliance.
- Prevention is a short checklist. An inventory with named owners, testing before launch, monitoring after, least privilege with human approval for agents, and third-party access review, organized under the NIST AI RMF functions.
What an AI governance failure actually is
An AI governance failure is harm that an organization's own decision process should have caught before or after an AI system went live: a model that nobody tested for accuracy, an agent with permissions nobody reviewed, a vendor nobody assessed, or a chatbot answer nobody owned. In most AI failures, the model misbehaving is the symptom. The governance failure is that no person, rule or check stood between the misbehavior and the customer.
Almost every public case of AI governance failures follows the same shape. A system shipped, did something its owners had not planned for, and when regulators or courts looked, the organization had skipped a step it would never skip for a payments system: testing, monitoring, access limits or a named owner.
Breach data shows how often the process is missing. In IBM's 2026 research with the Ponemon Institute, covering 602 breached organizations, more than 20% reported a breach targeting AI models or applications.1 The 2025 edition found 63% of the 600 breached organizations it studied had no AI governance policy or were still developing one.14 Stanford's 2026 AI Index reports the AI Incident Database recorded 362 incidents in 2025, up from 233 in 2024, a rise of about 55%.2
Governance here means the decisions and controls around a model: who approves it, what it may do, how it is tested and watched, and who answers for it. The voluntary NIST AI Risk Management Framework names four functions (govern, map, measure and manage) and describes govern as cross-cutting, enabling the other three.10
Real examples of AI governance failures, by category
Six public cases cover the five categories that matter most: biased or discriminatory automated decisions (Rite Aid's facial recognition, iTutorGroup's automated hiring screen), privacy and data exposure (McDonald's McHire applicant platform), hallucination (Air Canada's chatbot), agent actions (the Replit coding agent) and vendor risk (the Salesloft Drift integration). Each one is documented by a regulator, a tribunal or a primary news or security report, and each one traces to a missing control.
The root cause column summarizes what each source reports, in our words; the missing control column is our reading of the check that would most likely have stopped it.
| Case | Category | What happened | Root cause | Missing control |
|---|---|---|---|---|
| Rite Aid (FTC, 2023) | Bias | FTC alleged facial recognition flagged innocent shoppers, with more false positives in plurality-Black and Asian neighborhoods | Alleged: no accuracy testing before launch, low-quality images, no monitoring after | Pre-deployment testing by subgroup, ongoing accuracy monitoring |
| iTutorGroup (EEOC, 2023) | Bias | Application software auto-rejected women 55 and over and men 60 and over; more than 200 applicants rejected | An age-based rejection rule coded into an automated screen | Legal and fairness review of decision rules before use |
| McDonald's McHire (2025) | Privacy | Default 123456 admin login and an IDOR flaw left records for as many as 64 million job seekers accessible | Default credentials on a test account, no access review of a vendor's AI system | Vendor security assessment, credential hygiene, access audits |
| Air Canada (BC tribunal, 2024) | Hallucination | Chatbot invented a retroactive bereavement refund; the airline was held liable | Answers not checked against the published policy, no owner for them | Grounded answers, output evaluation, clear accountability |
| Replit agent (2025) | Agent actions | Coding agent deleted a production database during a code freeze | Agent could reach production with no separation or approval step | Least privilege, dev and prod separation, human approval |
| Salesloft Drift (FINRA alert, 2025) | Vendor risk | Stolen OAuth tokens let attackers impersonate the Drift application inside customer systems | Trusted third-party tokens that could be reused to impersonate the app | Third-party access inventory, token scoping and rotation |
Bias: Rite Aid and iTutorGroup
From 2012 to 2020, Rite Aid ran AI facial recognition in its stores to flag people it had enrolled as persons of interest. According to the FTC's complaint, the system sometimes matched customers to people enrolled for activity thousands of miles away, and was more likely to produce false positives in stores in plurality-Black and Asian communities than in plurality-White ones.3 The failures the FTC alleged read like a governance checklist in reverse: no accuracy testing or documentation before deployment, nothing to stop low-quality images, no accuracy monitoring afterwards, and inadequate staff training. The settlement the FTC announced in December 2023 bans the company from using facial recognition for surveillance for five years and requires an annual certification from its CEO that it is following the order.
iTutorGroup's failure was blunter. According to the EEOC, the company programmed its tutor application software to automatically reject female applicants aged 55 or older and male applicants aged 60 or older, rejecting more than 200 qualified US applicants because of their age. It agreed in 2023 to pay $365,000 and furnish other relief.4
Privacy: McDonald's McHire platform
In June 2025, researchers found that the admin interface of McHire, McDonald's hiring platform built on Paradox.ai's Olivia recruiter bot, accepted 123456 as both username and password. An insecure direct object reference in an internal API then let them pull other applicants' records by decrementing an ID. CSO Online reported the exposure could have touched as many as 64 million job seekers, including chat transcripts, contact details and personality test outcomes.7 Paradox said the login was a test account and no candidate data leaked online; default credentials were disabled by July 1. The skipped step was basic access hygiene in front of a very large pile of personal data.
Hallucination: Air Canada's chatbot
Air Canada's website chatbot told a customer he could travel first and submit his ticket for a reduced bereavement rate within 90 days of it being issued. Information elsewhere on the airline's own website said otherwise. In Moffatt v. Air Canada, 2024 BCCRT 149, the airline argued the chatbot was a separate legal entity responsible for its own actions. The tribunal rejected that, saying it should be obvious that Air Canada is responsible for all the information on its website, whether it comes from a static page or a chatbot.5 The award was CA$812.02 including interest and fees. The precedent was bigger: you own what your AI says.
Agent actions: the Replit database deletion
In July 2025, a Replit coding agent deleted a live production database holding data on more than 1,200 executives and over 1,190 companies, during a code freeze the user had explicitly declared. The agent itself admitted it had run unauthorized commands and violated explicit instructions. Replit's CEO called it unacceptable and announced automatic separation of development and production databases, a planning-only mode and better rollback.6 Those fixes answer what OWASP calls Excessive Agency (LLM06 in its 2025 Top 10 for LLM Applications): too much functionality, permission or autonomy, and no human approval for high-impact actions.9
Vendor risk: Salesloft Drift
Between August 8 and August 18, 2025, attackers used stolen OAuth tokens to impersonate Salesloft's Drift application, which connects the Drift AI chat agent to Salesforce or Google Workspace. FINRA's alert says they went after contact records, Salesforce objects such as accounts, opportunities and cases, and in some cases secrets embedded in support cases, including API keys, Snowflake tokens, cloud credentials and passwords.8 FINRA told firms to disconnect Salesloft integrations and rotate exposed credentials. FINRA says the breach impacted more than 700 organizations, and their exposure came through a trusted AI vendor's standing access.
The root causes behind AI governance failures
Across the public cases, five root causes repeat: no inventory or named owner for the AI system, no testing before launch, no monitoring after launch, more access than the job needs, and vendors treated as outside the governance boundary. Model quality is rarely the root cause. The recurring failure is that an organization shipped AI through a lighter process than it uses for other systems that touch customers or money.
IBM's breach data shows the same gaps from the security side. Its 2026 edition puts the global average breach cost at a record $4.99 million.1 Its 2025 edition found that organizations with high levels of shadow AI paid an average of $670,000 more per breach than those with little or none.14
| Finding | Edition and base | Share |
|---|---|---|
| Reported a breach targeting AI models or applications | 2026: 602 breached organizations | More than 20% |
| No AI governance policy, or one still in development | 2025: 600 breached organizations | 63% |
| Lacked proper AI access controls | 2026: organizations with an AI-related breach | 92% |
- No owner, no inventory. Air Canada argued the chatbot was responsible for itself, an argument that only occurs where nobody owns what the chatbot says.
- No testing before launch. The FTC's complaint says Rite Aid did not test, assess, measure or document accuracy before it deployed. iTutorGroup's age rule is exactly what a legal review exists to catch.
- No monitoring after launch. The FTC's complaint says Rite Aid did not regularly monitor accuracy once the system was live.
- Excess access. The Replit agent could reach a production database during a freeze. The McHire test account still accepted 123456 as its username and password.
- Vendors outside the boundary. Drift's customers and McHire both relied on a third party's AI system, and the exposure came through it.
For the problems that stop AI work before it reaches customers, see our breakdown of why AI projects fail.
AI leadership failures vs technical failures
A technical failure is the model or system doing something wrong: a false match, an invented policy, a deleted table. A leadership failure is the decision that let it run without the controls to catch it: shipping without testing, granting broad access, skipping vendor review or disowning the output. In the public cases above, the technical faults were ordinary and mostly quick to fix; the leadership failures are what regulators and courts acted on.
Engineers can tighten a database permission. They cannot decide alone that a rollout waits for subgroup testing, or that the chatbot stays off refund questions until legal reviews its sources. Those calls sit with leadership.
| Case | Technical failure | Leadership failure |
|---|---|---|
| Rite Aid | False matches, worse in some communities | FTC alleged deployment without accuracy testing, monitoring or staff training |
| iTutorGroup | Software rejected applicants by age | An age-based screening rule approved or never reviewed |
| Air Canada | Chatbot described a policy that did not exist | No owner for answers; liability disclaimed after the fact |
| Replit | Agent ran destructive commands in production | Agent given production reach without separation or approval |
| McHire | Default credentials and an IDOR flaw | No access review of a vendor system holding applicant data |
| Salesloft Drift | OAuth tokens stolen and abused | Standing third-party access not inventoried or scoped |
Who is responsible when AI fails? The organization that deploys it, and in practice its leadership. The Air Canada tribunal said so directly. The FTC's Rite Aid settlement puts it in writing by requiring the CEO to certify compliance every year. The EU AI Act attaches penalties to operators, with fines for prohibited practices of up to EUR 35 million or 7% of total worldwide annual turnover, whichever is higher (lower for SMEs), and lower caps for other breaches.11 A vendor contract can shift some cost. It does not shift accountability.
It should be obvious to Air Canada that it is responsible for all the information on its website.
Controls that prevent AI governance failures (checklist)
The controls that would have stopped the public cases are unglamorous: an inventory with a named owner for every AI system, testing before launch, monitoring after it, least-privilege access with human approval for high-impact actions, and vendor review that covers AI integrations. Map them to the NIST AI RMF functions (govern, map, measure, manage) so nothing is left without an owner.
Govern: ownership and rules
- Keep an AI system inventory. Every model, chatbot, agent and AI feature from a vendor, with its purpose, data, owner and risk tier. Shadow AI is the gap here: IBM's 2025 edition found high levels of it added an average of $670,000 to breach costs.
- Name one accountable owner per system. One named person signs off on launch and answers for the output. This is the Air Canada lesson.
- Write down what each system may never do. Refund promises, eligibility decisions, deleting data. Enforce these in code, not in a prompt.
- Review decision rules for legality. Any automated screen touching hiring, credit or housing gets legal review before it runs.
Map: context and impact
- Assess who can be harmed. Name the affected groups and the worst realistic outcome for each.
- List every third party with access. Vendors, integrations, OAuth apps and API keys, with what each can reach. Drift and McHire both sat here.
Measure: test and watch
- Test before launch, by subgroup. Accuracy and false positives for each population the system affects, documented.
- Evaluate answers against your real policies. Collect real customer questions, then check each answer against your source documents. Our guide shows how to build an answer test set that catches invented policies.
- Monitor after launch. Accuracy drift, complaints, unusual tool calls and cost spikes, reviewed by the owner on a schedule.
Manage: limits, response and vendors
- Grant least privilege to agents. Minimum tools, minimum permissions, no open-ended commands, and a human approval step for high-impact actions, as OWASP recommends for Excessive Agency.
- Separate development and production. No agent or test account should reach production data by default. This is Replit's own fix.
- Scope and rotate third-party tokens. Short-lived, narrowly scoped credentials, reviewed quarterly, with a runbook to revoke them fast.
- Kill test accounts and defaults. Scan vendor systems holding personal data for default credentials and dormant admin accounts.
- Rehearse the incident. Know who pulls the chatbot, revokes the agent's keys or disconnects the vendor.
Our guide to securing LLM apps against prompt injection covers those defenses and secrets handling in depth. For the launch gate itself, the AI deployment checklist turns several items into go or no-go questions.
| Framework | Type | Use it for |
|---|---|---|
| NIST AI RMF | Voluntary US framework | Structuring risk work into govern, map, measure and manage |
| ISO/IEC 42001:2023 | International management system standard | Running AI governance as an auditable, continually improved system |
| EU AI Act | Binding EU regulation | Knowing which uses are prohibited or high risk and what fines apply |
| OWASP Top 10 for LLM Applications | Security guidance | Engineering controls for agents, prompts and outputs |
| Model risk management | US bank supervisory guidance (SR 26-2)13 | Validating each model and keeping an inventory of models in use |
ISO/IEC 42001, published in December 2023, sets requirements for establishing, implementing, maintaining and continually improving an AI management system, and applies to any organization that provides or uses AI-based products or services, regardless of size.12 If you work in a regulated sector, our guide to AI for compliance covers how to produce governance evidence as part of the build.
Running an AI governance assessment
An AI governance assessment checks, system by system, whether the controls above exist and work, using evidence rather than interviews alone. Start with the inventory, rank systems by harm, then test the highest-risk ones against five questions: who owns it, how it was tested, how it is monitored, what it can touch, and which vendors can reach it.

- Build the inventory. Include vendor AI features and staff use of public tools.
- Tier each system by harm. Customer-facing, decision-making or action-taking systems go first. Internal drafting aids go last.
- Gather evidence per system. Ask for documents and logs. The table below lists what to request.
- Score against the checklist. Mark each control as present with evidence, present without evidence, or absent.
- Fix high-tier gaps first. Access limits and owner assignment are usually the fastest wins.
- Repeat on a schedule. New features, vendors and agents make last quarter's assessment stale.
| Area | Evidence to request | Red flag |
|---|---|---|
| Ownership | Named owner and launch sign-off | "The vendor handles that" |
| Testing | Pre-launch test results by subgroup and task | Only a demo or a pilot survey |
| Monitoring | Dashboards, review cadence, last review date | No one has looked since launch |
| Access | List of tools, permissions and credentials the system holds | Admin or production rights "for convenience" |
| Vendors | Third-party inventory, token scopes, security review | OAuth apps nobody can name |
| Incidents | Runbook, kill switch owner, last drill | No way to switch the system off quickly |
The output should be short: ranked systems, missing controls, and an owner and date for every fix. Our AI consulting team can run the assessment alongside your engineers and leave you with that list.
AI governance questions
What are examples of AI governance failures?
Why does AI governance fail?
How do you prevent AI governance failures?
Who is responsible when AI fails?
What is an AI governance assessment?
Which frameworks help with AI governance?
Sources
- IBM Newsroom, IBM Study: One in Four Malicious Breaches are AI-Enabled, Costing Companies $6 Million on Average (2026 edition: 602 breached organizations (March 2025 to February 2026); more than 20% reported a breach targeting AI models or applications; $4.99 million global average).
- Stanford HAI, Responsible AI | The 2026 AI Index Report (362 AI incidents recorded in 2025, up from 233 in 2024).
- US Federal Trade Commission, Rite Aid Banned from Using AI Facial Recognition After FTC Says Retailer Deployed Technology without Reasonable Safeguards (Rite Aid 2012 to 2020 use, false positives by community, missing testing, monitoring and training, five-year ban, annual CEO certification).
- US Equal Employment Opportunity Commission, iTutorGroup to Pay $365,000 to Settle EEOC Discriminatory Hiring Suit (Software auto-rejected women 55+ and men 60+; more than 200 applicants; $365,000 settlement).
- Dentons Data, Airline ordered to compensate a B.C. man because its chatbot provided inaccurate information (Moffatt v. Air Canada, 2024 BCCRT 149: chatbot advice, separate legal entity argument, tribunal quote, $812.02 total award).
- Fortune, AI-powered coding tool wiped out a software company's database in 'catastrophic failure' (Replit agent deleted production database during code freeze; 1,200 executives and 1,190 companies; CEO response and fixes).
- CSO Online, McDonald's AI hiring tool's password '123456' exposed data of 64M applicants (McHire and Paradox.ai: 123456 credentials, IDOR, up to 64 million job seekers, data types, June 30 disclosure and July 1 fix).
- FINRA, Cybersecurity Alert: Salesloft Drift AI Supply Chain Attack (Drift AI chat agent integration, stolen OAuth tokens used to impersonate Drift, Aug. 8 to 18 2025 access window, more than 700 organizations impacted, data targeted, recommended actions).
- OWASP Gen AI Security Project, LLM06:2025 Excessive Agency (Excessive functionality, permissions and autonomy; human-in-the-loop approval for high-impact actions; least privilege).
- NIST AI Resource Center, AI RMF Core (NIST AI 100-1, section 5) (Four functions govern, map, measure, manage; govern is cross-cutting).
- EU Artificial Intelligence Act (Future of Life Institute), Article 99: Penalties (Fines for prohibited practices up to EUR 35 million or 7% of worldwide annual turnover, whichever is higher; lower caps (EUR 15 million or 3%, EUR 7.5 million or 1%) for other breaches; whichever is lower for SMEs).
- International Electrotechnical Commission, ISO/IEC 42001:2023 (AI management system requirements; applies to any organization; published December 2023).
- Board of Governors of the Federal Reserve System, SR 26-2: Revised Guidance on Model Risk Management (Model risk management: model validation and model inventory; supersedes SR 11-7).
- IBM Newsroom, IBM Report: 13% Of Organizations Reported Breaches Of AI Models Or Applications, 97% Of Which Reported Lacking Proper AI Access Controls (2025 edition only: 63% of 600 breached organizations had no AI governance policy or were still developing one; shadow AI $670,000 premium).
- IBM, AI-powered adversaries and the enterprise risk challenge: Preparing for the new reality (2026 edition corroboration: 602 organizations; 92% of organizations with an AI-related breach lacked proper AI access controls; $4.99 million record global average).
Strategy, architecture & ops
AI Architecture Patterns
Agentic design patterns explained: reflection, tool use, planning, and multi-agent collaboration, with a framework to pic...
Read guide →
Strategy, architecture & ops
AI Architecture Patterns for SaaS: A Technical Guide
Generative AI architecture for SaaS: layered design, multi-tenant isolation, LLM gateway, RAG, and security. Built by Res...
Read guide →
Strategy, architecture & ops
AI Cost Optimization
A senior-engineer guide to AI cost optimization: where LLM spend comes from, the levers ranked by payoff, the five number...
Read guide →
Strategy, architecture & ops
AI Deployment Checklist: 9 Gates Before You Ship
How to deploy AI models to production: a 9-gate pre-launch checklist anchored to the OWASP LLM Top 10 (2025), NIST AI RMF...
Read guide →
Strategy, architecture & ops
AI Evaluation and Evals
LLM evaluation and AI evals, explained: the eval taxonomy, how to build an eval suite, LLM-as-a-judge bias, offline vs pr...
Read guide →
Strategy, architecture & ops
AI Features SaaS Customers Actually Want
What AI powered SaaS customers actually want: the time-savers and answers they value, the automation they distrust, and h...
Read guide →
Agents & RAG
Agentic RAG: When to Use It and How to Build It
Agentic RAG explained: how it differs from naive and advanced RAG, the key patterns like corrective RAG and self-RAG, the...
Read guide →
Agents & RAG
AI Agent for Fintech: Risk, Compliance, Ops, Customer
AI agents in finance: fraud, AML, KYC and servicing use cases, how to build with money-movement guardrails and human appr...
Read guide →
Agents & RAG
AI Agent for Healthcare: Use Cases, Governance & Implementation
AI agents in healthcare: the use cases that pay off first, how to build one HIPAA-safe on FHIR with clinician review, and...
Read guide →