AI project recovery: how to rescue a stalled AI project with a 10-point takeover audit
AI project recovery starts with a takeover audit, before anyone writes new code. Score the stalled build on ten points, from the success metric and evals to serving cost, security and who owns the accounts, then decide keep, fix or rebuild for each component and run a 30-60-90 day plan to production. On inherited AI systems we typically cut serving cost 40-70%, and this guide shows the levers behind it.

The short version
- Stalled AI projects usually still demo. MIT NANDA found generative AI implementation falling short for 95% of companies in its dataset, and it blames flawed integration more than model quality (reported by Fortune, 2025).
- Audit before you fix. Score ten areas red, amber or green from evidence: success metric, evals, data, code, serving cost, latency, security (OWASP Top 10 for LLM Applications 2025), governance (NIST AI RMF), MLOps and ownership.
- Decide per component. Keep what passes its evals at the planned cost, fix what lacks engineering, rebuild what rests on the wrong problem or unusable data. Anthropic says 20-50 tasks from real failures is a great first eval set.
- Own everything at handover. OMB memo M-25-22 (April 2025) names the protections: knowledge transfer, data and model portability, rights to code and models, clear licensing and pricing.
- Recovery often pays back in serving cost. On inherited AI systems we typically cut it 40-70%; cache reads at 0.1 times the input price (Anthropic) and a 50% batch discount (OpenAI) show the levers that make cuts like that possible.
7 signs your AI project has stalled
A stalled AI project usually still demos well. The warning signs sit around the model: no numeric success target, no eval suite, a cost per request nobody can state, a release date that moves every sprint, a vendor who holds the accounts, and no monitoring once real users arrive. If three or more of the seven signs below apply, stop adding features and run a takeover audit before anyone writes new code.
MIT's NANDA initiative, drawing on 150 leader interviews, a survey of 350 employees and 300 public deployments, found that for 95% of companies in its dataset generative AI implementation is falling short, and only about 5% of pilots reach rapid revenue acceleration3. The report points at flawed enterprise integration rather than model quality, so swapping the model is rarely the fix.
- It demos, but it has never carried real traffic. This is the AI prototype to production gap. Google's MLOps guidance describes the most common starting point as a manual process driven by notebook code, and notes that in practice models often break when they meet the real world5.
- Nobody can state the success target as a number. Accuracy on a named task, p95 latency and cost per request each need a written threshold.
- There is no eval suite. Anthropic describes teams without evals as "flying blind": wait for complaints, reproduce by hand, fix, and hope nothing else regressed6. Every prompt change becomes a guess.
- Cost per request is unknown or climbing. Nobody instrumented token usage per call.
- The date moves every sprint while scope grows. Features keep landing on a system that has not met its first target.
- The vendor holds the keys. The repository, the cloud account and the model provider keys sit in the vendor's name, and so does the knowledge of why things were built that way.
- You learn about failures from users. Google lists lack of active performance monitoring as a defining trait of the manual stage: predictions are not logged, so drift goes unseen5.
A skills gap usually sits underneath. In Deloitte's 2025 tech value survey of nearly 550 leaders, lack of internal technical expertise rose from 52% to 58% of respondents in a year, the largest increase of any barrier4. The root causes behind these symptoms are covered in our analysis of why AI projects fail and how often; this guide is about what to do once you are already stuck.
The 10-point takeover audit for AI project recovery
The takeover audit scores ten areas red, amber or green from evidence, not interviews: the success metric, evals, data, code, serving cost, latency and reliability, security against the OWASP Top 10 for LLM Applications, governance against the NIST AI RMF, MLOps and release, and ownership of accounts and IP. The output is a gap list and a keep, fix or rebuild verdict per area.
Run it before quoting any fix; a recovery priced before the audit is priced on the old vendor's status report. Keep it time-boxed. If the audit stalls because nobody can get access, that is your first red.
| Audit point | Evidence to ask for | Red flag |
|---|---|---|
| 1. Success metric | Written targets for accuracy, p95 latency and cost per request, tied to a business outcome | Targets exist only in slides, or not at all |
| 2. Evals | A versioned test set drawn from real failures, pass rates per release, a regression suite | No evals, or a one-off score the vendor ran once |
| 3. Data | Sources, freshness, lineage, access rights, the same feature logic in training and serving | Data prepared by hand in a notebook; training-serving skew |
| 4. Code and architecture | Repo builds from a clean checkout; prompts, configs and dependencies versioned | Critical code on one laptop or only in the vendor's repo |
| 5. Serving cost | Cost per request with a token breakdown, caching, model choice per task | Nobody can state cost per request |
| 6. Latency and reliability | p95 latency and error rate under realistic load, timeouts, fallbacks | Only tested at demo volume |
| 7. Security | Threat review against the OWASP Top 10 for LLM Applications 2025 | An agent with write access and no limits |
| 8. Governance and risk | Named owner, risk register, incident response mapped to the NIST AI RMF | No named owner for what the system says or does |
| 9. MLOps and release | CI/CD, staged rollout, rollback path, production monitoring | Manual deploys, no rollback, no logging of predictions |
| 10. Ownership and handover | Who holds the repo, cloud, API keys, data, model artifacts and licenses | The vendor owns the accounts you pay for |

What the ten points are really testing
- Points 1 and 2 decide whether anything else can be judged. Without a target and an eval set, "is it better now?" has no answer. Anthropic's guidance is that 20-50 simple tasks drawn from real failures is a great start, and that once evals exist you get latency, token usage, cost per task and error rates tracked on a fixed bank of tasks6. Build it in week one; our guide to building evals that gate releases covers the method.
- Points 3, 4 and 9 are the MLOps layer. Google defines MLOps as automation and monitoring at every step of building an ML system, and notes that only a small fraction of a real-world ML system is ML code5.
- Point 7 uses a public list. The OWASP 2025 list runs from prompt injection and sensitive information disclosure through excessive agency to unbounded consumption7. Unbounded consumption links security to cost: an endpoint with no usage limits is a budget risk as well as an attack surface.
- Point 8 uses the NIST AI RMF. Its core has four functions, Govern, Map, Measure and Manage, with governance cutting across the other three8. At minimum it asks for a named owner and a way to switch the system off. The public cases in our review of AI governance failures and their missing controls show what happens without either.
- Point 10 blocks everything else. Without the accounts there is no audit, fix or handover.
A useful rule: treat a red on point 7, 8 or 10 as a blocker whatever the other scores say, because each is a legal, security or access problem. Clear those first.
On day one, request repository access (including infrastructure code and prompt templates), the cloud and model provider consoles with recent invoices, any eval results, a sample of logs with inputs and outputs, and the contract with its license terms.
Keep, fix or rebuild: the decision matrix
Decide per component, not for the whole project. Keep what passes its evals at the planned cost, and fix what has a sound design but lacks engineering such as evals, monitoring, caching or access limits. Rebuild what rests on a wrong problem definition, unusable data or an architecture that cannot meet the latency or cost target at any tuning.
| Component | Keep when | Fix when | Rebuild when | First test |
|---|---|---|---|---|
| Problem and metric | A numeric target exists and users want the outcome | The goal is right but unmeasured | The system solves a problem nobody owns | Can the sponsor state the target? |
| Data | Sources are reliable and you hold the rights | Quality gaps, manual prep, skew | The needed data does not exist or cannot be used | Rebuild the dataset from source |
| Model and prompts | Passes the eval set at target | Close to target; prompts or retrieval need work | Misses by a wide margin after tuning | Run the eval set on two alternatives |
| Serving layer | Meets latency and cost under load | No caching, routing or cost caps yet | The design cannot hit latency or cost | Load test at expected volume |
| Evals and monitoring | Versioned, run on every change | Partial or manual | Absent (build new; there is nothing to fix) | Does a bad change fail the build? |
| Security and access | Reviewed against OWASP; least privilege | Missing limits on tools or usage | Agent design needs broad write access | List every action the system can take |
| Code and ownership | Builds cleanly, in your accounts | Undocumented but readable | Unreadable, or you cannot obtain it | Build from a clean checkout |
Expect a mixed verdict; the data pipeline might stay while the serving layer is replaced. Three rules keep it honest:
- Ignore sunk cost. The only question is the cheapest path to the target from here.
- Test before you swap models. Given MIT's finding that integration, not model quality, is the usual gap3, run the eval set on the current model with better retrieval or prompts before paying to change it. With evals in place, Anthropic notes, teams can assess a new model and upgrade in days rather than weeks6.
- Rebuild small. If the verdict is rebuild, treat it as a new proof of concept with a written success target, not a rewrite of every feature the old build promised. Our playbook on moving an AI proof of concept into production covers that path.
A 30-60-90 day recovery plan
Days 1 to 30 stabilize and measure: take control of the accounts, freeze new features, build the eval baseline and a cost dashboard, and close any security blocker. Days 31 to 60 fix the red items in order of risk, add monitoring and a rollback path, and bring serving cost down. Days 61 to 90 ship to real users in stages against the written targets, then hand over to the team that will own the system. Ninety days is also our median from kickoff to production.
| Phase | Goal | Main work | Deliverable | Exit gate |
|---|---|---|---|---|
| Days 1 to 30 | Stabilize and measure | Account transfer, feature freeze, audit, eval baseline, cost per request, security blockers | Audit report, baseline scores, verdict per component | Targets signed off; decide continue, rescope or stop |
| Days 31 to 60 | Fix | Remediate reds by risk, add monitoring and rollback, caching and cost caps, data fixes | Eval pass rates at or near target in staging | Regression suite green; load test passed |
| Days 61 to 90 | Ship and hand over | Staged rollout to real users, on-call, runbooks, paired handover | System in production, handover pack | Targets held in production; owner accepts |
Each exit gate is a real decision. The NIST AI RMF expects organizations to assign responsibility to "supersede, disengage, or deactivate" AI systems whose performance is inconsistent with intended use8. A recovery that cannot stop at day 30 is the old project with a new vendor.
- Measure before you fix. The eval baseline and cost per request from the first month are how you will prove the recovery worked. Without them, the day 90 review is opinion.
- Ship in stages. Move traffic in steps with an automatic rollback when a target is breached. NIST's MEASURE function expects the system's behavior to be monitored when in production, and MANAGE 4.1 adds incident response, recovery and change management after deployment8. Our pre-launch gates for AI releases list what to check before each step.
- Plan the handover from day one. The future owners should sit in the reviews from the first month.
How to get a clean vendor handover
A clean handover means you own everything needed to run, change and rebuild the system without the old vendor: code and prompts, cloud and model accounts, data and derived artifacts, eval sets, and the reasoning behind the design. The US government's own AI buying rules, OMB memo M-25-22 of April 2025, list the same protections: knowledge transfer, data and model portability, rights to code and models, and clear licensing and pricing.
OMB's reasoning applies to any AI vendor takeover: without those terms, switching vendors can become cost-prohibitive10. The memo also asks for documentation of development decisions, coding languages and testing scripts so that AI tools can move from one vendor to the next. Use it as your checklist:
- Code. Every repository, including infrastructure as code, prompt templates and eval harnesses, moved into your organization.
- Accounts and keys. Cloud and model accounts in your name, keys rotated.
- Data and derived assets. Raw and processed data, embeddings and vector indexes, with an agreed format. OMB's closeout guidance asks for a mutual understanding of the format and usability of data and a plan for transferring derived assets10.
- Models and licenses. Fine-tuned weights, model cards and the license for each third-party model, dataset and library. NIST's generative AI profile flags untraceable third-party components and weak supplier vetting as a distinct risk, and notes it can be hard to attribute a fault to any one component9.
- Evals. Test sets, past results and known failure cases.
- Knowledge. Architecture decision records, runbooks and a paid transition window with the engineers who built it.
Working with the outgoing vendor
Agree the handover terms before you announce the change, and pay for a short transition window rather than asking for favors. Frame the exit around the audit's red items rather than blame: a list of what failed its eval, cost or security checks is easier to agree than a dispute over fault.
Who can fix an AI project your vendor could not finish?
Look for a team that audits before it quotes, shows eval and cost numbers from systems it runs in production, names the senior engineer who will lead, and plans the handover back to you from the start. Bringing everything in-house is not automatically safer. MIT's research found that buying AI tools from specialized vendors and building partnerships succeeded about 67% of the time, while internal builds succeeded one-third as often3. Choose the partner on evidence and own the assets yourself.
What AI project recovery costs
On the ranges we publish, a takeover audit is priced like AI consulting at $100 to $450 an hour, fixing an inherited system compares with the $40k to $250k mid-complexity band, and a rebuild starts again at a proof of concept of about $75k to $150k. Mapping recovery work to those bands is ours; the audit sets the real number. Serving cost is where recovery often pays back: on inherited AI systems we typically cut it 40-70%.
| Recovery path | Closest published range | What moves it |
|---|---|---|
| Takeover audit only | AI consulting, $100 to $450 an hour | Number of systems, access, log quality |
| Fix and ship | Mid-complexity custom AI, $40k to $250k | Count of red items, integrations, eval build |
| Rebuild | Well-scoped POC, $75k to $150k over about 8 to 12 weeks; production-grade builds climb into the mid six figures | Scope held to the first target |
| Enterprise platform recovery | $500k to $1M plus | Compute and MLOps staffing |
| Ongoing operation | Retainers around $5k to $25k a month | Monitoring, on-call, retraining |
The same guide notes that the real cost drivers are data quality, integration, accuracy needs and inference volume, rarely the model, and gated work lets a bad assumption surface in a $90k phase instead of a $900k build2. The full breakdown of AI build pricing explains each band.
The levers behind serving cost cuts
Inherited systems are rarely tuned for cost. The levers are ordinary engineering, and provider price sheets show why they move the bill:
| Pricing mode | Source | Price vs standard input |
|---|---|---|
| Standard input token | Provider base rate | 100% |
| Prompt cache write, 5-minute lifetime | Anthropic docs | 125% |
| Batch API request, 24-hour turnaround | OpenAI docs | 50% |
| Prompt cache read, standard rate | Anthropic docs | 10% |
- Prompt caching. Anthropic prices a cache read at 0.1 times the base input price (its standard rate) and a 5-minute cache write at 1.25 times11, so a long system prompt or document reused across requests is paid for once at a premium and then at a tenth.
- Batch what can wait. OpenAI's Batch API costs 50% less than synchronous calls, with results within 24 hours12. Nightly classification, enrichment and eval runs rarely need an instant answer.
- Right-size the model. Route simple requests to a smaller model, then confirm on the eval set that accuracy held.
- Cap usage. Per-request token limits and rate limits close OWASP's unbounded consumption risk and put a ceiling on the bill7.
None of these is safe without evals, which is why the audit builds them first. We cover the levers that move an AI bill in more depth separately.
Evals get harder to build the longer you wait.
If your build is stuck, our AI integration and recovery team will run the audit above against your system and give you a keep, fix or rebuild verdict for each component.
AI project recovery questions
How do you rescue a failing AI project?
What should you do when an AI vendor can't deliver?
Should you keep, fix or rebuild a stalled AI project?
Who can fix an AI project our vendor couldn't finish?
How long does AI project recovery take?
How much does AI project recovery cost?
Sources
- Resourcifi, Resourcifi homepage, FAQ: Can you take over a project another vendor stalled? (on inherited AI systems we typically cut serving cost 40-70% while bringing latency and accuracy inside the locked targets; 90-day median from kickoff to first production release; 600+ projects).
- Resourcifi, AI development cost: how to scope, quote, and price an AI build (published ranges: well-scoped POC about $75k to $150k over about 8 to 12 weeks; mid-complexity custom AI $40k to $250k; enterprise AI platform $500k to $1M plus; AI consulting $100 to $450 an hour, retainers around $5k to $25k a month; production-grade builds climb into the mid six figures; cost drivers; the $90k phase versus $900k build line).
- Fortune, MIT report: 95% of generative AI pilots at companies are failing (MIT NANDA, The GenAI Divide 2025: 95% of companies in the dataset falling short; about 5% of pilots reach rapid revenue acceleration; 150 interviews, 350 employees surveyed, 300 public deployments; flawed enterprise integration rather than model quality; buying AI tools from specialized vendors and building partnerships succeed about 67% of the time, internal builds one-third as often).
- Deloitte Insights, AI is capturing the digital dollar. What's left for the rest of the tech estate? (2025 tech value survey of nearly 550 leaders; lack of internal technical expertise rose from 52% to 58%, the largest year-over-year increase).
- Google Cloud, MLOps: Continuous delivery and automation pipelines in machine learning (definition of MLOps; only a small fraction of a real-world ML system is ML code; level 0 manual, notebook-driven process; lack of active performance monitoring; training-serving skew; models often break when deployed in the real world).
- Anthropic, Demystifying evals for AI agents (20-50 simple tasks drawn from real failures is a great start; teams without evals are flying blind; evals track latency, token usage, cost per task and error rates; upgrade models in days rather than weeks; pull quote).
- OWASP Gen AI Security Project, OWASP Top 10 for LLM Applications 2025 (the 2025 risk list, including prompt injection, sensitive information disclosure, excessive agency and unbounded consumption).
- US National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (four functions Govern, Map, Measure and Manage with governance cross-cutting; MEASURE 2.4 production monitoring; MANAGE 2.4 supersede, disengage or deactivate; MANAGE 4.1 post-deployment monitoring, incident response, recovery and change management).
- US National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (value chain and component integration risk: untraceable third-party components, improper supplier vetting, difficulty attributing issues to one component).
- US Office of Management and Budget, M-25-22, Driving Efficient Acquisition of Artificial Intelligence in Government (dated April 3, 2025; vendor lock-in protections: knowledge transfer, data and model portability, rights to code and models, transparency in licensing and pricing; switching vendors could become cost-prohibitive; documentation of coding languages and testing scripts; contract closeout data format and transfer plan).
- Anthropic, Prompt caching (cache reads 0.1 times and 5-minute cache writes 1.25 times the base input token price).
- OpenAI, Batch API (50% cost discount compared to synchronous APIs; 24-hour turnaround).
Strategy, architecture & ops
AI Architecture Patterns
Agentic design patterns explained: reflection, tool use, planning, and multi-agent collaboration, with a framework to pic...
Read guide →
Strategy, architecture & ops
AI Architecture Patterns for SaaS: A Technical Guide
Generative AI architecture for SaaS: layered design, multi-tenant isolation, LLM gateway, RAG, and security. Built by Res...
Read guide →
Strategy, architecture & ops
AI Cost Optimization
A senior-engineer guide to AI cost optimization: where LLM spend comes from, the levers ranked by payoff, the five number...
Read guide →
Strategy, architecture & ops
AI Deployment Checklist: 9 Gates Before You Ship
How to deploy AI models to production: a 9-gate pre-launch checklist anchored to the OWASP LLM Top 10 (2025), NIST AI RMF...
Read guide →
Strategy, architecture & ops
AI Evaluation and Evals
LLM evaluation and AI evals, explained: the eval taxonomy, how to build an eval suite, LLM-as-a-judge bias, offline vs pr...
Read guide →
Strategy, architecture & ops
AI Features SaaS Customers Actually Want
What AI powered SaaS customers actually want: the time-savers and answers they value, the automation they distrust, and h...
Read guide →
Agents & RAG
Agentic RAG: When to Use It and How to Build It
Agentic RAG explained: how it differs from naive and advanced RAG, the key patterns like corrective RAG and self-RAG, the...
Read guide →
Agents & RAG
AI Agent for Fintech: Risk, Compliance, Ops, Customer
AI agents in finance: fraud, AML, KYC and servicing use cases, how to build with money-movement guardrails and human appr...
Read guide →
Agents & RAG
AI Agent for Healthcare: Use Cases, Governance & Implementation
AI agents in healthcare: the use cases that pay off first, how to build one HIPAA-safe on FHIR with clinician review, and...
Read guide →
