AI-first software development: what it means and how to adopt it
AI-first software development is a delivery model where AI agents do the first pass on most work and humans direct it, gate it and answer for what ships. It is more than giving developers coding assistants: 90% of technology professionals surveyed by DORA already use AI at work (2025), yet the same research links adoption to lower delivery stability. This guide defines the model, shows the team, gates and metrics it needs, and sets out a 90-day adoption plan.

The short version
- AI-first is an operating model, not a tool list. Agents do the first draft of most tickets; humans frame the work, own the review gates and answer for production. The repository, team shape and metrics change to fit.
- Adoption is near universal; agent workflows lag. 90% of nearly 5,000 technology professionals use AI at work (DORA, 2025), but 52% of developers do not use agents or stick to simpler tools (Stack Overflow, 2025).
- Speed without gates buys instability. DORA links AI adoption to higher throughput and to lower delivery stability, and in Veracode's tests with no security-specific prompting, roughly 44% of AI code generation tasks introduced a risky vulnerability (2026).
- Measure delivery against a baseline. In METR's 2025 trial, experienced developers took 19% longer with AI while believing they were 20% faster. Baseline the five DORA metrics before rollout and add cost and latency for AI features.
- Ninety days is enough to know. Baseline and policy in month one, a gated pilot in month two, measure and scale in month three. It matches our own 90-day median from kickoff to production.
What AI-first software development means
AI-first software development is a delivery model in which AI agents and assistants do the first pass on most of the work (specs, code, tests, reviews, documentation) and humans set the direction, own the review gates and answer for what reaches production. The difference from ordinary tool use is structural: the workflow, the repository, the team shape and the metrics are designed around AI doing the volume, so speed is less likely to come at the cost of stability.
The phrase gets used loosely: a chat box in a product, or a coding assistant licence for every developer. An AI-first approach in software engineering goes further. It changes who, or what, does the first draft of each task, then changes everything downstream so that draft can be trusted.
Three tests tell you whether a team is actually working AI-first:
- The default first draft is machine-made. When a ticket opens, the first attempt at the plan, the code and the tests comes from an agent working against the repository, not from a person typing into an empty file.
- The repository is written for agents as well as people. Coding standards, architecture decisions and review checklists live in files the agent reads at the start of every session. Claude Code, for example, reads a
CLAUDE.mdfile from the project root, and can also read anAGENTS.mdwritten for other coding agents11. - The gates are explicit and measured. Every AI-produced change passes the same tests, security scans and human review as any other change, and the team tracks delivery stability alongside speed. If only speed is measured, the team is still running an experiment.
The term also covers what gets built. AI features a team ships (a support agent, a retrieval system, a copilot) get the same discipline: agreed latency, cost and accuracy targets, and evaluation suites that decide what ships. That is the core of Production-First AI, the approach we use on every build.
AI-first vs AI-assisted vs AI-enabled development
AI-assisted development means individual developers use AI tools on their own terms, mostly autocomplete and chat. AI-enabled means the organization has sanctioned tools, a policy and some shared setup, but the workflow is still built around humans writing code. AI-first means the workflow itself assumes AI does the first draft, with humans directing and gating. Most teams today sit between the first two.
The adoption numbers show how far the first stage has spread. In the 2025 DORA research, 90% of nearly 5,000 technology professionals reported using AI at work1. Stack Overflow's 2025 survey found 84% of respondents using or planning to use AI tools, up from 76% a year earlier, and about half of professional developers (50.6%) using them daily4. Agents are a different story: 52% of developers either do not use agents or stick to simpler AI tools, and 38% have no plans to adopt them4. Tool use is near universal, and agent-based workflows lag well behind it.
| Dimension | AI-assisted | AI-enabled | AI-first |
|---|---|---|---|
| Who writes the first draft | The developer, with suggestions | The developer, with approved tools | An agent, directed by a human |
| Unit of AI work | A line or a function | A function or a file | A ticket, end to end, as a pull request |
| Policy | None, or informal | Written tool and data policy | Written policy plus task routing rules |
| Repository context | Whatever the developer pastes in | Some shared prompts | Standards, architecture and checklists in agent-readable files |
| Review | Normal code review | Normal code review | Tiered gates: tests, scans, evals, then a named human |
| Batch size | Unchanged | Often grows | Deliberately kept small |
| What is measured | Nothing, or seat usage | Adoption and perceived speed | DORA throughput and instability, plus cost and latency |
| Typical failure | Uneven quality between people | Faster output, same or worse stability | Over-trust in agents if gates slip |
Moving from assisted to enabled is a procurement decision. Moving from enabled to AI-first is an operating model change, and that is where most of the value and most of the risk sit.
Why businesses are moving to AI-first delivery
Three things changed at once. The tools got far more capable (Stanford's 2026 AI Index reports SWE-bench Verified scores rising from 60% to near 100% in one year), adoption became the norm, and the 2025 DORA research found a positive link between AI adoption and delivery throughput. The catch: the same research found a continued negative link with delivery stability.
SWE-bench Verified, the benchmark behind that jump9, asks a model to patch a real codebase to resolve a described issue. GitHub's Copilot cloud agent takes an assigned issue and prepares a pull request12, and Claude Code edits files, runs commands and opens pull requests from the terminal, the IDE or CI11.
Gartner predicted in July 2025 that 90% of enterprise software engineers will use AI code assistants by 2028, up from less than 14% in early 2024, and that the developer's role "will shift from implementation to orchestration"8. Gartner also expects at least 55% of software engineering teams to be building LLM-based features by 20278, which would mean most teams are shipping AI as well as using it.
AI doesn't fix a team; it amplifies what's already there.
The evidence that AI makes every developer faster is mixed. The stronger business case is that a team organized to direct and verify AI output can clear more of its backlog while keeping control of quality. Bolt tools onto an unchanged process and you get the speed and the instability together: AI "accelerates software development, but that acceleration can expose weaknesses downstream"1.
The AI-first operating model: team, workflow and QA gates
An AI-first software delivery team keeps fewer people writing code and more people framing work and checking it. The workflow runs as a loop: a senior person frames the task and sets the standard, agents generate the plan, code and tests, automated gates and a named human review every output, and the team ships in small batches and feeds what it learns back into the repository's instructions.

Team shape
Each familiar role stays, with its weight shifted toward framing and checking.
- Tech lead or architect. Owns the agent instructions, the architecture decisions agents must respect, and the rules for which tasks agents may take alone. Reviews the riskiest changes.
- Engineers. Spend more time writing specifications, splitting work into small tickets, reviewing agent pull requests and fixing the part of a change an agent gets wrong. Stack Overflow respondents' biggest single frustration, cited by 66%, is "AI solutions that are almost right, but not quite"4.
- QA and evaluation. Moves from manual test passes to owning test suites and, for AI features, evaluation sets. When agents write tests in bulk, the valuable skill is knowing which ones matter. Our guide to evals for AI systems covers how to build them.
- Platform or DevOps. Owns the paved road: CI, agent sandboxes, secrets, permissions and delivery dashboards. DORA found that 90% of organizations have adopted at least one platform and ties platform quality directly to getting value from AI1.
For how these roles combine in teams of different sizes, see our guide on how to structure an AI engineering team.
Workflow
Our own method runs the same four steps: Direct (a senior expert frames the problem and sets the bar), Generate (AI agents do the volume work under that direction), Review (a human checks every output against the standard, and nothing moves on unchecked) and Ship and learn (measure against the goal and feed it back). The detail is on how we work.
Two practices matter more than any tool choice. First, keep batches small. DORA's AI Capabilities Model names working in small batches as one of seven capabilities that amplify AI's benefit, because "AI can easily generate massive blocks of code, which are hard to review and test"2. Second, invest in context. DORA calls it AI-accessible internal data, or context engineering: connecting tools securely to internal documentation and codebases2. Every correction a reviewer makes twice should become a line in the agent's instruction file.
QA gates
Every agent-produced change passes the same gates in the same order, with nothing skipped because it looks fine.
| Gate | What it checks | Who or what owns it |
|---|---|---|
| 1. Spec check | The ticket states the goal, constraints and acceptance tests before an agent starts | Tech lead or engineer |
| 2. Build and tests | Compiles, existing tests pass, new tests cover the change | CI, automatic |
| 3. Static and security scans | Lint, dependency and secret scanning, SAST for injection and XSS patterns | CI, automatic |
| 4. Evals (AI features only) | Accuracy, latency and cost against agreed thresholds on a fixed test set | QA or evaluation owner |
| 5. Human review | Design fit, edge cases, readability, whether the change should exist | A named engineer, never the agent |
| 6. Staged release | Canary or feature flag, with a tested rollback | Platform or DevOps |
Mind the repository rules. GitHub notes that a rule allowing only specific commit authors can stop its cloud agent from opening pull requests, and that administrators can add Copilot as a bypass actor12. Decide deliberately what agents may bypass. For merges to main, nothing.
Metrics that prove AI-first software delivery works
Measure AI-first delivery with the five DORA software delivery metrics, split into throughput (change lead time, deployment frequency, failed deployment recovery time) and instability (change fail rate, deployment rework rate). Add the cost and latency of any AI features you ship. Take a baseline before the rollout, because developers' own sense of speed is unreliable.
The perception problem is well documented. In METR's randomized trial, 16 experienced open-source developers completed 246 tasks. Before starting they expected AI to speed them up by 24%. With AI allowed they took 19% longer, and afterwards they still believed AI had sped them up by 20%5. METR says this does not show AI fails to speed up most developers, and its February 2026 update judges developers likely more sped up in early 2026 than in early 2025, though METR calls its new data only very weak evidence of how large the change is6. The measurement lesson stands: self-reports are not delivery data.
| Metric | Type | Definition | What to watch after AI-first |
|---|---|---|---|
| Change lead time | Throughput (DORA) | Commit to running in production | Should fall; if not, review is the bottleneck |
| Deployment frequency | Throughput (DORA) | Deployments per period | Should rise with smaller batches |
| Failed deployment recovery time | Throughput (DORA) | Time to recover from a failed deployment | Must not rise; tests rollback discipline |
| Change fail rate | Instability (DORA) | Share of deployments needing immediate intervention | The first place over-trust shows up |
| Deployment rework rate | Instability (DORA) | Unplanned deployments caused by production incidents | Rising rework means gates are too loose |
| p95 latency | AI feature | Response time at the 95th percentile | Agreed before build, checked by evals |
| Cost per call | AI feature | Model and infrastructure cost per request | Agreed before build, tracked in production |
| Accuracy floor | AI feature | Minimum score on a fixed evaluation set | A release that drops below it does not ship |
The five DORA definitions come from DORA's own guide3. The last three rows come from the five numbers we lock before any AI build under Production-First AI; the other two are serving throughput and recovery time.
Do not count lines of code or the share of code written by AI; both reward the batch sizes that hurt stability. Read throughput and instability together. A team that ships twice as often and breaks production twice as often has not improved.
Risks and guardrails
The main risks of AI-first software development are insecure generated code, over-trust in output that is almost right, instability from large AI-sized changes, and, once agents have tool access, prompt injection and excessive agency. Each has a known guardrail: automated security scanning, a named human reviewer, small batches with tested rollback, and least-privilege permissions for every agent.
| Survey answer | Source | Share |
|---|---|---|
| Use or plan to use AI tools in development | Stack Overflow, 2025 | 84% |
| Biggest frustration: AI solutions that are almost right | Stack Overflow, 2025 | 66% |
| Distrust the accuracy of AI tools | Stack Overflow, 2025 | 46% |
| Trust the accuracy of AI tools | Stack Overflow, 2025 | 33% |
| Highly trust the output | Stack Overflow, 2025 | 3% |
The security data shows why that distrust is healthy. In Veracode's 2026 tests, with no security-specific prompting, AI-generated code had an average security pass rate of 56%, meaning roughly 44% of generation tasks introduced a risky vulnerability714. Across more than 100 models tracked since testing began, that average has barely moved from 55% in the first report, and Java had the lowest mean pass rate, at 30%1314. Generated code needs at least the same scanning as human code.
| Risk | Evidence | Guardrail |
|---|---|---|
| Insecure generated code | About 44% of tasks introduced a vulnerability (Veracode, 2026) | SAST, dependency and secret scanning on every pull request |
| Almost-right output accepted | 66% cite it as their top frustration (Stack Overflow, 2025) | Acceptance tests written before generation; named human reviewer |
| Instability from large changes | Negative link between AI adoption and stability (DORA, 2025) | Small batches, feature flags, tested rollback |
| Prompt injection | LLM01 in the OWASP Top 10 for LLM Applications 2025 | Treat repository, issue and web content as untrusted input to agents |
| Excessive agency | LLM06 in the OWASP Top 10 for LLM Applications 2025 | Least-privilege tokens; no agent merges or deploys alone |
| Data leaking into prompts | LLM02, sensitive information disclosure (OWASP, 2025) | Written AI policy on which data and repositories tools may see |
An agent that reads issues, web pages and pull request comments can be steered by text planted in any of them, which is why OWASP ranks prompt injection first and lists excessive agency and unbounded consumption in the same Top 1010. Give each agent minimal permissions, log its actions and keep production credentials out of reach. Before an AI feature ships to users, run it through a structured AI deployment checklist.
Write the policy down, too. The first capability in DORA's model is a clear and communicated AI stance, because "ambiguity creates risk"2.
A 90-day plan to adopt AI-first software development
Adopt AI-first software development in three 30-day phases: set the baseline, policy and repository context; pilot agent-first delivery on one team with full QA gates; then scale what the data supports and retire what it does not. Ninety days is enough to see whether change lead time falls without change fail rate rising, and it matches the median time our own projects take from kickoff to production.
Days 1 to 30: baseline and foundations
- Pull three months of the five DORA metrics for the pilot team as the baseline.
- Write the AI policy: approved tools (for example GitHub Copilot, Claude Code or Cursor), which repositories and data they may see, and which tasks agents may take without a human pairing.
- Create the agent instruction files in the pilot repository: coding standards, architecture decisions, test commands and the review checklist.
- Wire the gates: tests, security scans and branch rules that stop any agent from merging to main.
Days 31 to 60: pilot one team
- Route small, well-specified tickets to agents first: tests for untested code, bug fixes with a clear reproduction, dependency updates, internal tools.
- Keep every change small enough to review in one sitting. Repeated review corrections go into the instruction files weekly.
- If the team ships AI features, agree the latency, cost and accuracy targets and build the evaluation set before any feature work starts.
Days 61 to 90: measure and scale
- Compare the pilot's DORA numbers with the baseline. Scale only if throughput improved and change fail rate and rework rate held.
- Widen the routing rules to larger tickets where the data supports it, and add a second team.
- Publish the playbook internally: policy, instruction file template, gates and scorecard.
If you run the pilot with a partner, check the record. Ours is below, and our guide to vetting an AI-ready development team lists the questions to ask any vendor.
| Measure | Figure | What it tells you |
|---|---|---|
| Median time, kickoff to production | 90 days | Typical time to first production release |
| Projects delivered since 2017 | 600+ | Breadth of delivery experience |
| Repeat clients | 95% | Client retention |
| Team | 200+ in-house experts | Engineering teams in Noida, India; based in Wilmington, Delaware |
| Serving cost on inherited AI systems | Typically cut 40 to 70% | Typical result on AI systems we take over |
| Client rating | 4.9 on Clutch | Independent client reviews |
An AI-first software engineering transformation is mostly a management change. If you want the same 90-day structure applied to your own backlog, our custom software development team runs it human-led and AI-driven, with the gates and the scorecard above from the first week.
AI-first development questions
What is AI-first software development?
How is AI-first different from using AI coding tools?
What does an AI-first delivery team look like?
How do you measure AI-first delivery?
What does an AI-first software company actually mean?
How do we move our dev team to an AI-first delivery model?
Is AI-generated code safe to ship?
Sources
- Google Cloud (DORA), Announcing the 2025 DORA Report (90% of respondents use AI at work; more than 80% perceive higher productivity; 30% little or no trust; nearly 5,000 respondents; AI as amplifier quote; positive link to throughput, negative link to stability; 90% of organizations have at least one platform).
- Google Cloud (DORA), From adoption to impact: Putting the DORA AI Capabilities Model to work (Seven DORA AI capabilities, including clear and communicated AI stance, AI-accessible internal data (context engineering) and working in small batches; quotes on large AI code blocks and ambiguity creating risk).
- DORA, DORA's software delivery performance metrics (Definitions of change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate, grouped into throughput and instability).
- Stack Overflow, AI | 2025 Stack Overflow Developer Survey (84% use or plan to use AI (76% prior year); 50.6% of professional developers use AI daily; 46% distrust vs 33% trust accuracy, 3% highly trust; 66% almost-right frustration; 52% do not use agents or use simpler tools, 38% no plans).
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (16 developers, 246 tasks; 19% longer with AI; expected 24% speedup, believed 20% speedup afterwards; caveats on generalization).
- METR, We are Changing our Developer Productivity Experiment Design (February 2026 update: selection effects in the new study; METR judges developers likely more sped up in early 2026 than early 2025, while calling its new data only weak evidence of the size of the change).
- Veracode, 2026 GenAI Code Security Report: AI Is Writing More of Your Code but Security Hasn't Caught Up (Average security pass rate of 56%, barely changed from 55% in the first report; roughly 44% of AI code generation tasks introduced a risky security vulnerability in tests (28-Jul-2026)).
- DEVOPSdigest (reporting Gartner), Gartner: Top Strategic Trends in Software Engineering for 2025 and Beyond (Gartner: 90% of enterprise software engineers will use AI code assistants by 2028, up from less than 14% in early 2024; developer role shifts from implementation to orchestration; at least 55% of teams building LLM-based features by 2027).
- Stanford HAI, The 2026 AI Index Report (SWE-bench Verified performance rose from 60% to near 100% in a single year).
- OWASP Gen AI Security Project, OWASP Top 10 for LLM Applications 2025 (LLM01 prompt injection, LLM02 sensitive information disclosure, LLM06 excessive agency, LLM10 unbounded consumption).
- Anthropic, Overview, Claude Code Docs (Claude Code reads the codebase, edits files, runs commands, opens pull requests and runs in CI; CLAUDE.md read every session; can read AGENTS.md).
- GitHub, About GitHub Copilot cloud agent (Issues can be assigned to Copilot cloud agent, which works in the background and prepares a pull request; commit-author rules can block it; admins can add Copilot as a bypass actor).
- Veracode, 2026 GenAI Code Security Report: AI Is Writing More of Your Code but Security Hasn't Caught Up (Primary research: average security pass rate of 56% across more than 100 models tested over four years, which has not moved; Java last with a mean security pass rate of 30%).
- The Next Web (Ana Maria Constantin), AI writes half our code now. It still fails security tests 44% of the time. (Editorial coverage (28-Jul-2026) of Veracode's 2026 report: 56% average pass rate against 55% at the start; nearly 44% security failure given no security-specific prompting; Java 30%; more than 100 models tested).
Strategy, architecture & ops
AI Architecture Patterns
Agentic design patterns explained: reflection, tool use, planning, and multi-agent collaboration, with a framework to pic...
Read guide →
Strategy, architecture & ops
AI Architecture Patterns for SaaS: A Technical Guide
Generative AI architecture for SaaS: layered design, multi-tenant isolation, LLM gateway, RAG, and security. Built by Res...
Read guide →
Strategy, architecture & ops
AI Cost Optimization
A senior-engineer guide to AI cost optimization: where LLM spend comes from, the levers ranked by payoff, the five number...
Read guide →
Strategy, architecture & ops
AI Deployment Checklist: 9 Gates Before You Ship
How to deploy AI models to production: a 9-gate pre-launch checklist anchored to the OWASP LLM Top 10 (2025), NIST AI RMF...
Read guide →
Strategy, architecture & ops
AI Evaluation and Evals
LLM evaluation and AI evals, explained: the eval taxonomy, how to build an eval suite, LLM-as-a-judge bias, offline vs pr...
Read guide →
Strategy, architecture & ops
AI Features SaaS Customers Actually Want
What AI powered SaaS customers actually want: the time-savers and answers they value, the automation they distrust, and h...
Read guide →
Agents & RAG
Agentic RAG: When to Use It and How to Build It
Agentic RAG explained: how it differs from naive and advanced RAG, the key patterns like corrective RAG and self-RAG, the...
Read guide →
Agents & RAG
AI Agent for Fintech: Risk, Compliance, Ops, Customer
AI agents in finance: fraud, AML, KYC and servicing use cases, how to build with money-movement guardrails and human appr...
Read guide →
Agents & RAG
AI Agent for Healthcare: Use Cases, Governance & Implementation
AI agents in healthcare: the use cases that pay off first, how to build one HIPAA-safe on FHIR with clinician review, and...
Read guide →
