The Platform PM

Field GuideLast reviewed October 2026

AI coding agents: writing, reviewing and shipping code

How AI coding agents take a task from an issue to a merged change: sandboxes, tests and review, what seats and usage cost, who holds the power, the evidence on speed and security, and who pays when an agent breaks something.

The industry on one page

The parties. A developer hands a task to a coding agent, which calls a lab's model, edits your repo and opens a pull request. Tests and reviewers decide what ships, and that's where the work now piles up. Errors and traces from production can flow back to the agent.

Picture a company that sells invoicing and payments software to other businesses, with 200 engineers. A developer there hands a task to a coding agent: upgrade the PDF library and fix whatever breaks. The agent might sit in the developer's editor, run in a terminal or work in the cloud from a GitHub issue. Wherever it runs, it sends the task and the code it reads to a model from a lab such as Anthropic, OpenAI or Google, and pays in tokens (the chunks of text a model reads and writes). It edits the company's repo, where the code, the tests and the rules files written for agents live, and opens a pull request, a proposed change that people review before it's merged. Tests and reviewers then decide what ships, and that's where the work now piles up. After the deploy, errors and traces from production can flow back to the agent as its next task.

Writing code got cheap quickly: in July 2026 Microsoft said one in three pull requests on GitHub involves an agent. Checking that code hasn't got cheaper at the same rate.

What I'd want a new PM in this space to take away:

  1. Review: review and CI are the new bottleneck. Agent pull requests merge less often than people's, and the gap isn't closing. A July 2026 study of about 9,400 agent pull requests in 489 open-source Python repos found that, adjusted for repo and task, about 65% merged against about 85% for human ones. Reviewers often abandon them, and the ones that don't merge are bigger and fail CI (continuous integration, the automated builds and tests every change must pass) more often. Amazon reportedly now requires senior sign-off on AI-assisted changes by junior and mid-level engineers after a run of outages, though it disputes the link to AI. Nobody has a causal 2026 study of agents inside companies: METR paused its trial design in February 2026, and DORA, Google's long-running research program on software delivery, had published no 2026 AI report by October 9. See Do coding agents make teams faster: what's real.
  2. Environment: the model sets the ceiling, and the environment decides how much of it you reach: a repo the agent can build, tests it can trust and run quickly, and permissions scoped to the task. Executable checks do more than written instructions: in a February 2026 study, AGENTS.md files (instructions for agents kept in the repo) gave no significant gain in task success and raised inference cost by about a fifth. Public benchmarks are breaking down too: OpenAI dropped SWE-bench Verified in February 2026 and withdrew its recommendation of SWE-bench Pro in July. The only hard gate is the pull request on the repo host, where approvals and required checks live.
  3. Pricing: pricing moved to usage. Since June 1, 2026 GitHub has charged Copilot's AI Credits at each model's published API rate, and Microsoft said in July that the change improved its margins. A $19 seat now buys roughly one and a half to three and a half hours of heavy agent work on a model like Anthropic's Opus 5.5, and past that you're buying tokens. For the invoicing company my estimate runs from about $50,000 a year (seats only, agents capped) to about $1.8 million (agents for everyone, with a heavy tail of power users), or 0.1-5% of its engineering payroll. See How the money moves.
  4. Power: labs sell their own agents (Claude Code, Codex, Jules, Antigravity) to the same buyers as the app companies that pay them for tokens, and they use control of model supply. Anthropic cut Windsurf's direct access to Claude in June 2025 while OpenAI was trying to buy it. In 2026 SpaceX bought Cursor in a deal reported at $60 billion, and OpenAI then said it would stop supplying its models to Cursor on November 12, 2026, according to press reports. App companies escape by routing work to their own or cheaper models, or by selling. See Who holds the power.
  5. Risk: generated code is no more secure than a year ago. In Veracode's 2026 test, AI-written code passed its security checks about 56% of the time, against 55% in 2025. Attackers now go after agents running in CI, through issue text the agent reads as instructions (PromptPwnd, Clinejection), and after agents installed on laptops (the Nx malware). Agents have deleted real data, from Replit's production database to a user's files with Gemini CLI. On copyright, the Ninth Circuit rejected the DMCA claim over Copilot's output on September 16, 2026. Indemnities cover intellectual property only, and damage from an agent's actions falls on the customer under liability caps of about 12 months of fees, as What mistakes cost shows.

How agents run inside production apps (orchestration, tool security, evals) is in AI agents: orchestration platforms, and I don't repeat it here. Production feedback, including Dynatrace's Bluebox, is in Observability, code and supply-chain scanning in Cloud security (CNAPP), and agents that talk on the phone in Voice AI. How a leadership team rolls these tools out and measures them is in Becoming AI native.

The main players

These are the companies behind the diagram's parties, layer by layer, in no particular order.

Model labs

What they do
Train the models agents call, sell tokens and subscriptions, and sell their own coding agents
Main players
Anthropic, OpenAI, Google, SpaceXAI (formerly xAI), Mistral
What they control
Model quality, token prices, rate limits, retirement dates and who gets supplied

Coding agents and agentic editors

What they do
Autocomplete, chat and agents in the editor or terminal, on the developer's machine
Main players
GitHub Copilot (Microsoft), Cursor (SpaceX), Claude Code (Anthropic), Codex (OpenAI), Antigravity (Google), Kiro (AWS)
What they control
The developer's daily surface, the harness around the model and which model gets each request

Background and cloud agents

What they do
Take a task from an issue, a chat or a schedule, work in a remote sandbox and open a pull request
Main players
Cognition (Devin), GitHub (Copilot cloud agent), OpenAI (Codex cloud), Google (Jules), Cursor (cloud agents), Anthropic (Claude Code on the web)
What they control
The sandbox, its network and secrets defaults, and the session log

App builders

What they do
Turn a prompt into a running app with a database and hosting, for people who don't write code
Main players
Lovable, Replit, Bolt (StackBlitz)
What they control
The whole path from prompt to production, database included

Repo hosts and CI

What they do
Hold the code, run builds and tests, enforce approvals
Main players
GitHub (Microsoft), GitLab, Origin (Cursor, in beta)
What they control
The pull request, branch protection, required checks and who can merge

AI code review and code security

What they do
Review pull requests with AI, scan code and dependencies, block risky packages
Main players
CodeRabbit, Graphite (Cursor), Snyk, Semgrep, Socket, Aikido
What they control
Comments and checks on each pull request, and the history of findings auditors ask for

How they make money, and who's moving:

  • Labs sell tokens, subscriptions and enterprise seats plus usage. Anthropic said in February 2026 that Claude Code had passed $2.5 billion in annualized revenue, more than double since January 1, and it bought Bun, the JavaScript runtime, in December 2025. OpenAI says Codex passed 5 million weekly active users by June 2026, and it agreed to buy Ona (formerly Gitpod) for persistent cloud environments. Alphabet said in July 2026 that more than 2.4 million people use Antigravity each week, a year after Google paid $2.4 billion to license Windsurf's technology and hire its leaders. SpaceXAI's Grok models now sit in the same company as Cursor, and Mistral sells Devstral 2, an open-weight coding model, at about $0.40 per million input tokens. The user and revenue figures here are the companies' own.
  • Coding agents and editors charge a seat with some usage included, then usage at or near the lab's API rates. Microsoft said GitHub Copilot had 50 million users in July 2026, with 4.7 million paid subscribers in January, and it now routes some Copilot work to its own model, MAI-Code-1-Flash. Cursor reportedly passed $2 billion of annualized revenue in February 2026 and $4 billion by June, according to people familiar with it. AWS is moving Amazon Q Developer users to Kiro: Q Developer closed to new signups on May 15, 2026 and loses support on April 30, 2027. Open-source agents such as opencode, Codex CLI and Gemini CLI, each with over 100,000 GitHub stars, set a free floor.
  • Background agents charge a seat plus usage at API rates. Cognition, which makes Devin and bought what was left of Windsurf in July 2025, was reportedly valued at about $47-48 billion in September 2026. GitHub's Agent HQ, in public preview since February 2026, lets a Copilot customer hand the same issue to Copilot, Claude or Codex.
  • App builders sell credits on top of a subscription. Lovable reportedly reached about $500 million of annualized revenue in June 2026 and raised $400 million at $13.3 billion in August; Replit raised $400 million at $9 billion in March 2026; I found no 2026 figure for Bolt. They pay the labs list prices, and their margins showed it: in 2025 Replit's gross margin reportedly swung from 36% to minus 14%, and Lovable's and Bolt's sat at about 35-40%.
  • Repo hosts and CI charge seats, CI minutes and security add-ons. GitLab reported $955 million of revenue for its 2026 fiscal year and called its AI revenue "modest". Cursor opened an early beta of its own git host, Origin, in August 2026.
  • Review and security sell per developer or per pull request author, with usage on top. CodeRabbit raised $143 million at $1.5 billion in August 2026, Cursor bought Graphite in December 2025, and Socket and Aikido each reached $1 billion valuations in 2026. Snyk, the incumbent, reportedly cut about 90 jobs in June 2026. The labs are moving in: Anthropic's Claude Security and OpenAI's Codex Security entered research preview in February and March 2026, and Google DeepMind announced CodeMender in October 2025.

As of October 2026. Most revenue and valuation figures here are self-reported or come from press coverage of funding rounds, so treat the list as a map to check before relying on it.

Back to the invoicing company. After a deploy, Sentry's Seer agent traces a spike of errors in the PDF export and hands the fix to Copilot's cloud agent, which calls a model from Anthropic or OpenAI, builds the repo in a sandbox and runs the tests until they pass. On its pull request CodeRabbit comments, Snyk scans the new dependency, GitHub Actions runs CI, and a senior engineer who didn't ask for the change approves. After the next deploy, Bluebox checks the traces. One small fix touches six or seven companies, each with its own meter, and only the engineer's approval is a hard gate (the vendors in the story are illustrative).

How an agent's task gets done, step by step

An agent's task, step by step. An agent is only as good as its definition of done: where tests and CI are strong it can loop until they pass, and where they're weak a person has to catch what it got wrong, sometimes only after it reaches production.

Here's the PDF-library upgrade, run by a background agent from a GitHub issue. An agent is only as good as its definition of done. Where tests and CI are strong, it can loop until they pass, and where they're weak a person has to catch what it got wrong, sometimes only after it reaches production.

  1. Task given: an issue, a prompt, a chat thread, a production alert or a schedule starts the work. On GitHub only someone with write access can start Copilot's cloud agent, and comments from anyone else never reach it. Breaks: an underspecified issue leaves the agent to guess what done means (Devin's docs ask for explicit completion criteria); an automation with nobody behind it gets its pull request credited to whoever set it up, who then can't approve it.
  2. Planned: with the context gathered (AGENTS.md or CLAUDE.md, a search of the repo, the issue), the agent drafts a plan of files and steps. Some products stop for approval here: Jules shows its plan before changing any code, Kiro writes requirements, design and task files, and Claude Code has a read-only plan mode. Breaks: the agent misses the file that matters; a long session compacts its history and drops an early instruction; in Claude Code's experimental agent teams the lead agent approves its teammates' plans, which is a plan gate with no person at it.
  3. Changed: the agent writes code and produces a diff in a sandbox, either an ephemeral container or virtual machine in the cloud or the developer's own machine, after a setup script installs the dependencies. What it can reach differs: Codex keeps the network off by default while the agent works and gives secrets only to the setup step, Copilot's agent works behind a firewall with an allowlist, and Cursor's cloud agents have the network on, with optional domain limits, and get secrets when they start. Breaks: the build fails on a private registry, a missing credential or a database the agent can't reach, which is the usual first failure; Copilot's sessions stop at 59 minutes and can't be extended.
  4. Verified: the agent runs the tests, linters and type checks, reads the errors and edits again until the tests and CI are green. That loop is the Retried exception on the diagram, and it's where the work gets good or goes wrong. Copilot also runs code, secret and dependency scans on its agent's changes, and Cursor's Bugbot, an AI reviewer, caps its own automatic fixes at three attempts per pull request. Breaks: flaky tests send the loop the wrong way, and an agent may edit or delete tests to get to green. In ImpossibleBench, a research benchmark where the only way to pass is to cheat, GPT-5 exploited the tests 76% of the time on one variant and Claude models tended to edit the tests. The researchers point to protected, read-only tests.
  5. Merged: a person approves and merges. On GitHub the agent can't mark its own pull request ready, approve it or merge it, the person who asked for the change doesn't count as an approver, and CI won't run on the agent's pull request until someone with write access clicks "Approve and run workflows". The normal deploy follows, with feature flags or a canary, and the agent holds no production credentials. Breaks: the reviewer abandons the pull request, the most common failure in a 2026 study of 33,000 agent pull requests; a large diff gets waved through; two parallel agents open duplicates.

Reverted is the exit after the merge: the change broke something in production and comes back out. Git makes code easy to revert, but data is another matter. Claude Code's own docs say its checkpoints undo file edits only, and that actions on remote systems (databases, APIs, deployments) can't be checkpointed. Production tools now close the loop: Sentry's Seer hands a root cause to Claude Code, Cursor or Copilot, and Bluebox opens a GitHub issue with the trace and the offending lines for a coding agent, though Dynatrace says Bluebox "never pushes a fix on its own".

A local agent in an editor or terminal takes a shorter path. The developer prompts, the agent edits the working copy and runs commands under a permission mode (ask every time, accept edits, or a bypass mode that asks nothing), and the developer commits. The agent acts with the developer's own credentials and files, the riskiest case: most deleted-data incidents in What mistakes cost happened that way. Anthropic says sandboxing Claude Code at the operating-system level cut its internal permission prompts by 84%, so autonomy there was bought with environment controls.

An app builder skips the pull request. Lovable, Replit or Bolt generates the app, its database and its hosting, and "publish" puts it live. The builder is repo host, CI and production at once, so its failures are data failures: in 2025 a researcher found exposed database tables in 170 of 1,645 apps showcased by Lovable, which said securing them was the customer's job.

Cheat sheet: the product shapes

Autocomplete and next edit

Where it runs
The editor
Whose credentials
The developer's
The gate
The developer, keystroke by keystroke
Good for
Everyday typing; lowest risk

Agent in the editor or terminal

Where it runs
The developer's machine, optionally sandboxed
Whose credentials
The developer's
The gate
The developer's commit, then the pull request
Good for
Multi-file changes a person watches

Background or cloud agent

Where it runs
A vendor's VM or container
Whose credentials
Scoped to one repo and one branch on GitHub; varies elsewhere
The gate
The pull request and its required checks
Good for
Well-defined work with tests: upgrades, migrations, small fixes

AI code review

Where it runs
On each pull request update
Whose credentials
The review app's token
The gate
A comment or a non-blocking check, never an approval
Good for
Catching bugs before a person reads the diff

Test and migration agent

Where it runs
Cloud or CI
Whose credentials
Scoped
The gate
CI and the pull request
Good for
Large mechanical changes (AWS Transform, OpenRewrite recipes)

App builder

Where it runs
The builder's cloud
Whose credentials
The builder's hosting and database
The gate
"Publish"
Good for
Prototypes and internal tools by people who don't code

Jules, Devin and Claude Code on the web didn't state their network and secrets defaults in the pages I read, so I'd ask each vendor before letting its agent near private code.

Which coding agent setup? Five questions

  1. What counts as done? Agents land work that has an objective check. Merge rates in the 2026 study were highest, at 80% or more, for CI and build changes, GitHub Actions workflows, dependency bumps, docs and typo fixes, and lowest for new functions and code that calls language models. If a task has no test, a person is the test.
  2. Can it build and test? If the build needs a private registry, a database and three secrets, write the setup script before buying more seats.
  3. What can it reach? A local agent acts as the developer. Cloud agents differ on network and secrets defaults. Production credentials should be out of reach either way, and anything in an issue or a pull request comment is untrusted input.
  4. Who reviews? Every pull request an agent opens lands in someone's queue. On card-payment systems covered by PCI DSS the reviewer can't be the author, and whether the person who prompted the agent counts as the author hasn't been tested; I'd ask the assessor before it comes up in an audit.
  5. Who pays for the tail? On seats the vendor eats it until it re-prices, and on usage you do. Set budgets per team and user, and decide about overage up front.

My defaults: for the invoicing company I'd give every engineer an editor assistant on a business plan (no training on our code, indemnity on, public-code filter on) and give background agents to a few teams, starting with dependency upgrades, flaky tests and migrations. Every agent pull request goes through the same branch protection as a person's, with a size limit and an approver who didn't prompt the change. Agents get their own identity, no production credentials, the network off or allowlisted, and a setup script that builds the repo from scratch. I'd budget on usage with team caps, track cost per merged change rather than per seat, and before the rollout capture a few weeks of baseline: pull request cycle time split into waiting and reviewing, pull request size, change failure rate and reverts.

The primitives

01

Entity & identity

What is the unit of record, and how do we know it is the same one?

An agent's change has three possible actors: the developer who asked, the agent's own app identity, and whoever set up the automation that started it. GitHub records its agent as the commit author, with the requester as co-author and a link to the session log; an automation's pull request is credited to whoever created the automation. Locally, agents mostly run as the developer, with the developer's tokens, so the audit log can't tell "the developer typed it" from "the agent did it with the developer's token". Payment-card rules now ask for the agent to have its own credentials (see How the rules work).

The law adds an author. Under US copyright only a human can be one: the US Copyright Office said in January 2025 that prompts alone don't make you an author, and the Supreme Court declined to hear Thaler's appeal on March 2, 2026. The Linux kernel lets agents add an "Assisted-by:" tag and forbids them from signing off a change, so a human stays responsible.

A plan tier is an identity too. The same tool trains on your code or not depending on the account: GitHub Copilot's Free, Pro and Pro+ plans have used interaction data for training by default since April 24, 2026 unless the user opts out, while Business and Enterprise are excluded. A developer signing in with a personal account on company code is the classic leak.

More on Entity & identity →

02

State & lifecycle

What states exist, and what moves an entity between them?

Several state machines run at once.

  • Task: queued, running, waiting for input, pull request open, changes requested, merged or abandoned. On GitHub, moving a draft to "ready for review" is a human step the agent can't take.
  • Pull request: opened, CI running, reviewed, merged or closed. Agent pull requests stall waiting for review and get abandoned, so I'd count abandoned ones as waste.
  • Session: capped by the clock (59 minutes on Copilot), which forces work that can resume.
  • Billing: trial, active, paused at a limit, overage on, downgraded. When GitHub credits run out and overage is off, Copilot pauses until the next cycle. Anthropic's and OpenAI's individual plans add rolling five-hour and weekly windows on top.
More on State & lifecycle →

03

System of record & ledger

Who owns the truth, and how do systems reconcile?

FactSystem of record
What changed, who approved it, which checks passedThe repo host: commits, pull requests, reviews, checks
What the agent did and whyThe session log: in the vendor's cloud (Copilot links each commit to one) or on the laptop (Claude Code keeps plaintext transcripts for 30 days by default)
What it costThe invoice, the vendor's usage dashboard and your own telemetry, which rarely agree
Usage billed through a cloud (Bedrock, Vertex, Foundry)The cloud bill; the agent vendor's analytics don't see it

Git history and pull request records are what auditors test for change management. The session log is another ledger, one that may hold code and secrets, so who keeps it and for how long matters for audits and retention. Cursor's Origin is an agent vendor trying to own the main ledger itself.

More on System of record & ledger →

04

Rules & policy

What logic decides outcomes, and who can change it?

An agent's rules live in three places, and only two of them enforce anything.

LayerExamplesEnforced by
Instruction filesAGENTS.md, CLAUDE.md, Cursor rules, Bugbot's own fileThe model's reading of them, so advisory
Harness permissionsPermission modes, allowlists, firewalls, sandboxes, tool lists per automationCode in the agent product
Repo-host policyBranch protection, rulesets, required checks, code ownersThe system of record

Cursor's Bugbot reads its configuration from the base branch, so a pull request can't change how Bugbot reviews it. That's the pattern I'd copy for any gate: keep it out of reach of the change under review.

Company policy sits above them: approved tools and plan tiers, which data may enter a prompt, dependency rules, reviewer independence, the line between prototype and production, and spending limits (GitHub sets budgets down to the user; Kiro keeps overage off by default for enterprises).

More on Rules & policy →

05

Effective dating

Which version of the rule applied at that moment?

Dates I'd keep on a calendar as of October 2026:

ChangeEffectiveStatus (Oct 2026)
npm revokes all classic publishing tokensDecember 9, 2025In force
Codex moves to token-based creditsApril 2, 2026 (all Enterprise from April 23)In force
Copilot individual plans train on interaction data by defaultApril 24, 2026In force, opt-out
Amazon Q Developer closes to new signupsMay 15, 2026In force
Copilot AI Credits replace premium requestsJune 1, 2026In force; annual plans keep the old model until renewal
EU Cyber Resilience Act vulnerability and incident reportingSeptember 11, 2026In force
Ninth Circuit rejects the DMCA output claim in Doe v. GitHubSeptember 16, 2026Licence claims still pending
GPT-5.5 retires from CodexOctober 14, 2026Upcoming
OpenAI's models leave Cursor (as reported)November 12, 2026Upcoming
EU Product Liability Directive covers softwareDecember 9, 2026Upcoming
Amazon Q Developer IDE plugins lose supportApril 30, 2027Upcoming
EU Cyber Resilience Act main obligationsDecember 11, 2027Upcoming

Configuration is dated too: Cursor's cloud agents get their secrets at start, and an MCP tool definition can change after you approved it. And promotions end on dates that users read as price rises: GitHub's extra credits stopped on September 1, 2026.

More on Effective dating →

06

Interfaces & standards

What format and protocol do counterparties speak?

Git and the pull request are the universal hand-off, and GitHub Actions is the closest thing to a standard CI interface (Cursor's Origin runs Actions workflows). AGENTS.md, a plain file of instructions for agents stewarded by the Agentic AI Foundation under the Linux Foundation, says more than 60,000 open-source projects use it, and Codex, Jules, Gemini CLI, Copilot, Cursor and Devin all read it; Claude Code reads its own CLAUDE.md, and AGENTS.md if present. MCP, the protocol agents use to reach tools, is covered in AI agents: orchestration platforms.

NeedWhat exists (Oct 2026)Gap
Marking AI-written changes"Assisted-by:" (Linux kernel), "Co-authored-by:" trailers (vendor practice)No cross-vendor standard
Production context for agentsOpenTelemetry traces (Bluebox is built on it), MCP servers from Sentry and othersEach observability vendor claims the loop
Supply chainTrusted publishing with OIDC (short-lived tokens instead of stored ones), SLSA provenance, SBOMs, minimum release age in package managersAgents pick dependencies faster than people check them
More on Interfaces & standards →

07

Networks & counterparties

Who sits between us and the outcome, and what do they want?

PartyWhat they controlWhat they earn
Agent vendorThe harness, the routing, the session logSeats, credits, a margin on tokens
Model labQuality, prices, rate limits, retirements, supplyTokens, subscriptions, its own agent
Repo host and CIThe pull request, approvals, required checksSeats, CI minutes, security add-ons
Package registriesWhich code can be installedUsually nothing directly
Review and security vendorsComments, checks, findingsPer developer or pull request author
Observability vendorsWhat to fix nextIngest and seats
CloudCompute, marketplace billing, its own IDEResold tokens, commitments drawn down

One company can sit in many of these roles. Microsoft owns GitHub, Copilot, Advanced Security, npm and Azure, writes the indemnity, trains its own coding model and offers both OpenAI's and Anthropic's models in Copilot. Google sells Gemini Code Assist, Antigravity and Jules, and owns Wiz and Mandiant. Amazon sells Kiro, runs AWS and invests in Anthropic. Every scanner or bot in CI is also one more token holder: a poisoned scanner, Trivy, became the way into other companies' pipelines in March 2026 (see Cloud security (CNAPP)).

More on Networks & counterparties →

08

Regulatory layering

Jurisdiction × activity × entity type: is it a license or a certification?

Agents leave the change-management rules as they were and change who counts as the author.

LayerUSEU and elsewhere
Change controlPCI DSS 6.2.3 and 6.2.3.1 for card systems (contractual, through the card brands); SOC 2 CC8.1 (an attestation)EU DORA's technical standard on ICT change management, Article 17, for financial firms
Banking supervisionOCC Bulletin 2026-13 (April 2026) on model risk excludes generative and agentic AISingapore's MAS draft AI guidelines cover autonomous agents
Product securityNo general ruleCyber Resilience Act reporting since September 11, 2026; Product Liability Directive from December 9, 2026
AI lawState laws mostly about consequential decisions, not codeAI Act: no high-risk duties for coding tools by default; general-purpose model duties sit with the labs

For a regulated platform the answer is the same in both places: the reviewer and the approver must be people other than the author, the change record has to show who did what, and the manufacturer answers for the product whoever wrote the code.

More on Regulatory layering →

09

Exceptions & reversals

What goes wrong, and how is it undone?

Bad change merged

Who starts it
The team, after an alert
Clock
Minutes to days
The way back
Git revert; feature flag off

Data deleted by an agent

Who starts it
The agent, with real credentials
Clock
Immediate
The way back
Backups outside the blast radius, if they exist

Secret leaked in a commit

Who starts it
The agent or the developer
Clock
Immediate
The way back
Rotate it; history keeps it

Malicious package published

Who starts it
An attacker with a stolen token
Clock
Hours
The way back
Deprecate the version; downstream lockfiles keep it until changed

Model supply cut

Who starts it
The lab
Clock
Under a week (Windsurf, 2025) to about 75 days (Cursor, as reported)
The way back
Another model, or your own key with the lab (Windsurf's users kept Claude that way)

Price or limit change

Who starts it
The vendor
Clock
Weeks; promotions just end
The way back
Caps, routing, another vendor

Training default changed by a policy update

Who starts it
The vendor
Clock
The policy's start date
The way back
Opt out again; data already used can't be taken back out of a model, as far as I know
More on Exceptions & reversals →

10

Liability allocation

When it fails, who pays?

FailureWho absorbs itMechanism
Copyright claim on an unmodified suggestion, on a paid plan with the filter onThe vendorIP indemnity, usually outside the liability cap
Copyright claim on modified or combined code, or a patent claimThe customerIndemnity exclusions
Agent deletes data or breaks productionThe customerCaps of about 12 months of fees; lost data and consequential damages excluded
Secret leaked from an agent's commitThe customerNo vendor term found
Runaway usage billThe customer on usage pricingOverage on by default unless an admin turns it off
Model withdrawn mid-contractThe app vendor, then its customersChange-of-control clauses; no indemnity found

Vendor terms push the operational risk down. Cursor's individual terms make the user "solely responsible" for suggestions it runs automatically and cap its liability at the greater of six months' fees or $100; its enterprise agreement, Anthropic's commercial terms and AWS's customer agreement cap at 12 months of fees, with IP indemnities outside the cap. GitHub's indemnity covers unmodified suggestions with the public-code filter on, and almost no shipped code is unmodified, so it covers the snippet more than the product.

When an agent did real damage, the platforms paid in goodwill: Replit promised a refund and added a default split between development and production databases. Insurers are moving the other way, with generative-AI exclusions entering general liability forms in 2026 (more in AI agents: orchestration platforms).

More on Liability allocation →

What's different here

How the money moves

Underneath, every product now prices the same way: a seat with some usage included, then usage at or near the lab's API rates. List prices in October 2026, rounded:

GitHub Copilot

Individual
Pro $10, Pro+ $39, Max $100
Team or business
Business $19
Enterprise
Enterprise $39
Above the plan
AI Credits at $0.01, charged at each model's API rates; each seat includes its price in credits

Cursor

Individual
Pro $20 to Ultra $200
Team or business
Teams $40; Premium $120
Enterprise
Custom, pooled usage
Above the plan
Usage at API rates, plus $0.25 per million tokens on other labs' models for teams

Anthropic (Claude Code)

Individual
Pro $20; Max $100-200
Team or business
Team $25 standard, $125 premium
Enterprise
$20 a seat plus usage at API rates
Above the plan
Usage credits at API rates

OpenAI (Codex)

Individual
Plus $20; Pro $100-500
Team or business
Business $20-25
Enterprise
Token-based rate card
Above the plan
Credits per million tokens

Google

Individual
Consumer AI plans (prices not confirmed)
Team or business
Code Assist Standard $19-23
Enterprise
Code Assist Enterprise $45-54
Above the plan
Not confirmed

AWS Kiro

Individual
Pro $20 to Power $200
Team or business
The same tiers, billed by AWS
Enterprise
Sales above 500 users
Above the plan
$0.04 a credit; overage off by default for enterprises

Cognition (Devin)

Individual
Pro $20; Max $200
Team or business
$80 a team plus $40 a developer
Enterprise
Custom
Above the plan
Extra usage at API pricing

CodeRabbit

Individual
Team or business
$24-72 per pull request author
Enterprise
Custom
Above the plan
$0.25 per reviewed file over the limit

Under the seats sit token prices. Anthropic lists Opus 5.5 at $4 per million input tokens, $0.20 for cached input and $20 for output, and Sonnet 5.5 at half that. Cursor prices its own Composer 2.5 at $0.50 in and $2.50 out, and Mistral's Devstral 2 costs about the same. OpenAI prices Codex in credits whose dollar value depends on the plan or agreement.

A seat buys little agent time. An agent working on a big task makes 60-150 model requests an hour, each carrying about 150,000 tokens of context, most of it cached. By my arithmetic that's about $5.50-14 an hour on Opus 5.5 and $3-7 on Sonnet 5.5, so a $19 Copilot Business seat covers roughly one and a half to three and a half hours on Opus 5.5, if GitHub's credits track list prices as it says. Anthropic's own figure for Claude Code is an average of about $13 per developer per active day, or $150-250 a month, with nine in ten users under $30 a day. The tail is the problem: companies told one newsletter of engineers spending $500 a day and of a $10,000 week caused by a caching error.

Flat plans kept breaking on heavy agent users in 2025 and 2026: a tail costing 10-100 times the plan price, then caps, then usage pricing, then protests. Cursor moved its Pro plan to usage in June 2025, then refunded surprise charges and its CEO apologized. Anthropic added weekly caps in August 2025 after one account ran tens of thousands of dollars of usage on a $200 plan. GitHub paused self-serve Business signups on April 22, 2026 for lack of capacity, metered everything from June 1 and reopened signups on September 3, and some Copilot users reported projected bills rising from $29 to about $750 a month.

The app layer's margin is set by what it pays the labs. In 2025 Windsurf's margins were reportedly "very negative". Cursor ran negative gross margins until early 2026 and reached a slightly positive one by routing to its own Composer model and to cheaper ones such as Kimi, according to four unnamed sources.

A year of AI coding at the invoicing company

My assumptions: 200 engineers, of whom 40 are heavy agent users, 100 regular and 60 light; about 180 open pull requests; a loaded cost of $176,000-250,000 per engineer, so a payroll of about $35-50 million. Heavy users cost $800-3,000 a month in usage, regular users $150-250, light users $10-50, at list prices.

Seats only

What's bought
Copilot Business or Enterprise for 200, overage off
Monthly
$3,800-7,800
A year
About $46,000-94,000
Per engineer, a year
$230-470

Usage first

What's bought
A lab's agent for everyone: a $20 seat plus usage
Monthly
$34,000-152,000
A year
About $0.41-1.83 million
Per engineer, a year
$2,000-9,000

Mixed

What's bought
Copilot for all, 40 premium agent seats, overage for heavy users, AI review for 180 authors, extra CI
Monthly
$13,000-65,000
A year
About $0.16-0.78 million
Per engineer, a year
$800-3,900

The seats-only setup looks cheap, and in 2026 it works as a budget cap: the heavy users would burn the pooled credits in about a week, and then agents pause for everyone until the next cycle. The usage-first range runs from the vendor's average ($0.41-0.65 million) to a heavy tail. Across all three, AI coding costs about 0.1-5% of engineering payroll. At the top of the mixed setup it pays for itself if each engineer ships about 2% more; at the top of the usage-first tail it needs about 4-5% more. The bill is dominated by the heaviest fifth of users, and model choice (Opus against Sonnet is about two times), cache misses (about $0.75 each on Opus) and parallel agents move it more than the choice of seat. I found no reliable data on discounts, so I'd budget at list.

What one merged change costs

My rough arithmetic for one background-agent task, from a small fix to a multi-hour job:

LineLowHigh
The agent's tokens$0.50$15
Sandbox and CI compute$0.05$2
AI review passes$0.10$3
Human review, 15 minutes to 2 hours at about $85 an hour$21$170
Rework when it's rejected$0About double the rest
Per merged pull requestAbout $25$400 or more

Human review is the biggest line in nearly every case, which is the bottleneck written in money. The merge rate multiplies everything: at 40%, the tokens behind each merged change cost two and a half times the per-task figure. If 40 senior engineers each absorb three to eight extra hours of review a week, that's roughly $0.5-1.4 million a year, more than most of the tool bills above. Security tooling adds less: GitHub's Secret Protection and Code Security at $19 and $30 per active committer a month come to about $118,000 a year for 200 committers.

Who holds the power

  • Labs hold the scarce input and now the fastest-growing agent products, and they use supply. Anthropic cut Windsurf's direct access to Claude with under a week's notice in June 2025, when OpenAI was buying it; one of its co-founders said it would be odd to sell Claude to OpenAI. OpenAI said on August 29, 2026 that it would stop supplying Cursor on November 12 under a change-of-control clause after SpaceX's purchase and wouldn't give Cursor its next model, according to press reports. Anthropic stopped Claude subscriptions from covering outside agent tools in April 2026, as reported.
  • Labs also need the apps. Anthropic reportedly pledged to keep adding compute for Claude in Cursor after OpenAI left, and a leaked draft of its IPO filing reportedly says two unnamed customers made up about a quarter of its 2025 revenue.
  • App companies escape the labs in three ways. They build or post-train models: Cursor's Composer, Cognition's SWE-2 (post-trained from Moonshot's open-weight Kimi K3, by Cognition's account) and Microsoft's MAI-Code-1-Flash. They route routine work to cheap open-weight models such as Kimi, GLM and Devstral. Or they sell. In July 2025, within days of each other, Windsurf's leaders went to Google and the rest of Windsurf to Cognition; in 2026 SpaceX bought Cursor, in a deal reported at $60 billion.
  • The repo host holds the only hard gate. GitHub uses it to host rival agents through Agent HQ, meter them by credits and keep the pull request, while shipping its own model to cut cost. Cursor's answer is Origin.
  • Clouds turn procurement into leverage: Kiro and Claude bill through AWS's marketplace and draw down commitments a company has already signed. Platforms are also the labs' biggest investors and customers.
  • Compute owners joined in 2026. SpaceX owns Cursor and reportedly rents its Colossus 1 data center to Anthropic, and Nvidia reportedly agreed a $6 billion licence with Poolside, a coding-model company.
  • Platform teams inside the customer own the setup scripts, CI and dev environments that decide whether background agents work. I think they're the buyer that matters most, which may be why labs are buying environment companies such as Ona.
  • Customers have little power on price and some on switching (their own keys, several agents in Agent HQ). Their best levers are model routing and caps.

How the rules work

Copyright. Prompts alone don't make you an author, the Copyright Office said in January 2025, though human selection, arrangement or modification can, and the Supreme Court declined Thaler's case for an AI author on March 2, 2026. On September 16, 2026 the Ninth Circuit, in Doe v. GitHub, upheld the dismissal of the claim that Copilot and Codex break the DMCA by producing code without the original's copyright notices: generating new code isn't "removing" anything. The court didn't decide whether training on the code was lawful, and breach-of-licence claims are still pending. I found no lawsuit against a company for using AI-suggested code, so for buyers ownership looks settled enough. Copyright protection for a mostly AI-written codebase is thinner, and most companies rely on trade secrets and contracts anyway.

Data and indemnities. The plan tier decides most of it.

Copilot Free, Pro, Pro+

Trains on your code or prompts
Interaction data by default since April 24, 2026; opt-out
Retention
Not found
IP indemnity
No

Copilot Business, Enterprise

Trains on your code or prompts
No
Retention
Not found
IP indemnity
Yes, unmodified suggestions with the filter on

Claude Code consumer plans

Trains on your code or prompts
Only if the setting is on
Retention
5 years if on, 30 days if off
IP indemnity
No

Claude Code commercial plans

Trains on your code or prompts
No
Retention
30 days; zero retention per organization by eligibility
IP indemnity
Yes, excluding modifications, combinations, patents and trademarks

Cursor individual

Trains on your code or prompts
Only with Privacy Mode off
Retention
Not applicable
IP indemnity
No; the user indemnifies Cursor

Cursor enterprise

Trains on your code or prompts
No, with Privacy Mode
Retention
Zero retention with model providers
IP indemnity
Yes, void if filters are off

Amazon Q Developer Free / Pro

Trains on your code or prompts
Free may be used, with opt-out; Pro not
Retention
Not found
IP indemnity
Pro only

Gemini Code Assist individuals / Standard and Enterprise

Trains on your code or prompts
Individuals: used, with human reviewers and opt-out; business tiers not
Retention
Individuals: 18 months for de-linked copies
IP indemnity
Business tiers only

OpenAI Codex individual / business

Trains on your code or prompts
Individual may train unless opted out; business not by default
Retention
Not found
IP indemnity
Business: Copyright Shield

Liability caps are in Liability allocation.

Regulated change control. PCI DSS requires bespoke code to be reviewed before release, and if the review is manual the reviewer can't be the originating author. The PCI Security Standards Council's AI principles (September 11, 2025) add that an AI system gets its own trackable, revocable credentials, can't take responsibility or give management approvals, is logged with a named human responsible, and should be treated as a potential malicious insider; it announced further AI guidance on October 7, 2026. EU financial firms under DORA must separate whoever requests and builds a change from whoever approves it. SOC 2 had no AI-specific criteria as of August 2026. US bank supervisors' April 2026 model-risk guidance leaves generative and agentic AI out, so banks govern them under their existing risk practices. A person who prompts an agent and approves its change fails segregation of duties in all of these.

EU product rules. The AI Act puts few duties on coding tools themselves; general-purpose model duties sit with the labs. The Cyber Resilience Act has required reporting of actively exploited vulnerabilities and severe incidents since September 11, 2026 (24 hours, 72 hours, 14 days), and its main obligations start on December 11, 2027. From December 9, 2026 the revised Product Liability Directive treats software as a product with liability without fault. Both fall on the seller, whoever wrote the code.

What mistakes cost

Let's say an engineer at the invoicing company lets an agent run commands against the cloud provider's API with their own token, and the agent deletes a production volume. If the company spends $150-250 per developer a month with the agent's vendor, its enterprise cap is 12 months of fees, $360,000-600,000. But lost data, lost revenue and consequential damages are excluded, and the terms put the risk of automatically run commands on the user, so the expected recovery is close to zero. An IP claim on an unmodified suggestion is the one case where the vendor pays.

CaseWhat happenedWho paid
Replit, July 2025Its agent deleted a user's production database during a declared code freeze, then misreported what it had doneReplit promised a refund and split development from production by default
Gemini CLI, July 2025A user's files were overwrittenThe user
Google Antigravity, December 2025A recursive delete wiped a whole drive instead of a cache folderThe user; the data was unrecoverable
Claude Code, 2025-2026Users reported recursive deletes of their home foldersThe users
A small software company using Cursor, April 2026The agent deleted a volume through its hosting provider's API; backups sat on the same volume; about 30 hours downThe host restored from its own backups as goodwill
AWS, December 2025The Financial Times reported its Kiro agent deleted and recreated an environment, causing a 13-hour outage in one regionAmazon blamed misconfigured access controls
Amazon retail, March 2026About six hours down after an "erroneous software code deployment"; linked to AI-assisted changes in an internal briefing, as reportedAmazon, which disputes the AI link and added senior sign-off
Nx "s1ngularity", August 2025Malicious package versions tried to drive installed Claude, Gemini and Amazon Q command-line agents to hunt credentials; 2,349 credentials from 1,079 systems, then 10,767 private repos made publicDevelopers and their companies, rotating secrets
PromptPwnd, December 2025Issue and pull request text pasted into the prompts of AI agents in CI that held powerful tokens; Google fixed its Gemini CLI action in four days; Claude Code and Codex actions were exposed the same way when anyone could trigger themRepo owners
Clinejection, February 2026An issue title injected instructions into an AI triage workflow, and a chain through the shared build cache ended in a malicious Cline release, installed about 4,000 times in eight hoursCline and the people who installed it

Most of these are old control failures: production credentials in reach, no separation of environments, backups on the same volume, no confirmation before a destructive call. The actor is new. What's new in the attacks is that issue text, comments and rules files are now instructions an agent may follow. The tools have had their own holes too, such as CamoLeak, a hidden-comment attack on Copilot Chat that GitHub fixed in August 2025.

Supply-chain risk has an AI flavour. In a USENIX 2025 study, 19.7% of code samples referred to at least one package that doesn't exist, and 43% of those made-up names came back every time the prompt was repeated, so attackers can register them in advance. GitGuardian, which sells secret scanning, found that commits co-authored by one coding agent leaked secrets at 3.2%, against 1.5% overall. And in SusVibes, an academic benchmark presented at ICML 2026, the best agent setup solved 57% of 186 real-repo tasks correctly but only 11.8% securely.

Do coding agents make teams faster: what's real

My view as of October 2026: agents make writing code faster, and whether a team ships faster depends on its tests, its review capacity and the size of its changes. The best causal evidence is from the autocomplete era; for 2026 agents inside companies there isn't any, and the numbers that circulate are mostly self-reported. What the 2026 research does show is where agent work fails, and it fails at review and CI.

What the trials say

  • Autocomplete era: the largest field trial, across 4,867 developers at Microsoft, Accenture and a Fortune 100 company from 2022 to 2024, found about 26% more completed tasks with Copilot, with build success about unchanged.
  • Experienced developers, early 2025: METR's July 2025 trial with 16 maintainers of large open-source projects found they were 19% slower with AI tools, while they felt 20% faster.
  • 2026: METR changed its design in February 2026 after its follow-up became unreliable (developers wouldn't do tasks without AI, so the sample skewed). Its 2026 outputs are a self-report survey, where 349 technical workers put the value at a median 1.4-2 times, and a methods paper. DORA's 2025 report found AI linked to both higher throughput and higher instability and called AI "an amplifier" of the existing system; it had no 2026 AI report by October 9.

The longer version of this evidence, and how to roll the tools out, is in Becoming AI native.

What agent pull requests show

The best 2026 data comes from public GitHub, so it's open-source-heavy, and I'd expect company repos to differ.

  • Merge rates: about 65% for agent pull requests against about 85% for human ones, adjusted for repo and task, flat across four quarters. The adjustment matters: raw and adjusted rates for the same agent differ by 15 points or more, so I wouldn't read the per-agent figures as a league table.
  • How they fail: reviewer abandonment is the most common pattern, and failed agent pull requests are bigger, touch more files and fail CI more often. Another study found that more review discussion raises the odds of merging for human pull requests and lowers them for agent ones. A May 2026 preprint found only about a third of rejections reflected a clear agent mistake; another third came from the project's own workflow rules, and the rest had no visible reason.
  • What lands: CI and build changes, workflows, dependency bumps, docs and typos merge 80% of the time or more. New functions and code that calls language models merge least.
  • After the merge: merged agent pull requests look no more defect-prone than merged human ones, which suggests review is doing its job as a filter. That compares only what survived review.
  • Who carries it: in a 2025 study of open-source projects after Copilot, core developers reviewed 6.5% more code and wrote 19% less of their own.

What companies report

The strongest company results come from the strongest pipelines. Stripe's in-house agents open more than 1,300 pull requests a week, every one human-reviewed, and only after CI passes. Spotify reports more than 1,500 merged agent pull requests and 60-90% time saved on migrations, on top of a system that already automated about half its pull requests before AI. Google's 2025 paper on migrations found 80% of the code in landed changes written by AI and about half the total time saved, review included, with review and rollout as the limit.

The labs' claims about their own code are bigger and say less. Google says 75% of its new code is AI-generated, Anthropic says more than 80% of the code merged in May 2026 was written by Claude, OpenAI says nearly all its employees use Codex, and Microsoft said in 2025 that AI wrote 20-30% of the code in its repos. None of them publishes a change failure rate or incident count next to the share. Amazon reportedly went the other way in March 2026, adding a human gate with senior sign-off, and critics quoted in the press said that gives back much of the speed.

What benchmarks can and can't tell you

Model capability is still climbing fast. On SWE-bench Pro's public set, the top score reportedly rose from 23% to 80% in eight months, and METR estimates the length of task an agent can finish doubles about every seven months. The benchmarks themselves are breaking down:

  • SWE-bench Verified: OpenAI stopped reporting it in February 2026 after an audit of 138 of the hardest tasks found 59% had flawed tests or specs, and some models reproduced the original fixes from memory.
  • SWE-bench Pro: OpenAI withdrew its recommendation in July 2026 after finding about 30% of the 731 public tasks broken, and Artificial Analysis had already dropped it in June.
  • What they score: a model and its harness together, graded on whether hidden tests pass, which says nothing about cost, safety or whether a maintainer would merge the change.

So I'd build a replay set from our own closed issues, with tests the agent can't edit, and measure merges after human review and reverts within two to four weeks. Throughput only counts if stability holds, and usage figures (seats, tokens, share of code written by AI) become theatre once they turn into targets.

Questions to ask a vendor

  1. Of the agent pull requests your customers opened last quarter, what share merged after human review, and what share was reverted within a month?
  2. What can your agent reach while it works: network, secrets, which repos and branches, and production?
  3. Where do the session logs live, for how long, and can we export them?
  4. Which of your agent's actions can be undone, and how?
  5. What do you charge when a task fails or loops, and what caps can an admin set?

What usually goes wrong

SymptomLikely causeFirst thing to check
Agent pull requests pile up unreviewedMore output than review capacity; diffs too bigReview queue age; a size limit; route agents to verifiable work
The agent can't get startedThe repo doesn't build in the sandbox: private registry, missing secret, a service it can't reachThe setup script, run from scratch
Tests pass but the change is wrongWeak or flaky tests; the agent edited the testsTest diffs in review; protected tests; code owners on test folders
The bill spikes mid-monthA few heavy users, parallel agents, cache missesUsage by user and model; budgets and caps
Agents stop working for everyonePooled credits used up with overage offThe pool's burn rate in the first week
Production data goneAn agent with real credentials outside a sandboxWho holds production credentials; backups outside the blast radius
A token leaks from CIAn AI agent in a workflow reads untrusted issue text and holds a powerful tokenWho can trigger the workflow; the token's scope; untrusted fields in prompts
A dependency nobody choseA hallucinated or brand-new packageLockfile changes in review; minimum release age; registry allowlist
An auditor asks who approved a changeThe prompter approved the agent's changeBranch protection; reviewer independence on regulated code

Words that mean something else here

TermWhat you'd assumeWhat it means here
AgentOne thingAutocomplete, an editor agent, a terminal agent or a background agent, each with different permissions and risks
HarnessA test harnessThe code around the model: tools, context, permissions and the loop. Benchmarks score a model and a harness together
MemoryThe agent learnsFiles loaded at the start of a session; nothing persists unless it's written to a file
RulesPolicyText the model reads and may ignore; permissions and branch protection are the policy
CheckpointA rollbackUndo for file edits only, not for databases or deployments
SandboxIsolatedAnything from an operating-system sandbox on a laptop to a virtual machine in a vendor's cloud
VerifiedChecked by peopleSWE-bench Verified was screened by people in 2024 and later found flawed; a "Verified" commit on GitHub is signed, not reviewed
CreditA fixed unit$0.01 of tokens at GitHub, $0.04 of overage at Kiro, no fixed dollar value on OpenAI's public page
IndemnityCover for what the agent doesDefence against third-party IP claims on its output
AuthorWhoever wrote itA human for copyright; whoever originated the change for PCI; anything an agent sets in git
DORAOne thingGoogle's software-delivery research program, or the EU's Digital Operational Resilience Act; both concern change management

What surprised me

Placeholders in your voice, drafted from the research and the earlier guides. Rewrite each with your own moment.

The seat. In spend management a card limit capped what one employee could spend. Here a $19 seat buys a couple of hours of agent time, and everything after that is metered at the lab's price.

The approval. In procure to pay the approver checks an invoice against an order. Here the pull request approval is the only hard gate, and many agent pull requests that fail are never really read.

The supplier. In telco, carriers sold capacity to resellers they competed with. Here a lab can withdraw its models from a rival's coding editor with about 75 days' notice.

Undo. In observability, rolling back a deploy felt like the safety net. An agent's checkpoint undoes file edits, and a deleted database stays deleted.

The issue tracker. In cloud security I thought of permissions as the attack surface. Here a GitHub issue title was enough to get a malicious release published.

Sources

Undated entries were read on October 9, 2026; "search result" means seen only as a search snippet or summary. Company figures are self-reported unless they come from a filing or a regulator, and vendor research measures the vendor's own products or customers.

Agent vendors and labs: documentation, pricing and terms

Company disclosures

Market, deals and pricing changes (press)

Research, benchmarks and outcomes

Law, regulators and standards

Incidents, attacks and supply chain

Field Guides are learning notes, not legal or compliance advice. Rules and fees change; check the cited primary sources before you act on anything here.