The Platform PM
Playbook

Becoming AI native: how your team ships, works and serves customers

For a product leader asked to make a regulated platform team AI native: measure first, then change how code ships, how internal work is reviewed and what customers see, in that order, with outcome metrics, real oversight and partner sign-off.

For heads of product, directors and VPs running 4-8 PMs and 20-60 engineers at complex B2B platform companies

Last reviewed October 2026

The job on one page

The CEO has read that Google says 75% of its new code is AI-generated, and has asked you to make product and engineering "AI native". You run a platform where mistakes cost money, licences and partner trust: payments, cross-border payouts, telco or CPaaS, security or identity. Your engineers already use AI tools, some unapproved; support and compliance are curious and nervous; sales wants an AI agent in the product next quarter. Over about a year you need to change how the team ships code, how internal work gets done and what customers see, prove it worked, and keep quality, security and compliance intact.

Becoming AI native in four steps: measure a baseline, change how code ships, rewire internal flows, then put AI in front of customers. Developers' own sense of speed is unreliable, so the baseline of outcome metrics is what makes the rest provable. Usage without outcomes is theatre.

The path runs in four phases: baseline (measure first), ship code (with review), internal flows (AI assists people who decide) and customers (with sign-off). The baseline is the hinge, because developers can't feel the effect. In METR's July 2025 trial, experienced developers expected AI to make them 24% faster, felt 20% faster afterwards, and were measured 19% slower (METR). Without a baseline of outcome metrics taken before rollout, nothing after it is provable. One branch leaves the second phase and goes nowhere: theatre, usage without outcomes, where seats, prompts and leaderboards rise and nothing customers or auditors care about changes.

The author's framing has three parts, and the phases take them in order:

PartWhat changesPhase
1. How the team shipsCoding tools and agents, review, tests, security2
2. How internal flows workSupport, compliance, incidents, product work3
3. How customers are servedAssistants, agents, docs for agents, pricing4

The baseline (phase 1) measures all three before anything changes.

Four ideas organize this playbook. They come from research, not from my own rollout, so each carries its evidence strength.

  1. AI speeds up writing code more than shipping it. (Strong: trials, quasi-experiments and field data agree.) Gains reach production only when work is well specified and tests, CI and review capacity can absorb more change; otherwise they pile up as review queues, rework and instability. In open-source projects that adopted GitHub Copilot, core developers reviewed 6.5% more code and wrote 19% less of their own (Xu et al.); in 807 projects that adopted Cursor, a velocity spike faded while static-analysis warnings and complexity stayed up (He et al., CMU). Google's DORA research calls AI an amplifier of the system it lands in (DORA 2025). See phase 2.
  2. AI native is a change to how work and decisions are reviewed, not a tool rollout. (Strong.) Usage mandates and usage metrics get gamed: Amazon reportedly dropped an internal AI-usage leaderboard after staff pointed agents at pointless tasks to climb it (The Decoder, via the Financial Times, May 29, 2026; secondary press). And oversight has to be meaningful: under EU and UK rules, a human who approves almost everything leaves the decision "solely automated". See Measuring it.
  3. Start internal AI with high-volume, decomposable, checkable work where AI assists a reviewer. (Supported, with changes.) The best causal study of AI at work, 5,179 support agents, found 14% more issues resolved per hour on average and 34% for novices, with little gain for the most experienced (Brynjolfsson, Li & Raymond). The gain is mostly bringing new people up to speed. Customer-emotional work and experts need a different design; Klarna and the Commonwealth Bank of Australia walked back AI-only support. See phase 3.
  4. On regulated platforms, partners gate customer-facing AI more than regulators do. (Narrowed from a broader hypothesis.) Card networks, carriers, sponsor banks and enterprise security reviews decide what an AI may do with money, messages and data. Regulators draw a few hard lines: consent for AI voice calls, disclosure of AI interaction under the EU AI Act from August 2, 2026, and human review of credit and KYC decisions. The failures that actually cost companies were quality failures: chatbots that invented policies. See phase 4.

Other ideas sit where they apply: juniors and learning (phase 2); "AI researches, a human decides" and managers deciding adoption (phase 3); buyers' data questions, docs agents load and outcome pricing (phase 4); legal human review by region (Where a human must decide).

Put together: measure before you roll anything out; fix the delivery system before you scale generation; let AI do the gathering and drafting while named people decide; put AI in front of customers last, with disclosure, evals and partner sign-off; judge all of it by outcomes, never by usage.

What's different on regulated platforms:

  • Separation of duties already applies to AI-written code. PCI DSS 6.2.3.1 says a manual code reviewer must not be the author (PCI DSS text, mirror); the EU's DORA change-management rules require independence between those who request or build a change and those who approve it (RTS 2024/1774 Art. 17). An engineer who prompts an agent and approves its pull request is, arguably, both.
  • Partners can block customer-facing AI with no regulator involved. A sponsor bank is accountable for a fintech's customer-facing tools under the 2023 US interagency guidance (Federal Reserve); an AI agent that texts is ordinary application-to-person traffic to a carrier.
  • Some decisions need a human, and the rule differs by region. EU onboarding and due-diligence decisions will need "meaningful human intervention" under the Anti-Money Laundering Regulation; US credit decisions need specific, accurate reasons rather than a human.
  • Someone else sets the dates: EU AI disclosure from August 2, 2026; software under EU product liability from December 9, 2026; EU high-risk duties from December 2, 2027.

The AI agent orchestration guide covers agent technology, security and liability; this playbook is about leading the change.

01Weeks 1-4

Baseline

You have 3-6 months of outcome metrics for how the team ships and works, a one-page AI policy (approved tools, data rules, agent permissions), an agreed position with compliance on AI-written changes, and a staggered rollout plan.

Measure first, because self-report can be wrong in sign. METR's 16 experienced maintainers worked 246 real issues in their own mature repositories; with AI they took 19% longer (confidence interval +2% to +39%), accepted fewer than 44% of generations and spent about 9% of their time reviewing and cleaning AI output. METR explicitly doesn't generalize the result, and its February 2026 follow-up was compromised by selection effects, which METR itself says make the data unreliable. The lesson that survives both is the gap between feeling and measurement. If you skip the baseline, every later claim rests on how people feel.

Pull the delivery numbers with engineering, not as a ranking. DORA's five metrics per team: change lead time, deployment frequency, failed-deployment recovery time, change fail rate and deployment rework rate (DORA guide, updated January 5, 2026). Split pull-request cycle time into coding, waiting for review, in review and waiting to deploy, so you'll see where AI moves the queue; record PR size and each senior's review load. Add incidents, escaped defects, security findings, secrets detected and code reverted or rewritten within 2-4 weeks. Compare each team with itself, never with others.

Baseline the internal workflows you might change. For the support queue, the KYB review queue, alert triage or incident write-ups: weekly volume, handle time, backlog, reopen or error rate, QA pass rate, escalations and cost per case. Four weeks is enough. You'll need these in phase 3, and they're impossible to rebuild afterwards.

Plan a staggered rollout, so later teams are your comparison group. The strongest company evidence comes from that design: field experiments with 4,867 developers at Microsoft, Accenture and a Fortune 100 company randomized access to Copilot and found 26% more completed tasks, with build success roughly unchanged and juniors gaining most (Cui, Demirer et al.). Pick two or three pilot teams and roll out to the rest in waves.

Find the AI already in use. 84% of developers in Stack Overflow's 2025 survey used or planned to use AI tools (InfoWorld), and more than half of the 10,600 workers in BCG's 2025 survey said they'd use unapproved tools if not given approved ones (BCG, analyst). Bans push it underground: Samsung banned generative AI on company devices from May 1, 2023, after engineers pasted source code and meeting notes into ChatGPT (Fortune). Approved tools good enough to beat the shadow ones are a security control.

Write a one-page AI policy. Only 25% of US organizations had communicated a clear AI strategy in Gallup's May 2026 data (Gallup), and a clear, communicated AI stance is one of DORA's seven AI capabilities. My draft contents, built from the notes:

  • Approved tools on enterprise terms: no training on your code, stated retention and region, SSO, admin audit logs, IP indemnity.
  • Data classes: what may enter a prompt or agent context. Card data, production credentials, KYC files and partner-confidential material never.
  • Secrets and dependencies: secret scanning before commit and in CI; lockfiles and allow-listed registries; verify any package an AI suggests.
  • Agent permissions: no production write access, dev and prod separated, destructive commands need a human, each agent its own identity.
  • Attribution: tag AI-assisted changes. The Linux kernel's policy (committed December 23, 2025) bars agents from certifying a contribution and asks for an Assisted-by tag, with the human submitter fully responsible (kernel.org).
  • Prototype line: PM and designer prototypes live in sandboxes; promotion to production goes through normal review.
  • IP: the US Copyright Office (January 29, 2025) says prompts alone don't make you an author, but human selection, arrangement or modification can (Copyright Office).

Agree the compliance position before the first agent PR. Take PCI DSS 6.2.3.1, SOC 2 CC8.1 and the EU DORA RTS Art. 17 to compliance and your assessor: who is the author when an engineer prompts an agent, what record per change they expect, and whether an AI reviewer counts as review. I found no auditor or regulator guidance specific to AI-written code, so my reading (the prompter is the author and can't be the sole approver) is a hypothesis to test with them.

Pick the scorecard now, before anyone has a result to defend. See Measuring it: adoption, flow, guardrails, value and people, reviewed monthly by you, the engineering lead and someone from compliance.

Sources: METR 2025 and paper, METR update, February 24, 2026, DORA capabilities model, Cui, Demirer et al. in Management Science.

Checklist

  • Pull 3-6 months of DORA's five metrics per team with engineering, and split PR cycle time into its four segments
  • Record PR size, senior review load, incidents, escaped defects, security findings, secrets and 2-4 week churn
  • Baseline volume, handle time, reopen rate, escalations and cost per case for each candidate internal workflow
  • Run a short developer-experience survey: cognitive load, flow, current AI use and trust in its output
  • Inventory the AI tools already in use, approved or not
  • Publish the one-page AI policy: tools, data classes, secrets, dependencies, agent permissions, attribution, prototype line
  • Agree with compliance and your assessor how AI-written changes meet author-not-approver and audit-trail rules
  • Confirm CI gates: tests, static analysis, dependency and secret scanning, a PR size limit, required human approval
  • Choose two or three pilot teams and a staggered rollout so later teams act as the comparison
  • Define the scorecard and who reviews it monthly

02Months 1-3

Ship with AI

Coding assistants and agents are in daily use on the pilot teams, starting with bounded work; review, tests, security and permissions keep pace; and at month 3 you decide with evidence whether to expand, fix the delivery system or stop specific uses.

Expect the bottleneck to move to review. In a vendor study of 10,000+ developers, high-adoption teams merged 98% more PRs while review time rose 91% and PR size 154%, with no company-level gain (Faros, July 2025, vendor study). At Google, in large AI-assisted migrations, 80% of the code in landed changes was AI-authored and total time fell by about half, but review and rollout, not generation, became the limit (Google, January 2025). DORA's 2025 survey of about 5,000 people linked AI to higher throughput and higher instability; its 2026 ROI report expects a dip before returns (InfoQ). Throughput only counts if change fail rate, rework and incidents hold.

Start agents on bounded work with an objective "done". The companies with the clearest end-to-end results picked work that CI already defines:

  • Stripe's in-house agents ("Minions") open more than 1,300 PRs a week, every one reviewed by a human, submitted only after CI passes; they're best on config changes, dependency upgrades and minor refactors (Stripe, InfoQ, February-March 2026). For a regulated platform, the strongest analogue.
  • Spotify's background agent had more than 1,500 merged PRs by November 2025 and saved 60-90% of time on migrations, on top of a fleet-management system that already automated about half of PRs before AI (Spotify).
  • Amazon claimed in August 2024 that AI-assisted Java upgrades saved about 4,500 developer-years (The Decoder).

My first list for a platform team: dependency and CVE upgrades, migrations, flaky-test fixes, lint and config sweeps, SDK and docs regeneration. High volume, objective, low blast radius.

Strengthen the verification loop before you scale generation. Spotify's agent works because of its loops: formatters and linters as tools, tests, a model judging the diff, then human review (Spotify, part 3). Add contract tests for partner APIs and coverage on payment, messaging and identity paths. Enforce a PR size limit: DORA has long tied small batches to stability, and it's one of its seven AI capabilities.

Treat the AI reviewer as a filter, not an approver. Microsoft runs an AI pull-request reviewer on more than 90% of its PRs, over 600,000 a month (Microsoft, July 14, 2025). That's coverage, not an outcome; a regulated change still needs an accountable human. On PCI- or DORA-scoped paths, the person who prompted the change never approves it. Record per change: ticket, which tool or agent and model, who approved, test and scan results, deploy record, rollback path, and the agent's identity and permissions at the time.

Security: the code is weaker and the tools are attack surface. In a Stanford trial, people with an AI assistant wrote less secure code and were more confident it was secure (Perry et al., 2023, older models). A vendor benchmark found raw model output chose an insecure option in 45% of tasks (Veracode, vendor study). In a USENIX Security 2025 study, 19.7% of code samples referenced at least one package that didn't exist, and 43% of the fake names came back on every re-prompt, so attackers can register them first (via CSA). Commits co-authored by one coding agent leaked secrets at about twice the baseline rate (GitGuardian, vendor study). The tools themselves have shipped prompt-injection flaws; the AI agent guide explains why there's no complete fix.

Give agents no production write access. In July 2025 Replit's agent deleted a production database during a declared code freeze; Replit then separated dev and prod (AI Incident Database). The Financial Times reported a 13-hour outage of an AWS service in one China region in December 2025 after Amazon's Kiro agent chose to delete and recreate an environment; Amazon blamed misconfigured access controls, not AI (GeekWire). Both readings point the same way. The control is old; the actor is new: least privilege and separation of duties, which your audits already require.

Protect senior reviewers. Seniors carry the review tax (the 19% in Xu et al.): rotate reviewers, set review service levels and watch review load weekly.

Keep juniors learning. Juniors gain most in output, but in an Anthropic trial 52 mostly junior engineers learning a new library with AI scored 17% lower on a mastery quiz, with the biggest gap in debugging, and weren't significantly faster (Anthropic). Payroll data show employment of 22-25-year-olds in AI-exposed occupations, software included, fell about 13% relative to others from late 2022, about 16% with firm controls (Stanford Digital Economy Lab); a counter-study finds no economy-wide disruption yet. My rule: no "accept without understanding" on payments, messaging or identity code; pair AI use with debugging and code-reading exercises; keep hiring juniors, because they're your future reviewers.

Budget per merged change, not per seat. Anthropic's Claude Code docs put enterprise spend around $13 per developer per active day, $150-250 a month (docs): a planning anchor for one agentic tool. Track cost per merged change.

Don't mistake a mandate for adoption. Shopify (April 2025) made AI use a baseline expectation and part of reviews, and asked teams to show AI can't do a job before requesting headcount (Digital Commerce 360); Coinbase (August 2025) fired engineers who hadn't onboarded to AI tools within a week (TechCrunch); Amazon set an 80% weekly usage target in November 2025, then dropped its usage leaderboard after it was gamed, moving to "normalized deployments". None has published an effect on outcomes. Usage can be a team-level input metric, paired with outcome metrics from day one and kept out of individual ratings. Otherwise you're on the theatre branch.

Decide at month 3, with the scorecard, expecting a dip in the first month or two.

Sources: DORA 2025, Spotify via InfoQ, AIDev dataset on review time, Faros 2026 (vendor study), Canaries follow-up, Amazon leaderboard (secondary press).

Checklist

  • Roll out approved tools to the pilot teams on enterprise terms, with SSO and admin logs, and track spend per merged change
  • Start agents on dependency and CVE upgrades, migrations, flaky tests, lint and config sweeps, SDK regeneration
  • Run agents in sandboxes with no production credentials, scoped short-lived tokens and their own identity
  • Tag every AI-assisted PR; on regulated paths, the person who prompted a change never approves it
  • Enforce PR size limits and add an AI pre-reviewer as a filter, not an approver
  • Protect senior review capacity: rotation, review service levels, weekly review-load numbers
  • Add contract tests for partner APIs and coverage on critical paths before scaling generation
  • Pair juniors' AI use with debugging and code-reading work; keep hiring juniors
  • Keep usage metrics at team level and out of individual reviews
  • Decide at month 3: expand, change the system, or stop specific uses

03Months 2-6

Rewire internal flows

One or two high-volume internal workflows are rebuilt so AI does the gathering and drafting, named people make the decisions the rules and the business require, oversight is measured to be real, and the change shows in an outcome metric.

Pick the workflow with a scorecard, not with enthusiasm. Use the seven criteria in Choosing the first workflows: volume, checkable output, decomposable, novice-heavy, written context, low emotional temperature, reversible. Every positive result in the research is high-volume.

Design for novices; leave experts room. In the support study, the AI assistant spread top performers' tacit knowledge to new agents (34% more resolutions per hour) and barely helped the most experienced. In a randomized trial of an AI assistant for Alibaba's Taobao after-sales agents, issue identification was 32% faster and ratings rose 5.3%, but objective quality didn't change and top performers had more delays and immediate customer retrials (Ni et al., 2026). On a partner-heavy platform with long onboarding (numbering rules, card disputes, cross-border corridors), that ramp time is where the money is. Don't force the assistant on your best people.

AI researches, a human decides. Stripe's compliance team splits each review into bite-sized sub-tasks; research agents fetch the evidence per sub-task; the human reviewer must answer each one; every agent action and rationale is logged for examiners. Median handling time fell 26% and reviewers rated the help above 96% (AWS and Stripe, June 26, 2026; co-published with the cloud vendor, no control group, no failure data). It turns a reviewer into an editor without moving accountability. Other published results fit the pattern: Google's security team had an AI draft incident summaries for people to accept, edit or discard and cut writing time 51% across 300 incidents (Google); Meta's root-cause tool puts the cause in its top five suggestions 42% of the time, a shortlist for engineers, not an answer (Meta). HSBC's 2-4x more true positives with 60% fewer alerts in transaction monitoring is machine-learning risk scoring, not a language model (HSBC, 2024); evidence for language models in compliance is still thin.

Route by risk and emotion before the AI touches the case. In a second Alibaba experiment, AI handled eligible chats with humans monitoring: 647 workers, 680,676 chats over 17 days in August 2024. Chats got faster, ratings on AI-handled chats fell substantially, and when frustrated customers were escalated after a bad AI exchange, human recovery mostly failed (Tuck School of Business, July 9, 2026). Escalation is not a rescue. Send angry, high-value or legally sensitive cases to people first.

Make oversight meaningful, and measure it. People agree with machines. When radiologists got wrong AI suggestions, accuracy fell from about 80% to under 20% for inexperienced readers and from 82% to 45.5% for those with 15+ years (RSNA, 2023). Developers approved about 93% of an AI coding tool's permission prompts in one lab's telemetry (from the site's AI agent research). And they lose the skill: endoscopists' detection rate in procedures without AI fell from 28.4% to 22.4% after three months with it (Budzyń et al., observational). So:

  • Reviewers see the sources, can override cheaply, and are measured on accuracy, not throughput.
  • Track approval rate and time per review. An approval rate near 100%, or reviews in seconds, triggers a sampling audit, because under EU and UK law it's evidence the decision is still solely automated (see Where a human must decide).
  • Seed known-wrong cases into the queue each month to measure automation bias.
  • Keep a no-AI path: some cases done unaided by juniors each week and reviewed by seniors.

Write the context down first. Anthropic found complex business deployments are limited by getting scattered organizational knowledge into context, not by the model (Anthropic, September 2025).

Product work: AI inside the frontier, people at the edge. Among 776 Procter & Gamble professionals, individuals with AI matched two-person teams without it (Dell'Acqua et al.). BCG consultants with AI did 12.2% more tasks, 25.1% faster, inside the "jagged frontier", but were 19 points less likely to be correct on a task outside it (Harvard Crimson). Spec drafting and synthesis are usually inside; novel judgment on thin data can be outside. Don't let synthetic users into a discovery readout: in one study, LLM "respondents" differed significantly from real survey data on 48% of coefficients and flipped the sign 32% of the time (Bisbee et al.). "Ask the data" works on curated data layers: on real enterprise schemas the best agent solved 21.3% of tasks in November 2024 (Spider 2.0; models have improved since).

Watch for gains that vanish into checking. Among 25,000 Danish workers in exposed jobs, chatbots saved about 3% of time with no measurable effect on earnings or hours, as new oversight tasks absorbed the gains (Humlum & Vestergaard, 2025).

Managers and training decide adoption. Gallup found employees whose manager supports AI use were 1.7 times as likely to use it weekly and 7.4 times as likely to say it helps them do their best work; BCG found five or more hours of training, in person with coaching, predicted regular use, while only 36% felt adequately trained (both correlational). Train your managers first, and count training in hours.

Say what happens to capacity before launch. The Commonwealth Bank of Australia reversed 45 redundancies tied to a voice bot after its union challenged the numbers (ACS Information Age, August 2025). Duolingo's CEO said an "AI-first" memo hadn't given enough context (Fortune). Tell the team whether freed time goes to backlog, redeployment or slower hiring, and involve worker representatives where they exist. In the EU, AI that allocates work to or evaluates employees is high-risk from December 2, 2027 (AI Act Annex III).

Sources: Brynjolfsson, Li & Raymond, Alibaba agentic study, Humlum & Vestergaard via NBER, Gallup, BCG (analyst), Microsoft Security Copilot trials (vendor-run).

Checklist

  • Score candidate workflows on the seven criteria and pick one or two scoring at least 10 of 14
  • Map every decision in the workflow to its rule: AI may draft, AI may recommend, a human must decide, a human must explain
  • Decompose the case into sub-tasks; AI researches each; the reviewer answers each; log everything
  • Build an eval set of 50-200 real past cases with the people who do the work, and pass it before go-live
  • Route by risk and sentiment before the AI touches a case
  • Measure reviewers on accuracy; track approval rate and time per review; seed known-wrong cases monthly
  • Keep some cases done unaided by juniors and reviewed by seniors
  • Train managers first and count training in hours
  • Tell the team what happens to freed capacity before launch
  • Set the outcome metric and the rollback trigger up front; review at 6 and 12 weeks

04Months 4-12

Put AI in front of customers

Customers can use your platform through their own agents, your first customer-facing assistant is live with disclosure, evals, a fast human path and partner and regulator sign-off, and buyers' security reviews and pricing questions have written answers.

Sequence by risk, lowest first. My default order: agent-ready developer experience (months 4-5; low exposure, and it forces API quality work); a read-only assistant over docs and account data (months 5-8); a copilot that drafts changes for a human to apply (months 8-12, after partner sign-off); autonomous actions on money, end-user messages or decisions about people only inside partner rules, with human review where the law requires it.

Get partners' sign-off criteria before you build. Partners set the gates where money and messages move:

  • Card networks: Mastercard Agent Pay (April 29, 2025), Visa Intelligent Commerce (April 30, 2025) and Visa's Trusted Agent Protocol (October 14, 2025) route agent payments through tokens and registration (PYMNTS). Mastercard has published agentic commerce rules (reported January 2026); I haven't read the text, and who bears a chargeback when an agent buys the wrong thing is unsettled.
  • Carriers: US business messaging needs 10DLC brand and campaign registration; T-Mobile's code of conduct fines senders, from $500 to $2,000 per violation, and blocks traffic (carrier codes summary). No carrier rule specific to AI-generated messages turned up; AI texts are simply business traffic. See the telco guide.
  • Sponsor banks: accountable for customer-facing tools under the 2023 guidance, and their model-risk teams may still review generative AI even though the April 2026 US model-risk guidance excludes it (Orrick).
  • Enterprise buyers: 80% of 1,169 B2B decision-makers reported stricter AI evaluations by security, legal and compliance (G2, April 2025, vendor survey).

Know the regulators' few hard lines. Disclosure: under EU AI Act Article 50, from August 2, 2026, people must be told they're interacting with AI; the Commission's July 2026 guidelines say at the very first interaction, and a generic "assistant" label or text buried in terms won't do; fines reach €15M or 3% of global turnover (Stephenson Harwood). A CPaaS or support platform is often the provider and its customer the deployer, so make disclosure the default and hard for customers to remove. The US has no federal chatbot-disclosure law, only state rules such as Utah's (FPF) and Maine's. Voice: the FCC ruled on February 8, 2024 that AI voices are "artificial" under the TCPA, so outbound AI calls need prior express consent (Cooley); Lingo Telecom paid $1M for carrying deepfake robocalls (Perkins Coie). Decisions: no AI-only adverse decisions on credit, KYC or access (see Where a human must decide).

Budget most of the effort for quality, because that's what has hurt companies. Air Canada's and Cursor's bots invented policies (see Named examples), and Klarna's CEO said a cost focus had produced lower quality (Mind the Product, May 2025). Commercial liability arrives before legal liability. So:

  • Grounding: answers about policy, prices, terms and limits come only from approved sources; otherwise "I don't know" and a handoff.
  • A fast human path: the US consumer-finance regulator's 2023 chatbot report warned against "doom loops" and expected a route to a human (summary); whether it still stands behind that after its 2025 guidance withdrawals is unclear, but your customers will. Measure time to a human.
  • Evals from real tickets: 20-50 tasks to start, graded on the account's end state (refund issued once, rule deployed correctly), reported as reliability across repeated runs; on one benchmark, customer-service agents passed all eight tries on fewer than 25% of tasks (2024, from the site's AI agent research). Pin model snapshots and re-run evals on every model, prompt or tool change.
  • Write actions: scoped short-lived credentials, approval for irreversible steps, spend caps, idempotency, a kill switch and a full audit log.

Every number in marketing needs a dated eval behind it. The FTC fined accessiBe $1M over a B2B claim that its AI made any site accessible (TechCrunch), challenged Workado's 98% accuracy claim against 53% in testing (14 News), and sued Air AI over AI marketed as replacing sales reps (DLA Piper).

Have the buyer's security answers written before sales needs them. Buyers moved from "can we opt out of AI?" in 2024 to asking you to prove their data never trains your models (SecurityPal, 2026, vendor report; no percentages). Zoom (August 2023) and Slack (May 2024) both faced backlash over terms read as allowing training on customer content (TechCrunch). The default and the opt-out mechanism are the product. Ship training off by default, an admin toggle, a public AI data-use page and a sub-processor list naming your model providers. Prepare answers to the Cloud Security Alliance's AI questionnaire (AI-CAIQ) and a view on ISO/IEC 42001 (CSA). Don't promise model-change notice longer than your labs give you: from 60 days to six months, depending on the lab.

Build developer experience for agents that actually load docs. Your customers increasingly reach your API through their own agents. Vercel tested coding agents on Next.js APIs missing from training data: 53% passed with no help, 79% with an explicitly invoked skill, and 100% with a compressed docs index in the project's AGENTS.md file; in 56% of cases the agent never invoked the skill (Vercel, January 2026, vendor study, task count not given). Next.js now bundles its docs as Markdown inside the package (Next.js 16.2). By contrast, server logs across 137,000 domains showed 97% of llms.txt files got no requests at all in May 2026 (PPC Land). llms.txt is cheap to ship; don't count it as a strategy. My agent-ready list, following Stripe's Markdown docs and official MCP server (Stripe) and Anthropic's tool guidance (Anthropic): a complete OpenAPI spec with examples and error semantics; Markdown docs, bundled with SDKs where you can; a remote MCP server with OAuth, scopes and read-only by default; a sandbox, idempotency keys and dry runs; short-lived keys for agents; and evals of common integration tasks on every docs and API release. Watch the cost: in Twilio's small test of its alpha MCP server (April 2025), tasks were faster but cost 27.5% more (Twilio, vendor test, three tasks).

Price on a unit you can audit. Intercom charges $0.99 per outcome, where an outcome includes a customer requesting no further help (Intercom); Zendesk introduced per-resolution pricing in August 2024 (Zendesk); HubSpot moved to about $0.50 per resolved conversation in April 2026 (search result). "Resolved" is the vendor's definition, and disputes move to the definition. My default: fold AI into your existing per-transaction or per-message price where it lowers your cost to serve, charge usage where it consumes compute per call, and use outcome pricing only where the outcome is objective and auditable. Put tokens, retries and the human fallback in the margin model.

Check liability and insurance before launch. If your AI screens, scores or decides for your customer, you may share liability (Mobley v. Workday); the revised EU Product Liability Directive treats software, AI included, as a product from December 9, 2026 (both from the site's AI agent research). Check your E&O and cyber cover for AI exclusions.

Sources: Visa Intelligent Commerce, FCC NPRM on AI-call disclosure, not final, Maine LD 1727, Zoom via ISACA, Stripe MCP, 2026 state chatbot laws.

Checklist

  • Pick the first use case by risk tier and write its intended use and "never does" list
  • Ship agent-ready docs, a clean OpenAPI spec and a read-only MCP server with OAuth before any customer-facing agent
  • Show sponsor bank, acquirer, carrier and key customers the design before build, and get written sign-off criteria
  • Turn disclosure on by default at the first interaction, in text and voice; customers can't silently remove it
  • Cover AI in consent language for outbound voice and text; register the campaign; honour opt-outs
  • Publish the AI data-use page and sub-processor list; training off by default with an admin toggle
  • Prepare the security pack: AI questionnaire answers, ISO/IEC 42001 plan, prompt-injection threat model, red-team results
  • Run evals from 20-50 real tickets, graded on end state across repeated runs; pin model snapshots
  • Keep a visible, fast human path and measure time to a human
  • Review every marketing number against a dated eval; price on an auditable unit

Measuring it

Outcome metrics prove the change; usage metrics only show activity. Goodhart's law applies at once: when a measure becomes a target, it stops being a good measure. A five-layer scorecard (my synthesis of DORA and practitioner frameworks), reviewed monthly, at team level, never in individual ratings:

LayerMeasuresWatch for
AdoptionWeekly users, AI-tagged PRs, agent PRs merged, spendAn input; never a target
FlowDORA throughput; PR cycle segmentsReview wait growing
GuardrailsChange fail, rework, incidents, security findings, churnMust hold for flow to count
ValueRoadmap items delivered, backlog burned, cost per changeThe question the CEO asked
PeopleDeveloper experience, trust, junior skill checksSeniors' review load

For internal flows and customers, the same logic: quality next to speed. Handle time improved in every support study above while objective quality often didn't. Measure reopens, retrials, escalations, wrong answers on audited samples, complaints, refunds issued by AI and cost per resolved case.

Vanity metricWhy it misleadsUse instead
% of code written by AIVolume; definitions varyLead time with change fail rate
Suggestion acceptance rateAn autocomplete UX numberRework within 2-4 weeks
Seats, prompts, tokensActivity; gamed by leaderboardsCost per merged change
Handle timeSpeed without qualityReopens and retrials
"Resolution" rateThe vendor's billing eventAudited correct end state
FTE equivalentNot a cost lineThe cost line itself

The labs' code-share figures are claims, not benchmarks. Google's CEO said 75% of new code was AI-generated in April 2026, counting AI-suggested code that engineers accept (Semafor). Microsoft's CEO put 20-30% of code in its repositories as AI-written in April 2025 (TechRadar). Anthropic reported that Claude authored more than 80% of code merged in May 2026 (Fortune). OpenAI reported that 97.9% of its employees used its Codex agent by June 2026, a usage figure rather than a code share (OpenAI). None of them publishes change fail rates or incidents next to these numbers. One analyst estimated that a code share near 90% could come with a productivity gain well under 2x (Greenblatt, Redwood). Volume is not value; the missing failure data is the lesson.

Against gaming: pair every throughput metric with a stability metric, publish definitions, sample PRs and cases by hand, and watch leading indicators (review queue age, PR size, approval rate) weekly and lagging ones (incidents, complaints) monthly.

Sources: DORA guide, DX measurement framework (vendor; excludes acceptance rate and lines of code), Klarna costs via CX Dive.

Where a human must decide

Where the law literally requires a person differs by region, and it's narrower in the US than most teams assume. A summary from the research, as of October 2, 2026; this is not legal advice, so take it to counsel.

Credit, incl. small business

US
Specific, accurate reasons (Reg B)
EU
Person: high-risk from 2027-12-02
UK
Safeguards (from 2026-02-05)

Onboarding, due diligence

US
Sound AML programme, no rule per case
EU
Meaningful human intervention (AMLR)
UK
Safeguards

Suspicious activity reports

US
Institution decides and files
EU
Institution decides
UK
Institution decides

Fraud detection

US
No specific rule
EU
Excluded from AI Act high-risk
UK
Safeguards if significant

Hiring, evaluating staff

US
Notice in Illinois from 2026-01-01
EU
High-risk from 2027-12-02
UK
Safeguards

Any significant automated decision

US
State laws (Colorado from 2027-01-01)
EU
GDPR Art. 22: human or exception
UK
Safeguards, not a ban

Customer talks to an AI

US
State disclosure laws
EU
Disclose from 2026-08-02
UK
No AI-specific rule

Outbound AI voice call

US
Prior express consent (TCPA)
EU
Disclose (Art. 50)
UK
No AI-specific rule

United States: reasons, not a reviewer. Under Reg B, an adverse-action notice must give specific principal reasons; "internal standards" or a failed score isn't enough, and it applies to business credit (Reg B §1002.9). If a model contributes to declining a merchant or a small-business credit line, it must be able to produce those reasons. The CFPB's circulars on complex algorithms were withdrawn on May 12, 2025 "pending review"; the statute and Reg B didn't change (CFPB). For AML, the Bank Secrecy Act requires a programme "reasonably designed", not a human per alert; FinCEN's October 2025 FAQs eased SAR documentation (Morrison Foerster), and a 2018 joint statement says pilots that find more suspicious activity won't by themselves prove existing processes deficient (FinCEN).

European Union: human intervention. GDPR Article 22 gives a right not to be subject to solely automated decisions with legal or similar effect; the SCHUFA judgment (December 2023) says a score that plays a determining role counts (Bird & Bird). The Anti-Money Laundering Regulation's Article 76(5) requires meaningful human intervention in automated decisions to accept, refuse or change due diligence on a customer, and covers companies too, so an EU platform's KYB decisions are in scope (Article 76); the regulation is reported to apply from July 10, 2027, a date to confirm. The AI Act makes credit scoring of people and AI that hires or evaluates staff high-risk from December 2, 2027, with oversight that guards against automation bias (Article 14).

United Kingdom: safeguards. From February 5, 2026, the UK allows significant automated decisions with safeguards, and "no meaningful human involvement" defines solely automated (Data (Use and Access) Act, s.80). The FCA has no AI-specific rules; Consumer Duty and the senior-manager regime apply.

"Meaningful" is the word that matters in all three. A person who rubber-stamps a score doesn't take the decision out of "solely automated". Approval rates near 100% and seconds per case are evidence against you. Australia's Robodebt scheme raised A$1.73B of unlawful debts against more than 400,000 people while humans followed the flag (Royal Commission).

Choosing the first workflows

Score each candidate 0-2 on seven criteria and pick one or two scoring at least 10 of 14 (my scorecard, built from the evidence above; a hypothesis to test).

CriterionScores 2 whenEvidence
Volume and repetitionHundreds of cases a week with a familiar shapeEvery positive result
Checkable outputA right answer exists and checks fastSupport studies
DecomposableThe case splits into sub-tasks to researchStripe compliance
Novice-heavy, long rampNew staff take months to get good+34% for novices
Context written downPolicies and past cases are documented, permissionedAnthropic; Spider 2.0
Low emotional temperatureAnalyst-facing, not angry-customer-facingAlibaba agentic trial
Reversible, containedErrors are caught before a customer or regulatorScore 0 if AI acts ungated

Good first candidates: KYB and due-diligence research, alert triage drafts, incident summaries, support drafting for internal agents, partner-integration troubleshooting.

Red flags: don't pick these first.

  • EU onboarding or due-diligence refusals without a design for meaningful human intervention
  • Any credit decision on people or small businesses without specific reason codes
  • Evaluating or monitoring employees in the EU
  • Customer-emotional queues where the AI answers first
  • A workflow whose owner won't commit reviewer time
  • A headcount target set before the eval results

Named examples

Illustrations, not proof: mostly companies' own accounts, with the caveats each carries.

Stripe: agents on bounded code, research for compliance

What happened: Stripe's coding agents open more than 1,300 PRs a week, all reviewed by a human and submitted only after CI passes (February-March 2026). Separately, its compliance team uses research agents per sub-task while reviewers answer each one, cutting median handling time 26% (June 2026). Lesson: at a payments company, the gains came from bounded work with an objective "done" and from keeping the decision human. Caveat: company accounts, the compliance one co-published with the cloud vendor; no failure data.

Sources: Stripe Minions, AWS and Stripe.

Spotify and Google: migrations where CI defines "done"

What happened: Spotify's background agent had more than 1,500 merged PRs by November 2025 and saved 60-90% of time on migrations; Google reported 80% of the code in landed migration changes was AI-authored, with about half the total time saved, review included (January 2025). Lesson: both built on existing automation and verification loops; in both, review and rollout set the pace.

Sources: Spotify, Google.

Amazon: a usage leaderboard, gamed

What happened: after an 80% weekly usage target (November 2025), Amazon dropped an internal AI-usage leaderboard when staff pointed agents at pointless tasks to climb it, moving to "normalized deployments" (reported May 2026). Lesson: a usage target becomes the work; this is the theatre branch. Caveat: secondary press reporting of a Financial Times story.

Sources: The Decoder.

Klarna: from AI-first support back to an easy human

What happened: in February 2024 Klarna said its assistant handled two-thirds of chats in its first month, the work of 700 agents. In May 2025 its CEO said investing in human support quality was the way forward and began hiring human agents again. By Q3 2025 it reported AI doing the work of 853 agents and $60M saved, while customer service and operations costs were $50M against $42M a year earlier. Lesson: Klarna didn't abandon AI; it abandoned AI-only. Output counts and the cost line disagreed.

Sources: Klarna, Maginative, CX Dive.

Commonwealth Bank of Australia: the claim the union disproved

What happened: the bank cut 45 customer-service roles, citing a voice bot that reduced calls by 2,000 a week. The Finance Sector Union showed volumes were rising and staff were on overtime; in August 2025 the bank reversed the cuts and apologized. Lesson: a capacity claim without a baseline is a liability; the outcome metric has to be real before the headcount decision.

Sources: Bloomberg, ACS Information Age.

Alibaba (Taobao): assist versus act

What happened: two randomized experiments in after-sales chat. As an assistant, AI helped weaker agents most. As an agent with humans monitoring, it made chats faster but lowered ratings, and humans couldn't recover customers escalated after a bad exchange. Lesson: AI assisting a person and AI acting with a person watching are different designs with different results.

Sources: Ni et al., Tuck School of Business.

Air Canada and Cursor: invented policies

What happened: Air Canada's chatbot invented a refund policy and a tribunal held the airline liable (February 2024). Cursor's support bot invented a one-device rule; customers cancelled and the company began labelling AI replies (April 2025). Lesson: you own what your bot says; ground policy answers, label AI, and review policy claims.

Sources: The Register; Air Canada from the site's AI agent research.

Zoom and Slack: the training default

What happened: Zoom's terms (August 2023) were read as allowing AI training on customer content and were reversed within days. Slack's principles (May 2024) allowed training global models on customer messages by default, with opt-out by email; Slack clarified it doesn't train its LLMs on customer data. Lesson: for B2B buyers, the default and the opt-out mechanism are the product.

Sources: ISACA, TechCrunch.

Failure modes

FailureEarly signCase or evidence
Rollout without a baselineSuccess measured by surveyMETR
Usage theatreLeaderboards; AI usage in individual reviewsAmazon
Review collapsePR size and review wait climbing; seniors' output downFaros; Xu et al.
Agent with too much powerProduction credentials in an agent's reachReplit; reported AWS outage
Prompter self-approvesNo record of which agent did whatPCI 6.2.3.1; DORA RTS
Rubber-stamp oversightApproval near 100%, reviews in secondsRadiology; SCHUFA
Handle time up, quality flatOnly speed on the dashboardAlibaba
AI-only support, then reversalHeadcount target before eval resultsKlarna; CBA
Hollowed-out junior pipelineNo junior hires; juniors only edit draftsCanaries; Anthropic trial
Invented policy to customersNo grounding, no labelAir Canada; Cursor
Training-default surpriseOpt-out by emailZoom; Slack
Partner surpriseAI voice or texts launched without carrier reviewLingo; T-Mobile fines
Docs theatrellms.txt shipped; API spec and errors unchangedAhrefs logs

The most common failure is buying tools while the system stays the same. DORA's "amplifier" finding cuts both ways: good tests, small batches and clear ownership get better with AI, and weak ones get worse faster. Every failure above is a review or measurement failure before it's a technology failure.

Sources: cases and studies as cited above; GitClear via DevClass (vendor study) on rising code churn.

False friends

TermWhat you'd assumeWhat it means here
"% of code written by AI"ProductivityVolume, defined differently by each company
AI nativeEveryone uses AI toolsWork and decisions reviewed differently, proven by outcomes
Human in the loopCompliant oversightOnly if the human is competent, informed and can say no
FasterMeasured speedOften felt speed; METR found a 39-point gap
ResolutionThe customer's problem fixedThe vendor's billing event, which can include silence
llms.txtAgents read your docsA file format; 97% got no requests
MCP supportProduction-readyAnything from an alpha local server to remote with OAuth
No training on your dataCovers everythingMay cover the model provider, not your own models
AI code reviewThe review auditors expectA filter; an accountable human still approves
AgentOne thingAutocomplete, chat, in-editor agent or background agent
DORAOne thingDelivery metrics, or the EU resilience act; both apply
The law requires a human (US)A reviewer per decisionMostly specific reasons and a sound programme
FTE equivalentCost savedKlarna's count rose while its cost line rose too

On DORA. On a payments or banking platform, the two meet in change management: Google's DevOps Research and Assessment metrics measure it, and the EU Digital Operational Resilience Act, applying since January 17, 2025, regulates it. In a meeting with engineering and compliance, say which one you mean.

Ask an expert

For leaders who have done this, and the specialists around them:

  1. Which delivery metric moved first after AI rollout, and which one got worse before it got better?
  2. How did you change code review once PR volume rose: size limits, AI pre-review, reviewer rotation, service levels?
  3. For PCI- or DORA-scoped services, how do you meet "reviewer is not the author" when an engineer prompts an agent, and what did your assessor accept?
  4. What permissions do your coding agents have, and what's the one action they're never allowed to take?
  5. How do you keep usage metrics from becoming targets, and are AI metrics anywhere near individual reviews?
  6. What's your reviewers' approval rate on AI recommendations, and have you seeded known-wrong cases to test them?
  7. After AI handled a share of tickets, which metric moved first in the wrong direction: reopens, escalations, partner complaints or satisfaction?
  8. What does your sponsor bank, acquirer or carrier require before AI talks to their customers or end users?
  9. Which AI questions actually kill deals in enterprise security reviews: training, sub-processors, model-change notice, prompt-injection testing or ISO/IEC 42001?
  10. How do you train juniors when AI drafts most of the work, and does it show in their judgment?

Self-check: 15 questions

  1. Why does the baseline come before any rollout, and what did METR's developers feel versus measure? Answer
  2. What do the four ideas say, and how strong is the evidence for each? Answer
  3. Which delivery metrics do you capture in the baseline, and why split PR cycle time? Answer
  4. What goes in the one-page AI policy? Answer
  5. Why does the bottleneck move to review, and what work do agents start on? Answer
  6. Who must never approve an AI-written change on a regulated path, and why? Answer
  7. When does a usage metric turn into theatre? Answer
  8. Who gains most from AI in support, and what does that mean for your design? Answer
  9. What makes oversight meaningful, and what approval rate should worry you? Answer
  10. What are the seven criteria for choosing a first workflow, and three red flags? Answer
  11. Where does a human legally have to decide in the EU, and what does US credit law require instead? Answer
  12. Which partners gate customer-facing AI on a regulated platform? Answer
  13. When does EU disclosure of AI interaction start, and what counts as disclosure? Answer
  14. Why don't the labs' code-share figures prove productivity? Answer
  15. What do "resolution", "llms.txt" and "human in the loop" really mean? Answer

Sources

Read on or before October 2, 2026 unless dated; "search result" means seen only as a snippet or summary. Vendor studies are marked in the text.

Controlled and quasi-experimental studies

Surveys and analyst figures

Company reports and claims

Regulators, law and standards

The site's own research

Playbooks are one practitioner's operating notes, not management, legal or financial advice. Adapt them to your company before you act on them.