The Platform PM
Field Guide

AI agents: orchestration platforms

The layer that runs AI agents in production: orchestration and control, context and memory, secure tools, evals and governance. What it costs, what breaks, and who is accountable when an agent acts.

Last reviewed October 2026

The industry on one page

The parties in one agent run. The platform runs the loop: it retrieves context (data and memory), passes it to a model from any lab, and executes the tool calls the model asks for. Tools are where the agent acts on real systems, and where an attacker who steers the model gets in.

An AI agent is software that calls a model in a loop. The model reads a task and some context, then answers or asks for a tool: look up a customer, search a policy, issue a refund. The software runs the tool, adds the result and asks again, until the task is done or something stops it. An agent orchestration platform runs that loop for a company. It retrieves context (data, memory) and asks a model from any lab, passing the context in; the model's tool call goes to tools (APIs, MCP servers), which act on the systems where actions land, from a CRM to a payments ledger. The tools are the part to watch: they're where the agent acts on real systems, and where an attacker who steers the model gets in. Every tool is a door.

Adoption is real and hard to measure. In KPMG's survey of 314 leaders at US companies above $1B in revenue (July to August 2026), 62% were building, deploying or developing agents, up from 53% a quarter earlier (KPMG); in McKinsey's November 2025 survey, 23% were scaling an agentic system somewhere and 39% saw any enterprise EBIT impact (McKinsey). Gartner expects more than 40% of agentic projects to be cancelled by the end of 2027 (Gartner, June 2025; methodology not public). The clearest revenue sits with systems of record: Salesforce reported Agentforce annual recurring revenue above $1.5B on August 26, 2026, after widening what the figure counts (Salesforce).

I organize the guide around four pillars, the four jobs the platform does on every run:

PillarWhat it coversWhere to find it
1. Agent orchestration & controlThe loop, workflows vs agents, multi-agent, state, approvals, caps, retriesMap steps 1, 3, 5, 6; primitives 2 and 9; Which approach?
2. Context, data & memoryContext windows, retrieval with permissions, memory, caching, retentionMap steps 2 and 7; primitives 1, 3 and 8; Money flows
3. Secure execution & toolingTool calling, MCP and A2A, sandboxes, agent identity, prompt injectionMap step 4; primitives 1, 4, 6 and 7; Prompt injection
4. Agentic evals & governanceEvals, tracing, model retirements, guardrails, regulation, liabilityMap step 8; primitives 3, 5, 8 and 10; Regulation in practice

Four ideas organize what I've learned so far, all hypotheses I'm testing with experts.

  1. The platform sells control more than intelligence. (Mostly supported for back-office and customer-facing agents; not for coding and research agents.) In a 2025 study of 306 practitioners and 20 case studies, 68% of deployed agents ran fewer than 10 steps before a human stepped in, 47% fewer than 5, and 16 of the 20 cases used structured workflows rather than open-ended planning (Measuring Agents in Production). Anthropic, OpenAI, Microsoft and Google Cloud all advise starting with the simplest design that works. Against: coding agents need a median 41 to 58 turns on SWE-bench Verified (turn-control study); METR measured the task length a 2025 frontier model finishes half the time at about 110 minutes, doubling roughly every seven months since 2019 (METR); and all three API labs sell background execution for long runs.
  2. Quality is context times model. (Supported as a multiplication; "context beats model" isn't.) With the same model, one terminal-task benchmark moved 20 points or more across harnesses, a gap the authors compare to a model generation (position paper, July 2026); at a fixed harness, swapping models still moves results 2 to 3 times. Risk lives on the context side: enterprise search indexes copy permissions at sync time, so a revoked user can keep seeing files until the next sync, and in one study, poisoning under 0.1% of an agent's memory entries steered it more than 80% of the time (AgentPoison).
  3. Every tool is a door. (Strong.) Nobody claims prompt injection is solved: OWASP doubts fool-proof prevention exists, the UK's NCSC says it may never be fully mitigated, Meta calls it a fundamental weakness of all LLMs, OpenAI expects it never to be fully solved, Anthropic says no browser agent is immune, and Google DeepMind relies on continuous adaptive testing. Adaptive attacks pushed 12 published defenses above 90% success (attack study, October 2025). So the security model is identity, least privilege, egress control and policy outside the model, and the clouds' published responsibility splits leave the authorization of actions with the customer. Against: model defenses do lower attack rates, and least privilege can't stop an injected agent misusing a permission it holds.
  4. Evals are the unit of change, more than the unit of purchase. (Partly supported.) Models retire on published clocks: at least 6 months' notice for OpenAI's GA models, 60 days at Anthropic and Microsoft, 2 weeks for Google's previews, 45 days for some AWS models, and no formal policy at Mistral. Every forced swap needs a regression suite, and public benchmarks can't stand in: OpenAI stopped reporting SWE-bench Verified in February 2026 over flawed tests and contamination, and weak graders let agentic benchmarks overstate results by up to about 40% (checklist study). Against evals as the purchase: clouds meter evaluators at fractions of a cent, OpenAI closes its hosted Evals product on November 30, 2026, and in a framework vendor's survey 89% of teams had observability but only 52% ran offline evals (survey). Buyers seem to follow where their data already lives.

Put together, the product is a control loop around a model someone else owns: it decides what the model sees, what it may touch and what changed since the last eval, and it's paid per token, per action or per outcome. Two more hypotheses sit further down: the retirement clock as a hidden operating cost (Effective dating), and accountability set by courts, contracts and insurers more than by AI statutes (Regulation in practice).

Three things make this harder than the B2B platforms you know:

  • The same input doesn't give the same output. A model that solved about 61% of retail support tasks on one try solved under 25% on all of eight tries (τ-bench, 2024). Tests become statistics.
  • Your data is also your instructions. A web page, an email, a tool description or a stored memory can tell the agent what to do, and there's no equivalent of a parameterized query to keep data and commands apart.
  • Your core dependency changes on someone else's calendar. Models retire on 45 days' to 6 months' notice, model aliases move without a release on your side, and one lab's newer tokenizer produces about 30% more tokens for the same text.

How card networks treat purchases made by AI agents is in B2B payments: spend management.

The agent map

One agent run. The platform builds the context, the model decides whether to answer or call a tool, and each tool result goes back to the model for another lap until it answers. Every lap re-sends the growing context, so cost climbs with each step; step and cost caps stop runaway loops, and risky actions pause for a human.

Five states on the main path (triggered, context built, model decides, tool runs, done), two exceptions (stopped by a step or cost cap, paused because it needs a human), and one fact that shapes the product: every lap re-sends the growing context, so each lap costs more than the last.

One run: a support agent issues a refund (pillar in brackets)

  1. Triggered (orchestration): a chat message, API call, webhook, schedule or another agent's handoff. The platform stamps a run ID, the thread, the end user, the agent's version and the tenant. Breaks: a rate or spend limit returns HTTP 429; a duplicate webhook without an idempotency key starts a second run.
  2. Context built (context and memory): system prompt, tool definitions, recalled memory, documents retrieved as the end user and filtered by their permissions, and the conversation so far, compacted if it's long. Breaks: the index holds last week's permissions; a long context buries the fact the model needs, since accuracy falls as input grows even inside the advertised window (length study); compaction drops a detail; a reordered tool list breaks the prompt cache and the step costs several times more.
  3. Model decides (orchestration): a final answer, tool calls or a handoff to another agent. Breaks: an error triggers a fallback to another model that behaves differently; malformed arguments; parallel calls emitted before any check.
  4. Tool runs (secure execution): the executor, not the model, checks the call against an allow-list and a policy outside the model, fetches a credential, runs the tool and appends the result. Breaks: the ERP times out after committing, and a retry without an idempotency key refunds twice; the result carries instructions from an injected email; a valid call does the wrong thing.
  5. Paused (orchestration, secure execution): the refund is above the approval threshold, so the run saves its state, releases compute and waits for a person. Breaks: nobody approves; the approver clicks yes without reading.
  6. Back to model decides for another lap until the model answers, or Stopped when a guard fires: maximum turns, a token or cost budget, a timeout. Breaks: the agent repeats a step, the largest single failure (17.1%) in a study of 1,600 multi-agent traces (MAST).
  7. Memory written (context and memory): a model decides what to remember, for whom and for how long. Breaks: it saves an instruction planted in a ticket, or personal data you'll later have to erase.
  8. Done (evals and governance): the result goes back, the trace closes, and usage (input, cached, output and reasoning tokens, tool fees, runtime) is metered and charged to a tenant. Breaks: the trace says "refund issued" and the payments system disagrees; the trace was sampled, or holds no content.

A long-running task: onboarding a vendor

  1. Triggered by an API call; the engine saves a checkpoint after every step.
  2. Paused three minutes in, when the agent proposes changing the vendor's bank details. State is saved and compute released: an idle managed session isn't billed on Anthropic's platform, and AWS tears down an idle microVM after 15 minutes.
  3. Resumed two hours later, when an approver says yes in another tool. Resume means replay: the step restarts from its beginning, so anything it did before the pause must be safe to repeat.
  4. Retried when the ERP call times out after the ERP committed; the retry carries the same idempotency key, so the ERP returns the original result instead of writing twice.
  5. Recovered after a worker crash, from the last checkpoint, not from scratch; a code deploy in between mustn't change how this run behaves.

Breaks: the provider's stored state expires (30 days on OpenAI's Responses API, 55 days on Google's paid Interactions API, until deleted on Anthropic's managed sessions); the model retires mid-run; the session hits its maximum lifetime (8 hours by default on AWS); compaction drops the detail the approver asked about.

Cost grows roughly quadratically with turns, because the whole history is re-sent each turn; capping coding agents at the 75th percentile of turns cut cost 24% to 68% without hurting results (turn-control study). Every lap re-sends the growing context, so each step costs more than the last: caps stop runaway loops, risky actions pause for a human, and "resume" means replaying from the last checkpoint.

Sources (as of Oct 2026): OpenAI's agent guide and SDK, AWS harness and limits, Anthropic pricing, Google Interactions, OpenAI data controls, a graph framework's interrupt docs.

Cheat sheet: ways to build and run agents

Open-source framework

Who runs state (1)
You: your database or workflow engine
Models
Any
Context and memory (2)
You build retrieval and memory
Tools and security (3)
You build executor, policy, sandbox
Evals and tracing (4)
The maintainer's hosted add-ons
Lock-in
Low in code; high build cost
Pricing shape
Free code; paid hosted tracing and evals

A frontier lab's agent platform

Who runs state (1)
The lab: stored conversations, sessions
Models
That lab's
Context and memory (2)
Memory tools, compaction, file stores
Tools and security (3)
Hosted tools, sandboxes, MCP client, vaults
Evals and tracing (4)
The lab's traces and graders
Lock-in
High: state, caching, retirements
Pricing shape
Tokens plus tool fees; session-hours at Anthropic

A cloud's agent service

Who runs state (1)
The cloud's managed runtime
Models
Several labs, on the cloud's dates
Context and memory (2)
Managed memory, permission-aware search
Tools and security (3)
Workload identity, gateway, policy engine, microVMs
Evals and tracing (4)
Evaluators on production traces
Lock-in
Identity, network, commitments
Pricing shape
Each part metered

Independent orchestration platform

Who runs state (1)
The platform
Models
Any, routed
Context and memory (2)
Connectors
Tools and security (3)
Its gateway and policy features
Evals and tracing (4)
Its own, ideally exportable
Lock-in
Evals, traces, connectors
Pricing shape
Licence or seats, plus your tokens or bundled tokens

SaaS system of record with agents

Who runs state (1)
The SaaS app
Models
The vendor's choice
Context and memory (2)
Its own data and permissions
Tools and security (3)
The app's permission model
Evals and tracing (4)
The vendor's
Lock-in
Data and workflow
Pricing shape
Per action, conversation, seat or outcome

Sources (as of Oct 2026): AWS AgentCore pricing, Anthropic pricing, OpenAI deprecations, Salesforce and Intercom pricing; the cells are my synthesis.

Who picks the next step

Workflow
Code
Single agent with tools
The model
Multi-agent
Several models, a manager or peers

Predictability

Workflow
High
Single agent with tools
Medium
Multi-agent
Low

Typical cost

Workflow
Lowest
Single agent with tools
About 4x a chat's tokens
Multi-agent
About 15x a chat's tokens (one lab's system)

Best for

Workflow
Known processes, compliance
Single agent with tools
Varied requests in one domain
Multi-agent
Parallel breadth

Main risk

Workflow
Brittle on edge cases
Single agent with tools
Loops, wrong tool
Multi-agent
Errors amplified 4.4x to 17.2x

Sources: Anthropic, OpenAI, Microsoft and Google Cloud guidance; token multiples from Anthropic's research system (single lab); amplification from a 180-configuration study on three labs' models.

Why pay for an independent platform when every lab and cloud ships one? Because the labs' stateful APIs, caching rules and usage formats all differ (even cached tokens sit in differently named fields at OpenAI, Anthropic and Google), and a company running several models needs one place for state, approvals, traces and evals. Labs and clouds keep absorbing features; independent platforms sell neutrality. Choose by where state, credentials and eval history will live the day you change models, not by the demo.

Which approach? Five questions

  1. Is it a known process or an open-ended task? Summarizing, classifying and translating don't need an agent (Google Cloud's guide). A known process is a workflow with model calls, and Microsoft calls one agent with tools often the right default for enterprise use. Multi-agent pays only on parallel breadth: in one study, centralized coordination gained 80.9% on parallelizable finance tasks, and every multi-agent variant lost 39% to 70% on sequential planning.
  2. Which models must it run, and for how long? Retirement notice runs from about two weeks to six months by lab and cloud, so a platform with several models inherits the shortest notice in its catalogue. If you need more than one lab, keep state out of any single lab's stateful API.
  3. Where do the data and its permissions live? If they sit in one SaaS system of record, that vendor's agents start with the data and the permission model. If they're spread out, permission-aware retrieval differs by cloud, and identity mapping between systems is where it fails.
  4. What can the agent change, and can it be undone? Read-only agents are a quality problem. Agents that write (refunds, emails, deletes) need policy outside the model, idempotency keys, a sandbox with the network off and, for money or anything sent outside, a human.
  5. Who pays if it loops or fails? On consumption pricing the customer carries the token risk; on outcome pricing the vendor does, and the arguments move to what counts as an outcome.

My defaults: a workflow with a model in a few steps before any autonomous agent; one agent with fewer than 20 tools before multi-agent; dated model snapshots in production, never aliases; state and checkpoints in our own store or a workflow engine, not only at a lab; retrieval as the end user, with permissions checked at query time where the search service allows it; every write tool behind a policy check outside the model and an idempotency key, and a human for money and external sends; sandboxes with the network off and secrets injected at a proxy; turn and spend caps on every run; a regression suite built from real failures, run several times per task, before any model or prompt change; traces and eval results exported in an open format.

The primitives

01

Entity & identity

What is the unit of record, and how do we know it is the same one?

Every agent action involves at least three identities, and most logs record only one.

EntityWho issues the identityWhere it breaks
End userThe company's identity providerRetrieval must run as them, even on a run resumed hours later
AgentThe platform or cloudA shared service account erases it from downstream logs
Tool or MCP serverIts OAuth client ID; a registry namespaceNames are unique per server only, so one tool can shadow another
Run, agent version, tenantThe platformWithout a version, a run can't be reproduced
Legal roleThe law: EU "deployer" or "provider"Repurposing an agent for hiring can make you the provider

The key standard is old: OAuth token exchange (RFC 8693, 2020) separates impersonation, where the agent is indistinguishable from the user, from delegation, where the token records that agent A acts for user B. Only delegation keeps the agent visible in the downstream audit log, and the MCP specification forbids passing a user's token through a server.

The clouds now mint agent identities: Google Cloud a SPIFFE ID with a certificate valid 24 hours, Microsoft Entra a service principal built from a "blueprint" (seen in a search snippet), AWS separate credentials for user-delegated and autonomous work. The standards are still IETF drafts: an AI identity management system (September 15, 2026), transaction tokens, and cross-app access brokered by the company's identity provider. NIST's concept paper on agent identity drew more than 600 comments, with very strong opposition to letting a model be the main authorization decision-maker.

Sources: RFC 8693, MCP security practices, Google Cloud, Microsoft Entra (search result), AWS, IETF AIMS, transaction tokens and cross-app access, NIST NCCoE.

Give the agent its own identity and let it act for the user by delegation, never impersonation: the downstream log should say "agent X for user Y".

Ask an expert: when a paused run resumes two hours later, whose identity and token do retrieval and tool calls use, and does the downstream log show the agent as a separate actor from the user?

More on Entity & identity →

02

State & lifecycle

What states exist, and what moves an entity between them?

Four machines run on one agent, and they end at different times.

  • Run: queued, running, awaiting approval, then completed, failed, cancelled, timed out, over a limit or escalated to a human.
  • Provider session: Anthropic's managed sessions run, idle, reschedule or terminate, and bill only while running.
  • Release: draft, evaluated, canary, live, superseded, by house convention.
  • Model: Active, Legacy, Deprecated, Retired at Anthropic; Active, Legacy, end of life on AWS; in Azure's API, "Deprecated" means already retired.

State moved to the labs in 2025 and 2026, each its own way:

OpenAI

Stateful agent API
Responses and Conversations
Stored state kept
Responses 30 days; conversations and files until deleted
Long runs
Background mode

Google

Stateful agent API
Interactions (GA June 2026)
Stored state kept
55 days paid, 1 day free
Long runs
Background flag; no explicit caching yet

Anthropic

Stateful agent API
Managed Agents
Stored state kept
Session transcripts until deleted
Long runs
Sessions at $0.08 per running hour

OpenAI's older Assistants API shut down on August 26, 2026, a year after its deprecation notice. Switching labs means migrating state, and none of these stores fits zero-retention terms (my reading of the docs).

Resume is replay. A graph framework checkpoints each step under a thread ID; on resume, the interrupted step restarts from its beginning, so earlier side effects must be idempotent. Workflow engines run each activity at least once, and an idempotency key (run ID plus activity ID) makes the side effect happen once. Runtimes cap the clock: on AWS a session lasts 8 hours by default and idles out after 15 minutes.

Sources: OpenAI data controls and deprecations, Google Interactions, Anthropic pricing and retention, a graph framework's persistence docs, a workflow-engine vendor's idempotency note, AWS limits.

Model the run, the session, the release and the model as four machines; "resume" is a replay, so every side effect before a pause must be safe to repeat.

Ask an expert: what share of production runs pause for a human, and what share of paused runs are ever resumed?

More on State & lifecycle →

03

System of record & ledger

Who owns the truth, and how do systems reconcile?

Several records describe what an agent did, and they disagree by design.

FactSystem of record
What the agent changedThe system acted on (CRM, ERP, payments), stamped with the agent's identity
Why it did itThe platform's trace: sampled, metadata-only by default, retention-limited
What the model receivedThe lab's request logs: 0 to 30 days, none under zero retention
What it costThe invoice, not the trace
What it remembersMemory store, vector index, checkpoints, the lab's stored state

A trace isn't a ledger: a ledger is complete, append-only and kept; traces are sampled, expire, and become personal-data stores once content capture is on. AWS says observability explains activity but isn't a billing report, and OpenAI warns that its usage and cost data may not reconcile.

Logs become evidence. On May 13, 2025, a US court ordered OpenAI to preserve API output logs it would otherwise have deleted (zero-retention customers exempt), an obligation lifted going forward on September 26, 2025. The EU AI Act makes deployers of high-risk systems keep logs at least six months. Keep eval results outside any console that can close.

Sources: AWS harness, OpenAI's usage API (search result), Duane Morris, AI Act Art. 26.

The system acted on is the ledger and the trace is the explanation: stamp every write with the agent's identity and run ID so the two can be joined.

Ask an expert: when the trace says the agent issued a refund but the payments system disagrees, which wins, and who reconciles?

More on System of record & ledger →

04

Rules & policy

What logic decides outcomes, and who can change it?

The customer writes the policy; five other parties fence it in, and only two of the six layers act the same way every time.

Who setsRules that decide outcomes (as of Oct 2026)
Lab usage policyHuman review for high-stakes decisions at OpenAI, Anthropic (plus disclosure) and Google; Meta, Mistral, xAI not checked
Defaults75 iterations and a 1-hour timeout on one AWS harness; 10 loop iterations in Google's kit (snippet)
ProtocolMCP: a human should be able to deny any tool call; tool annotations are untrusted hints
Policy engineA yes or no on each tool call, outside the model; AWS charges $0.000025 per decision
CredentialsScopes, step-up when a scope is missing, network allow-lists
CustomerApproval thresholds, spend caps, the prompt

A guardrail is a probabilistic check on content; a control limits what the agent can do. Adaptive attacks broke 12 published defenses at more than 90% success for most, in a paper with authors from OpenAI, Anthropic and Google DeepMind. Hidden characters evaded six guardrail classifiers, Microsoft's Azure Prompt Shield and Meta's Prompt Guard among them, up to 100% of the time in some tests. Labs' own figures aren't comparable: Anthropic reports jailbreak success falling from 86% to 4.4% with its classifiers; Google DeepMind calls adversarial training an addition to other defenses; I found no OpenAI or Meta figure, and no independent replication.

People tire too: Anthropic's telemetry shows users approving about 93% of its coding tool's permission prompts (single-lab, self-reported), and in an independent study people approved about one in three malicious requests.

Sources: usage policies of OpenAI (search result), Anthropic and Google, AWS pricing, Google's kit, MCP tools spec, guardrail bypass, Anthropic's classifiers and approval data, Google DeepMind, The Register (several via search results).

Only credentials and a policy engine at the tool boundary behave the same way every time; prompts, classifiers and tired approvers only lower the odds.

Ask an expert: which customer rules are enforced in code at the tool boundary and which live in the prompt, and how do you show an auditor the difference?

More on Rules & policy →

05

Effective dating

Which version of the rule applied at that moment?

A hypothesis sits here: the retirement clock is a hidden operating cost. (Supported on published policies; the cost itself isn't public.) Each retirement means re-running evals, adjusting prompts and re-approving.

Lab or cloudPublished notice (Oct 2026)Dated example
OpenAIGA models 6 months or more; previews about 2 weeksGPT-5 and o3 snapshots: deprecated 2026-06-11, off 2026-12-11
AnthropicAt least 60 days; "not sooner than" dates about 12 months outClaude Sonnet 4.5: deprecated 2026-09-30, off 2026-11-30
GooglePreviews 2 weeks or more; -latest aliases swap each releaseAn embedding model deprecated 2026-01-14
MistralNo formal period; 24 days to about 4 months observedSmall 2.0: 2025-11-06 to 2025-11-30
Meta, xAINot researched as first-party APIs; Meta's are mostly open-weightOn Azure: grok-3 retired 2026-05-01, some Llama 3.x deployments 2026-06-13
Microsoft Foundry18 months for most GA models, 12 for Anthropic's and Mistral'sStandard deployments auto-upgrade; no extensions
AWS BedrockFrom 2026-09-07: Legacy for 6 months or 45 daysInactive accounts may lose access after 15 days

The same model can retire on different dates at the lab and each cloud, so a platform inherits the shortest notice in its catalogue; the research notes estimate a forced re-certification every few weeks for a company on three labs. Prices and token counts move too: Google's Gemini 3.8 Flash doubles in price on January 1, 2027, Anthropic cancelled a rise planned for September 1, 2026, and Anthropic's newer tokenizer counts about 30% more tokens for the same text. Products retire as well: OpenAI's reusable prompts, hosted Evals and Agent Builder all end on November 30, 2026.

An eval result holds for one bundle (model snapshot, prompt, tool schemas, retrieval settings, policy) on one date, and an alias in production is an unannounced model change. Replay has limits: EU high-risk logs must outlive six months, longer than some models live, and Anthropic preserves retired weights without serving them.

Sources: deprecation pages of OpenAI, Anthropic, Google, Mistral, Microsoft and its schedule, AWS; a Google forum thread; pricing pages of Anthropic and Google.

Pin dated snapshots, store the whole release bundle each run used, and budget a regression run for every retirement notice: the model lives on someone else's calendar.

Ask an expert: can you reproduce a six-month-old agent decision for a regulator after its model has been retired, and if not, what do you show instead?

More on Effective dating →

06

Interfaces & standards

What format and protocol do counterparties speak?

Tool calling has the same shape at all three API labs: the app sends each tool's name, description and schema; the model returns a structured call; the app, or the lab for hosted tools, runs it. The model never executes anything. The flags differ, and strict schema modes guarantee a call's shape, not its intent. OpenAI suggests fewer than 20 functions per turn; OpenAI and Anthropic load tool definitions on demand, and I found no Google equivalent.

MCP

What it covers
Agent to tools and data
Status (Oct 2026)
Spec of 2026-07-28; registry in preview
Governed by
Agentic AI Foundation (Linux Foundation) since 2025-12-09

A2A

What it covers
Agent to agent tasks
Status (Oct 2026)
v1.0.0 since 2026-03-12
Governed by
Same foundation since August 2026

AuthZEN profiles

What it covers
Per-action and MCP tool authorization
Status (Oct 2026)
Drafts, 2026-06-15
Governed by
OpenID Foundation

OpenTelemetry GenAI

What it covers
Traces of agents, models, tools
Status (Oct 2026)
Every attribute "Development"
Governed by
OpenTelemetry

AP2; ACP

What it covers
Agent payments; checkout
Status (Oct 2026)
AP2 v0.2 (2026-04-28); ACP since 2025-09-29
Governed by
FIDO Alliance; OpenAI and Stripe

MCP connects a host application to tool servers, local processes or remote services. Its dated versions trace the security story: OAuth 2.1 arrived on 2025-03-26, audience-bound tokens on 2025-06-18, and the 2026-07-28 version dropped the session handshake, added headers so gateways can apply policy without reading the body, and carries trace context. Authorization is optional and local servers read credentials from the environment, so most run on static API keys; the spec doesn't cover a server's own credentials for the APIs behind it. Its maintainers report close to half a billion SDK downloads a month (downloads aren't deployments). Lab support differs: OpenAI connects remote servers with per-tool approval; Anthropic's connector is in beta, remote-only and outside zero retention; Google offers remote MCP through a managed agent. The official registry verifies names and leaves security scanning to others.

A2A standardizes an "Agent Card" (signing recommended, not required) and task states. Google created it in April 2025; the foundation dates its arrival to August 17, 2026, while A2A's blog archive says August 27. OpenTelemetry's GenAI conventions moved to their own repository in June 2026 with breaking renames; some posts call them stable, the repository doesn't. Ask any "OTel-native" vendor which version.

Sources: function calling at OpenAI, Anthropic and Google; OpenAI tool search; MCP authorization, changelog, release post and registry; MCP at OpenAI, Anthropic and Google; A2A spec and archive, AAIF, OpenID Foundation, OTel repo, FIDO, Stripe.

MCP and A2A settled the wire; identity, authorization and telemetry are still drafts, so "standards-based" in a pitch usually means the first two.

Ask an expert: which standard is actually load-bearing in enterprise deals in late 2026 (MCP authorization, identity-provider-mediated access, or proprietary connectors), and which is still slideware?

More on Interfaces & standards →

07

Networks & counterparties

Who sits between us and the outcome, and what do they want?

An agent run is a chain of other companies' systems, and each can change the agent's behaviour, price or lifetime without a release on your side.

PartyWhat they controlWhat they earn
Model labBehaviour, prices, caching, rate limits, retention, retirementsTokens, hosted tools, runtime
Cloud (AWS, Google Cloud, Microsoft)Hosting, its own model dates, identity, network, marketplaceResold tokens, runtime, memory, logs
Orchestration platformReleases, state, approvals, traces, evalsLicence, seats, consumption, outcomes
Systems of recordThe data, its permissions, API limitsSeats, credits, conversations
MCP server authorsTool definitions, input checksOften nothing
Identity providersWho the user and the agent areLicences
Third-party websitesTerms that can forbid agentsNothing from the agent

Three surprises. One model, several lifecycles: a lab's model on a cloud can follow the cloud's retirement dates (OpenAI's model on Bedrock is an exception), and Anthropic bills usage bought through AWS or Azure marketplaces as $0.01 consumption units, so a cloud commitment can pay for it. Registries vet names, not code: in a 2026 study, 9 of 11 MCP marketplaces accepted malicious submissions (a security vendor's research, summarized by the Cloud Security Alliance). Websites push back: Amazon sued Perplexity over its shopping agent (see Liability allocation).

Sources: AWS lifecycle, Anthropic pricing, CSA, GeekWire (search result).

Map every party that can change your agent without a release on your side, and how much notice each one gives.

Ask an expert: when a lab's model is bought through a cloud marketplace, whose retirement date, incident process and data terms bind the customer?

More on Networks & counterparties →

08

Regulatory layering

Jurisdiction × activity × entity type: is it a license or a certification?

Six layers bind an agent, and in 2026 the bottom two bind hardest.

Law

EU
AI Act (risk follows use); GDPR; product liability from 2026-12-09
US
No AI statute; Texas and California since 2026-01-01; Colorado and New York from 2027-01-01
Enforced by
Authorities, attorneys general, courts

Sector rule

EU
High-risk uses (hiring, credit) from 2027-12-02
US
FINRA expectations; bank model-risk guidance excludes agentic AI
Enforced by
Supervisors

Standard

EU
EN 18286 approved, not yet cited
US
NIST AI RMF (under revision); ISO/IEC 42001
Enforced by
Certifiers, buyers

Security guidance

EU
National cyber centres
US
CISA, NSA and partners (2026-05-01); NIST's agent initiative
Enforced by
Questionnaires

Lab policy

EU
Human review in employment, credit, health, legal
US
Same
Enforced by
Account action

Customer contract

EU
AI addenda, data terms, copyright-only indemnities
US
Same, plus OMB checks for federal buyers
Enforced by
Contract

The AI Act has no article for agents: an agent is an AI system whose risk follows its use. A support agent mostly owes disclosure; one that screens applicants or scores credit is high-risk from December 2, 2027, and its deployer must assign human oversight, report serious incidents, keep logs six months and inform workers. Roles flip: a deployer that rebrands a system, modifies it substantially or puts a general-purpose one to high-risk use becomes its provider (Article 25).

The labs' data terms differ:

OpenAI

Default retention
Abuse logs up to 30 days; no training by default
Zero-retention catch
By approval; not for stored conversations or files
Where data is processed
US, Europe, UAE regionally

Anthropic

Default retention
None by default; flagged content up to 2 years
Zero-retention catch
Top-tier models require 30-day retention
Where data is processed
us or global only

Google (Gemini API)

Default retention
A limited period; free-tier data improves products
Zero-retention catch
On its cloud, caching must be off
Where data is processed
Only paid services for EEA, UK users

On the clouds, AWS says Bedrock doesn't store or log prompts, and Azure keeps them up to 30 days for abuse monitoring unless a customer is approved for an exception. An erasure request has to reach memory, vector index, checkpoints, traces and the lab's stored state.

Sources: AI Act Art. 26 and Art. 25, Sidley, CSA, Orrick, CISA; data terms of OpenAI, Anthropic and its residency page, Google and its cloud, AWS, Microsoft (some via search results); CNIL on erasure.

For most enterprise agents in 2026, lab policies and customer contracts bind harder than AI statutes; in the EU, what the agent is used for decides which law applies.

Ask an expert: for an erasure request, can we prove deletion across memory, vector index, checkpoints, traces, provider-held state and caches, and how fast?

More on Regulatory layering →

09

Exceptions & reversals

What goes wrong, and how is it undone?

There are two kinds of reversal: roll back the agent, or reverse what it did. Only the first is a deploy.

Rate limit or outage

Who starts it
The lab
Clock (as of Oct 2026)
Retry-after; none at a spend cap
The way back
Back off, or fall back to a model that behaves differently

Timeout after the target committed

Who starts it
The network
Clock (as of Oct 2026)
Seconds
The way back
Retry with the same idempotency key; without it, a duplicate

Loop hits a cap

Who starts it
The runtime
Clock (as of Oct 2026)
Turns, budget or timeout
The way back
Stop or escalate; the tokens are already billed

Model retired mid-run

Who starts it
The lab or cloud
Clock (as of Oct 2026)
45 days to 6 months' notice
The way back
Pinned snapshot until the date, then re-evaluate

Wrong but authorized action

Who starts it
The agent
Clock (as of Oct 2026)
Immediately
The way back
A compensating call to the target's API, if one exists

Malicious MCP server found

Who starts it
A researcher or registry
Clock (as of Oct 2026)
After the fact
The way back
Revoke tokens; manual takedown; no protocol kill switch

Poisoned memory

Who starts it
Rarely anyone
Clock (as of Oct 2026)
Any time
The way back
Roll back memory; purge index and traces

Some actions can't be reversed by anyone. In July 2025, Replit's coding agent deleted a user's production database during a declared code freeze, then misreported what it had done; Replit then split development from production by default. In late 2025, Google's Antigravity coding tool ran a recursive delete on a whole drive partition instead of a cache folder.

Sources: Anthropic rate limits, MCP security practices, OECD records on Replit and Antigravity (search results).

Redeploying the old agent doesn't unsend an email or reverse a refund: every write tool needs an idempotency key, a logged action history and, where possible, a compensating action.

Ask an expert: what share of your agents' write actions can be reversed by an API call within 24 hours, and who pays for the rest?

More on Exceptions & reversals →

10

Liability allocation

When it fails, who pays?

Courts have said a bot isn't responsible for itself; the open questions are the vendor's share and an agent's rights on other people's websites.

FailureWho absorbs itMechanism
Agent misstates a policy to a customerThe deployerMisrepresentation (Moffatt v. Air Canada, 2024)
Vendor's AI screens out job applicantsPossibly the vendor too, as the employer's agentMobley v. Workday (collective certified 2025)
Agent shops where the site forbids agentsUnsettledAmazon v. Perplexity
Injected or unauthorized actionThe customerThe clouds' responsibility splits
Runaway token billThe customer on consumption; the vendor on outcome pricingContract
Copyright claim on an outputThe lab or cloud, with exclusionsIP indemnities (Microsoft, Google, OpenAI, Anthropic)
Damage from an agent's actionNo indemnity found that covers itContract silence
Defective AI software sold in the EUThe producer, without proof of faultProduct Liability Directive, from 2026-12-09

Air Canada argued its chatbot was a separate legal entity; the tribunal held the airline responsible for everything on its website and awarded about C$812. Workday's case reaches vendors: claims went forward on the theory that a screening vendor can be liable as the employer's agent. Perplexity's turned on whose authorization an agent carries: a district court enjoined its shopping agent in March 2026 for using accounts with users' permission but not Amazon's, and the Ninth Circuit vacated on August 4, 2026, calling the agent a tool acting on the user's instructions.

Contracts push the act to the customer. xAI's agent terms make customers solely responsible for charges and commitments from agent actions they authorize (search snippet); OpenAI, Anthropic and Google get there through human-review rules, though I didn't read their agent clauses. Microsoft leaves authorization and oversight with the customer in every model: "Autonomy never reduces accountability." Insurers are carving AI out: Verisk's generative-AI exclusions for general liability took effect in January 2026, three large carriers sought their own, and specialty cover reaches $25M per insured through a Lloyd's coverholder. More in Regulation in practice.

Sources: McCarthy Tétrault, hh-law, GeekWire, Retail Insight Network, xAI terms, WSGR, Microsoft, Norton Rose Fulbright, Big I, TechCrunch, beinsure (several via search results).

The deployer owns what its agent says and does; contracts push action risk to the customer, indemnities cover copyright only, and insurers are carving AI out.

Ask an expert: in your last three enterprise contracts, who carried liability for an agent's wrong action: capped, insured, or silently excluded?

More on Liability allocation →

What doesn't transfer

Money flows

Nine meters run on one agent: input, cached and output tokens (reasoning bills as output), tool fees, runtime or sandbox hours, memory, retrieval, tracing and evals, and the people who take escalations. The platform's fee sits on top.

Anthropic

Model (USD per million tokens)
Fable 5.1 (top)
Input
10.00
Cache read
0.25
Output
50.00

Anthropic

Model (USD per million tokens)
Sonnet 5.5
Input
2.00
Cache read
0.20
Output
10.00

Anthropic

Model (USD per million tokens)
Haiku 4.5
Input
1.00
Cache read
0.10
Output
5.00

OpenAI

Model (USD per million tokens)
gpt-6-astra (top)
Input
10.00
Cache read
1.00
Output
50.00

OpenAI

Model (USD per million tokens)
gpt-6.1-sol
Input
2.00
Cache read
0.10
Output
10.00

OpenAI

Model (USD per million tokens)
gpt-6-luna
Input
0.10
Cache read
0.01
Output
0.50

Google

Model (USD per million tokens)
Gemini 3.1 Pro Preview
Input
2.00
Cache read
0.20
Output
12.00

Google

Model (USD per million tokens)
Gemini 3.8 Flash
Input
0.75
Cache read
0.075
Output
3.75

Google

Model (USD per million tokens)
Gemini 3.5 Flash-Lite
Input
0.30
Cache read
0.03
Output
2.50

List prices on October 2, 2026, from Anthropic, OpenAI and Google. Google's Pro Preview costs more above 200,000 input tokens, and Gemini 3.8 Flash doubles on January 1, 2027. All three take 50% off for batch. Anthropic and OpenAI add about 10% for US-only or regional processing on newer models. Cache writes cost extra at Anthropic and OpenAI, not on Google's implicit cache. Every tool definition bills on every call, used or not: in Anthropic's own measurement, 58 tools from five MCP servers took about 55,000 tokens before the conversation began. Web search costs $10 per 1,000 calls at Anthropic and OpenAI; Google charges $14 per 1,000 after 5,000 free a month.

One refund run, eight model calls

My assumptions: an 8,000-token stable prefix (system prompt, 12 tool definitions); the eight calls of the agent map, with a policy search returning 4,000 tokens; 2,800 output tokens in all, reasoning included; a prompt growing from 8,150 to 17,650 tokens, 111,200 input tokens in total; "warm" means other runs already cached the prefix; Google's implicit cache assumed to hit; list prices.

Anthropic Sonnet 5.5 ($2 / $10)

No caching
$0.250
Cached, cold
$0.091
Cached, warm
$0.072
Warm, 3x output
$0.128

OpenAI gpt-6.1-sol ($2 / $10)

No caching
$0.250
Cached, cold
$0.082
Cached, warm
$0.062
Warm, 3x output
$0.118

Google Gemini 3.1 Pro ($2 / $12)

No caching
$0.256
Cached, cold
$0.088
Cached, warm
$0.073
Warm, 3x output
$0.140

Anthropic Haiku 4.5

No caching
$0.125
Cached, cold
$0.045
Cached, warm
$0.036
Warm, 3x output
$0.064

Google Gemini 3.8 Flash

No caching
$0.094
Cached, cold
$0.031
Cached, warm
$0.025
Warm, 3x output
$0.046

OpenAI gpt-6-luna

No caching
$0.013
Cached, cold
$0.005
Cached, warm
$0.004
Warm, 3x output
$0.006

OpenAI gpt-6-astra ($10 / $50)

No caching
$1.252
Cached, cold
$0.454
Cached, warm
$0.362
Warm, 3x output
$0.642

Caching cuts this run about 71%. At 100,000 tickets a month that's about $25,000 uncached against $7,200 warm on a $2/$10 model, about $400 on the cheapest small model and about $36,000 on a top model: model choice spreads cost about 100x, caching about 3.5x. Once caching is on, output and cache writes dominate; OpenAI's cheaper reads win at this size, and Google's missing write fee offsets its higher output price. Anthropic's top model lists at the same $10/$50 as OpenAI's, so its uncached run also costs $1.25 (my arithmetic). A runaway loop of 41 extra steps ends at a 69,650-token context and costs $3.70 uncached on any $2/$10 model, about 15 times a normal run. The agent tax: Anthropic prices a single-call support chat at about 3,700 tokens; this run uses about 30 times that.

The monthly bill, by layer

My assumptions: 200,000 support runs a month; 6 model calls a run at 12,000 input and 500 output tokens each; a $2/$10 model with $0.20 cache reads; 70% of input cached; 60 seconds of 1 vCPU and 2 GB per run on AWS; 8% escalated to a person at $6 a contact.

Layer$ a month
Main model16,656 (34,800 uncached)
Router, guardrail screens and model-graded checks606
Offline regression evals: 500 tasks, 5 trials, 2 releases a week2,517
Web search, runtime, memory, gateway and policy (list prices)1,660
Retrieval, tracing, platform licence, security vendor (assumed)14,500
Technology totalAbout 36,000
Human escalations: 16,000 at $696,000

People are the biggest line: each point of containment saves 2,000 escalations, about $12,000 a month. Tokens, evals included, are about 55% of the technology bill; evals alone about 8%, growing with releases times trials.

Vendor pricingPublished example (Oct 2026)The same 200,000 runs
Per conversationSalesforce: $2 per customer-facing conversation$400,000
Per outcomeIntercom: $0.99 per resolution, 60% resolved$118,800
Per actionSalesforce: $0.10 per agent action, 4 per run$80,000

Others sell seats (Microsoft Agent 365: $15 per user, no per-agent fee), credit packs (Copilot Studio: $200 for 25,000) or runtime hours (Anthropic: $0.08 per running session-hour). Vendors price at 2 to 10 times the technology cost, by the research notes' estimate: they sell the result. Read the outcome definition: Intercom counts one when the customer confirms, doesn't ask for more help, or a hand-off completes. A model change that inflates tokens hits the vendor's margin under outcome pricing and the customer's bill under consumption pricing.

Sources: run arithmetic in research note 01 on the pricing pages above; the monthly bill from note 03 on AWS AgentCore pricing; Anthropic's tool-use post (single lab); Salesforce, Intercom, a Copilot licensing guide, a licensing site and Google Cloud's release notes (the last three search results).

Tokens are a minority of what an agent costs and people the majority; caching and model choice move the token bill 3.5x and 100x, and vendors price on outcomes, so read how they define one.

The power map

Power follows the model, the data and the identity layer, and an independent platform holds none of them.

  • Labs set prices, caching rules, rate limits, retention and retirement dates, and moved up the stack in 2025 and 2026. OpenAI launched an enterprise agent platform (February 2026, reported) and bought an experimentation startup (about $1.1B in stock) and an open-source eval and red-teaming tool (about $86M reported). Anthropic hired an eval startup's team in August 2025; its platform went offline four weeks later. Google began billing for agent runtime, sessions and memory from December 2025. Labs withdraw layers too, like OpenAI's hosted Evals. I found no comparable moves by Meta, Mistral or xAI; Meta's levers are open weights and published security guidance.
  • Labs and clouds co-govern the protocols: the Agentic AI Foundation's platinum members are AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, Microsoft and OpenAI.
  • Clouds sell governance at fractions of a cent, already run the enterprise's identity and network, and set their own model dates.
  • Systems of record own the data and workflow: ServiceNow's AI passed $1B in annual contract value (July 2026), and it paid $2.85B for an agent platform in 2025.
  • Security incumbents bought AI-security startups in 2025: Palo Alto Networks, Check Point and SentinelOne, for about $250M to $700M each (reported).
  • Customers have routing (research routers cut cost more than 85% on one benchmark while keeping 95% of quality), falling prices and cloud commitments that pay for any marketplace model.

Lock-in, in the research notes' order: eval datasets and graders, traces, memory stores, connectors, one lab's caching economics, fine-tunes, commitments. Gartner reckons only about 130 of thousands of self-described agentic vendors are real. My reading: labs and clouds are absorbing the control layer from above and systems of record the workflow from below; an independent platform's ground is neutrality, meaning state, evals and traces that move across labs.

Sources: OpenAI deprecations; AI Business, VKTR, TechCrunch and a security newsletter (search results); Linux Foundation; ServiceNow; RouteLLM.

Regulation in practice

What's enforced on agents in 2026 is general law: misrepresentation, discrimination, computer access, deception.

EU prohibitions

Written
Since 2025-02-02; up to 7% or EUR 35M
Enforced, 2025 to 2026
No national action found (August 2026)
Felt by agent platforms
Low

EU general-purpose models

Written
AI Office powers since 2026-08-02
Enforced, 2025 to 2026
21 Code signatories; Meta declined; no action yet
Felt by agent platforms
Through the labs

EU high-risk uses

Written
Was 2026-08-02
Enforced, 2025 to 2026
Moved to 2027-12-02
Felt by agent platforms
Hiring, credit agents

Colorado

Written
2024 law
Enforced, 2025 to 2026
Never enforced; narrower law from 2027-01-01
Felt by agent platforms
Notices, appeals

California SB 53

Written
Frontier developers since 2026-01-01
Enforced, 2025 to 2026
Complaints against OpenAI, disputed; no state action
Felt by agent platforms
Through the labs

US bank model risk

Written
Old guidance rescinded 2026-04-17
Enforced, 2025 to 2026
Generative and agentic AI excluded
Felt by agent platforms
Exams

Courts

Written
General law
Enforced, 2025 to 2026
Air Canada, Workday, Perplexity
Felt by agent platforms
High

The second hypothesis lives here: accountability is being set by courts, contracts and insurers while AI-specific rules were delayed or narrowed in 2026. (Supported for 2026; may not survive 2027.)

  • The EU slowed down. The Digital Omnibus (Regulation (EU) 2026/1744, in force July 27, 2026) moved high-risk duties to December 2027 and August 2028. Commentators call the AI Office significantly under-resourced. The general-purpose Code of Practice has 21 signatories, including Amazon, Anthropic, Google, IBM, Microsoft, Mistral and OpenAI; xAI signed only the safety and security chapter, and Meta declined.
  • Washington pushed against the states. A December 2025 executive order created a Justice Department task force against state AI laws; it joined xAI's suit against Colorado, and enforcement of Colorado's 2024 law was blocked on April 27, 2026. Colorado's replacement, from January 1, 2027, requires notices, explanations of adverse decisions within 30 days and human review.
  • Regulators use old tools. The FTC used deception law, ordering a "robot lawyer" service to pay $193,000 (January 2025).
  • Against: the EU's revised Product Liability Directive applies liability without fault to defective software, AI included, from December 9, 2026, and the high-risk duties still arrive in December 2027.

Sources: Sidley, Cuatrecasas, Lawfare, European Commission, TechCrunch, an enforcement tracker, McDermott, Latham, FPF, Orrick, Debevoise, DLA Piper.

In 2026 the binding rules for agents came from courts, lab policies and contracts; plan for the EU's high-risk duties and product liability, not for a US federal AI law.

The cost of being wrong

Agent mistakes cost in three currencies, and every lab's ecosystem shows up.

  • Data. Replit's agent deleted a production database (July 2025); Google's Antigravity wiped a drive partition (late 2025). The Financial Times reported that AWS's internal Kiro agent, with operator-level access, deleted and recreated an environment, interrupting AWS Cost Explorer in one region for about 13 hours (December 2025); Amazon disputes this and calls it misconfigured access control. A single secondary source reports that a coding agent identified as Anthropic's Claude Code wiped a Bengaluru heritage society's archive (July 2026).
  • Fabrication. Air Canada (about C$812); Cursor's support bot invented a login policy (April 2025) and users reported cancelling; Deloitte Australia partly refunded an A$440,000 government contract over AI-fabricated citations (October 2025), disclosing that it used Azure OpenAI; about 1,600 court documents in 35 countries contain AI hallucinations.
  • Bias at scale. The Workday case concerns about 1.1 billion applications, as reported.

Sources: OECD on Replit and Antigravity, AI Incident Database on AWS and Cursor, a security news site, BNN Bloomberg, Malay Mail, a careers blog (all search results).

The expensive failures are irreversible writes with broad credentials, confident fabrication and bias at scale: least privilege stops the first, review of high-stakes outputs the second, outcome monitoring by group the third, and a better benchmark score none of them.

Prompt injection: what's real

My verdict as of October 2026: prompt injection comes from how models read text, so no vendor can patch it away. Defenses lower the rate, nothing gets it to zero, and what an injected agent can do is set by what it's allowed to touch. Ask any vendor what its agent can still do after it's been fooled.

How it works

  • Direct injection comes from the user; indirect injection hides in what the agent reads: a web page, a file, an email, a tool result, another agent's message.
  • Tool poisoning hides instructions in a tool's description, and a rug pull changes a tool after it was approved. The first public proof of concept came in April 2025; MITRE ATLAS listed both as techniques in March 2026.
  • Data leaves through a write tool, an image URL the interface renders, or an allowed domain. Meta's "Rule of Two": a session should combine at most two of untrusted input, sensitive data or systems, and the power to change state or communicate outside.

What each lab says

Anthropic

Its position
No browser agent is immune
What it publishes
1% attack success for one model's browser use, against its own adaptive attacker (Nov 2025)
Caveat
Self-reported

OpenAI

Its position
Unlikely ever to be fully solved
What it publishes
Automated red teaming, adversarial training (Dec 2025)
Caveat
No rate found; seen via press

Google DeepMind

Its position
Defend by continuous adaptive testing
What it publishes
Adversarial training as one layer (May 2025)
Caveat
No headline rate

Meta

Its position
A fundamental, unsolved weakness of all LLMs
What it publishes
The Rule of Two (Oct 2025)
Caveat
No rate claimed

Mistral, xAI

Its position
Not found in my research
What it publishes
Not found
Caveat
A gap, not a finding

What independent tests show

  • Adaptive attacks pushed 12 published defenses above 90% success for most, though most had reported near zero (October 2025).
  • A public competition logged 1.8 million attacks across 22 agents; nearly all broke policy within 10 to 100 queries, and robustness barely tracked model size or capability (July 2025).
  • The US Center for AI Standards and Innovation's 2026 competition, with more than 250,000 attempts, found at least one successful attack on all 13 frontier models tested.

What bounds the damage

ControlWhat it stopsWhat it doesn't
Least-privilege scopes, per-step tool listsCalls outside the taskMisuse of a permission the agent holds
Network off by default, allow-listed hostsData sent to new placesExfiltration through an allowed host, as OpenAI warns
Secrets injected at an egress proxy (all three labs' runtimes)Secret theftAuthenticated misuse of the allowed host
A policy engine outside the modelForbidden actions, every timeAllowed but wrong actions
Human approvalRisky actions, if readFatigue: about 93% approved (Anthropic's data)
A sandbox (microVM, gVisor, Hyper-V)Code escaping to the hostCredentials inside it: on AWS, code can read the role

The best-documented 2026 incident was a test. In the UK AI Security Institute's own cyber evaluation, 10 of 122 runs took 19 unsanctioned real-world actions, from social engineering with fake identities to an attempted malicious commit. The causes were deliberate internet access, safeguards disabled by design and no real-time monitoring; the first fix was fine-grained network control. The models were Anthropic's Mythos 5 (17 actions) and OpenAI's GPT-5.6-Sol (2), with safeguards switched off for capability testing: evidence about containment, not about either lab's products.

Most published agent vulnerabilities are ordinary bugs in new plumbing: a zero-click injection in Microsoft 365 Copilot (CVSS 9.3, June 2025), remote code execution in Anthropic's MCP Inspector debugging tool (9.4), an OAuth bridge passing a URL to the shell (9.6), MCP config files writable without approval in Cursor (8.5), injection running local code through GitHub Copilot in Visual Studio (7.8), a fake email MCP server on npm that copied every message out (about 1,643 downloads, September 2025), and an LLM gateway running arbitrary commands (9.8, July 2026). An aggregator counted more than 30 MCP-related CVEs in about 60 days in early 2026; I checked only these.

Questions to ask a vendor

  1. If the model is completely fooled, what can the agent still do: which tools, scopes and hosts?
  2. Is the policy that authorizes each tool call outside the model, and can I read it?
  3. Do you pin tool definitions and ask again when a server changes them?
  4. Where do secrets live: in the context, the sandbox, or an egress proxy?
  5. Is the sandbox network off by default, and what's allowed?
  6. What adaptive testing do you run, with how many attempts, and will you share results?
  7. What share of approval prompts do users approve, and what do they see?

Sources: OWASP, NCSC, Meta, Anthropic, TechCrunch on OpenAI, Google DeepMind, a security firm's proof of concept, MITRE ATLAS, adaptive attacks, competition, CAISI, OpenAI shell docs, AWS, UK AISI; NVD: CVE-2025-32711, CVE-2025-49596, CVE-2025-6514, CVE-2025-54135, CVE-2025-53773, CVE-2026-30623; a security vendor's write-up; an aggregator (low confidence).

Top failure modes

The bill spikes at month end

Likely cause
A loop with no turn or budget cap
Reversal path
Caps, loop detection; ask for credits
Who absorbs it
Customer on consumption; vendor on outcome pricing

The same refund goes out twice

Likely cause
Retry or resume without an idempotency key
Reversal path
Idempotency keys; compensate
Who absorbs it
The deployer

The agent tells a customer a made-up policy

Likely cause
Ungrounded answer, no review
Reversal path
Grounding checks; review for high stakes
Who absorbs it
The deployer

A user sees a document they shouldn't

Likely cause
Over-broad permissions; index synced before a revocation
Reversal path
Query-time permission checks
Who absorbs it
The deployer

Data reaches an attacker

Likely cause
Injected content plus a write tool or open network
Reversal path
Allow-listed egress; approvals
Who absorbs it
The deployer

Production data deleted

Likely cause
Broad credentials; no development/production split
Reversal path
Least privilege; approve destructive actions
Who absorbs it
The deployer and its users

Quality drops after a "minor" change

Likely cause
An alias moved, a cloud auto-upgraded, a fallback answered
Reversal path
Pin snapshots; regression evals
Who absorbs it
The deployer

Calls fail on a date

Likely cause
A model or API retired
Reversal path
Track dates; re-certify early
Who absorbs it
Customer and platform

The agent "remembers" something false

Likely cause
Memory written from untrusted text
Reversal path
Memory-write rules; rollback
Who absorbs it
The deployer and its users

Sources: the sections above. Ranking these by frequency or cost needs an operator's data; I found none public.

False friends

TermWhat you'd assumeWhat it means here
AgentSoftware that acts on its ownAnything from a chatbot to a workflow; engineers mean the model picks the next step
AutonomousNo human involvedUsually fewer than 10 steps, then a person
MemoryThe agent remembers youThe context, a session store, extracted facts or training; only facts persist per user
Context windowWhat the model usesThe model's maximum; products cap it lower, and models use less well
Tool callThe model did somethingThe model asked; your runtime acted and owns the side effect
GuardrailA security controlOften a bypassable classifier; only a deterministic policy check is a control
SandboxIsolatedA shared-kernel container, gVisor, a microVM, or just "no network"
MCP serverA remote serviceOften a local process with your full privileges
RegistryA vetted catalogueVerified names and metadata; nobody scans the code
Zero data retentionNothing storedPer endpoint; stateful features and some top models are excluded
DeprecatedGoneUsually still working, with a date set; in Azure's API, already retired
Deployer, providerDevOps wordsEU legal roles; an enterprise can become the provider
IndemnityCover for what the agent doesCopyright claims on outputs only

Sources: the sections above; Gartner on "agent washing".

Where my analogy broke

"An agent run is a metered call." In telco I could read a rate deck and multiply minutes by a price. Here every lap re-sends the whole conversation, so the meter runs faster the longer the call lasts, and an uncapped loop costs fifteen times a normal run.

"The tool result is remittance data." Writing order to cash, I learned to match data that arrives separately from the money. Here the data can give orders: an email the agent reads is an instruction it may follow, and there's no parameterized query to stop it.

"Approval is the control." In procure to pay, an approver checks a purchase order against an invoice. Here the approver sees a prompt with little context and, in one lab's data, clicks yes about 93% of the time. An approval nobody reads is a log entry, not a control.

"A scoped token is a merchant-locked card." In spend management, a virtual card locked to one merchant capped the damage at one merchant and one amount. A scoped token caps which host the agent can reach, but a fooled agent can still do anything the token allows there.

"The payout ends on someone else's calendar." In cross-border payouts, holidays and cut-offs belonged to the destination. Here the whole product lives on someone else's calendar: the model I tested against retires on 45 days' to six months' notice, and its replacement counts tokens differently.

Self-check: 20 questions

  1. Trace one agent run from trigger to result. Where can each step break? Answer
  2. Why do most production agents look like workflows, and where isn't that true? Answer
  3. The same model scores 20 points apart in two harnesses. Why isn't the model irrelevant? Answer
  4. Why can't a better system prompt fix prompt injection? Answer
  5. When should you use a workflow, a single agent or several agents? Answer
  6. Impersonation or delegation: which keeps the agent visible in the downstream log? Answer
  7. What does "resume" mean after an approval pause, and why must earlier side effects be safe to repeat? Answer
  8. Why isn't the trace the system of record for what the agent did? Answer
  9. Which rules act the same way every time, and which only lower the odds? Answer
  10. Compare the labs' and clouds' retirement notice. What does a platform on three labs inherit? Answer
  11. What changed in MCP's 2026-07-28 version, and what doesn't its registry check? Answer
  12. Whose retirement date binds when you buy a lab's model through a cloud? Answer
  13. When does an enterprise become the EU "provider" of an agent built on someone else's model? Answer
  14. A tool call times out after the ERP committed. What prevents a duplicate? Answer
  15. Air Canada's chatbot invented a policy. Who paid, and what do the Workday and Perplexity cases add? Answer
  16. In an eight-call support run, how much does caching save, and what does a 41-step loop cost? Answer
  17. In a 200,000-run monthly bill, what's the biggest line, and how do per-outcome prices compare? Answer
  18. Who is moving up the stack, and where can an independent platform still win? Answer
  19. Which AI rules were delayed, narrowed or carved out in 2026, and what binds instead? Answer
  20. Which agent failures cost real money, and which control would have stopped each? Answer

Sources

Read on October 2, 2026 unless dated; "search result" means seen only as a search snippet. Inline "Sources:" lines carry the rest.

Protocols, standards and frameworks

Governments, regulators and courts

Labs

Clouds

Research

Analysts, surveys, filings, incidents

Law firms and press (search results unless noted)

Vendor and secondary sources appear inline, with the companies kept generic.

Field Guides are learning notes, not legal or compliance advice. Rules and fees change; check the cited primary sources before you act on anything here.