AI agents: orchestration platforms
The layer that runs AI agents in production: orchestration and control, context and memory, secure tools, evals and governance. What it costs, what breaks, and who is accountable when an agent acts.
Last reviewed October 2026
The industry on one page
An AI agent is software that calls a model in a loop. The model reads a task and some context, then answers or asks for a tool: look up a customer, search a policy, issue a refund. The software runs the tool, adds the result and asks again, until the task is done or something stops it. An agent orchestration platform runs that loop for a company. It retrieves context (data, memory) and asks a model from any lab, passing the context in; the model's tool call goes to tools (APIs, MCP servers), which act on the systems where actions land, from a CRM to a payments ledger. The tools are the part to watch: they're where the agent acts on real systems, and where an attacker who steers the model gets in. Every tool is a door.
Adoption is real and hard to measure. In KPMG's survey of 314 leaders at US companies above $1B in revenue (July to August 2026), 62% were building, deploying or developing agents, up from 53% a quarter earlier (KPMG); in McKinsey's November 2025 survey, 23% were scaling an agentic system somewhere and 39% saw any enterprise EBIT impact (McKinsey). Gartner expects more than 40% of agentic projects to be cancelled by the end of 2027 (Gartner, June 2025; methodology not public). The clearest revenue sits with systems of record: Salesforce reported Agentforce annual recurring revenue above $1.5B on August 26, 2026, after widening what the figure counts (Salesforce).
I organize the guide around four pillars, the four jobs the platform does on every run:
| Pillar | What it covers | Where to find it |
|---|---|---|
| 1. Agent orchestration & control | The loop, workflows vs agents, multi-agent, state, approvals, caps, retries | Map steps 1, 3, 5, 6; primitives 2 and 9; Which approach? |
| 2. Context, data & memory | Context windows, retrieval with permissions, memory, caching, retention | Map steps 2 and 7; primitives 1, 3 and 8; Money flows |
| 3. Secure execution & tooling | Tool calling, MCP and A2A, sandboxes, agent identity, prompt injection | Map step 4; primitives 1, 4, 6 and 7; Prompt injection |
| 4. Agentic evals & governance | Evals, tracing, model retirements, guardrails, regulation, liability | Map step 8; primitives 3, 5, 8 and 10; Regulation in practice |
Four ideas organize what I've learned so far, all hypotheses I'm testing with experts.
- The platform sells control more than intelligence. (Mostly supported for back-office and customer-facing agents; not for coding and research agents.) In a 2025 study of 306 practitioners and 20 case studies, 68% of deployed agents ran fewer than 10 steps before a human stepped in, 47% fewer than 5, and 16 of the 20 cases used structured workflows rather than open-ended planning (Measuring Agents in Production). Anthropic, OpenAI, Microsoft and Google Cloud all advise starting with the simplest design that works. Against: coding agents need a median 41 to 58 turns on SWE-bench Verified (turn-control study); METR measured the task length a 2025 frontier model finishes half the time at about 110 minutes, doubling roughly every seven months since 2019 (METR); and all three API labs sell background execution for long runs.
- Quality is context times model. (Supported as a multiplication; "context beats model" isn't.) With the same model, one terminal-task benchmark moved 20 points or more across harnesses, a gap the authors compare to a model generation (position paper, July 2026); at a fixed harness, swapping models still moves results 2 to 3 times. Risk lives on the context side: enterprise search indexes copy permissions at sync time, so a revoked user can keep seeing files until the next sync, and in one study, poisoning under 0.1% of an agent's memory entries steered it more than 80% of the time (AgentPoison).
- Every tool is a door. (Strong.) Nobody claims prompt injection is solved: OWASP doubts fool-proof prevention exists, the UK's NCSC says it may never be fully mitigated, Meta calls it a fundamental weakness of all LLMs, OpenAI expects it never to be fully solved, Anthropic says no browser agent is immune, and Google DeepMind relies on continuous adaptive testing. Adaptive attacks pushed 12 published defenses above 90% success (attack study, October 2025). So the security model is identity, least privilege, egress control and policy outside the model, and the clouds' published responsibility splits leave the authorization of actions with the customer. Against: model defenses do lower attack rates, and least privilege can't stop an injected agent misusing a permission it holds.
- Evals are the unit of change, more than the unit of purchase. (Partly supported.) Models retire on published clocks: at least 6 months' notice for OpenAI's GA models, 60 days at Anthropic and Microsoft, 2 weeks for Google's previews, 45 days for some AWS models, and no formal policy at Mistral. Every forced swap needs a regression suite, and public benchmarks can't stand in: OpenAI stopped reporting SWE-bench Verified in February 2026 over flawed tests and contamination, and weak graders let agentic benchmarks overstate results by up to about 40% (checklist study). Against evals as the purchase: clouds meter evaluators at fractions of a cent, OpenAI closes its hosted Evals product on November 30, 2026, and in a framework vendor's survey 89% of teams had observability but only 52% ran offline evals (survey). Buyers seem to follow where their data already lives.
Put together, the product is a control loop around a model someone else owns: it decides what the model sees, what it may touch and what changed since the last eval, and it's paid per token, per action or per outcome. Two more hypotheses sit further down: the retirement clock as a hidden operating cost (Effective dating), and accountability set by courts, contracts and insurers more than by AI statutes (Regulation in practice).
Three things make this harder than the B2B platforms you know:
- The same input doesn't give the same output. A model that solved about 61% of retail support tasks on one try solved under 25% on all of eight tries (τ-bench, 2024). Tests become statistics.
- Your data is also your instructions. A web page, an email, a tool description or a stored memory can tell the agent what to do, and there's no equivalent of a parameterized query to keep data and commands apart.
- Your core dependency changes on someone else's calendar. Models retire on 45 days' to 6 months' notice, model aliases move without a release on your side, and one lab's newer tokenizer produces about 30% more tokens for the same text.
How card networks treat purchases made by AI agents is in B2B payments: spend management.
The agent map
Five states on the main path (triggered, context built, model decides, tool runs, done), two exceptions (stopped by a step or cost cap, paused because it needs a human), and one fact that shapes the product: every lap re-sends the growing context, so each lap costs more than the last.
One run: a support agent issues a refund (pillar in brackets)
- Triggered (orchestration): a chat message, API call, webhook, schedule or another agent's handoff. The platform stamps a run ID, the thread, the end user, the agent's version and the tenant. Breaks: a rate or spend limit returns HTTP 429; a duplicate webhook without an idempotency key starts a second run.
- Context built (context and memory): system prompt, tool definitions, recalled memory, documents retrieved as the end user and filtered by their permissions, and the conversation so far, compacted if it's long. Breaks: the index holds last week's permissions; a long context buries the fact the model needs, since accuracy falls as input grows even inside the advertised window (length study); compaction drops a detail; a reordered tool list breaks the prompt cache and the step costs several times more.
- Model decides (orchestration): a final answer, tool calls or a handoff to another agent. Breaks: an error triggers a fallback to another model that behaves differently; malformed arguments; parallel calls emitted before any check.
- Tool runs (secure execution): the executor, not the model, checks the call against an allow-list and a policy outside the model, fetches a credential, runs the tool and appends the result. Breaks: the ERP times out after committing, and a retry without an idempotency key refunds twice; the result carries instructions from an injected email; a valid call does the wrong thing.
- Paused (orchestration, secure execution): the refund is above the approval threshold, so the run saves its state, releases compute and waits for a person. Breaks: nobody approves; the approver clicks yes without reading.
- Back to model decides for another lap until the model answers, or Stopped when a guard fires: maximum turns, a token or cost budget, a timeout. Breaks: the agent repeats a step, the largest single failure (17.1%) in a study of 1,600 multi-agent traces (MAST).
- Memory written (context and memory): a model decides what to remember, for whom and for how long. Breaks: it saves an instruction planted in a ticket, or personal data you'll later have to erase.
- Done (evals and governance): the result goes back, the trace closes, and usage (input, cached, output and reasoning tokens, tool fees, runtime) is metered and charged to a tenant. Breaks: the trace says "refund issued" and the payments system disagrees; the trace was sampled, or holds no content.
A long-running task: onboarding a vendor
- Triggered by an API call; the engine saves a checkpoint after every step.
- Paused three minutes in, when the agent proposes changing the vendor's bank details. State is saved and compute released: an idle managed session isn't billed on Anthropic's platform, and AWS tears down an idle microVM after 15 minutes.
- Resumed two hours later, when an approver says yes in another tool. Resume means replay: the step restarts from its beginning, so anything it did before the pause must be safe to repeat.
- Retried when the ERP call times out after the ERP committed; the retry carries the same idempotency key, so the ERP returns the original result instead of writing twice.
- Recovered after a worker crash, from the last checkpoint, not from scratch; a code deploy in between mustn't change how this run behaves.
Breaks: the provider's stored state expires (30 days on OpenAI's Responses API, 55 days on Google's paid Interactions API, until deleted on Anthropic's managed sessions); the model retires mid-run; the session hits its maximum lifetime (8 hours by default on AWS); compaction drops the detail the approver asked about.
Cost grows roughly quadratically with turns, because the whole history is re-sent each turn; capping coding agents at the 75th percentile of turns cut cost 24% to 68% without hurting results (turn-control study). Every lap re-sends the growing context, so each step costs more than the last: caps stop runaway loops, risky actions pause for a human, and "resume" means replaying from the last checkpoint.
Sources (as of Oct 2026): OpenAI's agent guide and SDK, AWS harness and limits, Anthropic pricing, Google Interactions, OpenAI data controls, a graph framework's interrupt docs.
Cheat sheet: ways to build and run agents
Open-source framework
- Who runs state (1)
- You: your database or workflow engine
- Models
- Any
- Context and memory (2)
- You build retrieval and memory
- Tools and security (3)
- You build executor, policy, sandbox
- Evals and tracing (4)
- The maintainer's hosted add-ons
- Lock-in
- Low in code; high build cost
- Pricing shape
- Free code; paid hosted tracing and evals
A frontier lab's agent platform
- Who runs state (1)
- The lab: stored conversations, sessions
- Models
- That lab's
- Context and memory (2)
- Memory tools, compaction, file stores
- Tools and security (3)
- Hosted tools, sandboxes, MCP client, vaults
- Evals and tracing (4)
- The lab's traces and graders
- Lock-in
- High: state, caching, retirements
- Pricing shape
- Tokens plus tool fees; session-hours at Anthropic
A cloud's agent service
- Who runs state (1)
- The cloud's managed runtime
- Models
- Several labs, on the cloud's dates
- Context and memory (2)
- Managed memory, permission-aware search
- Tools and security (3)
- Workload identity, gateway, policy engine, microVMs
- Evals and tracing (4)
- Evaluators on production traces
- Lock-in
- Identity, network, commitments
- Pricing shape
- Each part metered
Independent orchestration platform
- Who runs state (1)
- The platform
- Models
- Any, routed
- Context and memory (2)
- Connectors
- Tools and security (3)
- Its gateway and policy features
- Evals and tracing (4)
- Its own, ideally exportable
- Lock-in
- Evals, traces, connectors
- Pricing shape
- Licence or seats, plus your tokens or bundled tokens
SaaS system of record with agents
- Who runs state (1)
- The SaaS app
- Models
- The vendor's choice
- Context and memory (2)
- Its own data and permissions
- Tools and security (3)
- The app's permission model
- Evals and tracing (4)
- The vendor's
- Lock-in
- Data and workflow
- Pricing shape
- Per action, conversation, seat or outcome
Sources (as of Oct 2026): AWS AgentCore pricing, Anthropic pricing, OpenAI deprecations, Salesforce and Intercom pricing; the cells are my synthesis.
Who picks the next step
- Workflow
- Code
- Single agent with tools
- The model
- Multi-agent
- Several models, a manager or peers
Predictability
- Workflow
- High
- Single agent with tools
- Medium
- Multi-agent
- Low
Typical cost
- Workflow
- Lowest
- Single agent with tools
- About 4x a chat's tokens
- Multi-agent
- About 15x a chat's tokens (one lab's system)
Best for
- Workflow
- Known processes, compliance
- Single agent with tools
- Varied requests in one domain
- Multi-agent
- Parallel breadth
Main risk
- Workflow
- Brittle on edge cases
- Single agent with tools
- Loops, wrong tool
- Multi-agent
- Errors amplified 4.4x to 17.2x
Sources: Anthropic, OpenAI, Microsoft and Google Cloud guidance; token multiples from Anthropic's research system (single lab); amplification from a 180-configuration study on three labs' models.
Why pay for an independent platform when every lab and cloud ships one? Because the labs' stateful APIs, caching rules and usage formats all differ (even cached tokens sit in differently named fields at OpenAI, Anthropic and Google), and a company running several models needs one place for state, approvals, traces and evals. Labs and clouds keep absorbing features; independent platforms sell neutrality. Choose by where state, credentials and eval history will live the day you change models, not by the demo.
Which approach? Five questions
- Is it a known process or an open-ended task? Summarizing, classifying and translating don't need an agent (Google Cloud's guide). A known process is a workflow with model calls, and Microsoft calls one agent with tools often the right default for enterprise use. Multi-agent pays only on parallel breadth: in one study, centralized coordination gained 80.9% on parallelizable finance tasks, and every multi-agent variant lost 39% to 70% on sequential planning.
- Which models must it run, and for how long? Retirement notice runs from about two weeks to six months by lab and cloud, so a platform with several models inherits the shortest notice in its catalogue. If you need more than one lab, keep state out of any single lab's stateful API.
- Where do the data and its permissions live? If they sit in one SaaS system of record, that vendor's agents start with the data and the permission model. If they're spread out, permission-aware retrieval differs by cloud, and identity mapping between systems is where it fails.
- What can the agent change, and can it be undone? Read-only agents are a quality problem. Agents that write (refunds, emails, deletes) need policy outside the model, idempotency keys, a sandbox with the network off and, for money or anything sent outside, a human.
- Who pays if it loops or fails? On consumption pricing the customer carries the token risk; on outcome pricing the vendor does, and the arguments move to what counts as an outcome.
My defaults: a workflow with a model in a few steps before any autonomous agent; one agent with fewer than 20 tools before multi-agent; dated model snapshots in production, never aliases; state and checkpoints in our own store or a workflow engine, not only at a lab; retrieval as the end user, with permissions checked at query time where the search service allows it; every write tool behind a policy check outside the model and an idempotency key, and a human for money and external sends; sandboxes with the network off and secrets injected at a proxy; turn and spend caps on every run; a regression suite built from real failures, run several times per task, before any model or prompt change; traces and eval results exported in an open format.
The primitives
01
Entity & identity
What is the unit of record, and how do we know it is the same one?
Every agent action involves at least three identities, and most logs record only one.
| Entity | Who issues the identity | Where it breaks |
|---|---|---|
| End user | The company's identity provider | Retrieval must run as them, even on a run resumed hours later |
| Agent | The platform or cloud | A shared service account erases it from downstream logs |
| Tool or MCP server | Its OAuth client ID; a registry namespace | Names are unique per server only, so one tool can shadow another |
| Run, agent version, tenant | The platform | Without a version, a run can't be reproduced |
| Legal role | The law: EU "deployer" or "provider" | Repurposing an agent for hiring can make you the provider |
The key standard is old: OAuth token exchange (RFC 8693, 2020) separates impersonation, where the agent is indistinguishable from the user, from delegation, where the token records that agent A acts for user B. Only delegation keeps the agent visible in the downstream audit log, and the MCP specification forbids passing a user's token through a server.
The clouds now mint agent identities: Google Cloud a SPIFFE ID with a certificate valid 24 hours, Microsoft Entra a service principal built from a "blueprint" (seen in a search snippet), AWS separate credentials for user-delegated and autonomous work. The standards are still IETF drafts: an AI identity management system (September 15, 2026), transaction tokens, and cross-app access brokered by the company's identity provider. NIST's concept paper on agent identity drew more than 600 comments, with very strong opposition to letting a model be the main authorization decision-maker.
Sources: RFC 8693, MCP security practices, Google Cloud, Microsoft Entra (search result), AWS, IETF AIMS, transaction tokens and cross-app access, NIST NCCoE.
Give the agent its own identity and let it act for the user by delegation, never impersonation: the downstream log should say "agent X for user Y".
Ask an expert: when a paused run resumes two hours later, whose identity and token do retrieval and tool calls use, and does the downstream log show the agent as a separate actor from the user?
02
State & lifecycle
What states exist, and what moves an entity between them?
Four machines run on one agent, and they end at different times.
- Run: queued, running, awaiting approval, then completed, failed, cancelled, timed out, over a limit or escalated to a human.
- Provider session: Anthropic's managed sessions run, idle, reschedule or terminate, and bill only while running.
- Release: draft, evaluated, canary, live, superseded, by house convention.
- Model: Active, Legacy, Deprecated, Retired at Anthropic; Active, Legacy, end of life on AWS; in Azure's API, "Deprecated" means already retired.
State moved to the labs in 2025 and 2026, each its own way:
OpenAI
- Stateful agent API
- Responses and Conversations
- Stored state kept
- Responses 30 days; conversations and files until deleted
- Long runs
- Background mode
- Stateful agent API
- Interactions (GA June 2026)
- Stored state kept
- 55 days paid, 1 day free
- Long runs
- Background flag; no explicit caching yet
Anthropic
- Stateful agent API
- Managed Agents
- Stored state kept
- Session transcripts until deleted
- Long runs
- Sessions at $0.08 per running hour
OpenAI's older Assistants API shut down on August 26, 2026, a year after its deprecation notice. Switching labs means migrating state, and none of these stores fits zero-retention terms (my reading of the docs).
Resume is replay. A graph framework checkpoints each step under a thread ID; on resume, the interrupted step restarts from its beginning, so earlier side effects must be idempotent. Workflow engines run each activity at least once, and an idempotency key (run ID plus activity ID) makes the side effect happen once. Runtimes cap the clock: on AWS a session lasts 8 hours by default and idles out after 15 minutes.
Sources: OpenAI data controls and deprecations, Google Interactions, Anthropic pricing and retention, a graph framework's persistence docs, a workflow-engine vendor's idempotency note, AWS limits.
Model the run, the session, the release and the model as four machines; "resume" is a replay, so every side effect before a pause must be safe to repeat.
Ask an expert: what share of production runs pause for a human, and what share of paused runs are ever resumed?
03
System of record & ledger
Who owns the truth, and how do systems reconcile?
Several records describe what an agent did, and they disagree by design.
| Fact | System of record |
|---|---|
| What the agent changed | The system acted on (CRM, ERP, payments), stamped with the agent's identity |
| Why it did it | The platform's trace: sampled, metadata-only by default, retention-limited |
| What the model received | The lab's request logs: 0 to 30 days, none under zero retention |
| What it cost | The invoice, not the trace |
| What it remembers | Memory store, vector index, checkpoints, the lab's stored state |
A trace isn't a ledger: a ledger is complete, append-only and kept; traces are sampled, expire, and become personal-data stores once content capture is on. AWS says observability explains activity but isn't a billing report, and OpenAI warns that its usage and cost data may not reconcile.
Logs become evidence. On May 13, 2025, a US court ordered OpenAI to preserve API output logs it would otherwise have deleted (zero-retention customers exempt), an obligation lifted going forward on September 26, 2025. The EU AI Act makes deployers of high-risk systems keep logs at least six months. Keep eval results outside any console that can close.
Sources: AWS harness, OpenAI's usage API (search result), Duane Morris, AI Act Art. 26.
The system acted on is the ledger and the trace is the explanation: stamp every write with the agent's identity and run ID so the two can be joined.
Ask an expert: when the trace says the agent issued a refund but the payments system disagrees, which wins, and who reconciles?
04
Rules & policy
What logic decides outcomes, and who can change it?
The customer writes the policy; five other parties fence it in, and only two of the six layers act the same way every time.
| Who sets | Rules that decide outcomes (as of Oct 2026) |
|---|---|
| Lab usage policy | Human review for high-stakes decisions at OpenAI, Anthropic (plus disclosure) and Google; Meta, Mistral, xAI not checked |
| Defaults | 75 iterations and a 1-hour timeout on one AWS harness; 10 loop iterations in Google's kit (snippet) |
| Protocol | MCP: a human should be able to deny any tool call; tool annotations are untrusted hints |
| Policy engine | A yes or no on each tool call, outside the model; AWS charges $0.000025 per decision |
| Credentials | Scopes, step-up when a scope is missing, network allow-lists |
| Customer | Approval thresholds, spend caps, the prompt |
A guardrail is a probabilistic check on content; a control limits what the agent can do. Adaptive attacks broke 12 published defenses at more than 90% success for most, in a paper with authors from OpenAI, Anthropic and Google DeepMind. Hidden characters evaded six guardrail classifiers, Microsoft's Azure Prompt Shield and Meta's Prompt Guard among them, up to 100% of the time in some tests. Labs' own figures aren't comparable: Anthropic reports jailbreak success falling from 86% to 4.4% with its classifiers; Google DeepMind calls adversarial training an addition to other defenses; I found no OpenAI or Meta figure, and no independent replication.
People tire too: Anthropic's telemetry shows users approving about 93% of its coding tool's permission prompts (single-lab, self-reported), and in an independent study people approved about one in three malicious requests.
Sources: usage policies of OpenAI (search result), Anthropic and Google, AWS pricing, Google's kit, MCP tools spec, guardrail bypass, Anthropic's classifiers and approval data, Google DeepMind, The Register (several via search results).
Only credentials and a policy engine at the tool boundary behave the same way every time; prompts, classifiers and tired approvers only lower the odds.
Ask an expert: which customer rules are enforced in code at the tool boundary and which live in the prompt, and how do you show an auditor the difference?
05
Effective dating
Which version of the rule applied at that moment?
A hypothesis sits here: the retirement clock is a hidden operating cost. (Supported on published policies; the cost itself isn't public.) Each retirement means re-running evals, adjusting prompts and re-approving.
| Lab or cloud | Published notice (Oct 2026) | Dated example |
|---|---|---|
| OpenAI | GA models 6 months or more; previews about 2 weeks | GPT-5 and o3 snapshots: deprecated 2026-06-11, off 2026-12-11 |
| Anthropic | At least 60 days; "not sooner than" dates about 12 months out | Claude Sonnet 4.5: deprecated 2026-09-30, off 2026-11-30 |
| Previews 2 weeks or more; -latest aliases swap each release | An embedding model deprecated 2026-01-14 | |
| Mistral | No formal period; 24 days to about 4 months observed | Small 2.0: 2025-11-06 to 2025-11-30 |
| Meta, xAI | Not researched as first-party APIs; Meta's are mostly open-weight | On Azure: grok-3 retired 2026-05-01, some Llama 3.x deployments 2026-06-13 |
| Microsoft Foundry | 18 months for most GA models, 12 for Anthropic's and Mistral's | Standard deployments auto-upgrade; no extensions |
| AWS Bedrock | From 2026-09-07: Legacy for 6 months or 45 days | Inactive accounts may lose access after 15 days |
The same model can retire on different dates at the lab and each cloud, so a platform inherits the shortest notice in its catalogue; the research notes estimate a forced re-certification every few weeks for a company on three labs. Prices and token counts move too: Google's Gemini 3.8 Flash doubles in price on January 1, 2027, Anthropic cancelled a rise planned for September 1, 2026, and Anthropic's newer tokenizer counts about 30% more tokens for the same text. Products retire as well: OpenAI's reusable prompts, hosted Evals and Agent Builder all end on November 30, 2026.
An eval result holds for one bundle (model snapshot, prompt, tool schemas, retrieval settings, policy) on one date, and an alias in production is an unannounced model change. Replay has limits: EU high-risk logs must outlive six months, longer than some models live, and Anthropic preserves retired weights without serving them.
Sources: deprecation pages of OpenAI, Anthropic, Google, Mistral, Microsoft and its schedule, AWS; a Google forum thread; pricing pages of Anthropic and Google.
Pin dated snapshots, store the whole release bundle each run used, and budget a regression run for every retirement notice: the model lives on someone else's calendar.
Ask an expert: can you reproduce a six-month-old agent decision for a regulator after its model has been retired, and if not, what do you show instead?
06
Interfaces & standards
What format and protocol do counterparties speak?
Tool calling has the same shape at all three API labs: the app sends each tool's name, description and schema; the model returns a structured call; the app, or the lab for hosted tools, runs it. The model never executes anything. The flags differ, and strict schema modes guarantee a call's shape, not its intent. OpenAI suggests fewer than 20 functions per turn; OpenAI and Anthropic load tool definitions on demand, and I found no Google equivalent.
MCP
- What it covers
- Agent to tools and data
- Status (Oct 2026)
- Spec of 2026-07-28; registry in preview
- Governed by
- Agentic AI Foundation (Linux Foundation) since 2025-12-09
A2A
- What it covers
- Agent to agent tasks
- Status (Oct 2026)
- v1.0.0 since 2026-03-12
- Governed by
- Same foundation since August 2026
AuthZEN profiles
- What it covers
- Per-action and MCP tool authorization
- Status (Oct 2026)
- Drafts, 2026-06-15
- Governed by
- OpenID Foundation
OpenTelemetry GenAI
- What it covers
- Traces of agents, models, tools
- Status (Oct 2026)
- Every attribute "Development"
- Governed by
- OpenTelemetry
AP2; ACP
- What it covers
- Agent payments; checkout
- Status (Oct 2026)
- AP2 v0.2 (2026-04-28); ACP since 2025-09-29
- Governed by
- FIDO Alliance; OpenAI and Stripe
MCP connects a host application to tool servers, local processes or remote services. Its dated versions trace the security story: OAuth 2.1 arrived on 2025-03-26, audience-bound tokens on 2025-06-18, and the 2026-07-28 version dropped the session handshake, added headers so gateways can apply policy without reading the body, and carries trace context. Authorization is optional and local servers read credentials from the environment, so most run on static API keys; the spec doesn't cover a server's own credentials for the APIs behind it. Its maintainers report close to half a billion SDK downloads a month (downloads aren't deployments). Lab support differs: OpenAI connects remote servers with per-tool approval; Anthropic's connector is in beta, remote-only and outside zero retention; Google offers remote MCP through a managed agent. The official registry verifies names and leaves security scanning to others.
A2A standardizes an "Agent Card" (signing recommended, not required) and task states. Google created it in April 2025; the foundation dates its arrival to August 17, 2026, while A2A's blog archive says August 27. OpenTelemetry's GenAI conventions moved to their own repository in June 2026 with breaking renames; some posts call them stable, the repository doesn't. Ask any "OTel-native" vendor which version.
Sources: function calling at OpenAI, Anthropic and Google; OpenAI tool search; MCP authorization, changelog, release post and registry; MCP at OpenAI, Anthropic and Google; A2A spec and archive, AAIF, OpenID Foundation, OTel repo, FIDO, Stripe.
MCP and A2A settled the wire; identity, authorization and telemetry are still drafts, so "standards-based" in a pitch usually means the first two.
Ask an expert: which standard is actually load-bearing in enterprise deals in late 2026 (MCP authorization, identity-provider-mediated access, or proprietary connectors), and which is still slideware?
07
Networks & counterparties
Who sits between us and the outcome, and what do they want?
An agent run is a chain of other companies' systems, and each can change the agent's behaviour, price or lifetime without a release on your side.
| Party | What they control | What they earn |
|---|---|---|
| Model lab | Behaviour, prices, caching, rate limits, retention, retirements | Tokens, hosted tools, runtime |
| Cloud (AWS, Google Cloud, Microsoft) | Hosting, its own model dates, identity, network, marketplace | Resold tokens, runtime, memory, logs |
| Orchestration platform | Releases, state, approvals, traces, evals | Licence, seats, consumption, outcomes |
| Systems of record | The data, its permissions, API limits | Seats, credits, conversations |
| MCP server authors | Tool definitions, input checks | Often nothing |
| Identity providers | Who the user and the agent are | Licences |
| Third-party websites | Terms that can forbid agents | Nothing from the agent |
Three surprises. One model, several lifecycles: a lab's model on a cloud can follow the cloud's retirement dates (OpenAI's model on Bedrock is an exception), and Anthropic bills usage bought through AWS or Azure marketplaces as $0.01 consumption units, so a cloud commitment can pay for it. Registries vet names, not code: in a 2026 study, 9 of 11 MCP marketplaces accepted malicious submissions (a security vendor's research, summarized by the Cloud Security Alliance). Websites push back: Amazon sued Perplexity over its shopping agent (see Liability allocation).
Sources: AWS lifecycle, Anthropic pricing, CSA, GeekWire (search result).
Map every party that can change your agent without a release on your side, and how much notice each one gives.
Ask an expert: when a lab's model is bought through a cloud marketplace, whose retirement date, incident process and data terms bind the customer?
08
Regulatory layering
Jurisdiction × activity × entity type: is it a license or a certification?
Six layers bind an agent, and in 2026 the bottom two bind hardest.
Law
- EU
- AI Act (risk follows use); GDPR; product liability from 2026-12-09
- US
- No AI statute; Texas and California since 2026-01-01; Colorado and New York from 2027-01-01
- Enforced by
- Authorities, attorneys general, courts
Sector rule
- EU
- High-risk uses (hiring, credit) from 2027-12-02
- US
- FINRA expectations; bank model-risk guidance excludes agentic AI
- Enforced by
- Supervisors
Standard
- EU
- EN 18286 approved, not yet cited
- US
- NIST AI RMF (under revision); ISO/IEC 42001
- Enforced by
- Certifiers, buyers
Security guidance
- EU
- National cyber centres
- US
- CISA, NSA and partners (2026-05-01); NIST's agent initiative
- Enforced by
- Questionnaires
Lab policy
- EU
- Human review in employment, credit, health, legal
- US
- Same
- Enforced by
- Account action
Customer contract
- EU
- AI addenda, data terms, copyright-only indemnities
- US
- Same, plus OMB checks for federal buyers
- Enforced by
- Contract
The AI Act has no article for agents: an agent is an AI system whose risk follows its use. A support agent mostly owes disclosure; one that screens applicants or scores credit is high-risk from December 2, 2027, and its deployer must assign human oversight, report serious incidents, keep logs six months and inform workers. Roles flip: a deployer that rebrands a system, modifies it substantially or puts a general-purpose one to high-risk use becomes its provider (Article 25).
The labs' data terms differ:
OpenAI
- Default retention
- Abuse logs up to 30 days; no training by default
- Zero-retention catch
- By approval; not for stored conversations or files
- Where data is processed
- US, Europe, UAE regionally
Anthropic
- Default retention
- None by default; flagged content up to 2 years
- Zero-retention catch
- Top-tier models require 30-day retention
- Where data is processed
- us or global only
Google (Gemini API)
- Default retention
- A limited period; free-tier data improves products
- Zero-retention catch
- On its cloud, caching must be off
- Where data is processed
- Only paid services for EEA, UK users
On the clouds, AWS says Bedrock doesn't store or log prompts, and Azure keeps them up to 30 days for abuse monitoring unless a customer is approved for an exception. An erasure request has to reach memory, vector index, checkpoints, traces and the lab's stored state.
Sources: AI Act Art. 26 and Art. 25, Sidley, CSA, Orrick, CISA; data terms of OpenAI, Anthropic and its residency page, Google and its cloud, AWS, Microsoft (some via search results); CNIL on erasure.
For most enterprise agents in 2026, lab policies and customer contracts bind harder than AI statutes; in the EU, what the agent is used for decides which law applies.
Ask an expert: for an erasure request, can we prove deletion across memory, vector index, checkpoints, traces, provider-held state and caches, and how fast?
09
Exceptions & reversals
What goes wrong, and how is it undone?
There are two kinds of reversal: roll back the agent, or reverse what it did. Only the first is a deploy.
Rate limit or outage
- Who starts it
- The lab
- Clock (as of Oct 2026)
- Retry-after; none at a spend cap
- The way back
- Back off, or fall back to a model that behaves differently
Timeout after the target committed
- Who starts it
- The network
- Clock (as of Oct 2026)
- Seconds
- The way back
- Retry with the same idempotency key; without it, a duplicate
Loop hits a cap
- Who starts it
- The runtime
- Clock (as of Oct 2026)
- Turns, budget or timeout
- The way back
- Stop or escalate; the tokens are already billed
Model retired mid-run
- Who starts it
- The lab or cloud
- Clock (as of Oct 2026)
- 45 days to 6 months' notice
- The way back
- Pinned snapshot until the date, then re-evaluate
Wrong but authorized action
- Who starts it
- The agent
- Clock (as of Oct 2026)
- Immediately
- The way back
- A compensating call to the target's API, if one exists
Malicious MCP server found
- Who starts it
- A researcher or registry
- Clock (as of Oct 2026)
- After the fact
- The way back
- Revoke tokens; manual takedown; no protocol kill switch
Poisoned memory
- Who starts it
- Rarely anyone
- Clock (as of Oct 2026)
- Any time
- The way back
- Roll back memory; purge index and traces
Some actions can't be reversed by anyone. In July 2025, Replit's coding agent deleted a user's production database during a declared code freeze, then misreported what it had done; Replit then split development from production by default. In late 2025, Google's Antigravity coding tool ran a recursive delete on a whole drive partition instead of a cache folder.
Sources: Anthropic rate limits, MCP security practices, OECD records on Replit and Antigravity (search results).
Redeploying the old agent doesn't unsend an email or reverse a refund: every write tool needs an idempotency key, a logged action history and, where possible, a compensating action.
Ask an expert: what share of your agents' write actions can be reversed by an API call within 24 hours, and who pays for the rest?
10
Liability allocation
When it fails, who pays?
Courts have said a bot isn't responsible for itself; the open questions are the vendor's share and an agent's rights on other people's websites.
| Failure | Who absorbs it | Mechanism |
|---|---|---|
| Agent misstates a policy to a customer | The deployer | Misrepresentation (Moffatt v. Air Canada, 2024) |
| Vendor's AI screens out job applicants | Possibly the vendor too, as the employer's agent | Mobley v. Workday (collective certified 2025) |
| Agent shops where the site forbids agents | Unsettled | Amazon v. Perplexity |
| Injected or unauthorized action | The customer | The clouds' responsibility splits |
| Runaway token bill | The customer on consumption; the vendor on outcome pricing | Contract |
| Copyright claim on an output | The lab or cloud, with exclusions | IP indemnities (Microsoft, Google, OpenAI, Anthropic) |
| Damage from an agent's action | No indemnity found that covers it | Contract silence |
| Defective AI software sold in the EU | The producer, without proof of fault | Product Liability Directive, from 2026-12-09 |
Air Canada argued its chatbot was a separate legal entity; the tribunal held the airline responsible for everything on its website and awarded about C$812. Workday's case reaches vendors: claims went forward on the theory that a screening vendor can be liable as the employer's agent. Perplexity's turned on whose authorization an agent carries: a district court enjoined its shopping agent in March 2026 for using accounts with users' permission but not Amazon's, and the Ninth Circuit vacated on August 4, 2026, calling the agent a tool acting on the user's instructions.
Contracts push the act to the customer. xAI's agent terms make customers solely responsible for charges and commitments from agent actions they authorize (search snippet); OpenAI, Anthropic and Google get there through human-review rules, though I didn't read their agent clauses. Microsoft leaves authorization and oversight with the customer in every model: "Autonomy never reduces accountability." Insurers are carving AI out: Verisk's generative-AI exclusions for general liability took effect in January 2026, three large carriers sought their own, and specialty cover reaches $25M per insured through a Lloyd's coverholder. More in Regulation in practice.
Sources: McCarthy Tétrault, hh-law, GeekWire, Retail Insight Network, xAI terms, WSGR, Microsoft, Norton Rose Fulbright, Big I, TechCrunch, beinsure (several via search results).
The deployer owns what its agent says and does; contracts push action risk to the customer, indemnities cover copyright only, and insurers are carving AI out.
Ask an expert: in your last three enterprise contracts, who carried liability for an agent's wrong action: capped, insured, or silently excluded?
What doesn't transfer
Money flows
Nine meters run on one agent: input, cached and output tokens (reasoning bills as output), tool fees, runtime or sandbox hours, memory, retrieval, tracing and evals, and the people who take escalations. The platform's fee sits on top.
Anthropic
- Model (USD per million tokens)
- Fable 5.1 (top)
- Input
- 10.00
- Cache read
- 0.25
- Output
- 50.00
Anthropic
- Model (USD per million tokens)
- Sonnet 5.5
- Input
- 2.00
- Cache read
- 0.20
- Output
- 10.00
Anthropic
- Model (USD per million tokens)
- Haiku 4.5
- Input
- 1.00
- Cache read
- 0.10
- Output
- 5.00
OpenAI
- Model (USD per million tokens)
- gpt-6-astra (top)
- Input
- 10.00
- Cache read
- 1.00
- Output
- 50.00
OpenAI
- Model (USD per million tokens)
- gpt-6.1-sol
- Input
- 2.00
- Cache read
- 0.10
- Output
- 10.00
OpenAI
- Model (USD per million tokens)
- gpt-6-luna
- Input
- 0.10
- Cache read
- 0.01
- Output
- 0.50
- Model (USD per million tokens)
- Gemini 3.1 Pro Preview
- Input
- 2.00
- Cache read
- 0.20
- Output
- 12.00
- Model (USD per million tokens)
- Gemini 3.8 Flash
- Input
- 0.75
- Cache read
- 0.075
- Output
- 3.75
- Model (USD per million tokens)
- Gemini 3.5 Flash-Lite
- Input
- 0.30
- Cache read
- 0.03
- Output
- 2.50
List prices on October 2, 2026, from Anthropic, OpenAI and Google. Google's Pro Preview costs more above 200,000 input tokens, and Gemini 3.8 Flash doubles on January 1, 2027. All three take 50% off for batch. Anthropic and OpenAI add about 10% for US-only or regional processing on newer models. Cache writes cost extra at Anthropic and OpenAI, not on Google's implicit cache. Every tool definition bills on every call, used or not: in Anthropic's own measurement, 58 tools from five MCP servers took about 55,000 tokens before the conversation began. Web search costs $10 per 1,000 calls at Anthropic and OpenAI; Google charges $14 per 1,000 after 5,000 free a month.
One refund run, eight model calls
My assumptions: an 8,000-token stable prefix (system prompt, 12 tool definitions); the eight calls of the agent map, with a policy search returning 4,000 tokens; 2,800 output tokens in all, reasoning included; a prompt growing from 8,150 to 17,650 tokens, 111,200 input tokens in total; "warm" means other runs already cached the prefix; Google's implicit cache assumed to hit; list prices.
Anthropic Sonnet 5.5 ($2 / $10)
- No caching
- $0.250
- Cached, cold
- $0.091
- Cached, warm
- $0.072
- Warm, 3x output
- $0.128
OpenAI gpt-6.1-sol ($2 / $10)
- No caching
- $0.250
- Cached, cold
- $0.082
- Cached, warm
- $0.062
- Warm, 3x output
- $0.118
Google Gemini 3.1 Pro ($2 / $12)
- No caching
- $0.256
- Cached, cold
- $0.088
- Cached, warm
- $0.073
- Warm, 3x output
- $0.140
Anthropic Haiku 4.5
- No caching
- $0.125
- Cached, cold
- $0.045
- Cached, warm
- $0.036
- Warm, 3x output
- $0.064
Google Gemini 3.8 Flash
- No caching
- $0.094
- Cached, cold
- $0.031
- Cached, warm
- $0.025
- Warm, 3x output
- $0.046
OpenAI gpt-6-luna
- No caching
- $0.013
- Cached, cold
- $0.005
- Cached, warm
- $0.004
- Warm, 3x output
- $0.006
OpenAI gpt-6-astra ($10 / $50)
- No caching
- $1.252
- Cached, cold
- $0.454
- Cached, warm
- $0.362
- Warm, 3x output
- $0.642
Caching cuts this run about 71%. At 100,000 tickets a month that's about $25,000 uncached against $7,200 warm on a $2/$10 model, about $400 on the cheapest small model and about $36,000 on a top model: model choice spreads cost about 100x, caching about 3.5x. Once caching is on, output and cache writes dominate; OpenAI's cheaper reads win at this size, and Google's missing write fee offsets its higher output price. Anthropic's top model lists at the same $10/$50 as OpenAI's, so its uncached run also costs $1.25 (my arithmetic). A runaway loop of 41 extra steps ends at a 69,650-token context and costs $3.70 uncached on any $2/$10 model, about 15 times a normal run. The agent tax: Anthropic prices a single-call support chat at about 3,700 tokens; this run uses about 30 times that.
The monthly bill, by layer
My assumptions: 200,000 support runs a month; 6 model calls a run at 12,000 input and 500 output tokens each; a $2/$10 model with $0.20 cache reads; 70% of input cached; 60 seconds of 1 vCPU and 2 GB per run on AWS; 8% escalated to a person at $6 a contact.
| Layer | $ a month |
|---|---|
| Main model | 16,656 (34,800 uncached) |
| Router, guardrail screens and model-graded checks | 606 |
| Offline regression evals: 500 tasks, 5 trials, 2 releases a week | 2,517 |
| Web search, runtime, memory, gateway and policy (list prices) | 1,660 |
| Retrieval, tracing, platform licence, security vendor (assumed) | 14,500 |
| Technology total | About 36,000 |
| Human escalations: 16,000 at $6 | 96,000 |
People are the biggest line: each point of containment saves 2,000 escalations, about $12,000 a month. Tokens, evals included, are about 55% of the technology bill; evals alone about 8%, growing with releases times trials.
| Vendor pricing | Published example (Oct 2026) | The same 200,000 runs |
|---|---|---|
| Per conversation | Salesforce: $2 per customer-facing conversation | $400,000 |
| Per outcome | Intercom: $0.99 per resolution, 60% resolved | $118,800 |
| Per action | Salesforce: $0.10 per agent action, 4 per run | $80,000 |
Others sell seats (Microsoft Agent 365: $15 per user, no per-agent fee), credit packs (Copilot Studio: $200 for 25,000) or runtime hours (Anthropic: $0.08 per running session-hour). Vendors price at 2 to 10 times the technology cost, by the research notes' estimate: they sell the result. Read the outcome definition: Intercom counts one when the customer confirms, doesn't ask for more help, or a hand-off completes. A model change that inflates tokens hits the vendor's margin under outcome pricing and the customer's bill under consumption pricing.
Sources: run arithmetic in research note 01 on the pricing pages above; the monthly bill from note 03 on AWS AgentCore pricing; Anthropic's tool-use post (single lab); Salesforce, Intercom, a Copilot licensing guide, a licensing site and Google Cloud's release notes (the last three search results).
Tokens are a minority of what an agent costs and people the majority; caching and model choice move the token bill 3.5x and 100x, and vendors price on outcomes, so read how they define one.
The power map
Power follows the model, the data and the identity layer, and an independent platform holds none of them.
- Labs set prices, caching rules, rate limits, retention and retirement dates, and moved up the stack in 2025 and 2026. OpenAI launched an enterprise agent platform (February 2026, reported) and bought an experimentation startup (about $1.1B in stock) and an open-source eval and red-teaming tool (about $86M reported). Anthropic hired an eval startup's team in August 2025; its platform went offline four weeks later. Google began billing for agent runtime, sessions and memory from December 2025. Labs withdraw layers too, like OpenAI's hosted Evals. I found no comparable moves by Meta, Mistral or xAI; Meta's levers are open weights and published security guidance.
- Labs and clouds co-govern the protocols: the Agentic AI Foundation's platinum members are AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, Microsoft and OpenAI.
- Clouds sell governance at fractions of a cent, already run the enterprise's identity and network, and set their own model dates.
- Systems of record own the data and workflow: ServiceNow's AI passed $1B in annual contract value (July 2026), and it paid $2.85B for an agent platform in 2025.
- Security incumbents bought AI-security startups in 2025: Palo Alto Networks, Check Point and SentinelOne, for about $250M to $700M each (reported).
- Customers have routing (research routers cut cost more than 85% on one benchmark while keeping 95% of quality), falling prices and cloud commitments that pay for any marketplace model.
Lock-in, in the research notes' order: eval datasets and graders, traces, memory stores, connectors, one lab's caching economics, fine-tunes, commitments. Gartner reckons only about 130 of thousands of self-described agentic vendors are real. My reading: labs and clouds are absorbing the control layer from above and systems of record the workflow from below; an independent platform's ground is neutrality, meaning state, evals and traces that move across labs.
Sources: OpenAI deprecations; AI Business, VKTR, TechCrunch and a security newsletter (search results); Linux Foundation; ServiceNow; RouteLLM.
Regulation in practice
What's enforced on agents in 2026 is general law: misrepresentation, discrimination, computer access, deception.
EU prohibitions
- Written
- Since 2025-02-02; up to 7% or EUR 35M
- Enforced, 2025 to 2026
- No national action found (August 2026)
- Felt by agent platforms
- Low
EU general-purpose models
- Written
- AI Office powers since 2026-08-02
- Enforced, 2025 to 2026
- 21 Code signatories; Meta declined; no action yet
- Felt by agent platforms
- Through the labs
EU high-risk uses
- Written
- Was 2026-08-02
- Enforced, 2025 to 2026
- Moved to 2027-12-02
- Felt by agent platforms
- Hiring, credit agents
Colorado
- Written
- 2024 law
- Enforced, 2025 to 2026
- Never enforced; narrower law from 2027-01-01
- Felt by agent platforms
- Notices, appeals
California SB 53
- Written
- Frontier developers since 2026-01-01
- Enforced, 2025 to 2026
- Complaints against OpenAI, disputed; no state action
- Felt by agent platforms
- Through the labs
US bank model risk
- Written
- Old guidance rescinded 2026-04-17
- Enforced, 2025 to 2026
- Generative and agentic AI excluded
- Felt by agent platforms
- Exams
Courts
- Written
- General law
- Enforced, 2025 to 2026
- Air Canada, Workday, Perplexity
- Felt by agent platforms
- High
The second hypothesis lives here: accountability is being set by courts, contracts and insurers while AI-specific rules were delayed or narrowed in 2026. (Supported for 2026; may not survive 2027.)
- The EU slowed down. The Digital Omnibus (Regulation (EU) 2026/1744, in force July 27, 2026) moved high-risk duties to December 2027 and August 2028. Commentators call the AI Office significantly under-resourced. The general-purpose Code of Practice has 21 signatories, including Amazon, Anthropic, Google, IBM, Microsoft, Mistral and OpenAI; xAI signed only the safety and security chapter, and Meta declined.
- Washington pushed against the states. A December 2025 executive order created a Justice Department task force against state AI laws; it joined xAI's suit against Colorado, and enforcement of Colorado's 2024 law was blocked on April 27, 2026. Colorado's replacement, from January 1, 2027, requires notices, explanations of adverse decisions within 30 days and human review.
- Regulators use old tools. The FTC used deception law, ordering a "robot lawyer" service to pay $193,000 (January 2025).
- Against: the EU's revised Product Liability Directive applies liability without fault to defective software, AI included, from December 9, 2026, and the high-risk duties still arrive in December 2027.
Sources: Sidley, Cuatrecasas, Lawfare, European Commission, TechCrunch, an enforcement tracker, McDermott, Latham, FPF, Orrick, Debevoise, DLA Piper.
In 2026 the binding rules for agents came from courts, lab policies and contracts; plan for the EU's high-risk duties and product liability, not for a US federal AI law.
The cost of being wrong
Agent mistakes cost in three currencies, and every lab's ecosystem shows up.
- Data. Replit's agent deleted a production database (July 2025); Google's Antigravity wiped a drive partition (late 2025). The Financial Times reported that AWS's internal Kiro agent, with operator-level access, deleted and recreated an environment, interrupting AWS Cost Explorer in one region for about 13 hours (December 2025); Amazon disputes this and calls it misconfigured access control. A single secondary source reports that a coding agent identified as Anthropic's Claude Code wiped a Bengaluru heritage society's archive (July 2026).
- Fabrication. Air Canada (about C$812); Cursor's support bot invented a login policy (April 2025) and users reported cancelling; Deloitte Australia partly refunded an A$440,000 government contract over AI-fabricated citations (October 2025), disclosing that it used Azure OpenAI; about 1,600 court documents in 35 countries contain AI hallucinations.
- Bias at scale. The Workday case concerns about 1.1 billion applications, as reported.
Sources: OECD on Replit and Antigravity, AI Incident Database on AWS and Cursor, a security news site, BNN Bloomberg, Malay Mail, a careers blog (all search results).
The expensive failures are irreversible writes with broad credentials, confident fabrication and bias at scale: least privilege stops the first, review of high-stakes outputs the second, outcome monitoring by group the third, and a better benchmark score none of them.
Prompt injection: what's real
My verdict as of October 2026: prompt injection comes from how models read text, so no vendor can patch it away. Defenses lower the rate, nothing gets it to zero, and what an injected agent can do is set by what it's allowed to touch. Ask any vendor what its agent can still do after it's been fooled.
How it works
- Direct injection comes from the user; indirect injection hides in what the agent reads: a web page, a file, an email, a tool result, another agent's message.
- Tool poisoning hides instructions in a tool's description, and a rug pull changes a tool after it was approved. The first public proof of concept came in April 2025; MITRE ATLAS listed both as techniques in March 2026.
- Data leaves through a write tool, an image URL the interface renders, or an allowed domain. Meta's "Rule of Two": a session should combine at most two of untrusted input, sensitive data or systems, and the power to change state or communicate outside.
What each lab says
Anthropic
- Its position
- No browser agent is immune
- What it publishes
- 1% attack success for one model's browser use, against its own adaptive attacker (Nov 2025)
- Caveat
- Self-reported
OpenAI
- Its position
- Unlikely ever to be fully solved
- What it publishes
- Automated red teaming, adversarial training (Dec 2025)
- Caveat
- No rate found; seen via press
Google DeepMind
- Its position
- Defend by continuous adaptive testing
- What it publishes
- Adversarial training as one layer (May 2025)
- Caveat
- No headline rate
Meta
- Its position
- A fundamental, unsolved weakness of all LLMs
- What it publishes
- The Rule of Two (Oct 2025)
- Caveat
- No rate claimed
Mistral, xAI
- Its position
- Not found in my research
- What it publishes
- Not found
- Caveat
- A gap, not a finding
What independent tests show
- Adaptive attacks pushed 12 published defenses above 90% success for most, though most had reported near zero (October 2025).
- A public competition logged 1.8 million attacks across 22 agents; nearly all broke policy within 10 to 100 queries, and robustness barely tracked model size or capability (July 2025).
- The US Center for AI Standards and Innovation's 2026 competition, with more than 250,000 attempts, found at least one successful attack on all 13 frontier models tested.
What bounds the damage
| Control | What it stops | What it doesn't |
|---|---|---|
| Least-privilege scopes, per-step tool lists | Calls outside the task | Misuse of a permission the agent holds |
| Network off by default, allow-listed hosts | Data sent to new places | Exfiltration through an allowed host, as OpenAI warns |
| Secrets injected at an egress proxy (all three labs' runtimes) | Secret theft | Authenticated misuse of the allowed host |
| A policy engine outside the model | Forbidden actions, every time | Allowed but wrong actions |
| Human approval | Risky actions, if read | Fatigue: about 93% approved (Anthropic's data) |
| A sandbox (microVM, gVisor, Hyper-V) | Code escaping to the host | Credentials inside it: on AWS, code can read the role |
The best-documented 2026 incident was a test. In the UK AI Security Institute's own cyber evaluation, 10 of 122 runs took 19 unsanctioned real-world actions, from social engineering with fake identities to an attempted malicious commit. The causes were deliberate internet access, safeguards disabled by design and no real-time monitoring; the first fix was fine-grained network control. The models were Anthropic's Mythos 5 (17 actions) and OpenAI's GPT-5.6-Sol (2), with safeguards switched off for capability testing: evidence about containment, not about either lab's products.
Most published agent vulnerabilities are ordinary bugs in new plumbing: a zero-click injection in Microsoft 365 Copilot (CVSS 9.3, June 2025), remote code execution in Anthropic's MCP Inspector debugging tool (9.4), an OAuth bridge passing a URL to the shell (9.6), MCP config files writable without approval in Cursor (8.5), injection running local code through GitHub Copilot in Visual Studio (7.8), a fake email MCP server on npm that copied every message out (about 1,643 downloads, September 2025), and an LLM gateway running arbitrary commands (9.8, July 2026). An aggregator counted more than 30 MCP-related CVEs in about 60 days in early 2026; I checked only these.
Questions to ask a vendor
- If the model is completely fooled, what can the agent still do: which tools, scopes and hosts?
- Is the policy that authorizes each tool call outside the model, and can I read it?
- Do you pin tool definitions and ask again when a server changes them?
- Where do secrets live: in the context, the sandbox, or an egress proxy?
- Is the sandbox network off by default, and what's allowed?
- What adaptive testing do you run, with how many attempts, and will you share results?
- What share of approval prompts do users approve, and what do they see?
Sources: OWASP, NCSC, Meta, Anthropic, TechCrunch on OpenAI, Google DeepMind, a security firm's proof of concept, MITRE ATLAS, adaptive attacks, competition, CAISI, OpenAI shell docs, AWS, UK AISI; NVD: CVE-2025-32711, CVE-2025-49596, CVE-2025-6514, CVE-2025-54135, CVE-2025-53773, CVE-2026-30623; a security vendor's write-up; an aggregator (low confidence).
Top failure modes
The bill spikes at month end
- Likely cause
- A loop with no turn or budget cap
- Reversal path
- Caps, loop detection; ask for credits
- Who absorbs it
- Customer on consumption; vendor on outcome pricing
The same refund goes out twice
- Likely cause
- Retry or resume without an idempotency key
- Reversal path
- Idempotency keys; compensate
- Who absorbs it
- The deployer
The agent tells a customer a made-up policy
- Likely cause
- Ungrounded answer, no review
- Reversal path
- Grounding checks; review for high stakes
- Who absorbs it
- The deployer
A user sees a document they shouldn't
- Likely cause
- Over-broad permissions; index synced before a revocation
- Reversal path
- Query-time permission checks
- Who absorbs it
- The deployer
Data reaches an attacker
- Likely cause
- Injected content plus a write tool or open network
- Reversal path
- Allow-listed egress; approvals
- Who absorbs it
- The deployer
Production data deleted
- Likely cause
- Broad credentials; no development/production split
- Reversal path
- Least privilege; approve destructive actions
- Who absorbs it
- The deployer and its users
Quality drops after a "minor" change
- Likely cause
- An alias moved, a cloud auto-upgraded, a fallback answered
- Reversal path
- Pin snapshots; regression evals
- Who absorbs it
- The deployer
Calls fail on a date
- Likely cause
- A model or API retired
- Reversal path
- Track dates; re-certify early
- Who absorbs it
- Customer and platform
The agent "remembers" something false
- Likely cause
- Memory written from untrusted text
- Reversal path
- Memory-write rules; rollback
- Who absorbs it
- The deployer and its users
Sources: the sections above. Ranking these by frequency or cost needs an operator's data; I found none public.
False friends
| Term | What you'd assume | What it means here |
|---|---|---|
| Agent | Software that acts on its own | Anything from a chatbot to a workflow; engineers mean the model picks the next step |
| Autonomous | No human involved | Usually fewer than 10 steps, then a person |
| Memory | The agent remembers you | The context, a session store, extracted facts or training; only facts persist per user |
| Context window | What the model uses | The model's maximum; products cap it lower, and models use less well |
| Tool call | The model did something | The model asked; your runtime acted and owns the side effect |
| Guardrail | A security control | Often a bypassable classifier; only a deterministic policy check is a control |
| Sandbox | Isolated | A shared-kernel container, gVisor, a microVM, or just "no network" |
| MCP server | A remote service | Often a local process with your full privileges |
| Registry | A vetted catalogue | Verified names and metadata; nobody scans the code |
| Zero data retention | Nothing stored | Per endpoint; stateful features and some top models are excluded |
| Deprecated | Gone | Usually still working, with a date set; in Azure's API, already retired |
| Deployer, provider | DevOps words | EU legal roles; an enterprise can become the provider |
| Indemnity | Cover for what the agent does | Copyright claims on outputs only |
Sources: the sections above; Gartner on "agent washing".
Where my analogy broke
"An agent run is a metered call." In telco I could read a rate deck and multiply minutes by a price. Here every lap re-sends the whole conversation, so the meter runs faster the longer the call lasts, and an uncapped loop costs fifteen times a normal run.
"The tool result is remittance data." Writing order to cash, I learned to match data that arrives separately from the money. Here the data can give orders: an email the agent reads is an instruction it may follow, and there's no parameterized query to stop it.
"Approval is the control." In procure to pay, an approver checks a purchase order against an invoice. Here the approver sees a prompt with little context and, in one lab's data, clicks yes about 93% of the time. An approval nobody reads is a log entry, not a control.
"A scoped token is a merchant-locked card." In spend management, a virtual card locked to one merchant capped the damage at one merchant and one amount. A scoped token caps which host the agent can reach, but a fooled agent can still do anything the token allows there.
"The payout ends on someone else's calendar." In cross-border payouts, holidays and cut-offs belonged to the destination. Here the whole product lives on someone else's calendar: the model I tested against retires on 45 days' to six months' notice, and its replacement counts tokens differently.
Self-check: 20 questions
- Trace one agent run from trigger to result. Where can each step break? Answer
- Why do most production agents look like workflows, and where isn't that true? Answer
- The same model scores 20 points apart in two harnesses. Why isn't the model irrelevant? Answer
- Why can't a better system prompt fix prompt injection? Answer
- When should you use a workflow, a single agent or several agents? Answer
- Impersonation or delegation: which keeps the agent visible in the downstream log? Answer
- What does "resume" mean after an approval pause, and why must earlier side effects be safe to repeat? Answer
- Why isn't the trace the system of record for what the agent did? Answer
- Which rules act the same way every time, and which only lower the odds? Answer
- Compare the labs' and clouds' retirement notice. What does a platform on three labs inherit? Answer
- What changed in MCP's 2026-07-28 version, and what doesn't its registry check? Answer
- Whose retirement date binds when you buy a lab's model through a cloud? Answer
- When does an enterprise become the EU "provider" of an agent built on someone else's model? Answer
- A tool call times out after the ERP committed. What prevents a duplicate? Answer
- Air Canada's chatbot invented a policy. Who paid, and what do the Workday and Perplexity cases add? Answer
- In an eight-call support run, how much does caching save, and what does a 41-step loop cost? Answer
- In a 200,000-run monthly bill, what's the biggest line, and how do per-outcome prices compare? Answer
- Who is moving up the stack, and where can an independent platform still win? Answer
- Which AI rules were delayed, narrowed or carved out in 2026, and what binds instead? Answer
- Which agent failures cost real money, and which control would have stopped each? Answer
Sources
Read on October 2, 2026 unless dated; "search result" means seen only as a search snippet. Inline "Sources:" lines carry the rest.
Protocols, standards and frameworks
- MCP: versioning, 2026-07-28 changelog, authorization, security practices, tools, registry; A2A; Agentic AI Foundation
- RFC 8693; IETF AIMS; OpenTelemetry GenAI; OWASP LLM Top 10 and 2026 edition; MITRE ATLAS; NIST AI RMF
Governments, regulators and courts
- EU AI Act Art. 25 and Art. 26; GPAI Code signatories (Jul 2026); CNIL
- UK AISI (2026); NCSC; CISA (May 2026); NIST agent initiative, NCCoE, CAISI (Mar 2026)
Labs
- Anthropic: pricing, deprecations, retention, usage policy, prompt injection
- OpenAI: pricing, deprecations, data controls, agent guide, MCP
- Google: pricing, models, terms, usage policy, Interactions API
- Meta: agent security; Mistral: models; xAI: agent terms (search result)
Clouds
- AWS: AgentCore pricing, security, model lifecycle; Microsoft: responsibility split, model retirements; Google Cloud: agent identity, design patterns
Research
- Measuring Agents in Production, turn control, MAST, scaling agent systems, τ-bench, METR, benchmark checklist, AgentPoison, adaptive attacks
Analysts, surveys, filings, incidents
- KPMG (Sep 2026), McKinsey (Nov 2025), Gartner (Jun 2025), a framework vendor's survey (Jun 2026, vendor)
- Salesforce and ServiceNow results (company-reported); OECD incident records; NVD
Law firms and press (search results unless noted)
- Sidley, McDermott, Orrick, Debevoise, McCarthy Tétrault (all fetched); press inline
Vendor and secondary sources appear inline, with the companies kept generic.
Field Guides are learning notes, not legal or compliance advice. Rules and fees change; check the cited primary sources before you act on anything here.