Models and inference: what AI runs on
How a request becomes tokens and a bill: model labs, open-weight models, clouds, inference providers, GPU clouds and chips, what a feature really costs, why capacity is scarce, and who pays when a model is throttled, retired or wrong.
Last updated October 2026
The industry on one page
Picture a company that sells help-desk software to other businesses. It adds a button that summarizes a support ticket and drafts a reply, and customers press it about a million times a month, each time sending roughly 2,000 tokens in and getting 500 back. A token is the unit a model reads, writes and bills, a piece of a word: on Anthropic's current tokenizer a million tokens hold about 555,000 English words. Your app sends requests, often through a gateway that picks a model and caches answers, and that also keeps a budget per team and switches models when a call fails. Behind it, a lab's API, a cloud platform or an inference provider runs the model on GPUs in someone's data center: Anthropic, OpenAI or Google selling its own closed model; Amazon Bedrock, Microsoft Foundry or Google Vertex AI reselling several labs' models; or Fireworks or Together serving an open-weight model such as DeepSeek or OpenAI's gpt-oss. Most rent their GPUs from a GPU cloud such as CoreWeave or from a hyperscaler, and the chips come from Nvidia, Google's TPUs or custom designs such as Amazon's Trainium. At the bottom, chips and power decide how much capacity there is, which is why rate limits exist.
If you only remember a few things:
- Prices: the price of a fixed level of capability falls roughly 10-40 times a year, so the button that would have cost about $90,000 a month on GPT-4 in 2023 costs a few hundred dollars on a small 2026 model. The price of the best model doesn't fall, because labs keep adding a new top tier: OpenAI's flagship input price went from $1.25 per million tokens for GPT-5 to $10 for gpt-6-astra. Cost per task rises, because prompts grow, reasoning models bill thousands of hidden tokens per answer, agents loop, and demand barely responds to price. Some prices went up outright: retired Claude models kept on Bedrock at twice their launch price, Gemini Flash doubling on January 1, 2027, and GPU rental rates. See How the money moves.
- The bill: for the help-desk button, a frontier workhorse model such as Anthropic's Opus 5.5, OpenAI's gpt-6.1-sol or Google's Gemini 3.1 Pro costs about $6,000-18,000 a month at list prices, a mid-tier model $2,400-9,000, and an open-weight model at an inference provider $300-5,000. Reasoning multiplies the bill by three to four, and a 10-step agent by about 24 on the same model. Self-hosting loses money at this volume. Buying the same model a different way (batch, residency, a gateway's fee) moves the bill by 10-50%, while model tier and reasoning tokens move it by orders of magnitude.
- Compute: every big cloud said in 2026 that it was capacity constrained, and rate limits follow the compute: Anthropic raised its Opus API limits in May 2026 when new capacity from SpaceX's Colossus 1 data center arrived. Contracted capacity runs two to three times ahead of what's switched on, the biggest US power market has cleared at its price cap three auctions running, and Nvidia, Amazon, Microsoft and Google all invest in the labs that buy their compute. See Who holds the power.
- Routing: open-weight models trail the best closed model by about 6-12 points on Artificial Analysis's index and cost a half to a sixth as much, and gateways make switching a configuration change. But usage barely responds to price, the same open model behaves differently at different hosts, and buying speed costs throughput. See Open-weight vs closed models: what's real.
- Operations: in Datadog's customer data, rate limits were 60% of failed LLM calls in February 2026. Priority capacity is hard to buy, retirement notices are short (60-62 days at Anthropic in 2026, 20 days for one specialized OpenAI model), outage credits are small, and I found no SLA (service level agreement) that covers the quality of answers: Anthropic's postmortem of its 2025 serving bugs mentioned no refunds. See What mistakes cost.
- Rules: the EU AI Act's duties for general-purpose model providers have been enforceable since August 2, 2026, and the Commission sent its opening information requests on September 1. A builder becomes a "provider" only if fine-tuning uses more than a third of the original model's training compute. US export controls work as an allocation tool: no replacement for the rescinded diffusion rule, case-by-case H200 sales to China, and China blocking imports itself. Copyright costs land on the labs (Anthropic's $1.5 billion Bartz settlement; Thomson Reuters v. Ross), while Chinese open-weight models carry jurisdiction risk and can mean losing your vendor's IP indemnity, as How the rules work explains.
This guide is the supply side. How agents use these models (orchestration, memory, tool security, evals) is in AI agents: orchestration platforms, and I don't repeat it. How labs use model supply against coding tools, and what seats and usage cost there, is in AI coding agents; speech and realtime models are in Voice AI; and measuring tokens, latency and cost per request is in Observability.
The main players
These are the companies behind the diagram's parties, layer by layer, in no particular order.
Model labs (closed)
- What they do
- Train frontier models; sell tokens, subscriptions and enterprise deals
- Main players
- Anthropic, OpenAI, Google DeepMind, xAI (now part of SpaceX), Meta, Mistral
- What they control
- Quality, prices, rate limits, retention and retirement dates
Open-weight model makers
- What they do
- Publish weights anyone can download and serve
- Main players
- DeepSeek, Alibaba (Qwen), Moonshot (Kimi), Xiaomi (MiMo), OpenAI (gpt-oss), Google (Gemma)
- What they control
- The price floor for capable models, licences, and whether hosts serve their model faithfully
Cloud AI platforms
- What they do
- Resell many labs' models under one cloud contract
- Main players
- Amazon Bedrock, Microsoft Foundry, Google Vertex AI
- What they control
- Enterprise commitments, regions, their own SLAs, dates and indemnities
Inference providers
- What they do
- Serve open-weight models per token or per GPU-hour
- Main players
- Fireworks, Together AI, Baseten, Cerebras, Groq
- What they control
- Speed, price and serving quality for open models
GPU clouds
- What they do
- Rent GPUs and data centers on multi-year contracts
- Main players
- CoreWeave, Oracle (OCI), Nebius, Crusoe
- What they control
- Energized capacity, power contracts and the debt that pays for them
Chip makers
- What they do
- Design and sell the accelerators
- Main players
- Nvidia, AMD, Google (TPU, with Broadcom), Amazon (Trainium), Broadcom, Cerebras
- What they control
- Allocation of the scarcest chips and the highest margin in the chain
Gateways and routing
- What they do
- One API in front of many models: routing, fallback, budgets, caching
- Main players
- OpenRouter, LiteLLM, Portkey (Palo Alto Networks), Cloudflare AI Gateway
- What they control
- Switching, spend controls and a second ledger of every call
How they make money, and who's moving:
- Closed labs sell tokens, subscriptions and enterprise deals, directly and through clouds. Every lab revenue figure here is unaudited, and most are unofficial. Anthropic's come from a leaked draft of its IPO filing, as reported by Reuters on September 28, 2026: about $4.6 billion of 2025 revenue against $7.3 billion of compute and infrastructure costs, then $11.5 billion in April to June 2026 with an adjusted operating profit. OpenAI's are just as unofficial: Reuters and The Information, citing anonymous sources and internal documents, reported about $13 billion of 2025 revenue and $5.7 billion for January to March 2026, and its own draft filing was submitted confidentially. People familiar with each company put Anthropic's run-rate above $65 billion at the end of July and OpenAI's near $70 billion in September. The two book cloud resale differently (Anthropic counts the whole dollar, OpenAI only its share of some partner sales), so their numbers don't compare cleanly. xAI is the only lab with figures in a public filing: SpaceX's showed $3.2 billion of 2025 revenue and a $6.4 billion operating loss. Google doesn't break out Gemini revenue. Meta released its debut closed model in April 2026 and a paid API at about $1.25 per million input tokens in July. Mistral, Europe's regional leader, raised €3 billion at more than €21 billion on September 8, 2026.
- Open-weight makers earn little directly; their influence is price. DeepSeek sells its own API at half price off-peak. Chinese labs held all of the top ten open-weight spots on Artificial Analysis's index in April 2026, and Qwen overtook Meta's Llama in Hugging Face downloads in February, according to coverage of Mozilla's open-source AI report. Prices aren't only falling: Zhipu raised overseas API fees for its GLM models by 67-100% in February 2026. I found no revenue figure for DeepSeek, Alibaba's Qwen or Moonshot.
- Cloud platforms resell models at roughly the lab's price and earn on the commitment and the rest of the cloud bill. In the quarter to June 2026 AWS grew 37%, Azure 43% and Google Cloud 82%, to $24.8 billion. Bedrock has carried OpenAI's models since late April 2026, after Microsoft's API exclusivity ended (as reported), and Claude runs on all three clouds.
- Inference providers sell per token on shared ("serverless") endpoints and per GPU-hour on dedicated ones. Fireworks says it passed $1 billion of annualized revenue, and it raised about $1.5 billion at $17.5 billion in July 2026. Together AI raised $800 million at $8.3 billion, and Baseten $1.5 billion at $13 billion after its run-rate reportedly tripled to about $600 million in one quarter. Cerebras went public on May 14, 2026, and reported a GAAP gross margin of 14% in the quarter to June. Nvidia paid about $20 billion in December 2025 to license Groq's technology and hire its leaders, and GroqCloud carries on.
- GPU clouds sign multi-year take-or-pay contracts (the customer pays whether or not it uses the capacity), borrow against them and buy GPUs. CoreWeave's quarterly revenue more than doubled to $2.58 billion, but $640 million of interest left it with a $626 million net loss. Oracle's backlog reached $664 billion in September 2026, Nebius grew revenue 454%, and Crusoe, which develops OpenAI's Abilene site, was valued at about $30 billion.
- Chip makers keep the highest margins in the chain. Nvidia's revenue for the quarter to July 2026 was $96.2 billion, $89 billion of it from data centers, at a gross margin of about 74%. AMD's data center revenue roughly doubled to $6.7 billion, Amazon says its Trainium and Graviton business runs above $25 billion a year, and Google supplies TPUs, built with Broadcom, to other labs, Anthropic among them.
- Gateways earn a thin fee or nothing. OpenRouter takes 5.5% on credit purchases (nothing if you bring your own keys, up to $25,000 a month) and reportedly raised $113 million at about $1.3 billion in May 2026. LiteLLM is the most-used open-source proxy, Palo Alto Networks announced it would buy Portkey on April 30, 2026, Cloudflare includes its AI Gateway on all plans, and the clouds have their own routers.
As of October 2026. Most private-company figures here are self-reported or come from press coverage of rounds and leaks, so treat the list as a map to check.
Back to the help-desk company. Its app calls LiteLLM, which it runs itself, and the gateway checks the team's monthly budget and sends the ticket to Claude Sonnet 5.5 on Bedrock, so the spend draws down the company's existing AWS commitment. Bedrock runs the model in AWS's data centers, filled with Nvidia GPUs and Amazon's own Trainium chips. When Bedrock answers with a rate-limit error, the gateway sends the ticket to gpt-oss-120b at Fireworks instead, which serves it from rented Nvidia GPUs. One summary can touch five or six companies, each with its own meter and retirement calendar (the vendors in the story are illustrative).
How a request becomes a bill, step by step
Here's one press of the help-desk button. The model reads the whole prompt before the first word comes back and then writes the answer token by token. The bill counts input, output and cached tokens separately.
- Request sent. The app assembles the prompt: system instructions, tool definitions, the ticket history and any retrieved documents. Caching only works on an identical prefix, so the stable parts go at the start. In Datadog's data, system prompts make up 69% of input tokens. What can go wrong: the prompt exceeds the context window and the call fails; a new model generation counts the same text as about 30-35% more tokens, so a "same price" upgrade can cost more.
- Queued. The provider admits the request after checking the account's tier, spend cap and rate limits (requests and tokens per minute, per model), then assigns a service tier (standard, priority, flex or batch) and a region. What can go wrong: the request is Throttled, the exception on the diagram: a 429 error that says "retry later". Anthropic returns the same 429 at its monthly spend cap, where retrying won't help until the 1st of next month. Failed requests still count toward OpenAI's per-minute limit, so retries without a random delay keep the bucket empty.
- First token. In prefill, the model reads every input token in one parallel pass and builds a working memory of the prompt (the KV cache). Prefill is bound by raw compute and sets the time to the first token. A cached prefix is skipped, which is faster and cheaper. What can go wrong: a cache miss after someone edits the system prompt brings back the full price and wait; a reasoning model at its highest effort can think for minutes, so client timeouts set lower kill work you still pay for.
- Streamed. In decode, the model writes one token per step, re-reading its weights and the cache from memory each time, which is why decode is limited by memory bandwidth and output tokens cost more. A reasoning model writes hidden "thinking" tokens before the visible answer, and they're billed as output. What can go wrong: the answer hits
max_tokensand comes back incomplete, possibly after you've paid for thousands of reasoning tokens; the connection drops. The other exception, Failed over, is the gateway switching to another model or host, which it can do cleanly only before the first token arrives. - Billed. The provider meters uncached input, cache writes, cache reads, output (reasoning included), service tier and region, prices each line separately and bills prepaid credits, an invoice or a cloud marketplace. What can go wrong: the gateway, the provider's console and the cloud invoice disagree.
That's where the diagram's note comes from. On OpenRouter, the average prompt grew from about 1,500 tokens to more than 6,000 between 2024 and late 2025, and completions from about 150 to 400. Reasoning models and agents use many more tokens per task, which matters more than the price per token.
A few terms carry most of the explanations in this space:
- Why output costs more. Most labs price output at four to six times input (five times at Anthropic, OpenAI's gpt-6.1-sol and Gemini 3.8 Flash). Decode is slower per token than prefill, but the exact multiple is a pricing choice and isn't a measured hardware ratio.
- Reasoning tokens. OpenAI's docs say they're invisible, billed as output and range from a few hundred to tens of thousands per request. By late 2025 reasoning models produced more than half of OpenRouter's tokens.
- Batching and the speed trade. Serving engines such as vLLM, SGLang and Nvidia's TensorRT-LLM pack many users' requests onto the same GPUs. More users per GPU means cheaper tokens and slower answers for each. In SemiAnalysis's InferenceX benchmark, serving one open model about three times faster per user cost about four times as much per token, which is the economics behind "fast" tiers at twice the list price.
- Prompt caching. A repeated prefix costs a tenth of the input price or less to read, and on most Claude models cached reads don't count toward the per-minute input limit either. Yet only 28% of the LLM calls in Datadog's data showed any cache reads.
- Quantization. Storing a model's numbers at lower precision (8 or 4 bits instead of 16) makes it smaller and faster. On Nvidia's newest GPUs, 4-bit DeepSeek-R1 ran 2.2-3.5 times faster than 8-bit. Nvidia reports accuracy losses of about 1% or less on most tasks. It's also why "the same model" differs between hosts.
Cheat sheet: ways to buy inference
A lab's own API
- Good for
- The newest models first, fast modes and caching variants
- Watch for
- Tier caps; short retirement notices; the lab is your data processor
- Who sets the clock
- The lab
A cloud AI platform
- Good for
- One contract and commitment, private networking, residency, an SLA with credits
- Watch for
- Its own retirement dates; features arrive later; auto-upgrades on some deployments
- Who sets the clock
- The cloud
An inference provider (serverless)
- Good for
- Open-weight models, speed, low prices
- Watch for
- Quality and speed vary by host; quantization; models come and go
- Who sets the clock
- The provider and the model maker
A dedicated endpoint or provisioned throughput
- Good for
- Predictable latency and capacity
- Watch for
- You pay for idle hours (about $5.50-13 per GPU-hour at providers)
- Who sets the clock
- The contract
Self-hosting on a GPU cloud or your own hardware
- Good for
- Control: residency, fine-tuned weights, no retirement clock
- Watch for
- Utilization, operations, redundancy
- Who sets the clock
- You
A gateway in front of any of these
- Good for
- Failover, budgets, multi-model routing, one log
- Watch for
- A fee or your own ops, an extra hop, supply-chain risk
- Who sets the clock
- You
Which way to buy inference? Five questions
- Frontier or not? Classification, extraction, summaries and chat usually don't need it, and a small or open-weight model at a tenth of the price or less is close enough. Hard agentic and coding work still does.
- Does it need to reason? Reasoning multiplies the bill by three to four on a chat-shaped task. Set the effort per feature, cap
max_tokensand measure tokens per task as well as the price per token. - Can it wait? If an answer can come back within 24 hours, batch halves the bill. If it must be fast, priority or fast tiers cost 1.75-2 times list, and capacity for them may not be for sale.
- What does the contract need? Residency (about 10% extra), zero data retention, an uptime SLA with credits and an IP indemnity each narrow the routes. A cloud platform usually adds an SLA with credits, which a lab's standard tier doesn't have; check that the indemnity covers the specific model.
- When does it go away? Every route except self-hosting has a retirement clock, some as short as 60 days. With three labs in use, expect a forced re-test every few weeks.
My defaults: for the help-desk company I'd start on a mid-tier closed model through the cloud it already has a commitment with, with caching designed in and reasoning off unless an eval shows it helps. I'd put a gateway in front from day one, with per-team budgets, a fallback to another provider, and an open-weight model for the easy half of the traffic once our evals say it's good enough. I'd pin dated snapshots, keep a regression set of a few hundred real tickets that runs on every model change, and track cost per resolved ticket. I wouldn't self-host until volume is ten times higher or a customer requires it.
The primitives
01
Entity and identity
What is the unit of record, and how do we know it is the same one?
"The model" has at least three parts: a family (Sonnet, gpt-6.1-sol), a snapshot ID (the exact version), and the host or deployment that serves it. On Azure you call a deployment name you created, and on Bedrock the same Claude model has its own ID and its own retirement date. Open-weight models add the precision and serving stack, which is why OpenRouter lets you filter hosts by quantization. The billing identity is a tree of organization, workspace or project, and API key, and rate limits attach to the organization (Anthropic, OpenAI) or the project (Google), per model, so teams in one organization share its headroom.
The law names its own entities. Under the EU AI Act the "provider" places a model on the market and the "deployer" uses it, and a company that fine-tunes with more than a third of the original training compute becomes the provider of the result. An IP indemnity keys on plan, service and model together, so one API key can be covered for one model and not another. US export rules now follow a company's ultimate parent, wherever its subsidiary is registered (guidance of May 31, 2026).
02
State and lifecycle
What states exist, and what moves an entity between them?
A request goes from admitted to queued, prefilling, decoding and streaming, and ends complete, incomplete or failed. A model goes from preview to generally available, legacy, deprecated (still works, retirement date set) and retired. A prompt cache is written, stays warm while used, and dies when it expires or when the prefix, tools or reasoning settings change. An account climbs tiers with spend history and pauses at its spend cap until the next month. Capacity goes from contracted to built, energized and allocated.
Retirement clocks differ by vendor and are getting shorter at some:
| Vendor | Typical notice before retirement | Notes |
|---|---|---|
| Anthropic | At least 60 days; 60-62 days for every 2026 notice | Down from 181-189 days in 2025; Opus 4.1 lived exactly 12 months |
| OpenAI | At least 6 months for generally available models, 3 for specialized ones, about 2 weeks for previews | A restricted cyber model got 20 days in September 2026 |
| Google (Gemini API) | Dates are "earliest possible"; previews retire 2-4 months after release | Some previews run in production for months |
| Microsoft Foundry | 18-month lifecycle for most generally available models, at least 60 days' notice | Standard deployments auto-upgrade; "retirement dates aren't extendable" |
| AWS Bedrock | "End of life no sooner than" a stated date; Legacy for 6 months or 45 days | Legacy blocks new customers and provisioned throughput |
| Mistral | No formal policy; about 24 days observed for one model | |
| Self-hosted open weights | None | Hosted versions at clouds still retire |
The same model can retire on different dates at the lab and at each cloud. Products retire too: OpenAI ends its Assistants API, hosted evals and self-serve fine-tuning jobs between August 2026 and January 2027.
03
System of record and ledger
Who owns the truth, and how do systems reconcile?
| Fact | System of record |
|---|---|
| What one call used | The provider's usage record per request: input, cache write, cache read, output, tier, region |
| What the company spent | The provider's usage and cost API, the gateway's log and the cloud invoice, which rarely agree |
| How much headroom is left | Rate-limit headers on every response |
| Whether the provider was up | Its status page, which the provider writes and which can lag |
| What a lab has contracted to buy | Backlog at the clouds; take-or-pay schedules in filings |
On cloud marketplaces Anthropic converts usage into consumption units of $0.01 and reports them to AWS or Azure, so a discount shows up as fewer units. A private lab's "run-rate" is a recent month times twelve. Your own gateway log is the only record of errors from your side, and it's the evidence an SLA claim needs.
04
Rules and policy
What logic decides outcomes, and who can change it?
Access is rationed by rules the provider writes and can change:
Anthropic
- Tiers
- Evaluation, Start, Build, Scale, Custom
- Monthly spend cap
- $500 / $1,000 / $200,000
- Example limit at the top self-serve tier
- Opus 5.5: 10,000 requests and 10 million input tokens a minute; Fable 5.x: 4,000 and 4 million
OpenAI
- Tiers
- Free, Build, Launch, Grow (after $5, $100 and $500 of purchases)
- Monthly spend cap
- $100 / $500 / $5,000 / $200,000
- Example limit at the top self-serve tier
- gpt-6 Astra, Sol and Terra: 15,000 requests and 40 million tokens a minute
Google (Gemini API)
- Tiers
- Free, Tiers 1-3 (by spend and account age)
- Monthly spend cap
- $250 to $100,000 or more, plus a 10-minute spend limit
- Example limit at the top self-serve tier
- Per project, per model
The top model gets the tightest limit, and Anthropic calls its limits "maximum allowed usage, not guaranteed minimums". New organizations start lower "to prevent fraud and abuse", and fast ramps hit acceleration limits. Gateways add the customer's own policy (route by price or speed, cap the price per token, use only zero-retention hosts, set budgets per team), and the labs' usage policies restrict high-stakes uses without human review.
05
Effective dating
Which version of the rule applied at that moment?
Dates I'd keep on a calendar as of October 2026:
| Change | Effective | Status (Oct 2026) |
|---|---|---|
| EU AI Act duties for general-purpose model providers | August 2, 2025 | In force |
| California SB 53 for frontier developers | January 1, 2026 | In force |
| H200-class chips to China move to case-by-case review | January 13, 2026 | In force; China limits imports |
| EU AI Office enforcement powers, including fines | August 2, 2026 | In force; first requests September 1 |
| Claude Sonnet 4.5 retires | November 30, 2026 | Upcoming |
| EU Product Liability Directive covers software | December 9, 2026 | Upcoming |
| OpenAI GPT-5 and o3 snapshots retire | December 11, 2026 | Upcoming |
| Gemini 3.6-3.8 Flash prices double | January 1, 2027 | Upcoming |
| New York RAISE Act | January 1, 2027 (some sources say July 1) | Upcoming |
| Older general-purpose models must comply with the AI Act | August 2, 2027 | Upcoming |
Some dates float. Retirement dates are "not sooner than" (Anthropic, AWS) or "earliest possible" (Google), and introductory prices can end or be made permanent: Anthropic kept Sonnet 5's and cancelled a rise planned for September 1, 2026. DeepSeek even prices by time of day.
06
Interfaces and standards
What format and protocol do counterparties speak?
OpenAI's Chat Completions and Responses formats are the de facto interface: Bedrock exposes them ("change the base URL and API key"), most inference providers do, and Anthropic runs an OpenAI-compatible endpoint next to its own Messages API. That makes switching mechanically cheap. Everything around the call is unstandardized: usage fields, rate-limit headers, deprecation metadata and reasoning controls differ per provider. Tokenizers differ too, so the same text can be about 30-35% more tokens on Anthropic's newest models than on its older ones.
Measurement is its own layer. Artificial Analysis measures time to the first token and output speed from one server over 72 hours, MLPerf audits hardware results that vendors submit, and SemiAnalysis's InferenceX turns benchmarks into cost per million tokens. Model makers have started checking hosts: Moonshot publishes a Kimi Vendor Verifier and OpenAI added a compatibility test for gpt-oss. Below the token the units change again: GPU-hours for renting chips, gigawatts of power capacity for compute deals, and megawatt-days in power markets.
07
Networks and counterparties
Who sits between us and the outcome, and what do they want?
| Party | What they control | What they earn |
|---|---|---|
| Lab | Quality, prices, limits, retention, retirements | Tokens, subscriptions, enterprise deals |
| Cloud platform | Procurement, regions, its own SLA and dates | Resold tokens; commitments drawn down |
| Inference provider | Speed and price for open models | Tokens, dedicated endpoints, GPU rental |
| GPU cloud | Energized capacity | Multi-year take-or-pay contracts |
| Chip maker | Allocation of new chips | Hardware at about 74% gross margin (Nvidia) |
| Gateway | Routing and budgets | A small fee, or nothing (open source) |
| Grid operator and utility | Power and connection dates | Capacity payments, set by auction |
Money flows down the chain and equity flows back up it: Nvidia invests in OpenAI and in GPU clouds and providers, Amazon and Google in Anthropic, Microsoft owns about 27% of OpenAI and has invested in Anthropic, and AMD gave OpenAI warrants for up to 160 million of its shares. Amazon sells Claude on Bedrock, holds Anthropic notes, supplies it Trainium and sells its own Nova models, all at once.
08
Regulatory layering
Jurisdiction × activity × entity type: is it a license or a certification?
Model makers
- US
- California SB 53 (frontier developers above 10^26 operations); New York RAISE from 2027
- EU
- AI Act duties for general-purpose models; Code of Practice
- China and elsewhere
- Chinese labs mostly outside the EU Code
Chips and weights
- US
- Export controls: case-by-case H200 licences for China, approved entities in the UAE
- EU
- China and elsewhere
- Beijing approves or blocks imports
Data
- US
- Sector rules; court preservation orders
- EU
- GDPR; residency demands
- China and elsewhere
- Data stays in-country in many sovereign deals
Products
- US
- Deployer liability (case law)
- EU
- Product Liability Directive from December 9, 2026
- China and elsewhere
The same model reached through the lab, Bedrock or Vertex can come with different retention, residency, SLA and indemnity terms, so the route is a legal choice too. For most builders the law arrives through vendor terms and procurement questionnaires.
09
Exceptions and reversals
What goes wrong, and how is it undone?
Rate limit or spend cap hit
- Who starts it
- The provider's limiter
- Clock
- Seconds, or until the 1st of next month
- The way back
- Back off with jitter; fail over; ask for a tier increase
Model retired
- Who starts it
- The lab or cloud
- Clock
- 60 days to 6 months; 20 days seen
- The way back
- Re-test on the replacement; pin the next snapshot
Bad serving change
- Who starts it
- The provider
- Clock
- Hours to weeks to detect
- The way back
- Provider rollback (GPT-4o's sycophantic update came out about four days after release); your own evals
Price step-up
- Who starts it
- The provider
- Clock
- A dated announcement, or a promotion ending
- The way back
- Route elsewhere; renegotiate
Capacity product withdrawn
- Who starts it
- The provider
- Clock
- Immediate for new buyers
- The way back
- Existing contracts honoured; buy elsewhere
Compute deal cancelled or reshaped
- Who starts it
- Either party
- Clock
- 90 days for xAI's capacity to Anthropic; years for take-or-pay
- The way back
- Other sites, other chips
Rule reversed
- Who starts it
- A government
- Clock
- Months
- The way back
- Switchable architecture; exit rights in contracts
Reversals run both ways. OpenAI's "$1.4 trillion" of compute ambitions became about $600 billion through 2030 in reporting by February 2026, and the expansion of the Abilene site beyond 1.2 GW was dropped. Export rules went from ban to unban to a revenue share that was never codified, all within 2025.
10
Liability allocation
When it fails, who pays?
| Failure | Who absorbs it | Mechanism |
|---|---|---|
| Outage | The builder, minus a small credit where an SLA exists | Bedrock: 99.9% per region, credits of 10-100% of that region's charges as future credit; lab standard tiers: best effort, no credit |
| Silent quality drop | The builder | No SLA found that counts wrong answers as downtime |
| Retirement and migration | The builder | Notice periods; no compensation |
| Copyright claim on an output | The vendor, if the indemnity's conditions are met | Paid plan, covered model, output unmodified, filters on; usually outside the liability cap |
| Copyright claim over training data | The lab | Bartz, the New York Times case; reaches builders only through prices |
| Harm from an answer | The deployer by default | Air Canada's chatbot case; consumer cases against model makers are pending or settled |
| Unused compute | The lab, then its investors | Take-or-pay contracts: about 80% of Anthropic's commitments, as reported |
| GPU cloud failure | Lenders, then customers mid-contract | Debt against contracts |
Builders carry price and capacity risk, labs carry volume risk, GPU clouds carry financing and obsolescence risk, and household electricity customers carry part of the grid cost. Liability caps in the labs' and clouds' terms are typically 12 months of fees.
What's different here
How the money moves
The builder pays per token, and the money flows down to whoever owns the chips, the buildings and the power. List prices per million tokens in October 2026, input / output:
| Tier | Examples |
|---|---|
| Top tier | Anthropic Fable 5.1 $10 / $50; OpenAI gpt-6-astra $10 / $50; gpt-5.5-pro $30 / $180 |
| Frontier workhorse | Anthropic Opus 5.5 $4 / $20; OpenAI gpt-5.5 $5 / $30 and gpt-6.1-sol $2 / $10; Google Gemini 3.1 Pro Preview $2 / $12; xAI grok-4.7 $2 / $6 |
| Mid tier | Anthropic Sonnet 5.5 $2 / $10; Google Gemini 3.8 Flash $0.75 / $3.75 (doubling on January 1, 2027); Meta Muse Spark 1.1 about $1.25 / $4.25; Mistral Large $0.50 / $1.50 |
| Small | OpenAI gpt-6-luna and Anthropic Haiku 5.5 $0.10 / $0.50; Google Gemini 3.5 Flash-Lite $0.30 / $2.50 |
| Open-weight at a provider | gpt-oss-120b $0.15 / $0.60; DeepSeek V4.1 Flash $0.30 / $1.20; GLM 5.3 $1.40 / $4.40; Qwen 3.8 Max $2 / $6; Kimi K3 $3 / $15 |
And the modifiers that apply on top:
| Modifier | Typical effect | Where |
|---|---|---|
| Cached input | 0.025-0.1 times the input price; writing the cache costs 1.25-2 times at Anthropic | OpenAI, Anthropic, Google (which also charges storage per hour), providers |
| Batch | Half price, answers within about 24 hours | OpenAI, Anthropic, Google, Bedrock, Fireworks, Mistral |
| Flex | Half price, synchronous but can be pre-empted | OpenAI, Google, Bedrock |
| Priority or fast | +25% (Fireworks), +75% (Bedrock), 1.8 times (Google), 2 times (OpenAI Fast, Anthropic fast mode), 6 times (OpenAI Ultrafast) | Capacity permitting |
| Residency | +10%; +50% for Fireworks's US-hosted variants | OpenAI regional processing, Anthropic US-only inference, Claude regional endpoints on clouds |
Under these sheets, prices move in different directions. The price of a fixed level of capability falls: a16z measured about 10 times a year at a constant benchmark score, and Epoch AI found GPT-4-level performance on a hard science benchmark getting 40 times cheaper a year. The top of the ladder stays put, because each generation adds a rung: OpenAI's flagship launch prices fell from $30 per million input tokens for GPT-4 to $1.25 for GPT-5 and climbed back to $10 for gpt-6-astra, while Anthropic cut Opus from $15 to $4 and added Fable at $10. And the cost per task rises: on Artificial Analysis's index Opus 4.6 used twice the output tokens of Opus 4.5 on the same tests. In OpenRouter's 100-trillion-token study a 10% price cut brought only 0.5-0.7% more usage, so spending grows from new uses such as agents and coding far more than from cheaper tokens.
Prices also went up. Bedrock keeps retired Claude 3.5 Sonnet under "extended access" at twice its launch price, Gemini 3.6-3.8 Flash list a 2027 price twice the 2026 one, and Zhipu, Moonshot and MiniMax raised some API prices. Compute got dearer too: one-year H100 rental contracts rose about 40% between October 2025 and March 2026, and B200 spot prices more than doubled in six weeks in spring 2026.
A month of the help-desk button
My assumptions, from list prices on October 9, 2026: 1 million requests a month, each with 2,000 input tokens (1,500 of them a shared, cacheable system prompt) and 500 output tokens, so 2 billion tokens in and 500 million out. The reasoning variant adds 2,000 thinking tokens per request, billed as output. The agentic variant turns each press into a 10-step agent task of about 200,000 input tokens (90% cache hits) and 8,000 output.
Top tier (Fable 5.1, gpt-6-astra)
- Base month
- $30,000-45,000
- With reasoning on
- About $130,000
- As a 10-step agent
- Not computed; likely far higher
Frontier workhorse (Opus 5.5, gpt-6.1-sol, Gemini 3.1 Pro)
- Base month
- $6,000-18,000
- With reasoning on
- $26,000-58,000
- As a 10-step agent
- $140,000-300,000
Mid tier (Sonnet 5.5, Gemini 3.8 Flash, grok-4.7)
- Base month
- $2,400-9,000
- With reasoning on
- About $10,000-26,000
- As a 10-step agent
Small (gpt-6-luna, Haiku 5.5)
- Base month
- $300-450
- With reasoning on
- About $1,300
- As a 10-step agent
Open-weight at a provider (gpt-oss-120b to GLM 5.3)
- Base month
- $300-5,000
- With reasoning on
- About $1,800-14,000
- As a 10-step agent
- About $17,000 (DeepSeek V4.1 Flash)
Top open-weight (Kimi K3)
- Base month
- $13,500; about $20,000 US-hosted
- With reasoning on
- About $43,000
- As a 10-step agent
Self-hosted small open model on rented H100s
- Base month
- $7,000-30,000, people included
- With reasoning on
- Similar; bound by GPUs
- As a 10-step agent
Self-hosted large open model on B200 nodes
- Base month
- $58,000-152,000 plus people
- With reasoning on
- Similar
- As a 10-step agent
2023: GPT-4
- Base month
- $90,000
- With reasoning on
- As a 10-step agent
What the table says:
- Model tier dominates. The same feature costs about $300 or about $45,000 depending on the tier.
- Reasoning: 3-4 times. On Opus 5.5 with caching, the bill goes from about $12,300 to about $52,300.
- Agents multiply by about 24. On Opus 5.5 an agent task costs about $0.30, or $300,000 a month, against $12,300 for the chat-shaped button. Without caching the agent's input alone would cost about $1 million a month, so cache hygiene is worth three to four times on agent work.
- The route moves it less. Caching cuts the base bill by about 30%, batch halves any line if answers can wait, residency adds 10%, and OpenRouter's fee adds 5.5%. Priority doubles it.
- Self-hosting loses here. Two to four rented H100s at $2-4 an hour plus a quarter to one engineer come to $7,000-30,000 a month against about $600 at a provider, 10 to 50 times more, because the provider runs newer chips at high utilization. Self-hosting pays at tens of billions of tokens a month with steady load, or when control is the requirement: residency, fine-tuned weights, or a model that never retires.
- Against 2023. For a fixed task the bill fell from $90,000 to a few hundred dollars on a small model, but a team that moved to the top tier with reasoning pays more than it did in 2023.
Where the dollar goes
- Lab direct: customer to lab, at list less any negotiated discount (I found few public ones), then to the lab's compute. Gross margins are unaudited estimates for every lab: The Information, as reported by Reuters, put OpenAI's adjusted gross margin at 33% in 2025, down from 40% as inference costs rose, and I found no reliable figure for Anthropic, Google's Gemini or xAI.
- Cloud marketplace: customer to AWS, Azure or Google, drawing down a commitment, then to the lab. Marketplaces were 47% of Anthropic's 2025 revenue in its leaked draft filing, as reported.
- Inference provider: customer to Together, Fireworks or Baseten, then to GPU clouds or their own clusters, then to Nvidia. Cerebras, the one with public numbers, reported a GAAP gross margin of 14% in the quarter to June.
- GPU cloud and chips: lab or hyperscaler to CoreWeave, Nebius or Oracle, then to lenders, utilities and Nvidia, which earned $59.7 billion of net income on $96.2 billion of revenue in one quarter, the richest seat in the chain.
Who holds the power
Compute is the binding constraint, and power follows whoever controls a piece of it.
- Clouds and the grid. Alphabet said in April 2026 that it was "compute constrained in the near-term", Microsoft's CFO said in July that it "remains capacity constrained", and AWS called its Trainium2 chips "largely sold out". The binding input is moving from chips to power: CoreWeave had about 1.5 GW active against 3.7 GW contracted in June 2026. PJM, the grid operator for much of the eastern US, cleared its capacity auction at the price cap three times running and ended about 6,800 MW short of its reliability target in July 2026, with data centers the main driver of new load.
- Nvidia allocates the scarcest chips, keeps the highest margin and invests in labs, GPU clouds and providers. TPUs, Trainium, AMD and Broadcom's custom chips check it in part.
- Labs over builders. Labs set prices, limits, retention and retirement dates, and ration when capacity is short: peak-hour cuts in Claude Code in March 2026, Priority Tier withdrawn, OpenAI limiting GPT-4.5 to its top plan in 2025 because it was "out of GPUs". Downstream, GitHub paused Copilot Business signups in April 2026 for lack of capacity. When capacity lands, limits rise, as Anthropic's did on May 6, 2026.
- Clouds and labs. The clouds are the labs' suppliers, investors, distributors and competitors, and the labs are the anchor tenants behind the clouds' backlogs.
- Contracts. Anthropic's commitments of at least $518 billion run about a decade, around 80% non-cancelable, per the leaked draft filing as reported. OpenAI's are reported at about $600 billion through 2030, including more than $300 billion at Oracle and a 750 MW deal that Cerebras's own filing values at more than $20 billion. Meta signed about $21 billion more with CoreWeave in April 2026. xAI builds its own Colossus sites, which is why Anthropic could rent capacity from SpaceX on 90 days' notice.
- Suppliers as investors. OpenAI's $122 billion round at $852 billion, closed around April 2026, included $30 billion from Nvidia and $15 billion from Amazon (plus $35 billion conditional) tied to 2 GW of Trainium. Amazon marked its Anthropic notes up from $42.2 billion to $97.9 billion in one quarter, according to an analysis of the leaked filing. These loops prop up demand signals and concentrate risk.
- Providers and gateways compete on speed, price and switching, with thin margins; model makers can shame hosts (Moonshot found gaps between third-party and official Kimi APIs "widespread"), and gateways are being folded into security companies.
- Builders hold demand and little contractual power. Their leverage is spend commitments, a second provider and moving easy traffic to cheaper models.
How the rules work
The EU AI Act. Duties for providers of general-purpose AI models have applied since August 2, 2025, and the Commission's enforcement powers, including fines of up to 3% of worldwide turnover or €15 million, since August 2, 2026; models already on the market have until August 2, 2027. Every provider owes technical documentation, a copyright policy that honours opt-outs and a public summary of its training data, and models presumed to carry systemic risk (trained with more than 10^25 operations) add evaluations, incident reporting and security. As of October 7, 2026 the Code of Practice, the accepted way to show compliance, had 22 signatories including Amazon, Anthropic, Google, IBM, Microsoft, Mistral and OpenAI; xAI signed only the safety and security chapter, Meta declined, and DeepSeek, Alibaba and Moonshot aren't on the list. On September 1, 2026 the Commission confirmed requests for information to more than 30 unnamed AI companies, and no fines have been issued. For a builder the trigger is narrow: more than a third of the original model's training compute spent on modifying it.
Export controls. The AI Diffusion Rule was rescinded on May 13, 2025, and in July 2026 the Commerce Department's under secretary told Congress there would be no replacement. Policy now works through licences. Since January 13, 2026 Nvidia's H200 and AMD's equivalent can go to China case by case, capped at half of US volume per chip, with a 25% fee collected as a tariff and Blackwell excluded. China then blocked H200 imports in mid-January and later let in about 10,000 units each to ByteDance and Tencent, as reported, and Nvidia's outlook assumes no China data center revenue. The UAE got licence-free access for approved entities from July 10, 2026. Builders feel export rules as capacity and prices, and directly only if they give restricted parties remote access.
US states. California's SB 53 has applied to frontier developers since January 1, 2026, with safety frameworks, incident reporting and penalties up to $1 million per violation; I found no enforcement yet. New York's RAISE Act follows in 2027, and no federal preemption had passed by October 9, 2026.
Copyright. The costs land on the labs. Anthropic's $1.5 billion Bartz settlement covered about 482,000 books at about $3,100 each and got final approval in July 2026; the underlying ruling held that training on lawfully bought books was fair use and pirated libraries weren't. On September 29, 2026 the Third Circuit held in Thomson Reuters v. Ross that training a non-generative legal AI on Westlaw's headnotes wasn't fair use. It was the first appellate ruling on AI training and left the generative question open. Meta won on fair use in Kadrey in 2025, and the New York Times' case against OpenAI and Microsoft reached summary judgment briefing in September 2026. In Europe a Munich court held in November 2025 that OpenAI's memorizing song lyrics infringed, and an English court held that model weights aren't infringing copies; both are on appeal. Settlements can force retirements downstream: Warner Music's settlement with Suno requires it to retire its current models.
Indemnities. Paid tiers come with IP indemnities, each with conditions. Google's list covers its own models on Vertex, and third-party models in its catalogue aren't listed. AWS gives an uncapped indemnity for its own Nova models and narrower cover for others. Microsoft requires its content filters to stay on. Anthropic's commercial terms carry an uncapped IP indemnity, and OpenAI offers Copyright Shield. Routing to an open-weight or third-party model can drop you outside the list.
Chinese open-weight models. They're cheap and good, and they carry jurisdiction risk. NIST's AI standards center found in September 2025 that agents built on DeepSeek R1 were about 12 times more likely to follow hijacking instructions than US models. Several states ban DeepSeek on state devices, Congress opened a probe in July 2026, and the administration was reported to be weighing procurement bans and sanctions. If that happens, a cost-driven routing choice becomes a compliance problem.
Availability and data. Bedrock promises 99.9% a month per region, with credits of 10%, 25% or 100% of that region's charges as future credit. Anthropic's standard tier is "best-effort", its Priority Tier only "targets" 99.5% and isn't sold to new buyers, and OpenAI's 99.9% SLA sits in an enterprise tier with a minimum commitment. OpenAI keeps abuse-monitoring logs for up to 30 days, Anthropic keeps nothing by default except for its top models, which require 30 days, and Bedrock doesn't store prompts. A court can override any of it: in May 2025 a preservation order in the New York Times case made OpenAI keep output logs.
Sovereignty. "Sovereign" can mean an EU-run region of a US company (the AWS European Sovereign Cloud, opened January 15, 2026), an EU-owned lab (Mistral) or public compute (the EU's AI gigafactories, with bids due November 12, 2026). These regions tend to launch with open-weight models: the AWS one's debut model family on Bedrock, in September 2026, was Google's Gemma 4.
What mistakes cost
Let's say the help-desk company runs its button on a mid-tier closed model and pays $4,500-13,500 a month, about $54,000-162,000 a year. My rough arithmetic for what goes wrong and who pays:
- A three-hour outage in a 720-hour month cuts uptime to 99.58% and loses about 4,200 requests. On a lab's standard tier the credit is zero; on Bedrock it's 10% of that region's charges, about $450-1,350, as a future credit. With a gateway failing over to another provider, most of the loss becomes about $20-60 of extra tokens and some quality variance.
- A silent quality drop like Anthropic's 2025 bugs: if 5% of requests give worse answers for three weeks, that's about 37,500 degraded summaries, which no SLA treats as downtime.
- A retirement with 60 days' notice costs 10-40 engineer-days of re-testing ($8,000-60,000 loaded), and up to 30% more tokens if the new tokenizer counts more, about $16,000-49,000 a year on this bill.
- A third-party copyright claim is the vendor's to defend if the output qualifies for the indemnity; anything else is capped at 12 months of fees.
- A ban on a Chinese open model for one customer segment would push that traffic back to a closed model at three to ten times the price, plus the re-testing.
The operational events cost far more, in expectation, than the legal ones. The record so far:
| Case | What happened | Who paid |
|---|---|---|
| Overlapping outages, September 3, 2026 | OpenAI's ChatGPT and Codex were down about 34 minutes, Claude for 3 hours 6 minutes, and xAI's Grok at the same time | Builders and users; no credits reported |
| Anthropic serving bugs, August-September 2025 | Three infrastructure bugs; in the worst hour 16% of Sonnet 4 requests were misrouted, and about 30% of Claude Code users had at least one bad message; benchmarks, safety evals and canaries missed it | Builders; the postmortem mentions no refunds |
| OpenAI GPT-4o update, April 2025 | A sycophantic update was rolled back about four days after release | Users; OpenAI's reputation |
| AWS us-east-1, October 20, 2025 | About 15 hours of disruption from a DNS failure | Customers; Moody's put insured losses at a mean of $22 million, and a parametric insurer paid claims |
| LiteLLM on PyPI, March 24, 2026 | Compromised releases of the gateway package stole credentials | Users, who rotated every key |
| OpenAI and Mixpanel, November 2025 | A vendor breach exposed API users' names, emails and organization IDs, though no prompts or keys | Users, at risk of phishing |
| DeepSeek, January 2025 | An open database exposed more than a million log lines, chat history and keys among them | Users |
| Nvidia H20, April 2025 | A new licence requirement for China | Nvidia: a $4.5 billion charge |
Several of these are old failures in shared dependencies: DNS, a package registry, a sub-processor. The new kind is the quality drop that no contract counts.
Open-weight vs closed models: what's real
My view as of October 2026: below the frontier, an open-weight model is good enough for most chat, summarization and extraction work, at a half to a sixth of the price, and a gateway makes moving to it a configuration change. At the frontier, especially agentic coding, the gap reopens with each closed release and most enterprise money stays closed. The real switching cost is evals, prompt rework, checking each host and reading the contract, more than changing the API call.
How big the gap is
- On the main composite index: in Artificial Analysis's April 30, 2026 snapshot, the best open-weight models (Moonshot's Kimi K2.6 and Xiaomi's MiMo V2.5 Pro) scored 54 against 60 for OpenAI's GPT-5.5 at its highest effort, a 6-point gap, and open models held 9 of the 13 spots on its frontier of intelligence against price. By October, on a newer version of the index, Xiaomi's MiMo-V2.6-Pro scored 46 against 58 for Anthropic's Opus 5.5 and 53 for the best of OpenAI and Google, about 12 points. Index versions rescale, so compare only within one.
- On hard tasks the gap is wider: 43-46% for the best open models against 61% for closed ones on Artificial Analysis's hard terminal-coding tasks in April.
- Who makes them: the top six open-weight models in October were all Chinese. The US ones are OpenAI's gpt-oss, Google's Gemma and Nvidia's Nemotron, and Meta has moved its frontier work to closed models.
What open costs
The price gap is large at the small and mid end: gpt-oss-120b at $0.15 / $0.60 and DeepSeek V4.1 Flash at $0.30 / $1.20 per million tokens, against $2 / $10 for a closed mid-tier model. It narrows at the top, where Kimi K3 lists at $3 / $15, close to Opus 5.5's $4 / $20. Serving a big open model yourself takes a full node of eight B200s per replica.
Same model, different host
- Quality varies: on a 2025 maths test run 32 times, gpt-oss-120b scored 93.3% at seven hosts, 86.7% at Groq, 83.3% at Amazon, 80% at Azure and 36.7% at one small host. Hosts differ in precision, chat templates and serving bugs.
- Speed varies more: about 37 times between the slowest and fastest gpt-oss host in October 2026, and up to 8.9 times on price.
So "route to the cheapest host of the same model" needs a per-host eval, and a failover to the same model on another host can still change your answers.
Where the money actually goes
Usage share and money share point different ways. On OpenRouter, open-weight models were about a third of tokens by late 2025. Datadog found more than 70% of organizations using three or more models. But a 10% price cut moved usage by only 0.5-0.7%, and four in ten users of one 2025 Claude model were still on it five months later. In Menlo Ventures' late-2025 survey (Menlo is an investor in Anthropic), Anthropic, OpenAI and Google took 88% of enterprise LLM spend between them, and by mid-2025 the open-source share of enterprise workloads had fallen from 19% to 13%. Quality and habit decide the vendor; price decides the long tail.
Jurisdiction and contracts
Open weights move the contract risk to you. Self-hosted, there's no retirement clock and no indemnity; at a provider, the model maker isn't your counterparty, and most IP indemnities name only the vendor's own models. Chinese labs aren't Code of Practice signatories and US restrictions were under discussion in 2026, while sovereign regions tend to launch with open models. For a team selling to US government-adjacent buyers I'd keep Chinese-model routes behind a switch I can turn off per customer.
Questions to ask a provider
- Which precision and serving engine do you run this model on, and do you pass the model maker's own verification test?
- What are your median and 95th-percentile time to the first token and output speed on our prompt sizes, measured by someone else?
- What notice do you give before dropping a model from serverless, and what happens to our dedicated endpoints?
- Is our data retained or used, by you or by any sub-processor, and can you give zero retention in writing?
- Which of your models, if any, come with an IP indemnity?
What usually goes wrong
| Symptom | Likely cause | First thing to check |
|---|---|---|
| Bursts of 429 errors | Rate limits per model, retries without jitter, a fast ramp | Rate-limit headers, retry logic, a second provider |
| The feature stops until the 1st | Monthly spend cap reached | Tier caps and alerts at 50% and 80% |
| The bill jumps with no traffic change | Reasoning effort raised, cache broken by a prompt edit, a new tokenizer | Tokens per request by type; cache hit rate |
| Answers got worse with no release | A serving change at the provider, or a failover to another host | Online evals by model and host |
| A model is going away | Retirement notice | The regression set; pinned snapshot IDs |
| Legal blocks a route | Retention, residency or indemnity doesn't cover that model | The data and indemnity terms per model and venue |
| A gateway leaks keys | Supply-chain compromise | Pinned package versions; key rotation |
| Self-hosting costs more than planned | Low utilization, idle redundancy, people | Tokens per GPU-hour against the provider's price |
Words that mean something else here
| Term | What you'd assume | What it means here |
|---|---|---|
| Token | A word | A piece of text in one model's tokenizer; the same text can be about a third more tokens on a newer model |
| Latency | One number | Time to the first token (for reasoning models, the first thinking token), time to the first visible token, or end to end |
| Throughput | Speed | Tokens a second one user sees, or tokens a second a GPU produces; they trade against each other |
| Context window | What the model can use | The maximum input and output; quality falls well before it |
| Priority | Same thing everywhere | Committed capacity (Anthropic, no longer sold), a 2x pay-as-you-go tier (OpenAI, now "Fast"), 1.8x per request (Google), +75% (Bedrock) |
| Fast mode | A faster model | The same model served faster at about twice the price; at OpenAI, the renamed priority tier |
| Deprecated | Gone | Still works, with a retirement date set; in Azure's API, the value Deprecated means retired |
| Open source | Open | Usually open weights under a licence, without training data or code; the AI Act's open-source exemption needs a free licence and no monetization |
| Provider | A vendor | An inference company, or the AI Act's party that places a model on the EU market, which a heavy fine-tuner can become |
| Uptime | The model works | The endpoint accepts requests; nothing about the quality of the tokens |
| Run-rate | Revenue | Recent revenue times twelve; usage businesses aren't contracted for it |
| Gigawatt | Chips | Power capacity in a compute deal; contracted, energized and in use are three different numbers |
What surprised me
Placeholders in your voice, drafted from the research and the earlier guides. Rewrite each with your own moment.
The price sheet. In telco a minute had a price and a carrier sold it. Here the price per token fell every year and the help-desk button could still cost more than in 2023, because the model started thinking before it answered.
The rate limit. In observability I thought of 429s as a client bug. Here they're how a supplier with too few GPUs decides who gets served.
The SLA. In cloud security an outage was the worst case. Here the model can stay up and quietly answer worse for weeks, and no contract I found counts that.
The supplier. In the AI coding agents guide a lab could cut off a rival. Here the chip maker, the cloud and the lab invest in one another, so the supplier is also the customer's shareholder.
Sources
Undated entries were read on October 9, 2026; "search result" means seen only as a search snippet or summary. Company figures are self-reported unless they come from a filing or a regulator. Lab revenue and margin figures from leaks are unaudited.
Labs and clouds: documentation, pricing and terms
- Anthropic: rate limits, service tiers, data residency, data retention, prompt caching, models overview, model deprecations, pricing
- OpenAI: rate limits, pricing, data controls, deprecations, reasoning, Scale Tier (search result), gpt-oss-120b model card (Aug 2025)
- Google: Gemini API rate limits, pricing and deprecations; indemnified services (Apr 2026, search result)
- AWS: Bedrock pricing, Bedrock SLA, Bedrock model lifecycle, open-weight models in the European Sovereign Cloud (Sep 2026), Nova indemnity (search result), European Sovereign Cloud launch (Jan 2026, search result)
- Microsoft: Foundry deployment types (Aug 2026), model retirements (Jul 2026), Customer Copyright Commitment (search result)
- Other labs: xAI models and pricing; Mistral pricing and models; DeepSeek API pricing
- Providers and GPU prices: Together AI pricing; Fireworks pricing and serverless pricing; Akash, H100 rental prices (Sep 2026); Tomasz Tunguz, B200 pricing (Apr 2026)
- Gateways: OpenRouter FAQ and provider selection; Cloudflare AI Gateway; LiteLLM proxy reliability
- Serving: Nvidia, inference optimization (Nov 2023) and NVFP4 (search result); Kwon and colleagues, vLLM (2023)
Company filings and disclosures
- Alphabet, Q2 2026 results (Jul 2026); TechCrunch, Google Cloud capacity constrained (Apr 2026)
- Microsoft, 10-Q on the OpenAI agreement (Oct 2025, search result); Nasdaq, FY26 Q4 call highlights (Jul 2026, search result)
- Amazon, Q2 2026 results (Jul 2026); Nasdaq, Trainium commitments (2026, search result)
- Oracle, Q1 FY27 results (Sep 2026, search result)
- Nvidia, Q2 FY27 results (Aug 2026); Data Center Dynamics, AMD's Q2 2026 results (Aug 2026, search result)
- CoreWeave, Q2 2026 results (Aug 2026); Nebius, Q2 2026 6-K (Aug 2026, search result)
- Cerebras, Q2 2026 8-K and Q1 2026 8-K (2026); Let's Data Science, Cerebras IPO (May 2026, search result)
- xAI: TechCrunch via Yahoo Finance, xAI in SpaceX's filing (May 2026); Reuters via New Straits Times, SpaceX acquires xAI (Feb 2026)
- Anthropic, higher limits and SpaceX compute (May 2026)
Labs' finances and deals (press and leaks)
- Anthropic's leaked draft filing: Fortune, the income statement (Sep 2026); Implicator, $518 billion of commitments (2026); Morningstar, the leaked financials (Oct 2026)
- Run-rates and accounting: Bloomberg Government, Anthropic run-rate (Aug 2026); Axios, OpenAI nears $70B (Sep 2026, search result) and gross vs net revenue (Sep 2026, search result)
- OpenAI: Reuters via Yahoo Finance, compute spend through 2030 (Feb 2026); AI Weekly, Q1 2026 revenue (2026); Bloomberg Law, $122 billion round (2026, search result); StorageNewsletter, confidential S-1 (Jun 2026, search result); TweakTown, out of GPUs (Feb 2025, search result); Techwyse, OpenAI models on Bedrock (Apr 2026, search result)
- Compute deals: Data Center Knowledge, Anthropic's TPU deal (2026, search result); Telecoms.com, Anthropic and SpaceX compute (May 2026, search result); The Next Web, Meta and CoreWeave (Apr 2026, search result); Epoch AI, Stargate Abilene (Jul 2026, search result)
- Meta: Reuters via The Star, Muse Spark 1.1 (Jul 2026)
- Mistral: TechCrunch, Mistral raises €3B (Sep 2026, search result)
- China and open weights: TrendForce, Zhipu's price rise (Feb 2026); UC Today, Mozilla's State of Open Source AI (Jul 2026); Moonshot, Kimi Vendor Verifier (search result)
Inference providers, GPU clouds and gateways (press)
- IntelligentCIO, Fireworks Series D (Jul 2026, search result); Noqta, Together AI Series C (2026, search result); Dealroom, Baseten Series F (Jun 2026, search result); GovConWire, Nvidia and Groq (Dec 2025, search result)
- AI Weekly, Crusoe's Series F (Sep 2026, search result)
- Shopifreaks, OpenRouter Series B (May 2026, search result); Palo Alto Networks, to acquire Portkey (Apr 2026, search result)
Research, benchmarks and telemetry
- Artificial Analysis: performance methodology, leaderboard, open-weights models, recent open-weights launches (Apr 2026), gpt-oss-120b providers, Opus 4.6 (Feb 2026)
- Prices: Epoch AI, LLM inference price trends (Mar 2025); a16z, LLMflation (Nov 2024, search result); Wing VC, who makes money when inference gets 10x cheaper (Aug 2026)
- Usage: a16z and OpenRouter, State of AI, 100 trillion tokens (Jan 2026); Datadog, State of AI Engineering (2026, vendor telemetry); Menlo Ventures via TechCrunch, enterprise model preferences (Jul 2025, search result) and via BigDATAwire, enterprise LLM spend (2025, search result)
- Hardware and hosts: SemiAnalysis InferenceX, MiniMax M2.7 on H100 and FP4 vs FP8 (search result); Simon Willison, inconsistent gpt-oss performance (Aug 2025, search result); Morph, context rot (search result)
- Capacity and power: Scientific American, the AI compute crunch (May 2026); PJM, 2028/29 capacity auction (Jul 2026, search result); E&E News, auction hits the cap again (Jul 2026, search result)
Law, regulators and courts
- EU AI Act: European Commission, guidelines for general-purpose AI providers (Apr 2026) and Code of Practice signatories (Oct 2026); WilmerHale, the guidelines' thresholds (Jul 2025, search result); Allegiance, first information requests (Sep 2026); Help Net Security, enforcement powers (Aug 2026, search result); TechCrunch, Meta declines the Code (Jul 2025, search result); Euronews, AI gigafactories call (Jul 2026, search result)
- Export controls: BIS, diffusion rule rescinded (May 2025, search result); IAPS, H200 licensing policy (Jan 2026); Morgan Lewis, UAE rule and testimony (Jul 2026); TechCrunch, H20 licensing charge (May 2025, search result); GIGAZINE, China blocks H200 (Jan 2026, search result); Asia Economy, limited H200 entry (Aug 2026, search result); Expertlancing, ultimate-parent guidance (2026, search result)
- Chinese models: NIST, CAISI evaluation of DeepSeek (Sep 2025, search result); Just Security, regulate, don't ban (2026, search result); CNBC, congressional probe (Jul 2026, search result)
- US states: FPF, California SB 53 (search result); Wiley, New York RAISE Act (2026, search result); Tech Policy Press, September 2026 roundup (Sep 2026)
- Copyright: Writer Beware, Bartz final approval (Jul 2026); Authors Alliance, Thomson Reuters v. Ross (Oct 2026, search result); ChatGPT Is Eating the World, OpenAI MDL summary judgment (Sep 2026) and Kadrey v. Meta (Jun 2026, search result); Bird & Bird, GEMA v. OpenAI-on-copyright-and-ai-training) (Nov 2025, search result); IPKat, Getty appeal permitted (Jan 2026, search result); Bloomberg Law, Warner Music and Suno (Nov 2025, search result)
- Retention: Duane Morris, OpenAI preservation orders (Oct 2025)
Incidents
- The Register, ChatGPT, Claude and Grok outages (Sep 2026); Anthropic, postmortem of three recent issues (Sep 2025); OpenAI, expanding on sycophancy (May 2025, search result) and the Mixpanel incident (Nov 2025, search result)
- AWS outage losses: Moody's, insured losses (2025, search result); Artemis, parametric claims paid (2025, search result)
- Outlook Business, Cloudflare outage (Nov 2025, search result); PyPI Stats, the LiteLLM compromise (Mar 2026, search result); Wiz, DeepSeek's exposed database (Jan 2025, search result)
Field Guides are learning notes, not legal or compliance advice. Rules and fees change; check the cited primary sources before you act on anything here.