The Platform PM
Field Guide

Models and inference: what AI runs on

How a request becomes tokens and a bill: model labs, open-weight models, clouds, inference providers, GPU clouds and chips, what a feature really costs, why capacity is scarce, and who pays when a model is throttled, retired or wrong.

Last updated October 2026

The industry on one page

The parties. Your app sends requests, often through a gateway that picks a model and caches answers. A lab's API, a cloud platform or an inference provider runs the model on GPUs in someone's data center, and chips and power decide how much capacity there is, which is why rate limits exist.

Picture a company that sells help-desk software to other businesses. It adds a button that summarizes a support ticket and drafts a reply, and customers press it about a million times a month, each time sending roughly 2,000 tokens in and getting 500 back. A token is the unit a model reads, writes and bills, a piece of a word: on Anthropic's current tokenizer a million tokens hold about 555,000 English words. Your app sends requests, often through a gateway that picks a model and caches answers, and that also keeps a budget per team and switches models when a call fails. Behind it, a lab's API, a cloud platform or an inference provider runs the model on GPUs in someone's data center: Anthropic, OpenAI or Google selling its own closed model; Amazon Bedrock, Microsoft Foundry or Google Vertex AI reselling several labs' models; or Fireworks or Together serving an open-weight model such as DeepSeek or OpenAI's gpt-oss. Most rent their GPUs from a GPU cloud such as CoreWeave or from a hyperscaler, and the chips come from Nvidia, Google's TPUs or custom designs such as Amazon's Trainium. At the bottom, chips and power decide how much capacity there is, which is why rate limits exist.

If you only remember a few things:

  1. Prices: the price of a fixed level of capability falls roughly 10-40 times a year, so the button that would have cost about $90,000 a month on GPT-4 in 2023 costs a few hundred dollars on a small 2026 model. The price of the best model doesn't fall, because labs keep adding a new top tier: OpenAI's flagship input price went from $1.25 per million tokens for GPT-5 to $10 for gpt-6-astra. Cost per task rises, because prompts grow, reasoning models bill thousands of hidden tokens per answer, agents loop, and demand barely responds to price. Some prices went up outright: retired Claude models kept on Bedrock at twice their launch price, Gemini Flash doubling on January 1, 2027, and GPU rental rates. See How the money moves.
  2. The bill: for the help-desk button, a frontier workhorse model such as Anthropic's Opus 5.5, OpenAI's gpt-6.1-sol or Google's Gemini 3.1 Pro costs about $6,000-18,000 a month at list prices, a mid-tier model $2,400-9,000, and an open-weight model at an inference provider $300-5,000. Reasoning multiplies the bill by three to four, and a 10-step agent by about 24 on the same model. Self-hosting loses money at this volume. Buying the same model a different way (batch, residency, a gateway's fee) moves the bill by 10-50%, while model tier and reasoning tokens move it by orders of magnitude.
  3. Compute: every big cloud said in 2026 that it was capacity constrained, and rate limits follow the compute: Anthropic raised its Opus API limits in May 2026 when new capacity from SpaceX's Colossus 1 data center arrived. Contracted capacity runs two to three times ahead of what's switched on, the biggest US power market has cleared at its price cap three auctions running, and Nvidia, Amazon, Microsoft and Google all invest in the labs that buy their compute. See Who holds the power.
  4. Routing: open-weight models trail the best closed model by about 6-12 points on Artificial Analysis's index and cost a half to a sixth as much, and gateways make switching a configuration change. But usage barely responds to price, the same open model behaves differently at different hosts, and buying speed costs throughput. See Open-weight vs closed models: what's real.
  5. Operations: in Datadog's customer data, rate limits were 60% of failed LLM calls in February 2026. Priority capacity is hard to buy, retirement notices are short (60-62 days at Anthropic in 2026, 20 days for one specialized OpenAI model), outage credits are small, and I found no SLA (service level agreement) that covers the quality of answers: Anthropic's postmortem of its 2025 serving bugs mentioned no refunds. See What mistakes cost.
  6. Rules: the EU AI Act's duties for general-purpose model providers have been enforceable since August 2, 2026, and the Commission sent its opening information requests on September 1. A builder becomes a "provider" only if fine-tuning uses more than a third of the original model's training compute. US export controls work as an allocation tool: no replacement for the rescinded diffusion rule, case-by-case H200 sales to China, and China blocking imports itself. Copyright costs land on the labs (Anthropic's $1.5 billion Bartz settlement; Thomson Reuters v. Ross), while Chinese open-weight models carry jurisdiction risk and can mean losing your vendor's IP indemnity, as How the rules work explains.

This guide is the supply side. How agents use these models (orchestration, memory, tool security, evals) is in AI agents: orchestration platforms, and I don't repeat it. How labs use model supply against coding tools, and what seats and usage cost there, is in AI coding agents; speech and realtime models are in Voice AI; and measuring tokens, latency and cost per request is in Observability.

The main players

These are the companies behind the diagram's parties, layer by layer, in no particular order.

Model labs (closed)

What they do
Train frontier models; sell tokens, subscriptions and enterprise deals
Main players
Anthropic, OpenAI, Google DeepMind, xAI (now part of SpaceX), Meta, Mistral
What they control
Quality, prices, rate limits, retention and retirement dates

Open-weight model makers

What they do
Publish weights anyone can download and serve
Main players
DeepSeek, Alibaba (Qwen), Moonshot (Kimi), Xiaomi (MiMo), OpenAI (gpt-oss), Google (Gemma)
What they control
The price floor for capable models, licences, and whether hosts serve their model faithfully

Cloud AI platforms

What they do
Resell many labs' models under one cloud contract
Main players
Amazon Bedrock, Microsoft Foundry, Google Vertex AI
What they control
Enterprise commitments, regions, their own SLAs, dates and indemnities

Inference providers

What they do
Serve open-weight models per token or per GPU-hour
Main players
Fireworks, Together AI, Baseten, Cerebras, Groq
What they control
Speed, price and serving quality for open models

GPU clouds

What they do
Rent GPUs and data centers on multi-year contracts
Main players
CoreWeave, Oracle (OCI), Nebius, Crusoe
What they control
Energized capacity, power contracts and the debt that pays for them

Chip makers

What they do
Design and sell the accelerators
Main players
Nvidia, AMD, Google (TPU, with Broadcom), Amazon (Trainium), Broadcom, Cerebras
What they control
Allocation of the scarcest chips and the highest margin in the chain

Gateways and routing

What they do
One API in front of many models: routing, fallback, budgets, caching
Main players
OpenRouter, LiteLLM, Portkey (Palo Alto Networks), Cloudflare AI Gateway
What they control
Switching, spend controls and a second ledger of every call

How they make money, and who's moving:

  • Closed labs sell tokens, subscriptions and enterprise deals, directly and through clouds. Every lab revenue figure here is unaudited, and most are unofficial. Anthropic's come from a leaked draft of its IPO filing, as reported by Reuters on September 28, 2026: about $4.6 billion of 2025 revenue against $7.3 billion of compute and infrastructure costs, then $11.5 billion in April to June 2026 with an adjusted operating profit. OpenAI's are just as unofficial: Reuters and The Information, citing anonymous sources and internal documents, reported about $13 billion of 2025 revenue and $5.7 billion for January to March 2026, and its own draft filing was submitted confidentially. People familiar with each company put Anthropic's run-rate above $65 billion at the end of July and OpenAI's near $70 billion in September. The two book cloud resale differently (Anthropic counts the whole dollar, OpenAI only its share of some partner sales), so their numbers don't compare cleanly. xAI is the only lab with figures in a public filing: SpaceX's showed $3.2 billion of 2025 revenue and a $6.4 billion operating loss. Google doesn't break out Gemini revenue. Meta released its debut closed model in April 2026 and a paid API at about $1.25 per million input tokens in July. Mistral, Europe's regional leader, raised €3 billion at more than €21 billion on September 8, 2026.
  • Open-weight makers earn little directly; their influence is price. DeepSeek sells its own API at half price off-peak. Chinese labs held all of the top ten open-weight spots on Artificial Analysis's index in April 2026, and Qwen overtook Meta's Llama in Hugging Face downloads in February, according to coverage of Mozilla's open-source AI report. Prices aren't only falling: Zhipu raised overseas API fees for its GLM models by 67-100% in February 2026. I found no revenue figure for DeepSeek, Alibaba's Qwen or Moonshot.
  • Cloud platforms resell models at roughly the lab's price and earn on the commitment and the rest of the cloud bill. In the quarter to June 2026 AWS grew 37%, Azure 43% and Google Cloud 82%, to $24.8 billion. Bedrock has carried OpenAI's models since late April 2026, after Microsoft's API exclusivity ended (as reported), and Claude runs on all three clouds.
  • Inference providers sell per token on shared ("serverless") endpoints and per GPU-hour on dedicated ones. Fireworks says it passed $1 billion of annualized revenue, and it raised about $1.5 billion at $17.5 billion in July 2026. Together AI raised $800 million at $8.3 billion, and Baseten $1.5 billion at $13 billion after its run-rate reportedly tripled to about $600 million in one quarter. Cerebras went public on May 14, 2026, and reported a GAAP gross margin of 14% in the quarter to June. Nvidia paid about $20 billion in December 2025 to license Groq's technology and hire its leaders, and GroqCloud carries on.
  • GPU clouds sign multi-year take-or-pay contracts (the customer pays whether or not it uses the capacity), borrow against them and buy GPUs. CoreWeave's quarterly revenue more than doubled to $2.58 billion, but $640 million of interest left it with a $626 million net loss. Oracle's backlog reached $664 billion in September 2026, Nebius grew revenue 454%, and Crusoe, which develops OpenAI's Abilene site, was valued at about $30 billion.
  • Chip makers keep the highest margins in the chain. Nvidia's revenue for the quarter to July 2026 was $96.2 billion, $89 billion of it from data centers, at a gross margin of about 74%. AMD's data center revenue roughly doubled to $6.7 billion, Amazon says its Trainium and Graviton business runs above $25 billion a year, and Google supplies TPUs, built with Broadcom, to other labs, Anthropic among them.
  • Gateways earn a thin fee or nothing. OpenRouter takes 5.5% on credit purchases (nothing if you bring your own keys, up to $25,000 a month) and reportedly raised $113 million at about $1.3 billion in May 2026. LiteLLM is the most-used open-source proxy, Palo Alto Networks announced it would buy Portkey on April 30, 2026, Cloudflare includes its AI Gateway on all plans, and the clouds have their own routers.

As of October 2026. Most private-company figures here are self-reported or come from press coverage of rounds and leaks, so treat the list as a map to check.

Back to the help-desk company. Its app calls LiteLLM, which it runs itself, and the gateway checks the team's monthly budget and sends the ticket to Claude Sonnet 5.5 on Bedrock, so the spend draws down the company's existing AWS commitment. Bedrock runs the model in AWS's data centers, filled with Nvidia GPUs and Amazon's own Trainium chips. When Bedrock answers with a rate-limit error, the gateway sends the ticket to gpt-oss-120b at Fireworks instead, which serves it from rented Nvidia GPUs. One summary can touch five or six companies, each with its own meter and retirement calendar (the vendors in the story are illustrative).

How a request becomes a bill, step by step

One request, step by step. The model reads the whole prompt before the first word comes back, then writes the answer token by token. The bill counts input, output and cached tokens separately, and reasoning models and agents use many more tokens per task, which matters more than the price per token.

Here's one press of the help-desk button. The model reads the whole prompt before the first word comes back and then writes the answer token by token. The bill counts input, output and cached tokens separately.

  1. Request sent. The app assembles the prompt: system instructions, tool definitions, the ticket history and any retrieved documents. Caching only works on an identical prefix, so the stable parts go at the start. In Datadog's data, system prompts make up 69% of input tokens. What can go wrong: the prompt exceeds the context window and the call fails; a new model generation counts the same text as about 30-35% more tokens, so a "same price" upgrade can cost more.
  2. Queued. The provider admits the request after checking the account's tier, spend cap and rate limits (requests and tokens per minute, per model), then assigns a service tier (standard, priority, flex or batch) and a region. What can go wrong: the request is Throttled, the exception on the diagram: a 429 error that says "retry later". Anthropic returns the same 429 at its monthly spend cap, where retrying won't help until the 1st of next month. Failed requests still count toward OpenAI's per-minute limit, so retries without a random delay keep the bucket empty.
  3. First token. In prefill, the model reads every input token in one parallel pass and builds a working memory of the prompt (the KV cache). Prefill is bound by raw compute and sets the time to the first token. A cached prefix is skipped, which is faster and cheaper. What can go wrong: a cache miss after someone edits the system prompt brings back the full price and wait; a reasoning model at its highest effort can think for minutes, so client timeouts set lower kill work you still pay for.
  4. Streamed. In decode, the model writes one token per step, re-reading its weights and the cache from memory each time, which is why decode is limited by memory bandwidth and output tokens cost more. A reasoning model writes hidden "thinking" tokens before the visible answer, and they're billed as output. What can go wrong: the answer hits max_tokens and comes back incomplete, possibly after you've paid for thousands of reasoning tokens; the connection drops. The other exception, Failed over, is the gateway switching to another model or host, which it can do cleanly only before the first token arrives.
  5. Billed. The provider meters uncached input, cache writes, cache reads, output (reasoning included), service tier and region, prices each line separately and bills prepaid credits, an invoice or a cloud marketplace. What can go wrong: the gateway, the provider's console and the cloud invoice disagree.

That's where the diagram's note comes from. On OpenRouter, the average prompt grew from about 1,500 tokens to more than 6,000 between 2024 and late 2025, and completions from about 150 to 400. Reasoning models and agents use many more tokens per task, which matters more than the price per token.

A few terms carry most of the explanations in this space:

  • Why output costs more. Most labs price output at four to six times input (five times at Anthropic, OpenAI's gpt-6.1-sol and Gemini 3.8 Flash). Decode is slower per token than prefill, but the exact multiple is a pricing choice and isn't a measured hardware ratio.
  • Reasoning tokens. OpenAI's docs say they're invisible, billed as output and range from a few hundred to tens of thousands per request. By late 2025 reasoning models produced more than half of OpenRouter's tokens.
  • Batching and the speed trade. Serving engines such as vLLM, SGLang and Nvidia's TensorRT-LLM pack many users' requests onto the same GPUs. More users per GPU means cheaper tokens and slower answers for each. In SemiAnalysis's InferenceX benchmark, serving one open model about three times faster per user cost about four times as much per token, which is the economics behind "fast" tiers at twice the list price.
  • Prompt caching. A repeated prefix costs a tenth of the input price or less to read, and on most Claude models cached reads don't count toward the per-minute input limit either. Yet only 28% of the LLM calls in Datadog's data showed any cache reads.
  • Quantization. Storing a model's numbers at lower precision (8 or 4 bits instead of 16) makes it smaller and faster. On Nvidia's newest GPUs, 4-bit DeepSeek-R1 ran 2.2-3.5 times faster than 8-bit. Nvidia reports accuracy losses of about 1% or less on most tasks. It's also why "the same model" differs between hosts.

Cheat sheet: ways to buy inference

A lab's own API

Good for
The newest models first, fast modes and caching variants
Watch for
Tier caps; short retirement notices; the lab is your data processor
Who sets the clock
The lab

A cloud AI platform

Good for
One contract and commitment, private networking, residency, an SLA with credits
Watch for
Its own retirement dates; features arrive later; auto-upgrades on some deployments
Who sets the clock
The cloud

An inference provider (serverless)

Good for
Open-weight models, speed, low prices
Watch for
Quality and speed vary by host; quantization; models come and go
Who sets the clock
The provider and the model maker

A dedicated endpoint or provisioned throughput

Good for
Predictable latency and capacity
Watch for
You pay for idle hours (about $5.50-13 per GPU-hour at providers)
Who sets the clock
The contract

Self-hosting on a GPU cloud or your own hardware

Good for
Control: residency, fine-tuned weights, no retirement clock
Watch for
Utilization, operations, redundancy
Who sets the clock
You

A gateway in front of any of these

Good for
Failover, budgets, multi-model routing, one log
Watch for
A fee or your own ops, an extra hop, supply-chain risk
Who sets the clock
You

Which way to buy inference? Five questions

  1. Frontier or not? Classification, extraction, summaries and chat usually don't need it, and a small or open-weight model at a tenth of the price or less is close enough. Hard agentic and coding work still does.
  2. Does it need to reason? Reasoning multiplies the bill by three to four on a chat-shaped task. Set the effort per feature, cap max_tokens and measure tokens per task as well as the price per token.
  3. Can it wait? If an answer can come back within 24 hours, batch halves the bill. If it must be fast, priority or fast tiers cost 1.75-2 times list, and capacity for them may not be for sale.
  4. What does the contract need? Residency (about 10% extra), zero data retention, an uptime SLA with credits and an IP indemnity each narrow the routes. A cloud platform usually adds an SLA with credits, which a lab's standard tier doesn't have; check that the indemnity covers the specific model.
  5. When does it go away? Every route except self-hosting has a retirement clock, some as short as 60 days. With three labs in use, expect a forced re-test every few weeks.

My defaults: for the help-desk company I'd start on a mid-tier closed model through the cloud it already has a commitment with, with caching designed in and reasoning off unless an eval shows it helps. I'd put a gateway in front from day one, with per-team budgets, a fallback to another provider, and an open-weight model for the easy half of the traffic once our evals say it's good enough. I'd pin dated snapshots, keep a regression set of a few hundred real tickets that runs on every model change, and track cost per resolved ticket. I wouldn't self-host until volume is ten times higher or a customer requires it.

The primitives

01

Entity and identity

What is the unit of record, and how do we know it is the same one?

"The model" has at least three parts: a family (Sonnet, gpt-6.1-sol), a snapshot ID (the exact version), and the host or deployment that serves it. On Azure you call a deployment name you created, and on Bedrock the same Claude model has its own ID and its own retirement date. Open-weight models add the precision and serving stack, which is why OpenRouter lets you filter hosts by quantization. The billing identity is a tree of organization, workspace or project, and API key, and rate limits attach to the organization (Anthropic, OpenAI) or the project (Google), per model, so teams in one organization share its headroom.

The law names its own entities. Under the EU AI Act the "provider" places a model on the market and the "deployer" uses it, and a company that fine-tunes with more than a third of the original training compute becomes the provider of the result. An IP indemnity keys on plan, service and model together, so one API key can be covered for one model and not another. US export rules now follow a company's ultimate parent, wherever its subsidiary is registered (guidance of May 31, 2026).

More on Entity and identity →

02

State and lifecycle

What states exist, and what moves an entity between them?

A request goes from admitted to queued, prefilling, decoding and streaming, and ends complete, incomplete or failed. A model goes from preview to generally available, legacy, deprecated (still works, retirement date set) and retired. A prompt cache is written, stays warm while used, and dies when it expires or when the prefix, tools or reasoning settings change. An account climbs tiers with spend history and pauses at its spend cap until the next month. Capacity goes from contracted to built, energized and allocated.

Retirement clocks differ by vendor and are getting shorter at some:

VendorTypical notice before retirementNotes
AnthropicAt least 60 days; 60-62 days for every 2026 noticeDown from 181-189 days in 2025; Opus 4.1 lived exactly 12 months
OpenAIAt least 6 months for generally available models, 3 for specialized ones, about 2 weeks for previewsA restricted cyber model got 20 days in September 2026
Google (Gemini API)Dates are "earliest possible"; previews retire 2-4 months after releaseSome previews run in production for months
Microsoft Foundry18-month lifecycle for most generally available models, at least 60 days' noticeStandard deployments auto-upgrade; "retirement dates aren't extendable"
AWS Bedrock"End of life no sooner than" a stated date; Legacy for 6 months or 45 daysLegacy blocks new customers and provisioned throughput
MistralNo formal policy; about 24 days observed for one model
Self-hosted open weightsNoneHosted versions at clouds still retire

The same model can retire on different dates at the lab and at each cloud. Products retire too: OpenAI ends its Assistants API, hosted evals and self-serve fine-tuning jobs between August 2026 and January 2027.

More on State and lifecycle →

03

System of record and ledger

Who owns the truth, and how do systems reconcile?

FactSystem of record
What one call usedThe provider's usage record per request: input, cache write, cache read, output, tier, region
What the company spentThe provider's usage and cost API, the gateway's log and the cloud invoice, which rarely agree
How much headroom is leftRate-limit headers on every response
Whether the provider was upIts status page, which the provider writes and which can lag
What a lab has contracted to buyBacklog at the clouds; take-or-pay schedules in filings

On cloud marketplaces Anthropic converts usage into consumption units of $0.01 and reports them to AWS or Azure, so a discount shows up as fewer units. A private lab's "run-rate" is a recent month times twelve. Your own gateway log is the only record of errors from your side, and it's the evidence an SLA claim needs.

More on System of record and ledger →

04

Rules and policy

What logic decides outcomes, and who can change it?

Access is rationed by rules the provider writes and can change:

Anthropic

Tiers
Evaluation, Start, Build, Scale, Custom
Monthly spend cap
$500 / $1,000 / $200,000
Example limit at the top self-serve tier
Opus 5.5: 10,000 requests and 10 million input tokens a minute; Fable 5.x: 4,000 and 4 million

OpenAI

Tiers
Free, Build, Launch, Grow (after $5, $100 and $500 of purchases)
Monthly spend cap
$100 / $500 / $5,000 / $200,000
Example limit at the top self-serve tier
gpt-6 Astra, Sol and Terra: 15,000 requests and 40 million tokens a minute

Google (Gemini API)

Tiers
Free, Tiers 1-3 (by spend and account age)
Monthly spend cap
$250 to $100,000 or more, plus a 10-minute spend limit
Example limit at the top self-serve tier
Per project, per model

The top model gets the tightest limit, and Anthropic calls its limits "maximum allowed usage, not guaranteed minimums". New organizations start lower "to prevent fraud and abuse", and fast ramps hit acceleration limits. Gateways add the customer's own policy (route by price or speed, cap the price per token, use only zero-retention hosts, set budgets per team), and the labs' usage policies restrict high-stakes uses without human review.

More on Rules and policy →

05

Effective dating

Which version of the rule applied at that moment?

Dates I'd keep on a calendar as of October 2026:

ChangeEffectiveStatus (Oct 2026)
EU AI Act duties for general-purpose model providersAugust 2, 2025In force
California SB 53 for frontier developersJanuary 1, 2026In force
H200-class chips to China move to case-by-case reviewJanuary 13, 2026In force; China limits imports
EU AI Office enforcement powers, including finesAugust 2, 2026In force; first requests September 1
Claude Sonnet 4.5 retiresNovember 30, 2026Upcoming
EU Product Liability Directive covers softwareDecember 9, 2026Upcoming
OpenAI GPT-5 and o3 snapshots retireDecember 11, 2026Upcoming
Gemini 3.6-3.8 Flash prices doubleJanuary 1, 2027Upcoming
New York RAISE ActJanuary 1, 2027 (some sources say July 1)Upcoming
Older general-purpose models must comply with the AI ActAugust 2, 2027Upcoming

Some dates float. Retirement dates are "not sooner than" (Anthropic, AWS) or "earliest possible" (Google), and introductory prices can end or be made permanent: Anthropic kept Sonnet 5's and cancelled a rise planned for September 1, 2026. DeepSeek even prices by time of day.

More on Effective dating →

06

Interfaces and standards

What format and protocol do counterparties speak?

OpenAI's Chat Completions and Responses formats are the de facto interface: Bedrock exposes them ("change the base URL and API key"), most inference providers do, and Anthropic runs an OpenAI-compatible endpoint next to its own Messages API. That makes switching mechanically cheap. Everything around the call is unstandardized: usage fields, rate-limit headers, deprecation metadata and reasoning controls differ per provider. Tokenizers differ too, so the same text can be about 30-35% more tokens on Anthropic's newest models than on its older ones.

Measurement is its own layer. Artificial Analysis measures time to the first token and output speed from one server over 72 hours, MLPerf audits hardware results that vendors submit, and SemiAnalysis's InferenceX turns benchmarks into cost per million tokens. Model makers have started checking hosts: Moonshot publishes a Kimi Vendor Verifier and OpenAI added a compatibility test for gpt-oss. Below the token the units change again: GPU-hours for renting chips, gigawatts of power capacity for compute deals, and megawatt-days in power markets.

More on Interfaces and standards →

07

Networks and counterparties

Who sits between us and the outcome, and what do they want?

PartyWhat they controlWhat they earn
LabQuality, prices, limits, retention, retirementsTokens, subscriptions, enterprise deals
Cloud platformProcurement, regions, its own SLA and datesResold tokens; commitments drawn down
Inference providerSpeed and price for open modelsTokens, dedicated endpoints, GPU rental
GPU cloudEnergized capacityMulti-year take-or-pay contracts
Chip makerAllocation of new chipsHardware at about 74% gross margin (Nvidia)
GatewayRouting and budgetsA small fee, or nothing (open source)
Grid operator and utilityPower and connection datesCapacity payments, set by auction

Money flows down the chain and equity flows back up it: Nvidia invests in OpenAI and in GPU clouds and providers, Amazon and Google in Anthropic, Microsoft owns about 27% of OpenAI and has invested in Anthropic, and AMD gave OpenAI warrants for up to 160 million of its shares. Amazon sells Claude on Bedrock, holds Anthropic notes, supplies it Trainium and sells its own Nova models, all at once.

More on Networks and counterparties →

08

Regulatory layering

Jurisdiction × activity × entity type: is it a license or a certification?

Model makers

US
California SB 53 (frontier developers above 10^26 operations); New York RAISE from 2027
EU
AI Act duties for general-purpose models; Code of Practice
China and elsewhere
Chinese labs mostly outside the EU Code

Chips and weights

US
Export controls: case-by-case H200 licences for China, approved entities in the UAE
EU
China and elsewhere
Beijing approves or blocks imports

Data

US
Sector rules; court preservation orders
EU
GDPR; residency demands
China and elsewhere
Data stays in-country in many sovereign deals

Products

US
Deployer liability (case law)
EU
Product Liability Directive from December 9, 2026
China and elsewhere

The same model reached through the lab, Bedrock or Vertex can come with different retention, residency, SLA and indemnity terms, so the route is a legal choice too. For most builders the law arrives through vendor terms and procurement questionnaires.

More on Regulatory layering →

09

Exceptions and reversals

What goes wrong, and how is it undone?

Rate limit or spend cap hit

Who starts it
The provider's limiter
Clock
Seconds, or until the 1st of next month
The way back
Back off with jitter; fail over; ask for a tier increase

Model retired

Who starts it
The lab or cloud
Clock
60 days to 6 months; 20 days seen
The way back
Re-test on the replacement; pin the next snapshot

Bad serving change

Who starts it
The provider
Clock
Hours to weeks to detect
The way back
Provider rollback (GPT-4o's sycophantic update came out about four days after release); your own evals

Price step-up

Who starts it
The provider
Clock
A dated announcement, or a promotion ending
The way back
Route elsewhere; renegotiate

Capacity product withdrawn

Who starts it
The provider
Clock
Immediate for new buyers
The way back
Existing contracts honoured; buy elsewhere

Compute deal cancelled or reshaped

Who starts it
Either party
Clock
90 days for xAI's capacity to Anthropic; years for take-or-pay
The way back
Other sites, other chips

Rule reversed

Who starts it
A government
Clock
Months
The way back
Switchable architecture; exit rights in contracts

Reversals run both ways. OpenAI's "$1.4 trillion" of compute ambitions became about $600 billion through 2030 in reporting by February 2026, and the expansion of the Abilene site beyond 1.2 GW was dropped. Export rules went from ban to unban to a revenue share that was never codified, all within 2025.

More on Exceptions and reversals →

10

Liability allocation

When it fails, who pays?

FailureWho absorbs itMechanism
OutageThe builder, minus a small credit where an SLA existsBedrock: 99.9% per region, credits of 10-100% of that region's charges as future credit; lab standard tiers: best effort, no credit
Silent quality dropThe builderNo SLA found that counts wrong answers as downtime
Retirement and migrationThe builderNotice periods; no compensation
Copyright claim on an outputThe vendor, if the indemnity's conditions are metPaid plan, covered model, output unmodified, filters on; usually outside the liability cap
Copyright claim over training dataThe labBartz, the New York Times case; reaches builders only through prices
Harm from an answerThe deployer by defaultAir Canada's chatbot case; consumer cases against model makers are pending or settled
Unused computeThe lab, then its investorsTake-or-pay contracts: about 80% of Anthropic's commitments, as reported
GPU cloud failureLenders, then customers mid-contractDebt against contracts

Builders carry price and capacity risk, labs carry volume risk, GPU clouds carry financing and obsolescence risk, and household electricity customers carry part of the grid cost. Liability caps in the labs' and clouds' terms are typically 12 months of fees.

More on Liability allocation →

What's different here

How the money moves

The builder pays per token, and the money flows down to whoever owns the chips, the buildings and the power. List prices per million tokens in October 2026, input / output:

TierExamples
Top tierAnthropic Fable 5.1 $10 / $50; OpenAI gpt-6-astra $10 / $50; gpt-5.5-pro $30 / $180
Frontier workhorseAnthropic Opus 5.5 $4 / $20; OpenAI gpt-5.5 $5 / $30 and gpt-6.1-sol $2 / $10; Google Gemini 3.1 Pro Preview $2 / $12; xAI grok-4.7 $2 / $6
Mid tierAnthropic Sonnet 5.5 $2 / $10; Google Gemini 3.8 Flash $0.75 / $3.75 (doubling on January 1, 2027); Meta Muse Spark 1.1 about $1.25 / $4.25; Mistral Large $0.50 / $1.50
SmallOpenAI gpt-6-luna and Anthropic Haiku 5.5 $0.10 / $0.50; Google Gemini 3.5 Flash-Lite $0.30 / $2.50
Open-weight at a providergpt-oss-120b $0.15 / $0.60; DeepSeek V4.1 Flash $0.30 / $1.20; GLM 5.3 $1.40 / $4.40; Qwen 3.8 Max $2 / $6; Kimi K3 $3 / $15

And the modifiers that apply on top:

ModifierTypical effectWhere
Cached input0.025-0.1 times the input price; writing the cache costs 1.25-2 times at AnthropicOpenAI, Anthropic, Google (which also charges storage per hour), providers
BatchHalf price, answers within about 24 hoursOpenAI, Anthropic, Google, Bedrock, Fireworks, Mistral
FlexHalf price, synchronous but can be pre-emptedOpenAI, Google, Bedrock
Priority or fast+25% (Fireworks), +75% (Bedrock), 1.8 times (Google), 2 times (OpenAI Fast, Anthropic fast mode), 6 times (OpenAI Ultrafast)Capacity permitting
Residency+10%; +50% for Fireworks's US-hosted variantsOpenAI regional processing, Anthropic US-only inference, Claude regional endpoints on clouds

Under these sheets, prices move in different directions. The price of a fixed level of capability falls: a16z measured about 10 times a year at a constant benchmark score, and Epoch AI found GPT-4-level performance on a hard science benchmark getting 40 times cheaper a year. The top of the ladder stays put, because each generation adds a rung: OpenAI's flagship launch prices fell from $30 per million input tokens for GPT-4 to $1.25 for GPT-5 and climbed back to $10 for gpt-6-astra, while Anthropic cut Opus from $15 to $4 and added Fable at $10. And the cost per task rises: on Artificial Analysis's index Opus 4.6 used twice the output tokens of Opus 4.5 on the same tests. In OpenRouter's 100-trillion-token study a 10% price cut brought only 0.5-0.7% more usage, so spending grows from new uses such as agents and coding far more than from cheaper tokens.

Prices also went up. Bedrock keeps retired Claude 3.5 Sonnet under "extended access" at twice its launch price, Gemini 3.6-3.8 Flash list a 2027 price twice the 2026 one, and Zhipu, Moonshot and MiniMax raised some API prices. Compute got dearer too: one-year H100 rental contracts rose about 40% between October 2025 and March 2026, and B200 spot prices more than doubled in six weeks in spring 2026.

A month of the help-desk button

My assumptions, from list prices on October 9, 2026: 1 million requests a month, each with 2,000 input tokens (1,500 of them a shared, cacheable system prompt) and 500 output tokens, so 2 billion tokens in and 500 million out. The reasoning variant adds 2,000 thinking tokens per request, billed as output. The agentic variant turns each press into a 10-step agent task of about 200,000 input tokens (90% cache hits) and 8,000 output.

Top tier (Fable 5.1, gpt-6-astra)

Base month
$30,000-45,000
With reasoning on
About $130,000
As a 10-step agent
Not computed; likely far higher

Frontier workhorse (Opus 5.5, gpt-6.1-sol, Gemini 3.1 Pro)

Base month
$6,000-18,000
With reasoning on
$26,000-58,000
As a 10-step agent
$140,000-300,000

Mid tier (Sonnet 5.5, Gemini 3.8 Flash, grok-4.7)

Base month
$2,400-9,000
With reasoning on
About $10,000-26,000
As a 10-step agent

Small (gpt-6-luna, Haiku 5.5)

Base month
$300-450
With reasoning on
About $1,300
As a 10-step agent

Open-weight at a provider (gpt-oss-120b to GLM 5.3)

Base month
$300-5,000
With reasoning on
About $1,800-14,000
As a 10-step agent
About $17,000 (DeepSeek V4.1 Flash)

Top open-weight (Kimi K3)

Base month
$13,500; about $20,000 US-hosted
With reasoning on
About $43,000
As a 10-step agent

Self-hosted small open model on rented H100s

Base month
$7,000-30,000, people included
With reasoning on
Similar; bound by GPUs
As a 10-step agent

Self-hosted large open model on B200 nodes

Base month
$58,000-152,000 plus people
With reasoning on
Similar
As a 10-step agent

2023: GPT-4

Base month
$90,000
With reasoning on
As a 10-step agent

What the table says:

  • Model tier dominates. The same feature costs about $300 or about $45,000 depending on the tier.
  • Reasoning: 3-4 times. On Opus 5.5 with caching, the bill goes from about $12,300 to about $52,300.
  • Agents multiply by about 24. On Opus 5.5 an agent task costs about $0.30, or $300,000 a month, against $12,300 for the chat-shaped button. Without caching the agent's input alone would cost about $1 million a month, so cache hygiene is worth three to four times on agent work.
  • The route moves it less. Caching cuts the base bill by about 30%, batch halves any line if answers can wait, residency adds 10%, and OpenRouter's fee adds 5.5%. Priority doubles it.
  • Self-hosting loses here. Two to four rented H100s at $2-4 an hour plus a quarter to one engineer come to $7,000-30,000 a month against about $600 at a provider, 10 to 50 times more, because the provider runs newer chips at high utilization. Self-hosting pays at tens of billions of tokens a month with steady load, or when control is the requirement: residency, fine-tuned weights, or a model that never retires.
  • Against 2023. For a fixed task the bill fell from $90,000 to a few hundred dollars on a small model, but a team that moved to the top tier with reasoning pays more than it did in 2023.

Where the dollar goes

  • Lab direct: customer to lab, at list less any negotiated discount (I found few public ones), then to the lab's compute. Gross margins are unaudited estimates for every lab: The Information, as reported by Reuters, put OpenAI's adjusted gross margin at 33% in 2025, down from 40% as inference costs rose, and I found no reliable figure for Anthropic, Google's Gemini or xAI.
  • Cloud marketplace: customer to AWS, Azure or Google, drawing down a commitment, then to the lab. Marketplaces were 47% of Anthropic's 2025 revenue in its leaked draft filing, as reported.
  • Inference provider: customer to Together, Fireworks or Baseten, then to GPU clouds or their own clusters, then to Nvidia. Cerebras, the one with public numbers, reported a GAAP gross margin of 14% in the quarter to June.
  • GPU cloud and chips: lab or hyperscaler to CoreWeave, Nebius or Oracle, then to lenders, utilities and Nvidia, which earned $59.7 billion of net income on $96.2 billion of revenue in one quarter, the richest seat in the chain.

Who holds the power

Compute is the binding constraint, and power follows whoever controls a piece of it.

  • Clouds and the grid. Alphabet said in April 2026 that it was "compute constrained in the near-term", Microsoft's CFO said in July that it "remains capacity constrained", and AWS called its Trainium2 chips "largely sold out". The binding input is moving from chips to power: CoreWeave had about 1.5 GW active against 3.7 GW contracted in June 2026. PJM, the grid operator for much of the eastern US, cleared its capacity auction at the price cap three times running and ended about 6,800 MW short of its reliability target in July 2026, with data centers the main driver of new load.
  • Nvidia allocates the scarcest chips, keeps the highest margin and invests in labs, GPU clouds and providers. TPUs, Trainium, AMD and Broadcom's custom chips check it in part.
  • Labs over builders. Labs set prices, limits, retention and retirement dates, and ration when capacity is short: peak-hour cuts in Claude Code in March 2026, Priority Tier withdrawn, OpenAI limiting GPT-4.5 to its top plan in 2025 because it was "out of GPUs". Downstream, GitHub paused Copilot Business signups in April 2026 for lack of capacity. When capacity lands, limits rise, as Anthropic's did on May 6, 2026.
  • Clouds and labs. The clouds are the labs' suppliers, investors, distributors and competitors, and the labs are the anchor tenants behind the clouds' backlogs.
  • Contracts. Anthropic's commitments of at least $518 billion run about a decade, around 80% non-cancelable, per the leaked draft filing as reported. OpenAI's are reported at about $600 billion through 2030, including more than $300 billion at Oracle and a 750 MW deal that Cerebras's own filing values at more than $20 billion. Meta signed about $21 billion more with CoreWeave in April 2026. xAI builds its own Colossus sites, which is why Anthropic could rent capacity from SpaceX on 90 days' notice.
  • Suppliers as investors. OpenAI's $122 billion round at $852 billion, closed around April 2026, included $30 billion from Nvidia and $15 billion from Amazon (plus $35 billion conditional) tied to 2 GW of Trainium. Amazon marked its Anthropic notes up from $42.2 billion to $97.9 billion in one quarter, according to an analysis of the leaked filing. These loops prop up demand signals and concentrate risk.
  • Providers and gateways compete on speed, price and switching, with thin margins; model makers can shame hosts (Moonshot found gaps between third-party and official Kimi APIs "widespread"), and gateways are being folded into security companies.
  • Builders hold demand and little contractual power. Their leverage is spend commitments, a second provider and moving easy traffic to cheaper models.

How the rules work

The EU AI Act. Duties for providers of general-purpose AI models have applied since August 2, 2025, and the Commission's enforcement powers, including fines of up to 3% of worldwide turnover or €15 million, since August 2, 2026; models already on the market have until August 2, 2027. Every provider owes technical documentation, a copyright policy that honours opt-outs and a public summary of its training data, and models presumed to carry systemic risk (trained with more than 10^25 operations) add evaluations, incident reporting and security. As of October 7, 2026 the Code of Practice, the accepted way to show compliance, had 22 signatories including Amazon, Anthropic, Google, IBM, Microsoft, Mistral and OpenAI; xAI signed only the safety and security chapter, Meta declined, and DeepSeek, Alibaba and Moonshot aren't on the list. On September 1, 2026 the Commission confirmed requests for information to more than 30 unnamed AI companies, and no fines have been issued. For a builder the trigger is narrow: more than a third of the original model's training compute spent on modifying it.

Export controls. The AI Diffusion Rule was rescinded on May 13, 2025, and in July 2026 the Commerce Department's under secretary told Congress there would be no replacement. Policy now works through licences. Since January 13, 2026 Nvidia's H200 and AMD's equivalent can go to China case by case, capped at half of US volume per chip, with a 25% fee collected as a tariff and Blackwell excluded. China then blocked H200 imports in mid-January and later let in about 10,000 units each to ByteDance and Tencent, as reported, and Nvidia's outlook assumes no China data center revenue. The UAE got licence-free access for approved entities from July 10, 2026. Builders feel export rules as capacity and prices, and directly only if they give restricted parties remote access.

US states. California's SB 53 has applied to frontier developers since January 1, 2026, with safety frameworks, incident reporting and penalties up to $1 million per violation; I found no enforcement yet. New York's RAISE Act follows in 2027, and no federal preemption had passed by October 9, 2026.

Copyright. The costs land on the labs. Anthropic's $1.5 billion Bartz settlement covered about 482,000 books at about $3,100 each and got final approval in July 2026; the underlying ruling held that training on lawfully bought books was fair use and pirated libraries weren't. On September 29, 2026 the Third Circuit held in Thomson Reuters v. Ross that training a non-generative legal AI on Westlaw's headnotes wasn't fair use. It was the first appellate ruling on AI training and left the generative question open. Meta won on fair use in Kadrey in 2025, and the New York Times' case against OpenAI and Microsoft reached summary judgment briefing in September 2026. In Europe a Munich court held in November 2025 that OpenAI's memorizing song lyrics infringed, and an English court held that model weights aren't infringing copies; both are on appeal. Settlements can force retirements downstream: Warner Music's settlement with Suno requires it to retire its current models.

Indemnities. Paid tiers come with IP indemnities, each with conditions. Google's list covers its own models on Vertex, and third-party models in its catalogue aren't listed. AWS gives an uncapped indemnity for its own Nova models and narrower cover for others. Microsoft requires its content filters to stay on. Anthropic's commercial terms carry an uncapped IP indemnity, and OpenAI offers Copyright Shield. Routing to an open-weight or third-party model can drop you outside the list.

Chinese open-weight models. They're cheap and good, and they carry jurisdiction risk. NIST's AI standards center found in September 2025 that agents built on DeepSeek R1 were about 12 times more likely to follow hijacking instructions than US models. Several states ban DeepSeek on state devices, Congress opened a probe in July 2026, and the administration was reported to be weighing procurement bans and sanctions. If that happens, a cost-driven routing choice becomes a compliance problem.

Availability and data. Bedrock promises 99.9% a month per region, with credits of 10%, 25% or 100% of that region's charges as future credit. Anthropic's standard tier is "best-effort", its Priority Tier only "targets" 99.5% and isn't sold to new buyers, and OpenAI's 99.9% SLA sits in an enterprise tier with a minimum commitment. OpenAI keeps abuse-monitoring logs for up to 30 days, Anthropic keeps nothing by default except for its top models, which require 30 days, and Bedrock doesn't store prompts. A court can override any of it: in May 2025 a preservation order in the New York Times case made OpenAI keep output logs.

Sovereignty. "Sovereign" can mean an EU-run region of a US company (the AWS European Sovereign Cloud, opened January 15, 2026), an EU-owned lab (Mistral) or public compute (the EU's AI gigafactories, with bids due November 12, 2026). These regions tend to launch with open-weight models: the AWS one's debut model family on Bedrock, in September 2026, was Google's Gemma 4.

What mistakes cost

Let's say the help-desk company runs its button on a mid-tier closed model and pays $4,500-13,500 a month, about $54,000-162,000 a year. My rough arithmetic for what goes wrong and who pays:

  • A three-hour outage in a 720-hour month cuts uptime to 99.58% and loses about 4,200 requests. On a lab's standard tier the credit is zero; on Bedrock it's 10% of that region's charges, about $450-1,350, as a future credit. With a gateway failing over to another provider, most of the loss becomes about $20-60 of extra tokens and some quality variance.
  • A silent quality drop like Anthropic's 2025 bugs: if 5% of requests give worse answers for three weeks, that's about 37,500 degraded summaries, which no SLA treats as downtime.
  • A retirement with 60 days' notice costs 10-40 engineer-days of re-testing ($8,000-60,000 loaded), and up to 30% more tokens if the new tokenizer counts more, about $16,000-49,000 a year on this bill.
  • A third-party copyright claim is the vendor's to defend if the output qualifies for the indemnity; anything else is capped at 12 months of fees.
  • A ban on a Chinese open model for one customer segment would push that traffic back to a closed model at three to ten times the price, plus the re-testing.

The operational events cost far more, in expectation, than the legal ones. The record so far:

CaseWhat happenedWho paid
Overlapping outages, September 3, 2026OpenAI's ChatGPT and Codex were down about 34 minutes, Claude for 3 hours 6 minutes, and xAI's Grok at the same timeBuilders and users; no credits reported
Anthropic serving bugs, August-September 2025Three infrastructure bugs; in the worst hour 16% of Sonnet 4 requests were misrouted, and about 30% of Claude Code users had at least one bad message; benchmarks, safety evals and canaries missed itBuilders; the postmortem mentions no refunds
OpenAI GPT-4o update, April 2025A sycophantic update was rolled back about four days after releaseUsers; OpenAI's reputation
AWS us-east-1, October 20, 2025About 15 hours of disruption from a DNS failureCustomers; Moody's put insured losses at a mean of $22 million, and a parametric insurer paid claims
LiteLLM on PyPI, March 24, 2026Compromised releases of the gateway package stole credentialsUsers, who rotated every key
OpenAI and Mixpanel, November 2025A vendor breach exposed API users' names, emails and organization IDs, though no prompts or keysUsers, at risk of phishing
DeepSeek, January 2025An open database exposed more than a million log lines, chat history and keys among themUsers
Nvidia H20, April 2025A new licence requirement for ChinaNvidia: a $4.5 billion charge

Several of these are old failures in shared dependencies: DNS, a package registry, a sub-processor. The new kind is the quality drop that no contract counts.

Open-weight vs closed models: what's real

My view as of October 2026: below the frontier, an open-weight model is good enough for most chat, summarization and extraction work, at a half to a sixth of the price, and a gateway makes moving to it a configuration change. At the frontier, especially agentic coding, the gap reopens with each closed release and most enterprise money stays closed. The real switching cost is evals, prompt rework, checking each host and reading the contract, more than changing the API call.

How big the gap is

  • On the main composite index: in Artificial Analysis's April 30, 2026 snapshot, the best open-weight models (Moonshot's Kimi K2.6 and Xiaomi's MiMo V2.5 Pro) scored 54 against 60 for OpenAI's GPT-5.5 at its highest effort, a 6-point gap, and open models held 9 of the 13 spots on its frontier of intelligence against price. By October, on a newer version of the index, Xiaomi's MiMo-V2.6-Pro scored 46 against 58 for Anthropic's Opus 5.5 and 53 for the best of OpenAI and Google, about 12 points. Index versions rescale, so compare only within one.
  • On hard tasks the gap is wider: 43-46% for the best open models against 61% for closed ones on Artificial Analysis's hard terminal-coding tasks in April.
  • Who makes them: the top six open-weight models in October were all Chinese. The US ones are OpenAI's gpt-oss, Google's Gemma and Nvidia's Nemotron, and Meta has moved its frontier work to closed models.

What open costs

The price gap is large at the small and mid end: gpt-oss-120b at $0.15 / $0.60 and DeepSeek V4.1 Flash at $0.30 / $1.20 per million tokens, against $2 / $10 for a closed mid-tier model. It narrows at the top, where Kimi K3 lists at $3 / $15, close to Opus 5.5's $4 / $20. Serving a big open model yourself takes a full node of eight B200s per replica.

Same model, different host

  • Quality varies: on a 2025 maths test run 32 times, gpt-oss-120b scored 93.3% at seven hosts, 86.7% at Groq, 83.3% at Amazon, 80% at Azure and 36.7% at one small host. Hosts differ in precision, chat templates and serving bugs.
  • Speed varies more: about 37 times between the slowest and fastest gpt-oss host in October 2026, and up to 8.9 times on price.

So "route to the cheapest host of the same model" needs a per-host eval, and a failover to the same model on another host can still change your answers.

Where the money actually goes

Usage share and money share point different ways. On OpenRouter, open-weight models were about a third of tokens by late 2025. Datadog found more than 70% of organizations using three or more models. But a 10% price cut moved usage by only 0.5-0.7%, and four in ten users of one 2025 Claude model were still on it five months later. In Menlo Ventures' late-2025 survey (Menlo is an investor in Anthropic), Anthropic, OpenAI and Google took 88% of enterprise LLM spend between them, and by mid-2025 the open-source share of enterprise workloads had fallen from 19% to 13%. Quality and habit decide the vendor; price decides the long tail.

Jurisdiction and contracts

Open weights move the contract risk to you. Self-hosted, there's no retirement clock and no indemnity; at a provider, the model maker isn't your counterparty, and most IP indemnities name only the vendor's own models. Chinese labs aren't Code of Practice signatories and US restrictions were under discussion in 2026, while sovereign regions tend to launch with open models. For a team selling to US government-adjacent buyers I'd keep Chinese-model routes behind a switch I can turn off per customer.

Questions to ask a provider

  1. Which precision and serving engine do you run this model on, and do you pass the model maker's own verification test?
  2. What are your median and 95th-percentile time to the first token and output speed on our prompt sizes, measured by someone else?
  3. What notice do you give before dropping a model from serverless, and what happens to our dedicated endpoints?
  4. Is our data retained or used, by you or by any sub-processor, and can you give zero retention in writing?
  5. Which of your models, if any, come with an IP indemnity?

What usually goes wrong

SymptomLikely causeFirst thing to check
Bursts of 429 errorsRate limits per model, retries without jitter, a fast rampRate-limit headers, retry logic, a second provider
The feature stops until the 1stMonthly spend cap reachedTier caps and alerts at 50% and 80%
The bill jumps with no traffic changeReasoning effort raised, cache broken by a prompt edit, a new tokenizerTokens per request by type; cache hit rate
Answers got worse with no releaseA serving change at the provider, or a failover to another hostOnline evals by model and host
A model is going awayRetirement noticeThe regression set; pinned snapshot IDs
Legal blocks a routeRetention, residency or indemnity doesn't cover that modelThe data and indemnity terms per model and venue
A gateway leaks keysSupply-chain compromisePinned package versions; key rotation
Self-hosting costs more than plannedLow utilization, idle redundancy, peopleTokens per GPU-hour against the provider's price

Words that mean something else here

TermWhat you'd assumeWhat it means here
TokenA wordA piece of text in one model's tokenizer; the same text can be about a third more tokens on a newer model
LatencyOne numberTime to the first token (for reasoning models, the first thinking token), time to the first visible token, or end to end
ThroughputSpeedTokens a second one user sees, or tokens a second a GPU produces; they trade against each other
Context windowWhat the model can useThe maximum input and output; quality falls well before it
PrioritySame thing everywhereCommitted capacity (Anthropic, no longer sold), a 2x pay-as-you-go tier (OpenAI, now "Fast"), 1.8x per request (Google), +75% (Bedrock)
Fast modeA faster modelThe same model served faster at about twice the price; at OpenAI, the renamed priority tier
DeprecatedGoneStill works, with a retirement date set; in Azure's API, the value Deprecated means retired
Open sourceOpenUsually open weights under a licence, without training data or code; the AI Act's open-source exemption needs a free licence and no monetization
ProviderA vendorAn inference company, or the AI Act's party that places a model on the EU market, which a heavy fine-tuner can become
UptimeThe model worksThe endpoint accepts requests; nothing about the quality of the tokens
Run-rateRevenueRecent revenue times twelve; usage businesses aren't contracted for it
GigawattChipsPower capacity in a compute deal; contracted, energized and in use are three different numbers

What surprised me

Placeholders in your voice, drafted from the research and the earlier guides. Rewrite each with your own moment.

The price sheet. In telco a minute had a price and a carrier sold it. Here the price per token fell every year and the help-desk button could still cost more than in 2023, because the model started thinking before it answered.

The rate limit. In observability I thought of 429s as a client bug. Here they're how a supplier with too few GPUs decides who gets served.

The SLA. In cloud security an outage was the worst case. Here the model can stay up and quietly answer worse for weeks, and no contract I found counts that.

The supplier. In the AI coding agents guide a lab could cut off a rival. Here the chip maker, the cloud and the lab invest in one another, so the supplier is also the customer's shareholder.

Sources

Undated entries were read on October 9, 2026; "search result" means seen only as a search snippet or summary. Company figures are self-reported unless they come from a filing or a regulator. Lab revenue and margin figures from leaks are unaudited.

Labs and clouds: documentation, pricing and terms

Company filings and disclosures

Labs' finances and deals (press and leaks)

Inference providers, GPU clouds and gateways (press)

Research, benchmarks and telemetry

Law, regulators and courts

Incidents

Field Guides are learning notes, not legal or compliance advice. Rules and fees change; check the cited primary sources before you act on anything here.