The Platform PM

Field GuideLast reviewed October 2026

Voice AI: real-time voice agents

How software holds a phone call in real time: turn-taking, speech and language models, what a minute costs next to a person, the rules for AI calls, who pays for mistakes, and what's real about replacing human agents.

The industry on one page

The parties in one call. The caller's audio arrives over the phone network or an app, and the voice platform runs the conversation: speech models hear and speak and a language model decides what to say, or a single speech-to-speech model does both, and tools reach your systems. The platform has about a second per turn before the pause feels wrong.

Picture a company that sells payroll software to small businesses. Its support line takes about 100,000 calls a month, mostly from owners whose payroll didn't go out or who want to change a bank account, and it puts a voice agent on that line. A caller dials from a mobile phone, and the call crosses the phone network (mobile carriers, then SIP, the protocol that sets up internet phone calls) to a voice platform. The platform runs the conversation: speech models hear the caller and speak the reply, a language model decides what to say, and tools reach the company's own systems: the CRM, the payroll records and payments (for a clinic or a repair business, the booking calendar). Callers in the company's app reach the same agent over WebRTC, the standard for live audio in browsers and apps.

That's the chained way to build it: speech to text, a language model, then text to speech. A speech-to-speech model can do the hearing, thinking and speaking in one model, and keeps the tone a transcript drops; many teams now run a hybrid, with a fast voice model on the line handing harder reasoning and tool calls to a separate model. Each choice moves latency, control, cost and what you can audit, which I compare in Chained, speech-to-speech or hybrid.

The platform has about a second per turn before the pause feels wrong. Across ten languages, people typically answer within 200 milliseconds of the other person stopping (a 2009 study by Stivers and colleagues). Twilio's target for a chained agent (speech to text, then a language model, then text to speech) is about 1.1 seconds from the caller's mouth to their ear, with 1.4 seconds as its upper limit. Those are a vendor's numbers for its own product, but they set the order of magnitude: a voice agent answers four to ten times slower than a person, so every hundred milliseconds counts.

What I'd want a new PM in this space to take away:

  1. Speed: latency is necessary but not sufficient. On τ-Voice, a benchmark Sierra's research team published in March 2026, the best voice agents completed roughly 30-50% of customer-service tasks on clean audio and 26-38% on phone-quality audio, while a text model completed 85%. Roughly four in five failures sat on the agent's side, mostly wrong logic, mishearing and invented answers. The trade between speed and reasoning now exists inside speech models too: in Artificial Analysis's June 2026 tests, turning up a speech model's reasoning bought accuracy and cost one to two seconds. Builders are converging on a hybrid, a fast voice layer that hands harder thinking to a separate model. See Chained, speech-to-speech or hybrid.
  2. Cost: a voice agent costs about $0.06-0.31 a minute at list prices in October 2026. By my arithmetic, an in-house US agent costs about $0.60-1.45 per handled minute and an offshore agent $0.16-0.40, so the AI minute is far cheaper than a US agent and roughly level with an offshore one once failed calls and fixed costs count. Containment, the share of calls the agent finishes without a person, decides the bill: in How the money moves, doubling it saves five to ten times more than halving the AI price. The biggest line in a mid-priced stack is the platform fee, about $0.05 a minute. Most prices fell from 2024 to 2026, with exceptions: Bland raised its prices in December 2025, and Google lists double prices for some Gemini models from January 1, 2027.
  3. Market: speech to text and text to speech are commoditizing (one platform resells six voice providers at the same $0.015 a minute), yet the biggest voice company by revenue is a model company, ElevenLabs, at about $600 million of annual recurring revenue (its CEO's figure) and a $22 billion valuation in a September 2026 share sale. Investors and buyers pay most for finished agents (Sierra, Decagon, Parloa) and for incumbents that own the contact center (NICE, Genesys).
  4. Rules: since February 2024 the FCC treats AI voices as "artificial" under the Telephone Consumer Protection Act (TCPA), so an outbound AI call without consent risks $500-1,500 per call in private suits, and TCPA class actions nearly doubled in the first half of 2025. Plaintiffs do more enforcing than regulators, with California wiretap suits against AI call vendors, a suit naming OpenAI and Twilio, and biometric suits naming ElevenLabs. After the fake-Biden robocall of January 2024, the carrier paid $1 million while the consultant behind it was acquitted and hasn't paid a $6 million FCC fine. Telling callers they're talking to AI has a measured sales cost, and Maine and the EU AI Act now require it. Platform contracts push the liability onto the customer. See How the rules work.
  5. Risk: OpenAI says its own cloning model needs only 15 seconds of audio, and clones have passed bank voice-ID checks (Lloyds in 2023, Santander and Halifax in 2024). An agent that authenticates callers or changes bank details is a target itself, so I wouldn't let a voiceprint be the only check (see What mistakes cost).

A voice agent is an AI agent with a clock on it, so orchestration, tool security and evals live in AI agents: orchestration platforms and I don't repeat them here. Numbers, caller ID, STIR/SHAKEN and robocall consent are in Phone numbers, messaging and voice, tracing a call across services is in Observability: metrics, logs and traces, and how a team reorganizes around agents is in Becoming AI native. Contact centers appear here only as buyers and channels; their routing and staffing deserve a guide of their own later.

The main players

These are the companies behind the diagram's parties, layer by layer, in no particular order.

Telephony and communications platforms

What they do
Rent numbers, carry calls to and from the phone network, sign caller ID, stream call audio to your code
Main players
Twilio, Telnyx, Bandwidth, Vonage (Ericsson)
What they control
Numbers, caller-ID attestation, the media path and a per-minute toll on every call

Voice agent platforms and open-source frameworks

What they do
Run the turn loop: end-of-turn detection, interruptions, tool calls, transfers, testing
Main players
Vapi, Retell, Bland, LiveKit, Pipecat (open source, from Daily)
What they control
The builder, the call logs and the turn-taking settings

Speech models

What they do
Speech to text, text to speech, voices, and bundled voice agent APIs
Main players
ElevenLabs, Deepgram, AssemblyAI, Cartesia, Gladia (OVHcloud)
What they control
Voices, accents, languages and the per-minute or per-character price

Language model labs

What they do
Text models for chained stacks, and speech-to-speech models that hear and speak in one step
Main players
OpenAI, Google, xAI, Meta
What they control
The ceiling on task success, preview terms, retirements and the price of a minute

Finished voice agents and contact center vendors

What they do
Sell a working agent per conversation or resolution, or add AI to the contact center
Main players
Sierra, Decagon, Parloa, PolyAI, NICE, Genesys
What they control
The enterprise buyer, what counts as a resolution, and the handoff to human agents

Voice security and fraud detection

What they do
Voice authentication, deepfake detection and the analytics behind spam labels
Main players
Pindrop, Nuance (Microsoft), Reality Defender, Hiya
What they control
Whether a caller is trusted and whether an outbound call gets answered

How they make money, and who's moving:

  • Telephony charges per minute and per number. Twilio lists US inbound calls to a local number at $0.0085 a minute, numbers at $1.15-2.15 a month and ConversationRelay, its managed speech layer, at $0.07 a minute. Its revenue reached $1.50 billion in the quarter to June 2026, up 22%, with voice up more than 20% (earnings coverage), and two AI-native customers grew into $6 million and $9 million a year accounts within about 18 months. Bandwidth grew 22% in the same quarter, and all five of its new million-dollar deals involved AI. Telnyx sells its own voice engine at $0.05 a minute including speech. Vonage's unit inside Ericsson is shrinking.
  • Voice agent platforms charge a fee per minute and pass model costs through: Vapi $0.05 with models at cost, Retell $0.055, and Bland $0.12-0.14 all-in except telephony after its December 2025 increase. Vapi raised $50 million in May 2026 at a reported $500 million valuation, a research firm estimates Retell's revenue at about $60 million a year, and LiveKit raised $100 million at $1 billion in January 2026 and says it carries ChatGPT's voice mode. Pipecat is free, with about 16,000 GitHub stars. All of them are squeezed, because model companies, telephony providers and labs now sell bundled agents too.
  • Speech model companies charge per minute of audio or per thousand characters. ElevenLabs gets more than 55% of revenue from enterprises, sold employee shares at $22 billion on September 30, 2026, and cut text-to-speech prices 20% on October 5. Deepgram raised $130 million at $1.3 billion in January 2026, Cartesia about $100 million in late 2025, and I found no AssemblyAI round since December 2023. OVHcloud bought Gladia, a Paris speech-to-text company, in July 2026. All four specialists sell a bundled voice agent API at $0.06-0.08 a minute, which is what a layer does when its own unit gets cheap.
  • Labs charge per audio token or per minute. OpenAI's GPT-Live-1 reached its API on September 10, 2026 at $0.05 a minute for the voice layer, with reasoning billed separately. Google lists Gemini Live audio at half a cent a minute in and under 2 cents out, but its Live API is a preview with no service level agreement. xAI's Grok Voice scored highest on the task-success part of Artificial Analysis's index in June 2026, at 52%. Labs also buy teams: Meta took PlayAI's team in July 2025 and shut its API, and Google DeepMind licensed Hume AI's technology and hired its CEO in January 2026.
  • Finished agents charge per conversation or per resolution: Fin lists $0.99 per outcome, and others reportedly charge $0.50-2.50. Sierra raised $950 million at $15.8 billion in May 2026 on about $150 million of recurring revenue, Decagon was valued at $4.5 billion and Parloa at $3 billion in January 2026, and PolyAI raised $86 million in December 2025. Salesforce agreed to buy Fin for about $3.6 billion in June 2026.
  • Contact center vendors sell AI on top of seats. NICE reported $362 million of AI recurring revenue, 15% of its cloud revenue, in the quarter to June 2026, after buying Cognigy for about $955 million, and Genesys reports more than $400 million. Five9's August 2026 filing says it must replace license revenue lost to AI with revenue from selling AI.
  • Voice security sells per call or per agent seat, at prices that aren't public. Pindrop took $100 million of debt financing in July 2024, Reality Defender raised $33 million in October 2024, and Pindrop and Nuance both won 2026 appeals that kept voice authentication for financial firms outside Illinois's biometric law. Hiya runs analytics behind carrier spam labels and launched branded calling with Deutsche Telekom in January 2026.

As of October 2026. Most figures here are self-reported or come from press coverage of funding rounds, so treat the list as a map to check before relying on it.

Back to the payroll company. Its numbers are rented from Twilio, which carries the call and streams the audio to an agent built on LiveKit. Deepgram transcribes the caller, a language model from OpenAI or Google decides the next step, and ElevenLabs speaks the reply. A tool call checks the payroll run and finds the bank rejected the file. When the caller asks to change the company's bank account, the agent transfers the call to a person in the company's NICE contact center, and Pindrop scores the caller's voice for signs of cloning on the way. One four-minute call touches six or seven companies, each with its own meter (the vendors in the story are illustrative).

How one turn of a call works, step by step

One turn of a call. The hardest call is often the first: deciding the caller has finished speaking, without cutting them off or leaving a long silence. Then the agent answers or calls a tool, hands off to a person when it should, and stops talking when the caller interrupts.

Here's one turn of that payroll call, on a chained pipeline, from a mobile phone, using Twilio's target timings. The hardest decision often comes at the very start: deciding the caller has finished speaking, without cutting them off or leaving a long silence. Then the agent answers or calls a tool, hands off to a person when it should, and stops talking when the caller interrupts.

  1. Caller speaks: "My payroll didn't go out on Friday." The audio streams in across the mobile network and the carriers to the communications platform's edge, where it's buffered and decoded (about 95 ms in Twilio's diagram). Phone audio is narrowband G.711 at 8 kHz, which carries roughly 300-3,400 Hz, so a model that wants 16 kHz audio gets a resampled signal with nothing above about 3.4 kHz. Breaks: background noise, a TV or another person talking; one noise-cancellation vendor claims background voices push transcription errors from 5% to over 30%.
  2. Turn ended: the caller pauses, and the platform has to decide whether they're done. Voice activity detection (VAD) only says whether someone is speaking; end-of-turn detection decides they've finished. A plain silence timer waits about 500 ms by default. Turn models also weigh the words and the intonation: Deepgram claims about 260 ms for its Flux model, and Pipecat's open Smart Turn model runs in about 10 ms on some CPUs. Some stacks start the language model early on a likely end of turn and cancel if the caller keeps going, which saves hundreds of milliseconds and costs 50-70% more model calls (Deepgram's estimate). Breaks: too eager and the agent cuts off a caller reading out an account number; too patient and the line goes quiet. On τ-Voice, realistic turn-taking cut retail task success by 7 points overall and by 11 for Google's model.
  3. Understood: a streaming speech-to-text model sends partial transcripts and then a final one, with a target of 350 ms and an upper limit of 500. In a speech-to-speech model the audio goes to the model directly. Breaks: names, emails, addresses and strings of letters and digits fail most, and OpenAI listed "better alphanumeric recognition" as an improvement in gpt-realtime-2.1. A 2020 study of five commercial recognizers found almost twice the word error rate for Black speakers as for white speakers, and τ-Voice found accents cut task success by 10 points.
  4. Answered: the language model chooses a reply or calls a tool. Time to first token is the largest and most variable slice of the budget, with a target of 375 ms and an upper limit of 750. A tool call, such as looking up the payroll run, adds the tool's round trip and another model call, and the tool definitions and the history are re-sent every turn. Breaks: a slow backend leaves dead air; a speech-to-speech model promises something before the tool returns; a guardrail that checks a transcript trips after the words are spoken.
  5. Spoken: text to speech streams the first audio chunk before the sentence is finished. Twilio targets 100 ms; Vapi's tests in June 2026 measured 159-197 ms for the fastest voices, network included, about twice what vendors quote for the model alone. Add about 10 ms for each of roughly eight hops between services and 40-100 ms back to the caller's ear. Breaks: the 95th-percentile turn is much slower than the median, and a dashboard that leaves out end-of-turn time and the network (Vapi's own estimate does) understates what the caller hears.

Two exits sit off that path. Handed off is the move to a person. A cold transfer sends a SIP REFER through the trunk and the agent drops out; a warm transfer puts the caller on hold, briefs a human, connects them and leaves, and comes back to the caller if nobody answers. The payroll agent hands off whenever someone asks to change bank details. Transfers fail when the carrier rejects the REFER, when the caller ID shown to the human changes unexpectedly, or when a cold transfer loses the context and the caller starts over.

Interrupted is the caller talking over the agent, called barge-in. Stopping takes two jobs: flush the audio queued for playback (Twilio's clear message), then cut the agent's turn in the history down to what the caller actually heard (Twilio's mark messages report what played). Skip the second job and the model believes it said things the caller never heard. The opposite failure is a false stop on a cough, a "mm-hmm" or a TV. τ-Voice also measured interruption rates under realistic audio, 14% for OpenAI's model, 21% for Google's and 84% for xAI's, which shows how far apart the models still are on turn-taking.

Caller to the platform's edge (network, buffer, decode)

Target
About 95 ms
Upper limit
100 ms or more
Whose number
Twilio's diagram

End-of-turn decision

Target
260 ms (a turn model's claim) to 500 ms (a silence timer)
Upper limit
800 ms or more
Whose number
Deepgram; Twilio

Speech to text, final transcript

Target
350 ms
Upper limit
500 ms
Whose number
Twilio

Language model, first token

Target
375 ms
Upper limit
750 ms
Whose number
Twilio

Text to speech, first audio

Target
100 ms target; 159-197 ms measured
Upper limit
250 ms
Whose number
Twilio; Vapi's tests

Hops between services

Target
About 10 ms each, about eight of them
Upper limit
Whose number
Twilio

Total, mouth to ear

Target
About 1.1 s
Upper limit
1.4 s
Whose number
Twilio

A person answering

Target
About 100-200 ms
Upper limit
About 300 ms (the slowest languages' median)
Whose number
Stivers and colleagues, 2009

The stages overlap, so they don't add up exactly. On a phone call, the network alone takes a real share: the international telecom standard treats one-way delay under 150 ms as unnoticeable, and mobile networks are designed to about 200 ms one way.

Chained, speech-to-speech or hybrid

A chained stack (also called cascaded, or a pipeline) runs speech to text, then a language model, then text to speech. Every hop is text you can read, log and check, and you can swap any part. OpenAI's own guidance says to choose it when you need to inspect or transform the text in the middle, for example to run a policy check before the agent speaks.

A speech-to-speech model (sold as "realtime" or "native audio") hears and speaks in one step and keeps the tone that a transcript drops. It isn't automatically faster: Artificial Analysis measured 0.44-0.82 seconds to first audio at low reasoning and 2.3-3.0 seconds at high reasoning, and τ-Voice measured 0.90-1.15 seconds under phone-like audio, the same range as Twilio's chained target. I found no independent test of the two on the same tasks and audio, which is the biggest gap in the evidence. Cost grows with call length, because OpenAI's Realtime API re-sends the whole conversation with each response. By my arithmetic from its prices, five minutes into a call an uncached session costs $0.77-1.15 a minute, against about a cent when the cache works.

A hybrid splits the fast talker from the slow thinker. OpenAI's GPT-Live-1 is full-duplex (it listens while it speaks) and hands reasoning and tool calls to a separate backend model, billed separately; OpenAI reports it cut turn-taking latency to 0.8 seconds from 1.41, a figure I saw only in press coverage. NVIDIA's open research model speaks an on-hold message while a tool runs, and LiveKit, Pipecat and Deepgram start the language model speculatively inside chained stacks.

Turn gap

Chained
About 1.1 s target (Twilio)
Speech-to-speech
0.4-0.8 s at low reasoning, 2.3-3.0 s at high; 0.9-1.15 s on phone-like audio
Hybrid
0.8 s, as reported by OpenAI

Control

Chained
Text at every hop; swap any part
Speech-to-speech
Transcripts are a side output
Hybrid
The backend is inspectable; the voice layer isn't

Guardrails

Chained
Before the agent speaks
Speech-to-speech
During or after speech
Hybrid
On the backend's output

Task success evidence

Chained
No τ-Voice data; text models reach 85%
Speech-to-speech
26-51% on τ-Voice
Hybrid
86% on a benchmark OpenAI cites, via press only

Cost shape

Chained
Sum of meters, roughly linear in minutes
Speech-to-speech
Audio tokens; grows with call length unless cached
Hybrid
A flat voice minute plus backend tokens when it delegates

Lock-in

Chained
Low
Speech-to-speech
High: one vendor's events and voices
Hybrid
High on the voice layer, low on the backend

My default: chained or hybrid for anything with a regulated script, a payment or a policy the agent must quote exactly, and speech-to-speech where tone matters more than exact words and calls are short. Whatever the architecture, I'd measure mouth to ear on real phone calls at the 95th percentile.

Which voice agent setup? Five questions

  1. Inbound or outbound? Outbound brings TCPA consent, caller-ID labels and disclosure, which decide whether a campaign is viable at all. The TCPA doesn't treat inbound calls as "made" by the business, so inbound comes down to whether the agent resolves calls well enough, plus recording notices and disclosure.
  2. What does a wrong answer cost? If the agent quotes prices, refunds or policy, the business is bound by what it says. Use a stack that can check the text before it's spoken, and read back anything the caller must get right.
  3. What data does it touch? Card numbers should go to keypad entry with masking or a secure link so the model never hears them; health data needs a business associate agreement (BAA) with every vendor in the chain; voiceprints bring biometric law. Each narrows which vendors and models you can use.
  4. How long are the calls? Long calls raise speech-to-speech costs unless caching holds, Google's Live API drops connections after about 10 minutes and has to resume them, and Vapi's default maximum call length is 600 seconds.
  5. Who pays when it fails? On per-minute pricing you pay for a failed six-minute call and then for the human minutes after it. On per-resolution pricing the vendor carries that, and the argument moves to what counts as resolved.

My defaults: for an inbound line like the payroll company's, I'd have the agent answer every call, verify the caller and handle the few most common requests, with a warm transfer for everything else and for any change to bank details. I'd say "AI assistant" in the greeting, announce recording and name the vendor that handles the audio, and keep card numbers on the keypad. I'd measure containment and resolution separately, with a person checking a sample of transcripts each week, and widen the agent's scope only once resolution holds. For outbound, I wouldn't dial without consent records I'd be happy to show a court.

The primitives

01

Entity & identity

What is the unit of record, and how do we know it is the same one?

Three identities meet on every call: the business the agent speaks for, the voice (synthetic, cloned or human) and the phone number. The caller's own number can be spoofed. STIR/SHAKEN attestation, when present, says how much the originating carrier vouches for it: an A means the carrier knows its customer may use the number, and says nothing about the voice or the words.

Inside the stack, one call collects IDs from the carrier, the communications platform, the voice platform, the model session and the trace. Joining them is real engineering work, and you need it before you can debug a bad call or reconcile a bill.

The law names its own entities, and one platform can be several at once:

Legal roleLawWhat it triggers
The "maker" or initiator of a callTCPAConsent, caller identification, opt-out
A third party listening inCalifornia's wiretap law$5,000 per violation if callers weren't told
Provider or deployerEU AI ActWatermarking for the provider; disclosure for both
Business associateHIPAAA BAA, and BAAs with every sub-vendor
A collector of voiceprintsIllinois's biometric law (BIPA) and othersWritten consent and a retention policy

A voiceprint is a biometric template that identifies a speaker. A plain transcription agent generally doesn't build one, though courts are testing that line in cases about meeting transcription tools.

More on Entity & identity →

02

State & lifecycle

What states exist, and what moves an entity between them?

A call carries several state machines that move at different speeds.

  • Turn: listening, caller speaking, end of turn pending (when the model starts early), thinking, tool running, agent speaking, interrupted, then back to listening.
  • Call: ringing, answered by a person or a machine, in progress, transferring, ended (completed, no answer, busy, failed or cancelled).
  • Model session: Google's Live API adds connected, "going away" (about 60 seconds' warning) and resumed, because a connection lasts about 10 minutes and an audio session about 15.
  • Consent: none, prior express, prior express written, then revoked, and a revocation has to be honoured within 10 business days.

On outbound calls, answering-machine detection takes about 4 seconds after pickup with Twilio's defaults. A model upgrade changes turn-taking behaviour as well as answers, so I'd put timing in every regression test.

More on State & lifecycle →

03

System of record & ledger

Who owns the truth, and how do systems reconcile?

Four records describe one call, and they disagree by design.

FactSystem of record
Minutes billed for the callThe carrier's and communications platform's call detail records
What the agent did, turn by turnThe voice platform's call log
What was saidThe transcript and the recording
What changed in the businessThe tool-call log in the business's own systems
Whether the caller agreed to be calledThe deployer's consent records

Two details matter. In a speech-to-speech model the transcript is a side product of a separate recognizer, so an audit or eval built on transcripts can miss what the model actually "heard". And after an interruption, what the model thinks it said and what the caller heard diverge unless the history was cut back.

Consent records are the defence in every TCPA case, and the platforms say plainly that they don't keep them for you: Retell's terms say it "does not obtain consent on behalf of Customer". Billing is its own reconciliation job, since five vendors can meter one call in carrier minutes, stream time, audio tokens, characters and voice minutes billed by the second.

More on System of record & ledger →

04

Rules & policy

What logic decides outcomes, and who can change it?

Turn-taking settings are policy decisions: how long a silence ends a turn, how eager the model is to answer, how long a sound has to last before it counts as an interruption. A longer timeout while the caller reads out digits protects accuracy and costs speed.

Guardrails behave differently by architecture. A chained stack can check the text before it's spoken. On a speech-to-speech stream the check runs alongside the speech, and an open issue on OpenAI's agent toolkit reports the agent kept talking after a guardrail tripped. Required disclosures and consent prompts are fixed text inside a probabilistic stream, and a chained or hybrid stack makes it easier to guarantee they're said word for word.

Above the builder's settings sit laws, card network rules, carrier analytics and the platform's and model provider's usage policies. Retell's terms, for example, require outbound agents to identify themselves and ban calls to emergency lines.

More on Rules & policy →

05

Effective dating

Which version of the rule applied at that moment?

Dates I'd keep on a calendar as of October 2026:

ChangeEffectiveStatus (Oct 2026)
FCC: AI voices are "artificial" under the TCPAFebruary 8, 2024In force, and applied to calls made before it in pending suits
California: prerecorded calls must say if the voice is artificialJanuary 1, 2025In force
Maine: AI chatbots, spoken ones included, need a clear noticeSeptember 2025In force
Carriers must sign caller ID with their own tokenSeptember 18, 2025In force
Carriers blocking on analytics must return a specific SIP code (603+)March 25, 2026In force
EU AI Act: tell people they're talking to AI; mark synthetic audioAugust 2, 2026In force
FCC comment deadline on letting political AI calls reach mobiles without consentOctober 19, 2026Upcoming
EU: marking deadline for audio systems already on the marketDecember 2, 2026Upcoming
Google's Gemini Flash and Flash speech models double in priceJanuary 1, 2027Upcoming

Some dates float. Deepgram's streaming prices are "limited-time promotional rates" with no end date, so I'd budget at the regular rate. Google's Live API is a preview with no deprecation guarantees, and some OpenAI model names point at a default snapshot that can change, so I'd pin dated snapshots where offered. The FCC's 2024 proposal to require AI disclosure at the start of calls hasn't been adopted, and Canada's regulator opened its own review in June 2026.

More on Effective dating →

06

Interfaces & standards

What format and protocol do counterparties speak?

The phone side is standardized: SIP for setting up and transferring calls, RTP and SRTP for the audio, G.711 at 8 kHz on the phone network, Opus in apps, E.164 numbers, STIR/SHAKEN for caller ID, and DTMF for keypad tones. The AI side has no shared standard. Every platform has its own WebSocket event protocol (Twilio's media streams, OpenAI's Realtime events, Google's Live messages, Deepgram's turn events), and nothing covers turn events, barge-in or tool calls across vendors. That's a lock-in point, and the reason open frameworks like Pipecat and LiveKit's agents exist.

WebRTC

Typical use
A browser or app talking to a model or platform
Strengths
Built for live audio: handles loss, jitter and echo; OpenAI's recommended client path
Weaknesses
Needs media servers and network traversal

WebSocket media stream

Typical use
A communications platform sending call audio to your server
Strengths
Simple, works everywhere
Weaknesses
Lost packets stall the stream; you handle buffering and barge-in

SIP trunk to a platform or model

Typical use
Phone numbers into LiveKit, OpenAI or a voice platform
Strengths
Fewer hops; native transfers; carrier-grade
Weaknesses
You still need a trunk provider; caller ID can change on transfer

The phone network end to end

Typical use
Any phone
Strengths
Reaches everyone
Weaknesses
8 kHz audio, extra delay, spam labels and blocking

OpenAI now accepts SIP calls directly, with a separate EU endpoint, and warns builders not to retry an outbound call automatically after an ambiguous timeout, because the retry can place a second call. For provenance, Google's SynthID and Meta's AudioSeal watermark synthetic audio and C2PA labels files, but a watermark only marks the provider's own output, and phone audio degrades it.

More on Interfaces & standards →

07

Networks & counterparties

Who sits between us and the outcome, and what do they want?

A call passes through the caller's carrier, transit carriers, a communications platform or SIP trunk, the voice platform, one to three model vendors (often on different clouds) and the business's own systems. Twilio counts at least ten network traversals per turn of a chained agent, and each boundary can add another encode, decode and buffer. Moving the pieces closer together is mostly about latency: Telnyx places GPUs next to its telephony sites, and speech-to-speech models collapse three vendors into one.

PartyWhat they controlWhat they earn
Communications platform or carrierNumbers, caller-ID attestation, the audio streamPer-minute and per-number fees, plus carrier surcharges
Terminating carriers and analytics engines (Hiya, TNS, First Orion)Spam labels and blockingBranded-calling services sold to businesses
Voice platformThe turn loop, logs, transfersA per-minute platform fee
Model vendorsAccuracy, voices, latency, retirementsPer token, minute or character

The analytics engines decide whether an outbound agent is heard at all. An outbound agent dials many short calls from few numbers, which is the pattern that earns a "Spam Likely" label, and a caller-ID company's 2026 survey found 86% of consumers say they won't answer an unidentified call (a vendor survey). Registration and branded calling are in Phone numbers, messaging and voice.

More on Networks & counterparties →

08

Regulatory layering

Jurisdiction × activity × entity type: is it a license or a certification?

CallLayers that apply
Outbound, USTCPA and the FCC's rules, state telemarketing laws, state AI-disclosure laws, state recording laws, carrier policy, the platform's terms
Inbound, USState AI-disclosure laws (Maine, Utah), recording and wiretap laws, biometric laws if voiceprints are used
Any call in the EUAI Act Article 50, GDPR, national telemarketing and recording rules
Health or card data, anywhereHIPAA and BAAs; PCI DSS through the card networks
CanadaThe CRTC's rules on automated calls (which already cover synthesized voices), privacy law

One agent usually serves callers in every state, so in practice the strictest layer wins: an all-party recording notice on every call, AI disclosure in every greeting.

More on Regulatory layering →

09

Exceptions & reversals

What goes wrong, and how is it undone?

Most reversals on a call are small and happen every minute: a speculative reply is cancelled when the caller keeps talking, the agent's turn is cut back after an interruption, a failed transfer comes back to the agent. The dangerous ones are side effects. A tool call cancelled mid-flight because the caller changed their mind may already have done something, and I couldn't find how vendors handle that. An outbound dial retried after an ambiguous timeout can ring the same person twice, a second TCPA exposure.

Caller revokes consent

Who starts it
The caller, by any reasonable means
Clock
Honour within 10 business days
The way back
Suppress the number everywhere

Transfer fails

Who starts it
Trunk, carrier or an unanswered human
Clock
Seconds
The way back
Return to the agent; offer a callback

Model or price change

Who starts it
The vendor
Clock
Days to months; previews without notice
The way back
Pinned snapshots; regression tests with timing

Vendor shuts down after an acquisition

Who starts it
The acquirer
Clock
Weeks: PlayHT's API went dark in July 2025
The way back
Re-integrate; cloned voices may be lost

Deployment pulled back

Who starts it
The business
Clock
After the damage
The way back
Rehire, apologize, rethink scope

Commercial terms reverse too. Bland raised its per-minute price by 22-56% in December 2025, and ElevenLabs said in February 2025 that it was absorbing language model costs and would eventually pass them on.

More on Exceptions & reversals →

10

Liability allocation

When it fails, who pays?

The deployer owns what its agent says. When Air Canada's chatbot invented a bereavement-fare rule in 2024, the tribunal rejected the airline's argument that the bot was a separate legal entity and made it pay about C$812, and the same principle applies to a voice agent that quotes a refund on a recorded call.

Platform contracts push the rest down to the customer:

ClauseRetell (June 2026)Vapi (September 2026)
Who gets consentThe customer onlyThe customer
AI disclosureRequired at the start of outbound callsNot explicit
Customer indemnityCovers regulatory fines, the TCPA included; uncappedCovers TCPA claims
Vendor's cap12 months of feesThe greater of $100 or 12 months of fees
AccuracyNot warranted; the customer assumes the riskAll warranties disclaimed

Plaintiffs are testing whether that holds. A December 2025 suit claims OpenAI and Twilio are liable as makers of AI robocalls that another company sent through them, two California rulings in 2025 let wiretap claims proceed against AI vendors able to use call audio for their own models, and nine biometric suits filed in May 2026 accuse ElevenLabs and big tech companies of building voiceprints from public recordings to train models. An indemnity doesn't stop a plaintiff from suing the vendor directly; it only moves the bill, and only if the customer can pay.

Carriers carry their own share through attestation: Lingo Telecom paid $1 million for signing spoofed AI calls. Fraud losses depend on the payment rail and the country. In the US the victim of a scam that tricks them into sending money usually bears it. Pindrop reportedly offers a deepfake warranty of up to $1 million a claim, the only explicit transfer of this risk from a security vendor I found.

More on Liability allocation →

What's different here

How the money moves

The business pays everyone, usually per minute, and the meters start at different moments. Here's the per-minute stack at US list prices in October 2026.

Streaming speech to text

Low
$0.0025
Typical
$0.005-0.008
High
$0.017

Text to speech

Low
About $0.01
Typical
$0.015-0.03
High
$0.08-0.10 (premium voices)

Language model, chained

Low
$0.002-0.003
Typical
$0.008-0.064
High
$0.32-0.64 (frontier, fast tier)

Speech-to-speech model

Low
$0.023 (Gemini Live, both directions)
Typical
$0.05 (GPT-Live-1 voice layer) plus backend
High
$0.35-0.38 (OpenAI's realtime model resold by Retell)

Telephony, per leg

Low
$0.003-0.004 (SIP)
Typical
$0.0085-0.015
High
$0.022 (toll-free inbound)

Platform fee

Low
$0.05
Typical
$0.055-0.08
High
$0.12-0.14 (all-in)

All-in, per call minute

Low
About $0.06
Typical
About $0.10-0.15
High
About $0.31-0.45

In a mid-priced chained stack at about $0.12 a minute, the platform takes about 45%, the language model and the voice 10-30% each, telephony 5-12% and speech to text about 5%. Model prices fell while platform fees seem to have held at about $0.05: OpenAI cut its realtime audio prices 60% in December 2024 and another 20% in August 2025, and ElevenLabs halved its agent prices in February 2025.

The traps sit in the billing units. AssemblyAI bills streaming by session time, silence included, and GPT-Live-1 bills the whole session, backend thinking included. Telnyx rounds each call up to the minute, so wrong numbers cost a full minute. Speech-to-speech models re-bill the conversation each turn, and changing instructions mid-call breaks the cache. And plans cap simultaneous calls, with extra lines at about $10 a month each at Vapi.

A month of calls on the payroll support line

My assumptions: 100,000 inbound calls a month at 4 minutes of talk each; a person needs 5 minutes per call including after-call work; calls the agent can't finish spend 1.5 AI minutes before a transfer and 4.5 human minutes after; the AI costs $0.06, $0.12 or $0.30 a minute; and fixed AI costs (platform minimums, one to three people designing and checking conversations, evals and integration) come to $15,000-50,000 a month. The human ranges are my arithmetic from the US median wage for customer service representatives ($21.53 an hour in May 2025) and outsourcing guides' loaded rates.

SetupWith US in-house staffWith offshore staff
All human (500,000 handled minutes)$295,000-720,000$80,000-200,000
Agent answers first, 30% contained$215,000-571,000 (about 21-27% less)$79,000-243,000 (1% less to 22% more)
Agent answers first, 60% contained$139,000-399,000 (about 45-53% less)$62,000-212,000 (22% less to 6% more)
Agent verifies and routes; people take every call$252,000-636,000 (about 12-15% less)$80,000-220,000 (up to 10% more)

Against US staff, the AI usage bill is 4-12% of the old labor bill and containment sets the savings. Each extra point of containment, 1,000 calls, is worth about $2,500-5,700 a month, so moving from 30% to 60% saves about $75,000-172,000 a month, while halving the AI price from $0.12 to $0.06 saves about $13,500-18,000. Against offshore staff at 30% containment, the program can cost more than it saves once fixed costs count. The last row is how contact center incumbents mostly sell AI: lower risk, smaller savings.

Priced per outcome instead, 60,000 resolutions a month would cost about $59,000 at Fin's $0.99 and $90,000-150,000 at the $1.50-2.50 reported for other vendors, against $14,000-72,000 of raw AI usage for the same calls. The vendor keeps that spread for carrying the failures and the model costs. I think buyers underrate the outcome definition: Fin bills handoffs that follow a procedure and self-serve routings as outcomes too, and a caller who hangs up angry still counts as contained.

Who holds the power

Power follows whoever owns the customer relationship, the model and the phone line.

  • Contact center incumbents (NICE, Genesys, Five9) own routing, the human agents' desktops and the handoff, and their AI revenue already exceeds most startups' total revenue. Their risk is that every contained call is a seat they no longer sell, as Five9's filing says.
  • Finished-agent vendors (Sierra, Decagon, Parloa, PolyAI, Fin) own the outcome definition. Investors value them highest, at roughly 45-105 times recurring revenue by my rough arithmetic, and CRMs are buying in: Salesforce is buying Fin. Sierra also wrote τ-Voice, the benchmark the industry cites, which I'd keep in mind when reading it.
  • Labs set the ceiling on task success and the price of the premium tier, and they're moving up the stack with full-duplex voice, direct phone connections and talent deals. They earn through every layer whoever wins the customer.
  • Speech model companies face commodity pricing on the basic tier and still earn a premium for expressive voices: on Retell, ElevenLabs voices cost 2.7-6.7 times the commodity set.
  • Communications platforms and carriers hold numbers, porting, attestation and the audio path. Telephony is only 5-12% of a mid-priced stack, but it's sticky, and Twilio, Telnyx and Bandwidth all sell their own agent layers.
  • Plaintiffs' lawyers and the carriers' analytics engines hold the most practical power over outbound calls. The lawyers price non-compliance and the analytics engines decide whether calls get answered, and both matter more as the FCC loosens its TCPA rules.
  • Developer platforms (Vapi, Retell, Bland) own the builder and are squeezed from below by model companies, from the side by telephony providers and from above by labs and contact center suites. Vapi's answer is enterprise work; Bland's is all-in pricing.

How the rules work

The FCC's ruling. On February 8, 2024 the FCC ruled unanimously that AI-generated and cloned voices are "artificial" under the TCPA. Outbound AI calls need prior express consent, written and naming the seller for telemarketing; they must identify the caller at the start and, for telemarketing, offer an automated opt-out. The FCC refused any carve-out for technology that claims to work like a live agent, and the 26 state attorneys general who asked for the ruling can sue under the TCPA too. Private damages are $500 per call, $1,500 if wilful, with a four-year limitation period. Since a June 2025 Supreme Court decision, courts read the TCPA for themselves without deferring to the FCC; an appeals court had already read "artificial voice" to include AI in 2023, so I'd treat the question as settled.

The FCC's direction has turned since. Its August 2024 proposal to require AI disclosure at consent and at the start of each call hasn't been adopted, the one-to-one consent rule was deleted in 2025, and in September 2026 it asked for comment on letting political AI-voice calls reach mobile phones without consent.

Plaintiffs do the enforcing. A litigation tracker counted 1,052 TCPA class actions from January to June 2025, against 539 a year earlier. The AI-voice suits I found turn on consent, ignoring "no", wrong numbers and state registration rather than model quality: a February 2026 suit against a mortgage lender alleges more than $5 million in damages, and a 2026 Texas suit says an AI agent kept selling after the person said no. The suit against OpenAI and Twilio tests whether model and communications providers can count as callers, and I found no ruling yet.

Fake-Biden robocall. On January 21, 2024, calls with a cloned Biden voice told New Hampshire Democrats not to vote in the primary. A political consultant commissioned them, paying a magician $150 to make the clip in under 20 minutes. Lingo Telecom, which signed the spoofed calls with the highest level of caller-ID attestation, paid $1 million under an August 2024 consent decree. The consultant was fined $6 million by the FCC, still refused to pay as of November 2025, and was acquitted of all 22 criminal counts in June 2025. The penalty that landed hit the carrier's attestation.

Disclosure. Maine's law, in force since September 2025, covers chatbots that talk as well as type and requires a clear notice wherever a reasonable consumer could think they're talking to a person. Utah requires disclosure when asked, and up front in high-risk interactions, until July 2027, and California requires prerecorded calls to say if the voice is artificial. Since August 2, 2026 the EU AI Act has required that people are told they're dealing with AI unless it's obvious, and that providers mark synthetic audio, with fines up to €15 million or 3% of turnover. Disclosure has a measured cost: in a 2019 field experiment with more than 6,200 customers of a financial firm, telling people up front that a sales call came from a bot cut purchases by more than 79.7%. That was pre-LLM technology in one market, and I found no newer replication.

Recording and wiretaps. About 11 states require every party's consent to record. The sharper risk is California's wiretap law. In February 2025 a court let a suit proceed against Google over its contact center AI on Verizon support calls, holding that a vendor capable of using call data for its own purposes counts as an eavesdropper whether or not it actually does, and an August 2025 ruling did the same for an AI pizza-ordering vendor. At $5,000 per violation, zero-retention and no-training settings and a greeting that names the vendor are worth their price.

Biometrics, health and cards. BIPA pays $1,000-5,000 per person for voiceprints collected without written consent; voice-authentication vendors serving financial firms won two appeals in 2026 under its exemption for financial institutions, which doesn't cover a vendor serving a clinic or a retailer. Under HIPAA a voice platform handling patient data is a business associate, and compliance becomes a price tier: Vapi charges $2,000 a month for HIPAA and $1,000 for zero data retention, and ElevenLabs signs BAAs only on its enterprise tier. Under PCI DSS, a card number spoken to the agent pulls the transcription, the model's context, logs, recordings and eval sets into scope, which is why card capture moves to the keypad.

Canada. The CRTC's rules on automated calls already cover "synthesized" voices and require express consent for telemarketing. A review opened in June 2026 asks whether to require AI disclosure and whether the rules should reach a consumer's own AI assistant making a booking.

TCPA and the FCC's 2024 ruling

Outbound
Yes
Inbound
No
Penalty
$500-1,500 per call
Status (Oct 2026)
In force

Maine's chatbot disclosure law

Outbound
Yes
Inbound
Yes
Penalty
Unfair trade practice
Status (Oct 2026)
In force since September 2025

California recording and wiretap law

Outbound
Yes
Inbound
Yes
Penalty
$5,000 per violation
Status (Oct 2026)
In force; vendor suits proceeding

BIPA and similar laws

Outbound
If voiceprints
Inbound
If voiceprints
Penalty
$1,000-5,000 per person
Status (Oct 2026)
In force

EU AI Act Article 50

Outbound
Yes
Inbound
Yes
Penalty
Up to €15 million or 3%
Status (Oct 2026)
In force since August 2, 2026

HIPAA; PCI DSS

Outbound
If health or card data
Inbound
If health or card data
Penalty
Regulator; card network fines
Status (Oct 2026)
In force

Canada's automated-call rules

Outbound
Yes
Inbound
No
Penalty
CRTC penalties
Status (Oct 2026)
Under review

What mistakes cost

Let's say a mortgage lender runs an outbound AI campaign to 100,000 leads. About 30-40% connect, for 2.5 minutes on average, so the AI minutes cost about $7,500-30,000 at $0.10-0.30 a minute. If the consent behind 5% of those dials is defective, that's 5,000 calls at $500 each, a $2.5 million statutory floor, or $7.5 million if a court finds it wilful, before defence costs and state claims. One bad source of leads can cost 100 to 1,000 times the AI minutes, and the platform's terms make sure the lender carries it.

CaseWhat happenedWho paid
Fake-Biden robocall, 2024Cloned voice, spoofed numbers, A-level attestationThe carrier, $1 million; the consultant's $6 million fine unpaid
Commonwealth Bank of Australia, 2025Cut 45 roles, claiming its voice bot reduced calls by 2,000 a week; the union showed volumes were risingThe bank reversed the cuts and apologized
McDonald's and IBM, 2024AI drive-thru ordering in about 100 restaurants; accuracy estimated in the low-to-mid 80s against a 95% targetThe test ended in July 2024
UK energy firm, 2019A cloned parent-company CEO's voice asked for a transferAbout €220,000, covered by insurance
UAE bank, 2020A cloned company director's voice plus forged emailsAbout $35 million
Arup, 2024A video call where every other participant was a deepfakeAbout $25 million

The fraud numbers keep growing. The FBI's crime complaint center counted $893 million of losses in 2025 complaints that cited AI, out of $20.9 billion in total, with no split for voice. And the bank cases show voice ID itself under pressure: a journalist used a cloned voice to get into a Lloyds account in 2023, the BBC passed Santander's and Halifax's voice ID with a clone in 2024, and OpenAI's chief executive told a Federal Reserve conference in July 2025 that it was "crazy" for institutions still to accept voiceprints.

Voice agents vs human agents: what's real

My view as of October 2026: a voice agent is cheap and fast enough for most inbound calls, and still fails too many tasks to replace the people behind it. The independent evidence on task success is recent and sobering, deployment numbers are almost all vendor-reported, and the public pull-backs came from quality and measurement more than from regulators.

What the benchmarks say

  • Task success: τ-Voice (278 tasks in airline, retail and telecom service, published in March 2026 and accepted at ICML) tested speech-to-speech agents on clean audio and on realistic audio with noise, 8 kHz phone quality, dropped packets and varied accents. The best completed about 30-50% of tasks on clean audio and 26-38% on realistic audio, against 85% for a text model. No chained stack was tested.
  • Who did best: in the paper, xAI's and OpenAI's agents led (51% and 49% on clean audio) and Google's trailed (31%); in Artificial Analysis's June 2026 index, xAI's Grok scored 52% on the same benchmark and OpenAI's top model at high reasoning 40%. Rankings move with each model version, so I'd test the current ones myself.
  • The vendors' own claims: OpenAI reports GPT-Live-1 at 86% on a benchmark it calls "Tau3 Voice", up from 46% for its previous model, a figure I saw only in press coverage.

What deployments report

  • Utilities: PG&E's line, run by PolyAI, contains 67% of calls, which PolyAI's own page says is 6 points above the old phone menu, and saves 35,000 agent hours.
  • Energy: Parloa says it resolves up to half of E.ON's inbound calls.
  • Customers: Gartner's survey of 5,728 customers found only 14% of service issues fully resolved through self-service, and 64% would prefer companies not use AI for service. Qualtrics found AI customer service fails at almost four times the rate of other AI uses, though neither survey is about voice alone.

Every named deployment figure is the vendor's own, none is audited, and containment, resolution and routing accuracy are different metrics. PG&E's case is the one I'd show a skeptic, because it says what the old phone menu already achieved.

Who pulled back

  • McDonald's ended its IBM drive-thru test in July 2024 after accuracy estimated in the low-to-mid 80s, with accents a reported problem.
  • Taco Bell said in August 2025 that it was rethinking where to use voice AI in its drive-thrus, after prank orders and trouble at peak times.
  • Commonwealth Bank of Australia reversed 45 job cuts in August 2025 when the union showed call volumes were rising.
  • Klarna, in chat, said in May 2025 that its cost focus had produced "lower quality" and began hiring people again.

Gartner predicted in June 2025 that half the companies planning big service headcount cuts because of AI would abandon those plans by 2027. In January 2026 it predicted, in a release I could see only by its title and press coverage, that generative AI's cost per resolution will exceed offshore human agents' by 2030. How a support organization changes around agents, beyond shrinking, is in Becoming AI native.

What containment is worth

The arithmetic in How the money moves makes reliability the bigger lever by roughly five to ten times. Containment also flatters: a caller who hangs up angry counts as contained, and the FTC sued a company in August 2025 for claiming its AI phone agent could replace sales staff. Disclosure cuts the other way on outbound sales, with the 2019 experiment's drop of more than 79.7% in purchases once callers knew. For inbound service, where Maine and the EU now require disclosure, I haven't seen data on its cost.

Questions to ask a vendor

  1. Of the calls your agent handled for customers like us last quarter, what share was resolved, as opposed to merely not transferred, and who checked?
  2. What's the mouth-to-ear turn gap at the 95th percentile on real phone calls, end-of-turn time included?
  3. How do you test with noise, accents, 8 kHz audio and callers who interrupt, and will you run our own scenarios?
  4. When the agent hands off, what does the human see, and what happens if nobody answers?
  5. Do you or your model providers retain or train on our call audio, and can we turn that off in writing?

What usually goes wrong

SymptomLikely causeFirst thing to check
Callers say the agent is slowLong model response time, tool round trips, high reasoning effort, hops across cloudsMouth-to-ear time at the 95th percentile, by stage
The agent talks over callersSilence timer too short; slow speakers; callers reading digitsEnd-of-turn settings during data capture; keypad entry
The agent stops for no reasonBackchannels, coughs, a TV, its own voice echoingFalse-interruption rate; interruption classifier
Wrong names, emails or account numbers8 kHz audio, accents, noiseRead-back confirmation; spelling and keypad fallbacks
Calls end mid-conversation at 10 minutesPlatform maximum duration or a model connection limitMax-duration settings; session resumption
Outbound calls go unanswered"Spam Likely" labels, weak attestation, number reputationRegistration with analytics engines; your own signing token; SIP 603+ responses
A customer got two callsAn ambiguous timeout retried automaticallyDial logic and retry rules
The bill grows with call lengthSpeech-to-speech context replay; cache broken by mid-call changesCached share of input; static instructions and tools
A lawyer's letter about recordingThe vendor can train on audio; no notice naming itRetention and training settings; the greeting
A cloned voice gets throughVoiceprint as the only checkA second factor for money movement and account changes

Words that mean something else here

TermWhat you'd assumeWhat it means here
LatencyOne numberAt least six: model inference, first audio, first token, end of turn, the platform's turn gap and mouth to ear. Ask which, and at which percentile
RealtimeInstantOpenAI's product name, any streaming speech-to-speech model, or loosely anything under a second
InterruptionThe caller cutting inEither direction: the caller interrupting the agent (wanted) or the agent cutting off the caller (not wanted); an "interrupt rate" can mean either
HandoffPassing to a personUsually a transfer, by SIP; in agent frameworks, a switch to another AI agent
AgentThe AIAn AI agent, a human contact center agent, or in SIP any endpoint. In "transfer to an agent", it's the human
MinuteSixty seconds of a callCarrier minutes, stream time with silence, audio tokens, or minutes of generated speech; "$0.05 a minute" can cover very different scope
ContainmentResolutionCalls not transferred to a person, abandoned calls included
Artificial voiceRobotic-soundingUnder the TCPA, any AI-generated or cloned voice, a live conversational agent included
Attestation AThe call is legitimateThe carrier vouches for the customer's right to the number, not the voice or the content
TranscriptWhat the model readIn a chained stack, yes; in speech-to-speech, a separate recognizer's guess at what was said

What surprised me

Placeholders in your voice, drafted from the research and the earlier guides. Rewrite each with your own moment.

The rate deck. In telco I could multiply minutes by a price per minute. Here one four-minute call is metered five ways by five vendors, and on a speech-to-speech model the fifth minute costs more than the first.

"Faster means better." The fastest agents in the benchmark still got about half the tasks wrong on phone audio. Speed got them into the conversation and logic decided the outcome.

"Inbound is the safe side." The TCPA barely touches calls the business answers, and then a California court treated the AI vendor on the line as an eavesdropper.

Caller ID. In the telco guide, attestation was a technical detail. In the Biden robocall it was the one penalty that got paid.

"The cheap part is the AI." The AI minute was a rounding error next to the human minutes that follow every call it can't finish.

Sources

Undated entries were read on October 8, 2026; "search result" means seen only as a search snippet or summary. Company figures are self-reported unless they come from a filing or a regulator, and vendor benchmarks measure the vendor's own products or customers.

Regulators, laws and courts

Labs and speech model companies: documentation and pricing

Telephony and voice agent platforms

Finished agents, contact centers and deals

Voice security and fraud

Benchmarks, research and outcomes

Field Guides are learning notes, not legal or compliance advice. Rules and fees change; check the cited primary sources before you act on anything here.