Field GuideLast reviewed October 2026
Voice AI: real-time voice agents
How software holds a phone call in real time: turn-taking, speech and language models, what a minute costs next to a person, the rules for AI calls, who pays for mistakes, and what's real about replacing human agents.
The industry on one page
Picture a company that sells payroll software to small businesses. Its support line takes about 100,000 calls a month, mostly from owners whose payroll didn't go out or who want to change a bank account, and it puts a voice agent on that line. A caller dials from a mobile phone, and the call crosses the phone network (mobile carriers, then SIP, the protocol that sets up internet phone calls) to a voice platform. The platform runs the conversation: speech models hear the caller and speak the reply, a language model decides what to say, and tools reach the company's own systems: the CRM, the payroll records and payments (for a clinic or a repair business, the booking calendar). Callers in the company's app reach the same agent over WebRTC, the standard for live audio in browsers and apps.
That's the chained way to build it: speech to text, a language model, then text to speech. A speech-to-speech model can do the hearing, thinking and speaking in one model, and keeps the tone a transcript drops; many teams now run a hybrid, with a fast voice model on the line handing harder reasoning and tool calls to a separate model. Each choice moves latency, control, cost and what you can audit, which I compare in Chained, speech-to-speech or hybrid.
The platform has about a second per turn before the pause feels wrong. Across ten languages, people typically answer within 200 milliseconds of the other person stopping (a 2009 study by Stivers and colleagues). Twilio's target for a chained agent (speech to text, then a language model, then text to speech) is about 1.1 seconds from the caller's mouth to their ear, with 1.4 seconds as its upper limit. Those are a vendor's numbers for its own product, but they set the order of magnitude: a voice agent answers four to ten times slower than a person, so every hundred milliseconds counts.
What I'd want a new PM in this space to take away:
- Speed: latency is necessary but not sufficient. On τ-Voice, a benchmark Sierra's research team published in March 2026, the best voice agents completed roughly 30-50% of customer-service tasks on clean audio and 26-38% on phone-quality audio, while a text model completed 85%. Roughly four in five failures sat on the agent's side, mostly wrong logic, mishearing and invented answers. The trade between speed and reasoning now exists inside speech models too: in Artificial Analysis's June 2026 tests, turning up a speech model's reasoning bought accuracy and cost one to two seconds. Builders are converging on a hybrid, a fast voice layer that hands harder thinking to a separate model. See Chained, speech-to-speech or hybrid.
- Cost: a voice agent costs about $0.06-0.31 a minute at list prices in October 2026. By my arithmetic, an in-house US agent costs about $0.60-1.45 per handled minute and an offshore agent $0.16-0.40, so the AI minute is far cheaper than a US agent and roughly level with an offshore one once failed calls and fixed costs count. Containment, the share of calls the agent finishes without a person, decides the bill: in How the money moves, doubling it saves five to ten times more than halving the AI price. The biggest line in a mid-priced stack is the platform fee, about $0.05 a minute. Most prices fell from 2024 to 2026, with exceptions: Bland raised its prices in December 2025, and Google lists double prices for some Gemini models from January 1, 2027.
- Market: speech to text and text to speech are commoditizing (one platform resells six voice providers at the same $0.015 a minute), yet the biggest voice company by revenue is a model company, ElevenLabs, at about $600 million of annual recurring revenue (its CEO's figure) and a $22 billion valuation in a September 2026 share sale. Investors and buyers pay most for finished agents (Sierra, Decagon, Parloa) and for incumbents that own the contact center (NICE, Genesys).
- Rules: since February 2024 the FCC treats AI voices as "artificial" under the Telephone Consumer Protection Act (TCPA), so an outbound AI call without consent risks $500-1,500 per call in private suits, and TCPA class actions nearly doubled in the first half of 2025. Plaintiffs do more enforcing than regulators, with California wiretap suits against AI call vendors, a suit naming OpenAI and Twilio, and biometric suits naming ElevenLabs. After the fake-Biden robocall of January 2024, the carrier paid $1 million while the consultant behind it was acquitted and hasn't paid a $6 million FCC fine. Telling callers they're talking to AI has a measured sales cost, and Maine and the EU AI Act now require it. Platform contracts push the liability onto the customer. See How the rules work.
- Risk: OpenAI says its own cloning model needs only 15 seconds of audio, and clones have passed bank voice-ID checks (Lloyds in 2023, Santander and Halifax in 2024). An agent that authenticates callers or changes bank details is a target itself, so I wouldn't let a voiceprint be the only check (see What mistakes cost).
A voice agent is an AI agent with a clock on it, so orchestration, tool security and evals live in AI agents: orchestration platforms and I don't repeat them here. Numbers, caller ID, STIR/SHAKEN and robocall consent are in Phone numbers, messaging and voice, tracing a call across services is in Observability: metrics, logs and traces, and how a team reorganizes around agents is in Becoming AI native. Contact centers appear here only as buyers and channels; their routing and staffing deserve a guide of their own later.
The main players
These are the companies behind the diagram's parties, layer by layer, in no particular order.
Telephony and communications platforms
- What they do
- Rent numbers, carry calls to and from the phone network, sign caller ID, stream call audio to your code
- Main players
- Twilio, Telnyx, Bandwidth, Vonage (Ericsson)
- What they control
- Numbers, caller-ID attestation, the media path and a per-minute toll on every call
Voice agent platforms and open-source frameworks
- What they do
- Run the turn loop: end-of-turn detection, interruptions, tool calls, transfers, testing
- Main players
- Vapi, Retell, Bland, LiveKit, Pipecat (open source, from Daily)
- What they control
- The builder, the call logs and the turn-taking settings
Speech models
- What they do
- Speech to text, text to speech, voices, and bundled voice agent APIs
- Main players
- ElevenLabs, Deepgram, AssemblyAI, Cartesia, Gladia (OVHcloud)
- What they control
- Voices, accents, languages and the per-minute or per-character price
Language model labs
- What they do
- Text models for chained stacks, and speech-to-speech models that hear and speak in one step
- Main players
- OpenAI, Google, xAI, Meta
- What they control
- The ceiling on task success, preview terms, retirements and the price of a minute
Finished voice agents and contact center vendors
- What they do
- Sell a working agent per conversation or resolution, or add AI to the contact center
- Main players
- Sierra, Decagon, Parloa, PolyAI, NICE, Genesys
- What they control
- The enterprise buyer, what counts as a resolution, and the handoff to human agents
Voice security and fraud detection
- What they do
- Voice authentication, deepfake detection and the analytics behind spam labels
- Main players
- Pindrop, Nuance (Microsoft), Reality Defender, Hiya
- What they control
- Whether a caller is trusted and whether an outbound call gets answered
How they make money, and who's moving:
- Telephony charges per minute and per number. Twilio lists US inbound calls to a local number at $0.0085 a minute, numbers at $1.15-2.15 a month and ConversationRelay, its managed speech layer, at $0.07 a minute. Its revenue reached $1.50 billion in the quarter to June 2026, up 22%, with voice up more than 20% (earnings coverage), and two AI-native customers grew into $6 million and $9 million a year accounts within about 18 months. Bandwidth grew 22% in the same quarter, and all five of its new million-dollar deals involved AI. Telnyx sells its own voice engine at $0.05 a minute including speech. Vonage's unit inside Ericsson is shrinking.
- Voice agent platforms charge a fee per minute and pass model costs through: Vapi $0.05 with models at cost, Retell $0.055, and Bland $0.12-0.14 all-in except telephony after its December 2025 increase. Vapi raised $50 million in May 2026 at a reported $500 million valuation, a research firm estimates Retell's revenue at about $60 million a year, and LiveKit raised $100 million at $1 billion in January 2026 and says it carries ChatGPT's voice mode. Pipecat is free, with about 16,000 GitHub stars. All of them are squeezed, because model companies, telephony providers and labs now sell bundled agents too.
- Speech model companies charge per minute of audio or per thousand characters. ElevenLabs gets more than 55% of revenue from enterprises, sold employee shares at $22 billion on September 30, 2026, and cut text-to-speech prices 20% on October 5. Deepgram raised $130 million at $1.3 billion in January 2026, Cartesia about $100 million in late 2025, and I found no AssemblyAI round since December 2023. OVHcloud bought Gladia, a Paris speech-to-text company, in July 2026. All four specialists sell a bundled voice agent API at $0.06-0.08 a minute, which is what a layer does when its own unit gets cheap.
- Labs charge per audio token or per minute. OpenAI's GPT-Live-1 reached its API on September 10, 2026 at $0.05 a minute for the voice layer, with reasoning billed separately. Google lists Gemini Live audio at half a cent a minute in and under 2 cents out, but its Live API is a preview with no service level agreement. xAI's Grok Voice scored highest on the task-success part of Artificial Analysis's index in June 2026, at 52%. Labs also buy teams: Meta took PlayAI's team in July 2025 and shut its API, and Google DeepMind licensed Hume AI's technology and hired its CEO in January 2026.
- Finished agents charge per conversation or per resolution: Fin lists $0.99 per outcome, and others reportedly charge $0.50-2.50. Sierra raised $950 million at $15.8 billion in May 2026 on about $150 million of recurring revenue, Decagon was valued at $4.5 billion and Parloa at $3 billion in January 2026, and PolyAI raised $86 million in December 2025. Salesforce agreed to buy Fin for about $3.6 billion in June 2026.
- Contact center vendors sell AI on top of seats. NICE reported $362 million of AI recurring revenue, 15% of its cloud revenue, in the quarter to June 2026, after buying Cognigy for about $955 million, and Genesys reports more than $400 million. Five9's August 2026 filing says it must replace license revenue lost to AI with revenue from selling AI.
- Voice security sells per call or per agent seat, at prices that aren't public. Pindrop took $100 million of debt financing in July 2024, Reality Defender raised $33 million in October 2024, and Pindrop and Nuance both won 2026 appeals that kept voice authentication for financial firms outside Illinois's biometric law. Hiya runs analytics behind carrier spam labels and launched branded calling with Deutsche Telekom in January 2026.
As of October 2026. Most figures here are self-reported or come from press coverage of funding rounds, so treat the list as a map to check before relying on it.
Back to the payroll company. Its numbers are rented from Twilio, which carries the call and streams the audio to an agent built on LiveKit. Deepgram transcribes the caller, a language model from OpenAI or Google decides the next step, and ElevenLabs speaks the reply. A tool call checks the payroll run and finds the bank rejected the file. When the caller asks to change the company's bank account, the agent transfers the call to a person in the company's NICE contact center, and Pindrop scores the caller's voice for signs of cloning on the way. One four-minute call touches six or seven companies, each with its own meter (the vendors in the story are illustrative).
How one turn of a call works, step by step
Here's one turn of that payroll call, on a chained pipeline, from a mobile phone, using Twilio's target timings. The hardest decision often comes at the very start: deciding the caller has finished speaking, without cutting them off or leaving a long silence. Then the agent answers or calls a tool, hands off to a person when it should, and stops talking when the caller interrupts.
- Caller speaks: "My payroll didn't go out on Friday." The audio streams in across the mobile network and the carriers to the communications platform's edge, where it's buffered and decoded (about 95 ms in Twilio's diagram). Phone audio is narrowband G.711 at 8 kHz, which carries roughly 300-3,400 Hz, so a model that wants 16 kHz audio gets a resampled signal with nothing above about 3.4 kHz. Breaks: background noise, a TV or another person talking; one noise-cancellation vendor claims background voices push transcription errors from 5% to over 30%.
- Turn ended: the caller pauses, and the platform has to decide whether they're done. Voice activity detection (VAD) only says whether someone is speaking; end-of-turn detection decides they've finished. A plain silence timer waits about 500 ms by default. Turn models also weigh the words and the intonation: Deepgram claims about 260 ms for its Flux model, and Pipecat's open Smart Turn model runs in about 10 ms on some CPUs. Some stacks start the language model early on a likely end of turn and cancel if the caller keeps going, which saves hundreds of milliseconds and costs 50-70% more model calls (Deepgram's estimate). Breaks: too eager and the agent cuts off a caller reading out an account number; too patient and the line goes quiet. On τ-Voice, realistic turn-taking cut retail task success by 7 points overall and by 11 for Google's model.
- Understood: a streaming speech-to-text model sends partial transcripts and then a final one, with a target of 350 ms and an upper limit of 500. In a speech-to-speech model the audio goes to the model directly. Breaks: names, emails, addresses and strings of letters and digits fail most, and OpenAI listed "better alphanumeric recognition" as an improvement in gpt-realtime-2.1. A 2020 study of five commercial recognizers found almost twice the word error rate for Black speakers as for white speakers, and τ-Voice found accents cut task success by 10 points.
- Answered: the language model chooses a reply or calls a tool. Time to first token is the largest and most variable slice of the budget, with a target of 375 ms and an upper limit of 750. A tool call, such as looking up the payroll run, adds the tool's round trip and another model call, and the tool definitions and the history are re-sent every turn. Breaks: a slow backend leaves dead air; a speech-to-speech model promises something before the tool returns; a guardrail that checks a transcript trips after the words are spoken.
- Spoken: text to speech streams the first audio chunk before the sentence is finished. Twilio targets 100 ms; Vapi's tests in June 2026 measured 159-197 ms for the fastest voices, network included, about twice what vendors quote for the model alone. Add about 10 ms for each of roughly eight hops between services and 40-100 ms back to the caller's ear. Breaks: the 95th-percentile turn is much slower than the median, and a dashboard that leaves out end-of-turn time and the network (Vapi's own estimate does) understates what the caller hears.
Two exits sit off that path. Handed off is the move to a person. A cold transfer sends a SIP REFER through the trunk and the agent drops out; a warm transfer puts the caller on hold, briefs a human, connects them and leaves, and comes back to the caller if nobody answers. The payroll agent hands off whenever someone asks to change bank details. Transfers fail when the carrier rejects the REFER, when the caller ID shown to the human changes unexpectedly, or when a cold transfer loses the context and the caller starts over.
Interrupted is the caller talking over the agent, called barge-in. Stopping takes two jobs: flush the audio queued for playback (Twilio's clear message), then cut the agent's turn in the history down to what the caller actually heard (Twilio's mark messages report what played). Skip the second job and the model believes it said things the caller never heard. The opposite failure is a false stop on a cough, a "mm-hmm" or a TV. τ-Voice also measured interruption rates under realistic audio, 14% for OpenAI's model, 21% for Google's and 84% for xAI's, which shows how far apart the models still are on turn-taking.
Caller to the platform's edge (network, buffer, decode)
- Target
- About 95 ms
- Upper limit
- 100 ms or more
- Whose number
- Twilio's diagram
End-of-turn decision
- Target
- 260 ms (a turn model's claim) to 500 ms (a silence timer)
- Upper limit
- 800 ms or more
- Whose number
- Deepgram; Twilio
Speech to text, final transcript
- Target
- 350 ms
- Upper limit
- 500 ms
- Whose number
- Twilio
Language model, first token
- Target
- 375 ms
- Upper limit
- 750 ms
- Whose number
- Twilio
Text to speech, first audio
- Target
- 100 ms target; 159-197 ms measured
- Upper limit
- 250 ms
- Whose number
- Twilio; Vapi's tests
Hops between services
- Target
- About 10 ms each, about eight of them
- Upper limit
- Whose number
- Twilio
Total, mouth to ear
- Target
- About 1.1 s
- Upper limit
- 1.4 s
- Whose number
- Twilio
A person answering
- Target
- About 100-200 ms
- Upper limit
- About 300 ms (the slowest languages' median)
- Whose number
- Stivers and colleagues, 2009
The stages overlap, so they don't add up exactly. On a phone call, the network alone takes a real share: the international telecom standard treats one-way delay under 150 ms as unnoticeable, and mobile networks are designed to about 200 ms one way.
Chained, speech-to-speech or hybrid
A chained stack (also called cascaded, or a pipeline) runs speech to text, then a language model, then text to speech. Every hop is text you can read, log and check, and you can swap any part. OpenAI's own guidance says to choose it when you need to inspect or transform the text in the middle, for example to run a policy check before the agent speaks.
A speech-to-speech model (sold as "realtime" or "native audio") hears and speaks in one step and keeps the tone that a transcript drops. It isn't automatically faster: Artificial Analysis measured 0.44-0.82 seconds to first audio at low reasoning and 2.3-3.0 seconds at high reasoning, and τ-Voice measured 0.90-1.15 seconds under phone-like audio, the same range as Twilio's chained target. I found no independent test of the two on the same tasks and audio, which is the biggest gap in the evidence. Cost grows with call length, because OpenAI's Realtime API re-sends the whole conversation with each response. By my arithmetic from its prices, five minutes into a call an uncached session costs $0.77-1.15 a minute, against about a cent when the cache works.
A hybrid splits the fast talker from the slow thinker. OpenAI's GPT-Live-1 is full-duplex (it listens while it speaks) and hands reasoning and tool calls to a separate backend model, billed separately; OpenAI reports it cut turn-taking latency to 0.8 seconds from 1.41, a figure I saw only in press coverage. NVIDIA's open research model speaks an on-hold message while a tool runs, and LiveKit, Pipecat and Deepgram start the language model speculatively inside chained stacks.
Turn gap
- Chained
- About 1.1 s target (Twilio)
- Speech-to-speech
- 0.4-0.8 s at low reasoning, 2.3-3.0 s at high; 0.9-1.15 s on phone-like audio
- Hybrid
- 0.8 s, as reported by OpenAI
Control
- Chained
- Text at every hop; swap any part
- Speech-to-speech
- Transcripts are a side output
- Hybrid
- The backend is inspectable; the voice layer isn't
Guardrails
- Chained
- Before the agent speaks
- Speech-to-speech
- During or after speech
- Hybrid
- On the backend's output
Task success evidence
- Chained
- No τ-Voice data; text models reach 85%
- Speech-to-speech
- 26-51% on τ-Voice
- Hybrid
- 86% on a benchmark OpenAI cites, via press only
Cost shape
- Chained
- Sum of meters, roughly linear in minutes
- Speech-to-speech
- Audio tokens; grows with call length unless cached
- Hybrid
- A flat voice minute plus backend tokens when it delegates
Lock-in
- Chained
- Low
- Speech-to-speech
- High: one vendor's events and voices
- Hybrid
- High on the voice layer, low on the backend
My default: chained or hybrid for anything with a regulated script, a payment or a policy the agent must quote exactly, and speech-to-speech where tone matters more than exact words and calls are short. Whatever the architecture, I'd measure mouth to ear on real phone calls at the 95th percentile.
Which voice agent setup? Five questions
- Inbound or outbound? Outbound brings TCPA consent, caller-ID labels and disclosure, which decide whether a campaign is viable at all. The TCPA doesn't treat inbound calls as "made" by the business, so inbound comes down to whether the agent resolves calls well enough, plus recording notices and disclosure.
- What does a wrong answer cost? If the agent quotes prices, refunds or policy, the business is bound by what it says. Use a stack that can check the text before it's spoken, and read back anything the caller must get right.
- What data does it touch? Card numbers should go to keypad entry with masking or a secure link so the model never hears them; health data needs a business associate agreement (BAA) with every vendor in the chain; voiceprints bring biometric law. Each narrows which vendors and models you can use.
- How long are the calls? Long calls raise speech-to-speech costs unless caching holds, Google's Live API drops connections after about 10 minutes and has to resume them, and Vapi's default maximum call length is 600 seconds.
- Who pays when it fails? On per-minute pricing you pay for a failed six-minute call and then for the human minutes after it. On per-resolution pricing the vendor carries that, and the argument moves to what counts as resolved.
My defaults: for an inbound line like the payroll company's, I'd have the agent answer every call, verify the caller and handle the few most common requests, with a warm transfer for everything else and for any change to bank details. I'd say "AI assistant" in the greeting, announce recording and name the vendor that handles the audio, and keep card numbers on the keypad. I'd measure containment and resolution separately, with a person checking a sample of transcripts each week, and widen the agent's scope only once resolution holds. For outbound, I wouldn't dial without consent records I'd be happy to show a court.
The primitives
01
Entity & identity
What is the unit of record, and how do we know it is the same one?
Three identities meet on every call: the business the agent speaks for, the voice (synthetic, cloned or human) and the phone number. The caller's own number can be spoofed. STIR/SHAKEN attestation, when present, says how much the originating carrier vouches for it: an A means the carrier knows its customer may use the number, and says nothing about the voice or the words.
Inside the stack, one call collects IDs from the carrier, the communications platform, the voice platform, the model session and the trace. Joining them is real engineering work, and you need it before you can debug a bad call or reconcile a bill.
The law names its own entities, and one platform can be several at once:
| Legal role | Law | What it triggers |
|---|---|---|
| The "maker" or initiator of a call | TCPA | Consent, caller identification, opt-out |
| A third party listening in | California's wiretap law | $5,000 per violation if callers weren't told |
| Provider or deployer | EU AI Act | Watermarking for the provider; disclosure for both |
| Business associate | HIPAA | A BAA, and BAAs with every sub-vendor |
| A collector of voiceprints | Illinois's biometric law (BIPA) and others | Written consent and a retention policy |
A voiceprint is a biometric template that identifies a speaker. A plain transcription agent generally doesn't build one, though courts are testing that line in cases about meeting transcription tools.
02
State & lifecycle
What states exist, and what moves an entity between them?
A call carries several state machines that move at different speeds.
- Turn: listening, caller speaking, end of turn pending (when the model starts early), thinking, tool running, agent speaking, interrupted, then back to listening.
- Call: ringing, answered by a person or a machine, in progress, transferring, ended (completed, no answer, busy, failed or cancelled).
- Model session: Google's Live API adds connected, "going away" (about 60 seconds' warning) and resumed, because a connection lasts about 10 minutes and an audio session about 15.
- Consent: none, prior express, prior express written, then revoked, and a revocation has to be honoured within 10 business days.
On outbound calls, answering-machine detection takes about 4 seconds after pickup with Twilio's defaults. A model upgrade changes turn-taking behaviour as well as answers, so I'd put timing in every regression test.
03
System of record & ledger
Who owns the truth, and how do systems reconcile?
Four records describe one call, and they disagree by design.
| Fact | System of record |
|---|---|
| Minutes billed for the call | The carrier's and communications platform's call detail records |
| What the agent did, turn by turn | The voice platform's call log |
| What was said | The transcript and the recording |
| What changed in the business | The tool-call log in the business's own systems |
| Whether the caller agreed to be called | The deployer's consent records |
Two details matter. In a speech-to-speech model the transcript is a side product of a separate recognizer, so an audit or eval built on transcripts can miss what the model actually "heard". And after an interruption, what the model thinks it said and what the caller heard diverge unless the history was cut back.
Consent records are the defence in every TCPA case, and the platforms say plainly that they don't keep them for you: Retell's terms say it "does not obtain consent on behalf of Customer". Billing is its own reconciliation job, since five vendors can meter one call in carrier minutes, stream time, audio tokens, characters and voice minutes billed by the second.
04
Rules & policy
What logic decides outcomes, and who can change it?
Turn-taking settings are policy decisions: how long a silence ends a turn, how eager the model is to answer, how long a sound has to last before it counts as an interruption. A longer timeout while the caller reads out digits protects accuracy and costs speed.
Guardrails behave differently by architecture. A chained stack can check the text before it's spoken. On a speech-to-speech stream the check runs alongside the speech, and an open issue on OpenAI's agent toolkit reports the agent kept talking after a guardrail tripped. Required disclosures and consent prompts are fixed text inside a probabilistic stream, and a chained or hybrid stack makes it easier to guarantee they're said word for word.
Above the builder's settings sit laws, card network rules, carrier analytics and the platform's and model provider's usage policies. Retell's terms, for example, require outbound agents to identify themselves and ban calls to emergency lines.
05
Effective dating
Which version of the rule applied at that moment?
Dates I'd keep on a calendar as of October 2026:
| Change | Effective | Status (Oct 2026) |
|---|---|---|
| FCC: AI voices are "artificial" under the TCPA | February 8, 2024 | In force, and applied to calls made before it in pending suits |
| California: prerecorded calls must say if the voice is artificial | January 1, 2025 | In force |
| Maine: AI chatbots, spoken ones included, need a clear notice | September 2025 | In force |
| Carriers must sign caller ID with their own token | September 18, 2025 | In force |
| Carriers blocking on analytics must return a specific SIP code (603+) | March 25, 2026 | In force |
| EU AI Act: tell people they're talking to AI; mark synthetic audio | August 2, 2026 | In force |
| FCC comment deadline on letting political AI calls reach mobiles without consent | October 19, 2026 | Upcoming |
| EU: marking deadline for audio systems already on the market | December 2, 2026 | Upcoming |
| Google's Gemini Flash and Flash speech models double in price | January 1, 2027 | Upcoming |
Some dates float. Deepgram's streaming prices are "limited-time promotional rates" with no end date, so I'd budget at the regular rate. Google's Live API is a preview with no deprecation guarantees, and some OpenAI model names point at a default snapshot that can change, so I'd pin dated snapshots where offered. The FCC's 2024 proposal to require AI disclosure at the start of calls hasn't been adopted, and Canada's regulator opened its own review in June 2026.
06
Interfaces & standards
What format and protocol do counterparties speak?
The phone side is standardized: SIP for setting up and transferring calls, RTP and SRTP for the audio, G.711 at 8 kHz on the phone network, Opus in apps, E.164 numbers, STIR/SHAKEN for caller ID, and DTMF for keypad tones. The AI side has no shared standard. Every platform has its own WebSocket event protocol (Twilio's media streams, OpenAI's Realtime events, Google's Live messages, Deepgram's turn events), and nothing covers turn events, barge-in or tool calls across vendors. That's a lock-in point, and the reason open frameworks like Pipecat and LiveKit's agents exist.
WebRTC
- Typical use
- A browser or app talking to a model or platform
- Strengths
- Built for live audio: handles loss, jitter and echo; OpenAI's recommended client path
- Weaknesses
- Needs media servers and network traversal
WebSocket media stream
- Typical use
- A communications platform sending call audio to your server
- Strengths
- Simple, works everywhere
- Weaknesses
- Lost packets stall the stream; you handle buffering and barge-in
SIP trunk to a platform or model
- Typical use
- Phone numbers into LiveKit, OpenAI or a voice platform
- Strengths
- Fewer hops; native transfers; carrier-grade
- Weaknesses
- You still need a trunk provider; caller ID can change on transfer
The phone network end to end
- Typical use
- Any phone
- Strengths
- Reaches everyone
- Weaknesses
- 8 kHz audio, extra delay, spam labels and blocking
OpenAI now accepts SIP calls directly, with a separate EU endpoint, and warns builders not to retry an outbound call automatically after an ambiguous timeout, because the retry can place a second call. For provenance, Google's SynthID and Meta's AudioSeal watermark synthetic audio and C2PA labels files, but a watermark only marks the provider's own output, and phone audio degrades it.
07
Networks & counterparties
Who sits between us and the outcome, and what do they want?
A call passes through the caller's carrier, transit carriers, a communications platform or SIP trunk, the voice platform, one to three model vendors (often on different clouds) and the business's own systems. Twilio counts at least ten network traversals per turn of a chained agent, and each boundary can add another encode, decode and buffer. Moving the pieces closer together is mostly about latency: Telnyx places GPUs next to its telephony sites, and speech-to-speech models collapse three vendors into one.
| Party | What they control | What they earn |
|---|---|---|
| Communications platform or carrier | Numbers, caller-ID attestation, the audio stream | Per-minute and per-number fees, plus carrier surcharges |
| Terminating carriers and analytics engines (Hiya, TNS, First Orion) | Spam labels and blocking | Branded-calling services sold to businesses |
| Voice platform | The turn loop, logs, transfers | A per-minute platform fee |
| Model vendors | Accuracy, voices, latency, retirements | Per token, minute or character |
The analytics engines decide whether an outbound agent is heard at all. An outbound agent dials many short calls from few numbers, which is the pattern that earns a "Spam Likely" label, and a caller-ID company's 2026 survey found 86% of consumers say they won't answer an unidentified call (a vendor survey). Registration and branded calling are in Phone numbers, messaging and voice.
08
Regulatory layering
Jurisdiction × activity × entity type: is it a license or a certification?
| Call | Layers that apply |
|---|---|
| Outbound, US | TCPA and the FCC's rules, state telemarketing laws, state AI-disclosure laws, state recording laws, carrier policy, the platform's terms |
| Inbound, US | State AI-disclosure laws (Maine, Utah), recording and wiretap laws, biometric laws if voiceprints are used |
| Any call in the EU | AI Act Article 50, GDPR, national telemarketing and recording rules |
| Health or card data, anywhere | HIPAA and BAAs; PCI DSS through the card networks |
| Canada | The CRTC's rules on automated calls (which already cover synthesized voices), privacy law |
One agent usually serves callers in every state, so in practice the strictest layer wins: an all-party recording notice on every call, AI disclosure in every greeting.
09
Exceptions & reversals
What goes wrong, and how is it undone?
Most reversals on a call are small and happen every minute: a speculative reply is cancelled when the caller keeps talking, the agent's turn is cut back after an interruption, a failed transfer comes back to the agent. The dangerous ones are side effects. A tool call cancelled mid-flight because the caller changed their mind may already have done something, and I couldn't find how vendors handle that. An outbound dial retried after an ambiguous timeout can ring the same person twice, a second TCPA exposure.
Caller revokes consent
- Who starts it
- The caller, by any reasonable means
- Clock
- Honour within 10 business days
- The way back
- Suppress the number everywhere
Transfer fails
- Who starts it
- Trunk, carrier or an unanswered human
- Clock
- Seconds
- The way back
- Return to the agent; offer a callback
Model or price change
- Who starts it
- The vendor
- Clock
- Days to months; previews without notice
- The way back
- Pinned snapshots; regression tests with timing
Vendor shuts down after an acquisition
- Who starts it
- The acquirer
- Clock
- Weeks: PlayHT's API went dark in July 2025
- The way back
- Re-integrate; cloned voices may be lost
Deployment pulled back
- Who starts it
- The business
- Clock
- After the damage
- The way back
- Rehire, apologize, rethink scope
Commercial terms reverse too. Bland raised its per-minute price by 22-56% in December 2025, and ElevenLabs said in February 2025 that it was absorbing language model costs and would eventually pass them on.
10
Liability allocation
When it fails, who pays?
The deployer owns what its agent says. When Air Canada's chatbot invented a bereavement-fare rule in 2024, the tribunal rejected the airline's argument that the bot was a separate legal entity and made it pay about C$812, and the same principle applies to a voice agent that quotes a refund on a recorded call.
Platform contracts push the rest down to the customer:
| Clause | Retell (June 2026) | Vapi (September 2026) |
|---|---|---|
| Who gets consent | The customer only | The customer |
| AI disclosure | Required at the start of outbound calls | Not explicit |
| Customer indemnity | Covers regulatory fines, the TCPA included; uncapped | Covers TCPA claims |
| Vendor's cap | 12 months of fees | The greater of $100 or 12 months of fees |
| Accuracy | Not warranted; the customer assumes the risk | All warranties disclaimed |
Plaintiffs are testing whether that holds. A December 2025 suit claims OpenAI and Twilio are liable as makers of AI robocalls that another company sent through them, two California rulings in 2025 let wiretap claims proceed against AI vendors able to use call audio for their own models, and nine biometric suits filed in May 2026 accuse ElevenLabs and big tech companies of building voiceprints from public recordings to train models. An indemnity doesn't stop a plaintiff from suing the vendor directly; it only moves the bill, and only if the customer can pay.
Carriers carry their own share through attestation: Lingo Telecom paid $1 million for signing spoofed AI calls. Fraud losses depend on the payment rail and the country. In the US the victim of a scam that tricks them into sending money usually bears it. Pindrop reportedly offers a deepfake warranty of up to $1 million a claim, the only explicit transfer of this risk from a security vendor I found.
What's different here
How the money moves
The business pays everyone, usually per minute, and the meters start at different moments. Here's the per-minute stack at US list prices in October 2026.
Streaming speech to text
- Low
- $0.0025
- Typical
- $0.005-0.008
- High
- $0.017
Text to speech
- Low
- About $0.01
- Typical
- $0.015-0.03
- High
- $0.08-0.10 (premium voices)
Language model, chained
- Low
- $0.002-0.003
- Typical
- $0.008-0.064
- High
- $0.32-0.64 (frontier, fast tier)
Speech-to-speech model
- Low
- $0.023 (Gemini Live, both directions)
- Typical
- $0.05 (GPT-Live-1 voice layer) plus backend
- High
- $0.35-0.38 (OpenAI's realtime model resold by Retell)
Telephony, per leg
- Low
- $0.003-0.004 (SIP)
- Typical
- $0.0085-0.015
- High
- $0.022 (toll-free inbound)
Platform fee
- Low
- $0.05
- Typical
- $0.055-0.08
- High
- $0.12-0.14 (all-in)
All-in, per call minute
- Low
- About $0.06
- Typical
- About $0.10-0.15
- High
- About $0.31-0.45
In a mid-priced chained stack at about $0.12 a minute, the platform takes about 45%, the language model and the voice 10-30% each, telephony 5-12% and speech to text about 5%. Model prices fell while platform fees seem to have held at about $0.05: OpenAI cut its realtime audio prices 60% in December 2024 and another 20% in August 2025, and ElevenLabs halved its agent prices in February 2025.
The traps sit in the billing units. AssemblyAI bills streaming by session time, silence included, and GPT-Live-1 bills the whole session, backend thinking included. Telnyx rounds each call up to the minute, so wrong numbers cost a full minute. Speech-to-speech models re-bill the conversation each turn, and changing instructions mid-call breaks the cache. And plans cap simultaneous calls, with extra lines at about $10 a month each at Vapi.
A month of calls on the payroll support line
My assumptions: 100,000 inbound calls a month at 4 minutes of talk each; a person needs 5 minutes per call including after-call work; calls the agent can't finish spend 1.5 AI minutes before a transfer and 4.5 human minutes after; the AI costs $0.06, $0.12 or $0.30 a minute; and fixed AI costs (platform minimums, one to three people designing and checking conversations, evals and integration) come to $15,000-50,000 a month. The human ranges are my arithmetic from the US median wage for customer service representatives ($21.53 an hour in May 2025) and outsourcing guides' loaded rates.
| Setup | With US in-house staff | With offshore staff |
|---|---|---|
| All human (500,000 handled minutes) | $295,000-720,000 | $80,000-200,000 |
| Agent answers first, 30% contained | $215,000-571,000 (about 21-27% less) | $79,000-243,000 (1% less to 22% more) |
| Agent answers first, 60% contained | $139,000-399,000 (about 45-53% less) | $62,000-212,000 (22% less to 6% more) |
| Agent verifies and routes; people take every call | $252,000-636,000 (about 12-15% less) | $80,000-220,000 (up to 10% more) |
Against US staff, the AI usage bill is 4-12% of the old labor bill and containment sets the savings. Each extra point of containment, 1,000 calls, is worth about $2,500-5,700 a month, so moving from 30% to 60% saves about $75,000-172,000 a month, while halving the AI price from $0.12 to $0.06 saves about $13,500-18,000. Against offshore staff at 30% containment, the program can cost more than it saves once fixed costs count. The last row is how contact center incumbents mostly sell AI: lower risk, smaller savings.
Priced per outcome instead, 60,000 resolutions a month would cost about $59,000 at Fin's $0.99 and $90,000-150,000 at the $1.50-2.50 reported for other vendors, against $14,000-72,000 of raw AI usage for the same calls. The vendor keeps that spread for carrying the failures and the model costs. I think buyers underrate the outcome definition: Fin bills handoffs that follow a procedure and self-serve routings as outcomes too, and a caller who hangs up angry still counts as contained.
Who holds the power
Power follows whoever owns the customer relationship, the model and the phone line.
- Contact center incumbents (NICE, Genesys, Five9) own routing, the human agents' desktops and the handoff, and their AI revenue already exceeds most startups' total revenue. Their risk is that every contained call is a seat they no longer sell, as Five9's filing says.
- Finished-agent vendors (Sierra, Decagon, Parloa, PolyAI, Fin) own the outcome definition. Investors value them highest, at roughly 45-105 times recurring revenue by my rough arithmetic, and CRMs are buying in: Salesforce is buying Fin. Sierra also wrote τ-Voice, the benchmark the industry cites, which I'd keep in mind when reading it.
- Labs set the ceiling on task success and the price of the premium tier, and they're moving up the stack with full-duplex voice, direct phone connections and talent deals. They earn through every layer whoever wins the customer.
- Speech model companies face commodity pricing on the basic tier and still earn a premium for expressive voices: on Retell, ElevenLabs voices cost 2.7-6.7 times the commodity set.
- Communications platforms and carriers hold numbers, porting, attestation and the audio path. Telephony is only 5-12% of a mid-priced stack, but it's sticky, and Twilio, Telnyx and Bandwidth all sell their own agent layers.
- Plaintiffs' lawyers and the carriers' analytics engines hold the most practical power over outbound calls. The lawyers price non-compliance and the analytics engines decide whether calls get answered, and both matter more as the FCC loosens its TCPA rules.
- Developer platforms (Vapi, Retell, Bland) own the builder and are squeezed from below by model companies, from the side by telephony providers and from above by labs and contact center suites. Vapi's answer is enterprise work; Bland's is all-in pricing.
How the rules work
The FCC's ruling. On February 8, 2024 the FCC ruled unanimously that AI-generated and cloned voices are "artificial" under the TCPA. Outbound AI calls need prior express consent, written and naming the seller for telemarketing; they must identify the caller at the start and, for telemarketing, offer an automated opt-out. The FCC refused any carve-out for technology that claims to work like a live agent, and the 26 state attorneys general who asked for the ruling can sue under the TCPA too. Private damages are $500 per call, $1,500 if wilful, with a four-year limitation period. Since a June 2025 Supreme Court decision, courts read the TCPA for themselves without deferring to the FCC; an appeals court had already read "artificial voice" to include AI in 2023, so I'd treat the question as settled.
The FCC's direction has turned since. Its August 2024 proposal to require AI disclosure at consent and at the start of each call hasn't been adopted, the one-to-one consent rule was deleted in 2025, and in September 2026 it asked for comment on letting political AI-voice calls reach mobile phones without consent.
Plaintiffs do the enforcing. A litigation tracker counted 1,052 TCPA class actions from January to June 2025, against 539 a year earlier. The AI-voice suits I found turn on consent, ignoring "no", wrong numbers and state registration rather than model quality: a February 2026 suit against a mortgage lender alleges more than $5 million in damages, and a 2026 Texas suit says an AI agent kept selling after the person said no. The suit against OpenAI and Twilio tests whether model and communications providers can count as callers, and I found no ruling yet.
Fake-Biden robocall. On January 21, 2024, calls with a cloned Biden voice told New Hampshire Democrats not to vote in the primary. A political consultant commissioned them, paying a magician $150 to make the clip in under 20 minutes. Lingo Telecom, which signed the spoofed calls with the highest level of caller-ID attestation, paid $1 million under an August 2024 consent decree. The consultant was fined $6 million by the FCC, still refused to pay as of November 2025, and was acquitted of all 22 criminal counts in June 2025. The penalty that landed hit the carrier's attestation.
Disclosure. Maine's law, in force since September 2025, covers chatbots that talk as well as type and requires a clear notice wherever a reasonable consumer could think they're talking to a person. Utah requires disclosure when asked, and up front in high-risk interactions, until July 2027, and California requires prerecorded calls to say if the voice is artificial. Since August 2, 2026 the EU AI Act has required that people are told they're dealing with AI unless it's obvious, and that providers mark synthetic audio, with fines up to €15 million or 3% of turnover. Disclosure has a measured cost: in a 2019 field experiment with more than 6,200 customers of a financial firm, telling people up front that a sales call came from a bot cut purchases by more than 79.7%. That was pre-LLM technology in one market, and I found no newer replication.
Recording and wiretaps. About 11 states require every party's consent to record. The sharper risk is California's wiretap law. In February 2025 a court let a suit proceed against Google over its contact center AI on Verizon support calls, holding that a vendor capable of using call data for its own purposes counts as an eavesdropper whether or not it actually does, and an August 2025 ruling did the same for an AI pizza-ordering vendor. At $5,000 per violation, zero-retention and no-training settings and a greeting that names the vendor are worth their price.
Biometrics, health and cards. BIPA pays $1,000-5,000 per person for voiceprints collected without written consent; voice-authentication vendors serving financial firms won two appeals in 2026 under its exemption for financial institutions, which doesn't cover a vendor serving a clinic or a retailer. Under HIPAA a voice platform handling patient data is a business associate, and compliance becomes a price tier: Vapi charges $2,000 a month for HIPAA and $1,000 for zero data retention, and ElevenLabs signs BAAs only on its enterprise tier. Under PCI DSS, a card number spoken to the agent pulls the transcription, the model's context, logs, recordings and eval sets into scope, which is why card capture moves to the keypad.
Canada. The CRTC's rules on automated calls already cover "synthesized" voices and require express consent for telemarketing. A review opened in June 2026 asks whether to require AI disclosure and whether the rules should reach a consumer's own AI assistant making a booking.
TCPA and the FCC's 2024 ruling
- Outbound
- Yes
- Inbound
- No
- Penalty
- $500-1,500 per call
- Status (Oct 2026)
- In force
Maine's chatbot disclosure law
- Outbound
- Yes
- Inbound
- Yes
- Penalty
- Unfair trade practice
- Status (Oct 2026)
- In force since September 2025
California recording and wiretap law
- Outbound
- Yes
- Inbound
- Yes
- Penalty
- $5,000 per violation
- Status (Oct 2026)
- In force; vendor suits proceeding
BIPA and similar laws
- Outbound
- If voiceprints
- Inbound
- If voiceprints
- Penalty
- $1,000-5,000 per person
- Status (Oct 2026)
- In force
EU AI Act Article 50
- Outbound
- Yes
- Inbound
- Yes
- Penalty
- Up to €15 million or 3%
- Status (Oct 2026)
- In force since August 2, 2026
HIPAA; PCI DSS
- Outbound
- If health or card data
- Inbound
- If health or card data
- Penalty
- Regulator; card network fines
- Status (Oct 2026)
- In force
Canada's automated-call rules
- Outbound
- Yes
- Inbound
- No
- Penalty
- CRTC penalties
- Status (Oct 2026)
- Under review
What mistakes cost
Let's say a mortgage lender runs an outbound AI campaign to 100,000 leads. About 30-40% connect, for 2.5 minutes on average, so the AI minutes cost about $7,500-30,000 at $0.10-0.30 a minute. If the consent behind 5% of those dials is defective, that's 5,000 calls at $500 each, a $2.5 million statutory floor, or $7.5 million if a court finds it wilful, before defence costs and state claims. One bad source of leads can cost 100 to 1,000 times the AI minutes, and the platform's terms make sure the lender carries it.
| Case | What happened | Who paid |
|---|---|---|
| Fake-Biden robocall, 2024 | Cloned voice, spoofed numbers, A-level attestation | The carrier, $1 million; the consultant's $6 million fine unpaid |
| Commonwealth Bank of Australia, 2025 | Cut 45 roles, claiming its voice bot reduced calls by 2,000 a week; the union showed volumes were rising | The bank reversed the cuts and apologized |
| McDonald's and IBM, 2024 | AI drive-thru ordering in about 100 restaurants; accuracy estimated in the low-to-mid 80s against a 95% target | The test ended in July 2024 |
| UK energy firm, 2019 | A cloned parent-company CEO's voice asked for a transfer | About €220,000, covered by insurance |
| UAE bank, 2020 | A cloned company director's voice plus forged emails | About $35 million |
| Arup, 2024 | A video call where every other participant was a deepfake | About $25 million |
The fraud numbers keep growing. The FBI's crime complaint center counted $893 million of losses in 2025 complaints that cited AI, out of $20.9 billion in total, with no split for voice. And the bank cases show voice ID itself under pressure: a journalist used a cloned voice to get into a Lloyds account in 2023, the BBC passed Santander's and Halifax's voice ID with a clone in 2024, and OpenAI's chief executive told a Federal Reserve conference in July 2025 that it was "crazy" for institutions still to accept voiceprints.
Voice agents vs human agents: what's real
My view as of October 2026: a voice agent is cheap and fast enough for most inbound calls, and still fails too many tasks to replace the people behind it. The independent evidence on task success is recent and sobering, deployment numbers are almost all vendor-reported, and the public pull-backs came from quality and measurement more than from regulators.
What the benchmarks say
- Task success: τ-Voice (278 tasks in airline, retail and telecom service, published in March 2026 and accepted at ICML) tested speech-to-speech agents on clean audio and on realistic audio with noise, 8 kHz phone quality, dropped packets and varied accents. The best completed about 30-50% of tasks on clean audio and 26-38% on realistic audio, against 85% for a text model. No chained stack was tested.
- Who did best: in the paper, xAI's and OpenAI's agents led (51% and 49% on clean audio) and Google's trailed (31%); in Artificial Analysis's June 2026 index, xAI's Grok scored 52% on the same benchmark and OpenAI's top model at high reasoning 40%. Rankings move with each model version, so I'd test the current ones myself.
- The vendors' own claims: OpenAI reports GPT-Live-1 at 86% on a benchmark it calls "Tau3 Voice", up from 46% for its previous model, a figure I saw only in press coverage.
What deployments report
- Utilities: PG&E's line, run by PolyAI, contains 67% of calls, which PolyAI's own page says is 6 points above the old phone menu, and saves 35,000 agent hours.
- Energy: Parloa says it resolves up to half of E.ON's inbound calls.
- Customers: Gartner's survey of 5,728 customers found only 14% of service issues fully resolved through self-service, and 64% would prefer companies not use AI for service. Qualtrics found AI customer service fails at almost four times the rate of other AI uses, though neither survey is about voice alone.
Every named deployment figure is the vendor's own, none is audited, and containment, resolution and routing accuracy are different metrics. PG&E's case is the one I'd show a skeptic, because it says what the old phone menu already achieved.
Who pulled back
- McDonald's ended its IBM drive-thru test in July 2024 after accuracy estimated in the low-to-mid 80s, with accents a reported problem.
- Taco Bell said in August 2025 that it was rethinking where to use voice AI in its drive-thrus, after prank orders and trouble at peak times.
- Commonwealth Bank of Australia reversed 45 job cuts in August 2025 when the union showed call volumes were rising.
- Klarna, in chat, said in May 2025 that its cost focus had produced "lower quality" and began hiring people again.
Gartner predicted in June 2025 that half the companies planning big service headcount cuts because of AI would abandon those plans by 2027. In January 2026 it predicted, in a release I could see only by its title and press coverage, that generative AI's cost per resolution will exceed offshore human agents' by 2030. How a support organization changes around agents, beyond shrinking, is in Becoming AI native.
What containment is worth
The arithmetic in How the money moves makes reliability the bigger lever by roughly five to ten times. Containment also flatters: a caller who hangs up angry counts as contained, and the FTC sued a company in August 2025 for claiming its AI phone agent could replace sales staff. Disclosure cuts the other way on outbound sales, with the 2019 experiment's drop of more than 79.7% in purchases once callers knew. For inbound service, where Maine and the EU now require disclosure, I haven't seen data on its cost.
Questions to ask a vendor
- Of the calls your agent handled for customers like us last quarter, what share was resolved, as opposed to merely not transferred, and who checked?
- What's the mouth-to-ear turn gap at the 95th percentile on real phone calls, end-of-turn time included?
- How do you test with noise, accents, 8 kHz audio and callers who interrupt, and will you run our own scenarios?
- When the agent hands off, what does the human see, and what happens if nobody answers?
- Do you or your model providers retain or train on our call audio, and can we turn that off in writing?
What usually goes wrong
| Symptom | Likely cause | First thing to check |
|---|---|---|
| Callers say the agent is slow | Long model response time, tool round trips, high reasoning effort, hops across clouds | Mouth-to-ear time at the 95th percentile, by stage |
| The agent talks over callers | Silence timer too short; slow speakers; callers reading digits | End-of-turn settings during data capture; keypad entry |
| The agent stops for no reason | Backchannels, coughs, a TV, its own voice echoing | False-interruption rate; interruption classifier |
| Wrong names, emails or account numbers | 8 kHz audio, accents, noise | Read-back confirmation; spelling and keypad fallbacks |
| Calls end mid-conversation at 10 minutes | Platform maximum duration or a model connection limit | Max-duration settings; session resumption |
| Outbound calls go unanswered | "Spam Likely" labels, weak attestation, number reputation | Registration with analytics engines; your own signing token; SIP 603+ responses |
| A customer got two calls | An ambiguous timeout retried automatically | Dial logic and retry rules |
| The bill grows with call length | Speech-to-speech context replay; cache broken by mid-call changes | Cached share of input; static instructions and tools |
| A lawyer's letter about recording | The vendor can train on audio; no notice naming it | Retention and training settings; the greeting |
| A cloned voice gets through | Voiceprint as the only check | A second factor for money movement and account changes |
Words that mean something else here
| Term | What you'd assume | What it means here |
|---|---|---|
| Latency | One number | At least six: model inference, first audio, first token, end of turn, the platform's turn gap and mouth to ear. Ask which, and at which percentile |
| Realtime | Instant | OpenAI's product name, any streaming speech-to-speech model, or loosely anything under a second |
| Interruption | The caller cutting in | Either direction: the caller interrupting the agent (wanted) or the agent cutting off the caller (not wanted); an "interrupt rate" can mean either |
| Handoff | Passing to a person | Usually a transfer, by SIP; in agent frameworks, a switch to another AI agent |
| Agent | The AI | An AI agent, a human contact center agent, or in SIP any endpoint. In "transfer to an agent", it's the human |
| Minute | Sixty seconds of a call | Carrier minutes, stream time with silence, audio tokens, or minutes of generated speech; "$0.05 a minute" can cover very different scope |
| Containment | Resolution | Calls not transferred to a person, abandoned calls included |
| Artificial voice | Robotic-sounding | Under the TCPA, any AI-generated or cloned voice, a live conversational agent included |
| Attestation A | The call is legitimate | The carrier vouches for the customer's right to the number, not the voice or the content |
| Transcript | What the model read | In a chained stack, yes; in speech-to-speech, a separate recognizer's guess at what was said |
What surprised me
Placeholders in your voice, drafted from the research and the earlier guides. Rewrite each with your own moment.
The rate deck. In telco I could multiply minutes by a price per minute. Here one four-minute call is metered five ways by five vendors, and on a speech-to-speech model the fifth minute costs more than the first.
"Faster means better." The fastest agents in the benchmark still got about half the tasks wrong on phone audio. Speed got them into the conversation and logic decided the outcome.
"Inbound is the safe side." The TCPA barely touches calls the business answers, and then a California court treated the AI vendor on the line as an eavesdropper.
Caller ID. In the telco guide, attestation was a technical detail. In the Biden robocall it was the one penalty that got paid.
"The cheap part is the AI." The AI minute was a rounding error next to the human minutes that follow every call it can't finish.
Sources
Undated entries were read on October 8, 2026; "search result" means seen only as a search snippet or summary. Company figures are self-reported unless they come from a filing or a regulator, and vendor benchmarks measure the vendor's own products or customers.
Regulators, laws and courts
- FCC: declaratory ruling 24-17 on AI voices (Feb 2024); notice of proposed rulemaking 24-84 (Aug 2024, search result); public notice DA 26-940 on political AI calls (Sep 2026); Lingo Telecom settlement (Aug 2024, search result)
- Fake-Biden robocall: Cyberscoop, Lingo agrees to $1 million (Aug 2024, search result); NBC Boston, the unpaid fine (Nov 2025); NHPR/AP, acquittal (Jun 2025, search result); NBC News, how the clip was made (Feb 2024, search result)
- TCPA courts and litigation: Steptoe, McLaughlin v. McKesson (Jun 2025, search result); TCPAWorld, class actions up 95% (Jul 2025, search result) and the suit naming OpenAI and Twilio (Dec 2025, search result); Inman, suit against a mortgage lender (Feb 2026, search result); Sheppard Mullin, AI robocalls and Texas (Jul 2026, search result)
- Consent rules (via the telco research): Federal Register, revocation rules (Oct 2024); Goodwin, one-to-one consent eliminated (Sep 2025); TransNexus, third-party signing rule (Aug 2025); TLP, SIP 603+ and blocking (Mar 2025); Hiya, FreeCallerRegistry
- State laws: Maine, 10 MRSA §1500-DD; California, Public Utilities Code §2874 (search result); Future of Privacy Forum, Utah's AI law (search result); Rev, recording laws by state (vendor, search result)
- California wiretap suits: Ambriz v. Google order (Feb 2025, search result) and case tracker (Sep 2026); Duane Morris, Taylor v. ConverseNow (Aug 2025, search result)
- Biometrics: Biometric Update, Pindrop exempt from BIPA (Aug 2026); National Law Review, voiceprints, vendors and the exemption (Sep 2026, search result); Capitol News Illinois, suits over training voices (May 2026, search result)
- EU AI Act: Jones Walker, August 2 still matters (Jul 2026)
- Health and cards: ElevenLabs, HIPAA (vendor, search result); Vapi support, HIPAA compliance (vendor, search result); PCI Security Standards Council, protecting telephone-based payment card data (Nov 2018, search result)
- Canada: CRTC, consultation 2026-132 (Jun 2026, search result); M3AAWG, comments on the consultation (Jul 2026)
- FTC: DLA Piper, the Air AI case (Aug 2025, search result)
- Liability: McCarthy Tétrault, Moffatt v. Air Canada (2024)
Labs and speech model companies: documentation and pricing
- OpenAI: gpt-realtime, gpt-realtime-2.1, GPT-Live-1, voice activity detection, realtime costs, SIP, WebRTC, voice agents, pricing, synthetic voices (Mar 2024, search result); agents SDK issue on guardrails (search result); AI Weekly, GPT-Live-1 in the API (Sep 2026, search result); DataNorth, gpt-realtime-2.1 (Jul 2026, search result); OpenAI community, December 2024 price cut (search result)
- Google: Gemini API pricing; Firebase, Live API limits (Oct 2026); TechCrunch, Google and Hume AI (Jan 2026, search result)
- Meta: Bloomberg, Meta acquires PlayAI (Jul 2025, search result); AudioStack, PlayHT deprecation notice (Jul 2025, search result)
- ElevenLabs: latency, API pricing, $22 billion tender (Sep 2026), price cut for agents (Feb 2025); AI Weekly, ARR and tender (Sep 2026, search result); usagepricing.com, ElevenLabs price tracker (Oct 2026)
- Deepgram: pricing, Flux; TechCrunch, Series C (Jan 2026, search result)
- AssemblyAI: pricing, Universal-Streaming (search result)
- Cartesia: pricing; Sacra, Cartesia (search result)
- Gladia: OVHcloud, Gladia acquisition (2026, search result)
Telephony and voice agent platforms
- Twilio: latency guide for AI voice agents (Nov 2025), Media Streams messages, US voice pricing (Aug 2026), ConversationRelay pricing (search result), answering machine detection (search result), Q2 2026 results (Aug 2026); Futurum, Twilio Q2 2026 (Aug 2026, search result); CMSWire, AI customers spending more (2026, search result)
- Telnyx: voice AI pricing, voice AI (search result)
- Bandwidth: Q2 2026 results (Jul 2026); MarketBeat, Q2 call highlights (Jul 2026, search result)
- Vonage: Light Reading, Ericsson's Vonage revival plan (Aug 2026)
- Vapi: pricing, terms of service (Sep 2026), understanding latency, call timeout settings (search result), Humanness Index (Jun 2026, vendor); TechCrunch, Vapi at $500 million (May 2026, search result)
- Retell: pricing, terms of service (Jun 2026); Sacra, Retell revenue estimate (Apr 2026)
- Bland: pricing, Series B (2025, search result); Lindy, Bland's pricing change (search result)
- LiveKit: Fortune, Series C (Jan 2026, search result); transfers documentation (search result)
- Pipecat: Smart Turn (search result) and its documentation (search result); GitHub star counts read through the GitHub API (Oct 2026)
- Noise and turn detection vendor: BusinessWire, noise and voice isolation launch (Jul 2025, vendor, search result)
Finished agents, contact centers and deals
- Sierra: Let's Data Science, Series E (May 2026)
- Decagon: Bloomberg, valued at $4.5 billion (Jan 2026, search result)
- Parloa: TechCrunch, $3 billion valuation (Jan 2026, search result); IT Brief, $50 million ARR (Jun 2026, search result); Parloa, customer examples (vendor, search result)
- PolyAI: Series D (Dec 2025, search result); PG&E case study (vendor)
- Fin and Salesforce: Fin pricing; The Next Web, Salesforce acquires Fin (Jun 2026, search result)
- NICE: Q2 2026 results (Aug 2026); Call Centre Helper, Cognigy acquisition completed (Sep 2025, search result)
- Genesys: Intelligent CIO, Q2 FY2027 (Sep 2026, search result)
- Five9: Q2 2026 results, 8-K exhibit 99.1 (Aug 2026)
Voice security and fraud
- Pindrop: Latham & Watkins, growth financing (Jul 2024, search result); CIO Influence, 2025 voice intelligence report (Jun 2025, vendor, search result)
- Reality Defender: FinTech Global, $33 million (Oct 2024, search result)
- Hiya: State of the Call 2026 (Mar 2026, vendor, search result); BusinessWire, branded calling with Deutsche Telekom (Jan 2026, vendor, search result)
- Bank voice ID: IT Brew, voice AI beats a bank's voice ID (Mar 2023, search result); Local10/AP, OpenAI's CEO on voice fraud (Jul 2025, search result)
- Fraud cases: CFO Dive, Arup (2024, search result); Sophos, the 2019 CEO voice case (search result); Senseient, the UAE bank case (search result); SecureWorld, FBI IC3 2025 AI losses (Apr 2026, search result)
Benchmarks, research and outcomes
- Ray and colleagues, τ-Voice (Mar 2026); Artificial Analysis, Speech to Speech Index (Jun 2026); NVIDIA, VoiceChat (2026, search result)
- Stivers and colleagues via Language Log, turn-taking across languages (2009 paper, search result); ITU-T, G.114 (2003, search result); Koenecke and colleagues, racial disparities in speech recognition (2020, search result)
- Luo and colleagues, disclosing chatbots in sales calls (2019, search result)
- Gartner: 14% resolved in self-service (Aug 2024, search result); 64% prefer no AI (Jul 2024, search result); half will abandon headcount cuts (Jun 2025, search result); cost per resolution by 2030 (Jan 2026, title only)
- Qualtrics, AI customer service fails at four times the rate (Oct 2025, search result)
- Pull-backs: ABC News/AP, McDonald's ends its AI drive-thru test (Jun 2024, search result); TechSpot, Taco Bell slows down (Aug 2025, search result); Bloomberg, Commonwealth Bank reverses job cuts (Aug 2025, search result); Fortune, Klarna hires humans again (May 2025, search result)
- Human cost: US Bureau of Labor Statistics, customer service representatives (May 2025 data); Outsource Accelerator, cost per call (search result)
Field Guides are learning notes, not legal or compliance advice. Rules and fees change; check the cited primary sources before you act on anything here.