All posts AI Call Answering

How to Build an AI Voice Agent with Twilio: Architecture, Costs and Pitfalls

I
iSolution Technologies
Sep 28, 2026 · 9 min read
How to Build an AI Voice Agent with Twilio: Architecture, Costs and Pitfalls

Photo: Christina Morillo on Pexels

ConversationRelay against Media Streams against a speech-to-speech model: how each call flows, the latency budget, the cost per minute line by line, and the pitfalls that take a demo four to eight weeks to become production.

There are two ways to build an AI voice agent on Twilio. ConversationRelay costs $0.07 a minute on top of the call and hands your server text over a WebSocket, with Twilio doing the speech recognition, the voice and the interruptions. Media Streams costs $0.004 a minute and hands you raw 8 kHz audio, and you run the speech stack yourself. Either way the raw vendor cost lands between $0.05 and $0.12 a minute, against the $0.12 to $0.25 most hosted platforms charge. A demo takes a weekend. A production agent takes four to eight weeks, and nearly all of that time goes on the pitfalls at the end of this post.

The three architectures

ConversationRelayMedia Streams, own pipelineSpeech-to-speech model
What your server receivesTranscribed textBase64 mu-law audio framesEvents from the model
Speech to textTwilio (Google or Deepgram)Yours, e.g. Deepgram Nova-3Inside the model
Text to speechTwilio (Google, Amazon, ElevenLabs)Yours, e.g. ElevenLabs Flash or CartesiaInside the model
Interruption handlingBuilt in, configurableYou write itBuilt in
Raw cost per minute$0.08 to $0.11$0.05 to $0.09$0.07 to $0.12 cached
Build effortLowestHighestMiddle

The cost rows are vendor list prices from September 2026 plus our estimate for language-model tokens. The breakdown is further down.

A sound waveform displayed on a dark screen
Photo by Egor Komarov on Pexels

How a ConversationRelay call flows

An inbound call hits your voice webhook. You answer with TwiML: a <Connect> verb containing <ConversationRelay> and the wss:// URL of your server. Twilio opens the WebSocket and sends a setup message, then a prompt message every time the caller says something. Each prompt carries the transcribed text and a last flag that tells you the caller has finished.

You send the text to your language model and stream the reply back as text messages, one token at a time, marking the final one last: true. Twilio's docs are explicit on two points: stream tokens as they arrive instead of waiting for the full reply, and don't trim them, because the spaces between tokens matter to the voice engine.

When the caller talks over the agent, you get an interrupt message with the words the caller actually heard before they cut in and the milliseconds played. When you want to hand the call back to Twilio (to transfer, take a payment, hang up), you send an end message with a handoffData string, and the call continues at the action URL on your <Connect> verb.

How a Media Streams call flows

Same webhook, different TwiML: <Connect><Stream> starts a bidirectional stream. Twilio sends media messages containing base64 audio encoded as audio/x-mulaw at 8000 Hz. You pipe that into your speech-to-text engine, run the model, synthesise the reply, convert it to the same format with no file header bytes, and send it back as media messages. Twilio buffers them and plays them in order.

Your server can send three message types: media, mark and clear. A mark tells you when a chunk of audio has finished playing. A clear empties the buffer, which is how you stop the agent mid-sentence. You get one bidirectional stream per call, and you must validate the X-Twilio-Signature header or anyone can connect to your socket.

The speech-to-speech option

Models such as OpenAI's gpt-realtime take audio in and produce audio out, skipping the separate transcription and voice steps. You connect through Media Streams or through OpenAI's SIP connector with a Twilio SIP trunk. Billing is per audio token: about 600 tokens for a minute of caller speech and 1,200 for a minute of agent speech, at $32 and $64 per million. With prompt caching working, measured sessions run $0.06 to $0.11 a minute. Without it, long calls reach $0.18 to $0.46, so caching is the first thing to verify in production.

The latency budget

Below about 800 milliseconds from the caller finishing to the agent starting, a call feels smooth. From 800 to 1,200 it's acceptable for business. Past 1,500 there's a pause the caller notices. Where the time goes:

  • End-of-speech detection. The wait to be sure the caller has stopped. ConversationRelay's speechTimeout accepts 600 to 5,000 milliseconds, so this is the largest single item and the first to tune.
  • Model first token. 200 to 500 milliseconds on a fast text model, longer with a large prompt or a tool call in the way.
  • Voice first audio. ElevenLabs quotes about 75 milliseconds for Flash v2.5.
  • Network. 50 to 100 milliseconds if your server sits in the same region as Twilio's media edge. Much more if it doesn't.

Streaming tokens is what makes the sum work. The voice starts on the first sentence while the model is still writing the third.

The cost stack per minute

ItemPrice
Twilio inbound voice, US local number$0.0085 a minute
ConversationRelay (includes speech to text and voice)$0.07 a minute
Media Streams$0.004 a minute
Deepgram Nova-3 streaming$0.0077 a minute
ElevenLabs Flash v2.5$0.05 per 1,000 characters, about $0.02 to $0.04 per call minute
Cartesia SonicAbout $0.03 a minute
Text model tokens$0.005 to $0.03 a minute, our estimate on a mid-range model
Phone numberAbout a dollar a month

So ConversationRelay totals $0.08 to $0.11 and an own-pipeline build $0.05 to $0.09. At 5,000 minutes a month the gap is $100 to $150. It takes a lot of minutes to pay back the extra engineering, which is why we start most builds on ConversationRelay.

Tools, payments and handoffs

The model needs functions for anything factual: check availability, create a booking, look up a balance. Two rules. Have the agent say something ("let me check that") before a tool call that may take a second, or the caller hears silence. And make the model answer from the tool result, never from memory, or it will confirm a slot that doesn't exist.

For payments, end the ConversationRelay session and let the action URL return Twilio's <Pay> verb, which captures the card by keypad and tokenises it with Stripe at $0.10 per successful payment. Twilio's docs warn against putting card data in handoffData, and the agent should never hear a card number at all. For a human transfer, the same action handler returns a <Dial>.

A laptop showing code in a dark room
Photo by Nemuel Sereti on Pexels

The pitfalls

  • End-of-speech set wrong. Too short and the agent cuts people off mid-thought. Too long and every reply feels slow. Tune it against recordings of your real callers, who pause more than you expect when reading a number.
  • Interruptions. On ConversationRelay, set interruptSensitivity and turn on ignoreBackchannel so "yeah" and "uh-huh" don't stop the agent. On Media Streams, send clear, stop your voice generator, and use marks to work out what was played.
  • History that lies. After an interruption, trim the agent's last turn in the model's history to what the caller heard. Otherwise the model believes it said things the caller never received.
  • A default that changed. reportInputDuringAgentSpeech defaulted to any before May 2025 and now defaults to none. Older tutorials assume the old behaviour.
  • Dropped sockets. When your WebSocket dies, Twilio raises error 64105 and the session ends. Give the <Connect> verb an action URL that falls back to a person or to voicemail, or the caller gets dead air.
  • Concurrency. Accounts have a limit on simultaneous ConversationRelay sessions (error 64109). Ask Twilio for yours before launch, and load test your own server at twice the peak you expect.
  • Names and numbers. Feed product names and place names to the hints attribute, and read back anything that matters.
  • Transcripts are personal data. They hold names, addresses and health or money details. Set a retention period, and play a recording notice where the law needs one.

A sensible build order

  1. Week one: ConversationRelay, one model, one tool. Get a call answering and booking into a test calendar.
  2. Weeks two and three: the real booking system, the handoff to a person, payments through Pay.
  3. Weeks four to six: replay recorded calls through it, tune end-of-speech and interruptions, write the fallbacks.
  4. Go live after hours only. Read every transcript for a fortnight. Then widen it.

Which one to pick

ConversationRelay if you want a working agent soon and your volume is under tens of thousands of minutes a month. Media Streams with your own pipeline if you need a specific speech engine, a custom voice, or control over every millisecond. Speech-to-speech if the conversation itself is the product and tone matters more than cost predictability. Most small-business receptionists belong in the first group.

If you'd rather have it built than build it, see how our AI call answering works and book a 30-minute call. For the payment side in detail, read how AI phone agents take payments and stay PCI compliant.

Questions developers ask before they start

What's the difference between ConversationRelay and Media Streams?

ConversationRelay gives your server transcribed text and turns your text replies into speech, with interruption handling built in, for $0.07 a minute. Media Streams gives you raw 8 kHz mu-law audio for $0.004 a minute and you supply the speech recognition, the voice and the interruption logic.

How much does a Twilio voice agent cost per minute?

In raw vendor cost, $0.08 to $0.11 on ConversationRelay and $0.05 to $0.09 on Media Streams with your own pipeline, including $0.0085 for the inbound call. A speech-to-speech model runs $0.07 to $0.12 with prompt caching working and far more without it.

How long does it take to build?

A demo that answers and talks takes a weekend. A production agent connected to a real booking system, with handoffs, payments and fallbacks, takes four to eight weeks. Most of that is tuning against real calls.

What response time should I aim for?

Under 800 milliseconds from the caller finishing to the agent starting feels smooth. 800 to 1,200 is acceptable for business calls. Past 1,500 the caller notices a pause. End-of-speech detection is usually the biggest part of the budget.

Can the agent transfer a call to a person?

Yes. On ConversationRelay, send an end message with handoff data and have the action URL return a Dial verb. Give the agent clear rules for when to do it, and always set a fallback so a failed WebSocket doesn't leave the caller in silence.

Which language model should I use?

A fast mid-range text model with reliable function calling. First-token speed matters more than reasoning depth for a receptionist. Test two or three against your own recorded calls, because the one that benchmarks best is not always the one that books correctly.

Building something like this?

Tell us what it needs to do, and we'll scope it with you.

Start a project →