← All posts
AI voice agentsMarket landscapeBuying guide

AI voice agent platforms in 2026: who's building what

16 September 2026 · 9 min read

If you've started looking into an AI phone answerer, you'll have hit a wall of near-identical websites. They all promise natural voices, 24/7 cover and instant setup. Underneath, they're built for very different jobs — and some of them aren't built for a small business's inbound line at all. This post maps the category plainly: what the main platforms are actually for, what a call really costs to run, and where the whole idea falls down.

First, what an "AI voice agent" actually is

Almost every product in this category is the same four pieces bolted together, whatever the branding says:

  1. Speech-to-text. Turning what the caller says into words the software can read.
  2. A language model. Deciding what to say back, and when to look something up or write something down.
  3. Text-to-speech. Turning that reply into a voice.
  4. Telephony. The actual phone line carrying the call.

Each layer costs money per minute and adds delay. That's the whole story of this market, really. The difference between a good agent and a bad one is usually how tightly those four pieces are stitched together, and how much thought has gone into the boring bit at the end — what happens to the booking, the message, the callback.

It matters because the layers are largely commodity. Anyone can rent them. So the interesting question when you look at a platform isn't "can it talk", it's "what has this company decided to build on top, and for whom".

The cost floor nobody puts on the homepage

Across the builder platforms in this category — Vapi, Retell, Bland, Synthflow — the real all-in running cost sits at roughly $0.13 to $0.33 a minute once you've paid for speech-to-text, the language model, text-to-speech and telephony. That's the raw cost of a call existing. It is not a price you pay a vendor; it's the floor underneath every price in the market.

Do the maths on your own line

Say you take 200 calls a month and they average two and a half minutes. That's 500 minutes.

At the bottom of the range, $0.13 a minute, the underlying compute costs about $65. At the top, $0.33, about $165. Convert at whatever the rate is the day you read this — at roughly 1.25 dollars to the pound that's about £52 to £132 of raw cost.

Anything a vendor charges you above that is covering support, setup, the software around the call, and margin. That's fair enough — but it tells you something useful. If someone quotes you a figure that works out well below the floor, ask what they've cut. If someone quotes a figure many times above it, ask what you're getting for the difference.

For a like-for-like look at what UK buyers actually pay at retail rather than at cost, we've broken that down separately in AI receptionist cost UK.

The platforms, and what each one is really for

Retell AI — the flexible builder

Retell is a custom-receptionist builder. Templates to start from, APIs and webhooks to wire it into your own systems, call analytics and transcripts on the way out. Its standout technical feature is responsiveness: a proprietary WebRTC stack that keeps latency in the 400 to 700 millisecond range, which is about as good as this category currently gets.

That matters more than it sounds. Latency is the gap between the caller finishing a sentence and the agent starting its reply. Push past roughly a second and callers start talking over the top, repeating themselves, or deciding they're on a bad line and hanging up. Retell is best suited to inbound work and local-service booking — which is to say, the job most readers of this blog actually have.

The trade-off is that a builder is a builder. You get the parts. Configuring intake questions, deciding what happens when the agent can't help, connecting the diary, testing edge cases and maintaining it as your business changes — that's your afternoon, or your developer's week.

Synthflow — no-code setup with enterprise telephony

Synthflow's pitch is that you don't need to write code to stand up a voice agent, and it brings enterprise-grade telephony with it. For a business that wants to configure something itself without hiring anyone, that's a genuine draw.

The weakness showed up in head-to-head testing on a specific scenario: mid-call rescheduling. When a caller changed the service type partway through the conversation, the agent looped back to the start of intake rather than adapting. Anyone who's answered a phone knows how common that is. "Actually, it's not the boiler, it's the radiator — and can we do Thursday instead?" A human absorbs that in half a second. An agent that restarts the script has just lost you the caller.

It's one test, not a verdict on the whole platform. But it's the exact failure mode worth checking on any product you're shown, because demos are almost always run on clean, linear calls.

Bland AI — API-first, built for outbound volume

Bland is developer-facing and honest about it. The pricing is clearer on a per-minute basis than most, and the architecture is built for high-volume outbound campaigns — the sort of scale that runs to millions of concurrent calls.

That is a serious piece of engineering and it is not aimed at you. A single small business with one inbound line has the opposite problem to a campaign operator: you don't need throughput, you need one call handled properly at 7pm on a Sunday. Bland is a good answer to a question most trades, salons and clinics aren't asking.

11x — enterprise B2B outbound

11x sits in the enterprise sales-development space: outbound B2B, aimed at companies with a sales team and a pipeline to feed. It isn't positioned for small-business inbound reception and doesn't pretend to be. Include it in your map of the category so you understand why the marketing language across these sites feels so different — half the market is selling to sales directors, not to owners who are missing calls from a loft.

Side by side

PlatformBuilt forNotable strengthWhere it's weak for a small inbound line
Retell AICustom inbound receptionists, local-service booking400–700ms latency on a proprietary WebRTC stack; templates, APIs, webhooks, transcriptsYou're assembling and maintaining it yourself
SynthflowNo-code agent setup with enterprise telephonyConfigurable without a developerFailed mid-call rescheduling in testing — looped back to start of intake
Bland AIHigh-volume outbound campaignsClear per-minute pricing; scales to millions of concurrent callsAPI-first and campaign-shaped; overkill for one inbound number
11xEnterprise B2B outbound salesFits a sales-development workflowNot aimed at small-business inbound reception at all

The split that actually matters: parts versus a finished thing

Once you've read four of these sites you notice the real division isn't between brands, it's between two business models.

One group sells you the parts. You get an agent builder, a dashboard, some documentation and a per-minute rate. You decide the script, the fallbacks, the integrations, the escalation rules. The ceiling is high and the floor is wherever your patience gives out.

The other group sells you a configured service. Someone else has already decided how intake works, where the job notes go, what happens when the agent is out of its depth, and how you get told. You give up flexibility and buy back your evenings. RedAgents sits in this second camp — an AI receptionist set up for you rather than a toolkit. Either model can be the right one; the mistake is buying a toolkit expecting a service, or a service expecting a toolkit.

If you want to see what the finished-service version does on a call, minute by minute, we've walked through it in what actually happens when an AI receptionist answers.

The 60–70% number, and the calls that aren't in it

Across this category, a well-configured agent on routine scheduling and FAQ work typically resolves 60 to 70% of calls without a human. That's a useful benchmark to hold any demo against — and the word "well-configured" is doing real work in that sentence. A hastily set up agent will sit below it.

Note what the figure implies. Somewhere between three and four calls in ten still need a person. The question you should be asking a vendor isn't "how good is your agent", it's "what happens on the other 30 to 40%". Does it take a proper message and get it to me fast? Does it transfer live? Does it say so honestly, or bluff? A platform that resolves 65% cleanly and hands over the rest gracefully beats one that resolves 70% and mangles the remainder.

Five things to test on any demo

  1. Change your mind mid-call. Switch service type or date halfway through. See whether the agent adapts or restarts.
  2. Interrupt it. Talk over the top the way real callers do. Listen to whether it stops.
  3. Time the gaps. If replies feel like they land after a beat of silence, callers will notice too.
  4. Ask something it can't answer. A price you don't publish, a scenario off-script. You're testing whether it admits the limit or invents an answer.
  5. Follow the output. Where did the booking, message or callback request end up, and how quickly did it reach a phone you actually look at?

Where this whole category isn't the right answer

Being even-handed cuts both ways, so: there are businesses where none of these platforms is the right buy in 2026.

If your call volume is genuinely low — a handful a week, all from people who'd happily leave a voicemail — the per-minute cost is irrelevant but the setup effort isn't. You'd spend more time configuring than answering.

If your calls are emotionally heavy or clinically sensitive, a synthetic voice is the wrong front door regardless of how good it sounds. Bereavement, safeguarding, distressed patients — those need a person from the first second.

If every call ends in a judgement only you can make — bespoke pricing, structural surveys, complex quoting — then an agent can qualify and book, but it can't do the job you're actually being rung about. That's still valuable, but be clear-eyed that you're buying triage, not delegation.

And if the real problem is that you want a human who knows your regulars by name, hire one, or use a human answering service. We've laid out that comparison honestly in AI receptionist vs answering service vs voicemail.

What's likely to differentiate these products next

This is opinion rather than fact, so treat it as such. The underlying layers are converging. Voices are broadly convincing across the board, and latency is heading towards the point where the difference stops being audible. When the parts become commodity, the parts stop being the product.

What's left to compete on is the unglamorous middle: handling a caller who changes their mind, knowing when to stop and fetch a human, getting the message into your hands in a form you can act on from a van or a salon chair. Those are workflow problems, not voice problems. When you're comparing platforms, weight your judgement accordingly — and if you want a structured way to interrogate any provider, the nine questions to ask a call handling service apply just as well to an AI one.