Voice AI Companies in 2026
Voice AI companies build software that understands and produces human speech. Their products include text-to-speech (TTS), speech-to-text (STT), real-time speech models, and platforms for running AI voice agents that handle phone calls or conversations. Customers are usually developers and businesses that put voice into products, contact centers, media and devices.
Voice AI Companies to know
We chose these companies for scale, funding, press coverage and adoption as of October 2026. Each one's main product is a speech model (TTS, STT or real-time voice) or a platform built mainly for voice agents. We favored companies we could check in independent reporting and funding databases, and we checked every funding figure against public reporting. We left out general AI agent vendors whose main channel is not voice, real-time infrastructure that fits AI Agent Infrastructure better, and companies that have moved into nearby categories such as voice evaluation.

ElevenLabs
Text-to-speech, voice cloning, dubbing, speech-to-text and voice agent platform
ElevenLabs develops AI voice models for text-to-speech, voice cloning, dubbing and speech-to-text (Scribe), along with a platform for conversational voice agents. Eleven v3 supports more than 70 languages. It sells to creators, media companies, developers and enterprises, and the company says it reached $500M in annual recurring revenue in 2026.
- Headquarters
- London, UK

Deepgram
Speech-to-text, text-to-speech and voice agent APIs for developers and enterprises
Deepgram sells speech recognition and voice APIs, including the Nova-3 speech-to-text model, the Aura-2 text-to-speech model, Flux conversational speech recognition and a Voice Agent API. Developers and enterprises use them for real-time voice applications. In January 2026 it acquired OfOne, a voice automation platform for restaurants, and folded it into Deepgram for Restaurants.
- Headquarters
- San Francisco, California, USA

PolyAI
Enterprise voice agents for customer service phone calls
PolyAI started as a spin-out from the University of Cambridge. It builds voice agents that answer customer service calls for large companies, using its own dialog models (Dialog-RSN-1, Raven) and its Agent Studio tooling. The company says it has more than 2,000 live deployments in 45 languages and over 100 enterprise customers.
- Headquarters
- London, UK

Sesame
Conversational voice AI companion and voice-first smart glasses
Sesame builds conversational speech models that generate speech directly rather than reading out text, with demo voices named Maya and Miles. It launched an iOS app beta in October 2025 and is building lightweight smart glasses around its voice companion. Former Oculus executives founded the company, which came out of stealth in February 2025.
- Headquarters
- San Francisco, California, USA

Cartesia
Low-latency text-to-speech and speech-to-text models built on state space models
Cartesia was founded by researchers from Stanford's AI Lab. It builds real-time voice models on state space model (SSM) architectures rather than transformers. Its products include the Sonic text-to-speech models (most recently Sonic-3.6, August 2026), the Ink-2 speech-to-text model (July 2026) and tools for voice agents.
- Headquarters
- San Francisco, California, USA

AssemblyAI
Speech-to-text and speech understanding APIs for developers
AssemblyAI trains its own speech recognition and speech understanding models and sells them as APIs. These cover pre-recorded and real-time transcription, audio intelligence, and a voice agent API. Its customers are mostly developers and product companies building on top of transcription.
- Headquarters
- San Francisco, California, USA

Vapi
Developer platform for building and deploying AI voice agents
Vapi offers developers infrastructure to build, test and run AI voice agents that handle inbound and outbound phone calls, and agents can use different LLM and speech providers. The company reports more than 1 billion calls handled on the platform. Amazon Ring routes all of its inbound calls through Vapi after evaluating more than 40 vendors.
- Headquarters
- San Francisco, California, USA

Speechmatics
Speech recognition for real-time and batch transcription in 55+ languages
Speechmatics is a speech recognition company that spun out of Cambridge AI research. It sells speech-to-text for real-time and batch use with speaker diarization and code-switching, deployable in the cloud, on-premises or on-device. It also has specialized models, including a bilingual Arabic-English medical model launched in 2026.
- Headquarters
- Cambridge, UK

Retell AI
Platform for building AI voice agents that answer and place phone calls
Retell AI offers an API and a no-code builder for AI voice agents that handle phone calls such as receptionist work, appointment setting, lead qualification and collections. Its 2026 releases include Retell Workflows for CRM and helpdesk automation and Conductor for voice agent monitoring. The company went through Y Combinator's Winter 2024 batch.
- Headquarters
- Redwood City, California, USA

Bland AI
Enterprise AI phone agents with in-house speech models
Bland AI builds AI phone agents that companies use for customer service, sales, lead qualification, receptionist and IVR-replacement calls. Its Conversational Pathways system constrains agent behavior, and it runs its own text-to-speech and transcription models to keep latency low. The company says it handles more than 3.5 million calls a week.
- Headquarters
- San Francisco, California, USA
Voice AI Companies at a glance
| Company | Headquarters | Focus |
|---|---|---|
| ElevenLabs | London, UK | text-to-speech, voice cloning |
| Deepgram | San Francisco, California, USA | speech-to-text, text-to-speech |
| PolyAI | London, UK | voice agents, contact center |
| Sesame | San Francisco, California, USA | conversational voice, speech generation |
| Cartesia | San Francisco, California, USA | text-to-speech, speech-to-text |
| AssemblyAI | San Francisco, California, USA | speech-to-text, API |
| Vapi | San Francisco, California, USA | voice agents, developer platform |
| Speechmatics | Cambridge, UK | speech-to-text, ASR |
| Retell AI | Redwood City, California, USA | voice agents, telephony |
| Bland AI | San Francisco, California, USA | voice agents, phone automation |
Frequently asked questions
Which companies are the most notable in voice AI in 2026?
Among the best-funded are ElevenLabs ($500M Series D at an $11B valuation, February 2026), Sesame ($250M Series B, October 2025), Deepgram ($130M Series C at a $1.3B valuation, January 2026), PolyAI ($86M Series D, December 2025), Vapi ($50M Series B at about a $500M valuation, May 2026) and Bland AI ($50M Series C, June 2026). Cartesia, AssemblyAI, Speechmatics and Retell AI are also widely used.
What is the difference between voice model companies and voice agent platforms?
Voice model companies train their own speech models and sell them through APIs. Examples are Deepgram (Nova-3 STT, Aura-2 TTS), AssemblyAI (speech-to-text), Speechmatics (ASR) and Cartesia (Sonic TTS, Ink-2 STT). Voice agent platforms such as Vapi, Retell AI and Bland AI connect speech models, LLMs and telephony so businesses can run phone agents. Some do both: ElevenLabs offers TTS and STT models plus a conversational agent platform, and Bland AI runs its own TTS and transcription models.
Which voice AI companies focus on speech-to-text?
Deepgram (Nova-3, Flux), AssemblyAI (Speech-to-Text and Speech Understanding APIs) and Speechmatics (real-time and batch ASR in 55+ languages, with on-premises and on-device options) focus mainly on speech recognition. ElevenLabs (Scribe) and Cartesia (Ink-2) also offer speech-to-text alongside their TTS models.
Which companies build voice agents for contact centers?
PolyAI builds enterprise customer-service voice agents and reports more than 2,000 live deployments in 45 languages. Vapi gives developers infrastructure for phone agents and handles all inbound calls for Amazon Ring. Retell AI and Bland AI offer platforms for AI phone agents that handle scheduling, lead qualification and customer service; Bland AI reports more than 3.5 million calls a week.
How is Sesame different from other voice AI companies?
Most companies here sell APIs or enterprise platforms. Sesame builds a consumer conversational companion with a speech model that generates speech directly, launched as an iOS app beta in October 2025, and it is developing smart glasses for it. Former Oculus executives, including Brendan Iribe and Nate Mitchell, founded it.
More AI company directories
- AI Agent CompaniesSierra, Harvey, Decagon and more
- AI Infrastructure CompaniesDatabricks, Scale AI, Surge AI and more
- Healthcare AI CompaniesOpenEvidence, Abridge, Tempus AI and more
- AI Coding CompaniesCursor (Anysphere), Cognition, GitHub and more
- AI GPU Cloud CompaniesCoreWeave, Lambda, Modal and more
- AI Model Hosting CompaniesHugging Face, Baseten, Modal and more
About this directory: OpenCurious curates this list for researchers, founders, operators and buyers. Company descriptions and headquarters are based on publicly reported information as of and change often; figures a company reports about itself are its own claims. This is an editorial selection, not a paid ranking: no company paid to be listed, and companies appear in no particular order. Company names and logos are trademarks of their respective owners and are used only to identify each company. See our editorial policy. To suggest a company or a correction, email hello@opencurious.com.
The OpenCurious Newsletter
Stay curious. Get the best of the open web.
New directories, researcher tools, and open-source finds — straight to your inbox. No spam, ever.
- New directories
- Researcher tools
- Open-source finds
Occasional emails from OpenCurious. Unsubscribe anytime. Privacy Policy