The archive · AI & Models · Product decision · 2024
Hume AI bets voice AI should read emotion; $50M Series B and EVI 2 in 2024
Ex-Google DeepMind researcher Alan Cowen's Hume turned emotional science into an API; a $50M Series B funded EVI, and EVI 2 cut latency 40% and price 30%.
Hume AI
What the business is
Hume AI, named after philosopher David Hume, trains voice models on cross-cultural voice recordings matched with self-reported emotion surveys. Its Empathic Voice Interface (EVI) is an end-to-end speech-to-speech API developers embed in apps, customer service lines, and assistants, alongside Expression Measurement and Custom Models APIs.
Starting capital:$50M Series B in spring 2024, reported by VentureBeat and announced alongside the EVI beta launch per ZDNET; earlier rounds not covered in cited sources.
How it started
Hume AI was co-founded and led by Alan Cowen, a computational scientist and former Google DeepMind researcher, to build AI optimized for human well-being; the company trains its models on cross-cultural voice recordings matched with self-reported emotional surveys, and it sells APIs that detect emotional expression as well as generate it.
What happened
After a $50M Series B in spring 2024, Hume launched the EVI beta and public demo in late March 2024, drawing immediate hands-on coverage (ZDNET) and, by July, named customers including Lawyer.com's 1-800 line. EVI 2 shipped in September 2024 with an end-to-end speech model, 40% lower latency (typical responses 500-800ms), customizable voices without cloning, in-conversation prompts, and a price cut to $0.072/min from $0.102/min — a 30% reduction — positioning Hume ahead of OpenAI's still-waitlisted GPT-4o Advanced Voice Mode and an Anthropic-Amazon Alexa project.
No ending yet — it is still running.
Background
Hume AI, named after the 18th-century Scottish philosopher David Hume, was co-founded and led by Alan Cowen, a former Google DeepMind researcher and computational scientist. The company trains its proprietary voice models on cross-cultural recordings matched with self-reported emotional surveys, and it sells APIs that both measure emotional expression and generate emotionally expressive speech.
In spring 2024 Hume raised a $50M Series B, then launched the beta of its Empathic Voice Interface (EVI) — which it billed as the first AI with emotional intelligence — with a public demo in late March. ZDNET tested it within days, reporting that when its writer spoke in a fake crying voice, EVI detected sadness and distress and responded with encouragement; Marketplace profiled it in July 2024, citing Lawyer.com's 1-800 line and unnamed big tech companies as customers.
EVI 2 arrived in September 2024 with an end-to-end speech model that takes audio in and audio out, cutting latency 40% (typical responses 500-800ms) and price 30% to $0.072/min from EVI 1's $0.102/min. New features included voice modulation for tunable personalities without voice cloning, in-conversation prompts, and plans for Spanish, French, and German; Hume said it would sunset EVI 1 in December 2024.
VentureBeat framed the timing as the bet paying off: EVI 2 launched while OpenAI's GPT-4o Advanced Voice Mode was still waitlist-only, and Cowen argued developers would embed empathic voice directly inside apps rather than routing users to phone lines. Experts interviewed by Marketplace were more cautious, questioning whether emotion inference is scientifically reliable or ethically safe to deploy at scale.
What has to be true
- Cowen's background made the bet plausible: a psychologist trained in measuring emotion, plus DeepMind experience, building on cross-cultural voice and self-report data rather than scraped audio alone.
- The product answered a concrete developer pain — voice agents that sound robotic and ignore tone — with an end-to-end model that was 40% faster and 30% cheaper than its own predecessor.
- Safety became a feature: refusing voice cloning and offering tunable personalities addressed the exact risk that made enterprises cautious about voice AI.
- Timing created buzz: EVI's public demo and EVI 2's launch both landed while OpenAI's GPT-4o voice mode was still waitlisted, letting Hume claim the lead in emotionally expressive, embeddable voice.
- The open risk is that the category's claims are contested: Marketplace's expert sources question whether emotion inference is reliable or safe, which could cap enterprise adoption.
What can be applied
Hume bet emotion, not raw speech quality, is the missing voice-AI layer — and that selling it safely (no cloning, tunable voices) could win developers before OpenAI's voice mode went wide.
Aftermath
As of September 2024 Hume is live: EVI 2 is in open beta through the API at $0.072/min with 500-800ms responses, following a $50M Series B in spring 2024 and an EVI 1 launch that drew first-person press tests and customers such as Lawyer.com. The company also sells Expression Measurement and Custom Models APIs, has no plans to offer voice cloning, and planned to sunset EVI 1 in December 2024 and add Spanish, French and German. Whether 'emotional AI' becomes a developer default or a niche depends on whether the science holds up; Marketplace's experts questioned its reliability and ethics.
Sources
- Who needs GPT-4o Advanced Voice Mode? Hume's EVI 2 is here with emotionally inflected voice AI and API
- An AI model with emotional intelligence? I cried, and Hume's EVI told me it cared
- A tour of 'emotionally intelligent' AI
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card