The archive · AI & Models · Product decision · 2023–2026
Boson AI bets open Higgs voice models beat cloud giants; $70M, 7.5K-star repo
Ex-AWS scientists Smola and Li Mu open-source Higgs TTS to seed a voice-AI stack; $70M raised as Higgs RealTime prepares to challenge OpenAI and Meta.
Boson AI
What the business is
Boson AI builds foundation audio models — Higgs TTS, speech-to-speech and avatar APIs — for conversational voice agents, targeting enterprises in finance, telecom, healthcare and insurance, with open weights for earlier models.
Starting capital:$70M raised across seed rounds; backed by Chinese entrepreneur Su Hua and Temasek's venture arm (Fortune, 2026-07-27).
How it started
Alex Smola, a former distinguished scientist at AWS, co-founded Boson AI in 2023 with Mu Li (李沐), his former student and a co-creator of the MXNet deep-learning framework, after concluding the industry was moving beyond text-only interfaces into multimodal audio and vision. The company began shipping Higgs-series models and open-sourced Higgs TTS 2 in May 2025.
What happened
Higgs TTS 2's open release built community traction — the boson-ai/higgs-audio repo hit 7.5K GitHub stars by 21 October 2025 (7.5x growth in Q4 2025, ROSS Index) — and the company kept shipping: Higgs TTS 3 on 04 June 2026 (100+ languages, zero-shot voice cloning, inline emotion control), a Higgs Avatar API in June 2026, and a Higgs Audio Hackathon in March 2026. Fortune reported $70M raised, backed by Su Hua and Temasek, with enterprise targets in finance, telecom, healthcare and insurance.
How it ended up
Live and scaling: as of September 2026 Boson AI is preparing Higgs RealTime, its first speech-to-speech model, claiming roughly one-tenth the cost of rivals like OpenAI; the Higgs Audio v3 release now lives on Hugging Face with a hosted API at boson.ai.
Background
Boson AI builds foundation audio models — text-to-speech, speech-to-speech and avatar generation — for conversational voice agents. Its Higgs line includes Higgs TTS 2 (open-sourced May 2025, trained on over 10 million hours of audio), Higgs TTS 3 (June 2026, 100+ languages with zero-shot cloning and inline emotion control), a Higgs Avatar API, and the upcoming Higgs RealTime speech-to-speech model.
The company was founded in 2023 by Alex Smola, a former distinguished scientist at AWS, and Mu Li (李沐), his former student and a co-creator of the MXNet deep-learning framework, on the belief that AI was moving beyond text-only interfaces. Boson's strategy pairs open-weight releases for developer adoption with hosted APIs and enterprise licenses for revenue.
The open-source bet showed in GitHub traction: the boson-ai/higgs-audio repo reached 7.5K stars by 21 October 2025 (7.5x growth, Runa Capital ROSS Index Q4 2025). Fortune reported in July 2026 that the startup had raised $70 million, backed by Su Hua and Temasek's venture arm, and was targeting finance, telecom, healthcare and insurance clients while claiming model costs around one-tenth of rivals like OpenAI.
The bet, in one line: voice is the next human-AI interface, and a full-stack, cheaper, open-weight alternative can carve a lane between OpenAI, Meta and ElevenLabs — even against incumbents that can bundle voice into existing platforms.
What has to be true
- Open-sourcing earlier Higgs models built developer goodwill and GitHub visibility without an ad budget, seeding the ecosystem the paid API later serves.
- The one-tenth-cost claim targeted the highest-volume pain in voice AI — inference cost — making price the wedge against well-funded incumbents.
- A full-stack bet (TTS, ASR, avatar, speech-to-speech) gave enterprises a single vendor for the whole voice loop, aligned with data-residency needs.
- Targeting finance, telecom, healthcare and insurance first gave the seed-stage company a narrow enterprise beachhead before consumer scale.
What can be applied
Open-weight release is cheap distribution for a model startup: giving away earlier models builds GitHub visibility and developer goodwill, while paid APIs and enterprise licenses capture value later.
Aftermath
As of 02 September 2026 Boson AI is live and scaling. Higgs TTS 3 is released (04 June 2026) with open weights on Hugging Face and a free hosted preview at boson.ai; the Higgs Avatar API shipped in June 2026; and Higgs RealTime, the company's first speech-to-speech model, was set to launch around August 2026 (Crypto Briefing, Fortune). The company reports roughly $70M raised with Su Hua and Temasek backing, and is selling enterprise deployments with on-premise data-residency options across finance, telecom, healthcare and insurance.
Sources
- How an ex-AWS scientist plans to take on OpenAI and Meta in voice AI models
- Boson AI founder Alex Smola targets voice AI market with Higgs RealTime model
- Top trending open-source startups in Q4 2025
- boson-ai/higgs-audio
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card