The archive · AI & Models · Strategic decision · 2025–2026
Fish Audio's $52M seed: hobby TTS project hits $21M ARR in a year
Voice AI grown from a single-GPU open-source TTS project to 8M users and $21M ARR in a year, then a $52M seed led by Coreline and Capital Today.
Fish Audio
What the business is
Fish Audio sells expressive real-time text-to-speech, voice cloning and voice agents: three open-source models plus a paid hosted API with 15,000+ natural-language controls.
Starting capital:$52M seed announced 2026-07-28, led by Coreline Ventures and Capital Today, with 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners and Alphalist Partners.
How it started
Fish Audio began as Fish Speech, an open-source project born in co-founder Shijia Liao's bedroom: frustrated by robotic synthetic voices, the former Nvidia video researcher trained a voice-generation model on a single gaming GPU and released it. The repo grew to more than 31,000 GitHub stars, drawing indie developers, video game designers and creators.
What happened
In its first year Fish Audio shipped five models — four speech-generation and one speech-to-text — open-sourcing three and keeping its flagship S2.1 Pro API-only, and added paid creator/team plans plus an enterprise platform used by HeyGen, Sanas, LiveKit, Retell and OpenArt. Growth reached 8M users and $21M ARR. A voice-consent controversy emerged when some creators alleged their voices were uploaded without permission and DMCA takedowns were slow; the company later automated removals to under three minutes.
How it ended up
On 2026-07-28 Fish Audio announced a $52M seed led by Coreline Ventures and Capital Today, to expand beyond TTS into a full audio-native stack — voice-native LLMs, speech-to-speech, and an audio-understanding model planned this year — plus an enterprise sales team.
Background
Fish Audio started as Fish Speech, an open-source text-to-speech project born in co-founder Shijia Liao's bedroom. A former Nvidia video researcher and lifelong anime fan frustrated by monotonous synthetic voices, Liao trained a voice-generation model on a single gaming GPU and released it. The repository grew to more than 31,000 GitHub stars, attracting indie developers, video game designers and creators.
In its first year the company shipped five models — four speech-generation and one speech-to-text — open-sourcing three and reserving its flagship S2.1 Pro for a paid API. It added monthly creator and team plans plus an enterprise platform, with customers including HeyGen, Sanas, LiveKit, Retell, OpenArt and Telnyx. By July 2026 it reported 8 million users across open-source and hosted versions and $21 million in annual recurring revenue.
On July 28, 2026 Fish Audio announced a $52 million seed round led by Coreline Ventures and Capital Today. The money funds expansion beyond text-to-speech into a full audio-native stack: voice-native LLMs, speech-to-speech and an audio-understanding model, plus an enterprise sales team. The round followed a consent controversy — creators alleged voices were uploaded without permission and takedowns were slow — which the company answered by automating removals to under three minutes.
What has to be true
- Open-sourcing the models built the funnel: 31k GitHub stars and a developer community gave Fish Audio distribution no paid marketing could match (TechCrunch).
- The seed came after, not before, the revenue: $21M ARR in year one let the founders stay efficient and raise from a position of traction (official PR).
- Community-driven voice data created the moat but also the risk: user-submitted voices powered the library, and slow consent takedowns threatened the trust the model depended on (TechCrunch).
- The same bet works in both directions: free open models win creators, while enterprise needs — latency, steerability, HIPAA, zero-data-retention — justify the paid API (official PR).
What can be applied
Giving the core model away built community, brand and demand data that let a bootstrapped project raise a $52M seed — but the same openness created a consent problem that forced product changes.
Aftermath
As of 2026-09-02 Fish Audio is scaling: the $52M seed (2026-07-28) funds an audio-native stack — voice-native LLMs, speech-to-speech, an audio-understanding model — plus an enterprise sales team and deeper integrations with LiveKit and Retell. S2.1 Pro was free via API through August 2026. It faces a crowded market (ElevenLabs, WellSaid, Cartesia, Speechify, Async, Krisp) and creator-trust questions around voice consent, now handled with automated sub-three-minute takedowns.
Sources
- Fish Audio raises $52M seed to build AI voice models for creators and enterprises
- Fish Audio Raises $52M in Seed Funding After Turning Passion Project Into One of Voice AI's Fastest-Growing Companies
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card