档案库 · AI 与模型 · 战略决策 · 2025–2026
Fish Audio 5200 万美元种子轮:业余 TTS 项目一年内 ARR 达 2100 万美元
语音 AI 从一个单 GPU 开源 TTS 项目成长为拥有 800 万用户和 2100 万美元年经常性收入的公司,随后由 Coreline 和今日资本领投 5200 万美元种子轮。
Fish Audio
做的是什么生意
Fish Audio sells expressive real-time text-to-speech, voice cloning and voice agents: three open-source models plus a paid hosted API with 15,000+ natural-language controls.
启动资金:$52M seed announced 2026-07-28, led by Coreline Ventures and Capital Today, with 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners and Alphalist Partners.
起因
Fish Audio began as Fish Speech, an open-source project born in co-founder Shijia Liao's bedroom: frustrated by robotic synthetic voices, the former Nvidia video researcher trained a voice-generation model on a single gaming GPU and released it. The repo grew to more than 31,000 GitHub stars, drawing indie developers, video game designers and creators.
经过
In its first year Fish Audio shipped five models — four speech-generation and one speech-to-text — open-sourcing three and keeping its flagship S2.1 Pro API-only, and added paid creator/team plans plus an enterprise platform used by HeyGen, Sanas, LiveKit, Retell and OpenArt. Growth reached 8M users and $21M ARR. A voice-consent controversy emerged when some creators alleged their voices were uploaded without permission and DMCA takedowns were slow; the company later automated removals to under three minutes.
结果
On 2026-07-28 Fish Audio announced a $52M seed led by Coreline Ventures and Capital Today, to expand beyond TTS into a full audio-native stack — voice-native LLMs, speech-to-speech, and an audio-understanding model planned this year — plus an enterprise sales team.
背景
Fish Audio 始于开源文本转语音项目 Fish Speech,诞生于联合创始人廖世佳的卧室。廖世佳是前英伟达视频研究员,也是一名终生的动漫迷,对单调的合成语音感到失望,他用一张游戏 GPU 训练了语音生成模型并将其发布。该仓库获得了超过 31000 个 GitHub 星标,吸引了独立开发者、电子游戏设计师和创作者。
在成立的第一年,公司发布了五个模型——四个语音生成和一个语音转文本——其中三个开源,旗舰产品 S2.1 Pro 仅通过付费 API 提供。它增加了月度创作者和团队套餐以及企业平台,客户包括 HeyGen、Sanas、LiveKit、Retell、OpenArt 和 Telnyx。截至 2026 年 7 月,它报告称开源和托管版本共有 800 万用户,年经常性收入达 2100 万美元。
2026 年 7 月 28 日,Fish Audio 宣布获得由 Coreline Ventures 和今日资本领投的 5200 万美元种子轮。这笔资金用于将业务扩展到文本转语音之外的完整音频原生技术栈:语音原生大语言模型、语音到语音和音频理解模型,以及企业销售团队。此轮融资发生在同意争议之后——创作者声称声音未经许可被上传,删除速度缓慢——公司通过自动化删除将时间缩短至三分钟以内来回应。
这件事要成立,得有什么
- 开源模型建立了漏斗:31000 个 GitHub 星标和开发者社区为 Fish Audio 提供了付费营销无法比拟的传播渠道(TechCrunch)。
- 种子轮发生在收入之后,而非之前:第一年 2100 万美元的 ARR 让创始人保持了效率,并凭借势头进行融资(官方新闻稿)。
- 社区驱动的语音数据创造了护城河,但也带来了风险:用户提交的声音支撑了库,而缓慢的同意删除威胁到了模型所依赖的信任(TechCrunch)。
- 同样的赌注双向适用:免费开放模型赢得创作者,而企业需求——低延迟、可控性、HIPAA 合规、零数据保留——证明了付费 API 的合理性(官方新闻稿)。
可借鉴之处
免费提供核心模型建立了社区、品牌和需求数据,使一个自力更生的项目能够筹集 5200 万美元种子轮——但同样的开放性也导致了同意问题,迫使产品做出改变。
后续进展
截至 2026-09-02,Fish Audio 正在规模化:5200 万美元种子轮(2026-07-28)资助音频原生技术栈——语音原生大语言模型、语音到语音、音频理解模型——以及企业销售团队和与 LiveKit 和 Retell 的更深集成。S2.1 Pro 在 2026 年 8 月之前通过 API 免费提供。它面临一个拥挤的市场(ElevenLabs、WellSaid、Cartesia、Speechify、Async、Krisp)以及围绕语音同意的创作者信任问题,现在通过自动化三分钟内的删除来处理。
资料来源
- Fish Audio raises $52M seed to build AI voice models for creators and enterprises
- Fish Audio Raises $52M in Seed Funding After Turning Passion Project Into One of Voice AI's Fastest-Growing Companies
发现哪里写错了?告诉我们。
轮到你了
你刚读完一家。说说你在做什么,看看谁在赌同一件事。
免费账号 · 3 次免费提问 · 不用绑卡