EN
返回档案库

档案库 · AI 与模型 · 产品决策 · 2025

Nari Labs 的开源 Dia TTS 两周下载破 10 万

两名韩国本科生用免费的谷歌 TPU 训练出 Dia——一个 16 亿参数、可控制笑声和咳嗽的开源 TTS,两周内下载量达到 10 万。

Nari Labs (나리랩스)

它在赌什么两人团队用免费的谷歌 TPU 训练一个 16 亿参数的开源 TTS,其可控的情感、笑声和咳嗽胜过 ElevenLabs 和 NotebookLM。已上线

做的是什么生意

Nari Labs is a two-person Seoul startup building open-source text-to-speech: its flagship Dia model (1.6B parameters, Apache-2.0) turns scripted multi-speaker dialogue into audio, letting users tag emotions and non-verbal cues like laughs and coughs.

启动资金None raised: founders say Dia was built with zero external funding, using Google's TPU Research Cloud and Hugging Face's ZeroGPU grant (VentureBeat, 2025-04).

起因

Toby Kim (Seoul National University) and Seong Jae-yong (KAIST) began learning speech AI about three months before launch. Inspired by NotebookLM's podcast feature, they wanted more control over voices and freedom in the script; after trying every commercial TTS API and finding none that sounded like real human conversation, they founded Nari Labs and trained on Google's free TPU Research Cloud.

经过

In April 2025 they released Dia under Apache-2.0: 1.6B parameters, speaker tags like [S1]/[S2], and cues such as (laughs) and (coughs) rendered as actual sound. It runs on PCs with about 10GB of VRAM. TechCrunch and VentureBeat covered it within a week; Hankyung TV reported downloads passed 100,000 in two weeks and Korean media framed the wave as 'K-deep voice', with praise from Hugging Face CEO Clem Delangue, Wharton professor Ethan Mollick and VC Deedy Das.

结果

No funding round has been announced; the repo reached 17.5K GitHub stars by 2025-07-21 (ROSS Q3-2025) and a 2026 roundup still lists Dia among the top five open-source TTS models. The founders' plan remains a consumer voice platform with a social layer.

背景

Nari Labs 是 2025 年初在首尔成立的两人创业公司,创始人是首尔大学的本科生 Toby Kim(金道烨)和韩国科学技术院的成在勇。赌注是:一个在谷歌免费 TPU 额度上训练的小团队,可以开源一个 16 亿参数的 TTS 模型,其对情感、节奏和非语言声音的控制能超过 ElevenLabs 和 NotebookLM。

结果是 Dia,一个 16 亿参数的模型,2025 年 4 月以 Apache-2.0 许可发布。它从脚本生成多说话人对话,用 [S1]/[S2] 等标签标记说话人,用 (laughs) 和 (coughs) 等提示生成真实声音,而不是说出“哈哈”。它需要约 10GB 显存,创始人在开始学语音 AI 后大约三个月就发布了。

报道和下载随之而来。TechCrunch 和 VentureBeat 都在 2025 年 4 月 21 日那周测试了 Dia,认为质量有竞争力;韩国经济电视台报道称 Hugging Face 下载量两周内突破 10 万,被韩国媒体称为“K-deep voice”。Runa Capital 的 ROSS 指数显示,GitHub 仓库星标从 1.5K 增长到 2025 年 7 月 21 日的 17.5K。

还没有公布融资轮次。公开的计划是在 Dia 之上做一个带社交层的消费级语音平台,2026 年 6 月的独立盘点仍把 Dia 列为前五开源 TTS 模型。悬而未决的问题是数据出处、语音克隆的滥用,以及模型周围是否能形成商业业务。

这件事要成立,得有什么

  • 免费 TPU 的限制迫使他们押注小模型:16 亿参数,能装进消费级显卡,差异点是可控性而不是规模。
  • 他们瞄准了现有玩家忽视的空白——脚本自由和非语言表达——而 NotebookLM 和 ElevenLabs 的输出听起来很生硬。
  • 采用 Apache-2.0 开源,让下载量和 GitHub 星标成了营销预算,带来了这两位学生买不到的报道。
  • 两人团队发货快:从未知到发布大约三个月,领先于大型实验室的发布周期。
  • 与 ElevenLabs 和 Sesame 的并排对比音频,把质量主张变成了记者和听众能自行验证的证据。

可借鉴之处

资金少的时候,挑现有玩家赚钱的瓶颈——表现力控制——然后把护城河送掉:开放权重让两个学生成了参照点。

后续进展

截至 2026 年 9 月 4 日,还没有公布融资轮次或付费平台上线。开源项目支撑了公司知名度:ROSS Q3-2025 记录到 2025 年 7 月 21 日 GitHub 星标 17.5K,2026 年 6 月的独立盘点仍把 Dia 列为前五开源 TTS 模型。TechCrunch 指出了未解决的风险——未披露的训练数据和可能助长诈骗的克隆——同时提到 Nari 的许可禁止滥用。创始人的计划仍是做一个带社交层的消费级语音平台,能否形成商业业务还未得到验证。

资料来源

发现哪里写错了?告诉我们。

轮到你了

你刚读完一家。说说你在做什么,看看谁在赌同一件事。

免费账号 · 3 次免费提问 · 不用绑卡

相关案例