EN
Back to the archive

The archive · AI & Models · Product decision · 2025

Nari Labs' open-source Dia TTS drew 100K downloads in two weeks

Two Korean undergrads trained Dia, a 1.6B-parameter open TTS with laugh and cough control, on free Google TPUs and hit 100K downloads in two weeks.

Nari Labs (나리랩스)

The betA two-person team training on free Google TPUs could open-source a 1.6B-parameter TTS whose controllable emotion, laughs and coughs beat ElevenLabs and NotebookLM.Live

What the business is

Nari Labs is a two-person Seoul startup building open-source text-to-speech: its flagship Dia model (1.6B parameters, Apache-2.0) turns scripted multi-speaker dialogue into audio, letting users tag emotions and non-verbal cues like laughs and coughs.

Starting capitalNone raised: founders say Dia was built with zero external funding, using Google's TPU Research Cloud and Hugging Face's ZeroGPU grant (VentureBeat, 2025-04).

How it started

Toby Kim (Seoul National University) and Seong Jae-yong (KAIST) began learning speech AI about three months before launch. Inspired by NotebookLM's podcast feature, they wanted more control over voices and freedom in the script; after trying every commercial TTS API and finding none that sounded like real human conversation, they founded Nari Labs and trained on Google's free TPU Research Cloud.

What happened

In April 2025 they released Dia under Apache-2.0: 1.6B parameters, speaker tags like [S1]/[S2], and cues such as (laughs) and (coughs) rendered as actual sound. It runs on PCs with about 10GB of VRAM. TechCrunch and VentureBeat covered it within a week; Hankyung TV reported downloads passed 100,000 in two weeks and Korean media framed the wave as 'K-deep voice', with praise from Hugging Face CEO Clem Delangue, Wharton professor Ethan Mollick and VC Deedy Das.

How it ended up

No funding round has been announced; the repo reached 17.5K GitHub stars by 2025-07-21 (ROSS Q3-2025) and a 2026 roundup still lists Dia among the top five open-source TTS models. The founders' plan remains a consumer voice platform with a social layer.

Background

Nari Labs is a two-person Seoul startup founded in early 2025 by undergrads Toby Kim (Kim Do-yeob) of Seoul National University and Seong Jae-yong of KAIST. The bet: a small team training on free Google TPU credits could open-source a text-to-speech model whose control over emotion, timing and non-verbal sounds beats ElevenLabs and NotebookLM.

The result was Dia, a 1.6-billion-parameter model released in April 2025 under Apache-2.0. From a script it generates multi-speaker dialogue, using tags like [S1]/[S2] for speakers and cues such as (laughs) and (coughs) that come out as actual sound instead of a spoken 'haha'. It runs on about 10GB of VRAM, roughly three months after the founders first started learning speech AI.

Coverage and downloads followed the release. TechCrunch and VentureBeat both tested Dia in the week of 2025-04-21 and rated the quality competitive, and Hankyung TV reported Hugging Face downloads passed 100,000 within two weeks, a moment Korean media called 'K-deep voice'. Runa Capital's ROSS index shows the GitHub repo growing from 1.5K to 17.5K stars by 2025-07-21.

No funding round has been announced. The stated plan is a consumer voice platform with a social layer on top of Dia, and a June 2026 independent roundup still lists Dia among the top five open-source TTS models. The open questions are data provenance, abuse of voice cloning, and whether a commercial business ever forms around the model.

What has to be true

  • The free-TPU constraint forced a small-model bet: 1.6B parameters that fit a consumer GPU, where controllability, not scale, is the differentiator.
  • They aimed at the gap incumbents neglected — script freedom and non-verbal delivery — where NotebookLM and ElevenLabs outputs sounded canned.
  • Open-sourcing under Apache-2.0 made downloads and GitHub stars the marketing budget, earning coverage the two students could not have bought.
  • A two-person team shipped fast: from first exposure to launch took about three months, staying ahead of bigger labs' release cycles.
  • Side-by-side comparison audio against ElevenLabs and Sesame turned a quality claim into evidence journalists and listeners could verify for themselves.

What can be applied

When capital is scarce, pick the bottleneck incumbents monetize — expressive control — and give the moat away: open weights turned two students into the reference point.

Aftermath

As of 2026-09-04 no funding round or paid platform launch has been announced. The open-source project carried the company's visibility: ROSS Q3-2025 logged 17.5K GitHub stars by 2025-07-21, and a June 2026 independent roundup still lists Dia among the top five open TTS models. TechCrunch flagged the unresolved risks — undisclosed training data and cloning that could aid scams — while noting Nari's license prohibits misuse. The founders' plan remains a consumer voice platform with a social layer; whether a commercial business forms is still unproven.

Sources

spotted an error? The archive wants to know.

Your turn

You just read one. Describe what you are building, and see who is betting on the same thing.

Free account · 3 free questions · no card

Related cases