档案库 · 开发与企业工具 · 产品决策 · 2026
Voicebox 的本地方向:4 个月 26.5K stars,100 万下载,不用云端
Jamie Pine 的开源工 Voicebox,在本地克隆声音、生成语音、听写文字;到 2026 年 5 月 19 日 GitHub 获得 26.5K stars,6 月 27 日下载超 100 万。
Voicebox
做的是什么生意
Voicebox is a free, open-source, local-first AI voice studio: clone a voice from a few seconds of audio, generate speech in 23 languages across seven TTS engines, dictate into any application with a global hotkey, and give MCP-aware AI agents a voice — with all models and audio staying on-device.
启动资金:No venture funding; donations brought in roughly $200–300/month until June 2026, when an optional $VOICEBOX token was launched to fund full-time development (locked liquidity, buybacks and burns); paid cloud sync (~$12/year) planned as the first revenue stream.
起因
Jamie Pine, the Canadian developer known for the open-source file manager Spacedrive, built the first version as a one-day experiment the day Qwen3-TTS was released, and open-sourced it three days later. He did no marketing; Reddit found the project, creators made tutorials, and it grew organically from late January 2026.
经过
Version 0.5.0 (April 2026) turned the app from a cloning studio into a full voice I/O platform: system-wide dictation with a floating overlay, 23 languages, and native integration with MCP-aware agents. By mid-May it passed 26,500 stars and drew mainstream coverage that flagged the consent gap — no technical mechanism verifies that a cloned voice's owner consented — set against a projected $40B US gen-AI fraud wave by 2027. In June 2026 Pine explained that open source 'doesn't pay rent' ($200–300/month in donations) and launched the optional $VOICEBOX token, saying it let him go full-time on the project overnight.
结果
Scaling: as of 2026-07-10 the repo was approaching 40K stars (~1,200 stars in one day), and Pine is building a mobile app plus encrypted paid cloud backup/sync (~$12/year, free for token holders), with the desktop app remaining free and local-first.
背景
Voicebox 是免费、开源的 AI 语音工作室,由 Spacedrive 文件管理器的开发者 Jamie Pine 开发。它从几秒音频克隆声音,用七个 TTS 引擎生成 23 种语言的语音,提供系统级听写,还能让 MCP 感知的代理(如 Claude Code 和 Cursor)用克隆声音说话——所有这些都在本地进行,不上传任何音频。Pine 在阿里巴巴发布 Qwen3-TTS 的当天花一天时间做了第一个版本,三天后开源,没做任何营销。
Reddit 发现了它,创作者制作了教程,还有"ElevenLabs 刚失去护城河"的帖子跟进:到 2026 年 5 月 19 日 GitHub 上超过 26,500 stars,到 6 月 27 日下载超百万——这个项目在增长过程中接近了他更出名的 Spacedrive 的 star 数量。主流关注带来审视:TechTimes 报道了应用在克隆声音方面没有任何同意验证机制,同时预计到 2027 年美国生成式 AI 诈骗浪潮将达 400 亿美元,欧盟 AI 法案的深度伪造标签期限是 2026 年 8 月。
业务问题随规模而来:捐款每月只有 200–300 美元,所以 2026 年 6 月 Pine 推出了可选 $VOICEBOX 代币——锁定流动性、回购和销毁——称这是他最接近工资的东西,让他能全职做 Voicebox。应用本身保持免费和开源;计划推出付费加密云同步(约 12 美元/年)作为真正的收入流。截至 2026 年 7 月,仓库接近 40K stars,移动应用在开发中。
这件事要成立,得有什么
- 赌的是隐私加开源能成为护城河:整个语音流程在设备端,用户就能享受与 ElevenLabs 相当的功能,但没有订阅成本,也没有数据保留风险。
- 楔入点是对称性——在一个本地应用中同时覆盖听写输入和文本转语音输出,这是 ElevenLabs 和 Wispr Flow 分开的两半语音回路。
- 分发是挣来的,不是买来的:没有营销支出,Reddit 发现、创作者教程和 star 增长(第四个月 26.5K)完成了本应靠订阅做的工作。
- 资金是为了适应模式而临时拼凑的:当捐款无法持续时,可选的社区代币(而不是付费墙)保持了开放核心的承诺。
可借鉴之处
当整个技术栈都在本地运行时,开源能把用户信任转化为免费的自然传播——真正的问题变成在保持承诺的前提下资助维护者。
后续进展
截至 2026 年 9 月 2 日,Voicebox 仍然免费、采用 MIT 许可证且本地优先,仓库在 2026 年 7 月接近 40K stars(一些追踪报告现在约 50K),发布频繁(7 月有 v0.4.x–v0.5.x 时代的功能)。Pine 表示他全职投入,移动应用原型和付费加密云备份/同步已宣布但截至最后验证更新尚未上线。代币是可选的支持机制;项目没有风险投资或公司出售。
资料来源
- Voicebox Clones Any Voice From 3 Seconds of Audio, Runs Locally for Free, and Has No Consent Lock
- Why Voicebox has a token
- The open-source AI voice studio (README)
发现哪里写错了?告诉我们。
轮到你了
你刚读完一家。说说你在做什么,看看谁在赌同一件事。
免费账号 · 3 次免费提问 · 不用绑卡