档案库 · AI 与模型 · 战略决策 · 2024–2026
Sesame押注逼真语音是AI的下一个界面:Maya演示爆红,获2.5亿美元B轮融资
Oculus创始人的Sesame押注富有表现力的语音而非文本聊天机器人是AI的下一个界面;Maya演示走红,随后获得红杉领投的2.5亿美元融资。
Sesame
做的是什么生意
A San Francisco conversational-AI startup building expressive voice agents (Maya, Miles, Simone, Charlie) in an iOS app, with its own AI glasses planned for 2027.
起因
Sesame emerged from stealth in February 2025, founded by Oculus co-founder and former CEO Brendan Iribe, former Ubiquity6 CTO Ankit Kumar, former Meta Reality Labs engineering director Ryan Brown and others, backed by a16z, Spark and Matrix — all early Oculus investors. Its bet was that humans evolved for voice: rather than converting LLM text output to audio, Sesame's Conversational Speech Model, trained on roughly one million hours of public audio, generates speech directly to capture rhythm, emotion and expressiveness.
经过
The Maya and Miles demo went viral within weeks: more than one million people engaged and generated over five million minutes of conversation, per investor Sequoia. On 2025-10-21 Sesame announced a $250 million Series B co-led by Sequoia and Spark and opened an early iOS beta. On 2026-05-28 it shipped a public iOS preview in 39 countries with four persistent agents — Maya, Miles, Simone and Charlie — featuring parallel search while speaking, notes, search cards, a texting mode and incognito memory; the company said agents will later take actions for users, and AI eyewear is targeted for 2027.
还没有结局,它还在跑。
背景
Sesame于2025年2月低调成立,其反主流押注是:计算的下一大转变是语音,仅靠模型是不够的——用户感受到的是智能的传递方式。公司由Oculus联合创始人Brendan Iribe和前Ubiquity6 CTO Ankit Kumar创立,得到a16z、Spark和Matrix的支持,构建了一个对话语音模型,直接生成语音而不是转换聊天机器人的文本,模型基于约一百万小时的公共音频训练。
证据来自演示。Maya和Miles两个语音在2025年2月发布后几周内走红:据投资方红杉称,超过一百万人参与其中,产生了超过五百万分钟的对话。The Verge称其“真的很有趣”,并表示这是其评测者第一次想多次交谈的语音助手。2025年10月21日,Sesame宣布获得由红杉和Spark共同领投的2.5亿美元B轮融资,并开放了早期iOS测试版。
2026年5月28日,Sesame在39个国家推出了公开iOS预览版,包含四个持久代理——Maya、Miles、Simone和Charlie——各自拥有独特的声音、个性和记忆。应用在说话时并行进行搜索,可以在句子中途转向,并加入了笔记、搜索卡片、短信模式和隐身记忆。Sesame表示,代理将来会代表用户行动,其设计为全天佩戴的AI眼镜预计在2027年推出。
这件事要成立,得有什么
- 人类为语音而进化,因此最自然的界面是对话式语音,而不是向聊天机器人输入提示词。
- 一个逼真到人们愿意持续交谈的演示在几周内吸引了超过100万用户和500万分钟对话——这种需求支撑了2.5亿美元的融资。
- 直接生成语音而非文本到语音转换,能捕捉节奏和情感,这是“听起来像人”和“听起来像机器人”的区别。
- 拥有代理和眼镜,Sesame就拥有了环境智能、始终可用的界面的完整技术栈,而不是在别人的设备上租用屏幕。
可借鉴之处
当竞争对手在模型智能方面竞赛时,界面可以成为切入点:一个感觉像人的语音吸引了百万人进入演示,并在真正产品发布前获得了2.5亿美元的融资。
后续进展
截至2026年9月2日,Sesame在39个国家运营免费的公开iOS预览版,注册时仍需排队,并计划推出Android预览版。该应用是公司宣称的最终目标的第一步:预计在2027年推出智能眼镜,代理将从与用户并排思考转变为代表用户采取行动。目前未披露收入数据或估值;公司仍由风险投资支持,且尚未推出硬件。
资料来源
- Sesame, the conversational AI startup from Oculus founders, launches its iOS app
- Sesame, the conversational AI startup from Oculus founders, raises $250M and launches beta
- Sesame is the first voice assistant I've ever wanted to talk to more than once
- Partnering with Sesame: A New Era for Voice
发现哪里写错了?告诉我们。
轮到你了
你刚读完一家。说说你在做什么,看看谁在赌同一件事。
免费账号 · 3 次免费提问 · 不用绑卡