档案库 · AI 与模型 · 技术决策 · 2023–2026
生数科技押注视频生成通往世界模型
清华团队打造Vidu,三年融资超25亿元人民币,对标Sora。
生数科技
做的是什么生意
Shengshu Technology (生数科技) is a Beijing-based multimodal generative AI company. Its core product is the text-to-video large model Vidu (including the Vidu Q3 series); it provides MaaS to enterprises and multimodal generative applications to consumers, and open-sourced the robot world model Motus at the end of 2025.
启动资金:June 2023 angel round valuation: USD 100 million; cumulative public financing exceeds RMB 2.5 billion (Qiming (启明), Baidu (百度), the Beijing AI Fund (北京市AI基金), Alibaba Cloud (阿里云), etc.).
起因
In March 2023, Zhu Jun (朱军), vice dean of the Institute for Artificial Intelligence at Tsinghua University (清华大学人工智能研究院), and Tsinghua alumni Tang Jiayu (唐家渝) and Bao Fan (鲍凡) founded Shengshu Technology (生数科技) in Beijing. The team had been researching diffusion model acceleration since 2021 (DPM-Solver was adopted by Stable Diffusion and DALL·E 2), judged that large models would move from language to multimodal fusion, and persisted with the U-ViT fusion architecture from the outset. In March 2023, it open-sourced UniDiffuser, the world's first multimodal diffusion model based on U-ViT; in June, its angel round was valued at USD 100 million; subsequently, it launched two tools, PixWeaver and VoxCraft, to test the consumer side.
经过
In April 2024, it released the text-to-video large model Vidu, raising the ceiling for domestic professional-grade video length to 16 seconds and fully benchmarking Sora; in June, it completed a Pre-A round of several hundred million yuan co-led by the Beijing Municipal Artificial Intelligence Industry Investment Fund (北京市人工智能产业投资基金) and Baidu (百度). In 2025, users and revenue both grew more than 10 times, with services covering more than 200 countries worldwide; in September, it completed a Series A of several hundred million yuan (led by Bohua Capital (博华资本), with Baidu Strategic Investment (百度战投) and Qiming Venture Partners (启明创投) following); in December, it open-sourced the Motus model, which controls robots based on multimodal data. In January 2026, it released Vidu Q3 (16-second synchronized audio and video, 1080P, precise shot switching); in February, it completed a Series A+ round exceeding RMB 600 million (co-led by Zhongguancun Science City (中关村科学城) and Xinglian Capital (星连资本), setting a new record for the largest single financing round in China's video generation sector); in April, it completed a Series B round of nearly RMB 2 billion, led by Alibaba Cloud (阿里云), with the entire Vidu series listed on the Alibaba Cloud Bailian (百炼) model marketplace on the same day, bringing cumulative public financing to more than RMB 2.5 billion.
还没有结局,它还在跑。
背景
生数的押注早于公司成立:2022年9月,团队发表U-ViT,首个Diffusion Transformer,比DiT早三个月;2023年3月开源UniDiffuser。2023年3月,朱军、唐家渝、鲍凡在北京创立生数科技,押注多模态融合。锁定U-ViT路线;2023年6月天使轮估值1亿美元。
2024年4月:发布Vidu,16秒视频,对标Sora。资本加速:6月Pre-A轮获北京AI基金和百度联合领投;2025年9月A轮由博华领投;2026年2月A+轮超6亿元,中关村和星连领投;4月B轮近20亿元,阿里云领投,Vidu上架百炼。累计超25亿元;股东包括启明、百度、蚂蚁、华为、智谱、阿里。
商业化:2025年用户和收入增长10倍,覆盖200多个国家,客户包括索尼、爱奇艺、芒果、字节、三星、好未来。2025年12月开源Motus用于机器人控制。Vidu Q3(2026年1月)支持16秒1080P音视频。竞争:快手可灵ARR超2.4亿美元;爱诗获阿里6000万美元;竞争已变成资本和生态之争。
这件事要成立,得有什么
- 学术先发不等于商业先发:U-ViT早于DiT,但Sora先发布;生数通过Vidu的工程和生态追赶。
- 单一视频生成缺乏护城河:融资纪录不断被打破;生数转向“通用世界模型+机器人”以支撑20亿元融资。
- 绑定大厂生态是生存之道:百度领投Pre-A,阿里领投B轮并上架百炼,用算力和渠道换依赖。
- 从生成走向理解是一场赌博:Motus通过多模态数据控制机器人,押注世界模型胜过LLM;周期和路径不确定。
可借鉴之处
学术先发不等于商业先发:视频生成变成资本密集型;护城河必须延伸到理解物理世界的应用。
后续进展
截至2026年4月10日:生数完成近20亿元B轮融资,由阿里云领投;累计融资超过25亿元;Vidu上架阿里云百炼。2025年用户和收入增长10倍,覆盖200多个国家,客户有索尼影业、爱奇艺、芒果TV、字节跳动、三星。下一阶段:通用世界模型;Vidu Q3支持16秒1080P音视频;Motus于2025年12月开源。压力:快手可灵ARR超2.4亿美元;爱诗获阿里6000万美元;大厂模型挤压;通过世界模型路径和阿里生态实现差异化。
资料来源
发现哪里写错了?告诉我们。
轮到你了
你刚读完一家。说说你在做什么,看看谁在赌同一件事。
免费账号 · 3 次免费提问 · 不用绑卡