The archive · AI & Models · Technical decision · 2023–2026
Shengshu Technology Bets Video Generation Leads to World Models
Tsinghua team builds Vidu, raises over RMB 2.5 billion in 3 years, benchmarking Sora.
Shengshu Technology (生数科技)
What the business is
Shengshu Technology (生数科技) is a Beijing-based multimodal generative AI company. Its core product is the text-to-video large model Vidu (including the Vidu Q3 series); it provides MaaS to enterprises and multimodal generative applications to consumers, and open-sourced the robot world model Motus at the end of 2025.
Starting capital:June 2023 angel round valuation: USD 100 million; cumulative public financing exceeds RMB 2.5 billion (Qiming (启明), Baidu (百度), the Beijing AI Fund (北京市AI基金), Alibaba Cloud (阿里云), etc.).
How it started
In March 2023, Zhu Jun (朱军), vice dean of the Institute for Artificial Intelligence at Tsinghua University (清华大学人工智能研究院), and Tsinghua alumni Tang Jiayu (唐家渝) and Bao Fan (鲍凡) founded Shengshu Technology (生数科技) in Beijing. The team had been researching diffusion model acceleration since 2021 (DPM-Solver was adopted by Stable Diffusion and DALL·E 2), judged that large models would move from language to multimodal fusion, and persisted with the U-ViT fusion architecture from the outset. In March 2023, it open-sourced UniDiffuser, the world's first multimodal diffusion model based on U-ViT; in June, its angel round was valued at USD 100 million; subsequently, it launched two tools, PixWeaver and VoxCraft, to test the consumer side.
What happened
In April 2024, it released the text-to-video large model Vidu, raising the ceiling for domestic professional-grade video length to 16 seconds and fully benchmarking Sora; in June, it completed a Pre-A round of several hundred million yuan co-led by the Beijing Municipal Artificial Intelligence Industry Investment Fund (北京市人工智能产业投资基金) and Baidu (百度). In 2025, users and revenue both grew more than 10 times, with services covering more than 200 countries worldwide; in September, it completed a Series A of several hundred million yuan (led by Bohua Capital (博华资本), with Baidu Strategic Investment (百度战投) and Qiming Venture Partners (启明创投) following); in December, it open-sourced the Motus model, which controls robots based on multimodal data. In January 2026, it released Vidu Q3 (16-second synchronized audio and video, 1080P, precise shot switching); in February, it completed a Series A+ round exceeding RMB 600 million (co-led by Zhongguancun Science City (中关村科学城) and Xinglian Capital (星连资本), setting a new record for the largest single financing round in China's video generation sector); in April, it completed a Series B round of nearly RMB 2 billion, led by Alibaba Cloud (阿里云), with the entire Vidu series listed on the Alibaba Cloud Bailian (百炼) model marketplace on the same day, bringing cumulative public financing to more than RMB 2.5 billion.
No ending yet — it is still running.
Background
Shengshu's bet predates founding: in Sep 2022, team published U-ViT, first Diffusion Transformer, 3 months before DiT; Mar 2023 open-sourced UniDiffuser. In Mar 2023, Zhu Jun, Tang Jiayu, Bao Fan founded Shengshu in Beijing, betting on multimodal fusion. Locked U-ViT path; June 2023 angel round valued at USD 100M.
Apr 2024: released Vidu, 16s video, benchmarking Sora. Capital accelerated: Jun Pre-A co-led by Beijing AI Fund and Baidu; Sep 2025 Series A led by Bohua; Feb 2026 A+ >RMB 600M led by Zhongguancun and Xinglian; Apr Series B ~RMB 2B led by Alibaba Cloud, Vidu on Bailian. Cumulative >RMB 2.5B; shareholders include Qiming, Baidu, Ant, Huawei, Zhipu, Alibaba.
Commercialization: 2025 users/revenue grew 10x, 200+ countries, clients Sony, iQIYI, Mango, ByteDance, Samsung, TAL. Dec 2025 open-sourced Motus for robot control. Vidu Q3 (Jan 2026) supports 16s 1080P audio-video. Competition: Kuaishou Kling ARR >USD 240M; Aishi got USD 60M from Alibaba; race is now capital and ecosystem.
What has to be true
- Academic first-mover does not equal commercial: U-ViT preceded DiT, but Sora shipped first; Shengshu caught up via Vidu's engineering and ecosystem.
- Single video generation lacks moat: financing records broken; Shengshu pivoted to 'general world model + robots' to justify RMB 2B funding.
- Tying to big-company ecosystems is survival: Baidu led Pre-A, Alibaba led Series B and listed on Bailian, trading compute and channels for dependence.
- Moving from generation to understanding is a gamble: Motus controls robots via multimodal data, betting world models beat LLMs; cycle and path uncertain.
What can be applied
Academic first-mover advantage does not equal commercial first-mover: video generation became capital-intensive; moat must extend to applications understanding the physical world.
Aftermath
As of April 10, 2026: Shengshu completed Series B of nearly RMB 2 billion led by Alibaba Cloud; cumulative financing exceeded RMB 2.5 billion; Vidu listed on Alibaba Cloud Bailian. In 2025, users and revenue grew 10x, covering 200+ countries, with clients like Sony Pictures, iQIYI, Mango TV, ByteDance, Samsung. Next stage: general world models; Vidu Q3 supports 16s 1080P audio-video; Motus open-sourced Dec 2025. Pressure: Kuaishou Kling ARR >USD 240M; Aishi got USD 60M from Alibaba; big companies' models squeeze; differentiation via world-model path and Alibaba ecosystem.
Sources
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card