EN
Back to the archive

The archive · AI & Models · Technical decision · 2021–2026

Twelve Labs bets video is AI's foundation; $100M Series B with Amazon and Nvidia

Korean-founded Twelve Labs bets native video foundation models — Marengo and Pegasus — will let AI watch, remember and reason like humans; $207M+ raised.

Twelve Labs (트웰브랩스)

The betVideo, not language, will be the foundation of AI: native video models (Marengo, Pegasus) let machines watch, remember and reason over footage like humans.Scaling

What the business is

Twelve Labs builds video-understanding foundation models — Marengo (index/retrieval) and Pegasus (video-language reasoning) — sold via API and Amazon Bedrock to enterprises searching and analyzing large video archives.

How it started

Founded in 2021 in San Francisco by five Koreans led by Jae Lee, with core AI research done in Seoul. The founding observation: 'the world does not happen in text. It happens in motion.' Seed funding reached about $17M, and in October 2023 Nvidia made its first investment in a Korean generative-AI startup, backing a roughly $10M pre-Series A.

What happened

June 2024: a $50M Series A co-led by Nvidia's NVentures and NEA, joined by Index Ventures, Radical Ventures, WndrCo and Korea Investment Partners. Products matured fast: Marengo-2.6 shipped in March 2024 for any-to-any video search, Pegasus-1 went to beta for video-language reasoning. December 2024: SK Telecom invested $3M to use the models in its AI agent and surveillance work. In July 2025 Amazon made Marengo 2.7 and Pegasus 1.2 fully managed in Bedrock; by 2026 GS Shop's video recommendation built on the models doubled click-through and raised purchase intent 9x, and U.S. government agencies were exploring video intelligence for defense and public safety.

How it ended up

Still running and scaling: a $100M Series B announced 2026-07-01 (Naver Ventures and NEA co-led, Amazon participating) brought total funding past $207M. CEO Jae Lee says the goal is 'video superintelligence' — AI that watches, remembers and reasons like humans — with New York and London offices planned and a first AI video-creation app, Rodeo, in preview.

Background

Twelve Labs, founded in 2021 in San Francisco by five Koreans led by Jae Lee, with core AI research in Seoul, makes video the native modality of AI rather than an add-on to text. Its founding observation: 'the world does not happen in text. It happens in motion.' The company builds two foundation models — Marengo, which embeds visual, audio, speech and on-screen text into one searchable representation, and Pegasus, which reasons over that representation to answer, summarize and ground descriptions in footage.

The contrarian bet attracted unusual backers early: in October 2023 Nvidia made its first investment in a Korean generative-AI startup with a ~$10M pre-Series A, and in June 2024 Nvidia's NVentures co-led a $50M Series A with NEA, joined by Index Ventures, Radical Ventures, WndrCo and Korea Investment Partners. In December 2024 SK Telecom invested $3M to fold the models into its AI agent and surveillance work.

Distribution followed: in July 2025 Amazon made Marengo 2.7 and Pegasus 1.2 fully managed in Amazon Bedrock, and by 2026 GS Shop's video recommendation system built on the models doubled click-through and lifted purchase intent ninefold, while U.S. government agencies explored video intelligence for defense and public safety. A significant share of revenue comes from U.S. and European enterprises, with developer revenue growing 80–120% annually.

In July 2026 Twelve Labs closed a $100M Series B co-led by Naver Ventures and NEA with Amazon participating — the first Korean AI startup backed by both Amazon and Nvidia — bringing total funding past $207M. CEO Jae Lee frames the destination as 'video superintelligence': AI that watches, remembers and reasons over recorded reality like a human, with offices planned in New York and London.

What has to be true

  • Twelve Labs chose video as the primary data type — not a second modality — arguing that text is a lossy compression of reality and causality lives in sequence.
  • The architecture made video machine-readable memory: footage is understood once, stored durably, and addressable to the exact second — what enterprises with huge archives actually need.
  • Nvidia's 2023 investment was the first validation and the wedge: it signaled the bet was credible before the models shipped, and three years of collaboration led to the Series B.
  • Commercial proof came from a narrow set — GS Shop's doubled click-through and 9x purchase intent — showing the platform could monetize before 'video superintelligence' arrived.

What can be applied

Betting on a modality most of AI ignored — video instead of text — pulled in Nvidia and Amazon as backers; the contrarian bet must still convert into enterprise revenue to compound.

Aftermath

As of 2026-09-02 Twelve Labs has raised more than $207M: the July 2026 Series B ($100M, co-led by Naver Ventures and NEA with Amazon) funds Marengo, Pegasus and its Video Cognition System. Enterprise revenue is concentrated in the U.S. and Europe, developer revenue grows 80–120% annually, adopters include GS Shop and U.S. agencies, and its models run fully managed in Amazon Bedrock. It plans New York and London offices, previewed its video-creation app Rodeo, and keeps pushing Korean-built video models as leaders of the emerging video-intelligence era.

Sources

spotted an error? The archive wants to know.

Your turn

You just read one. Describe what you are building, and see who is betting on the same thing.

Free account · 3 free questions · no card

Related cases