The archive · AI & Models · Strategic decision · 2025–2026
Standard Intelligence: $75M from Sequoia to train agents on raw video
A six-person lab of college dropouts bets pixel-level video pretraining, not language models, will produce agents that can use any computer.
Standard Intelligence
What the business is
Standard Intelligence is a six-person AI research lab in San Francisco, founded by Galen Mead and Devansh Pandey, building FDM-1: a foundation model pretrained directly on millions of hours of screen video that predicts mouse movements, clicks and keystrokes. Its goal is an agent that explores and operates software the way a person would, and it has also released an open conversational-speech model, hertz-dev.
Starting capital:$75M Series A announced April 2026, co-led by Sequoia Capital and Spark Capital, with angels including Andrej Karpathy, Stanley Druckenmiller and Milan Kovac; the team numbered six people at the time of the round.
How it started
Galen Mead and Devansh Pandey met in 2022 as teenagers in the Atlas Fellowship, a selective program for high-school students interested in AI alignment, then left their undergraduate programs to found Standard Intelligence. Sequoia's write-up calls their bet the 'Tesla FSD approach applied to knowledge work': pretrain on the raw stream of computer use and let generality emerge, instead of hand-engineering workflows.
What happened
FDM-1, the lab's first foundation model trained directly on computer-use video, can extrude a CAD gear in Blender, drive a car around a San Francisco block after an hour of fine-tuning and find bugs by exploring software, according to Sequoia. On April 30, 2026 Standard Intelligence announced a $75M Series A co-led by Sequoia and Spark Capital, with Andrej Karpathy among the angels; Pulse2 reported that the founders say FDM-1 moved computer use 'from a data-constrained regime to a compute-constrained one.'
How it ended up
Still a research lab: FDM-1 has demonstration videos but no public product or API as of April 2026. The round funds compute scaling and alignment research on 'general learners,' and the lab has released an open speech model, hertz-dev, while the six-person team hires researchers.
Background
Standard Intelligence is a six-person AI research lab in San Francisco founded by Galen Mead and Devansh Pandey, who met as teenagers in the Atlas Fellowship and left university to work on aligned AGI. Its thesis, per investor Sequoia, is that the best path to general computer agents runs through raw video: pretrain on the stream of computer use and let generality emerge, the way Tesla scaled self-driving from pixels.
The lab built an 11-million-hour computer-action dataset, which Sequoia calls the largest in the industry, plus a video encoder roughly 50x more token-efficient than competing approaches, letting hours of 30 FPS screen footage fit in a 1-million-token context window. FDM-1, its first model trained on that data, can extrude a CAD gear in Blender, drive a car after an hour of fine-tuning and find bugs by exploring software.
On April 30, 2026 Standard Intelligence announced a $75M Series A co-led by Sequoia Capital and Spark Capital, with angels including Andrej Karpathy, Stanley Druckenmiller and Milan Kovac. Pulse2 reported the founders' claim that FDM-1 moved computer use from a data-constrained to a compute-constrained regime, and that the capital unlocks several orders of magnitude more compute.
As of April 2026 FDM-1 is a research model with demos, not a product, and the open question is whether video pretraining will beat the language-model agent stacks that OpenAI, Anthropic and Google sell. The lab has also released an open speech model, hertz-dev, and published how it racked 30 petabytes of storage for under $500,000.
What has to be true
- Data scales where harnesses do not: 11 million hours of screen video is a pretraining set nobody else assembled, and it grows automatically as people use computers.
- Token efficiency is the unlock: a roughly 50x more efficient video encoder makes training on hours of 30 FPS footage affordable instead of impossible.
- Contrarian capital: Sequoia and Spark backed a six-person team with no product, betting a new pretraining regime, not wrappers, builds general agents.
- Alignment is part of the bet: the founders met in an AGI-alignment fellowship and treat safety research as a condition of scaling, not an afterthought.
What can be applied
When everyone scaffolds around an incumbent model, go downstream to the data it never used: raw screen video became this six-person lab's edge over every agent framework.
Aftermath
As of April 30, 2026, Standard Intelligence was spending its $75M Series A on compute scaling at 'several orders of magnitude' larger, per its founders, and hiring researchers for what it calls a science of alignment for general learners. FDM-1 remained a research model with demo videos rather than a product, and the lab had also released hertz-dev, an open 8.5B-parameter speech model, and written up the 30-petabyte storage cluster it built for under $500,000. Whether video-pretrained agents can beat language-model agent stacks on real computer work was still unproven.
Sources
- Standard Intelligence: Training General Intelligence in Pixel Space
- Standard Intelligence Raises $75 Million From Sequoia And Spark Capital To Scale AGI Research
- Founded & Funded: 2026-Q2 Portfolio News
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card