EN
Back to the archive

The archive · AI & Models · Strategic decision · 2025–2026

ShoFo bets an indexed 'Common Crawl for video' will feed AI labs; YC W26

ShoFo crawls billions of short-form videos, cleans and labels them, and sells custom datasets to AI labs - a bet that video data, not models, is the bottleneck.

ShoFo (Shofo)

The betVideo data, not models, will be AI's bottleneck: owning a pipeline that crawls, cleans and labels the public video web beats building the next model.Live

What the business is

ShoFo sells custom video datasets to AI labs: its crawlers index billions of short-form clips from TikTok, Instagram Reels and YouTube Shorts, then agents clean, dedupe and annotate matching clips for delivery in standard training formats.

How it started

Founded 2025 in San Francisco by a University of California, Santa Barbara-connected team that had built Correkt, an AI search engine for multimodal content that reached more than 40,000 users. While running Correkt they built crawling infrastructure and concluded the plumbing was worth more than the consumer product, so they kept the pipes and rebuilt the company around selling video data.

What happened

ShoFo went through Y Combinator's Winter 2026 batch and pivoted from Correkt to a data pipeline: crawl short-form video, screen and dedupe it, add object-detection and vision-language annotations, and deliver datasets in WebDataset, Hugging Face or raw formats. A 2026-07-29 profile said the four-person team had built the index, published a 58,000-video Hugging Face sample with 25,000+ downloads, and was selling datasets to labs training multimodal models, priced by how much annotation each clip needs.

No ending yet — it is still running.

Background

ShoFo is a San Francisco data-infrastructure startup from Y Combinator's Winter 2026 batch that calls itself 'The World's Largest Video Library' and, internally, 'Common Crawl for video'. Its bet is that video will matter to the next generation of AI models the way text did to LLMs, and that nobody has yet built the layer that collects, cleans and labels the public video web.

The pipeline runs crawl, sanitize, label, deliver: crawlers pull short-form clips from TikTok, Instagram Reels, YouTube Shorts and the wider web; NSFW screening, quality filtering and perceptual hashing remove unusable and duplicate footage; object detection plus vision-language reasoning add activity labels and segmentation; and results ship in formats labs already use. The company says the hard part is not the individual tools but running them at scale on video and keeping an index that turns a lab's niche request into a query instead of a months-long collection project.

ShoFo grew out of Correkt, a multimodal AI search engine the same founders built that reached more than 40,000 users. They noticed the crawling infrastructure they had built for the consumer app was worth more than the app itself, so they pivoted the company around the data layer. Its go-to-market gives a sample away first: a roughly 58,000-video Hugging Face dataset that had been downloaded more than 25,000 times by July 2026, intended to prove label quality to research labs before enterprise contracts.

The thesis is early and unproven at scale, as its own profile notes: funding is modest, the company had four people as of July 2026, and incumbents such as Scale AI and Labelbox already sell labeling services. ShoFo's difference is owning a pre-built index of public video before any customer shows up, and pointing it specifically at video rather than everything.

What has to be true

  • Multimodal labs genuinely lack labeled video at scale; a six-month collection bottleneck is the problem ShoFo sells against.
  • The pivot read its own metrics: Correkt's 40,000-user search engine proved the crawl worked before ShoFo existed as a company.
  • The free Hugging Face sample with 25,000+ downloads is a public proof-of-quality that also generates inbound demand.
  • Selling datasets rather than seats ties revenue to the buyer's training need, and the pre-built index makes marginal delivery cheap.
  • Anti-scraping defenses and private data partnerships give the team a head start a well-funded clone cannot quickly replicate.

What can be applied

When the plumbing built for one product becomes the scarce asset, rebuild the company around it; when selling data, giving away a quality sample first can be the cheapest way to win demanding buyers.

Aftermath

As of September 4, 2026, ShoFo is live and selling: it came out of Y Combinator's Winter 2026 batch, operates from San Francisco with a four-person team, indexes billions of short-form videos, and delivers custom cleaned and labeled datasets to AI labs, with a ~58,000-video Hugging Face sample downloaded more than 25,000 times. Whether the index becomes the infrastructure layer for video AI or an incumbent's future feature is still open; the company is early, modestly funded, and its bet depends on video becoming as central to AI as text has been.

Sources

spotted an error? The archive wants to know.

Your turn

You just read one. Describe what you are building, and see who is betting on the same thing.

Free account · 3 free questions · no card

Related cases