The archive · AI & Models · Strategic decision · 2023–2026
Inferact 2026: vLLM creators turn the top open-source serving engine into an $800M startup
Berkeley's vLLM, a default open engine for serving LLMs (~90.6k GitHub stars), spun out as Inferact in Jan 2026 with a $150M seed at $800M.
Inferact (vLLM)
What the business is
Inferact builds and sells managed serving infrastructure around vLLM, the open-source engine from UC Berkeley used to run large language models fast and cheaply on GPUs, while continuing to fund development of the upstream open-source engine.
Starting capital:About $300K before launch (including Sequoia backing, per Forbes reporting); then a $150M seed announced 2026-01-22 at an $800M valuation, co-led by a16z and Lightspeed with Databricks Ventures and the UC Berkeley Chancellor's Fund participating.
How it started
vLLM began in 2023 inside Ion Stoica's Sky Computing Lab at UC Berkeley; its PagedAttention paper (SOSP 2023) showed that paging the attention cache like virtual memory cut memory waste and lifted serving throughput. As AI spending shifted from training models to running them, the research engine became a default way many teams served open LLMs.
What happened
By late 2025 the project had roughly 80,000 stars and thousands of production users, but the maintainers had no company around it. Forbes reported in December 2025 that co-leader Simon Mo was pitching a raise of at least $160M despite minimal revenue and only about $300K raised to date. On 2026-01-22 the core team stepped out of stealth as Inferact with a $150M seed at an $800M valuation; co-founder Woosuk Kwon framed the goal as making serving AI 'as simple as spinning up a serverless database.'
How it ended up
Inferact is live as of 2026-09-02: the founding maintainers remain in charge with Simon Mo as CEO, the round is funding a managed serverless vLLM platform plus continued upstream work, and no shutdown, pivot, or further raise has been reported.
Background
vLLM is an open-source engine for running large language models efficiently, born in 2023 in Ion Stoica's Sky Computing Lab at UC Berkeley. Its PagedAttention technique stores the attention cache in non-contiguous memory, the same trick virtual memory uses, which cut memory waste and sharply raised serving throughput.
vLLM grew from a paper into a default way many teams served open models. By 2026-09-01 an auto-updated AI-repo ranking counted about 90,600 stars with more than 2,000 contributors, and Bloomberg reported that existing users include AWS and Amazon's shopping app.
The maintainers had no commercial vehicle, and in December 2025 Forbes reported they were raising at least $160M despite minimal revenue and roughly $300K raised to date. On 2026-01-22 they launched Inferact with a $150M seed at an $800M valuation, co-led by Andreessen Horowitz and Lightspeed with Databricks Ventures and the UC Berkeley Chancellor's Fund.
Inferact's bet is that the value sits above the engine, not inside it: it plans a paid serverless vLLM that automates provisioning, autoscaling, observability, and recovery on Kubernetes while upstream vLLM stays open and keeps adding model support. As of 2026-09-02 the company is live and spending the round on that platform.
What has to be true
- Serving, not training, became AI's recurring cost, so the engine that cuts serving cost became core infrastructure.
- vLLM's installed base gives Inferact distribution no rival can buy: ~90.6k stars, 2,000+ contributors, and AWS-scale users.
- The founding team includes the original creators and their Berkeley lab, the exact people who made the engine fast.
- Keeping the core open removes the trust objection that usually kills open-source commercialization.
- Enterprises pay for reliability, autoscaling, and operations, precisely the managed layer Inferact is building.
What can be applied
Open-source adoption alone is not a business; Inferact separated a free core from a paid managed layer, and investors priced the team and adoption at $800M before any product revenue.
Aftermath
As of 2026-09-02 Inferact remains live and in build-out. The $150M seed announced on 2026-01-22 funds a managed, serverless version of vLLM running on Kubernetes with observability, troubleshooting, and disaster recovery, while the company keeps investing in the upstream open-source engine, adding performance work and new model architectures. The founding maintainers lead the company and no product launch, pivot, shutdown, or follow-on round had been reported by the as-of date.
Sources
- Inference startup Inferact lands $150M to commercialize vLLM
- Inferact launches with $150M in funding to commercialize vLLM
- AI Inference Project vLLM Seeks US$160 Million With Minimal Revenue
- Github-Ranking-AI: most-starred AI repositories, auto-updated 2026-09-01
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card