EN
Back to the archive

The archive · Developer & Business Tools · Strategic decision · 2023–2026

RadixArk's SGLang bet: 33k GitHub stars, $100M seed at $400M

RadixArk bets open-source inference wins AI: SGLang's radix-tree caching cut LLM serving cost, hit 33k GitHub stars, and drew a $100M Accel-led seed in 2026.

RadixArk (SGLang)

The betOpen-source inference wins: SGLang's radix-tree prefix cache cuts serving cost, and RadixArk monetizes managed hosting while keeping the engine free.Scaling

What the business is

RadixArk commercializes SGLang, an open-source high-performance serving framework for LLM and multimodal models — free engine, paid managed hosting.

Starting capital$100M seed at $400M post-money, announced 2026-05-05; led by Accel with Spark Capital; NVentures, AMD, MediaTek, Databricks and others joined; angels include John Schulman, Soumith Chintala and Intel CEO Lip-Bu Tan.

How it started

In 2023 Ying Sheng — then xAI's inference lead — and collaborators created SGLang inside Berkeley's LMSYS research group, and the open-source engine spread through production users like Google, Microsoft and NVIDIA. In 2025 Sheng and Banghua Zhu (ex-NVIDIA) founded RadixArk in Palo Alto to bring it to market.

What happened

SGLang became the open-source inference default: deployed on 400k+ GPUs with trillions of tokens daily, and known for day-0 support of new model architectures — DeepSeek-V4 was served and RL-trained the day it launched in April 2026. In November 2025 the team also open-sourced Miles for large-scale RL training. On 2026-05-05 RadixArk launched with a $100M seed at a $400M valuation, with NVIDIA, AMD, MediaTek, Broadcom and Intel all in the round.

How it ended up

Still scaling: the seed funds support for more model types and hardware plus the managed platform at platform.radixark.com; SGLang remains Apache-2.0 open source under the sgl-project org.

Background

SGLang is an open-source serving framework for large language and multimodal models, created in 2023 inside Berkeley's LMSYS research group by Ying Sheng and collaborators. Its signature idea is the radix tree: shared prompt prefixes are computed once and reused, instead of recomputing context for every query, which cuts the per-token cost of chat and agent workloads.

The engine spread through technical merit alone — no marketing or sales team — and now serves trillions of tokens a day for Google, Microsoft, NVIDIA, xAI and others across more than 400,000 GPUs. It became known for day-0 support of new model architectures, including same-day inference and RL training for DeepSeek-V4 in April 2026, and its GitHub stars grew from 27k in May 2026 to 33k by September 2026, with repeated appearances on GitHub trending.

In 2025 Sheng and Banghua Zhu (ex-NVIDIA) founded RadixArk in Palo Alto to commercialize the project. On May 5, 2026 the company launched with a $100 million seed round at a $400 million valuation, led by Accel with Spark Capital and joined by NVentures, AMD, MediaTek, Databricks and more — plus angels including OpenAI co-founder John Schulman, PyTorch creator Soumith Chintala and Intel CEO Lip-Bu Tan.

The model is familiar from Databricks and Elastic: keep the engine open and free, and charge for managed hosting. RadixArk's stated mission is to make frontier AI infrastructure open and accessible, arguing that the next generation of AI will be built on shared systems rather than whoever owns the biggest private cluster.

What has to be true

  • The radix-tree prefix cache attacked the single biggest recurring cost in LLM serving — redundant context recomputation — giving SGLang a real technical moat.
  • Serving the same workloads for Google, Microsoft and NVIDIA gave the project production credibility that no marketing budget could buy.
  • Five major hardware makers — NVIDIA, AMD, MediaTek, Broadcom, Intel — all backed a hardware-agnostic open engine, betting on a neutral layer above their chips.
  • The open-source, managed-hosting model (like Databricks and Elastic) let RadixArk monetize adoption without breaking the community's trust.

What can be applied

A systems bet on the most expensive part of AI — inference — compounds quietly: make it cheaper for everyone, and the ecosystem adopts you before the market ever needs a pitch deck.

Aftermath

As of 2026-09-02 RadixArk is scaling: the seed funds expansion to more model types and hardware plus its managed platform at platform.radixark.com, and the team is hiring across systems, compilers and scheduling. SGLang remains Apache-2.0 under the sgl-project organization on GitHub with roughly 33k stars, and its competitors include vLLM, another Berkeley-born open-source engine that also turned into a funded startup — evidence that inference-layer consolidation is the open-source battleground of this AI cycle.

Sources

spotted an error? The archive wants to know.

Your turn

You just read one. Describe what you are building, and see who is betting on the same thing.

Free account · 3 free questions · no card

Related cases