EN
Back to the archive

The archive · Hardware & Devices · Technical decision · 2016–2025

Groq's LPU bet: viral 500 tok/s demos, $640M round, Meta deal

Inference-only chip startup, founded by Google's TPU co-inventor, goes viral in 2024, raises $640M at $2.8B, and powers Meta's official Llama API.

Groq

The betThat LLM inference — not training — becomes the bottleneck, and a purpose-built LPU chip with deterministic timing can beat GPUs at speed and cost per token.Scaling

What the business is

AI inference hardware and cloud: designs the LPU (language processing unit) chip, manufactures it at Samsung's foundry, and sells fast low-latency model inference via the GroqCloud API and data-center deployments.

Starting capitalOver $1B raised; the August 2024 Series D added $640M led by BlackRock at a $2.8B valuation, with Neuberger Berman, Cisco, KDDI and Samsung Catalyst Fund participating.

How it started

Jonathan Ross helped invent Google's TPU, then co-founded Groq in 2016 with Douglas Wightman. The bet was contrarian: skip the GPU training arms race and build a language processing unit whose only job is fast, predictable LLM inference.

What happened

Groq stayed quiet for years, then hit public consciousness on 19 February 2024, when the Hacker News post 'Groq runs Mixtral 8x7B-32k with 500 T/s' drew 847 points and 472 comments. In August 2024 it raised a $640M Series D led by BlackRock at a $2.8B valuation, taking total funding past $1B; GroqCloud had 356,000 developers and the company planned to deploy 108,000 LPUs by Q1 2025.

How it ended up

Still running and scaling: in April 2025 Groq announced a partnership with Meta to power the official Llama API at up to 625 tokens/sec, with more than 1.4 million developers on its platform.

Background

Groq builds custom chips for one job: running large language models fast. Its LPU (language processing unit) is designed for inference rather than training, and the company claims existing open models run on it at 10x the speed and a tenth of the energy of conventional processors.

Founder Jonathan Ross helped invent Google's TPU before co-founding Groq in 2016 with Douglas Wightman. The company labored in relative obscurity until 19 February 2024, when the post 'Groq runs Mixtral 8x7B-32k with 500 T/s' hit the top of Hacker News with 847 points and 472 comments, making LPU speed the AI industry's demo of the week.

Speed turned into capital and customers. In August 2024 Groq raised $640M led by BlackRock at a $2.8B valuation, passing $1B raised overall, with 356,000 developers on GroqCloud and a plan for 108,000 LPUs by Q1 2025. In April 2025 it announced a partnership to power Meta's official Llama API at up to 625 tokens/sec.

The bet was that inference — not training — becomes the recurring-cost bottleneck of AI, and that a purpose-built chip with predictable latency wins that market even against Nvidia's roadmap. Groq remains private and is scaling its cloud and enterprise deployments, while Nvidia still controls the bulk of the AI-chip market it is attacking.

What has to be true

  • Inference is the volume business: every deployed model generates tokens daily, so cost per token matters more than peak training performance.
  • The viral 500 T/s demos created demand before enterprise sales existed — GroqCloud reached 356,000 developers within months.
  • A single-purpose chip kept the design simple enough to reach production with a tiny team versus GPU-platform complexity.
  • Partnering with Meta put Groq inside the official Llama API, turning an open-model champion into its anchor distribution channel.

What can be applied

Pick the bottleneck, not the market: Groq didn't try to out-train Nvidia; it bet inference speed would matter more and made speed its product — turning a chip startup into Meta's inference partner.

Aftermath

As of late April 2025, Groq remains private and scaling: the Meta Llama API partnership (up to 625 tokens/sec) went into preview with select developers, GroqCloud passed 1.4 million developers, and the company was expanding its Samsung 4nm LPU supply toward its 108,000-chip deployment target. Its competition is still Nvidia's roadmap and the cloud giants' custom silicon; the open question is whether inference-only chips keep their cost edge as GPU suppliers add low-latency lines of their own.

Sources

spotted an error? The archive wants to know.

Your turn

You just read one. Describe what you are building, and see who is betting on the same thing.

Free account · 3 free questions · no card

Related cases