The archive · Hardware & Devices · Technical decision · 2016–2025
Groq's LPU bet: viral 500 tok/s demos, $640M round, Meta deal
Inference-only chip startup, founded by Google's TPU co-inventor, goes viral in 2024, raises $640M at $2.8B, and powers Meta's official Llama API.
Groq
What the business is
AI inference hardware and cloud: designs the LPU (language processing unit) chip, manufactures it at Samsung's foundry, and sells fast low-latency model inference via the GroqCloud API and data-center deployments.
Starting capital:Over $1B raised; the August 2024 Series D added $640M led by BlackRock at a $2.8B valuation, with Neuberger Berman, Cisco, KDDI and Samsung Catalyst Fund participating.
How it started
Jonathan Ross helped invent Google's TPU, then co-founded Groq in 2016 with Douglas Wightman. The bet was contrarian: skip the GPU training arms race and build a language processing unit whose only job is fast, predictable LLM inference.
What happened
Groq stayed quiet for years, then hit public consciousness on 19 February 2024, when the Hacker News post 'Groq runs Mixtral 8x7B-32k with 500 T/s' drew 847 points and 472 comments. In August 2024 it raised a $640M Series D led by BlackRock at a $2.8B valuation, taking total funding past $1B; GroqCloud had 356,000 developers and the company planned to deploy 108,000 LPUs by Q1 2025.
How it ended up
Still running and scaling: in April 2025 Groq announced a partnership with Meta to power the official Llama API at up to 625 tokens/sec, with more than 1.4 million developers on its platform.
Background
Groq builds custom chips for one job: running large language models fast. Its LPU (language processing unit) is designed for inference rather than training, and the company claims existing open models run on it at 10x the speed and a tenth of the energy of conventional processors.
Founder Jonathan Ross helped invent Google's TPU before co-founding Groq in 2016 with Douglas Wightman. The company labored in relative obscurity until 19 February 2024, when the post 'Groq runs Mixtral 8x7B-32k with 500 T/s' hit the top of Hacker News with 847 points and 472 comments, making LPU speed the AI industry's demo of the week.
Speed turned into capital and customers. In August 2024 Groq raised $640M led by BlackRock at a $2.8B valuation, passing $1B raised overall, with 356,000 developers on GroqCloud and a plan for 108,000 LPUs by Q1 2025. In April 2025 it announced a partnership to power Meta's official Llama API at up to 625 tokens/sec.
The bet was that inference — not training — becomes the recurring-cost bottleneck of AI, and that a purpose-built chip with predictable latency wins that market even against Nvidia's roadmap. Groq remains private and is scaling its cloud and enterprise deployments, while Nvidia still controls the bulk of the AI-chip market it is attacking.
What has to be true
- Inference is the volume business: every deployed model generates tokens daily, so cost per token matters more than peak training performance.
- The viral 500 T/s demos created demand before enterprise sales existed — GroqCloud reached 356,000 developers within months.
- A single-purpose chip kept the design simple enough to reach production with a tiny team versus GPU-platform complexity.
- Partnering with Meta put Groq inside the official Llama API, turning an open-model champion into its anchor distribution channel.
What can be applied
Pick the bottleneck, not the market: Groq didn't try to out-train Nvidia; it bet inference speed would matter more and made speed its product — turning a chip startup into Meta's inference partner.
Aftermath
As of late April 2025, Groq remains private and scaling: the Meta Llama API partnership (up to 625 tokens/sec) went into preview with select developers, GroqCloud passed 1.4 million developers, and the company was expanding its Samsung 4nm LPU supply toward its 108,000-chip deployment target. Its competition is still Nvidia's roadmap and the cloud giants' custom silicon; the open question is whether inference-only chips keep their cost edge as GPU suppliers add low-latency lines of their own.
Sources
- AI chip startup Groq lands $640M to challenge Nvidia
- Groq runs Mixtral 8x7B-32k with 500 T/s
- Meta and Groq Collaborate to Deliver Fast Inference for the Official Llama API
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card