EN
Back to the archive

The archive · AI & Models · Strategic decision · 2022–2026

Together AI's open-source cloud bet: $800M at $8.3B as neoclouds surge

Together AI bets enterprises choose open-source models on a fast GPU cloud over closed frontier APIs — $1.15B bookings, $800M Series C at $8.3B by July 2026.

Together AI

The betThat enterprises will run AI on open-source models on fast, cheap GPU clouds rather than closed frontier APIs — build the open-model neocloud and win inference.Scaling

What the business is

Together AI is an AI neocloud that rents Nvidia GPU clusters and serves 200+ open-source models for inference, training and fine-tuning, with proprietary kernels and a research lab; it claims over $1.15B annual bookings, 450,000+ developers and thousands of paying customers including Cursor, Cognition and Decagon.

Starting capital$102.5M Series A (2023, led by Kleiner Perkins), $305M Series B (February 2025, led by General Catalyst and co-led by Prosperity7), $800M Series C (July 2026, led by Aramco Ventures) at an $8.3B valuation — over $1.2B total.

How it started

Together AI was founded in 2022 by Vipul Ved Prakash (who sold Topsy to Apple for a reported $200M+), Stanford professor Percy Liang and Ce Zhang, raising a $102.5M Kleiner Perkins-led Series A in 2023. Its thesis: open-source models plus optimized infrastructure, not closed frontier APIs, would power most production AI.

What happened

In February 2025 the company raised a $305M Series B led by General Catalyst and Prosperity7, claiming 450,000+ developers and customers including Salesforce, Zoom, Cognition and The Washington Post, and deployed DeepSeek models with opt-out privacy controls. In July 2026 it raised an $800M Series C led by Aramco Ventures at an $8.3B valuation, reporting annual bookings above $1.15B and naming Cursor, Cognition and Decagon among thousands of paying customers.

How it ended up

Together AI's bet is validated by market position but still contested: it says open-source model usage tripled industry-wide in the past year as companies trade frontier-API premiums for open weights, yet hyperscalers and rival neoclouds (Upscale, TensorWave, Nscale, Volta) are raising record sums into the same thesis — the question is whether any neocloud builds a moat beyond GPUs and kernels.

Background

Together AI, founded in 2022 by Vipul Ved Prakash (who sold Topsy to Apple for a reported $200M+), Stanford professor Percy Liang and Ce Zhang, is an AI neocloud that rents Nvidia GPU clusters and serves more than 200 open-source models for inference, training and fine-tuning. Its thesis is that open weights plus optimized infrastructure — not closed frontier APIs — would power most production AI.

The company raised a $102.5M Kleiner Perkins-led Series A in 2023 and a $305M Series B in February 2025 led by General Catalyst and Prosperity7, claiming 450,000+ developers and customers including Salesforce, Zoom, Cognition and The Washington Post. Its pitch: proprietary inference engines and FlashAttention-class kernels deliver 2-3x faster inference than hyperscaler solutions, making open models the cheap, fast choice.

On July 1, 2026 Together AI announced an $800M Series C led by Aramco Ventures at an $8.3B valuation, reporting annual bookings above $1.15B and thousands of paying customers including Cursor, Cognition and Decagon. The company says open-source model usage tripled industry-wide over the past year, citing OpenRouter — the shift it bet on becoming measurable market data.

The bet is validated but contested: hyperscalers still push closed frontier models, and rival neoclouds (Upscale, TensorWave, Nscale, Volta) are raising record sums into the same thesis. Together's moat is speed and cost on open models — kernels, clusters and customer trust — rather than proprietary weights.

What has to be true

  • Together bet the future of AI inference is open-source models on purpose-built GPU clouds, not closed frontier APIs — a call that looked contrarian against the 2023 closed-model wave.
  • It differentiated on speed and cost: proprietary kernels and inference engines with FlashAttention heritage claim 2-3x faster inference than hyperscaler alternatives.
  • The $1.15B bookings figure and thousands of paying customers (Cursor, Cognition, Decagon) show the open-model shift became real demand, not just developer sentiment.
  • Raising $800M at $8.3B from Aramco, Vista and Nvidia in 2026 turned the thesis into one of the best-capitalized neoclouds — while rivals raise record sums into the same bet.

What can be applied

When the open-source alternative reaches parity, sell the commodity: Together won by making open models faster and cheaper than closed APIs — and monetizing the shift, not fighting it.

Aftermath

As of September 2026 Together AI is scaling: $1.15B+ annual bookings, an $800M Series C at $8.3B (July 2026), and thousands of paying customers from Cursor and Cognition to Decagon. The open question is durability: GPU clusters and kernels are replicable, hyperscalers keep cutting open-model prices, and rival neoclouds raised $1B+ in the same months, so Together must keep its speed/cost edge and research pipeline (Tri Dao's kernel work) to justify the valuation.

Sources

spotted an error? The archive wants to know.

Your turn

You just read one. Describe what you are building, and see who is betting on the same thing.

Free account · 3 free questions · no card

Related cases