The archive · Hardware & Devices · Technical decision · 2019–2026
d-Matrix's SRAM inference bet: $275M round, Corsair into volume production
Corsair runs LLM decode in on-chip SRAM instead of GPUs; a $275M Series C at $2B funded the ramp to full production in June 2026.
d-Matrix
What the business is
d-Matrix, founded in Santa Clara in 2019 by Sid Sheth and Sudeep Bhoja, builds a full-stack AI inference platform: Corsair accelerator cards based on SRAM digital in-memory computing, JetStream networking, Aviator software, and SquadRack rack-scale systems. Corsair is manufactured on TSMC's N6 process by Alchip, uses LPDDR5 instead of HBM, and is positioned to run the decode phase of LLM inference alongside GPU clusters. Microsoft's M12 venture fund is an investor.
How it started
d-Matrix was founded in 2019 by Sid Sheth and Sudeep Bhoja to build hardware for what they expected to follow the training boom: running generative models continuously at scale. When it exited stealth in November 2024 it had raised $154M from more than 25 investors, with Temasek leading the Series B and Microsoft's M12 among the backers, and was sampling Corsair to early-access customers with broad availability planned for Q2 2025.
What happened
The company spent 2025 converting claims into engineering disclosures and capital. At Hot Chips in August 2025, ServeTheHome detailed the Corsair board: two chips of four chiplets on TSMC N6, 2GB of SRAM, 256GB of LPDDR5X per card, PCIe 5.0, 38 TOPS per watt, and roughly 2ms per output token on Llama 3 70B. On 2025-11-12 d-Matrix closed a $275M Series C at a $2B valuation, bringing total funding to $450M, with Temasek, Bullhound Capital and Triatomic Capital co-leading and Qatar Investment Authority, EDBI and M12 participating, per technode.global. In April 2026 it acquired GigaIO's data center business for rack-scale integration, and on 2026-06-09 Converge Digest reported that Corsair had entered full production with volume shipments to hyperscalers, neoclouds and frontier AI labs starting that summer.
How it ended up
Full production separates d-Matrix from the wave of inference-chip startups that stopped at benchmarks: Corsair is shipping in volume to unnamed priority customers, SquadRack systems ship alongside it, and supply agreements cover the ramp. The company cites independent testing showing that pairing Corsair with GPUs cut speculative-decoding response times from about 24 seconds to under two seconds, but customer names, prices and revenue remain undisclosed, so the commercial test is still ahead.
Background
d-Matrix, founded in Santa Clara in 2019 by Sid Sheth and Sudeep Bhoja, designs silicon for the part of AI that runs after training: serving models. Its Corsair accelerator puts matrix math inside on-chip SRAM — digital in-memory computing — rather than fetching weights from HBM, and pairs that with LPDDR5 capacity memory, its own networking, and Aviator software.
The company exited stealth in November 2024 claiming 30,000 tokens per second at 2ms per token on Llama 3 70B from a single rack, then disclosed the full architecture at Hot Chips in August 2025: TSMC N6 chiplets, 2GB of SRAM per board, 256GB of LPDDR5X, PCIe 5.0, and 38 TOPS per watt. In November 2025 it closed a $275M Series C at a $2B valuation, bringing total funding to $450M, with Temasek, Bullhound Capital and Triatomic Capital co-leading and Qatar Investment Authority, EDBI and Microsoft's M12 participating.
On June 9, 2026 d-Matrix announced Corsair had entered full production, with volume shipments to hyperscalers, neoclouds and frontier AI labs starting that summer. Its pitch is complementary rather than adversarial: GPUs handle prefill while Corsair executes decode, and independent testing cited by the company cut speculative-decoding response times from about 24 seconds to under two seconds.
What has to be true
- Every output token is memory-bound, so a compute-in-memory part that cuts weight-fetch traffic attacks the real cost of serving LLMs rather than peak training flops.
- The design sidesteps scarce supply chains: SRAM chiplets with LPDDR5 need no HBM or CoWoS packaging, removing the bottleneck that GPU rivals wait in line for.
- Complementing GPUs made d-Matrix a low-risk buy: with prefill left to GPUs, independent testing cited by the company showed response times falling from about 24 seconds to under two.
- Proof accumulated in checkable steps: Hot Chips disclosures in August 2025, a $275M Series C at $2B in November 2025, and full production in June 2026.
What can be applied
Sell into the bottleneck, not the flagship: leave training to Nvidia, argue decode is memory-bound, and build an SRAM accelerator that plugs into GPU clusters — a wedge that reached full production.
Aftermath
As of June 2026 Corsair is in full production with volume shipments beginning that summer to unnamed hyperscalers, neoclouds and frontier AI labs, with SquadRack rack systems shipping alongside. The risk is market-side: Nvidia keeps accelerating decode in next-gen GPUs, hyperscalers keep building in-house silicon, and d-Matrix has disclosed neither customer names nor revenue. The GigaIO acquisition and Alchip supply agreements cover integration and capacity; whether measured latency wins convert into repeat orders is the open test.
Sources
- d-Matrix Emerges From Stealth With Strong AI Performance And Efficiency
- d-Matrix Corsair In-Memory Computing For AI Inference at Hot Chips 2025
- Temasek backs US AI chipmaker d-Matrix's $275M funding
- d-Matrix Ramps Corsair AI Inference Platform
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card