档案库 · 硬件与设备 · 技术决策 · 2019–2026
d-Matrix 的 SRAM 推理押注:2.75 亿美元融资,Corsair 进入量产
Corsair 在片上 SRAM 中运行 LLM 解码,而非 GPU;2 亿美元估值下的 2.75 亿美元 C 轮融资助力其在 2026 年 6 月全面量产。
d-Matrix
做的是什么生意
d-Matrix, founded in Santa Clara in 2019 by Sid Sheth and Sudeep Bhoja, builds a full-stack AI inference platform: Corsair accelerator cards based on SRAM digital in-memory computing, JetStream networking, Aviator software, and SquadRack rack-scale systems. Corsair is manufactured on TSMC's N6 process by Alchip, uses LPDDR5 instead of HBM, and is positioned to run the decode phase of LLM inference alongside GPU clusters. Microsoft's M12 venture fund is an investor.
起因
d-Matrix was founded in 2019 by Sid Sheth and Sudeep Bhoja to build hardware for what they expected to follow the training boom: running generative models continuously at scale. When it exited stealth in November 2024 it had raised $154M from more than 25 investors, with Temasek leading the Series B and Microsoft's M12 among the backers, and was sampling Corsair to early-access customers with broad availability planned for Q2 2025.
经过
The company spent 2025 converting claims into engineering disclosures and capital. At Hot Chips in August 2025, ServeTheHome detailed the Corsair board: two chips of four chiplets on TSMC N6, 2GB of SRAM, 256GB of LPDDR5X per card, PCIe 5.0, 38 TOPS per watt, and roughly 2ms per output token on Llama 3 70B. On 2025-11-12 d-Matrix closed a $275M Series C at a $2B valuation, bringing total funding to $450M, with Temasek, Bullhound Capital and Triatomic Capital co-leading and Qatar Investment Authority, EDBI and M12 participating, per technode.global. In April 2026 it acquired GigaIO's data center business for rack-scale integration, and on 2026-06-09 Converge Digest reported that Corsair had entered full production with volume shipments to hyperscalers, neoclouds and frontier AI labs starting that summer.
结果
Full production separates d-Matrix from the wave of inference-chip startups that stopped at benchmarks: Corsair is shipping in volume to unnamed priority customers, SquadRack systems ship alongside it, and supply agreements cover the ramp. The company cites independent testing showing that pairing Corsair with GPUs cut speculative-decoding response times from about 24 seconds to under two seconds, but customer names, prices and revenue remain undisclosed, so the commercial test is still ahead.
背景
d-Matrix 于 2019 年由 Sid Sheth 和 Sudeep Bhoja 在圣克拉拉创立,设计用于 AI 训练后运行阶段的芯片:模型服务。其 Corsair 加速器将矩阵运算置于片上 SRAM 中——即数字存内计算——而不是从 HBM 中获取权重,并搭配 LPDDR5 大容量内存、自研网络和 Aviator 软件。
公司于 2024 年 11 月脱离 stealth 模式,声称单个机架在 Llama 3 70B 上每秒可处理 30,000 个 token,每个 token 2 毫秒,随后于 2025 年 8 月在 Hot Chips 上披露了完整架构:台积电 N6 小芯片、每板 2GB SRAM、256GB LPDDR5X、PCIe 5.0 和每瓦 38 TOPS。2025 年 11 月,它以 20 亿美元估值完成 2.75 亿美元 C 轮融资,总融资额达 4.5 亿美元,Temasek、Bullhound Capital 和 Triatomic Capital 共同领投,Qatar Investment Authority、EDBI 和 Microsoft 的 M12 参投。
2026 年 6 月 9 日,d-Matrix 宣布 Corsair 进入全面量产,当年夏天开始向超大规模云、新型云和前沿 AI 实验室批量出货。其定位是互补而非对抗:GPU 处理预填充,Corsair 执行解码,公司引用的独立测试显示,推测解码的响应时间从约 24 秒降至不到 2 秒。
这件事要成立,得有什么
- 每个输出 token 都受内存限制,因此存内计算部件减少权重获取流量,直击 LLM 服务的真实成本,而非峰值训练算力。
- 设计绕过了稀缺的供应链:带有 LPDDR5 的 SRAM 小芯片无需 HBM 或 CoWoS 封装,消除了 GPU 竞争对手排队的瓶颈。
- 与 GPU 互补使 d-Matrix 成为低风险选择:由于预填充由 GPU 承担,公司引用的独立测试显示,响应时间从约 24 秒降至不到 2 秒。
- 证据以可验证的步骤积累:2025 年 8 月的 Hot Chips 细节披露、2025 年 11 月 2 亿美元估值的 2.75 亿美元 C 轮融资,以及 2026 年 6 月的全面量产。
可借鉴之处
切入瓶颈环节,而不是旗舰领域:将训练留给 Nvidia,主张解码受内存限制,并构建一个可插入 GPU 集群的 SRAM 加速器——这个楔子已进入全面量产。
后续进展
截至 2026 年 6 月,Corsair 已全面量产,当年夏天开始向未具名的超大规模云、新型云和前沿 AI 实验室批量出货,SquadRack 机架系统同步销售。风险在市场端:Nvidia 继续加速下一代 GPU 的解码性能,超大规模云继续自研芯片,而 d-Matrix 未披露客户名称或收入。GigaIO 收购和 Alchip 供应协议涵盖了集成和产能;测量到的延迟优势能否转化为重复订单,仍是待检验的问题。
资料来源
- d-Matrix Emerges From Stealth With Strong AI Performance And Efficiency
- d-Matrix Corsair In-Memory Computing For AI Inference at Hot Chips 2025
- Temasek backs US AI chipmaker d-Matrix's $275M funding
- d-Matrix Ramps Corsair AI Inference Platform
发现哪里写错了?告诉我们。
轮到你了
你刚读完一家。说说你在做什么,看看谁在赌同一件事。
免费账号 · 3 次免费提问 · 不用绑卡