The archive · Hardware & Devices · Technical decision · 2023–2026
Positron attacks Nvidia where GPUs waste memory, raises $230M before its ASIC ships
Founded by ex-Groq engineers, Positron built Atlas around memory bandwidth; Jump Trading measured 3x lower latency, then co-led its $230M Series B.
Positron AI
What the business is
Positron AI builds energy-efficient hardware and software for running AI models, selling an inference appliance called Atlas that plugs into cloud or on-premise data centers.
Starting capital:$23.5M seed by early 2025, then a $75M Series A in June 2025, then a $230M Series B at over $1B post-money in Feb 2026 — just over $300M total (TechCrunch; Business Wire).
How it started
Thomas Sohmers and Edward Kmett, both alumni of inference startup Groq, founded Positron in April 2023 on one observation: transformer inference is limited by memory bandwidth and capacity, not compute, and GPU architectures waste most of both. They chose FPGAs so they could iterate with paying customers before committing to an ASIC.
What happened
Atlas began shipping in summer 2024, and EE Times reported first deliveries of a multi-million-dollar order in February 2025. Positron raised a $75M Series A in June 2025 and began designing Asimov, its own custom silicon. On Feb 4, 2026 it announced a $230M Series B at over $1B valuation, co-led by ARENA Private Wealth, Jump Trading and Unless, with QIA, Arm and Helena joining and existing investors Valor Equity Partners, Atreides Management and DFJ Growth participating.
How it ended up
As of Feb 2026 Positron is selling Atlas systems built in the US while racing to tape out Asimov by late 2026 and reach production in early 2027 — a bet that a memory-first chip can beat Nvidia's Rubin on tokens per watt and performance per dollar.
Background
Thomas Sohmers and Edward Kmett, both veterans of inference chip startup Groq, founded Positron in April 2023 on a contrarian technical claim: running modern language models is bound by memory bandwidth and capacity, not raw compute, and Nvidia GPUs use under 30% of their theoretical memory bandwidth on transformer workloads. Instead of trying to out-FLOPS Nvidia, they built for the constraint that actually throttles inference.
They chose FPGAs over a from-scratch ASIC so they could ship and iterate with real customers first. Atlas, their turnkey inference appliance, began shipping in summer 2024; by February 2025 EE Times reported a multi-million-dollar order being delivered to a Tier 2 cloud provider, with about 20 more customers evaluating it. Positron claimed 70% faster tokens-per-second than Hopper-class systems at 3.5x the performance per watt and per dollar.
The traction convinced investors twice. After a $75M Series A in June 2025 funded the design of its own silicon, Asimov, Positron announced an oversubscribed $230M Series B on Feb 4, 2026 at a valuation above $1B. Jump Trading, which came in as a customer and measured roughly 3x lower end-to-end latency than a comparable H100 system on its inference workloads, stepped up to co-lead alongside ARENA Private Wealth and Unless.
Positron's next chip, Asimov, is designed to carry over 2,304 GB of RAM per device versus 384 GB on Nvidia's upcoming Rubin, with 5x more tokens per watt claimed for core workloads. Tape-out is targeted for late 2026 and production for early 2027 — about 16 months after the Series A that funded the design, a cadence the company treats as its main defense against Nvidia's release frequency.
What has to be true
- Transformer inference is memory-bound: GPUs sustain under 30% of their theoretical memory bandwidth on it, while Positron's FPGA design sustains 93% across use cases (EE Times).
- FPGAs let the startup prove demand cheaply — Atlas shipped to paying cloud customers before Positron committed the years and capital an ASIC requires (EE Times).
- A customer became the proof: Jump Trading measured roughly 3x lower end-to-end latency than an H100 system on its workloads, then stepped up to co-lead the round (Business Wire).
- Energy is the next bottleneck in AI deployment, and Atlas claims H100-class performance at under a third of the power, so inference buyers have a reason to switch (TechCrunch).
What can be applied
Attack incumbents where they are weak: Positron targeted memory bandwidth, where GPUs waste most capacity, and won paying customers before its own ASIC existed.
Aftermath
After the Feb 4, 2026 announcement, Positron said Atlas was shipping to frontier customers across cloud and advanced computing and that it expected large-scale commercial traction within about 2.5 years of launch. The next test is Asimov: tape-out targeted for late 2026 and production in early 2027, with 2 TB of memory per accelerator and 8 TB per Titan system at rack-scale memory capacity over 100 TB. As of early February 2026 nothing had been taped out — the $1B valuation rests on Atlas's shipping traction and Jump Trading's customer-side benchmark, not on Asimov silicon in hand.
Sources
- Exclusive: Positron raises $230M Series B to take on Nvidia's AI chips
- Positron AI raises $230M at over $1B valuation to build energy-efficient AI accelerator hardware
- Positron AI Raises $230 Million Series B at Over $1 Billion Valuation to Scale Energy-Efficient AI Inference
- Startup Positron Takes On Nvidia With FPGAs
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card