EN
Back to the archive

The archive · Hardware & Devices · Technical decision · 2023–2026

Taalas etches model weights into silicon; $219M raised, AMD acquires in 2026

Taalas bets hardwiring a model into silicon — no HBM reads — wins AI inference: 16,960 tok/s demo, $219M, then AMD acquisition.

Taalas

The betThat the model should be the computer: etch weights into chip logic as mask ROM so token generation never fetches from HBM — making fixed models 48x faster than GPUs.No longer exists

What the business is

Taalas builds model-specific AI chips ('Hardcore Models') that etch an LLM's weights directly into silicon, so inference skips the memory fetches that slow general-purpose GPUs.

Starting capital$219M total raised by Feb 2026 — a $169M round from Quiet Capital, Fidelity and Pierre Lamond plus earlier funding; AMD did not disclose deal terms.

How it started

Ljubisa Bajic, founder of AI-chip maker Tenstorrent, started Taalas in Toronto in 2023 with COO Lejla Bajic on the premise that 'the model is the computer': simulate nothing, hardwire the model. Its first TSMC 6nm test chip ran Llama 3.1 8B at a claimed 16,960 tokens/sec, about 48x an Nvidia GPU of the era.

What happened

Taalas emerged from stealth and in Feb 2026 raised $169M (about $219M total from Fidelity, Quiet Capital and Pierre Lamond), unveiling HC1, a chip optimized for Llama 3.1 8B, with a second-generation chip for roughly 20B-parameter models planned for 2026. The tradeoff — one model per chip, frozen at manufacturing — was the entire point.

How it ended up

On 2026-08-06 AMD announced a definitive agreement to acquire Taalas; terms were undisclosed. AMD plans to pair Taalas silicon with Instinct GPUs and Helios rackscale systems under its ROCm software, expecting the deal to close in Q4 2026 — framed as a business acquisition, not an acqui-hire.

Background

Taalas is a Toronto chip startup that builds 'Hardcore Models': AI inference chips in which a specific model's weights are etched into silicon as mask ROM during fabrication. Because weights never move from HBM to a processor, token generation skips the memory bottleneck that dominates modern inference cost.

Founder Ljubisa Bajic, who previously created AI-chip company Tenstorrent, started Taalas in 2023 with COO Lejla Bajic on the premise that 'the model is the computer.' A first TSMC 6nm test chip ran Llama 3.1 8B at a claimed 16,960 tokens/sec — about 48x an Nvidia GPU and 8.5x a Cerebras accelerator of the day, per the company's claims. Taalas raised $169M in Feb 2026 (about $219M total) from Fidelity, Quiet Capital and Pierre Lamond, and unveiled its HC1 chip optimized for Llama 3.1 8B.

The tradeoff was deliberate: each chip is frozen to one model at manufacturing, so it wins only where a single model serves massive real-time traffic. On 2026-08-06 AMD announced a definitive agreement to acquire Taalas (terms undisclosed), planning to integrate the technology with Instinct GPUs and Helios rackscale systems under ROCm, with closing expected in Q4 2026.

What has to be true

  • Etching weights into mask ROM eliminates the HBM weight-fetch round-trip that caps GPU inference speed — an order-of-magnitude architectural difference rather than an optimization.
  • The demo-first strategy — one chip for one model, Llama 3.1 8B at 16,960 tok/s — made the value visible and turned the HN thread into a 724-comment debate.
  • Raising $219M before shipping hardware kept the company alive through tape-outs while the bet matured.
  • The one-model-per-chip constraint limited the market to a handful of giant workloads, making acquisition by a platform owner the natural exit.

What can be applied

Betting against the industry's standard bottleneck — memory, not compute — can define a startup, but a chip frozen to one model only pays off when a platform owner buys it.

Aftermath

As of 2026-09-02, Taalas is being folded into AMD: the acquisition was expected to close in Q4 2026 pending regulatory approvals, with AMD planning to route prompt processing through Instinct GPUs and token generation through Taalas silicon in Helios rackscale systems. The startup's independent roadmap — a second-generation chip for roughly 20B-parameter models — continues inside AMD's accelerator roadmap. Deal terms were not disclosed.

Sources

spotted an error? The archive wants to know.

Your turn

You just read one. Describe what you are building, and see who is betting on the same thing.

Free account · 3 free questions · no card

Related cases