EN
Back to the archive

The archive · AI & Models · Strategic decision · 2026

PrismML bets 1-bit LLMs fit on phones; Bonsai 8B ships, Apple in talks

Caltech spin-out raises $16.25M to build 1-bit LLMs that fit on phones; Bonsai 27B runs on an iPhone and Apple is reportedly evaluating it.

PrismML

The betThat the next AI leap is intelligence density: 1-bit models can deliver near-full-precision reasoning on phones and edge devices, making on-device AI the default.Scaling

What the business is

An AI lab building 1-bit and ternary large language models — Bonsai 8B at 1.15GB and Bonsai 27B at 3.9GB — that run locally on smartphones, laptops and embedded systems.

Starting capital$16.25M in SAFE and seed funding from Khosla Ventures, Cerberus Ventures and Caltech, with compute grants from Google and Caltech (announced March 2026).

How it started

Founded by Caltech researchers around Professor Babak Hassibi after years of work on compressing neural networks without losing reasoning ability. The company emerged from stealth on 2026-03-31 with the $16.25M round and the open-source release of 1-bit Bonsai 8B (Apache 2.0), trained on Google TPUs.

What happened

Bonsai 8B hit the Hacker News front page (430 points, 153 comments) and drew broad coverage as the first commercially viable 1-bit LLM. On 2026-07-14 PrismML released Bonsai 27B, a 1-bit build of Qwen3.6 27B that occupies 3.9GB (ternary variant 5.9GB) and runs on an iPhone 17 Pro; edgen.tech reported Apple was in early talks to evaluate the technology for on-device Siri.

How it ended up

Scaling: Bonsai 27B is available under Apache 2.0 on Hugging Face, the team is hiring, and a compressed Gemma model is next in the pipeline; independent third-party benchmarks had not been published as of September 2026.

Background

PrismML is a Caltech spin-out building 1-bit and ternary large language models that run locally on phones, laptops and edge devices. Its flagship Bonsai 8B packs 8.2 billion parameters into 1.15GB — roughly 14x smaller than a 16-bit 8B model — while scoring competitively on standard benchmarks.

The bet is that the next big AI leap is intelligence density, not parameter count: if a model can run on-device, it removes the cloud latency, privacy and cost objections that keep most capable AI inside data centers. Founded by Caltech researchers around Professor Babak Hassibi, PrismML emerged from stealth on 2026-03-31 with $16.25M in SAFE and seed funding from Khosla Ventures, Cerberus Ventures and Caltech, plus Google TPU compute, and open-sourced Bonsai 8B under Apache 2.0.

The Show HN launch hit the front page with 430 points and 153 comments. On 2026-07-14 the company released Bonsai 27B, a 1-bit build of Qwen3.6 27B at 3.9GB that runs on an iPhone 17 Pro, and edgen.tech reported Apple in early talks to evaluate the compression for on-device Siri. As of September 2026 PrismML is hiring, has a Gemma compression in the pipeline, and still faces the open question of independent third-party benchmarks.

What has to be true

  • 1-bit quantization across the entire network cuts memory roughly 14x while benchmark averages stay competitive with 8B full-precision models.
  • On-device execution removes cloud latency, privacy and cost objections, opening phones, robots, wearables and secure enterprise deployments.
  • Caltech IP plus Khosla, Cerberus and Google backing gave a stealth lab credibility and TPU compute to train a full 1-bit family.
  • If Apple adopts the compression for Siri, on-device reasoning moves from developer novelty to a top OEM default overnight.

What can be applied

Pick the metric that flips the incumbent trade-off: by optimizing intelligence per GB instead of benchmark averages, PrismML turned 'runs on a phone' into a feature incumbents could not quickly copy.

Aftermath

As of 2026-09-02, PrismML has released Bonsai 27B (1-bit, 3.9GB) and 8B/4B/1.7B models, all Apache 2.0 on Hugging Face, with MLX and CUDA support. edgen.tech reported Apple in early talks to evaluate the technology for on-device Siri, and CEO Babak Hassibi told CNBC that Apple and other companies were evaluating it. The company is hiring and says a compressed Google Gemma model is next, followed by work above 27B parameters. The main caveat: benchmarks are self-reported, and independent third-party testing had not been published as of the as-of date.

Sources

spotted an error? The archive wants to know.

Your turn

You just read one. Describe what you are building, and see who is betting on the same thing.

Free account · 3 free questions · no card

Related cases