The archive · AI & Models · Product decision · 2025–2026
Cactus Compute bets tiny 26M tool-calling model can run phone AI agents
YC S25 startup distills Gemini tool-calling into a 14MB, 26M-parameter model at 6,000 tok/s — 776 HN points, 211 comments.
Cactus Compute
What the business is
Cactus Compute builds an open-source, cross-platform AI inference engine (Cactus) for phones, wearables and robots, plus Cactus Chat and the Needle model family — small tool-calling models that run fully on-device.
Starting capital:Not disclosed; company is Y Combinator Summer 2025-backed per its YC company page.
How it started
Co-founders Roman Shemet and Henry Ndubuaku met through YC co-founder matching in London; frustrated that nothing ran agentic AI on low-cost phones, they built Cactus, a cross-platform inference engine with Flutter, React Native and Kotlin bindings that runs Hugging Face models locally with cloud fallback. The company was founded in 2025 and went through YC Summer 2025.
What happened
Needle came from a bet inside the company: tool calling is retrieval-and-assembly, not reasoning, so massive models are overkill. The team pretrained a 26M-parameter 'Simple Attention Network' with no FFN layers on 200B tokens across 16 TPU v6e in 27 hours, then post-trained it for 45 minutes on Gemini-synthesized tool-calling data. The result was a 14MB model doing 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, beating FunctionGemma-270M, Qwen-0.6B, Granite-350M and LFM2.5-350M on single-shot function calling. It launched MIT-licensed on 2026-05-12, drew 776 points and 211 comments on HN, and Gigazine noted the wrinkle that Google's Gemini API terms prohibit distillation.
How it ended up
Still live and scaling: as of September 2026 the repo ships Needle 2, a 45M-parameter model in a single 14MB binary with 2-bit Cactus Quants, confidence-gated tool calls and an arXiv paper (2607.18363); YC lists an 8-person active SF team with Cactus Engine, Needle and Cactus Hybrid products.
Background
Cactus Compute is a YC Summer 2025 startup building on-device AI for phones, wearables and robots. Its founding observation: agentic AI on low-cost phones doesn't exist, and big cloud models can't run there. The company's answer was a cross-platform inference engine plus its own small models.
In May 2026 the team shipped Needle, a 26M-parameter model distilled from Gemini 3.1 Flash Lite, built on the claim that tool calling is retrieval-and-assembly rather than reasoning, so massive models are overkill for it. A no-FFN 'Simple Attention Network' pretrained on 200B tokens in 27 hours on 16 TPU v6e and post-trained for 45 minutes beat FunctionGemma-270M and Qwen-0.6B on single-shot function calling while running at 6,000 tok/s prefill.
The MIT release hit the HN front page on 2026-05-12 with 776 points and 211 comments, was covered by Gigazine, and put Cactus in front of developers while the commercial Cactus engine and Cactus Chat app remain the business. By September 2026 the project had evolved into Needle 2, a 45M-parameter model in a 14MB binary with an arXiv paper.
What has to be true
- Tool calling is retrieval-and-assembly, not reasoning: by isolating that job, Cactus showed a 26M model can beat 270M-parameter rivals on single-shot function calling — the whole bet.
- Open-source MIT release created instant distribution: 776 HN points and 211 comments in one day, while the proprietary Cactus engine stayed the commercial layer.
- Distillation made the bet cheap: 27 hours of pretraining plus a 45-minute post-train on synthesized data — though Gigazine noted Google's Gemini terms prohibit distillation.
- Cross-platform bindings (Flutter, React Native, Kotlin) aimed where phone apps are actually built, instead of betting on Apple and Google's platform-specific AI frameworks.
What can be applied
Narrow the problem until a tiny model wins: isolate tool calling from reasoning, ship a 26M model that beats 270M rivals, and turn a 776-point HN launch into distribution for the commercial engine.
Aftermath
As of 2026-09-02, Cactus Compute is live and scaling: the Needle repo ships Needle 2, a 45M-parameter model in a single 14MB binary with CQ2-bit quantization, ~28MB session memory, confidence-gated tool calls and an arXiv paper (2607.18363), and the team pitches commercial partnerships from the README. YC's directory lists an 8-person active San Francisco team with the Cactus Engine, Needle and Cactus Hybrid (cloud fallback) product line. Funding amounts are not disclosed beyond the YC Summer 2025 backing.
Sources
- Show HN: Needle: We Distilled Gemini Tool Calling into a 26M Model
- Needle, a lightweight version of Gemini's tool invocation functionality designed to run on smartphones, has been released
- Cactus Compute: Tiny Edge AI For Tiny Devices
- cactus-compute/needle
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card