档案库 · AI 与模型 · 产品决策 · 2025–2026
Cactus Compute押注2600万参数小模型可运行手机AI代理
YC S25初创公司将Gemini工具调用蒸馏为14MB、2600万参数模型,速度达6000 tok/s——HN获776分、211条评论。
Cactus Compute
做的是什么生意
Cactus Compute builds an open-source, cross-platform AI inference engine (Cactus) for phones, wearables and robots, plus Cactus Chat and the Needle model family — small tool-calling models that run fully on-device.
启动资金:Not disclosed; company is Y Combinator Summer 2025-backed per its YC company page.
起因
Co-founders Roman Shemet and Henry Ndubuaku met through YC co-founder matching in London; frustrated that nothing ran agentic AI on low-cost phones, they built Cactus, a cross-platform inference engine with Flutter, React Native and Kotlin bindings that runs Hugging Face models locally with cloud fallback. The company was founded in 2025 and went through YC Summer 2025.
经过
Needle came from a bet inside the company: tool calling is retrieval-and-assembly, not reasoning, so massive models are overkill. The team pretrained a 26M-parameter 'Simple Attention Network' with no FFN layers on 200B tokens across 16 TPU v6e in 27 hours, then post-trained it for 45 minutes on Gemini-synthesized tool-calling data. The result was a 14MB model doing 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices, beating FunctionGemma-270M, Qwen-0.6B, Granite-350M and LFM2.5-350M on single-shot function calling. It launched MIT-licensed on 2026-05-12, drew 776 points and 211 comments on HN, and Gigazine noted the wrinkle that Google's Gemini API terms prohibit distillation.
结果
Still live and scaling: as of September 2026 the repo ships Needle 2, a 45M-parameter model in a single 14MB binary with 2-bit Cactus Quants, confidence-gated tool calls and an arXiv paper (2607.18363); YC lists an 8-person active SF team with Cactus Engine, Needle and Cactus Hybrid products.
背景
Cactus Compute是一家YC Summer 2025初创公司,为手机、可穿戴设备和机器人构建设备端AI。其创见:低成本手机上不存在代理式AI,而云端大模型无法在那里运行。公司的解决方案是跨平台推理引擎加自有小模型。
2026年5月,团队发布了Needle——一个从Gemini 3.1 Flash Lite蒸馏出的2600万参数模型,基于“工具调用是检索与组装,而非推理,因此不需要大规模模型”的论断。一个无FFN的“简单注意力网络”,在16个TPU v6e上预训练200B tokens用时27小时,后训练45分钟,在单次函数调用上击败FunctionGemma-270M和Qwen-0.6B,同时预填充速度达6000 tok/s。
MIT许可的发布于2026年5月12日登上HN首页,获得776分和211条评论,被Gigazine报道,让Cactus在开发者中获得曝光,同时商业Cactus引擎和Cactus Chat应用仍是业务核心。到2026年9月,项目演变为Needle 2——一个4500万参数的模型,打包为14MB二进制文件,并有配套arXiv论文。
这件事要成立,得有什么
- 工具调用是检索与组装,不是推理:通过隔离这一工作,Cactus证明了2600万参数的模型在单次函数调用上能击败2.7亿参数的对手——这是核心赌注。
- MIT开源发布带来了即时分发:一天内HN获776分和211条评论,而专有Cactus引擎仍是商业层。
- 蒸馏使赌注成本低廉:27小时预训练加45分钟后训练,合成数据——尽管Gigazine指出谷歌Gemini条款禁止蒸馏。
- 跨平台绑定(Flutter、React Native、Kotlin)瞄准了手机应用的实际构建方式,而非押注苹果和谷歌的平台特定AI框架。
可借鉴之处
将问题缩小至小模型能胜出:将工具调用与推理分离,发布击败270M对手的26M模型,并把HN 776分的发布转化为商业引擎的分发渠道。
后续进展
截至2026年9月2日,Cactus Compute已上线并扩展:Needle仓库发布Needle 2——4500万参数模型,单14MB二进制文件,带CQ2位量化、约28MB会话内存、置信门控工具调用和arXiv论文(2607.18363),团队在README中推广商业合作。YC目录列出8人活跃旧金山团队,产品线包括Cactus Engine、Needle和Cactus Hybrid(云回退)。除YC Summer 2025支持外,融资金额未披露。
资料来源
- Show HN: Needle: We Distilled Gemini Tool Calling into a 26M Model
- Needle, a lightweight version of Gemini's tool invocation functionality designed to run on smartphones, has been released
- Cactus Compute: Tiny Edge AI For Tiny Devices
- cactus-compute/needle
发现哪里写错了?告诉我们。
轮到你了
你刚读完一家。说说你在做什么,看看谁在赌同一件事。
免费账号 · 3 次免费提问 · 不用绑卡