档案库 · AI 与模型 · 技术决策 · 2026
DeepGrove押注三元权重MoE,让推理AI跑在iPhone上
开源20B-A1B三元MoE(5.31GB)在Apple Silicon上跑出约120-218 token/s;HN首页173分,53条评论。
DeepGrove
做的是什么生意
DeepGrove is a US AI research lab building frontier-quality intelligence for any device; Maple-Preview is its open-source 20B-A1B ternary-weight MoE reasoning model (24 layers, 256 experts, 8 active, MIT) with an online chat demo.
起因
DeepGrove released Maple-Preview on Hugging Face under MIT in early August 2026, framed as a reasoning LLM designed from the start for efficient on-device inference: 24 layers, 256 experts (8 active), a 3:1 sliding-window-to-global attention ratio and 131,072-token context.
经过
DeepGrove reported 218 tokens/s on an M4 Mac mini (five to sixteen times faster than Gemma 4, Qwen3.5 and gpt-oss), about 120-127 tok/s on iPhone, and an IMO 2024 P1 solve of 7/7 at 281.5 tok/s on an M5 Pro MacBook Pro. The Show HN drew 173 points and 53 comments; commenters debated hallucination at two-bit sizes, and the llama.cpp project opened an issue requesting Maple support. The model card itself warns the preview received minimal post-training for agentic tasks.
结果
Still live as an open-source preview; DeepGrove is iterating toward a system that adapts models per user and conversation.
背景
DeepGrove是一家美国AI研究实验室,押注如果模型从一开始就为任何设备设计,前沿质量的智能就能跑起来。它的论点:量化现有的大模型是让本地AI工作的错误方式,因为这会降低质量,同时仍消耗内存和带宽。
Maple-Preview于2026年8月初在Hugging Face上以MIT许可发布,就是证据:一个20B参数的MoE,每个token只有约1.49B活跃,权重限制为三元值{-1,0,1},24层,256专家(8活跃),3:1滑动窗口与全局注意力比例,131072 token上下文,检查点5.31GB。
DeepGrove报告称,在M4 Mac mini上达到218 token/s,在iPhone上约120-127 token/s,并在M5 Pro MacBook Pro上以281.5 token/s解决了IMO 2024 P1的7/7——比Gemma 4、Qwen3.5和gpt-oss等高效模型快5到16倍。模型卡片警告,它面向代理任务的后期训练极少。
2026-08-04的Show HN登上首页,拿到173分和53条评论,llama.cpp项目开issue请求支持Maple。截至2026年9月,该模型仍是预览版:DeepGrove称正在迭代,并构建一个按用户和对话调整模型的系统。
这件事要成立,得有什么
- 三元权重逆转了本地AI的权衡:如果矩阵乘法变成加法,内存和带宽崩溃,而不缩小模型的推理能力。
- 在MIT下开源,将实验室声明变成生态基础设施——几天内就有llama.cpp支持请求和社区基准。
- 在人们拥有的硬件(M4 Mac mini、iPhone)上跑出数字,使速度声明变得具体,不同于排行榜表格。
- 这个赌注是结构性的,而非增量:如果5GB能装下一个推理模型,编码助手、数学导师和长上下文本地工具的经济学就会改变。
可借鉴之处
端侧AI的杠杆在于每权重的位数,而非参数数量:一个5.31GB的三元MoE能跑出交互速度,立刻引来生态关注。
后续进展
截至2026-09-02,Maple-Preview作为开源预览存活:Hugging Face上的MIT许可检查点,提供GGUF转换、社区运行时和在线聊天演示。DeepGrove表示将继续改进模型,并打造一个能适应每个对话和用户的系统;其既定方向是让前沿智能运行在任何设备上,所以这次发布是产品线的开局,而非成品。
资料来源
- Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone
- DeepGrove open-sources Maple-Preview, a 20B-A1B ternary MoE
- 苹果 iPhone 上最快本地 AI 模型:Maple-Preview-20B-A1B 登场,可达 127 tokens/s
发现哪里写错了?告诉我们。
轮到你了
你刚读完一家。说说你在做什么,看看谁在赌同一件事。
免费账号 · 3 次免费提问 · 不用绑卡