EN
返回档案库

档案库 · 开发与企业工具 · 战略决策 · 2025–2026

Oqoqo押注产品必须证明自己适合智能体——发布即登Product Hunt榜首

Oqoqo销售私有评测和定制基准,用以衡量AI智能体能否真正用上产品——发布当天以338票成为Product Hunt当日第一。

Oqoqo

它在赌什么每款产品现在都有第二个安静的用户:智能体。向这些公司销售测量层——针对他们自身真实工作流的私有评测和定制基准。已上线

做的是什么生意

Oqoqo runs real-world agentic tasks in sandboxed environments across models, agents and product versions, then reports pass rates, token spend and full step-by-step trajectories, so teams building agent-facing products can tell whether agents actually succeed on the workflows that matter to them.

起因

In November 2025 Bee Partners led Oqoqo's pre-seed out of Bee IV. Co-founders Renzo Viale and Haritha Nair, both Berkeley Haas, had started from the premise that developer documentation written for humans is the wrong substrate for coding agents; working alongside DevRel teams, they kept asking how anyone knows the docs and skills they ship actually help an agent. The answers were anecdotal, so they built experiments: running dev-tool products across different agents, with and without skills, each trial in its own sandbox. The results were unflattering — some skills made agent behavior worse, some inflated token consumption, and agents often worked around the product entirely.

经过

The experiment surface proved more valuable than the docs, so Viale and Nair opened it up. Oqoqo turns a real task — project state, files, data, tools and services plus a rubric defining success — into reproducible runs across models, agents and product versions, each trial in its own sandbox, and reports pass or fail with reasons, full trajectories, token spend and the points where an agent gave up on the product. The founders position it against public benchmarks, which they say rank models in curated environments that do not translate to the real world. The product launched on Product Hunt on 2026-08-10, taking #1 Product of the Day with 338 upvotes and #9 of the week in the developer-tools list, with the site live and free to try at oqoqo.ai.

结果

As of mid-August 2026 Oqoqo is live and pre-seed-funded: its homepage supports running real tasks across agents including Claude Code, Codex and Cursor, comparing a raw agent against MCP and tooling treatments, and reviewing trajectories and frictions with CI integration. No revenue, customer names or follow-on round are public yet.

背景

2025年11月,Bee Partners领投了Oqoqo的种子前轮,来自Bee IV基金。联合创始人Renzo Viale和Haritha Nair都毕业于加州大学伯克利分校哈斯商学院,他们最初的出发点是:给人类写的开发者文档不是编程智能体的正确载体,任何为开发者服务的公司都得重写文档,让智能体在无人介入的情况下也能使用。

在与DevRel团队合作时,他们反复追问一个没人答得上来的问题:你如何知道发布的文档和技能确实在帮助智能体?他们搭建了实验——在沙盒中跨智能体运行开发工具产品,有的带技能,有的不带——结果并不好看:有些技能让智能体行为更糟,有些推高token消耗,而且智能体常常完全绕开产品。实验层面本身更有价值,于是他们将其开放为Oqoqo。

Oqoqo在隔离沙盒中跨模型、智能体和产品版本运行真实任务,并报告通过率、token支出和完整轨迹,并接入CI,让破坏智能体工作流的改动像测试失败一样浮现。产品于2026年8月10日在Product Hunt发布,以338票成为当日第一,截至2026年8月中旬,网站上线且免费试用,网址为oqoqo.ai。

这件事要成立,得有什么

  • 人口结构变化肉眼可见:编程智能体被指向API、文档、命令行工具和技能,因此开发工具公司拥有了一类他们从未设计过的新用户。
  • 现有信号不起作用:请求日志显示流量,而非结果;公共基准在精选环境中排名模型;一次手动运行只是轶事——这为自主测量层留下真空。
  • 创始人握有第一手证据:他们自己的文档实验表明技能可能让智能体表现更差,因此转向开放平台是基于实测结果,而非理论推导。

可借鉴之处

衡量任务本身,而不是流量:从不抱怨的用户可能已经流失,而所有仪表盘仍是绿的——可衡量的单位是任务是否完成,以及没完成的原因。

后续进展

截至2026年8月16日,Oqoqo在oqoqo.ai上线且免费试用,支持跨智能体运行真实任务,如Claude Code、Codex和Cursor,支持原始智能体与MCP或工具处理对比,以及轨迹和摩擦审查,并提供CI集成。领投2025年11月种子前轮的Bee Partners称,Oqoqo评估“抵达的智能体”——当其他团队的智能体使用软件时,软件能否经得住考验。尚无收入、客户名称或后续轮次公开。

资料来源

发现哪里写错了?告诉我们。

轮到你了

你刚读完一家。说说你在做什么,看看谁在赌同一件事。

免费账号 · 3 次免费提问 · 不用绑卡

相关案例