EN
Back to the archive

The archive · Developer & Business Tools · Strategic decision · 2025–2026

Oqoqo bets products must prove they work for agents — launch goes #1 on Product Hunt

Oqoqo sells private evals and custom benchmarks that measure whether AI agents can actually use a product — #1 Product of the Day on launch with 338 upvotes.

Oqoqo

The betEvery product now has a second, silent user: agents. Sell those companies the measurement layer — private evals and custom benchmarks on their own real workflows.Live

What the business is

Oqoqo runs real-world agentic tasks in sandboxed environments across models, agents and product versions, then reports pass rates, token spend and full step-by-step trajectories, so teams building agent-facing products can tell whether agents actually succeed on the workflows that matter to them.

How it started

In November 2025 Bee Partners led Oqoqo's pre-seed out of Bee IV. Co-founders Renzo Viale and Haritha Nair, both Berkeley Haas, had started from the premise that developer documentation written for humans is the wrong substrate for coding agents; working alongside DevRel teams, they kept asking how anyone knows the docs and skills they ship actually help an agent. The answers were anecdotal, so they built experiments: running dev-tool products across different agents, with and without skills, each trial in its own sandbox. The results were unflattering — some skills made agent behavior worse, some inflated token consumption, and agents often worked around the product entirely.

What happened

The experiment surface proved more valuable than the docs, so Viale and Nair opened it up. Oqoqo turns a real task — project state, files, data, tools and services plus a rubric defining success — into reproducible runs across models, agents and product versions, each trial in its own sandbox, and reports pass or fail with reasons, full trajectories, token spend and the points where an agent gave up on the product. The founders position it against public benchmarks, which they say rank models in curated environments that do not translate to the real world. The product launched on Product Hunt on 2026-08-10, taking #1 Product of the Day with 338 upvotes and #9 of the week in the developer-tools list, with the site live and free to try at oqoqo.ai.

How it ended up

As of mid-August 2026 Oqoqo is live and pre-seed-funded: its homepage supports running real tasks across agents including Claude Code, Codex and Cursor, comparing a raw agent against MCP and tooling treatments, and reviewing trajectories and frictions with CI integration. No revenue, customer names or follow-on round are public yet.

Background

In November 2025 Bee Partners led Oqoqo's pre-seed out of Bee IV. Co-founders Renzo Viale and Haritha Nair, both Berkeley Haas, had started with the premise that developer documentation written for humans is the wrong substrate for coding agents, and that any company serving developers would need its docs rebuilt so an agent could use them without a person in the middle.

Working alongside DevRel teams, they kept asking a question nobody could answer: how do you know the documentation and skills you ship are actually helping an agent? They built experiments — running dev-tool products across agents, with and without skills, in sandboxes — and the results were unflattering: some skills made agent behavior worse, some inflated token consumption, and often the agent worked around the product entirely. The experiment surface was the more valuable thing, so they opened it up as Oqoqo.

Oqoqo runs real tasks in isolated sandboxes across models, agents and product versions, and reports pass rates, token spend and full trajectories, wired into CI so a change that breaks an agent workflow surfaces like a failing test. It launched on Product Hunt on 2026-08-10 as #1 Product of the Day with 338 upvotes, and was live and free to try at oqoqo.ai as of mid-August 2026.

What has to be true

  • The population shift is visible: coding agents are pointed at APIs, docs, CLIs and skills, so dev-tool companies have acquired a class of user they never designed for.
  • Existing signals fail: request logs show traffic, not outcomes; public benchmarks rank models in curated environments; one manual run is an anecdote — a real gap for an owned measurement layer.
  • Founders had first-hand evidence: their own docs experiments showed skills could make agents worse, so the pivot to an open platform came from measured results, not theory.

What can be applied

Instrument the task, not the traffic: a user that never complains can leave while every dashboard stays green — the measurable unit is whether the job finished, and why not.

Aftermath

As of August 16, 2026 Oqoqo is a live, free-to-try product at oqoqo.ai, supporting real tasks across agents such as Claude Code, Codex and Cursor, raw-agent versus MCP or tooling comparisons, and trajectory and friction review with CI integration. Bee Partners, which led the November 2025 pre-seed, describes Oqoqo as measuring 'the agents that arrive' — whether software holds up when other teams' agents use it. No revenue, customer names or follow-on round are public yet.

Sources

spotted an error? The archive wants to know.

Your turn

You just read one. Describe what you are building, and see who is betting on the same thing.

Free account · 3 free questions · no card

Related cases