The archive · AI & Models · Product decision · 2026
Weave bets most coding prompts don't need frontier models; 3.8k-star router
Router that sends Claude Code, Codex and Cursor prompts to the cheapest adequate model in <50ms; claims 40-70% savings, 3.8k GitHub stars.
Weave
What the business is
Weave sells a model router for AI coding agents: teams point Claude Code, Codex or Cursor at a proxy that classifies every request in under 50ms and sends routine prompts to cheaper models (Kimi, DeepSeek, Llama, Gemini, GPT variants) while reserving frontier models for hard work, with one policy across harnesses, automatic failover, and per-team quality-versus-cost settings.
How it started
Weave — which markets itself on weaveos.com as an engineering-intelligence platform that tracks code quality, keep rate and cost per merged PR — created the weave-os/router repository on 2026-04-27 and sells a hosted router. The pitch is the freight-truck analogy: most coding prompts don't need a frontier model, so the router 'reads each request, sends it to the smallest model that finishes the job at full quality', with classification in under 50ms at Weave's edge.
What happened
By the 2026-09-04 crawl the repo showed 3,849 stars. The product page, crawled the same day, promised 40-70% cost cuts with an endpoint change, an example calculator showing 53% lower monthly spend, 74% of requests routed to the cheaper pool, 99.9% completion with automatic failover and a 91/100 average code-quality score across routed traffic — all self-reported — and offered three policies (quality-first, balanced, cost-first) so teams set the quality floor the router may not cross.
How it ended up
Still live and early as of 2026-09-05: weave-os/router has 3,849 stars, the hosted sign-up and demo flow are open, and the source is available; the 40-70% savings and quality numbers remain Weave's own claims, with no independent revenue, funding or customer data in the record.
Background
Weave sells a model router for agentic coding tools: a drop-in proxy that sits between Claude Code, Codex or Cursor and the LLM providers, classifies each request in under 50ms, and sends routine prompts to cheaper models while keeping frontier models for the work that earns them. The marketing line is 'don't send a freight truck to deliver a postcard', and the page claims cost cuts of 40-70% with just an endpoint change.
The bet is that most coding prompts do not need a frontier model, so quality-per-token routing — the cheapest model that finishes the job at full quality — can roughly halve a team's token bill without measurable quality loss. Weave positions the router inside a product that already tracks code quality, keep rate and cost per merged PR, so savings are reported against metrics engineering teams already watch, net of cache misses.
The router launched in 2026: the weave-os/router repository was created on 2026-04-27 and reached 3,849 stars by the 2026-09-04 crawl. The product page offers three policies (quality-first, balanced and cost-first), one policy across every harness, automatic failover at 99.9% completion, and an example calculator showing 53% lower monthly spend on a $12,000 bill — figures that are Weave's own claims at this stage.
Adoption mechanics lean on low friction: an npx installer rewrites one environment variable per provider, the client keeps working unaware of the proxy, and the source is open so teams can inspect or self-host the router. As of 2026-09-05 no funding, revenue, customer or independent benchmark data appears in the record.
What has to be true
- Verifiable attention: weave-os/router went from creation on 2026-04-27 to 3,849 stars at the 2026-09-04 crawl — fast growth for a four-month-old developer tool.
- The bet is economically motivated and clearly stated: with a wide price spread across models, paying frontier prices for routine prompts is waste that routing can remove.
- The design removes adoption friction: one-command install, drop-in API compatibility and a single policy across Claude Code, Codex and Cursor, with no prompt changes required.
- The claims are structured to be checkable — open-source core, savings net of cache misses, quality measured on keep rate and cost per merged PR — though every number is still self-reported.
What can be applied
When model prices diverge widely, routing becomes a product: sell the measurable quality-per-token tradeoff with a floor the user sets, and open-source the core so cost claims can be checked.
Aftermath
As of 2026-09-05 Weave's router is live and early: weave-os/router carries 3,849 stars from the 2026-09-04 crawl, the weaveos.com page invites sign-up or a demo, and the source is public for self-hosting. The savings and quality figures (40-70% cost cuts, 53% on the example calculator, 91/100 quality score) are Weave's own marketing claims as crawled; no funding, revenue, customer list or independent evaluation appears in the material. The bet — routing routine coding prompts to cheaper models at half the cost without quality loss — remains unproven outside the company's own measurements.
Sources
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card