The Workweave authors argue that coding-agent traffic is dominated by trivial work — edits, lookups, formatting, status checks — that doesn't need a frontier reasoning model. They built and open-sourced the router because Opus 4.7's price made naive 'always use the smartest model' loops financially untenable at their own shop, where nearly all code is AI-written.
The editorial frames Workweave as part of a broader stratification: IDE, then agent loop, and now a routing tier underneath. It argues this layer is different from general-purpose routers like OpenRouter because it's tuned to the specific traffic shape coding agents produce — heavy tool-call follow-ups and a long tail of trivial completions.
The router exposes OpenAI- and Anthropic-compatible endpoints so existing agents need only a base-URL change in their config. The team's bet is that transparent insertion — the agent never knows the router is there — is what will let teams adopt cost-routing without rewriting their tooling.
A team at Weave pushed `workweave/router` to GitHub and posted it as a Show HN. It hit 158 points fast. The pitch is one line: a model router that plugs into coding agents — Claude Code, Codex, Cursor — and forwards each request to whichever model is cheapest-capable for the task at hand. The README ships with a local demo. The repo is small enough that you can read it in an afternoon.
The stated motivation is blunt. "At Weave, we write ~all our code with AI, and it's been getting more expensive. This came to a head when Opus 4.7 was released," the authors write — and they're not the only shop saying it out loud this quarter. Opus 4.7 is a real step up on hard reasoning tasks, but its per-token price puts naive 'always use the smartest model' agent loops into territory that finance can see from across the office. The router's bet is that most requests inside a coding agent are not hard reasoning tasks. They're edits, lookups, formatting, small refactors, file-tree traversals, status checks. A Haiku-class model can serve those for cents on the dollar. Opus comes off the bench only when the task warrants it.
The architectural move is to insert the router as an OpenAI- or Anthropic-compatible endpoint that the host agent already knows how to talk to. You change a base URL in your `claude` config or your Cursor settings, and the agent doesn't know anything has changed. The router does its triage and forwards.
The coding-agent stack has been quietly stratifying for a year. First there was the IDE (Cursor, Zed, VS Code). Then the agent loop (Claude Code, Codex, Aider, Cline). Now there's a third tier emerging underneath — a routing layer — and Workweave is one of several teams converging on it. OpenRouter has been doing a version of this for general inference for a while. What's new is routing tuned for the specific traffic shape that coding agents produce: lots of tool-call follow-ups, a long tail of trivial completions, and a few genuinely hard generations per session.
The economic argument writes itself: if 70% of an agent's calls are mechanical and 30% are reasoning, paying Opus prices on the 70% is a tax with no return. A router that hits 90% accuracy on that classification can cut bills by half without users noticing. That's a margin story dressed up as an infra story, which is the kind of story that gets adopted fast.
The community reaction in the HN thread tracks the usual fault lines. One camp wants the router to be invisible — auto-tune, no config, ship it. The other camp is wary of any layer that silently downgrades the model serving their request. Both camps are right, and the resolution is the same one we landed on for compilers and JIT: give power users a flag to pin the model and let everyone else trust the heuristic. The interesting question is not whether routing works. It's who owns the routing decision — the agent vendor (Anthropic, OpenAI, Cursor), or a neutral layer the user controls. If the agent vendors win that fight, routers become a feature flag. If neutral routers win, they become a market.
There's also a quieter implication. Routing layers need telemetry to improve — which classifications were right, which were wrong, which model the user retried with. That telemetry is enormously valuable, both for tuning the router and as a dataset on what coding agents actually do all day. Whoever runs the dominant router ends up with one of the better windows into developer workflow that's ever existed. That's worth more than the routing fees.
If you're paying real money for AI coding tools — meaning more than one seat, billed monthly — you should be running a router in front of your agents in Q3. Not because Workweave's is necessarily the right one, but because the savings are too obvious to leave on the floor. Benchmark it the way you'd benchmark a CDN: take a week of representative traffic, mirror it through the router, compare cost and task success against direct-to-Opus. If the router holds task success within a couple of percentage points at half the cost, ship it.
Two operational gotchas worth pre-empting. First: pin the router version in your team's config the way you pin a compiler version. Silent routing changes are the new silent model deprecations — the day your router decides Haiku can handle your migration scripts is the day three PRs land with subtly wrong SQL. Second: log the model that actually served each request. You want that field in your traces the first time a senior engineer says 'this used to work better.' Without it, you're debugging blind across an extra layer.
For agent vendors, this is a fork-in-the-road moment. Anthropic and OpenAI can either embrace routing — ship a first-party router that defaults to their stack but degrades gracefully — or pretend it isn't happening and watch the margin leak to third parties. The 'pretend' path didn't work for cloud providers when the multi-cloud abstractions showed up, and it won't work here.
Expect the routing layer to consolidate fast. Within two quarters, every serious coding-agent setup will have a router in front of it, and the interesting differentiation will move from 'does it route' to 'how good is its classifier.' Expect at least one frontier lab to ship its own first-party router and frame it as a feature. Expect at least one router to get acquired. And expect the next round of coding-agent benchmarks to start reporting cost-per-completed-task instead of raw pass@1 — because once routing is in the picture, the dollar figure is the only number that means anything.
We built a model router that plugs into coding agents (e.g. Claude Code, Codex, Cursor, etc.) and intelligently sends requests to the best model to serve them. Here's a quick demo of running it l
→ read on Hacker NewsI see a great tension in the market today. On one hand you want agents to work reliably and that needs a lot of harness, computer use, model routine, tasks running longer etc. And on other hand you simply want to reduce your dependencies and costs. Agent building is very nascent and all the frontier
The thing I do not get with these routers is that you will have more cache misses (5min ttl). And if there is one thing i’ve learned; using the cache is crucial.How does this router translate to $$$ when developing?
It's rather hard to do at the proxy level with agentic coding, such as Claude Code or similar. These are long-chained sessions of tool use that heavily rely on prompt caching. Changing mid-flight is costly.It looks like much more context is required to decide on the best model (e.g., summarizin
Man, I'm not so sure if I'd use something like this because the way I prompt already changes based upon what model I am using. I'm not convinced it would route to the right model based on my diction or whatever.
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.
I'm glad there are more attempts at solving model routing, as costs (at API rates) has really become an issue. Some feedback:1. Reiterate the cache issue from other comments already here. there is a lot of optimisation in harnesses around caching and a proxy model blows that up2. Coding agents