Weave Router 2.0 is an open-source model router that dynamically switches among specialist LLMs inside coding-agent sessions to assign each turn to the model best suited for the task. Benchmarks show it matches GPT-6 Astra’s pass rates on Terminal Bench 4.0 and SWE Atlas while cutting inference cost to 52% and 54% of Astra respectively, and completing sessions 2.2x and 2.5x faster. The result supports the core claim that a smart ensemble can equal frontier single models on coding tasks while materially reducing latency and cost; code and a hosted service are publicly available.
Three technical changes drove the gains. First, a new architecture replaces pure RL exploration with a hidden Markov model that traces session state and a classifier that maps sessions into buckets of candidate models, massively pruning the 10^100-scale routing search space. Second, a much larger training set was bootstrapped by using frontier LLMs to label diverse coding sessions, improving classifier and RL signals. Third, a cache-impact subsystem computes the expected value of switching models (one-time cache fill costs versus downstream benefit) to avoid costly unnecessary switches, which delivered most of the cost reduction. Developers plan further work to consistently outperform Astra/Fable.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.