How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
31 source pointsReal observations only. History before source connection is not reconstructed.
A few months ago we started building a model router for coding agents because we thought we could outperform any single model with an ensemble approach. Recently we’ve achieved that milestone and I want to talk about how we did it. First of all, a quick explanation: the Weave Router ( https://github.com/weave-os/router ) plugs into any coding agent (e.g. Claude Code or Codex) and intelligently switches between LLMs. So, for example, Astra handles tricky debugging or complex system design tasks, and Deepseek v4 Flash handles simple frontend updates. What we’re announcing today is our new routing model, which we’re calling Weave Router 2.0. We benchmarked 2.0 against GPT-6 Astra on Terminal Bench 4.0 and SWE Atlas. On both benchmarks, the router had equivalent pass rates. On Terminal Bench, the router hit 52% of Astra’s cost, and completed tasks 2.2x faster. On SWE Atlas, the router cost 54% as much as Astra and ran 2.5x faster. (Full results on our website at https://weaveos.com/router !) It turns out training a model to route effectively - taking into consideration model capabilities, costs, cache awareness, and more - is a really hard problem! I want to talk about three ways we were able to improve so much over the last few months: 1) a new architecture, 2) larger training data set size, and 3) smarter cache-eviction impact calculation. 1) a new architecture. Our initial approach used an RL model without many priors. While RL is still an important part of the story, the cost of fully exploring the space of routing decisions is very high, so we’ve taken some shortcuts that have significantly improved performance. Consider how large the search space for the routing problem is. Take a typical coding agent session, with ~100 agent turns (i.e. 100 LLM API calls). Technically there are 100 chances to select a model. If we assume a roster of ~10 models (of course there are lots more but we can remove any that are Pareto dominated), then there are 10^100 possible paths through that session. We simply cannot explore all of them! So that's why clever tricks to shrink this space are so important. In particular: we trained a hidden Markov model to trace the session state, then a classifier maps the session to one of a few buckets of similar models. Using the HMM allows us to evaluate not just where a session is currently, but how it got there . We've gotten significantly better performance on bucket selection by incorporating that information - we believe this is because two sessions that might look quite similar to a naive classifier are much better distinguished by this HMM approach. Using this HMM + classifier to select a bucket first significantly shrinks the space to explore, by throwing out most models that could not reasonably serve the given session. Th
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
31 source pointsReal observations only. History before source connection is not reconstructed.
Dots is an always-on agent system designed for continuous interaction and task execution.
2 days agoHacker NewsThe text describes how OpenAI agents were involved in a security incident at Hugging Face, but it does not provide specific technical details about the method or extent of the breach.
5 days agoHacker NewsThe project introduces Reladraw, a diagramming tool that allows users to define diagrams in a diagram language while maintaining control over the layout. It aims to combine the benefits of auto-placement languages like Mermaid and the flexibility of tools like Draw.io, with support for both human and agent use.
5 days agoHacker NewsThe title 'You said no MCP' suggests a response or rejection of an MCP (possibly a system or entity), but no further details are provided in the description.
yesterdayLessWrongThe article discusses how TeX, originally created for typesetting mathematical expressions, is now being used by language models to perform mathematical reasoning. When solving complex math problems, models like GLM-5.3 use TeX symbols as part of their reasoning process, even though TeX was never designed for computation. This shift highlights an unexpected application of TeX in modern AI systems.
2 days agoLessWrongThe transcript of an interview with Tristan Buckmaster discussing the Navier-Stokes controversy and his research. It includes his reflections on the experience of being overwhelmed by attention and the clash between the tech industry's fast-paced culture and mathematical research.
yesterday