If you are training on data labeled by frontier models, how do you expect to exceed the performance of frontier models, other than in the cost dimension by recognizing simpler problems and routing to cheaper models?
Have you compared this to using GPT-6.1 Sol instead of GPT 6 Astra + Deepseek? From my test, 6.1 Sol is a lot more token efficient than 6 Sol while being similar to Astra in performance, and I don't really find 6 Astra to be significantly better than 6/6.1 Sol for general coding as I feel 6 Astra is only noticeably better at spatial reasoning/vision compared to 6 Sol, and 6.1 Sol really closed the gap on that front.
To the point of using buckets of models: as long as there's >0 models available in a bucket, and we can order models in order of fit, we're resilient to different sets of models being available. With that said obviously cutting out some models has a much larger effect than others.
OpenRouter isn't strictly necessary but it does make it easier to not have to set up accounts/API keys with several different providers to get started.
You can think of buckets as models with similar capabilities. So for example Deepseek 4.1 Flash will not be in the same bucket as Astra.
The latter is an interesting question! In practice because the session continues, we can see it's going down the wrong path and escalate. Basically no decision we make when routing is entirely unsalvageable (but we do have a performance penalty for every incorrect decision we make so of course we try to avoid it).
Interesting work, and thanks for describing how your router works internally. It's definitely a fascinating subject. How would you say this compares to Cursor's auto mode?
I haven't tried it as recently but last time I checked they only route once per session (or subagent). Imo this is basically impossible to do correctly. Consider the case where you start with one prompt "rewrite this in rust". Trivial in a 1 day old repo, extremely difficult in e.g. the VSCode repo!
18 comments:
If you are training on data labeled by frontier models, how do you expect to exceed the performance of frontier models, other than in the cost dimension by recognizing simpler problems and routing to cheaper models?
Different frontier models are good at different things! We'll be the ones combining them optimally.
Can it route to locally or LAN hosted Qwen or some other open weights model?
Have you compared this to using GPT-6.1 Sol instead of GPT 6 Astra + Deepseek? From my test, 6.1 Sol is a lot more token efficient than 6 Sol while being similar to Astra in performance, and I don't really find 6 Astra to be significantly better than 6/6.1 Sol for general coding as I feel 6 Astra is only noticeably better at spatial reasoning/vision compared to 6 Sol, and 6.1 Sol really closed the gap on that front.
How do you handle provider variance on OpenRouter for the opensource models? Or do you use your own hosted version to mitigate this?
And for both opensource and closed source, does the router account for provider quality, or catch it when a provider degrades?
How does this choose which models to use with any arbitrary set of model providers to work from? And why is an openrouter necessary for self-hosting?
To the point of using buckets of models: as long as there's >0 models available in a bucket, and we can order models in order of fit, we're resilient to different sets of models being available. With that said obviously cutting out some models has a much larger effect than others.
OpenRouter isn't strictly necessary but it does make it easier to not have to set up accounts/API keys with several different providers to get started.
How do you define the model buckets, and what happens when a session genuinely needs a model that isn't in the bucket the HMM picked?
You can think of buckets as models with similar capabilities. So for example Deepseek 4.1 Flash will not be in the same bucket as Astra.
The latter is an interesting question! In practice because the session continues, we can see it's going down the wrong path and escalate. Basically no decision we make when routing is entirely unsalvageable (but we do have a performance penalty for every incorrect decision we make so of course we try to avoid it).
Interesting work, and thanks for describing how your router works internally. It's definitely a fascinating subject. How would you say this compares to Cursor's auto mode?
Absolutely!
Conceptually very similar to Cursor's auto mode. The key distinctions are:
- We plug into any harness (e.g. Claude Code, Codex, OpenCode, Pi)
- We aren't incentivized to route to our own model, we're incentivized to route to the best model whatever it may be
And similarly, Copilot’s Auto mode?
I haven't tried it as recently but last time I checked they only route once per session (or subagent). Imo this is basically impossible to do correctly. Consider the case where you start with one prompt "rewrite this in rust". Trivial in a 1 day old repo, extremely difficult in e.g. the VSCode repo!
Does this allow for a predefined budget?
Yes!
Is the model you trained available as open weights?
It is not sorry!
AGI is here!