Local models need to be tuned to work well so this looks useful. Seems to be for general purpose model serving. I’ve been using https://github.com/adrianco/retort to run experiments for coding models across 13 different programming languages to see which frontier and local models work.
show comments
anshad2u
Interesting approach. What does the cold-start phase look like for a new agent? How many traces or runs do you typically need before the router has enough signal to safely offload tasks from the frontier model??
jack_pp
Not sure I get it. The model you're improving is local? If so how do you even calculate cost compared to an API
show comments
surround
The title is misleading. This is model routing, not distillation.
show comments
digitaltrees
Cool project
show comments
yiyingzhang
Cool idea! How do you guarantee privacy?
show comments
rglover
Excited to play with this.
show comments
Art9681
The absolute best way to prove this works is by releasing a model that was fine-tuned with this method and then showing benchmarks depicting the improvement delta between the base model and the fine tuned one.
The work is not done. Then release it to the masses and wait a few days for the actual real world anecdotes.
Local models need to be tuned to work well so this looks useful. Seems to be for general purpose model serving. I’ve been using https://github.com/adrianco/retort to run experiments for coding models across 13 different programming languages to see which frontier and local models work.
Interesting approach. What does the cold-start phase look like for a new agent? How many traces or runs do you typically need before the router has enough signal to safely offload tasks from the frontier model??
Not sure I get it. The model you're improving is local? If so how do you even calculate cost compared to an API
The title is misleading. This is model routing, not distillation.
Cool project
Cool idea! How do you guarantee privacy?
Excited to play with this.
The absolute best way to prove this works is by releasing a model that was fine-tuned with this method and then showing benchmarks depicting the improvement delta between the base model and the fine tuned one.
The work is not done. Then release it to the masses and wait a few days for the actual real world anecdotes.
Until then, this is noise.