So Fable "won" but it cost $124.76 for marginal performance benefits over the $22.56 5.6 Sol run.
xnorswap
A really frustrating partial presentation, given an apparent lack of testing with a spread of efforts for each model.
Given that there's no reason to believe that Fable's xhigh is comparable to GPT-sol's xhigh, or Opus xhigh, for that matter, it would be far more useful to see the effort level where these tasks no longer achieved their goals.
show comments
DwarvenEngineer
Honestly, I just hate the term "physical AI". They're robots. It's unfortunate that we had to adopt a term with the words AI in it, just to get investor's attention.
giwook
Please forgive my naivety, but are world models (once they are in a consumer-ready form) expected to outperform any currently existing LLM on these sorts of tasks (i.e. of the physical world)?
show comments
hartator
It's kind of interesting this is already out of data as it's missing Kimi 3 and Opus 5.
show comments
effnorwood
Define "best" and "performs"
grim_io
I'd expect google to do well here, since they were historically strong at multimodal and physics.
show comments
jespinel
Nice! It is missing Codex in the agent harnesses comparison IMO.
arisAlexis
Google with apptronic should have good models soon
gizmodo59
Yet another "benchmark to promote their own harness"
So Fable "won" but it cost $124.76 for marginal performance benefits over the $22.56 5.6 Sol run.
A really frustrating partial presentation, given an apparent lack of testing with a spread of efforts for each model.
Given that there's no reason to believe that Fable's xhigh is comparable to GPT-sol's xhigh, or Opus xhigh, for that matter, it would be far more useful to see the effort level where these tasks no longer achieved their goals.
Honestly, I just hate the term "physical AI". They're robots. It's unfortunate that we had to adopt a term with the words AI in it, just to get investor's attention.
Please forgive my naivety, but are world models (once they are in a consumer-ready form) expected to outperform any currently existing LLM on these sorts of tasks (i.e. of the physical world)?
It's kind of interesting this is already out of data as it's missing Kimi 3 and Opus 5.
Define "best" and "performs"
I'd expect google to do well here, since they were historically strong at multimodal and physics.
Nice! It is missing Codex in the agent harnesses comparison IMO.
Google with apptronic should have good models soon
Yet another "benchmark to promote their own harness"