draginol

So Fable "won" but it cost $124.76 for marginal performance benefits over the $22.56 5.6 Sol run.

xnorswap

A really frustrating partial presentation, given an apparent lack of testing with a spread of efforts for each model.

Given that there's no reason to believe that Fable's xhigh is comparable to GPT-sol's xhigh, or Opus xhigh, for that matter, it would be far more useful to see the effort level where these tasks no longer achieved their goals.

show comments
DwarvenEngineer

Honestly, I just hate the term "physical AI". They're robots. It's unfortunate that we had to adopt a term with the words AI in it, just to get investor's attention.

giwook

Please forgive my naivety, but are world models (once they are in a consumer-ready form) expected to outperform any currently existing LLM on these sorts of tasks (i.e. of the physical world)?

show comments
hartator

It's kind of interesting this is already out of data as it's missing Kimi 3 and Opus 5.

show comments
effnorwood

Define "best" and "performs"

grim_io

I'd expect google to do well here, since they were historically strong at multimodal and physics.

show comments
jespinel

Nice! It is missing Codex in the agent harnesses comparison IMO.

arisAlexis

Google with apptronic should have good models soon

gizmodo59

Yet another "benchmark to promote their own harness"