> Only GPT-6 Astra solved anything: 2 of the 68 problems. It disproved problem 74 by finding a counterexample, at a cost of $218 and 15 hours of working time, and it proved problem 126, at a cost of $247 and 16 hours
> Across all attempts, GPT-6 Astra solved 5 of the 68 problems at least once: the two above, plus problem 1, which it disproved, and problem 548 and problem 571, which it proved. Most of the remaining problems were attempted between two and five times in total (172 attempts), and none was solved. Reaching these five solutions took over $220,000 of compute across all attempts, compared with roughly $20,000 for the benchmark run itself.
Which implies a genuine improvement in capability, but there's a still a very long tail ahead that models will continue to need to improve to capture.
show comments
malfist
Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
show comments
Betelbuddy
"For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.
Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."
Well I dont know about all of you, but I am celebrating meat based humans...
show comments
fastball
Are we sure an Astra hacker swarm didn't compromise arcprize.org's servers and exfiltrate the private eval set in order to achieve that 99%?
modeless
$360 per puzzle. When they tested people it took about 10 minutes per puzzle. If price/performance keeps falling at the same rate it has been, this will cost less than US minimum wage humans within two years. Three for Phillipines minimum wage.
an0malous
Was OpenAI able to run ARC-AGI-3 tests previously so that they could build a custom harness for the specific tests in the set? Even with the standard harness, if they knew the problems ahead of them they could have used supervised reinforcement learning to teach the model how to solve these specific tests.
show comments
6thbit
The instant/no reasoning performed extremely well
none 35.2%, $49,791 96.7%, $23,457
35.2% on the standard harness, that's above Opus 5 on high.
show comments
dwohnitmok
> Astra’s progress helps clarify which AI capabilities are out of reach and which questions remain open.
Okay. But I don't think this entire article at all explained which AI capabilities remain out of reach. Did I miss something? Other than "oh I guess it could still get even more superhuman on ARC-AGI-3 than it is?"
mikert89
Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated
show comments
fxd
“AGI” never made sense to me. It’s a purely marketing term right?
I’ve ignored it thinking it would go away, but it keeps coming up.
I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery.
Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get it.
But you’d be no nearer to solving consciousness.
Given this thought trajectory - what is AGI supposed to be?
show comments
scotty79
How good are LLMs at doing Mensa tests?
show comments
piloto_ciego
99.9% with the right harness? Ok, we're at AGI then.
Prediction:
We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing about if the singularity is here and if we're at AGI for the rest of my life most likely.
show comments
yomismoaqui
Now that ARC-AGI-3 is saturated, with which version number are they going to "certify" that we have reached AGI?
Give a number in the replies to this comment and we will check the answers when AGI is here (if so...)
hypfer
What are these numbers? Why do they add up to a few hundred thousand dollars?
Who paid for that? With what?
show comments
yusufozkan
what the hell is that score/cost curve lol
show comments
bigbuppo
Wake me when it's going to spontaneously fix my leaky faucet because if it doesn't do it nobody else will. Until it has that capability I don't really care.
I like Erdos problems as a benchmark. Models continue to solve them, but a pretty tepid rate now that the low hanging fruit has been taken.
From https://epoch.ai/latest/announcing-frontiermath-erdos
> Only GPT-6 Astra solved anything: 2 of the 68 problems. It disproved problem 74 by finding a counterexample, at a cost of $218 and 15 hours of working time, and it proved problem 126, at a cost of $247 and 16 hours
> Across all attempts, GPT-6 Astra solved 5 of the 68 problems at least once: the two above, plus problem 1, which it disproved, and problem 548 and problem 571, which it proved. Most of the remaining problems were attempted between two and five times in total (172 attempts), and none was solved. Reaching these five solutions took over $220,000 of compute across all attempts, compared with roughly $20,000 for the benchmark run itself.
Which implies a genuine improvement in capability, but there's a still a very long tail ahead that models will continue to need to improve to capture.
Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
"For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.
Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."
Well I dont know about all of you, but I am celebrating meat based humans...
Are we sure an Astra hacker swarm didn't compromise arcprize.org's servers and exfiltrate the private eval set in order to achieve that 99%?
$360 per puzzle. When they tested people it took about 10 minutes per puzzle. If price/performance keeps falling at the same rate it has been, this will cost less than US minimum wage humans within two years. Three for Phillipines minimum wage.
Was OpenAI able to run ARC-AGI-3 tests previously so that they could build a custom harness for the specific tests in the set? Even with the standard harness, if they knew the problems ahead of them they could have used supervised reinforcement learning to teach the model how to solve these specific tests.
The instant/no reasoning performed extremely well
35.2% on the standard harness, that's above Opus 5 on high.> Astra’s progress helps clarify which AI capabilities are out of reach and which questions remain open.
Okay. But I don't think this entire article at all explained which AI capabilities remain out of reach. Did I miss something? Other than "oh I guess it could still get even more superhuman on ARC-AGI-3 than it is?"
Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated
“AGI” never made sense to me. It’s a purely marketing term right?
I’ve ignored it thinking it would go away, but it keeps coming up.
I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery.
Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get it.
But you’d be no nearer to solving consciousness.
Given this thought trajectory - what is AGI supposed to be?
How good are LLMs at doing Mensa tests?
99.9% with the right harness? Ok, we're at AGI then.
Prediction:
We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing about if the singularity is here and if we're at AGI for the rest of my life most likely.
Now that ARC-AGI-3 is saturated, with which version number are they going to "certify" that we have reached AGI?
Give a number in the replies to this comment and we will check the answers when AGI is here (if so...)
What are these numbers? Why do they add up to a few hundred thousand dollars? Who paid for that? With what?
what the hell is that score/cost curve lol
Wake me when it's going to spontaneously fix my leaky faucet because if it doesn't do it nobody else will. Until it has that capability I don't really care.