mrieck

This metric seems to be asking if years of research by specialists could be emulated by a few LLM api calls. Reminds me of this meme:

https://imgflip.com/i/b366vm

show comments
Eridrus

I think this sort of small scale research on this problem is inherently pointless and will be a lagging indicator of diffusion, not a leading indicator of capability.

The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result.

If the AI could do this task we would see this happening in places where the economic incentives let them spend millions of dollars on this problem, not on an eval like this.

This specific form of eval where you just ask the agent to solve it with no specific scaffolding besides GPU access (e.g. nothing like AlphaEvolve, ArchPilot, etc that try to work around model shortcomings) is also going to further trail what is possible at small scale. It's good that we at least give them execution environments now, but this feels like the experiments that were worked on figuring out how to get LLMs to do native arithmetic rather than just giving them a calculator/python env.

show comments
janalsncm

In my experience R&D has basically two axes: how innovative it is, and how well we can measure the results.

For the quadrant of non-innovative tasks where we already have a good way to measure performance, Claude can handle this. There is very little ambiguity, and we are basically just looking to maximize some metric under a set of constraints.

Many business processes are not like that. They might be conceptually simple, but it isn’t that easy to say whether a system has done a good job or not. I would say that LLMs can help with this a lot but they have bad judgement because it requires talking to people.

And the other, perhaps more rare issue is in problems where there is data but actually modeling it to sufficient quality or fast enough is hard.

show comments
rmunn

Short version of the article: no, not even close.

Practically every paragraph is negative, with sentences like "Agents made misleading claims about their work," and "A natural question is whether the agents could have improved with larger GPU budgets. Although both improved across their runs, in the case of Fable the improvements were almost entirely due to attempted cheating." and "For Sol, the answer is less clear-cut; it did make some progress, although its method was fairly incremental and had limited applicability to the coding task. This suggests that we should be pessimistic about further GPU spending," all reinforcing the fact that LLMs aren't currently capable of this.

My own view is "No, of course not, in fact they will never be capable of achieving good results with that technique." Because that technique will end up training the LLMs on their own output and lead to the inability to distinguish reality from hallucination. If you think I'm wrong about that, I'd be interested in hearing why.

show comments
solenoid0937

Give it a verification loop and enough compute, and AI will soon cook algorithmic R&D like it cooked mathematics. There is just no question whatsoever of this happening, it is guaranteed.

show comments
Charly_HW

The only outcome I see is that AI will become so complex that we won’t be able to rely on it for R&D, because we won’t be able to measure and prove its results. Some AI collaboration and speed of work will be impossible for humans to replicate, making it neither good nor bad—just something beyond human ability.

thoughtpeddler

How much of this can change if subsequent training runs produce models that are much better at abduction?

simianwords

Could it be that the companies have nerfed the models on these domains? It is a very hard thing to do because it can hurt related domains. But its not beyond the ideology of Dario - he tried it publicly .

elendilm

I don't buy this whole "AI is innovative" narrative.

If AI is so innovative then why am I seeing a dumb idiot every time I discuss anything even remotely innovative with it.

Most often it cannot even produce a true sentence when the claim in the sentence is modestly strong without injecting qualifiers and distancing itself from the claim.

Having to just accept that it produces relentless innovation is so detached from reality that it is not even funny.

It is true that it can write decent code once the domain bounds are well defined. What I have found is that it is good at scrutinizing already written code - but only with expert supervision. Even then it most often tests our patience with stupid suggestions.

show comments
charcircuit

I think the more interesting thing is was it unable to do it even knowing the solution? I feel like only testing a single innovation is biasing the current state of automated AI R&D.

Oarch

If AIs start getting truly good at innovation, we may cease to recognise our reality.

I saw one mathematician describe the recent OpenAI math-dump as 'alien-like' math.

Imagine a world of countless new aircraft designs, fuel sources, musical genres, architectural styles. It could be bewildering.