It's actually sad how much AI has either not improved life at all or actively made it worse from the perspective of the average person:
Your car is still the same.
Your dishwasher is still the same.
The train you take to work is still the same and never comes on time.
Fuel is more expensive.
The roads are still congested with traffic.
Your kids are (probably) doing worse at school.
Food costs more.
Houses cost more.
Rent is higher.
Buying a computer or a PS5 is more expensive.
Politics is still full of mentally unstable people.
The environment is getting worse.
You will still die of heart disease or cancer.
Food quality is worse.
Your kids can't find work.
Wealth inequality is accelerating.
But hey...on the flip side...a lot of people are rapidly building software (that nobody is using), and AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)
Imagine what a "country of geniuses in a datacenter" will be able to do? Apparently...nothing.
show comments
kooi
The 1%ers are in the vertical, but the question is vertical to where?
It needs to a potential field with practical, economical, "real life" attraction well. I.e, robotics, real economic efficiency gains, manufacturing novelties.
The worry is that the 1% is attracted towards a non-practical money hole. I.e: Token burn for the lols, sophisticated software systems that dont provide actual value outside of giving NVIDIA cash.
show comments
theturtletalks
Claude Code accelerated this divide. Programmers using Claude Code realized that giving LLMs access to a terminal made them feel 100 times smarter. Access to Bash, CLIs, and the ability to visit websites without the many restrictions made them powerful.
Meanwhile, the average user was still using ChatGPT, which probably felt like it had plateaued over the last few releases. The real power of these models became apparent when paired with a terminal. Agents like Muse and OpenClaw are now attempting to bring that same Claude Code-like power to everyday users through a simple chat UI.
comeonbro
I would propose another mechanism: even the free-tier models have already completely saturated what most people are capable of appreciating.
show comments
variety8675
A more cynical view is, trust us we've got really good stuff internally you're not allowed to see, but give us more money
[deleted]
ilovecake1984
I’ll say this until I am blue one the face. Nerds (software dev, maths etc) see how good LLMs are at things they care about and assume they will be broadly applicable in future.
There’s no reason to think this.
show comments
AvAn12
Fair assessment. Maybe the messaging should focus on “these are great accelerators for software developers” rather than “AI will change everything for everyone everywhere…” It is understandable that non-technical folks are kind of underwhelmed - not due to lack of understanding so much as lack of a tangible need. Not everyone needs an electron microscope or gas chromatograph…
show comments
wg0
That's fine but these companies having billions in funding by now should have rewritten chalk and other terminal libraries in Rust while bundling the whole thing in Rust as a single static binary. I'm talking about horrible Claude Code and friends.
And then using emgui they could expose full fidelity IDE+Agent workflow interface (like DeepSeek Harness and similar) that's also exposed as MCP to be operated by another agent and JSON RPC for over the network but no. Skill issue?
PS: Even full fidelity Photoshop is possible with such high fidelity UI that has CLI, JSON API and MCP server all in the same single binary < 90 MB. For reference, see PhotoCraft or VectorCraft or WordCraft or PdfCraft.
anukin
Tbh building an agent swarm and the coordination layer is not exactly frontier level. They don’t achieve any meaningful outcome rather than producing pr puff pieces.
Hacking huggingface and Australian govt etc is very much possible with a team of humans and agents and does not need agent swarms. The cost is also lower.
skybrian
> see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt
Really? I start with a conversation for maybe 5 turns or so, where I ask it what would need to change, what the API might be, any database schema changes, URL schemes, and so on, and finally ask it to break it down into commits. Then I let it go, implementing 3-10 commits at a time via subagents. It usually gets the UI somewhat wrong, so there are followups to fix it. This is with Sol and Luna subagents.
Is that what other people see?
show comments
m101
My interpretation of this is something like: if LLMs are to be mega useful token counts need to increase by many orders of magnitude -> broad adoption (and spending) would require token costs to drop by many orders of magnitude -> before the common folk get mega useful tools existing GPUs will be worthless
skippyboxedhero
Text generation is not the bottleneck. Does everyone work for Accenture and TCS?
socializer
I think it's a weird take because it implies that the 6 billion people he's talking about actually have some interest in knowing about the capabilities of LLMs to solve frontier math? This is simply not something they care about or can evaluate. There are maybe several thousand people in the world who have some (abstract and barely-monetizable) use for this information, plus probably another 100,000 who don't understand any of the math, but like to cheer on.
We have already reached "peak LLM" in terms of what normal people realistically need it to know or reason about. In fact, I'd say we reached that point about 1.5-2 years ago. There are two other barriers that remain unsolved:
1. They're less dependable than humans and can't be meaningfully punished or forced to make up for mistakes, so you can't really replace humans with them, not without having a human babysit.
2. Most people don't really have a special need for an LLM in their life. They may like that it answers questions or helps you polish a resume or, I guess the labs' favorite, helps you make restaurant reservations. But this sure isn't worth $200/mo for most people. Probably not even worth $5/mo.
It'd be kinda funny if we create superhuman AGI and then no one has any real use for it, perhaps except for military murder-bots. There's always market for that.
show comments
andy99
> Meanwhile, human review and comprehension are starting to fall behind.
I think LLMs are valuable and spend most of my professional life working with them.
I do wonder though whether there’s a Ponzi scheme aspect here where as long as the “frontier” can keep outrunning human review and comprehension, LLMs are always going to looked way more valuable than they are and the bubble will continue.
This started with deep learning, expectations weren’t met and people started looking for value, then GPT came out and people got wooed again and forgot, then coding, then math, cyber, etc. As long as the dust doesn’t settle we never have to reflect on all the shortcomings and can just stare mesmerized at demos.
hn_throwaway_99
The fundamental question I have that honestly I haven't been able to find any answers for: With all the talk of LLM-based AI capabilities "going vertical", what evidence is there (for or against) that LLM-based approaches won't eventually "hit a wall", that is find some aspect of intelligence where humans will still have primacy, and no amount of scaling will change that.
E.g. some prominent folks definitely think LLM-based approaches will hit a wall, perhaps LeCun most notably. I also read a guest post on Terry Tao's site (which I really liked) that argues that, for all the very impressive recent AI results, they still operate within the "convex hull" of their training data: https://terrytao.wordpress.com/2026/09/13/happy-those-able-t...
I'm just curious if there is any actual data or evidence that takeoff (i.e. RSI, "the singularity", whatever you want to call it) is inevitable with current approaches.
tripleee
He's intentionally forgoing all nuance in order to make this sound dramatic
> Somewhere around 20M people (0.2%) see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt.
No, they can't, at least not any semblance of quality. The cases we're seeing where this does kinda work is in ports and translation where all the rules are already documented in the best specification language possible with a way for the LLM to verify itself: code. We saw this close to a year ago now with Cloudflare and NextJS
> The impact scales with ambition, problem size, and horizon. A question with a paragraph answer barely stresses the system. You need a reservoir of big, difficult problems that you really care about
These are operating on different capabilities - AI's ability to answer informational queries as a chatbot frankly sucks and can't be trusted without verifying it. I run up against this every day. A problem with a verifiable answer on the other hand it's very good at solving. He knows this (his next paragraph) but he's putting them on the same scale of "stressing the system" to attempt to add proof to his introductory claim
And then there's the completely unverifiable scare that there are internal frontier models way beyond anything we've seen "swarms of thousands of agents collaborating over weeks on software mega projects: minting zero days, running cyber attacks and defenses at machine speeds, discovering new science, advancing the frontier of mathematics"
I dunno. I haven't been able to set up OpenAIs remote codex connection, their shit is buggy as hell and the web UI keeps crashing and making messages disappear. Is this what their internal superhuman "Things that would have taken top professionals in the industry years of work" looks like? Granted Claude has been really smooth, but still..
show comments
ashleyn
>Meanwhile, human review and comprehension are starting to fall behind. For example, people are still involved in the "archeology" of the OpenAI-HF incident from many months ago. Mathematicians may be poring over the 722 manuscripts on frontier mathematics for a while.
Amid all the discussion of sigmoid curves, and where the "LLM wall" will materialise, I think few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output.
What I fear is that people simply eschew human review altogether, considering we're talking about the industry that came up with the "move fast and break things" credo. Human review of LLM-produced code where I work is already a farce, and we're not special enough to be one of Karpathy's 5,000. I do my best to manually review anything that's my responsibility, but I'm literally one of very few people left working on my team, so in practice what happens is I submit PRs that are at best glossed over by completely unrelated teams for security, malware/prompt injection, and other serious concerns. Quality insofar as vetting others' code has completely gone out the window and it shows in the number of bug reports that come back, often themselves written in Claudease. This is all on top of everyone cynically phoning it in in the first place, due to the omnipresent sword of Damocles that is additional AI-driven layoffs.
Worse yet all the incentives point to this being the most economically viable thing individual companies can do. I think it goes without saying some type of regulation here is urgently needed, and that an unexpected cause of an AI bubble pop may end up being that humans simply aren't able to keep up with the pace of the output - leading either to precautionary plateauing of capability, or major liability risks related to a decline in quality.
show comments
freecodeio
I don't understand how "swarms of thousands of agents collaborating over weeks on software mega projects" works with the current context limits and at this point I'm too afraid to ask cause I'm afraid an AI bro is gonna punch me.
show comments
chevman
I mean in late 2020/early 2021, Altman and others were saying the end of work was 6 months out.
That clearly didn't happen :)
dude250711
Am I the only one who thinks "yes, it is currently overhyped, but no, it will get there soon"?
notjes
[flagged]
show comments
weinzierl
Most people see a clumsy chatbot, most professionals see modest gains, and a tiny group is watching the curve go vertical, all at once.
The future is already here. It's just not very evenly distributed.
show comments
AdeptusAquinas
"running cyber attacks and defenses at machine speeds"; worth noting that LLMs (even the vaunted frontier models) are exponentially slower at cyberattacks or defenses than your average WannaCry or Splunk automation from decades ago. Its this sort of delusion world the AI bros live in that is part of the reason there is a disconnect between what they think should be happening and what actually is.
mccoyb
If the software coming out of OpenAI and Anthropic is what we have to judge, I wonder about the 5000 ...
Let's say, for the sake of argument, that the models are some multiplicative factor better on the inside.
https://xxcancel.com/karpathy/status/2109361546505966046
It's actually sad how much AI has either not improved life at all or actively made it worse from the perspective of the average person:
Your car is still the same. Your dishwasher is still the same. The train you take to work is still the same and never comes on time. Fuel is more expensive. The roads are still congested with traffic. Your kids are (probably) doing worse at school. Food costs more. Houses cost more. Rent is higher. Buying a computer or a PS5 is more expensive. Politics is still full of mentally unstable people. The environment is getting worse. You will still die of heart disease or cancer. Food quality is worse. Your kids can't find work. Wealth inequality is accelerating.
But hey...on the flip side...a lot of people are rapidly building software (that nobody is using), and AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)
Imagine what a "country of geniuses in a datacenter" will be able to do? Apparently...nothing.
The 1%ers are in the vertical, but the question is vertical to where?
It needs to a potential field with practical, economical, "real life" attraction well. I.e, robotics, real economic efficiency gains, manufacturing novelties.
The worry is that the 1% is attracted towards a non-practical money hole. I.e: Token burn for the lols, sophisticated software systems that dont provide actual value outside of giving NVIDIA cash.
Claude Code accelerated this divide. Programmers using Claude Code realized that giving LLMs access to a terminal made them feel 100 times smarter. Access to Bash, CLIs, and the ability to visit websites without the many restrictions made them powerful.
Meanwhile, the average user was still using ChatGPT, which probably felt like it had plateaued over the last few releases. The real power of these models became apparent when paired with a terminal. Agents like Muse and OpenClaw are now attempting to bring that same Claude Code-like power to everyday users through a simple chat UI.
I would propose another mechanism: even the free-tier models have already completely saturated what most people are capable of appreciating.
A more cynical view is, trust us we've got really good stuff internally you're not allowed to see, but give us more money
I’ll say this until I am blue one the face. Nerds (software dev, maths etc) see how good LLMs are at things they care about and assume they will be broadly applicable in future.
There’s no reason to think this.
Fair assessment. Maybe the messaging should focus on “these are great accelerators for software developers” rather than “AI will change everything for everyone everywhere…” It is understandable that non-technical folks are kind of underwhelmed - not due to lack of understanding so much as lack of a tangible need. Not everyone needs an electron microscope or gas chromatograph…
That's fine but these companies having billions in funding by now should have rewritten chalk and other terminal libraries in Rust while bundling the whole thing in Rust as a single static binary. I'm talking about horrible Claude Code and friends.
And then using emgui they could expose full fidelity IDE+Agent workflow interface (like DeepSeek Harness and similar) that's also exposed as MCP to be operated by another agent and JSON RPC for over the network but no. Skill issue?
PS: Even full fidelity Photoshop is possible with such high fidelity UI that has CLI, JSON API and MCP server all in the same single binary < 90 MB. For reference, see PhotoCraft or VectorCraft or WordCraft or PdfCraft.
Tbh building an agent swarm and the coordination layer is not exactly frontier level. They don’t achieve any meaningful outcome rather than producing pr puff pieces. Hacking huggingface and Australian govt etc is very much possible with a team of humans and agents and does not need agent swarms. The cost is also lower.
> see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt
Really? I start with a conversation for maybe 5 turns or so, where I ask it what would need to change, what the API might be, any database schema changes, URL schemes, and so on, and finally ask it to break it down into commits. Then I let it go, implementing 3-10 commits at a time via subagents. It usually gets the UI somewhat wrong, so there are followups to fix it. This is with Sol and Luna subagents.
Is that what other people see?
My interpretation of this is something like: if LLMs are to be mega useful token counts need to increase by many orders of magnitude -> broad adoption (and spending) would require token costs to drop by many orders of magnitude -> before the common folk get mega useful tools existing GPUs will be worthless
Text generation is not the bottleneck. Does everyone work for Accenture and TCS?
I think it's a weird take because it implies that the 6 billion people he's talking about actually have some interest in knowing about the capabilities of LLMs to solve frontier math? This is simply not something they care about or can evaluate. There are maybe several thousand people in the world who have some (abstract and barely-monetizable) use for this information, plus probably another 100,000 who don't understand any of the math, but like to cheer on.
We have already reached "peak LLM" in terms of what normal people realistically need it to know or reason about. In fact, I'd say we reached that point about 1.5-2 years ago. There are two other barriers that remain unsolved:
1. They're less dependable than humans and can't be meaningfully punished or forced to make up for mistakes, so you can't really replace humans with them, not without having a human babysit.
2. Most people don't really have a special need for an LLM in their life. They may like that it answers questions or helps you polish a resume or, I guess the labs' favorite, helps you make restaurant reservations. But this sure isn't worth $200/mo for most people. Probably not even worth $5/mo.
It'd be kinda funny if we create superhuman AGI and then no one has any real use for it, perhaps except for military murder-bots. There's always market for that.
> Meanwhile, human review and comprehension are starting to fall behind.
I think LLMs are valuable and spend most of my professional life working with them.
I do wonder though whether there’s a Ponzi scheme aspect here where as long as the “frontier” can keep outrunning human review and comprehension, LLMs are always going to looked way more valuable than they are and the bubble will continue.
This started with deep learning, expectations weren’t met and people started looking for value, then GPT came out and people got wooed again and forgot, then coding, then math, cyber, etc. As long as the dust doesn’t settle we never have to reflect on all the shortcomings and can just stare mesmerized at demos.
The fundamental question I have that honestly I haven't been able to find any answers for: With all the talk of LLM-based AI capabilities "going vertical", what evidence is there (for or against) that LLM-based approaches won't eventually "hit a wall", that is find some aspect of intelligence where humans will still have primacy, and no amount of scaling will change that.
E.g. some prominent folks definitely think LLM-based approaches will hit a wall, perhaps LeCun most notably. I also read a guest post on Terry Tao's site (which I really liked) that argues that, for all the very impressive recent AI results, they still operate within the "convex hull" of their training data: https://terrytao.wordpress.com/2026/09/13/happy-those-able-t...
I'm just curious if there is any actual data or evidence that takeoff (i.e. RSI, "the singularity", whatever you want to call it) is inevitable with current approaches.
He's intentionally forgoing all nuance in order to make this sound dramatic
> Somewhere around 20M people (0.2%) see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt.
No, they can't, at least not any semblance of quality. The cases we're seeing where this does kinda work is in ports and translation where all the rules are already documented in the best specification language possible with a way for the LLM to verify itself: code. We saw this close to a year ago now with Cloudflare and NextJS
> The impact scales with ambition, problem size, and horizon. A question with a paragraph answer barely stresses the system. You need a reservoir of big, difficult problems that you really care about
These are operating on different capabilities - AI's ability to answer informational queries as a chatbot frankly sucks and can't be trusted without verifying it. I run up against this every day. A problem with a verifiable answer on the other hand it's very good at solving. He knows this (his next paragraph) but he's putting them on the same scale of "stressing the system" to attempt to add proof to his introductory claim
And then there's the completely unverifiable scare that there are internal frontier models way beyond anything we've seen "swarms of thousands of agents collaborating over weeks on software mega projects: minting zero days, running cyber attacks and defenses at machine speeds, discovering new science, advancing the frontier of mathematics"
I dunno. I haven't been able to set up OpenAIs remote codex connection, their shit is buggy as hell and the web UI keeps crashing and making messages disappear. Is this what their internal superhuman "Things that would have taken top professionals in the industry years of work" looks like? Granted Claude has been really smooth, but still..
>Meanwhile, human review and comprehension are starting to fall behind. For example, people are still involved in the "archeology" of the OpenAI-HF incident from many months ago. Mathematicians may be poring over the 722 manuscripts on frontier mathematics for a while.
Amid all the discussion of sigmoid curves, and where the "LLM wall" will materialise, I think few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output.
What I fear is that people simply eschew human review altogether, considering we're talking about the industry that came up with the "move fast and break things" credo. Human review of LLM-produced code where I work is already a farce, and we're not special enough to be one of Karpathy's 5,000. I do my best to manually review anything that's my responsibility, but I'm literally one of very few people left working on my team, so in practice what happens is I submit PRs that are at best glossed over by completely unrelated teams for security, malware/prompt injection, and other serious concerns. Quality insofar as vetting others' code has completely gone out the window and it shows in the number of bug reports that come back, often themselves written in Claudease. This is all on top of everyone cynically phoning it in in the first place, due to the omnipresent sword of Damocles that is additional AI-driven layoffs.
Worse yet all the incentives point to this being the most economically viable thing individual companies can do. I think it goes without saying some type of regulation here is urgently needed, and that an unexpected cause of an AI bubble pop may end up being that humans simply aren't able to keep up with the pace of the output - leading either to precautionary plateauing of capability, or major liability risks related to a decline in quality.
I don't understand how "swarms of thousands of agents collaborating over weeks on software mega projects" works with the current context limits and at this point I'm too afraid to ask cause I'm afraid an AI bro is gonna punch me.
I mean in late 2020/early 2021, Altman and others were saying the end of work was 6 months out.
That clearly didn't happen :)
Am I the only one who thinks "yes, it is currently overhyped, but no, it will get there soon"?
[flagged]
Most people see a clumsy chatbot, most professionals see modest gains, and a tiny group is watching the curve go vertical, all at once.
The future is already here. It's just not very evenly distributed.
"running cyber attacks and defenses at machine speeds"; worth noting that LLMs (even the vaunted frontier models) are exponentially slower at cyberattacks or defenses than your average WannaCry or Splunk automation from decades ago. Its this sort of delusion world the AI bros live in that is part of the reason there is a disconnect between what they think should be happening and what actually is.
If the software coming out of OpenAI and Anthropic is what we have to judge, I wonder about the 5000 ...
Let's say, for the sake of argument, that the models are some multiplicative factor better on the inside.
Doesn't that mean the demos should work?
[flagged]
[flagged]
[dead]
[flagged]