Really dumb question from a software guy. Why aren't the labs burning their frontier models into chips already? Seems like the performance gains and cost per request would be worth it. That said, I understand neither the economics nor the physical challenges to doing this.
show comments
rcarmo
Well, as long as it doesn't start developing anatomically accurate metal skeletons with red glowing eyes...
show comments
athrowaway3z
I haven't really dug into the results yet, but my guess is
that a SOTA model has been able to produce an accelerator that runs a model since around December.
The obvious next step is to get enough memory throughput to run that SOTA model itself so that it develop its own hardware.
But perhaps the more interesting question is this: Can an AI be given a big FPGA and design a model architecture that takes advantage of the fabric being reconfigurable.
show comments
fsbonetto
After using AI to develop risc-v CPU cores, the same technique was used for developing openTPU. An open source AI inference engine. It's able to run most of the modern models like Qwen 3.5, Gemma 4, and many others. The TPU started able to produce only a few tokens per second and trough a recursive self improvement loop got to 80+ tok/sec on the smallers models.
xg15
"Recursive self-improvement will kill us all!"
Also: Here is our recursive self-improvement hard at work...
show comments
vatsachak
I feel like there is a lot to be gained from an experienced user pointing an LLM in a tasteful direction.
show comments
random__duck
Opened the RTL, looked at the floating point math, learned that apparently you don't need correct floating point operations for LLMs, closed the page.
This seems to be running on an FPGA board that costs ~$300? Anyone know more about the hardware?
show comments
bitwize
Colossus is building Colossus II.
show comments
AnimalMuppet
Can anyone comment on the performance of this hardware? How does it compare to state of the art, human-designed hardware? Is this actually an improvement? (To get to recursive self-improvement, you first have to improve at all.)
show comments
srameshc
This post brings me to question "What does it mean to be a software developer in future" ?
show comments
deepsun
Bulldozers, excavators and rollers are now capable of building roads.
gfalcao
The birth of SkyNet
mbgerring
> AI is now capable of developing its own inference hardware
No, it isn't.
A human prompted an LLM to build a software simulation environment for hardware design, enabling an LLM, when prompted by a human, to optimize hardware designs against constraints in the simulation.
show comments
rfgplk
Yep, 99.9% of people are completely oblivious to what LLMs can do. Just wait until the next gen of CPUs/GPUs designed by LLMs start coming out (fyi chip development tools have advanced centuries in the last few months) and you'll start seeing exponential gains in hardware.
Really dumb question from a software guy. Why aren't the labs burning their frontier models into chips already? Seems like the performance gains and cost per request would be worth it. That said, I understand neither the economics nor the physical challenges to doing this.
Well, as long as it doesn't start developing anatomically accurate metal skeletons with red glowing eyes...
I haven't really dug into the results yet, but my guess is that a SOTA model has been able to produce an accelerator that runs a model since around December.
The obvious next step is to get enough memory throughput to run that SOTA model itself so that it develop its own hardware.
But perhaps the more interesting question is this: Can an AI be given a big FPGA and design a model architecture that takes advantage of the fabric being reconfigurable.
After using AI to develop risc-v CPU cores, the same technique was used for developing openTPU. An open source AI inference engine. It's able to run most of the modern models like Qwen 3.5, Gemma 4, and many others. The TPU started able to produce only a few tokens per second and trough a recursive self improvement loop got to 80+ tok/sec on the smallers models.
"Recursive self-improvement will kill us all!"
Also: Here is our recursive self-improvement hard at work...
I feel like there is a lot to be gained from an experienced user pointing an LLM in a tasteful direction.
Opened the RTL, looked at the floating point math, learned that apparently you don't need correct floating point operations for LLMs, closed the page.
I've been vibecoding an open source hardware AV1 decoder: https://github.com/Jonahss/openav1
This seems to be running on an FPGA board that costs ~$300? Anyone know more about the hardware?
Colossus is building Colossus II.
Can anyone comment on the performance of this hardware? How does it compare to state of the art, human-designed hardware? Is this actually an improvement? (To get to recursive self-improvement, you first have to improve at all.)
This post brings me to question "What does it mean to be a software developer in future" ?
Bulldozers, excavators and rollers are now capable of building roads.
The birth of SkyNet
> AI is now capable of developing its own inference hardware
No, it isn't.
A human prompted an LLM to build a software simulation environment for hardware design, enabling an LLM, when prompted by a human, to optimize hardware designs against constraints in the simulation.
Yep, 99.9% of people are completely oblivious to what LLMs can do. Just wait until the next gen of CPUs/GPUs designed by LLMs start coming out (fyi chip development tools have advanced centuries in the last few months) and you'll start seeing exponential gains in hardware.