As the video points out, there is a hardware/software feedback loop at play here. Since the hardware is deeply pipelined SIMD FPU datapath, the software is hand tuned machine coded transformations. In addition to limiting flexibility (every variation has to be ground out in CUDA), this prevents sparse matrix type optimisations. I predict that a set of general purpose CPUs - non shared memory at this scale - would be a much better use of transistor/power. And you can code that at high level give a CSP style approach such as https://bil-lang.org
show comments
tolugenius
I wonder what would "compute" mean if cpus were more efficient at matrix multiplication say 15 years ago. And on the flipside, what it would take to say train a frontier model entirely on cpus in the future.
As the video points out, there is a hardware/software feedback loop at play here. Since the hardware is deeply pipelined SIMD FPU datapath, the software is hand tuned machine coded transformations. In addition to limiting flexibility (every variation has to be ground out in CUDA), this prevents sparse matrix type optimisations. I predict that a set of general purpose CPUs - non shared memory at this scale - would be a much better use of transistor/power. And you can code that at high level give a CSP style approach such as https://bil-lang.org
I wonder what would "compute" mean if cpus were more efficient at matrix multiplication say 15 years ago. And on the flipside, what it would take to say train a frontier model entirely on cpus in the future.