Why DuckDB 2.0 is faster

95 points22 comments4 hours ago
scythmic_waves

I also love the visualizations but I'm getting heavy LLM vibes from the prose:

> One setting drives this,...

> The cost is now about the rows you actually touch, not rounds times table size.

Etc.

I get the brain scramblies [1] from trying to parse this writing style at work so I hate to see it elsewhere. Apologies if I'm wrong. But if I'm not then OP don't use an LLM to write for you. It's hazardous to your reader's health [2].

[1]: https://www.youtube.com/watch?v=ipUJq-odt5Q

[2]: https://discourse.haskell.org/t/how-to-keep-enjoying-program...

show comments
stacktraceyo

Great visualization. Side note their new c++ extension api is also gonna be faster from the perspective of development / distribution of those extensions

jiggawatts

I wish more database engines used a Task-based design like Umbra / CedarDB.

Most of the DB engines out there still seem to use a "n-threads" style parallelism with exchange operations and poor async I/O management.

DuckDB is improving on this front, but in some sense is catching up to R&D (and implementation!) that is now decades old.

A "rhetorical challenge" I like to give software developers working on systems like this is the following: If I gave you a computer with 1,024 cores and matching network and storage bandwidth -- but with significant latency -- could you keep a system like this 100% utilised with one query?

The answer for almost all software is "no".

For example, SQL Server tops out at 64 hardware threads for any one query: https://learn.microsoft.com/en-us/sql/database-engine/config...

GPU codes are starting to get there, but CPU codes are way behind on this frontier of computer science.

It's not just databases! Can you (de)compress a file in parallel? Verify its hash in parallel? Upload/download from storage with CPU and I/O task parallelism? Can you overlap all of these operation so nothing is ever waiting on anything else it doesn't have to?

This matters! I ran some tests with bioinformatics codes and found that most got stuck in tar pits. Many could not scale to modern SSDs with millions of IOPS or modern networking with hundreds of gigabits of throughput, no matter how many CPU cores were thrown at them.

PS: AMD's Zen 6 era EPYC 9006 processors will have 512 cores and 1,024 threads per two-socket system, so this is not hypothetical: https://www.amd.com/en/products/processors/server/epyc/9006-...

larodi

...because Atlas and Fable took turns to go back and forth through it (source code and runtime) and track suspected bottlenecks.

show comments