The DuckDB team says one query in 2.0 runs about 40x faster. I reran that exact query on a 4-vCPU Linux VM and got 34x. Then I ran all 22 TPC-H queries at SF10 and six months of NYC taxi Parquet, and the gap shrank to roughly 20%.
Both numbers are real. Which one you get depends on whether your workload looks like a recursive graph walk or like ordinary joins and aggregates. This post has the DuckDB 2.0 vs 1.5.6 numbers, what the upgrade costs in memory, and what broke along the way: new database files, arrow lambdas and a community extension.
The setup: DuckDB 1.5.6 vs the 2.0 alpha, same machine
DuckDB 2.0 hasn't shipped yet. The official preview only says "this fall". What's on PyPI today are dev wheels, so "2.0" in this post means duckdb==2.0.0.dev2610011535, which reports itself as v2.0.0-alpha43763 (codename Cyanoptera). Alpha builds can change before release, so treat these numbers as a preview.
python3 -m venv v15 && v15/bin/pip install duckdb==1.5.6
python3 -m venv v20 && v20/bin/pip install duckdb==2.0.0.dev2610011535Hardware: a Linux VM capped at 4 vCPUs (AMD EPYC 7F32) and 8 GB of RAM, Python 3.13.5. That's closer to a mid-range laptop than to the M5 laptop used in MotherDuck's write-up.
Method:
- TPC-H SF10 (60M lineitem rows), generated once and exported to Parquet, then loaded into a native database by each version, so neither one reads a file the other wrote.
- I also ran 2.0 against the untouched 1.5.6 database file, to measure an "upgrade the package and change nothing else" scenario.
- Five NYC yellow taxi queries over January to June 2025 Parquet (24,083,384 rows, local disk).
- Every query gets a fresh process,
SET threads=4, one warm-up and 5 timed runs. I report the median time and the process's peak RSS.
TPC-H SF10 results: about 20% faster, with a few standouts
Summed medians across all 22 queries:
| | DuckDB 1.5.6 | 2.0 alpha, new file | 2.0 alpha, old 1.5 file |
|---|---|---|---|
| Total, 22 queries | 8.89 s | 7.08 s | 6.99 s |
| Geometric mean | 0.289 s | 0.238 s | 0.235 s |
That's about 1.26x on total time and 1.21x on geometric mean. Most of the gain comes from a handful of queries:
| Query | 1.5.6 | 2.0 alpha | Speedup |
|---|---|---|---|
| Q19 | 0.469 s | 0.160 s | 2.94x |
| Q02 | 0.082 s | 0.045 s | 1.81x |
| Q21 | 1.215 s | 0.689 s | 1.76x |
| Q22 | 0.235 s | 0.150 s | 1.57x |
| Q13 | 1.024 s | 0.666 s | 1.54x |
| Q04 | 0.324 s | 0.224 s | 1.45x |
| Q01 | 0.411 s | 0.300 s | 1.37x |
| Q09 | 1.383 s | 1.366 s | 1.01x |
| Q12 | 0.225 s | 0.219 s | 1.03x |
Q09 is the slowest query in the suite, and its time didn't change. Its memory did: peak RSS fell from 2,460 MB to 1,686 MB. Q14, Q20 and Q11 came out flat or a few milliseconds slower. Q06 looked like a regression in the first pass (0.094 s vs 0.117 s), so I reran it with 15 iterations and it came back at 0.096 s vs 0.092 s, which is noise. (DuckDB's built-in Q11 uses the SF1 threshold constant and returns zero rows at SF10, so ignore it.)
The column I didn't expect is the last one in the first table: 2.0 running on the old 1.5 file was as fast as on a freshly written 2.0 file. On this workload the new storage format bought no query speed at all, so the gain I measured comes from the engine.
The new format also isn't smaller here. Loading the same Parquet gave a 2,740,465,664-byte file on 1.5.6 and 2,870,226,944 bytes on 2.0, about 4.7% bigger. TPC-H is mostly numbers, though, and the preview credits the storage changes with faster opening of wide tables and big indexes, so your schema may look different.
The 40x claim holds up, for recursive CTEs only
The DuckDB preview's microbenchmark is a reachability query over a million-edge graph:
CREATE TABLE edges AS
SELECT (range % 100_000)::INTEGER AS src,
((range * 13 + 7) % 100_000)::INTEGER AS dst
FROM range(1_000_000);
WITH RECURSIVE reachable(node) AS (
SELECT 0
UNION
SELECT dst FROM edges, reachable WHERE src = node
)
SELECT count(*) FROM reachable;They report 4.90 s on 1.5.4 and 0.12 s on the preview. On my 4 vCPUs, 1.5.6 took 15.3 s (with runs anywhere from 14.5 to 22.2 s) and 2.0 took 0.449 s, with almost no variance. That's 34x. The recursive CTE rework keeps execution state between iterations, including the hash table built on edges, so the table is scanned once and each iteration only probes it with the new rows. Before, 1.5.6 rescanned the whole edge table on every pass.
If you walk lineage graphs or git histories in SQL, this alone justifies upgrading. If your hierarchies are three levels deep, you won't notice it.
NYC taxi Parquet: modest gains, a bit more memory
| Query (24M rows, local Parquet) | 1.5.6 | 2.0 alpha | Speedup | Peak RSS 1.5.6 → 2.0 |
|---|---|---|---|---|
| Hourly percentiles + distinct zones | 0.819 s | 0.620 s | 1.32x | 594 → 549 MB |
| Top routes, median distance | 0.772 s | 0.618 s | 1.25x | 544 → 653 MB |
| Trips by month | 0.529 s | 0.478 s | 1.11x | 119 → 219 MB |
| Top-10 per zone window | 0.136 s | 0.128 s | 1.07x | 107 → 153 MB |
| Selective filter | 0.095 s | 0.094 s | 1.01x | 115 → 234 MB |
On local files, async I/O didn't do much. The preview itself says "Local storage benefits a little too, but network storage is where you will see the big gains". MotherDuck measured a query over a 2.2 GB Parquet file on S3 going from 18.8 s on 1.5.5 to 7.7 s on the 2.0 alpha. I didn't test S3, so I can't confirm that. Their explanation is a separate pool of threads that only downloads and keeps row groups in flight, and the alpha does ship a new setting for it: read_ahead_depth, default -1 ("automatic, with the backlog being bound by a memory budget"; 0 disables it).
Memory is where 2.0 costs you something. On the three light scans it sat 46 to 119 MB higher at peak, and I'd guess the read-ahead buffers are part of that. I didn't test whether SET read_ahead_depth = 0 brings it back down. If you run DuckDB inside a tight Lambda or container memory limit, measure before you upgrade.
What breaks when you upgrade to DuckDB 2.0
I expected the file format to be the painful part, and it turned out to be the easiest.
Old database files open fine
2.0 opened my 1.5.6 file, queried it, wrote a new table and checkpointed. duckdb_databases() still reported storage_version: v1.0.0+ afterwards, and 1.5.6 could still read the file, new table included. One third-party summary says v2.0 can't open existing .duckdb files directly and that you need EXPORT/IMPORT. That didn't match what I saw with the alpha.
New files are one-way
A database created by 2.0 is unreadable to 1.5.6:
IO Error: Trying to read a database file with version number 69, but we can only read versions between 64 and 68.
The database file was created with a newer version of DuckDB.If one machine in your pipeline stays on 1.x, either keep the old file or pin the format when you attach. This worked, and 1.5.6 read the result:
ATTACH 'shared.duckdb' AS shared (STORAGE_VERSION 'v1.5.0');To move to the new format on purpose, COPY FROM DATABASE old TO new took 19.3 s for SF10.
Arrow lambdas are off by default
This is the one that will break your queries:
Binder Error: Deprecated lambda arrow (->) detected. Please transition to the new lambda syntax, i.e.., lambda x, i: x + i, before DuckDB's next release.
Use SET lambda_syntax='ENABLE_SINGLE_ARROW' to revert to the deprecated behavior.Rewrite list_transform(l, x -> x + 1) as list_transform(l, lambda x: x + 1). The docs say v2.1 removes the escape hatch, so grep your SQL now: grep -rn --include='*.sql' -e '->' . will flag JSON operators too, but it's a start.
Python package
duckdb.typing and duckdb.functional are gone, replaced by duckdb.sqltypes and duckdb.func. They were already missing from the 1.5.6 wheel, so if you're on 1.5.6 you've already dealt with this.
Extensions
httpfs, spatial, postgres_scanner, excel, delta and iceberg installed and loaded on the alpha. The community extension h3 returned HTTP 404 for the v2.0.0-alpha43763 build. Expect community extensions to arrive gradually after release, and check every one your code LOADs.
My upgrade checklist
-> lambdas and switch to lambda x:.STORAGE_VERSION 'v1.5.0'.My verdict: upgrade once the release is out, for the engine work rather than the new file format. The 40x headline is real but narrow. Most of us will get the 20% on ordinary joins and aggregates, which is still a decent improvement for a version bump.
Are you upgrading when 2.0 ships, or waiting for your extensions to catch up? If you've run the alpha on your own data, post your numbers in the comments, especially if they contradict mine.
Sources
- A Preview of DuckDB v2.0, DuckDB blog
- Why DuckDB 2.0 is faster, MotherDuck, and the HN discussion
- Lambda functions docs
- NYC TLC trip record data
