csvlite vs pandas, Polars & DuckDB · measured 29 September 2026

No code.
Faster than code.

Same files, same laptop, CSV to answer. Median of 5 runs.

vs DuckDB4.3× fastertypical · up to 12.3×
vs Polars2.2× fastertypical · up to 5.6× · out of memory at 10 GB
vs pandas57.3× fastertypical · up to 162× · out of memory at 10 GB
10 GB · 9 queries31 svs 111 s in DuckDB, the next fastest

Every query, every size

One dot per query and file size

How many times faster csvlite answers each query than DuckDB, Polars and pandas: one dot per query and file size, on a log scalecsvlite faster →0.5×2×5×10×20×50×100×200×same speed● DuckDB4.3×typicalfaster on 30/3610 MB100 MB1 GB10 GB● Polars2.2×typicalfaster on 34/3510 MB100 MB1 GB10 GB+1 out of memory● pandas57.3×typicalfaster on 27/2710 MB100 MB1 GB10 GBout of memory at 8 GB

The numbers

10 MB

53,596 rows · 15 columns · File in page cache

10 MB: median time from CSV to answer; the fastest tool in each row is bold
QuerycsvliteDuckDBPolarspandasvs fastest other
Text matchregion = north10 ms54 ms19 ms156 ms1.9× faster
Regex searchdescription has the word café14 ms58 ms21 ms190 ms1.4× faster
Number rangeamount 25,000–75,00010 ms56 ms18 ms157 ms1.9× faster
Date rangeoccurred_at in a range10 ms56 ms21 ms151 ms2.2× faster
Sort textevery row by region10 ms61 ms21 ms166 ms2.2× faster
Sort numbersevery row by amount10 ms64 ms20 ms162 ms2.0× faster
Sort datesevery row by occurred_at12 ms65 ms23 ms153 ms1.9× faster
Group + sumcount, sum(amount) by region, status12 ms64 ms20 ms167 ms1.7× faster
Group + distinct…plus distinct descriptions12 ms69 ms22 ms181 ms1.8× faster
Open once, keep askingTotal time as queries pile up. pandas loads the file once; Polars and DuckDB re-read it per query.
Total time after 9 queries, 10 MB: csvlite 40 ms, DuckDB 548 ms, Polars 184 ms, pandas 296 ms0 ms200 ms400 ms600 msopen123456789queries answered548 msDuckDB296 mspandas184 msPolars40 mscsvlite
Chart data
AftercsvliteDuckDBPolarspandas
Open10 ms0.0 ms0.0 ms148 ms
Text match12 ms54 ms19 ms151 ms
Regex search18 ms112 ms39 ms185 ms
Number range20 ms169 ms57 ms198 ms
Date range23 ms225 ms78 ms204 ms
Sort text25 ms286 ms99 ms222 ms
Sort numbers27 ms350 ms119 ms236 ms
Sort dates30 ms415 ms142 ms244 ms
Group + sum34 ms479 ms162 ms266 ms
Group + distinct40 ms548 ms184 ms296 ms

100 MB

532,923 rows · 15 columns · File in page cache

100 MB: median time from CSV to answer; the fastest tool in each row is bold
QuerycsvliteDuckDBPolarspandasvs fastest other
Text matchregion = north15 ms85 ms64 ms1.4 s4.4× faster
Regex searchdescription has the word café32 ms96 ms88 ms1.7 s2.7× faster
Number rangeamount 25,000–75,00021 ms88 ms79 ms1.6 s3.8× faster
Date rangeoccurred_at in a range19 ms89 ms90 ms1.5 s4.7× faster
Sort textevery row by region29 ms268 ms80 ms1.7 s2.7× faster
Sort numbersevery row by amount26 ms286 ms76 ms1.6 s2.9× faster
Sort datesevery row by occurred_at27 ms274 ms90 ms1.6 s3.4× faster
Group + sumcount, sum(amount) by region, status29 ms97 ms86 ms1.7 s3.0× faster
Group + distinct…plus distinct descriptions39 ms108 ms95 ms1.7 s2.5× faster
Open once, keep askingTotal time as queries pile up. pandas loads the file once; Polars and DuckDB re-read it per query.
Total time after 9 queries, 100 MB: csvlite 140 ms, DuckDB 1.4 s, Polars 746 ms, pandas 2.9 s0 s1 s2 s3 sopen123456789queries answered2.9 spandas1.4 sDuckDB746 msPolars140 mscsvlite
Chart data
AftercsvliteDuckDBPolarspandas
Open11 ms0.0 ms0.0 ms1.5 s
Text match14 ms85 ms64 ms1.5 s
Regex search35 ms180 ms152 ms1.8 s
Number range45 ms269 ms230 ms2.0 s
Date range53 ms358 ms320 ms2.0 s
Sort text69 ms625 ms400 ms2.2 s
Sort numbers83 ms911 ms475 ms2.4 s
Sort dates97 ms1.2 s565 ms2.5 s
Group + sum114 ms1.3 s651 ms2.7 s
Group + distinct140 ms1.4 s746 ms2.9 s

1 GB

5,300,205 rows · 15 columns · File in page cache

1 GB: median time from CSV to answer; the fastest tool in each row is bold
QuerycsvliteDuckDBPolarspandasvs fastest other
Text matchregion = north90 ms409 ms376 ms15 s4.2× faster
Regex searchdescription has the word café264 ms475 ms613 ms17 s1.8× faster
Number rangeamount 25,000–75,000172 ms408 ms473 ms16 s2.4× faster
Date rangeoccurred_at in a range139 ms395 ms779 ms15 s2.8× faster
Sort textevery row by region226 ms2.4 s535 ms17 s2.4× faster
Sort numbersevery row by amount197 ms2.4 s551 ms16 s2.8× faster
Sort datesevery row by occurred_at191 ms2.3 s766 ms15 s4.0× faster
Group + sumcount, sum(amount) by region, status236 ms439 ms620 ms16 s1.9× faster
Group + distinct…plus distinct descriptions362 ms563 ms716 ms18 s1.6× faster
Open once, keep askingTotal time as queries pile up. pandas loads the file once; Polars and DuckDB re-read it per query.
Total time after 9 queries, 1 GB: csvlite 1.4 s, DuckDB 9.9 s, Polars 5.4 s, pandas 29 s0 s10 s20 s30 sopen123456789queries answered29 spandas9.9 sDuckDB5.4 sPolars1.4 scsvlite
Chart data
AftercsvliteDuckDBPolarspandas
Open69 ms0.0 ms0.0 ms14 s
Text match99 ms409 ms376 ms14 s
Regex search303 ms884 ms989 ms18 s
Number range413 ms1.3 s1.5 s19 s
Date range491 ms1.7 s2.2 s20 s
Sort text654 ms4.1 s2.8 s22 s
Sort numbers791 ms6.5 s3.3 s24 s
Sort dates920 ms8.9 s4.1 s25 s
Group + sum1.1 s9.3 s4.7 s26 s
Group + distinct1.4 s9.9 s5.4 s29 s

10 GB

53,002,050 rows · 15 columns · Cold cache: evicted from memory before every run

10 GB: median time from CSV to answer; the fastest tool in each row is bold
QuerycsvliteDuckDBPolarspandasvs fastest other
Text matchregion = north4.6 s4.1 s7.6 sout of memory1.1× slower
Regex searchdescription has the word café6.2 s5.1 s10 sout of memory1.2× slower
Number rangeamount 25,000–75,0005.3 s4.4 s8.6 sout of memory1.2× slower
Date rangeoccurred_at in a range5.0 s4.7 s8.7 sout of memoryon par
Sort textevery row by region9.1 s29 s9.9 sout of memoryon par
Sort numbersevery row by amount6.4 s28 s9.7 sout of memory1.5× faster
Sort datesevery row by occurred_at6.3 s26 s9.1 sout of memory1.4× faster
Group + sumcount, sum(amount) by region, status5.7 s4.7 s11 sout of memory in 1 of 5 runsout of memory1.2× slower
Group + distinct…plus distinct descriptions7.2 s5.8 sout of memoryout of memory1.3× slower
Open once, keep askingTotal time as queries pile up. pandas loads the file once; Polars and DuckDB re-read it per query.
Total time after 9 queries, 10 GB: csvlite 31 s, DuckDB 111 s0 s50 s100 s150 sopen123456789queries answered111 sDuckDB31 scsvlite
Chart data
AftercsvliteDuckDBPolarspandas
Open3.1 s0.0 ms——
Text match4.7 s4.1 s——
Regex search7.8 s9.2 s——
Number range9.9 s14 s——
Date range12 s18 s——
Sort text18 s47 s——
Sort numbers21 s75 s——
Sort dates24 s101 s——
Group + sum27 s105 s——
Group + distinct31 s111 s——

Head to head

Machine
Intel Core i7-11850H · 32 GB RAM · NVMe SSD · Linux. Every run capped at 8 GB of memory.
Versions
csvlite 1.0.0 · DuckDB 1.5.5 · Polars 1.44.2 · pandas 3.0.6. All columns read as text.
Timing
CSV to answer, fresh process per run, startup excluded. ≤1 GB warm; 10 GB cold. Out of memory: one capped attempt.
Data
Generated analytics CSV. All 1064 raw runs (plus load-first) · csvlite-only benchmarks

Try it on your own CSV

Download csvlite