Dask读取CSV报错但Pandas可正常读取的问题及解决方法
Hey folks, I've dealt with this exact head-scratcher before—you've got a CSV that loads perfectly with Pandas, but Dask throws a frustrating ParserError when you try to process it. Let's break down what's happening and how to fix it.
The Problem Reproduced
First, let's confirm the scenario:
- Pandas works flawlessly:
import pandas as pd pdf = pd.read_csv("./tous_les_docs.csv") print(pdf.shape) # Output: (20140796, 7) - Dask throws an error:
The error you'll see:import dask.dataframe as dd df = dd.read_csv("./tous_les_docs.csv") df.describe().compute()ParserError: Error tokenizing data. C error: EOF inside string starting at line 192999
Why This Happens
The root cause is how Dask handles large files by default: it splits the file into smaller blocks to enable parallel processing. If one of those blocks cuts off mid-way through a quoted string (like a value with commas inside quotes), the parser hits the end of the block before the string is closed—hence the "EOF inside string" error.
Pandas doesn't have this issue because it reads the entire file in one go, so it never truncates a string mid-parsing.
The Fix
To resolve this, just add the blocksize=None parameter to dd.read_csv(). This tells Dask to read the entire file as a single block (instead of splitting it), which avoids the truncation issue:
import dask.dataframe as dd # This will work without parsing errors df = dd.read_csv("./tous_les_docs.csv", blocksize=None) df.describe().compute()
Note that this does mean you lose the parallel processing benefits for reading the file, but it's a necessary trade-off if your CSV has unbalanced quoted strings that break block-based parsing.
内容的提问来源于stack exchange,提问作者Romain Jouin

