Dask read_csv无法批量正常读取CSV但dask_cudf可行的技术咨询
read_csv() Not Reading All CSV Files Hey there! Let’s break down why your standard Dask DataFrames (dd) read_csv() call isn’t grabbing all your CSV files while dask_cudf works just fine. Here are the most likely culprits and fixes to get you sorted:
1. Verify File Path Matching First
Sometimes the glob pattern (*.csv) might not be matching all files as expected—Dask and dask_cudf can have subtle differences in how they resolve file paths. Start by confirming exactly which files your pattern is picking up:
import glob import os csv_files = glob.glob(os.path.join('data', '*.csv')) print(f"Total CSV files found: {len(csv_files)}") print("Files:", csv_files)
If the list is missing files, double-check:
- Are there files with uppercase
.CSVextensions? Use*.csvand*.CSV(or*.[Cc][Ss][Vv]) to cover both cases. - Is the
datadirectory relative to your script’s working directory? Try using an absolute path instead (e.g.,os.path.abspath('data')). - Do any filenames contain special characters or spaces? The glob should handle them, but testing individual file paths can rule out edge-case issues.
2. Standardize CSV Format Parameters
dask_cudf might be more lenient with inconsistent CSV formatting than standard Dask. Explicitly set parameters to match your files’ structure:
import dask.dataframe as dd da = dd.read_csv( os.path.join('data', '*.csv'), sep=',', # Replace with your actual delimiter (e.g., '\t' for tabs) encoding='utf-8', # Or 'latin-1' if files use that encoding on_bad_lines='skip', # Skip rows with formatting errors (adjust based on your Dask version) header=0, # Ensure all files use the first row as headers blocksize='64MB' # Force consistent partitioning (helps with small/large file mixes) ) print(da.head())
If some files have missing or mismatched headers, use the names parameter to define a universal column list:
da = dd.read_csv( os.path.join('data', '*.csv'), names=['col1', 'col2', 'col3'], # Replace with your actual column names header=None # Ignore existing headers if they’re inconsistent )
3. Update Dask to the Latest Version
Older Dask versions had bugs with glob pattern matching and CSV parsing that’ve been fixed in newer releases. Upgrade with:
pip install --upgrade dask[complete]
4. Test Individual Files
If the above steps don’t work, isolate the issue by testing one file at a time. Pick a file that you suspect isn’t being read and run:
test_df = dd.read_csv(os.path.join('data', 'problem_file.csv')) print(test_df.head())
Any error messages here will point directly to the problem (e.g., invalid encoding, malformed rows, or incompatible data types).
内容的提问来源于stack exchange,提问作者Krishna Roy

