导入超大CSV至Pandas DataFrame时如何处理company_name中的反斜杠
Hey there! Let's work through this problem with your 25M-row CSV file—those backslashes in the company_name column are definitely causing parsing headaches, and ditching rows with error_bad_lines=False isn't ideal. Here are a few targeted solutions to keep all your data intact:
1. Explicitly Define Escape Character (Most Reliable)
The root issue is that Pandas defaults to treating backslashes as escape characters, which messes up values like HPE\HPI. By explicitly setting escapechar='\\' (note the double backslash for Python string escaping) and using the Python engine for more flexible parsing, you can tell Pandas to handle backslashes correctly:
import pandas as pd import csv df = pd.read_csv( "your_large_file.csv", engine="python", # Python engine supports more nuanced escape handling than the default C engine escapechar="\\", # Treat backslashes as literal escape characters (preserves them in the data) usecols=["dest_profile", "first_name", "last_name", "id", "con", "company_name"], # Only load needed columns to save memory low_memory=False, # Avoids type-inference warnings with large datasets quoting=csv.QUOTE_MINIMAL # Adjust this if your CSV uses quotes around fields (e.g., csv.QUOTE_ALL if all fields are quoted) )
If your company_name values are wrapped in quotes, the default doublequote=True will handle any internal quotes alongside the escapechar, so you won't lose data.
2. Disable Quoting Entirely (For Simple CSVs)
If your CSV doesn't use quotes to wrap fields (and none of your columns contain the comma separator), you can disable quoting entirely. This tells Pandas to treat backslashes as regular characters:
import pandas as pd import csv df = pd.read_csv( "your_large_file.csv", engine="python", sep=",", quoting=csv.QUOTE_NONE, # No characters are treated as quotes usecols=["dest_profile", "first_name", "last_name", "id", "con", "company_name"], low_memory=False )
Note: Skip this if your CSV has fields with commas inside them (e.g., company_name="Doe, Inc.")—disabling quoting will split those fields incorrectly.
3. Chunked Reading (For Memory Constraints)
With 25 million rows, loading the entire file at once might strain your memory. Combine the escapechar fix with chunked reading to process the file in smaller batches:
import pandas as pd import csv chunk_size = 100000 # Adjust based on your available memory (100k rows per chunk works for most systems) chunk_list = [] # Iterate over chunks and collect them for chunk in pd.read_csv( "your_large_file.csv", engine="python", escapechar="\\", usecols=["dest_profile", "first_name", "last_name", "id", "con", "company_name"], chunksize=chunk_size, low_memory=False ): chunk_list.append(chunk) # Combine chunks into a single DataFrame df = pd.concat(chunk_list, ignore_index=True)
This approach keeps memory usage manageable while ensuring all rows (including those with backslashes) are imported correctly.
内容的提问来源于stack exchange,提问作者emie

