求助:fread读取5G CSV文件因引号导致数据结构识别错误
Hey there, I’ve run into this exact issue before with massive CSV files—fread’s speed is unbeatable, but it can get tripped up by quote formatting that read.csv handles more gracefully. Let’s break down how to fix this:
Why this happens
Chances are, your CSV has numeric values wrapped in quotes (either double " or single '). fread prioritizes speed over aggressive type guessing by default, so if it sees any quotes around a value, it’ll treat the entire column as character to avoid misinterpretation. read.csv, on the other hand, automatically strips quotes and converts quoted numbers to numeric types.
Solutions to try
1. Disable quote parsing entirely
Tell fread to ignore all quotes with the quote parameter. This is usually the quickest fix:
dt <- fread("data.csv", header = TRUE, quote = "")
This makes fread skip checking for quotes entirely, letting it correctly identify numeric columns that were wrapped in quotes.
2. Strip whitespace alongside quote handling
Sometimes extra spaces inside quotes can throw off type detection. Add strip.white = TRUE to clean up any leading/trailing whitespace in values:
dt <- fread("data.csv", header = TRUE, strip.white = TRUE)
3. Manually specify column types (if the above don’t work)
If the automatic fixes fail, you can first sample a small portion of the data to get the correct column types, then apply them to the full file:
# Read the first 100 rows to detect types sample_data <- fread("data.csv", header = TRUE, nrows = 100) # Extract the correct classes from the sample col_classes <- sapply(sample_data, class) # Read the full file with the pre-detected types dt <- fread("data.csv", header = TRUE, colClasses = col_classes)
This avoids the guesswork and ensures fread uses the right types from the start, without sacrificing speed.
After trying these, run str(dt) again—you should see numeric columns correctly identified just like with read.csv!
内容的提问来源于stack exchange,提问作者MLE

