使用data.table的fread读取CSV时卡在342637行,求解决方案
It sounds like you're hitting a stubborn edge case with fread where unbalanced unescaped quotes or malformed fields are throwing off the parser—especially frustrating when dealing with a massive dataset where manually checking the problematic line isn't feasible. Let's walk through some targeted fixes beyond the parameters you've already tested:
Recap of Your Attempts
You've already tried the most common quick fixes without luck:
fill=TRUEto handle column mismatches- Explicit
sep=","declaration blank.lines.skip=TRUEto skip empty rows
Alternative Solutions to Try
1. Disable Quoting Entirely
The error message specifically calls out unbalanced unescaped quotes as a potential culprit. Turning off quote handling can bypass this parser bottleneck:
user <- fread("user.csv", stringsAsFactors = FALSE, encoding = "UTF-8", quote='')
Note: This will treat all commas as separators, even those inside quoted fields. If your data has legitimate quoted commas, this might split fields incorrectly—but it's worth testing to see if it gets past line 342637.
2. Skip the Problematic Line Explicitly
Since you know the exact line number causing issues, split the file into two parts and combine them:
# Read all lines up to the problematic one user_part1 <- fread("user.csv", stringsAsFactors = FALSE, encoding = "UTF-8", nrows = 342636) # Read all lines after the problematic one user_part2 <- fread("user.csv", stringsAsFactors = FALSE, encoding = "UTF-8", skip = 342637) # Combine the two datasets user <- rbind(user_part1, user_part2)
Note: Double-check that the column structure is consistent between both parts to avoid binding errors.
3. Update data.table to the Latest Version
Your current version (1.10.4.3) is from 2018—fread has received dozens of bug fixes for quote handling and malformed lines since then. Updating might resolve the issue outright:
install.packages("data.table")
After updating, retry your original command—chances are the newer parser can handle the unbalanced quote scenario better.
4. Fall Back to readr for Better Error Handling
If fread still fails, use the readr package (part of the tidyverse) which has more robust error reporting and recovery for messy CSV files:
# Install if needed install.packages("readr") library(readr) user <- read_csv("user.csv", locale = locale(encoding = "UTF-8"), show_col_types = FALSE)
readr will flag exactly what's wrong with line 342637 and may auto-recover without skipping the entire line.
5. Inspect the Problematic Line Without Loading the Whole File
If you want to diagnose the issue directly without opening the massive CSV, use system commands to extract just line 342637:
# For Unix/macOS problem_line <- system("sed -n '342637p' user.csv", intern = TRUE) # For Windows (PowerShell) problem_line <- system('Get-Content user.csv | Select-Object -Index 342636', intern = TRUE) cat(problem_line)
This will show you exactly what's causing the column mismatch (e.g., an unclosed quote, extra commas) so you can fix it if needed.
Content of the question originates from Stack Exchange, question author joerna

