使用read.table与fread导入CSV文件的差异及特定列导入问题
data.table Great call switching to data.table::fread for your large file workflow—it’s hands down the fastest and most memory-efficient tool for this kind of task in R. Let’s fix up your code and address the common pitfalls you might be hitting:
Step 1: Set Up & Get Your File List
First, make sure you’re pointing to the right directory, and grab full file paths to avoid "file not found" errors:
library(data.table) # Get all CSV files (use full.names to include complete paths) my.files <- list.files(pattern = "\\.csv$", full.names = TRUE)
Step 2: Batch Read Only Your Target Columns
Your original read.table call skipped the first 6 rows (since your data starts at row 7) and used the 7th row as headers—don’t forget to replicate that in fread! Also, the select parameter needs explicit column positions or names (your c(1...) syntax won’t work):
# Define the columns you want to import (adjust positions/names to match your data) target_cols <- c(1, 2, 3, 4, 5) # Or use column names like c("X.run.number.", "scenario", "configuration", "col4", "col5") # Batch read each file, skip first 6 rows, keep only target columns my.data.list <- lapply(my.files, function(file) { fread( file, skip = 6, # Match your original read.table skip header = TRUE, # Use row 7 as column headers select = target_cols # Only import these columns ) })
Step 3: Merge All Files Into One Dataset
Use data.table’s rbindlist instead of do.call(rbind, ...)—it’s way faster for large datasets, and handles minor column mismatches with fill=TRUE:
# Combine all data.tables into one combined_dt <- rbindlist(my.data.list, fill = TRUE) # If all files have identical columns, you can omit fill=TRUE for a tiny speed boost
Common Issues to Troubleshoot
- Invalid
selectsyntax: Always list columns explicitly (e.g.,c(1,2,3)or named vectors). Never usec(1...)—that’s not valid R syntax. - Memory constraints: Even with just 5 columns, 40GB of raw data adds up. If you hit memory limits, consider:
- Processing files in smaller chunks (use
nrowsandskipinfreadto read partial files) - Exporting the combined dataset to a disk-based format like
fstorparquetmid-workflow to free up memory
- Processing files in smaller chunks (use
- Inconsistent headers: If some files have missing columns or different column names,
fill=TRUEwill fill missing values withNAinstead of throwing an error. - File encoding: If you see garbled text, add an
encodingparameter tofread(e.g.,encoding="UTF-8"orencoding="Latin-1").
内容的提问来源于stack exchange,提问作者Marijn

