You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用read.table与fread导入CSV文件的差异及特定列导入问题

Efficiently Importing & Merging Large CSV Files with data.table

Great call switching to data.table::fread for your large file workflow—it’s hands down the fastest and most memory-efficient tool for this kind of task in R. Let’s fix up your code and address the common pitfalls you might be hitting:

Step 1: Set Up & Get Your File List

First, make sure you’re pointing to the right directory, and grab full file paths to avoid "file not found" errors:

library(data.table)
# Get all CSV files (use full.names to include complete paths)
my.files <- list.files(pattern = "\\.csv$", full.names = TRUE)

Step 2: Batch Read Only Your Target Columns

Your original read.table call skipped the first 6 rows (since your data starts at row 7) and used the 7th row as headers—don’t forget to replicate that in fread! Also, the select parameter needs explicit column positions or names (your c(1...) syntax won’t work):

# Define the columns you want to import (adjust positions/names to match your data)
target_cols <- c(1, 2, 3, 4, 5) # Or use column names like c("X.run.number.", "scenario", "configuration", "col4", "col5")

# Batch read each file, skip first 6 rows, keep only target columns
my.data.list <- lapply(my.files, function(file) {
  fread(
    file,
    skip = 6,          # Match your original read.table skip
    header = TRUE,     # Use row 7 as column headers
    select = target_cols # Only import these columns
  )
})

Step 3: Merge All Files Into One Dataset

Use data.table’s rbindlist instead of do.call(rbind, ...)—it’s way faster for large datasets, and handles minor column mismatches with fill=TRUE:

# Combine all data.tables into one
combined_dt <- rbindlist(my.data.list, fill = TRUE)
# If all files have identical columns, you can omit fill=TRUE for a tiny speed boost

Common Issues to Troubleshoot

  • Invalid select syntax: Always list columns explicitly (e.g., c(1,2,3) or named vectors). Never use c(1...)—that’s not valid R syntax.
  • Memory constraints: Even with just 5 columns, 40GB of raw data adds up. If you hit memory limits, consider:
    • Processing files in smaller chunks (use nrows and skip in fread to read partial files)
    • Exporting the combined dataset to a disk-based format like fst or parquet mid-workflow to free up memory
  • Inconsistent headers: If some files have missing columns or different column names, fill=TRUE will fill missing values with NA instead of throwing an error.
  • File encoding: If you see garbled text, add an encoding parameter to fread (e.g., encoding="UTF-8" or encoding="Latin-1").

内容的提问来源于stack exchange,提问作者Marijn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:07:46