You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R中使用fread()读取含多表头的大数据集失败,如何修复?

问题背景

我有两个文件夹:

  • big_data:包含1个约2GB的文件
  • small_data:包含6个文件,总大小约150MB

需要将包含多行表头和空格的数据文件导入R,文件结构如下:

# File name
#
#@   1  "Some text"                                                   "aa"
#@   2  "Some text"                                                   "bb"
#@   3  "Some text"                                                   "cc"
#@   4  "Some text"                                                   "dd"
#@   5  "Some text"                                                   "ee"
#@   6  "Some text"                                                   "ff"

#
#
#
#

 0.000000e+00  0.000000e+00  0.000000e+00  0.000000e+00  0.000000e+00  0.000000e+00  
 1.000000e-03  3.727051e-04  2.532203e-04  4.736003e-04  3.727051e-07  0.000000e+00   
 2.000000e-03  2.266785e-03  1.540081e-03  2.880429e-03  2.639490e-06  0.000000e+00  
 3.000000e-03  7.538553e-03  5.121786e-03  9.579321e-03  1.017804e-05  0.000000e+00   
 4.000000e-03  1.838835e-02  1.249329e-02  2.336627e-02  2.856639e-05  0.000000e+00  
 5.000000e-03  3.703296e-02  2.516073e-02  4.705817e-02  6.559935e-05  0.000000e+00 
 6.000000e-03  2.266785e-03  1.540081e-03  2.880429e-03  2.639490e-06  0.000000e+00  
 7.000000e-03  7.538553e-03  5.121786e-03  9.579321e-03  1.017804e-05  0.000000e+00   
 8.000000e-03  1.838835e-02  1.249329e-02  2.336627e-02  2.856639e-05  0.000000e+00  
 9.000000e-03  3.703296e-02  2.516073e-02  4.705817e-02  6.559935e-05  0.000000e+00

该文件包含10行数据、6列,以#@开头的行是列名。

我编写了如下imp_data函数,根据指定文件夹选择read_table()(用于small_data)或fread()(用于big_data)导入数据,提取列名并移除NA列:

imp_data <- function(user_specified_folder){
  
  # path of a single file for extracting header name
  foo      <- here::here("data",species)
  p_header <- here::here("data",species, list.files(foo)[1])
  
  # find skip value
  
  skip <- tail(grep("\\#", readLines(p_header)),1)
  
  # Get a List of all`.txt` files
  filenames <- list.files(here::here("data",species), pattern = "*.txt", full.names = T)
  

  # Load data sets
  if(user_specified_folder == "big_data"){
    
    list.DFs  <- lapply(filenames, fread , skip = skip)
  }
  
  else{
    
    list.DFs  <- lapply(filenames, read_table, col_names = F, skip = skip)
  }
  
  # name the data frame the same as file name
  names(list.DFs) <- basename(filenames)

  # extract column names from a dataset for a particular type
  rl <- readLines(p_header)
  nms <- rl[grep('^#@', rl)] |>
    sub(pattern='[^"]*"(.*)"\\s*$', replacement='\\1') |> 
    gsub(pattern='[" ]+', replacement=' ')
  
  # remove NA columns
  list.DFs <- lapply(list.DFs, function(x) x[, colSums(is.na(x)) < nrow(x)])
  
  # Apply column names to all datasets of a particular user_specified_folder
  list.DFs <- lapply(list.DFs, setNames, nms)
  
  return(list.DFs)
}

调用data_big_data <- imp_data("big_data")时,fread()无报错,但仅导入列名(文件大小仅15kB),R会话内存占用约6GB。如何修改代码使fread()正常工作?


问题分析与解决方案

核心问题

  1. 未定义变量错误:函数中多处使用species变量,但该变量未在函数参数中定义,会导致路径指向错误,读取的文件不符合预期。
  2. skip参数计算逻辑错误:原代码取最后一个含#的行号作为skip值,但实际数据在该行之后还有空行,fread从该行开始读取会把空行当成数据起始,导致实际数据行被遗漏。
  3. fread自动识别表头干扰:fread默认会尝试识别表头,若skip后的行被误判为表头,会将后续数据行当成表头读取,最终只导入表头信息。

修改后的代码

imp_data <- function(user_specified_folder){
  
  # 修正路径变量,使用传入的参数替代未定义的species
  foo      <- here::here("data", user_specified_folder)
  p_header <- here::here("data", user_specified_folder, list.files(foo)[1])
  
  # 精准计算需要跳过的行数:找到第一个实际数据行的位置
  rl <- readLines(p_header)
  # 定位第一个非#开头且非空的行(实际数据起始行)
  data_start_line <- which(!grepl('^#', rl) & nchar(trimws(rl)) > 0)[1]
  skip <- data_start_line - 1
  
  # 获取所有txt文件路径
  filenames <- list.files(foo, pattern = "\\.txt$", full.names = TRUE)
  
  # 加载数据集
  if(user_specified_folder == "big_data"){
    # 指定col.names=FALSE,禁止fread自动识别表头
    list.DFs  <- lapply(filenames, fread, skip = skip, col.names = FALSE)
  } else {
    list.DFs  <- lapply(filenames, read_table, col_names = FALSE, skip = skip)
  }
  
  # 给数据框命名为文件名
  names(list.DFs) <- basename(filenames)
  
  # 提取列名:修正正则,同时提取两个引号内的内容
  nms <- rl[grep('^#@', rl)] |>
    sub(pattern='[^"]*"(.*)"\\s*"(.*)"$', replacement='\\1 - \\2') |>
    gsub(pattern='[" ]+', replacement=' ')
  
  # 移除全NA列
  list.DFs <- lapply(list.DFs, function(x) x[, colSums(is.na(x)) < nrow(x)])
  
  # 应用列名
  list.DFs <- lapply(list.DFs, setNames, nms)
  
  return(list.DFs)
}

关键修改点说明

  • 路径变量修正:将所有species替换为传入的user_specified_folder,确保路径指向目标文件夹。
  • 精准计算skip值:直接定位第一个实际数据行的位置,计算需要跳过的行数,避免空行和注释行干扰。
  • 禁用fread自动表头识别:添加col.names=FALSE参数,强制fread将skip后的所有行都作为数据读取。
  • 优化列名提取:修正正则表达式,同时提取#@行中两个引号内的内容,生成更准确的列名(例如Some text - aa)。

内容的提问来源于stack exchange,提问作者ACE

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 07:35:16