R中使用fread()读取含多表头的大数据集失败,如何修复?
问题背景
我有两个文件夹:
big_data:包含1个约2GB的文件small_data:包含6个文件,总大小约150MB
需要将包含多行表头和空格的数据文件导入R,文件结构如下:
# File name # #@ 1 "Some text" "aa" #@ 2 "Some text" "bb" #@ 3 "Some text" "cc" #@ 4 "Some text" "dd" #@ 5 "Some text" "ee" #@ 6 "Some text" "ff" # # # # 0.000000e+00 0.000000e+00 0.000000e+00 0.000000e+00 0.000000e+00 0.000000e+00 1.000000e-03 3.727051e-04 2.532203e-04 4.736003e-04 3.727051e-07 0.000000e+00 2.000000e-03 2.266785e-03 1.540081e-03 2.880429e-03 2.639490e-06 0.000000e+00 3.000000e-03 7.538553e-03 5.121786e-03 9.579321e-03 1.017804e-05 0.000000e+00 4.000000e-03 1.838835e-02 1.249329e-02 2.336627e-02 2.856639e-05 0.000000e+00 5.000000e-03 3.703296e-02 2.516073e-02 4.705817e-02 6.559935e-05 0.000000e+00 6.000000e-03 2.266785e-03 1.540081e-03 2.880429e-03 2.639490e-06 0.000000e+00 7.000000e-03 7.538553e-03 5.121786e-03 9.579321e-03 1.017804e-05 0.000000e+00 8.000000e-03 1.838835e-02 1.249329e-02 2.336627e-02 2.856639e-05 0.000000e+00 9.000000e-03 3.703296e-02 2.516073e-02 4.705817e-02 6.559935e-05 0.000000e+00
该文件包含10行数据、6列,以#@开头的行是列名。
我编写了如下imp_data函数,根据指定文件夹选择read_table()(用于small_data)或fread()(用于big_data)导入数据,提取列名并移除NA列:
imp_data <- function(user_specified_folder){ # path of a single file for extracting header name foo <- here::here("data",species) p_header <- here::here("data",species, list.files(foo)[1]) # find skip value skip <- tail(grep("\\#", readLines(p_header)),1) # Get a List of all`.txt` files filenames <- list.files(here::here("data",species), pattern = "*.txt", full.names = T) # Load data sets if(user_specified_folder == "big_data"){ list.DFs <- lapply(filenames, fread , skip = skip) } else{ list.DFs <- lapply(filenames, read_table, col_names = F, skip = skip) } # name the data frame the same as file name names(list.DFs) <- basename(filenames) # extract column names from a dataset for a particular type rl <- readLines(p_header) nms <- rl[grep('^#@', rl)] |> sub(pattern='[^"]*"(.*)"\\s*$', replacement='\\1') |> gsub(pattern='[" ]+', replacement=' ') # remove NA columns list.DFs <- lapply(list.DFs, function(x) x[, colSums(is.na(x)) < nrow(x)]) # Apply column names to all datasets of a particular user_specified_folder list.DFs <- lapply(list.DFs, setNames, nms) return(list.DFs) }
调用data_big_data <- imp_data("big_data")时,fread()无报错,但仅导入列名(文件大小仅15kB),R会话内存占用约6GB。如何修改代码使fread()正常工作?
问题分析与解决方案
核心问题
- 未定义变量错误:函数中多处使用
species变量,但该变量未在函数参数中定义,会导致路径指向错误,读取的文件不符合预期。 skip参数计算逻辑错误:原代码取最后一个含#的行号作为skip值,但实际数据在该行之后还有空行,fread从该行开始读取会把空行当成数据起始,导致实际数据行被遗漏。fread自动识别表头干扰:fread默认会尝试识别表头,若skip后的行被误判为表头,会将后续数据行当成表头读取,最终只导入表头信息。
修改后的代码
imp_data <- function(user_specified_folder){ # 修正路径变量,使用传入的参数替代未定义的species foo <- here::here("data", user_specified_folder) p_header <- here::here("data", user_specified_folder, list.files(foo)[1]) # 精准计算需要跳过的行数:找到第一个实际数据行的位置 rl <- readLines(p_header) # 定位第一个非#开头且非空的行(实际数据起始行) data_start_line <- which(!grepl('^#', rl) & nchar(trimws(rl)) > 0)[1] skip <- data_start_line - 1 # 获取所有txt文件路径 filenames <- list.files(foo, pattern = "\\.txt$", full.names = TRUE) # 加载数据集 if(user_specified_folder == "big_data"){ # 指定col.names=FALSE,禁止fread自动识别表头 list.DFs <- lapply(filenames, fread, skip = skip, col.names = FALSE) } else { list.DFs <- lapply(filenames, read_table, col_names = FALSE, skip = skip) } # 给数据框命名为文件名 names(list.DFs) <- basename(filenames) # 提取列名:修正正则,同时提取两个引号内的内容 nms <- rl[grep('^#@', rl)] |> sub(pattern='[^"]*"(.*)"\\s*"(.*)"$', replacement='\\1 - \\2') |> gsub(pattern='[" ]+', replacement=' ') # 移除全NA列 list.DFs <- lapply(list.DFs, function(x) x[, colSums(is.na(x)) < nrow(x)]) # 应用列名 list.DFs <- lapply(list.DFs, setNames, nms) return(list.DFs) }
关键修改点说明
- 路径变量修正:将所有
species替换为传入的user_specified_folder,确保路径指向目标文件夹。 - 精准计算
skip值:直接定位第一个实际数据行的位置,计算需要跳过的行数,避免空行和注释行干扰。 - 禁用
fread自动表头识别:添加col.names=FALSE参数,强制fread将skip后的所有行都作为数据读取。 - 优化列名提取:修正正则表达式,同时提取
#@行中两个引号内的内容,生成更准确的列名(例如Some text - aa)。
内容的提问来源于stack exchange,提问作者ACE
相关产品推荐
相关产品推荐

