如何将11GB的.csv文件转换为filebacked big.matrix对象?
从大CSV生成filebacked big.matrix的方法
因为11GB CSV无法直接加载到内存,核心思路是先创建空的磁盘备份big.matrix,再逐块读取CSV内容写入,全程避免加载完整文件。
步骤1:安装并加载必要包
install.packages(c("bigmemory", "data.table")) library(bigmemory) library(data.table)
步骤2:获取CSV的基础信息
先快速扫描CSV的表头、列数、列类型,同时估算总行数(留冗余避免不够):
# 读取表头 header <- fread("your_large_file.csv", nrows = 0) col_count <- ncol(header) col_types <- sapply(header, class) # 估算总行数(按样本比例计算) sample_rows <- fread("your_large_file.csv", nrows = 1000) sample_size <- object.size(sample_rows) total_file_size <- file.size("your_large_file.csv") estimated_rows <- as.integer((total_file_size / sample_size) * 1000) + 10000
步骤3:创建空的filebacked big.matrix
指定磁盘存储的路径和文件,匹配CSV的列类型:
# 创建磁盘备份的big.matrix fbm <- filebacked.big.matrix( nrow = estimated_rows, ncol = col_count, type = ifelse(all(col_types %in% c("integer", "numeric")), "double", "character"), backingfile = "my_fbm.bin", descriptorfile = "my_fbm.desc", backingpath = getwd() ) # 设置列名 colnames(fbm) <- colnames(header)
步骤4:逐块读取CSV并写入big.matrix
分批次读取写入,根据内存情况调整块大小:
# 设置每次读取的行数(按需调整,比如10万行) chunk_size <- 100000 start_row <- 1 while (TRUE) { # 读取当前块 chunk <- fread("your_large_file.csv", skip = start_row - 1, nrows = chunk_size, colClasses = col_types) # 无数据则退出循环 if (nrow(chunk) == 0) break # 计算写入范围 end_row <- start_row + nrow(chunk) - 1 # 写入big.matrix fbm[start_row:end_row, ] <- as.matrix(chunk) # 更新起始行 start_row <- end_row + 1 # 打印进度(可选) cat("已写入", end_row, "行\n") } # 裁剪多余空行(如果估算行数偏多) fbm <- fbm[1:(start_row - 1), ]
后续加载已生成的filebacked big.matrix
下次使用无需重新生成,直接加载描述文件:
fbm <- attach.big.matrix("my_fbm.desc")
注意事项
- 调整
chunk_size:内存充足可设大加快速度,内存紧张则设小。 - 列类型匹配:提前处理CSV中的混合类型,避免写入报错。
- 磁盘空间:备份文件大小和原CSV相近,确保磁盘有足够空间。
内容的提问来源于stack exchange,提问作者fil0607
相关产品推荐
相关产品推荐

