You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R data.table fread读取400万列CSV触发段错误的求助

问题背景

我尝试使用R的data.table包fread()函数读取一个包含约400万列、数百行的.csv文件,设置verbose=TRUE后触发段错误。当前运行环境为Linux机器,内存188Gb,目标文件大小约7-8Gb。需要确认该问题是否由内存导致,并寻求可行的解决思路。

报错信息
OpenMP version (_OPENMP)       201511
  omp_get_num_procs()            20
  R_DATATABLE_NUM_PROCS_PERCENT  unset (default 50)
  R_DATATABLE_NUM_THREADS        unset
  R_DATATABLE_THROTTLE           unset (default 1024)
  omp_get_thread_limit()         2147483647
  omp_get_max_threads()          20
  OMP_THREAD_LIMIT               unset
  OMP_NUM_THREADS                unset
  RestoreAfterFork               true
  data.table is using 10 threads with throttle==1024. See ?setDTthreads.
Input contains no \n. Taking this to be a filename to open
[01] Check arguments
  Using 10 threads (omp_get_max_threads()=20, nth=10)
  NAstrings = [<<NA>>]
  None of the NAstrings look like numbers.
  show progress = 0
  0/1 column will be read as integer
[02] Opening the file
  Opening file /q/combined.u.NA.ntwistbd.csv
  File opened, size = 7.265GB (7800965634 bytes).
  Memory mapped ok
[03] Detect and skip BOM
[04] Arrange mmap to be \0 terminated
  \n has been found in the input and different lines can end with different line endings (e.g. mixed \n and \r\n in one file). This is common and ideal.
[05] Skipping initial rows if needed
  Positioned on line 1 starting: <<gene,chr1.10469.10470.cpg_inte>>
[06] Detect separator, quoting rule, and ncolumns
  Detecting sep automatically ...
  sep=','  with 100 lines of 3986159 fields using quote rule 0
  Detected 3986159 columns on line 1. This line is either column names or first data row. Line starts as: <<gene,chr1.10469.10470.cpg_inte>>
  Quote rule picked = 0
  fill=false and the most number of columns found is 3986159
[07] Detect column types, good nrow estimate and whether first row is column names
  'header' changed by user from 'auto' to true
  Number of sampling jump points = 1 because (7800965633 bytes from row 1 to eof) / (2 * 3658680597 jump0size) == 1
  Type codes (jump 000)    : C7777777777777777777777777777777777775577755777777777777775772777755777777777777...2222222222  Quote rule 0
  Type codes (jump 001)    : C7777777777777777777777777777777777775577777777777777777777772777755777777777777...2222222222  Quote rule 0
  =====
  Sampled 153 rows (handled \n inside quoted fields) at 2 jump points
  Bytes from first data row on line 2 to the end of last row: 7456598287
  Line length: mean=33644907.61 sd=-nan min=18359868 max=54585593
  Estimated number of rows: 7456598287 / 33644907.61 = 222
  Initial alloc = 406 rows (222 + 82%) using bytes/max(mean-2*sd,min) clamped between [1.1*estn, 2.0*estn]
  =====
[08] Assign column names
[09] Apply user overrides on column types
  After 0 type and 0 drop user overrides : C7777777777777777777777777777777777775577777777777777777777772777755777777777777...2222222222
[10] Allocate memory for the datatable
  Allocating 3986159 column slots (3986159 - 0 dropped) with 406 rows
[11] Read the data
  jumps=[0..1), chunk_size=33644907614, total_size=7456598287
  2390 out-of-sample type bumps: C7777777777777777777777777777777777775577777777777777777777772777755777777777777...2222222222

 *** caught segfault ***
address (nil), cause 'unknown'
Segmentation fault (core dumped)
问题分析

188GB的物理内存理论上足以容纳该文件(文件仅7-8GB),但报错中的段错误(address (nil))更可能是内存分配逻辑问题而非物理内存不足:

  • fread在处理超宽表(400万列)时,内存分配的元数据管理可能出现指针异常;
  • 多线程环境下的内存竞争或类型检测时的临时内存溢出也可能触发段错误;
  • 自动类型检测过程中频繁的类型转换(日志中显示2390次类型提升)会增加内存波动,可能引发分配错误。
解决思路
  • 减少线程数:当前fread使用10线程,多线程在超宽表场景下可能引发内存竞争。先尝试单线程读取:
    setDTthreads(1)
    dt <- fread("your_file.csv", verbose=TRUE)
    
  • 手动指定列类型:禁用自动类型检测,提前指定列类型以避免频繁的内存重分配。例如第一列为字符型,其余为数值型:
    col_types <- c("character", rep("numeric", 3986158))
    dt <- fread("your_file.csv", colClasses=col_types, verbose=TRUE)
    
  • 转置文件格式:CSV为行优先存储,超宽表读取效率极低。先用shell工具将文件转置为行多列少的格式,再用fread读取:
    awk -F',' '{for(i=1;i<=NF;i++) a[i] = a[i] $i (NR==FNF?"":",")} END {for(i in a) print a[i]}' combined.u.NA.ntwistbd.csv > transposed.csv
    
  • 分块读取列:用select参数分批读取列,再合并结果。例如每次读取10万列:
    total_cols <- 3986159
    chunk_size <- 100000
    dt_list <- list()
    
    for (i in seq(1, total_cols, chunk_size)) {
      end_col <- min(i + chunk_size - 1, total_cols)
      dt_chunk <- fread("your_file.csv", select=i:end_col)
      dt_list[[length(dt_list)+1]] <- dt_chunk
    }
    
    dt <- do.call(cbind, dt_list)
    
  • 禁用内存映射:fread默认使用内存映射(mmap),超宽表下可能出现兼容性问题,尝试关闭:
    dt <- fread("your_file.csv", mmap=FALSE, verbose=TRUE)
    
  • 升级data.table版本:段错误可能是旧版本的已知bug,升级到最新版再尝试:
    install.packages("data.table")
    

内容的提问来源于stack exchange,提问作者719016

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 10:27:32