tidyr spread函数异常:24列后整列填充0的问题排查与解决
大长表转宽表时的数据异常与内存问题
数据概况
现有一个包含1×10^7条观测、3个变量的数据框df,结构如下:
sp_id sample reads <char> <char> <int> 1: sp1 sample1 255 2: sp1 sample2 1 3: sp1 sample3 1152 4: sp2 sample1 114 5: sp2 sample2 3 ---- ---- --- 10000000: sp42500 sample6700 4554
数据包含6700个唯一sample值和42500个唯一sp_id(物种ID)。
预期目标
将长表转换为宽表,以sample作为列名,reads作为对应值,缺失值填充为0。预期结果如下:
sp_id sample1 sample2 sample3 ... sample6700 <char> <int> <int> <int> <int> 1: sp1 255 1 1152 561 2: sp2 114 3 0 3 ---- ---- --- ---- ---- 42500: sp42500 715 0 0 4554
遇到的问题
1. spread函数结果异常
使用以下代码转换时:
data <- df %>% spread(key=sample, value=nreads, fill=0)
结果出现异常:第24列之后的所有sample列值全为0,但实际这些样本中存在物种观测值。实际输出如下:
sp_id sample24 sample25 sample26 ... sample6700 <char> <int> <int> <int> <int> 1: sp1 45 0 0 0 2: sp2 3 0 0 0 ---- ---- --- ---- ---- 42500: sp42500 715 0 0 0
2. pivot_wider函数报错
尝试使用pivot_wider时触发错误:
> data <- df %>% pivot_wider(names_from = sample, values_from = reads, values_fill=0) Error in `vec_rep_each()`:! `times` can't be missing. Location 1 is missing.Run `rlang::last_trace()` to see where the error occurred.Warning message:In nrow * ncol : NAs produced by integer overflow
3. dcast函数报错
尝试使用data.table的dcast时也报错:
> data <- dcast(setDT(df), sp_id ~ sample, value.var = "reads", fill=0) Error: Cross product of elements provided to CJ() would result in 2853905958 rows which exceeds .Machine$integer.max == 2147483647
临时解决方法
将数据框拆分为两部分分别转换后合并,操作成功:
df1 <- df[c(1:5000000),] df2 <- df[-c(1:5000000),] data1 <- df1 %>% spread(key=sample, value=nreads, fill=0) data2 <- df2 %>% spread(key=sample, value=nreads, fill=0) data <- bind_rows(data1,data2)
内容的提问来源于stack exchange,提问作者Laura Drh
相关产品推荐
相关产品推荐

