You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

tidyr spread函数异常:24列后整列填充0的问题排查与解决

大长表转宽表时的数据异常与内存问题

数据概况

现有一个包含1×10^7条观测、3个变量的数据框df,结构如下:

sp_id              sample       reads
         <char>             <char>       <int>
       1: sp1               sample1       255
       2: sp1               sample2       1
       3: sp1               sample3       1152
       4: sp2               sample1       114
       5: sp2               sample2       3
         ----               ----          ---
       10000000:  sp42500   sample6700    4554

数据包含6700个唯一sample值和42500个唯一sp_id(物种ID)。

预期目标

将长表转换为宽表,以sample作为列名,reads作为对应值,缺失值填充为0。预期结果如下:

sp_id              sample1       sample2      sample3    ...  sample6700
         <char>             <int>          <int>        <int>            <int> 
       1: sp1               255              1           1152            561
       2: sp2               114              3           0                3
          ----               ----          ---         ----             ----
       42500: sp42500        715             0           0               4554

遇到的问题

1. spread函数结果异常

使用以下代码转换时:

data <- df %>% spread(key=sample, value=nreads, fill=0)

结果出现异常:第24列之后的所有sample列值全为0,但实际这些样本中存在物种观测值。实际输出如下:

sp_id              sample24       sample25      sample26    ...  sample6700
         <char>             <int>          <int>        <int>            <int> 
       1: sp1               45              0              0                0
       2: sp2               3               0              0                0
          ----               ----          ---         ----             ----
       42500: sp42500        715             0           0                 0

2. pivot_wider函数报错

尝试使用pivot_wider时触发错误:

> data <- df %>%  pivot_wider(names_from = sample, values_from = reads, values_fill=0)
Error in `vec_rep_each()`:! `times` can't be missing. Location 1 is missing.Run `rlang::last_trace()` to see where the error occurred.Warning message:In nrow * ncol : NAs produced by integer overflow

3. dcast函数报错

尝试使用data.table的dcast时也报错:

> data <- dcast(setDT(df), sp_id ~ sample, value.var = "reads", fill=0)
Error: Cross product of elements provided to CJ() would result in 2853905958 rows which exceeds .Machine$integer.max == 2147483647

临时解决方法

将数据框拆分为两部分分别转换后合并,操作成功:

df1 <- df[c(1:5000000),]
df2 <- df[-c(1:5000000),]

data1 <- df1 %>% spread(key=sample, value=nreads, fill=0)
data2 <- df2 %>% spread(key=sample, value=nreads, fill=0)

data <- bind_rows(data1,data2)

内容的提问来源于stack exchange,提问作者Laura Drh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 21:40:14