You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中从汇总数据高效生成含0和1的二项式DataFrame?

高效展开二项式感染数据的R实现方法

针对你需要将汇总感染数据展开为个体级二项标记数据的需求,以下是两种优雅且高效的实现方式,避免手动硬编码rep()的繁琐:

方法1:Tidyverse工具链(简洁易读,推荐)

利用purrr的映射函数结合tidyr的行展开功能,代码逻辑清晰,适合后续维护:

library(tidyverse)

# 原始数据集
df <- data.frame(day=seq(1,3,1), tot.infected=c(1,2,4), tot.ind=5)

# 生成目标数据框
df2 <- df %>%
  # 为每行生成对应长度的感染标记向量
  mutate(infected = map2(tot.infected, tot.ind, ~c(rep(1, .x), rep(0, .y - .x)))) %>%
  # 将向量拆分为单独行
  unnest_longer(infected) %>%
  # 重命名day为year(匹配你的示例输出)
  rename(year = day) %>%
  # 保留需要的列
  select(year, infected)

# 查看结果
df2

方法2:Base R实现(无需额外依赖)

如果不想加载第三方包,可使用Base R的lapply批量处理行数据,效率同样出色:

# 原始数据集
df <- data.frame(day=seq(1,3,1), tot.infected=c(1,2,4), tot.ind=5)

# 遍历每行生成子数据框,再合并
df_list <- lapply(1:nrow(df), function(i) {
  data.frame(
    year = rep(df$day[i], df$tot.ind[i]),
    infected = c(rep(1, df$tot.infected[i]), rep(0, df$tot.ind[i] - df$tot.infected[i]))
  )
})

df2 <- do.call(rbind, df_list)

# 查看结果
df2

两种方法都能自动适配任意行数的原始数据集,无需手动调整rep()的参数,在大数据集上的表现远优于手动硬编码,既减少出错概率,又提升代码可维护性。

内容的提问来源于stack exchange,提问作者Alexander Grimaudo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 10:40:30