如何在R中为mdm分组生成含指定百分比的扩展数据集
解决方案
可以用dplyr包实现简洁的分组处理,或者用基础R代码完成,两种方法如下:
方法一:使用dplyr包
# 定义需要的perc序列 perc_vec <- c(50, 60, 70, 80, 85, 90, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 110, 115, 120, 130, 140, 150) # 加载dplyr(未安装的话先运行 install.packages("dplyr")) library(dplyr) # 处理数据 result <- dat %>% group_by(mdm) %>% summarise( perc = perc_vec, price = ifelse(perc == 100, first(price), NA), count = ifelse(perc == 100, first(count), NA), .groups = "drop" ) %>% select(-mdm) # 移除mdm列,若需要保留可删除此行
方法二:基础R实现
perc_vec <- c(50, 60, 70, 80, 85, 90, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 110, 115, 120, 130, 140, 150) # 生成所有mdm与perc的组合 result_base <- expand.grid(mdm = dat$mdm, perc = perc_vec) # 合并原数据中的price和count result_base <- merge(result_base, dat, by = "mdm", all.x = TRUE) # 将非100的perc对应的price和count设为NA result_base$price[result_base$perc != 100] <- NA result_base$count[result_base$perc != 100] <- NA # 排序并整理列顺序 result_base <- result_base[order(result_base$mdm, result_base$perc), c("perc", "price", "count")] # 重置行名 row.names(result_base) <- NULL
两种方法最终都会生成你需要的结果:每个mdm分组对应完整的perc序列,仅perc=100的行填充原数据的price和count值,其余行均为NA。
内容的提问来源于stack exchange,提问作者psysky
相关产品推荐
相关产品推荐

