如何用R函数将数据框中满足条件的行拆分为多行?
在R语言中拆分大数量行至每行数量小于1e6
示例原始数据
首先构造你提供的原始数据集:
df <- data.frame( X1 = c("a", "b", "c", "d"), X2 = c(1000000, 2000000, 2, 40) )
方法一:使用tidyverse工具(简洁高效)
借助tidyverse的函数可以快速实现需求,代码逻辑清晰且运行效率更高:
library(tidyverse) # 定义拆分逻辑的函数 split_large_values <- function(item, count, threshold = 1e6, chunk = 1e5) { if (count < threshold) { return(tibble(X1 = item, X2 = count)) } # 计算完整拆分块的数量和剩余值 full_chunks <- count %/% chunk remainder <- count %% chunk # 生成所有拆分后的行 chunks <- rep(chunk, full_chunks) if (remainder > 0) chunks <- c(chunks, remainder) tibble(X1 = rep(item, length(chunks)), X2 = chunks) } # 应用函数并合并结果 result <- df %>% rowwise() %>% group_split() %>% map_dfr(~split_large_values(.$X1, .$X2)) # 查看完整结果 print(result, n = Inf)
方法二:使用for循环(符合你的初始思路)
如果你倾向于用循环实现拆分逻辑,以下代码完全匹配你的需求:
# 初始化结果数据框 result <- data.frame(X1 = character(), X2 = integer(), stringsAsFactors = FALSE) # 设定阈值和拆分块大小 threshold <- 1e6 chunk_size <- 1e5 # 遍历每一行处理 for (i in 1:nrow(df)) { current_item <- df$X1[i] current_count <- df$X2[i] if (current_count < threshold) { # 直接保留原行 result <- rbind(result, data.frame(X1 = current_item, X2 = current_count)) } else { # 循环拆分直到剩余数量小于阈值 while (current_count >= threshold) { result <- rbind(result, data.frame(X1 = current_item, X2 = chunk_size)) current_count <- current_count - chunk_size } # 添加最后剩余的数量 if (current_count > 0) { result <- rbind(result, data.frame(X1 = current_item, X2 = current_count)) } } } # 查看完整结果 print(result, n = Inf)
说明
- 两种方法都会将
X2≥1e6的行拆分为多行,每行X2值小于1e6,且总数量与原数据一致。 - 若遇到无法被
1e5整除的数值(例如1234567),代码会自动拆分出若干个1e5的行,最后一行保留剩余的数值(34567),确保所有行都满足X2<1e6的要求。
内容的提问来源于stack exchange,提问作者Tracy
相关产品推荐
相关产品推荐

