You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R函数将数据框中满足条件的行拆分为多行?

在R语言中拆分大数量行至每行数量小于1e6

示例原始数据

首先构造你提供的原始数据集:

df <- data.frame(
  X1 = c("a", "b", "c", "d"),
  X2 = c(1000000, 2000000, 2, 40)
)

方法一:使用tidyverse工具(简洁高效)

借助tidyverse的函数可以快速实现需求,代码逻辑清晰且运行效率更高:

library(tidyverse)

# 定义拆分逻辑的函数
split_large_values <- function(item, count, threshold = 1e6, chunk = 1e5) {
  if (count < threshold) {
    return(tibble(X1 = item, X2 = count))
  }
  # 计算完整拆分块的数量和剩余值
  full_chunks <- count %/% chunk
  remainder <- count %% chunk
  # 生成所有拆分后的行
  chunks <- rep(chunk, full_chunks)
  if (remainder > 0) chunks <- c(chunks, remainder)
  tibble(X1 = rep(item, length(chunks)), X2 = chunks)
}

# 应用函数并合并结果
result <- df %>%
  rowwise() %>%
  group_split() %>%
  map_dfr(~split_large_values(.$X1, .$X2))

# 查看完整结果
print(result, n = Inf)

方法二:使用for循环(符合你的初始思路)

如果你倾向于用循环实现拆分逻辑,以下代码完全匹配你的需求:

# 初始化结果数据框
result <- data.frame(X1 = character(), X2 = integer(), stringsAsFactors = FALSE)

# 设定阈值和拆分块大小
threshold <- 1e6
chunk_size <- 1e5

# 遍历每一行处理
for (i in 1:nrow(df)) {
  current_item <- df$X1[i]
  current_count <- df$X2[i]
  
  if (current_count < threshold) {
    # 直接保留原行
    result <- rbind(result, data.frame(X1 = current_item, X2 = current_count))
  } else {
    # 循环拆分直到剩余数量小于阈值
    while (current_count >= threshold) {
      result <- rbind(result, data.frame(X1 = current_item, X2 = chunk_size))
      current_count <- current_count - chunk_size
    }
    # 添加最后剩余的数量
    if (current_count > 0) {
      result <- rbind(result, data.frame(X1 = current_item, X2 = current_count))
    }
  }
}

# 查看完整结果
print(result, n = Inf)

说明

  • 两种方法都会将X2≥1e6的行拆分为多行,每行X2值小于1e6,且总数量与原数据一致。
  • 若遇到无法被1e5整除的数值(例如1234567),代码会自动拆分出若干个1e5的行,最后一行保留剩余的数值(34567),确保所有行都满足X2<1e6的要求。

内容的提问来源于stack exchange,提问作者Tracy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 19:29:59