You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

优化DataFrame行间随机整数分配:dplyr加速方案问询

加速超大型DataFrame的R代码优化方案

场景与可复现代码

先定义核心函数及模拟超大型DataFrame的示例数据:

# 生成指定总和、长度的随机正整数向量的核心函数
generateIntegers <- function(sum_val, size) {
  if (size == 0 || sum_val < size) stop("sum_val必须大于等于size,且size不能为0")
  res <- diff(c(0, sort(sample(1:(sum_val - 1), size - 1)), sum_val))
  res
}

# 模拟超大型DataFrame结构
set.seed(123)
n_rows <- 10000  # 模拟大行数
D_T <- data.frame(
  N_h = sample(5:20, n_rows, replace = TRUE),
  s_aL = sample(10:100, n_rows, replace = TRUE)
)
D_H <- data.frame(
  id = 1:n_rows,
  s_aL = vector("list", n_rows)  # 用列表列存储向量结果
)

原慢速度的for循环实现

# 原低效循环代码
system.time({
  for (i in 1:nrow(D_T)) {
    D_H$s_aL[[i]] <- generateIntegers(D_T$s_aL[i], D_T$N_h[i])
  }
})

优化后的dplyr + purrr实现

用向量化映射替代R层面循环,大幅提升速度:

library(dplyr)
library(purrr)

system.time({
  D_H <- D_T %>%
    mutate(s_aL = map2(s_aL, N_h, ~generateIntegers(.x, .y))) %>%
    select(s_aL)  # 如需保留原D_H的其他列,可调整select逻辑
})

优化要点

  • map2是C级别的循环实现,逐行配对s_aL和N_h的元素调用函数,远快于R原生for循环
  • 必须用列表列存储结果:因为每个返回值是向量,普通向量列无法容纳嵌套结构
  • 可进一步优化generateIntegers:比如移除不必要的参数检查、改用更高效的采样逻辑,能进一步提升整体速度

内容的提问来源于stack exchange,提问作者Sophie Père

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 00:01:00