You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

移除for循环:基于权重列高效抽样生成新data.frame

Efficient Vectorized Sampling Without For Loops in R

Nice question—when you're dealing with 100k samples across multiple columns, ditching explicit for loops is absolutely the right call for speed and efficiency. Here are a few optimized approaches tailored to your needs:

First, let's recap your original data setup for context:

set.seed(100)
x <- data.frame("smp_1"=runif(20)*100,"smp_2"=runif(20)*99)
x["weight_1"] = x$smp_1/sum(x$smp_1)
x["weight_2"] = x$smp_2/sum(x$smp_2)

1. Use purrr for Clean Vectorized Iteration

The purrr package provides optimized map functions that avoid the overhead of manual for loops. map_dfc will automatically bind your sampled results into a single data frame column-wise:

library(purrr)

# Set your desired sample size (100k in your case)
n_samples <- 100000

# Iterate over column pairs and generate samples
sim_df <- map_dfc(1:2, function(k) {
  sample(x[[paste0("smp_", k)]], size = n_samples, replace = TRUE, prob = x[[paste0("weight_", k)]])
}) %>%
  setNames(paste0("sim_", 1:2))

For even more flexibility (if you add more columns later), you can auto-detect your smp_ and weight_ columns instead of hardcoding 1:2:

# Auto-identify smp and weight columns
smp_cols <- grep("^smp_", colnames(x), value = TRUE)
weight_cols <- gsub("smp_", "weight_", smp_cols)

# Use map2 to pair smp and weight columns directly
sim_df <- map2_dfc(smp_cols, weight_cols, function(smp_col, weight_col) {
  sample(x[[smp_col]], size = n_samples, replace = TRUE, prob = x[[weight_col]])
}) %>%
  setNames(gsub("smp_", "sim_", smp_cols))

2. Use data.table for Maximum Performance with Large Data

If you're working with very large datasets (both in rows of x and sample size), data.table offers superior memory efficiency and speed compared to base R data frames:

library(data.table)

# Convert your data frame to data.table
x_dt <- as.data.table(x)

n_samples <- 100000

# Generate samples with lapply, then convert to data.table
sim_dt <- as.data.table(lapply(1:2, function(k) {
  sample(x_dt[[paste0("smp_", k)]], size = n_samples, replace = TRUE, prob = x_dt[[paste0("weight_", k)]])
}))

# Rename columns to match your desired output
setnames(sim_dt, paste0("sim_", 1:2))

3. Optimize Further with sample.int

For a tiny extra speed boost, use sample.int to sample row indices first, then extract values from your smp_ columns. This avoids some overhead of sampling directly from vectors:

n_samples <- 100000

sim_df <- map_dfc(1:2, function(k) {
  # Sample row indices using the weight column
  idx <- sample.int(nrow(x), size = n_samples, replace = TRUE, prob = x[[paste0("weight_", k)]])
  # Extract values from the corresponding smp column
  x[[paste0("smp_", k)]][idx]
}) %>%
  setNames(paste0("sim_", 1:2))

Why These Methods Are Better Than For Loops

Your original for loop has a major inefficiency: every iteration modifies tempdf by adding a column, which requires copying the entire data frame each time. The methods above generate all columns first and then bind them once, drastically reducing memory overhead and speeding up execution—critical for 100k samples.

内容的提问来源于stack exchange,提问作者user33484

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:28:43