You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从data.table中抽取分组均匀的无重叠样本?

Split data.table into Balanced Non-Overlapping Subsets

Let's fix this problem properly—your current split with modulo doesn't account for group balance, so we need a stratified approach that ensures each subset has an even distribution of every group level, even when x is a non-divisible number like 7.

Step 1: Start with Your Data

First, let's confirm your original dataset:

library(data.table)
df = data.table(id = 1:100, group=rep(c('a','b','c','d'),25))

Step 2: Extract 80 Stratified Observations

We'll first pull 80 rows while preserving the original group distribution (each group has 25 rows in the full data, so we take 20 per group to reach 80 total):

# Stratified sampling to get 80 rows (20 per group)
df_80 = df[, .SD[sample(.N, 20)], by = group]

Step 3: Assign Balanced Subset IDs

Now we split these 80 rows into x subsets, making sure each group's rows are evenly spread across all subsets. For non-divisible numbers (like x=7), this means some subsets will have one extra row per group, but the overall balance stays intact:

# Define your target number of subsets (e.g., x=7 or x=20)
x <- 7

# For each group, assign subset IDs evenly (uses length.out to handle remainders)
df_80[, subset_id := rep(1:x, length.out = .N), by = group]

Step 4: Split into Subsets

Finally, split the data using the assigned subset IDs:

df_split = split(df_80, df_80$subset_id)

Verify the Balance

Check that each subset has a balanced group distribution:

# Print group counts for each subset
lapply(df_split, function(sub) table(sub$group))

For x=7, you'll see most subsets have 3 rows per group, and one subset has 2 rows per group (since 20 = 7*2 + 6—6 subsets get 3, 1 gets 2). For x=20, each subset will have exactly 1 row per group (perfect balance, since 20 rows per group ÷ 20 subsets = 1 each).

Why Your Original Method Failed

Using as.numeric(as.factor(df$id)) %% 7 assigns subsets based on ID order, which doesn't consider group membership. This leads to random group imbalances in each subset—our stratified approach fixes this by handling each group separately first.

内容的提问来源于stack exchange,提问作者T.Beil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:17:46