You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Base R从每对相差1的ID组中随机抽取1个样本?

用Base R实现配对组随机抽取样本

原始数据

首先定义你的DataFrame:

df <- data.frame (
  ID  = c("2330", "2331", "2333", "2334", "2336", "2337", "4430", "4431", "4510", "4511"),
  length = c(8.4,6,3,9,3,4,1,7,4,2)
)

需求说明

将数值差为1的ID划分为一组(如2330和2331、2333和2334等),从每个组中随机抽取1条记录,生成新的DataFrame。

实现代码

# 1. 将ID转为数值型,方便分组计算
df$num_id <- as.numeric(df$ID)

# 2. 对数据按num_id排序,生成组标识
sorted_df <- df[order(df$num_id), ]
sorted_df$group <- c(1, cumsum(diff(sorted_df$num_id) != 1) + 1)

# 3. 将组标识合并回原始数据
df <- merge(df, sorted_df[, c("num_id", "group")], by = "num_id", all.x = TRUE)

# 4. 按组随机抽取1条记录,合并结果
result <- do.call(rbind, by(df, df$group, function(x) x[sample(nrow(x), 1), ]))

# 5. 整理输出格式,保留需要的列并重置行名
result <- result[, c("ID", "length")]
rownames(result) <- NULL

运行后即可得到类似示例的随机抽取结果,例如:

> result
    ID length
1 2331    6.0
2 2333    3.0
3 2337    4.0
4 4431    7.0
5 4511    2.0

代码解释

  • 步骤2通过计算排序后ID的差值,将连续差1的ID归为同一组,确保分组逻辑准确,不受原始数据顺序影响。
  • 步骤4使用by函数按组处理,sample(nrow(x),1)实现每组随机抽取1行,最后用do.call(rbind,...)合并各组结果。

内容的提问来源于stack exchange,提问作者wooden05

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 08:30:50