You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于DataFrame的重估计Bootstrap:生成10个带新ID的重复样本

Julia DataFrame Bootstrap重采样实现

需求说明

对给定的DataFrame执行Bootstrap重采样(基于原始个体的有放回抽样),生成10个重复样本,并为最终数据集分配新ID。

原始数据集

Random.seed!(123)
df = DataFrame(
    id = ["1","1","1","2","2","2","3","3","3"],
    time = [0,0.5,2,0,0.5,2,0,0.5,2],
    cmt = [1,0,0,1,0,0,1,0,0],
    value = [0.01,0.02,0.03,0.02,0.03,0.05,0.01,0.05,0.10]
)

实现步骤与代码

核心思路

  1. 识别原始数据中的独立个体(即id列的唯一值)
  2. 对个体进行有放回抽样,每次抽样数量与原始个体数一致,生成单个Bootstrap样本
  3. 重复上述操作10次,合并所有样本
  4. 为每个Bootstrap样本中的个体分配新ID,格式为[样本编号]_[个体序号]

完整代码

using DataFrames, Random

Random.seed!(123)
df = DataFrame(
    id = ["1","1","1","2","2","2","3","3","3"],
    time = [0,0.5,2,0,0.5,2,0,0.5,2],
    cmt = [1,0,0,1,0,0,1,0,0],
    value = [0.01,0.02,0.03,0.02,0.03,0.05,0.01,0.05,0.10]
)

# 设置Bootstrap样本数量
n_bootstrap = 10

# 获取原始独立个体ID
original_subjects = unique(df.id)
n_subjects = length(original_subjects)

# 初始化存储所有Bootstrap样本的DataFrame
bootstrap_df = DataFrame()

for sample_idx in 1:n_bootstrap
    # 有放回抽取个体ID
    sampled_subjects = sample(original_subjects, n_subjects; replace=true)
    # 拼接抽样个体的所有行数据
    current_sample = vcat([df[df.id .== subj, :] for subj in sampled_subjects]...)
    # 生成新ID:绑定样本编号与个体在当前样本中的序号
    new_id_mapping = Dict(subj => "$sample_idx"*"_$(idx)" for (idx, subj) in enumerate(sampled_subjects))
    current_sample.new_id = [new_id_mapping[id] for id in current_sample.id]
    # 将当前样本追加到总数据集
    append!(bootstrap_df, current_sample)
end

# 查看前10行验证结果
first(bootstrap_df, 10)

代码说明

  • sample(original_subjects, n_subjects; replace=true):实现有放回抽样,确保每次生成的Bootstrap样本包含与原始数据相同数量的个体
  • new_id列:新ID由样本编号和个体在当前样本中的序号组成,例如1_2表示第1个Bootstrap样本中的第2个个体
  • append!:将每个Bootstrap样本合并到同一个DataFrame中,方便后续分析

内容的提问来源于stack exchange,提问作者Parsshava Mehta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 06:13:24