You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python按指定分布随机抽样生成含重复ID的DataFrame

解决方案

以下是基于 pandas 和 numpy 实现需求的完整代码,重点解决 age 列的分布抽样问题:

import pandas as pd
import numpy as np

# 生成id列:1-100重复3次,总长度300
id_col = np.tile(np.arange(1, 101), 3)

# 定义age的年龄段区间及对应占比
age_bins = [0, 14, 64, 100]
age_weights = [0.17, 0.65, 0.18]
total_samples = len(id_col)

# 第一步:按权重随机选择每个样本所属的年龄段分组
group_indices = np.random.choice([0, 1, 2], size=total_samples, p=age_weights)

# 第二步:根据分组生成对应区间内的随机年龄(这里生成整数年龄,若需要浮点数可替换为np.random.uniform)
age_col = np.where(
    group_indices == 0,
    np.random.randint(age_bins[0], age_bins[1]+1, size=total_samples),
    np.where(
        group_indices == 1,
        np.random.randint(age_bins[1]+1, age_bins[2]+1, size=total_samples),
        np.random.randint(age_bins[2]+1, age_bins[3]+1, size=total_samples)
    )
)

# 组合成DataFrame
df = pd.DataFrame({'id': id_col, 'age': age_col})

# 可选:验证年龄分布是否符合预期
print(df['age'].apply(lambda x: 0 if x <=14 else 1 if x <=64 else 2).value_counts(normalize=True))

关键步骤说明:

  • 用np.tile快速生成重复的id序列,避免低效的循环拼接
  • 通过np.random.choice按指定权重分配年龄段分组,保证整体占比符合设定要求
  • 用np.where嵌套实现分区间随机年龄生成,逻辑清晰且运行效率较高
  • 最后可选的验证代码可以快速检查生成的年龄分布是否接近设定的17%/65%/18%

内容的提问来源于stack exchange,提问作者Eisen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 01:54:22