You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Pandas DataFrame添加指定占比的分类列?

实现指定占比生成分类列的最优方法

直接利用np.random.choice的p参数指定每个类别的选取概率,这是最简洁高效的方案,无需额外复杂逻辑。

核心实现步骤

  • 定义目标类别和对应占比(注意占比总和必须为1)
  • 调用np.random.choice时传入p参数指定概率
  • 验证生成结果的占比是否符合预期

完整代码示例

import seaborn as sns
import numpy as np

# 加载数据集
df = sns.load_dataset("titanic")

# 定义类别和目标占比
country = ['UK', 'Ireland', 'France']
target_probs = [0.91, 0.06, 0.03]

# 生成指定占比的country列
df["country"] = np.random.choice(country, size=len(df), p=target_probs)

# 验证占比结果
print(df["country"].value_counts(normalize=True))

补充说明

  • 随机抽样会带来微小误差,但样本量越大(如泰坦尼克数据集的891条数据),结果越贴近目标占比。
  • 如果需要严格精确的数量(比如UK必须是811条,即891*0.91≈811),可以手动计算每个类别的数量后拼接生成,示例代码如下:
# 精确控制数量的实现方式
counts = [round(len(df)*p) for p in target_probs]
# 修正四舍五入后总和与总条数不一致的问题
counts[-1] = len(df) - sum(counts[:-1])

# 生成对应数量的类别列表
country_list = []
for c, cnt in zip(country, counts):
    country_list.extend([c]*cnt)

# 打乱顺序后赋值给数据集
np.random.shuffle(country_list)
df["country"] = country_list

# 验证精确占比
print(df["country"].value_counts(normalize=True))

内容的提问来源于stack exchange,提问作者Roy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 10:45:45