如何为Pandas DataFrame添加指定占比的分类列?
实现指定占比生成分类列的最优方法
直接利用np.random.choice的p参数指定每个类别的选取概率,这是最简洁高效的方案,无需额外复杂逻辑。
核心实现步骤
- 定义目标类别和对应占比(注意占比总和必须为1)
- 调用
np.random.choice时传入p参数指定概率 - 验证生成结果的占比是否符合预期
完整代码示例
import seaborn as sns import numpy as np # 加载数据集 df = sns.load_dataset("titanic") # 定义类别和目标占比 country = ['UK', 'Ireland', 'France'] target_probs = [0.91, 0.06, 0.03] # 生成指定占比的country列 df["country"] = np.random.choice(country, size=len(df), p=target_probs) # 验证占比结果 print(df["country"].value_counts(normalize=True))
补充说明
- 随机抽样会带来微小误差,但样本量越大(如泰坦尼克数据集的891条数据),结果越贴近目标占比。
- 如果需要严格精确的数量(比如UK必须是811条,即891*0.91≈811),可以手动计算每个类别的数量后拼接生成,示例代码如下:
# 精确控制数量的实现方式 counts = [round(len(df)*p) for p in target_probs] # 修正四舍五入后总和与总条数不一致的问题 counts[-1] = len(df) - sum(counts[:-1]) # 生成对应数量的类别列表 country_list = [] for c, cnt in zip(country, counts): country_list.extend([c]*cnt) # 打乱顺序后赋值给数据集 np.random.shuffle(country_list) df["country"] = country_list # 验证精确占比 print(df["country"].value_counts(normalize=True))
内容的提问来源于stack exchange,提问作者Roy
相关产品推荐
相关产品推荐

