如何基于Pandas分组按指定概率抽样并生成结果DataFrame?
高效实现基于分组离散概率分布的抽样(Pandas版)
针对你提出的需求——基于按NAME分组的离散概率分布生成抽样结果,我这里提供一个更简洁高效的Pandas实现方案,完全利用分组操作和向量化逻辑替代手动循环,性能和可读性都更优。
需求背景回顾
我们先明确原始数据结构:每个NAME对应一组离散概率分布,POSITION是取值,PROB是对应取值的概率:
import pandas as pd import numpy as np df = pd.DataFrame( [ ['X', 0, 0.5], ['X', 1, 0.5], ['Y', 0, 0.25], ['Y', 1, 0.3], ['Y', 2, 0.45], ['Z', 0, 0.6], ['Z', 1, 0.1], ['Z', 2, 0.3] ], columns=['NAME', 'POSITION', 'PROB'])
需要生成包含NAME、POSITION、PROB(0/1标记是否被抽样选中)、SAMPLE(抽样批次编号)的结果DataFrame。
高效实现代码
n_samples = 3 # 指定需要的抽样次数 # 第一步:为每个NAME组按概率抽取n_samples次POSITION,记录每个抽样批次的结果 sampled_results = ( df.groupby('NAME') .apply(lambda group: pd.Series( # 从当前组的POSITION中按PROB权重抽取n_samples次 np.random.choice(group['POSITION'], size=n_samples, p=group['PROB']), name='sampled_pos' ).reset_index().rename(columns={'index': 'SAMPLE'})) # 将索引转为抽样批次编号 .reset_index() .drop(columns='level_1') # 移除groupby生成的冗余索引列 ) # 第二步:将原始数据与抽样结果做笛卡尔积,标记每个POSITION是否被对应抽样选中 df_samples = ( df.merge(sampled_results, on='NAME', how='cross') .assign(PROB=lambda x: (x['POSITION'] == x['sampled_pos']).astype(int)) # 匹配则标记为1,否则0 .drop(columns='sampled_pos') # 移除中间辅助列 .sort_values(['NAME', 'SAMPLE']) # 按NAME和抽样批次排序,和示例结构对齐 .reset_index(drop=True) )
方案优势
- 无手动循环:利用
groupby.apply自动处理所有NAME分组,不需要手动遍历每个组,代码更简洁 - 向量化操作:抽样和匹配逻辑都是向量化实现,比Python循环快得多,尤其当分组数量多、抽样次数大时性能提升明显
- 自适应分组差异:自动处理不同
NAME组的POSITION数量差异,无需手动计算每组长度
单元验证(大数定律测试)
我们可以用大样本量验证抽样结果是否符合真实概率:
n_samples_large = 10000 # 大样本量保证统计显著性 # 生成大样本抽样结果 sampled_large = ( df.groupby('NAME') .apply(lambda g: pd.Series( np.random.choice(g['POSITION'], size=n_samples_large, p=g['PROB']), name='sampled_pos' ).reset_index().rename(columns={'index': 'SAMPLE'})) .reset_index() .drop(columns='level_1') ) df_samples_large = ( df.merge(sampled_large, on='NAME', how='cross') .assign(PROB=lambda x: (x['POSITION'] == x['sampled_pos']).astype(int)) .drop(columns='sampled_pos') ) # 计算估计概率和真实概率 p_est = df_samples_large.groupby(['NAME', 'POSITION'])['PROB'].mean() p_true = df.set_index(['NAME', 'POSITION'])['PROB'] # 计算4倍标准差的置信区间 z = 4 CI_lower = p_est - z * np.sqrt(p_est * (1 - p_est) / n_samples_large) CI_upper = p_est + z * np.sqrt(p_est * (1 - p_est) / n_samples_large) # 验证真实概率是否在置信区间内 assert (p_true < CI_upper).all() assert (p_true > CI_lower).all() print("所有分组的真实概率均在置信区间内,抽样结果有效!")
内容的提问来源于stack exchange,提问作者rwolst
相关产品推荐
相关产品推荐

