Python多条件按指定比例抽样的实现方案问询
抽样实现方案
核心思路
优先抽选同时满足多个条件的样本——这类样本能同时计入多个条件的配额,减少重复抽取的冗余;之后再补充单条件样本,最后抽取无任何条件的样本,确保总样本数100,各条件配额均达25。
具体实现步骤(Python + Pandas)
1. 给样本标记条件组合类型
先给每条数据打上标签,明确它满足哪些条件组合,方便后续按组抽样:
def get_condition_group(row): conditions = [] if row['Condition 1'] == 'Yes': conditions.append('1') if row['Condition 2'] == 'Yes': conditions.append('2') if row['Condition 3'] == 'Yes': conditions.append('3') return '+'.join(conditions) if conditions else 'none' # 新增列存储条件组合 df['condition_group'] = df.apply(get_condition_group, axis=1)
2. 给条件组合排序,优先选多重叠样本
把包含多个条件的组合排在前面,比如1+2+3 > 1+2 > 1+3 > 2+3 > 1 > 2 > 3 > none,确保抽样时先拿这类高价值样本:
# 定义组合优先级:条件越多优先级越高 group_priority = ['1+2+3', '1+2', '1+3', '2+3', '1', '2', '3', 'none'] df['priority'] = df['condition_group'].map(lambda x: group_priority.index(x)) # 按优先级升序排序,优先级数值越小越靠前 df_sorted = df.sort_values('priority')
3. 按配额逐步抽取样本
初始化剩余配额,遍历排序后的样本,逐个加入样本集并更新剩余配额,直到所有配额满足:
# 初始化各类型剩余配额 remaining = { '1': 25, '2': 25, '3': 25, 'none': 25 } sample_list = [] for _, row in df_sorted.iterrows(): groups = row['condition_group'].split('+') if row['condition_group'] != 'none' else ['none'] can_take = False # 判断当前样本是否能填补剩余配额 if groups == ['none']: can_take = remaining['none'] > 0 else: # 只要能填补至少一个剩余配额,就优先选取 for g in groups: if remaining[g] > 0: can_take = True break if can_take: sample_list.append(row) # 更新对应配额 if groups == ['none']: remaining['none'] -= 1 else: for g in groups: if remaining[g] > 0: remaining[g] -= 1 # 所有配额都满足时提前终止循环 if all(v <= 0 for v in remaining.values()): break # 转换为最终样本DataFrame final_sample = pd.DataFrame(sample_list) # 断言确保样本量达标(原数据集样本足够时生效) assert len(final_sample) == 100, "原数据集样本量不足,无法满足配额要求"
4. 验证抽样结果
可以加一段代码验证各条件计数是否符合要求:
# 统计各条件满足数量 cond1_count = final_sample[final_sample['Condition 1'] == 'Yes'].shape[0] cond2_count = final_sample[final_sample['Condition 2'] == 'Yes'].shape[0] cond3_count = final_sample[final_sample['Condition 3'] == 'Yes'].shape[0] none_count = final_sample[(final_sample['Condition 1'] == 'No') & (final_sample['Condition 2'] == 'No') & (final_sample['Condition 3'] == 'No')].shape[0] print(f"Condition 1 满足数: {cond1_count}") print(f"Condition 2 满足数: {cond2_count}") print(f"Condition 3 满足数: {cond3_count}") print(f"无任何条件数: {none_count}")
注意事项
- 如果原数据中多条件组合的样本量不足,代码会自动用单条件样本补充,确保配额完成。
- 如果原数据集整体样本量不够(比如无任何条件的样本不足25),会触发断言报错,需要提前检查数据分布。
内容的提问来源于stack exchange,提问作者user20743916
相关产品推荐
相关产品推荐

