You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多条件按指定比例抽样的实现方案问询

抽样实现方案

核心思路

优先抽选同时满足多个条件的样本——这类样本能同时计入多个条件的配额,减少重复抽取的冗余;之后再补充单条件样本,最后抽取无任何条件的样本,确保总样本数100,各条件配额均达25。

具体实现步骤(Python + Pandas)

1. 给样本标记条件组合类型

先给每条数据打上标签,明确它满足哪些条件组合,方便后续按组抽样:

def get_condition_group(row):
    conditions = []
    if row['Condition 1'] == 'Yes':
        conditions.append('1')
    if row['Condition 2'] == 'Yes':
        conditions.append('2')
    if row['Condition 3'] == 'Yes':
        conditions.append('3')
    return '+'.join(conditions) if conditions else 'none'

# 新增列存储条件组合
df['condition_group'] = df.apply(get_condition_group, axis=1)

2. 给条件组合排序,优先选多重叠样本

把包含多个条件的组合排在前面,比如1+2+3 > 1+2 > 1+3 > 2+3 > 1 > 2 > 3 > none,确保抽样时先拿这类高价值样本:

# 定义组合优先级:条件越多优先级越高
group_priority = ['1+2+3', '1+2', '1+3', '2+3', '1', '2', '3', 'none']
df['priority'] = df['condition_group'].map(lambda x: group_priority.index(x))
# 按优先级升序排序,优先级数值越小越靠前
df_sorted = df.sort_values('priority')

3. 按配额逐步抽取样本

初始化剩余配额,遍历排序后的样本,逐个加入样本集并更新剩余配额,直到所有配额满足:

# 初始化各类型剩余配额
remaining = {
    '1': 25,
    '2': 25,
    '3': 25,
    'none': 25
}
sample_list = []

for _, row in df_sorted.iterrows():
    groups = row['condition_group'].split('+') if row['condition_group'] != 'none' else ['none']
    can_take = False

    # 判断当前样本是否能填补剩余配额
    if groups == ['none']:
        can_take = remaining['none'] > 0
    else:
        # 只要能填补至少一个剩余配额,就优先选取
        for g in groups:
            if remaining[g] > 0:
                can_take = True
                break

    if can_take:
        sample_list.append(row)
        # 更新对应配额
        if groups == ['none']:
            remaining['none'] -= 1
        else:
            for g in groups:
                if remaining[g] > 0:
                    remaining[g] -= 1
        
        # 所有配额都满足时提前终止循环
        if all(v <= 0 for v in remaining.values()):
            break

# 转换为最终样本DataFrame
final_sample = pd.DataFrame(sample_list)
# 断言确保样本量达标(原数据集样本足够时生效)
assert len(final_sample) == 100, "原数据集样本量不足,无法满足配额要求"

4. 验证抽样结果

可以加一段代码验证各条件计数是否符合要求:

# 统计各条件满足数量
cond1_count = final_sample[final_sample['Condition 1'] == 'Yes'].shape[0]
cond2_count = final_sample[final_sample['Condition 2'] == 'Yes'].shape[0]
cond3_count = final_sample[final_sample['Condition 3'] == 'Yes'].shape[0]
none_count = final_sample[(final_sample['Condition 1'] == 'No') & 
                          (final_sample['Condition 2'] == 'No') & 
                          (final_sample['Condition 3'] == 'No')].shape[0]

print(f"Condition 1 满足数: {cond1_count}")
print(f"Condition 2 满足数: {cond2_count}")
print(f"Condition 3 满足数: {cond3_count}")
print(f"无任何条件数: {none_count}")

注意事项

  • 如果原数据中多条件组合的样本量不足,代码会自动用单条件样本补充,确保配额完成。
  • 如果原数据集整体样本量不够(比如无任何条件的样本不足25),会触发断言报错,需要提前检查数据分布。

内容的提问来源于stack exchange,提问作者user20743916

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 16:58:10