如何实现Pair without replacement:最大化处理组与唯一对照组匹配
无放回配对:最大化处理组与对照组匹配数量
需求说明
- 匹配规则:组(group)相同,且对照组分数(score)处于处理组分数的±10%区间内
- 核心约束:每个对照组仅能分配给一个处理组(无放回配对),目标是最大化获得匹配的处理组数量
当前问题
已通过merge关联两组数据并筛选出符合分数条件的配对df_matched,但使用df_matched.drop_duplicates(subset='id_control', keep='first')会优先保留单个处理组的所有配对,导致其他处理组无法匹配,无法实现最大化匹配的目标。
示例数据与初始代码
import pandas as pd df_treatment = pd.DataFrame({ 'id': ['a-1', 'a-2', 'a-3', 'b-1', 'b-2'], 'group': ['A', 'A', 'C', 'C', 'A'], 'score': [200, 230, 210, 390, 202]}) df_control = pd.DataFrame({ 'id': ['x-1', 'y-9', 'y-8', 'y-1', 'x-2', 'x-3', 'x-1'], 'group': ['A', 'A', 'A', 'C', 'C', 'C', 'C'], 'score': [180, 218, 210, 220, 400, 380, 210]}) # 生成同组的所有配对 df_merged = pd.merge(df_treatment, df_control, on='group', suffixes=('_treated', '_control'), how='left') # 筛选分数符合±10%条件的配对 df_matched = df_merged[ df_merged['score_control'].between(df_merged['score_treated'] * 0.9, df_merged['score_treated'] * 1.1) ][['id_treated', 'group', 'id_control']]
初始配对结果df_matched
| id_treated | group | id_control | |
|---|---|---|---|
| 0 | a-1 | A | x-1 |
| 1 | a-1 | A | y-9 |
| 2 | a-1 | A | y-8 |
| 4 | a-2 | A | y-9 |
| 5 | a-2 | A | y-8 |
| 6 | a-3 | C | y-1 |
| 9 | a-3 | C | x-1 |
| 11 | b-1 | C | x-2 |
| 12 | b-1 | C | x-3 |
| 15 | b-2 | A | y-9 |
| 16 | b-2 | A | y-8 |
期望结果(最大化匹配数量)
| id_treated | group | id_control | |
|---|---|---|---|
| 0 | a-1 | A | x-1 |
| 1 | a-2 | A | y-9 |
| 2 | a-3 | C | y-1 |
| 3 | b-1 | C | x-2 |
| 4 | b-1 | C | x-3 |
| 5 | b-2 | A | y-8 |
现有方法的不理想结果
使用df_matched.drop_duplicates(subset='id_control', keep='first')得到的结果,会让a-1占用所有A组可选对照组,导致a-2、b-2无法匹配:
| id_treated | group | id_control | |
|---|---|---|---|
| 0 | a-1 | A | x-1 |
| 1 | a-1 | A | y-9 |
| 2 | a-1 | A | y-8 |
| 6 | a-3 | C | y-1 |
| 11 | b-1 | C | x-2 |
| 12 | b-1 | C | x-3 |
解决方案:优先分配可选配对少的处理组
通过贪心算法,优先给可选对照组数量最少的处理组分配资源,避免单个处理组占用过多选项,从而最大化整体匹配数量:
# 统计每个处理组的可选配对数量 treatment_option_counts = df_matched.groupby('id_treated').size().reset_index(name='option_count') # 按可选数量升序排序配对数据,让可选少的处理组优先选择 df_sorted = df_matched.merge(treatment_option_counts, on='id_treated').sort_values(['option_count', 'id_treated']) # 遍历分配,标记已使用的对照组 used_controls = set() matched_pairs = [] for _, row in df_sorted.iterrows(): control_id = row['id_control'] if control_id not in used_controls: matched_pairs.append(row[['id_treated', 'group', 'id_control']]) used_controls.add(control_id) # 转换为最终结果DataFrame final_matched_df = pd.DataFrame(matched_pairs).reset_index(drop=True) print(final_matched_df)
方法说明
- 统计可选数量:先计算每个处理组能匹配的对照组总数,判断哪个处理组的可选资源最少
- 排序优先级:按可选数量升序排序配对数据,确保可选少的处理组先分配,避免资源被抢占
- 无放回分配:遍历排序后的配对,给每个处理组分配未被使用的对照组,直到所有可用资源耗尽或处理组都完成匹配
运行上述代码后,得到的结果与期望一致,实现了最大化处理组的匹配数量。
内容的提问来源于stack exchange,提问作者alig265
相关产品推荐
相关产品推荐

