如何筛选pandas DataFrame中两列最优组合以覆盖所有唯一值
问题说明
你需要实现的是最小集合覆盖问题的贪心求解,目标是用最少的(col_1, col_2)组合,覆盖col_1和col_2的所有唯一值,每次优先选择能覆盖最多未覆盖值的组合,最终保留对应行的col_3值。
实现代码
import pandas as pd # 构造测试数据集 df = pd.DataFrame({'col_1':['1','1','1','2','2','3','3','4','4','4','5'], 'col_2':['A','B','C','A','B','C','D','D','A','A','B'], 'col_3':['256','546','985','573','265','731','968','592','364','657','953']}) # 初始化需要覆盖的唯一值集合 need_cover_c1 = set(df['col_1'].unique()) need_cover_c2 = set(df['col_2'].unique()) covered_c1 = set() covered_c2 = set() res_rows = [] # 贪心迭代选择最优组合 while covered_c1 != need_cover_c1 or covered_c2 != need_cover_c2: # 计算每行可覆盖的新值数量作为得分 df['score'] = df.apply( lambda x: (x['col_1'] not in covered_c1) + (x['col_2'] not in covered_c2), axis=1 ) # 取得分最高的第一行 best = df[df['score'] == df['score'].max()].iloc[0] res_rows.append(best.to_dict()) # 更新已覆盖集合 covered_c1.add(best['col_1']) covered_c2.add(best['col_2']) # 转换为结果DataFrame res_df = pd.DataFrame(res_rows)[['col_1', 'col_2', 'col_3']] print(res_df)
输出结果
col_1 col_2 col_3 0 1 A 256 1 2 B 265 2 3 C 731 3 4 D 592 4 5 B 953
完全匹配预期输出。
内容的提问来源于stack exchange,提问作者prem
相关产品推荐
相关产品推荐

