如何在DataFrame中查找所有标题组合的共同门店?
解决方法:找出所有标题组合的共同门店
嘿,这个需求很明确,咱们用Pandas结合Python的itertools.combinations就能轻松实现。核心思路是先把每个标题对应的门店集合整理好,再生成所有可能的标题组合,最后计算每个组合下门店的交集(也就是所有标题都覆盖的门店)。
步骤1:准备数据(如果你的df已经存在,这步可以跳过)
先构造你提供的示例DataFrame:
import pandas as pd from itertools import combinations # 构造示例数据 data = { 'Title': ['T1', 'T1', 'T1', 'T1', 'T2', 'T2', 'T2', 'T3', 'T3', 'T4', 'T4'], 'Store': ['S1', 'S2', 'S3', 'S4', 'S1', 'S2', 'S4', 'S1', 'S4', 'S1', 'S2'] } df = pd.DataFrame(data)
步骤2:映射每个标题到对应的门店集合
通过分组把每个标题对应的所有门店存成集合(集合能方便后续计算交集):
# 按Title分组,将每个Title对应的Store转为集合,并存为字典 title_store_map = df.groupby('Title')['Store'].apply(set).to_dict()
执行后,title_store_map会是这样的结构:
{'T1': {'S1', 'S2', 'S3', 'S4'}, 'T2': {'S1', 'S2', 'S4'}, 'T3': {'S1', 'S4'}, 'T4': {'S1', 'S2'}}
步骤3:生成所有可能的标题组合
我们生成所有长度≥2的标题组合(如果需要包含单个标题的情况,把range(2, ...)改成range(1, ...)即可):
# 获取所有唯一的Title列表 all_titles = list(title_store_map.keys()) # 生成所有长度从2到标题总数的组合 title_combinations = [] for combo_length in range(2, len(all_titles) + 1): title_combinations.extend(combinations(all_titles, combo_length))
步骤4:计算每个组合的共同门店并整理结果
遍历每个组合,计算门店集合的交集,再整理成你需要的格式:
result_list = [] for combo in title_combinations: # 获取当前组合中每个标题对应的门店集合 store_sets = [title_store_map[title] for title in combo] # 计算所有集合的交集(共同门店) common_stores = set.intersection(*store_sets) # 把组合和共同门店转为逗号分隔的字符串 combo_str = ','.join(combo) stores_str = ','.join(sorted(common_stores)) if common_stores else '' # 添加到结果列表 result_list.append({ 'Title_combination': combo_str, 'Common_Store': stores_str }) # 转换为最终的DataFrame result_df = pd.DataFrame(result_list)
最终输出
运行后,result_df的结果如下:
| Title_combination | Common_Store |
|---|---|
| T1,T2 | S1,S2,S4 |
| T1,T3 | S1,S4 |
| T1,T4 | S1,S2 |
| T2,T3 | S1,S4 |
| T2,T4 | S1,S2 |
| T3,T4 | S1 |
| T1,T2,T3 | S1,S4 |
| T1,T2,T4 | S1,S2 |
| T1,T3,T4 | S1 |
| T2,T3,T4 | S1 |
| T1,T2,T3,T4 | S1 |
完全符合你预期的输出格式!如果需要调整组合的长度,只需要修改生成组合时的range参数就行。
内容的提问来源于stack exchange,提问作者Anubhav Dikshit
相关产品推荐
相关产品推荐

