如何按嵌套列表子项匹配pandas DataFrame,符合条件时输出对应索引与月份
Pandas嵌套列表匹配输出扩展实现
基础数据
嵌套列表
list_of_list = [["aa", "yy"], ["gg", "xx"]]
测试用DataFrame
month column1 column2 0 June xx aa 1 June gg xx 2 August xx yy
原有逻辑说明
原代码遍历嵌套列表的每个子项,检查子项是否存在于DataFrame任意列中,匹配时输出对应列的内容,现有代码如下:
for sub_list in list_of_lists: print("\n") print("Sub_list: ", sub_list) for item in sub_list: for col in df.columns: if item in df[col].values: print(df[col][df[col] == item].to_frame())
需求
在原有输出基础上,当且仅当子列表的各个匹配项不在DataFrame的同一索引行时,额外输出各匹配项对应的索引和月份值。示例:当输入list_of_list = [["aa", "yy"]]时,需额外输出:
0 June aa 2 August yy
实现方案
我们可以在原有遍历逻辑中增加匹配索引的收集与交集判断逻辑,判断是否存在同一行包含子列表的所有匹配项,无公共索引时触发额外输出:
import pandas as pd # 数据初始化 list_of_list = [["aa", "yy"], ["gg", "xx"]] df = pd.DataFrame({ 'month': ['June', 'June', 'August'], 'column1': ['xx', 'gg', 'xx'], 'column2': ['aa', 'xx', 'yy'] }) for sub_list in list_of_list: print("\n") print("Sub_list: ", sub_list) match_records = [] item_index_sets = [] for item in sub_list: current_item_indexes = set() for col in df.columns: if item in df[col].values: # 原有输出逻辑保留 matched_part = df[col][df[col] == item] print(matched_part.to_frame()) # 收集匹配到的索引、月份、匹配项信息 for idx, val in matched_part.items(): match_records.append((idx, df.loc[idx, 'month'], item)) current_item_indexes.add(idx) if current_item_indexes: item_index_sets.append(current_item_indexes) # 判断是否所有子项匹配结果没有公共索引 if len(item_index_sets) == len(sub_list): common_idx = set.intersection(*item_index_sets) if not common_idx: print("\n* 匹配项未出现在同一行,额外输出:") # 去重排序后输出 for idx, month, item in sorted(list(set(match_records)), key=lambda x:x[0]): print(f"{idx} {month} {item}")
效果说明
- 对于子列表
['aa', 'yy']:aa匹配索引为{0},yy匹配索引为{2},无公共索引,触发额外输出 - 对于子列表
['gg', 'xx']:gg匹配索引为{1},xx匹配索引包含1,存在公共索引,不触发额外输出
完全符合需求。
内容的提问来源于stack exchange,提问作者xavi
相关产品推荐
相关产品推荐

