Python实现同列行迭代:查找最匹配字符串并生成结果表
解决方案
我们可以通过将每行字符串转换为集合,利用集合交集操作快速计算匹配项数量,再遍历完成统计,代码实现如下:
import pandas as pd # 初始化输入DataFrame data = [[1, 'AA,AC,BC,DE'], [2, 'AA,AD,BC,D'], [3, 'A,C,BC,E'],[4, 'AA,AC,BC,DEEE'],[5, 'KK']] df = pd.DataFrame(data, columns=['rowid', 'Col1']) # 将Col1的逗号分隔字符串转为集合,方便后续计算交集 df['col_set'] = df['Col1'].str.split(',').apply(set) # 初始化结果列 df['max_matching#'] = 0 df['Rowid'] = 'NA' # 遍历每一行,计算与其他行的匹配数 for idx, row in df.iterrows(): current_set = row['col_set'] match_counts = {} # 跳过自身,遍历其他行 for other_idx, other_row in df.iterrows(): if idx == other_idx: continue intersection = current_set & other_row['col_set'] match_counts[other_row['rowid']] = len(intersection) if not match_counts: continue # 提取最大匹配数及对应的rowid max_count = max(match_counts.values()) max_rowids = [str(rid) for rid, cnt in match_counts.items() if cnt == max_count] # 更新结果 df.loc[idx, 'max_matching#'] = max_count df.loc[idx, 'Rowid'] = ','.join(max_rowids) # 整理为目标格式 df_out = df.drop('col_set', axis=1)[['rowid', 'Col1', 'max_matching#', 'Rowid']] print(df_out)
代码说明
- 转换集合:把
Col1的字符串转为集合,集合的交集操作能高效算出两行的匹配项数量。 - 遍历统计:对每行遍历其他所有行,记录每个rowid对应的匹配项数量。
- 提取结果:找出当前行的最大匹配数,收集所有达到该数值的rowid并格式化为逗号分隔字符串。
- 整理输出:移除辅助列,调整列顺序得到目标DataFrame。
内容的提问来源于stack exchange,提问作者san1
相关产品推荐
相关产品推荐

