如何统计两个DataFrame匹配字符串数量?解决contains返回错误问题
Pandas str.contains匹配失败的解决方法
问题场景
从df1提取指定区域的Titre列,用逗号拼接成字符串列表Shortlist:
temp = df1.loc[df1['quartier']=='quartier1', 'Titre'] Shortlist = ','.join(temp)
然后在df2同区域的Titre_test列(包含逗号分隔的字符串或NaN)中检查匹配:
temp_ilot = df2.loc[df2['quartier']=='quartier1', 'Titre_test'] temp_ilot.str.contains(Shortlist)
明明存在匹配项,但结果返回False。
原因分析
str.contains(Shortlist)会将整个逗号分隔的字符串作为完整正则表达式匹配,而非检查是否包含Shortlist中的任意单个元素。比如Shortlist = "A,B,C"时,它会查找是否有字符串包含"A,B,C"这个整体,而非A、B或C中的任意一个。
解决方案
步骤1:构建正则OR匹配模式
将Titre的唯一值用|分隔,生成能匹配任意单个元素的正则表达式:
# 提取唯一的Titre值,避免重复匹配 unique_titres = temp.unique() # 用|拼接成正则OR模式,同时转义特殊字符防止正则语法干扰 import re pattern = '|'.join(re.escape(titre) for titre in unique_titres)
步骤2:执行匹配并统计数量
# 执行匹配,na=False将空值转为False,方便统计 matches = temp_ilot.str.contains(pattern, na=False) # 统计匹配的行数 total_matches = matches.sum() # 如果需要查看具体匹配的内容,可以过滤出True的行 matched_rows = temp_ilot[matches]
替代方案:拆分字符串后匹配
如果Titre_test中的字符串是逗号分隔的,也可以拆分后逐个匹配:
# 拆分每个Titre_test的字符串为列表,排除空值 split_titres = temp_ilot.dropna().str.split(',') # 检查每个列表是否与unique_titres有交集 matches = split_titres.apply(lambda x: any(t in unique_titres for t in x)) # 统计数量 total_matches = matches.sum()
内容的提问来源于stack exchange,提问作者bravopapa
相关产品推荐
相关产品推荐

