如何统计DataFrame两列匹配单词数量?解决bad escape \o错误
问题解决与代码修正
报错原因
报错error: bad escape \o at position 1是因为你在split(',\o*')里用了无效的转义字符\o,这里应该用\s*来匹配逗号后0个或多个空格,实现按, 或,拆分people列的内容。
完整修正代码
结合你的需求(计算people列所有单词在original列中作为独立单词出现的总次数),同时处理NaN的情况,修正后的代码如下:
import pandas as pd import re # 构造示例DataFrame df = pd.DataFrame({ 'original': [ 'John is a good friend', 'Mary and Peter are going to marry', 'Bond just met the Bond girl', 'Chris is having dinner', 'All Marys are here' ], 'people': ['John, Mary', 'Peter, Mary', 'Bond', pd.NA, 'Mary'] }) def count_matches(original_text, people_str): # 处理people列为空的情况 if pd.isna(people_str): return 0 # 拆分people字符串为单词列表,处理逗号后带空格或不带的情况 people_list = [p.strip() for p in people_str.split(',')] total = 0 for person in people_list: # 使用正则匹配独立单词,避免部分匹配(比如Mary不匹配Marys) matches = re.findall(rf'\b{re.escape(person)}\b', original_text) total += len(matches) return total # 计算result列 df['result'] = df.apply(lambda row: count_matches(row['original'], row['people']), axis=1) print(df)
代码说明
- 用
pd.isna处理people列的空值,直接返回0 - 拆分people字符串时,用
split(',')再strip()去除每个单词的空格,比正则拆分更简洁可靠 - 用
re.escape()处理可能包含特殊字符的人名,避免正则语法错误 - 用
re.findall()统计每个单词的出现次数,求和得到总次数 - 最终输出结果完全匹配你给出的示例
内容的提问来源于stack exchange,提问作者MG Fern
相关产品推荐
相关产品推荐

