如何匹配pandas DataFrame行与字典键值并新增列存储匹配对应值
解决方案
你可以直接复用已经写好的正则匹配结果,只需要新增一行逻辑提取匹配到的具体文本即可,无需修改原有匹配逻辑:
import pandas as pd entertainment_dict = { "Food": ["McDonald", "Five Guys", "KFC"], "Music": ["Taylor Swift", "Jay Z", "One Direction"], "TV": ["Big Bang Theory", "Queen of South", "Ted Lasso"] } data = {'text':["Kevin Lee has bought a Taylor Swift's CD and eaten at McDonald.", "The best burger in McDonald is cheeze buger.", "Kevin Lee is planning to watch the Big Bang Theory and eat at KFC."]} df = pd.DataFrame(data) regex = '|'.join(f'(?P<{k}>{"|".join(v)})' for k,v in entertainment_dict.items()) # 先把匹配结果存下来,避免重复计算 matches = df['text'].str.extractall(regex) # 原有labels生成逻辑不变 df['labels'] = ((matches.notnull().groupby(level=0).max()*entertainment_dict.keys()) .apply(lambda r: ','.join([i for i in r if i]) , axis=1) ) # 新增words列:提取每行所有非空匹配值,用逗号加空格拼接,无匹配的行填充为空字符串 df['words'] = matches.stack().groupby(level=0).apply(', '.join).fillna('')
运行后得到的结果和你给出的预期输出完全一致。
逻辑说明
str.extractall返回的结果第一层索引是原DataFrame的行号,第二层索引是匹配次数,每一列对应正则里的命名分组(也就是字典的key),匹配到值就存对应文本,没匹配到就是NaNstack()会自动剔除所有NaN值,仅保留实际匹配到的文本- 按第一层索引(原行号)分组后拼接,就能得到每行所有匹配到的具体词汇
内容的提问来源于stack exchange,提问作者codedancer
相关产品推荐
相关产品推荐

