在Pandas DataFrame中匹配字典值并返回对应键至新列的语义匹配问题
解决Pandas DataFrame中同义词分组的问题
我明白你现在的需求——已经能找到特定词汇,但要把同义词都归到对应的主关键词组里对吧?其实你提到的字典匹配思路完全可行,可能是细节没处理好,我给你一套具体的实现方案,一步步来:
1. 梳理同义词映射关系
首先你需要把主关键词和对应的同义词整理成一个字典,主关键词作为键,同义词列表作为值:
import pandas as pd import re # 示例同义词字典,可根据你的实际需求修改 synonym_groups = { "car": ["car", "automobile", "vehicle", "motorcar"], "bike": ["bike", "bicycle", "cycle"], "laptop": ["laptop", "notebook", "portable computer"] }
2. 反转字典实现快速查找
为了高效匹配,我们把字典反转,让每个同义词直接映射到对应的主关键词,同时统一转小写避免大小写敏感问题:
# 构建「同义词→主关键词」的映射表 syn_to_main = {} for main_key, synonyms in synonym_groups.items(): for syn in synonyms: syn_to_main[syn.lower()] = main_key
3. 编写匹配函数处理文本
接下来写一个函数,检查字符串中是否包含任何同义词,并返回对应的主关键词。这里用正则的\b匹配完整单词,避免像"carpet"被误判成"car"这类部分匹配的问题:
def get_main_keyword(text): text_lower = text.lower() # 遍历所有同义词,匹配完整单词 for syn, main_key in syn_to_main.items(): if re.search(rf"\b{re.escape(syn)}\b", text_lower): return main_key # 若未匹配到任何同义词,可返回None或原文本,根据需求调整 return None
4. 应用到Pandas DataFrame
最后把这个函数用apply方法应用到你的目标列上:
# 示例DataFrame,替换成你自己的数据集 df = pd.DataFrame({ "content": [ "I bought a new automobile yesterday", "My old bike needs new tires", "This portable computer runs fast", "I love driving my motorcar", "The carpet is dirty" # 这个不会被误匹配为car ] }) # 添加分组后的主关键词列 df["main_keyword"] = df["content"].apply(get_main_keyword)
运行后会得到这样的结果:
content main_keyword 0 I bought a new automobile yesterday car 1 My old bike needs new tires bike 2 This portable computer runs fast laptop 3 I love driving my motorcar car 4 The carpet is dirty None
额外优化建议
- 如果文本包含标点符号,可以先预处理去除:
text = re.sub(r'[^\w\s]', '', text) - 如果需要匹配多个同义词并返回所有对应主关键词,可以修改函数返回列表而非单个值
这样应该就能解决你遇到的问题了,核心就是把同义词匹配转化为字典的快速查找,再结合正则确保匹配的是完整单词~
内容的提问来源于stack exchange,提问作者Bob Harris
相关产品推荐
相关产品推荐

