Pandas跨数据框匹配:先完全匹配再取前两词匹配的实现求助
解决Pandas中DataFrame列的分层字符串匹配问题
先构造示例数据方便测试:
import pandas as pd # 示例DataFrame df1 = pd.DataFrame({'company': ['Advanced Automation Co.', 'Tech Global Ltd.', 'Green Energy Inc.', 'Fast Logistics']}) df2 = pd.DataFrame({'entity': ['Advanced Automation', 'Tech Global', 'Green Energy Group', 'Fast Logistics']})
实现逻辑:先完整匹配,再用前两个词匹配
通过自定义函数结合apply方法实现需求,同时将df2的实体转成集合提升查询效率:
# 提取df2的实体集合(集合查询速度远快于列表) entity_set = set(df2['entity']) def match_entity(company_name): # 先尝试完整字符串匹配 if company_name in entity_set: return company_name # 拆分字符串取前两个词 word_list = company_name.split() if len(word_list) >= 2: first_two_words = ' '.join(word_list[:2]) if first_two_words in entity_set: return first_two_words # 无匹配返回空值 return None # 应用函数到df1的company列 df1['matched_entity'] = df1['company'].apply(match_entity)
运行结果
执行后df1的内容如下:
company matched_entity 0 Advanced Automation Co. Advanced Automation 1 Tech Global Ltd. Tech Global 2 Green Energy Inc. NaN 3 Fast Logistics Fast Logistics
可选优化:清洗字符串标点
如果遇到带标点的公司名(比如末尾的., ,),可以先做简单清洗再匹配,避免标点干扰:
import string def match_entity(company_name): # 移除末尾的标点符号 cleaned_name = company_name.rstrip(string.punctuation) if cleaned_name in entity_set: return cleaned_name word_list = cleaned_name.split() if len(word_list) >= 2: first_two_words = ' '.join(word_list[:2]) if first_two_words in entity_set: return first_two_words return None df1['matched_entity'] = df1['company'].apply(match_entity)
内容的提问来源于stack exchange,提问作者Anthony
相关产品推荐
相关产品推荐

