You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas跨数据框匹配:先完全匹配再取前两词匹配的实现求助

解决Pandas中DataFrame列的分层字符串匹配问题

先构造示例数据方便测试:

import pandas as pd

# 示例DataFrame
df1 = pd.DataFrame({'company': ['Advanced Automation Co.', 'Tech Global Ltd.', 'Green Energy Inc.', 'Fast Logistics']})
df2 = pd.DataFrame({'entity': ['Advanced Automation', 'Tech Global', 'Green Energy Group', 'Fast Logistics']})

实现逻辑:先完整匹配,再用前两个词匹配

通过自定义函数结合apply方法实现需求,同时将df2的实体转成集合提升查询效率:

# 提取df2的实体集合(集合查询速度远快于列表)
entity_set = set(df2['entity'])

def match_entity(company_name):
    # 先尝试完整字符串匹配
    if company_name in entity_set:
        return company_name
    # 拆分字符串取前两个词
    word_list = company_name.split()
    if len(word_list) >= 2:
        first_two_words = ' '.join(word_list[:2])
        if first_two_words in entity_set:
            return first_two_words
    # 无匹配返回空值
    return None

# 应用函数到df1的company列
df1['matched_entity'] = df1['company'].apply(match_entity)

运行结果

执行后df1的内容如下:

company      matched_entity
0  Advanced Automation Co.  Advanced Automation
1        Tech Global Ltd.        Tech Global
2       Green Energy Inc.                 NaN
3           Fast Logistics       Fast Logistics

可选优化:清洗字符串标点

如果遇到带标点的公司名(比如末尾的., ,),可以先做简单清洗再匹配,避免标点干扰:

import string

def match_entity(company_name):
    # 移除末尾的标点符号
    cleaned_name = company_name.rstrip(string.punctuation)
    if cleaned_name in entity_set:
        return cleaned_name
    word_list = cleaned_name.split()
    if len(word_list) >= 2:
        first_two_words = ' '.join(word_list[:2])
        if first_two_words in entity_set:
            return first_two_words
    return None

df1['matched_entity'] = df1['company'].apply(match_entity)

内容的提问来源于stack exchange,提问作者Anthony

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 09:30:37