You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现DataFrame关联:校验多子串匹配目标单词前缀

基于规则的前缀匹配实现方案

你这个场景的匹配规则是明确的硬逻辑,完全不需要使用fuzzywuzzy这类模糊匹配库——模糊匹配的阈值权衡问题本身就是因为规则不明确才会出现,直接按你定义的前缀校验规则实现即可,准确率100%,几百条数据规模下没有任何性能压力。

核心匹配规则

匹配判定逻辑完全对齐你的需求:

  • 所有文本统一转小写,规避大小写差异带来的匹配失败
  • 缩写按空格拆分为独立的前缀子串
  • 全称按空格拆分为独立单词
  • 仅当缩写拆分出的所有前缀子串,都能在全称的单词列表中找到至少一个以该子串为开头的单词时,判定二者匹配成功

按这个规则校验你给出的示例数据:

  • dis ijp:所有全称中不存在以ijp为前缀的单词,无匹配
  • dis inf:全称Disperzia na koncentrát na infúznu disperziu中,Disperzia前缀匹配dis、infúzna前缀匹配inf,匹配成功
  • dis inj:所有全称中不存在以inj为前缀的单词,无匹配
    结果完全符合预期。

实现代码

仅依赖pandas,无需安装额外第三方库:

import pandas as pd

def check_match(abbrev: str, full_term: str) -> bool:
    # 预处理:转小写、拆分去空
    abbrev_tokens = [t.strip() for t in abbrev.lower().split() if t.strip()]
    term_words = [w.strip().lower() for w in full_term.split() if w.strip()]
    
    # 逐个子串校验前缀匹配
    for token in abbrev_tokens:
        has_match = any(word.startswith(token) for word in term_words)
        if not has_match:
            return False
    return True

def abbrev_term_match(df1: pd.DataFrame, df2: pd.DataFrame, abbrev_col: str = "abrev", term_col: str = "term") -> pd.DataFrame:
    # 提前取出全称列表减少重复取数开销
    all_terms = df2[term_col].tolist()
    result = df1.copy()
    # 为每个缩写匹配符合规则的全称
    result["matches"] = result[abbrev_col].apply(
        lambda ab: "; ".join([term for term in all_terms if check_match(ab, term)])
    )
    return result

使用方式

直接传入两个DataFrame调用即可:

matched_result = abbrev_term_match(df1, df2)

性能说明

单条缩写和全称的校验是极轻量的字符串前缀判断,哪怕df1、df2各有1000条数据,总运算量也仅百万次级别的简单操作,毫秒级即可跑完,不存在模糊匹配的错配、漏配问题。如果后续数据量上涨到十万级,可以提前为全称构建前缀索引进一步提速,当前几百条的规模完全不需要额外优化。


内容的提问来源于stack exchange,提问作者Pedro Domingues

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 11:09:28