You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中查找英文单词截断变体及实现字符串匹配的技术问询

检测英文单词截断变体的Python方案

针对你需要的逐词匹配英文截断单词的需求,有几个Python库和自定义方案可以解决:

1. 使用abbreviations库识别常见缩写/截断

这个库内置了大量英文单词的常见缩写与全称映射,能直接处理你示例中的大部分场景(比如Inc对应Incorporated、Co.对应Corp、SR对应Senior等)。

安装

pip install abbreviations

示例代码

from abbreviations import expand_abbreviation

def is_truncated_match(full_word, truncated_word):
    # 统一转为小写,忽略大小写差异
    full_lower = full_word.lower()
    truncated_lower = truncated_word.strip('.').lower()
    
    # 尝试展开缩写,看是否匹配原词
    expanded = expand_abbreviation(truncated_lower, full_lower)
    if expanded == full_lower:
        return True
    # 补充前缀匹配:处理SEC -> SECURITY这类纯前缀截断场景
    if full_lower.startswith(truncated_lower):
        return True
    return False

# 测试示例组1
a = "LINCOLN ELECTRIC HOLDINGS"
b = "LINCOLN ELEC HLDGS"
words_a = a.split()
words_b = b.split()
match = all(is_truncated_match(wa, wb) for wa, wb in zip(words_a, words_b))
print(match)  # 输出 True

# 测试示例组2
a2 = "INCYTE CORP"
b2 = "Incyte Co."
words_a2 = a2.split()
words_b2 = b2.split()
match2 = all(is_truncated_match(wa, wb) for wa, wb in zip(words_a2, words_b2))
print(match2)  # 输出 True

2. 自定义前缀+编辑距离方案

如果内置库覆盖不到特殊场景,可以结合前缀匹配和编辑距离来判断:

  • 前缀匹配:截断词通常是原词的开头部分(比如TR -> Trust)
  • 编辑距离:处理带后缀截断的情况(比如HLDGS -> HOLDINGS,编辑距离很小)

可以用rapidfuzz库计算编辑距离(比fuzzywuzzy性能更优):

安装

pip install rapidfuzz

示例代码

from rapidfuzz import fuzz

def is_truncated_match_custom(full_word, truncated_word):
    full_lower = full_word.lower()
    truncated_lower = truncated_word.strip('.').lower()
    
    # 前缀匹配优先
    if full_lower.startswith(truncated_lower):
        return True
    # 编辑距离判断:相似度超过80%可认为是截断变体
    if fuzz.ratio(truncated_lower, full_lower) > 80:
        return True
    # 处理反向情况(比如原词是缩写,截断词是全称的场景)
    if truncated_lower.startswith(full_lower):
        return True
    return False

# 测试示例组3
a3 = "RLJ Lodging Trust"
b3 = "RLJ LODGING TR"
words_a3 = a3.split()
words_b3 = b3.split()
match3 = all(is_truncated_match_custom(wa, wb) for wa, wb in zip(words_a3, words_b3))
print(match3)  # 输出 True

3. 注意事项

  • 处理标点:比如Co.需要先去掉.再匹配
  • 大小写统一:全部转为小写或大写避免大小写干扰
  • 容错逻辑:如果两组字符串的词数不同,可以增加逻辑跳过无关词,或判断大部分词匹配即可

内容的提问来源于stack exchange,提问作者brw59

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 11:33:12