You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python自动检测并修正reportthatexplains这类连写英文单词?

解决方案

方法1:基于词典的动态规划分割

核心思路是将连写单词拆分为词典中存在的单词组合,用动态规划找出最优拆分方式,适合处理常规连写场景。

示例代码:

import nltk
from nltk.corpus import words
nltk.download('words')

word_dict = set(words.words())

def split_compound_word(compound):
    dp = [None] * (len(compound) + 1)
    dp[0] = []
    
    for i in range(1, len(compound)+1):
        for j in range(i):
            if dp[j] is not None and compound[j:i].lower() in word_dict:
                if dp[i] is None or len(dp[j]) + 1 < len(dp[i]):
                    dp[i] = dp[j] + [compound[j:i]]
    return ' '.join(dp[-1]) if dp[-1] else compound

# 测试
test_words = ["informationotherwise", "havebeen", "reportthatexplains"]
for word in test_words:
    print(f"{word} → {split_compound_word(word)}")

优势:轻量、速度快;不足:生僻词或专业术语需要额外扩充词典。

方法2:结合Hunspell拼写检查

Hunspell是LanguageTool等工具的底层依赖,支持词缀分析,能更好识别带时态、派生形式的连写单词。

先安装依赖:

pip install pyhunspell
# 系统依赖(按需安装):Ubuntu用sudo apt-get install libhunspell-dev;Mac用brew install hunspell

示例代码:

import hunspell

# 加载英文默认词典,可替换为专业领域词典
hobj = hunspell.HunSpell('/usr/share/hunspell/en_US.dic', '/usr/share/hunspell/en_US.aff')

def split_with_hunspell(compound):
    for i in range(1, len(compound)):
        left = compound[:i]
        right = compound[i:]
        if hobj.spell(left.lower()) and hobj.spell(right.lower()):
            right_split = split_with_hunspell(right)
            return f"{left} {right_split}"
    return compound

# 测试
test_words = ["informationotherwise", "havebeen", "reportthatexplains"]
for word in test_words:
    print(f"{word} → {split_with_hunspell(word)}")

优势:准确率高,能处理更多变形词;不足:需要额外安装系统依赖。

方法3:预训练语言模型辅助分割

针对上下文相关的复杂连写,用预训练模型结合语境判断最优拆分,解决词典覆盖不足的问题。

先安装依赖:

pip install transformers torch

示例代码:

from transformers import pipeline

fill_mask = pipeline("fill-mask", model="bert-base-uncased")

def split_with_ml(compound, context=""):
    best_score = -1
    best_split = compound
    for i in range(1, len(compound)):
        candidate = f"{compound[:i]} {compound[i:]}"
        mask_sent = context.replace(compound, f"{candidate[:i]} [MASK] {candidate[i+1:]}") if context else f"{candidate[:i]} [MASK] {candidate[i+1:]}"
        results = fill_mask(mask_sent)
        score = sum(res['score'] for res in results if res['token_str'] == candidate[i])
        if score > best_score:
            best_score = score
            best_split = candidate
    
    parts = best_split.split()
    final_split = []
    for part in parts:
        if len(part) > 5 and not fill_mask(f"{part[:3]} [MASK] {part[4:]}")[0]['token_str'] == part[3]:
            final_split.extend(split_with_ml(part).split())
        else:
            final_split.append(part)
    return ' '.join(final_split)

# 测试
test_words = ["informationotherwise", "havebeen", "reportthatexplains"]
for word in test_words:
    print(f"{word} → {split_with_ml(word)}")

优势:能处理复杂上下文相关连写;不足:速度较慢,适合小批量高精度场景。

补充建议

  • 多方法结合:先用动态规划处理简单连写,再用Hunspell或模型处理难拆分单词,兼顾效率与准确率。
  • 扩充专业词典:针对领域文本,添加对应专业词汇,提升拆分精度。
  • 批量优化:先提取长度超15且不在词典中的疑似连写单词,再集中处理,减少不必要的计算。

内容的提问来源于stack exchange,提问作者Martien Lubberink

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 06:51:18