如何用Python自动检测并修正reportthatexplains这类连写英文单词?
解决方案
方法1:基于词典的动态规划分割
核心思路是将连写单词拆分为词典中存在的单词组合,用动态规划找出最优拆分方式,适合处理常规连写场景。
示例代码:
import nltk from nltk.corpus import words nltk.download('words') word_dict = set(words.words()) def split_compound_word(compound): dp = [None] * (len(compound) + 1) dp[0] = [] for i in range(1, len(compound)+1): for j in range(i): if dp[j] is not None and compound[j:i].lower() in word_dict: if dp[i] is None or len(dp[j]) + 1 < len(dp[i]): dp[i] = dp[j] + [compound[j:i]] return ' '.join(dp[-1]) if dp[-1] else compound # 测试 test_words = ["informationotherwise", "havebeen", "reportthatexplains"] for word in test_words: print(f"{word} → {split_compound_word(word)}")
优势:轻量、速度快;不足:生僻词或专业术语需要额外扩充词典。
方法2:结合Hunspell拼写检查
Hunspell是LanguageTool等工具的底层依赖,支持词缀分析,能更好识别带时态、派生形式的连写单词。
先安装依赖:
pip install pyhunspell # 系统依赖(按需安装):Ubuntu用sudo apt-get install libhunspell-dev;Mac用brew install hunspell
示例代码:
import hunspell # 加载英文默认词典,可替换为专业领域词典 hobj = hunspell.HunSpell('/usr/share/hunspell/en_US.dic', '/usr/share/hunspell/en_US.aff') def split_with_hunspell(compound): for i in range(1, len(compound)): left = compound[:i] right = compound[i:] if hobj.spell(left.lower()) and hobj.spell(right.lower()): right_split = split_with_hunspell(right) return f"{left} {right_split}" return compound # 测试 test_words = ["informationotherwise", "havebeen", "reportthatexplains"] for word in test_words: print(f"{word} → {split_with_hunspell(word)}")
优势:准确率高,能处理更多变形词;不足:需要额外安装系统依赖。
方法3:预训练语言模型辅助分割
针对上下文相关的复杂连写,用预训练模型结合语境判断最优拆分,解决词典覆盖不足的问题。
先安装依赖:
pip install transformers torch
示例代码:
from transformers import pipeline fill_mask = pipeline("fill-mask", model="bert-base-uncased") def split_with_ml(compound, context=""): best_score = -1 best_split = compound for i in range(1, len(compound)): candidate = f"{compound[:i]} {compound[i:]}" mask_sent = context.replace(compound, f"{candidate[:i]} [MASK] {candidate[i+1:]}") if context else f"{candidate[:i]} [MASK] {candidate[i+1:]}" results = fill_mask(mask_sent) score = sum(res['score'] for res in results if res['token_str'] == candidate[i]) if score > best_score: best_score = score best_split = candidate parts = best_split.split() final_split = [] for part in parts: if len(part) > 5 and not fill_mask(f"{part[:3]} [MASK] {part[4:]}")[0]['token_str'] == part[3]: final_split.extend(split_with_ml(part).split()) else: final_split.append(part) return ' '.join(final_split) # 测试 test_words = ["informationotherwise", "havebeen", "reportthatexplains"] for word in test_words: print(f"{word} → {split_with_ml(word)}")
优势:能处理复杂上下文相关连写;不足:速度较慢,适合小批量高精度场景。
补充建议
- 多方法结合:先用动态规划处理简单连写,再用Hunspell或模型处理难拆分单词,兼顾效率与准确率。
- 扩充专业词典:针对领域文本,添加对应专业词汇,提升拆分精度。
- 批量优化:先提取长度超15且不在词典中的疑似连写单词,再集中处理,减少不必要的计算。
内容的提问来源于stack exchange,提问作者Martien Lubberink
相关产品推荐
相关产品推荐

