Python文本预处理如何设置规则保护指定专有名词不被错误加工
问题1:自动拼写纠错修正专有名词的解决方案
通过维护专有名词白名单的方式实现过滤,白名单内的词汇直接跳过纠错步骤即可:
import nltk from autocorrect import spell # 可根据业务需求扩展白名单条目 PROPER_NOUN_WHITELIST = {"skyrim"} def autospell(text): spells = [] for w in nltk.word_tokenize(text): if w.lower() in PROPER_NOUN_WHITELIST: spells.append(w) else: spells.append(spell(w)) return " ".join(spells) def main(): text = "skyrim" spell_text = autospell(text) print(spell_text)
运行后输出为skyrim,不会被错误修正为syria。
问题2:仅保留E3中的数字、其余文本移除数字的解决方案
通过临时占位替换的方式规避E3被规则误伤:
def remove_numbers(text): # 先替换E3为唯一占位符 text = text.replace("E3", "__TEMP_E3_MARKER__") # 执行全局数字移除逻辑 output = ''.join(c for c in text if not c.isdigit()) # 还原占位符为E3 output = output.replace("__TEMP_E3_MARKER__", "E3") return output def main(): text = "in 2000, I am watching the E3" print(remove_numbers(text))
运行后输出为in , I am watching the E3,符合预期要求。
内容的提问来源于stack exchange,提问作者Erica T
相关产品推荐
相关产品推荐

