You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python文本预处理如何设置规则保护指定专有名词不被错误加工

问题1:自动拼写纠错修正专有名词的解决方案

通过维护专有名词白名单的方式实现过滤,白名单内的词汇直接跳过纠错步骤即可:

import nltk
from autocorrect import spell

# 可根据业务需求扩展白名单条目
PROPER_NOUN_WHITELIST = {"skyrim"}

def autospell(text):
    spells = []
    for w in nltk.word_tokenize(text):
        if w.lower() in PROPER_NOUN_WHITELIST:
            spells.append(w)
        else:
            spells.append(spell(w))
    return " ".join(spells)

def main():
    text = "skyrim"
    spell_text = autospell(text)
    print(spell_text)

运行后输出为skyrim,不会被错误修正为syria。


问题2:仅保留E3中的数字、其余文本移除数字的解决方案

通过临时占位替换的方式规避E3被规则误伤:

def remove_numbers(text):
    # 先替换E3为唯一占位符
    text = text.replace("E3", "__TEMP_E3_MARKER__")
    # 执行全局数字移除逻辑
    output = ''.join(c for c in text if not c.isdigit())
    # 还原占位符为E3
    output = output.replace("__TEMP_E3_MARKER__", "E3")
    return output

def main():
    text = "in 2000, I am watching the E3"
    print(remove_numbers(text))

运行后输出为in , I am watching the E3,符合预期要求。


内容的提问来源于stack exchange,提问作者Erica T

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 01:06:06