使用NLTK进行对象标准化时缩写替换失效求助
问题分析与解决方案
先帮你拆解下代码没生效的核心问题,再给你修正后的实现:
1. 核心问题排查
你的代码没触发替换,主要是两个原因:
- 大小写不匹配:
lookup_dict里的键都是大写(比如'ECSC'),但代码里把输入单词转成小写后去查找——字典里根本没有小写键,自然匹配不到。 - 拼写错误:字典里的
"European Coal and Steel Commuinty"写错了,Commuinty应该是Community,就算匹配成功也会输出错误全称。 - 额外小问题:文本里的
ECSC.(带句号)、eec(小写)这类情况,原代码也处理不了。
修正后的代码
import string # 用于处理单词两端的标点 # 把字典键统一设为小写,兼容各种输入大小写 lookup_dict = { 'ec': 'European Commission', 'eu': 'European Union', "ecsc": "European Coal and Steel Community", # 修正拼写错误 "eec": "European Economic Community" } def _lookup_words(input_text): words = input_text.split() new_words = [] for word in words: # 先清理单词两端的标点,再转小写匹配字典 cleaned_word = word.strip(string.punctuation).lower() if cleaned_word in lookup_dict: # 保留原单词的标点(比如把ECSC.替换成全称加句号) word = lookup_dict[cleaned_word] + word[len(cleaned_word):] new_words.append(word) new_text = " ".join(new_words) print(new_text) return new_text # 测试输入文本 _lookup_words("The High Authority was the supranational administrative executive of the new European Coal and Steel Community ECSC. It took office first on 10 August 1952 in Luxembourg. In 1958, the Treaties of Rome had established two new communities alongside the ECSC: the eec and the European Atomic Energy Community (Euratom). However their executives were called Commissions rather than High Authorities")
关键修改说明
- 把字典键统一改为小写,不管输入是大写、小写还是混合大小写,都能匹配到对应的全称。
- 修正了
Community的拼写错误,保证输出内容准确。 - 添加了标点处理逻辑:先去掉单词两端的标点再匹配,匹配成功后把原标点加回全称末尾,避免因标点导致的匹配失败。
运行这段代码后,你就能看到所有目标缩写都被正确替换成全称了!
内容的提问来源于stack exchange,提问作者Nathan Cusack
相关产品推荐
相关产品推荐

