词形还原语料库时遇Word List Corpus Reader不可调用错误的解决咨询
错误原因与解决方法
核心错误点
你遇到的「Word List Corpus Reader not callable」错误,是因为**stopwords('english')的调用方式错误**:stopwords是NLTK库中的语料库读取器对象,不是可调用函数,不能直接通过括号传参调用。同时你的词形还原代码也存在导入和实例化的问题。
修正步骤
- 正确获取停用词
停用词的正确获取方式是stopwords.words('english'),需要先导入NLTK的stopwords模块。 - 正确配置词形还原工具
词形还原需先导入WordNetLemmatizer并实例化对象,再调用其lemmatize方法处理单词。
修正后的完整代码
import re from nltk.corpus import stopwords from nltk.stem import WordNetLemmatizer # 实例化词形还原器 lemmatizer = WordNetLemmatizer() corpus = [] for sentence in sentences: # 去除非字母字符 review = re.sub('[^a-zA-Z]', ' ', sentence) # 转为小写 review = review.lower() # 拆分单词 review = review.split() # 词形还原+过滤停用词 review = [lemmatizer.lemmatize(word) for word in review if word not in set(stopwords.words('english'))] # 拼接为字符串 review = ' '.join(review) corpus.append(review)
额外注意事项
- 提前确保已下载NLTK所需语料库,若未下载,先运行以下代码:
import nltk nltk.download('stopwords') nltk.download('wordnet') - 原代码中
range(len(sentences))的写法可简化为直接遍历sentences,更符合Python的简洁风格。
内容的提问来源于stack exchange,提问作者Prince Thakkar
相关产品推荐
相关产品推荐

