NLTK分词后移除停用词和标点失效的问题排查与解决方案咨询
问题原因
- 你的
tokenized_text是二维嵌套列表:外层每个元素是单句的分词结果列表,内层才是单个分词。你直接遍历外层列表判断word not in stopwords,实际是把整句的列表和单个停用词字符串做匹配,永远不会成立,所以没有任何内容被过滤。 - 现有逻辑没有添加标点过滤的规则,也无法实现你去除标点的需求。
可直接运行的实现方案
这里提供两种常用场景的写法,你可以根据后续训练ngram的需求选择:
方案1:保留句子边界(适合按句子粒度训练ngram)
import string from nltk.corpus import gutenberg, stopwords from nltk.tokenize import sent_tokenize, word_tokenize # 提前加载停用词和标点为集合,大幅提升查找效率 en_stopwords = set(stopwords.words('english')) punctuations = set(string.punctuation) text = gutenberg.raw('carroll-alice.txt')[1:600] # 原分词逻辑不变 tokenized_text = [list(map(str.lower, word_tokenize(sent))) for sent in sent_tokenize(text)] # 逐句过滤停用词和标点 clean_text = [ [word for word in sent if word not in en_stopwords and word not in punctuations] for sent in tokenized_text ] print(tokenized_text) print(clean_text)
方案2:打平为一维token列表(不需要保留句子结构时使用)
import string from nltk.corpus import gutenberg, stopwords from nltk.tokenize import word_tokenize en_stopwords = set(stopwords.words('english')) punctuations = set(string.punctuation) text = gutenberg.raw('carroll-alice.txt')[1:600] all_tokens = list(map(str.lower, word_tokenize(text))) clean_text = [word for word in all_tokens if word not in en_stopwords and word not in punctuations]
内容的提问来源于stack exchange,提问作者HHH
相关产品推荐
相关产品推荐

