You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLTK分词后移除停用词和标点失效的问题排查与解决方案咨询

问题原因

  • 你的tokenized_text是二维嵌套列表:外层每个元素是单句的分词结果列表,内层才是单个分词。你直接遍历外层列表判断word not in stopwords,实际是把整句的列表和单个停用词字符串做匹配,永远不会成立,所以没有任何内容被过滤。
  • 现有逻辑没有添加标点过滤的规则,也无法实现你去除标点的需求。

可直接运行的实现方案

这里提供两种常用场景的写法,你可以根据后续训练ngram的需求选择:

方案1:保留句子边界(适合按句子粒度训练ngram)

import string
from nltk.corpus import gutenberg, stopwords
from nltk.tokenize import sent_tokenize, word_tokenize

# 提前加载停用词和标点为集合,大幅提升查找效率
en_stopwords = set(stopwords.words('english'))
punctuations = set(string.punctuation)

text = gutenberg.raw('carroll-alice.txt')[1:600]
# 原分词逻辑不变
tokenized_text = [list(map(str.lower, word_tokenize(sent))) for sent in sent_tokenize(text)]
# 逐句过滤停用词和标点
clean_text = [
    [word for word in sent if word not in en_stopwords and word not in punctuations]
    for sent in tokenized_text
]

print(tokenized_text)
print(clean_text)

方案2:打平为一维token列表(不需要保留句子结构时使用)

import string
from nltk.corpus import gutenberg, stopwords
from nltk.tokenize import word_tokenize

en_stopwords = set(stopwords.words('english'))
punctuations = set(string.punctuation)

text = gutenberg.raw('carroll-alice.txt')[1:600]
all_tokens = list(map(str.lower, word_tokenize(text)))
clean_text = [word for word in all_tokens if word not in en_stopwords and word not in punctuations]

内容的提问来源于stack exchange,提问作者HHH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 16:24:00