You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CountVectorizer配置词形还原后停用词未被移除的技术问题

问题原因分析

当你给CountVectorizer指定自定义tokenizer后,它的内置停用词过滤机制会被跳过。这是因为默认情况下,停用词移除是在CountVectorizer的默认分词流程之后执行的;一旦你替换了分词器,整个分词后的处理逻辑(包括停用词过滤)就完全由你的自定义tokenizer接管了,所以原来的stop_words参数自然不会生效。

解决方案:将停用词过滤整合到自定义分词器中

你需要在LemmaTokenizer的逻辑里,手动加入停用词的过滤步骤。这里要注意西班牙语停用词列表是小写的,所以要先把分词后的token转小写再做匹配,避免因大小写不匹配导致过滤失效。

修改后的完整代码如下:

import nltk
from pattern.es import lemma
from nltk import word_tokenize
from nltk.corpus import stopwords
from sklearn.feature_extraction.text import CountVectorizer

# 确保nltk停用词和分词模型已下载
nltk.download('stopwords')
nltk.download('punkt')

class LemmaTokenizer(object):
    def __init__(self):
        # 加载西班牙语停用词并转为集合(提高查询效率),同时统一为小写
        self.stop_words = {word.lower() for word in stopwords.words('spanish')}
    
    def __call__(self, text):
        # 1. 对文本进行分词
        tokens = word_tokenize(text)
        # 2. 过滤停用词:排除转小写后属于停用词的token
        filtered_tokens = [token for token in tokens if token.lower() not in self.stop_words]
        # 3. 对剩余token进行词形还原
        return [lemma(token) for token in filtered_tokens]

# 初始化vectorizer时无需再传stop_words参数
vectorizer = CountVectorizer(tokenizer=LemmaTokenizer())

# 测试示例
sentences = ["EVOLUCIÓN es un proceso interesante, pero no es fácil de explicar", 
             "El perro corre por el parque, y el gato se esconde"]
vectorizer.fit(sentences)
print(vectorizer.get_feature_names_out())
额外注意事项
  • 如果你担心词形还原后的结果可能仍包含停用词(比如某些特殊形式的停用词还原后才匹配),可以调整顺序:先做词形还原,再过滤停用词。
  • pattern.es的lemma函数对西班牙语的支持虽然不错,但如果需要更精准的词形还原,也可以考虑使用spaCy的西班牙语模型(不过需要额外安装配置)。
  • 用集合存储停用词是为了让查询操作更快,相比列表的in操作,集合的时间复杂度是O(1),处理大文本时效率更高。

内容的提问来源于stack exchange,提问作者ambigus9

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:45:02