You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

sklearn中TfidfVectorizer停用词移除时机及自定义停用词生效节点

sklearn TfidfVectorizer: When Stopwords Are Removed (Including Custom Lists)

Great question! Let's break down how stopword removal works in TfidfVectorizer—it inherits most text processing logic from CountVectorizer, which is key to understanding the flow.

1. When does TfidfVectorizer perform stopword removal?

Stopword removal happens right after tokenization (splitting text into individual tokens) and before any further text normalization like stemming or lemmatization (if you’ve configured those steps).

The standard processing pipeline for TfidfVectorizer (with default analyzer='word') looks like this:

  • Lowercase the text (default behavior)
  • Split text into word tokens (tokenization)
  • Remove stopwords from the token list
  • (Optional) Apply stemming/lemmatization (if using a custom analyzer)
  • Count token frequencies and calculate TF-IDF scores

2. Custom stopword lists: Removal stage

Your initial推测 (speculation) is correct! When you pass a custom stopword list via the stop_words parameter, the removal occurs at the exact same stage as the built-in 'english' stopwords: post-tokenization, pre-normalization.

The official documentation confirms this behavior:

stop_words参数可选值为‘english’、列表或None(默认)……若为列表,该列表中的停用词将从生成的词元中移除,仅当analyzer='word'时生效。

The phrase "生成的词元" (generated tokens) here refers directly to the raw output of the tokenization step—unmodified word tokens, not stemmed or lemmatized versions.

What if my pipeline includes stemming?

If you’re using a stemmer (or lemmatizer) via a custom analyzer, stopword removal still happens before stemming. For example:

from sklearn.feature_extraction.text import TfidfVectorizer
from nltk.stem import PorterStemmer

stemmer = PorterStemmer()
def stemmed_analyzer(doc):
    # First run the default analyzer: tokenize → remove stopwords
    tokens = TfidfVectorizer().build_analyzer()(doc)
    # Then stem the filtered tokens
    return [stemmer.stem(token) for token in tokens]

vec = TfidfVectorizer(analyzer=stemmed_analyzer, stop_words=["your", "custom", "stopwords"])

The default analyzer handles stopword removal first, then you apply stemming to the remaining tokens. This makes practical sense because stopwords are full, unstemmed words (like "the", "and")—removing them before stemming avoids edge cases where a stemmed stopword might not match your custom list.

内容的提问来源于stack exchange,提问作者Eugenio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:26:42