sklearn中TfidfVectorizer停用词移除时机及自定义停用词生效节点
Great question! Let's break down how stopword removal works in TfidfVectorizer—it inherits most text processing logic from CountVectorizer, which is key to understanding the flow.
1. When does TfidfVectorizer perform stopword removal?
Stopword removal happens right after tokenization (splitting text into individual tokens) and before any further text normalization like stemming or lemmatization (if you’ve configured those steps).
The standard processing pipeline for TfidfVectorizer (with default analyzer='word') looks like this:
- Lowercase the text (default behavior)
- Split text into word tokens (tokenization)
- Remove stopwords from the token list
- (Optional) Apply stemming/lemmatization (if using a custom analyzer)
- Count token frequencies and calculate TF-IDF scores
2. Custom stopword lists: Removal stage
Your initial推测 (speculation) is correct! When you pass a custom stopword list via the stop_words parameter, the removal occurs at the exact same stage as the built-in 'english' stopwords: post-tokenization, pre-normalization.
The official documentation confirms this behavior:
stop_words参数可选值为‘english’、列表或None(默认)……若为列表,该列表中的停用词将从生成的词元中移除,仅当analyzer='word'时生效。
The phrase "生成的词元" (generated tokens) here refers directly to the raw output of the tokenization step—unmodified word tokens, not stemmed or lemmatized versions.
What if my pipeline includes stemming?
If you’re using a stemmer (or lemmatizer) via a custom analyzer, stopword removal still happens before stemming. For example:
from sklearn.feature_extraction.text import TfidfVectorizer from nltk.stem import PorterStemmer stemmer = PorterStemmer() def stemmed_analyzer(doc): # First run the default analyzer: tokenize → remove stopwords tokens = TfidfVectorizer().build_analyzer()(doc) # Then stem the filtered tokens return [stemmer.stem(token) for token in tokens] vec = TfidfVectorizer(analyzer=stemmed_analyzer, stop_words=["your", "custom", "stopwords"])
The default analyzer handles stopword removal first, then you apply stemming to the remaining tokens. This makes practical sense because stopwords are full, unstemmed words (like "the", "and")—removing them before stemming avoids edge cases where a stemmed stopword might not match your custom list.
内容的提问来源于stack exchange,提问作者Eugenio

