You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将自定义停用词列表传入TfidfVectorizer的自定义分词器?

解决自定义停用词传入TfidfVectorizer分词器的问题

你当前的分词器函数直接依赖全局的stop_words,可以通过以下两种常用方式将停用词列表显式传入分词器:

方法一:使用闭包生成带停用词的分词器

通过外层函数接收停用词,返回绑定了该停用词的分词器函数,内部函数可直接访问外层的停用词变量:

from nltk.stem.snowball import SnowballStemmer
import re
from sklearn.feature_extraction.text import TfidfVectorizer

def build_tokenizer(stop_words):
    def transformation_libelle(sentence):
        stemmer = SnowballStemmer("french")
        sentence_clean = re.compile(r'^[A-Z][A-Z][A-Z]\d ').sub('', sentence.replace(r'_', " ").replace(r'-', " "))
        return [stemmer.stem(token).upper() for token in re.split(r'\W+', sentence_clean) 
                if token not in stop_words and not all([char.isdigit() or char == '.' for char in token])]
    return transformation_libelle

# 你的自定义停用词列表
custom_stop_words = ["le", "la", "les", "un", "une"]

# 生成绑定停用词的分词器
custom_tokenizer = build_tokenizer(custom_stop_words)

tfidf_vectorizer = TfidfVectorizer(
    max_df=0.5,
    min_df=0,
    use_idf=True,
    tokenizer=custom_tokenizer,
    lowercase=False,
    ngram_range=(1,3),
    stop_words=None  # 设为None,避免TfidfVectorizer重复过滤停用词
)

方法二:使用functools.partial绑定参数

修改分词器函数使其接受stop_words参数,再用partial预先绑定该参数,生成符合TfidfVectorizer要求的单参数函数:

from functools import partial
from nltk.stem.snowball import SnowballStemmer
import re
from sklearn.feature_extraction.text import TfidfVectorizer

def transformation_libelle(sentence, stop_words):
    stemmer = SnowballStemmer("french")
    sentence_clean = re.compile(r'^[A-Z][A-Z][A-Z]\d ').sub('', sentence.replace(r'_', " ").replace(r'-', " "))
    return [stemmer.stem(token).upper() for token in re.split(r'\W+', sentence_clean) 
            if token not in stop_words and not all([char.isdigit() or char == '.' for char in token])]

custom_stop_words = ["le", "la", "les", "un", "une"]

# 绑定stop_words参数,生成仅需接收sentence的函数
custom_tokenizer = partial(transformation_libelle, stop_words=custom_stop_words)

tfidf_vectorizer = TfidfVectorizer(
    max_df=0.5,
    min_df=0,
    use_idf=True,
    tokenizer=custom_tokenizer,
    lowercase=False,
    ngram_range=(1,3),
    stop_words=None
)

注意事项

无论采用哪种方法,都要将TfidfVectorizer的stop_words参数设为None,避免双重过滤停用词导致结果偏差。

内容的提问来源于stack exchange,提问作者Lefloch Had

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 10:55:16