You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决TfidfVectorizer的stop_words参数类型错误?

问题场景

参考Jens Albrecht等人所著《Blueprints for text analysis using Python》(2020年第一版,第209页及以后)的说明,尝试对德语文本进行主题建模时,使用spaCy的德语停用词配置TfidfVectorizer,出现参数错误。

运行代码

# Load Data
import pandas as pd
# csv Datei über read_csv laden
xlsx = pd.ExcelFile("Priorisierung_der_Anforderungen.xlsx")
df = pd.read_excel(xlsx)

# Anforderungsbeschreibung in String umwandlen
df=df.astype({'Anforderungsbeschreibung':'string'})
df.info()

# "Ignore spaces after the stop..."
import re
df["paragraphs"] = df["Anforderungsbeschreibung"].map(lambda text:re.split('\.\s*\n', text))
df["number_of_paragraphs"] = df["paragraphs"].map(len)

%matplotlib inline
df.groupby('Title').agg({'number_of_paragraphs': 'mean'}).plot.bar(figsize=(24,12))


# Preparations
from sklearn.feature_extraction.text import TfidfVectorizer
from spacy.lang.de.stop_words import STOP_WORDS as stopwords

tfidf_text_vectorizer = TfidfVectorizer(stop_words=stopwords, min_df=5, max_df=0.7)
tfidf_text_vectors = tfidf_text_vectorizer.fit_transform(df['Anforderungsbeschreibung'])
tfidf_text_vectors.shape

报错信息

InvalidParameterError: The 'stop_words' parameter of TfidfVectorizer must be a str among {'english'}, an instance of 'list' or None.

完整报错栈:

InvalidParameterError                     Traceback (most recent call last)
Cell In[8], line 4
  1 #tfidf_text_vectorizer = = TfidfVectorizer(stop_words=stopwords.words('german'),)
  3 tfidf_text_vectorizer = TfidfVectorizer(stop_words=stopwords, min_df=5, max_df=0.7)
----> 4 tfidf_text_vectors = tfidf_text_vectorizer.fit_transform(df['Anforderungsbeschreibung'])
  5 tfidf_text_vectors.shape

InvalidParameterError: The 'stop_words' parameter of TfidfVectorizer must be a str among {'english'}, an instance of 'list' or None.
解决方案

问题原因

spaCy的STOP_WORDS是**集合(set)**类型,而scikit-learn的TfidfVectorizer对stop_words参数的要求仅为:

  • 字符串'english'
  • 列表(list)类型
  • None

集合类型不符合参数要求,因此抛出错误。

修改方法

将spaCy的停用词集合转换为列表,修改TfidfVectorizer的初始化代码:

tfidf_text_vectorizer = TfidfVectorizer(stop_words=list(stopwords), min_df=5, max_df=0.7)

完整修改后的关键代码片段

from sklearn.feature_extraction.text import TfidfVectorizer
from spacy.lang.de.stop_words import STOP_WORDS as stopwords

# 将stopwords集合转为列表传入
tfidf_text_vectorizer = TfidfVectorizer(stop_words=list(stopwords), min_df=5, max_df=0.7)
tfidf_text_vectors = tfidf_text_vectorizer.fit_transform(df['Anforderungsbeschreibung'])
tfidf_text_vectors.shape

内容的提问来源于stack exchange,提问作者SebastianS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 02:57:34