使用TfidfVectorizer自定义预处理移除二元组时遇TypeError的解决方法
问题描述
尝试移除TfidfVectorizer生成的二元组(bi-grams),自定义了预处理函数remove_bigrams,但执行fit_transform时抛出TypeError: expected string or bytes-like object错误。
测试代码如下:
doc2 = ['this is a test past performance here is another that has aa aa adding builing cat dog horse hurricane', 'another that has aa aa and start date and hurricane hitting south carolina'] def remove_bigrams(doc): gram_2 = ['past performance', 'start date', 'aa aa'] res = [] for record in doc: the_string = record for phrase in gram_2: the_string = the_string.replace(phrase, "") res.append(the_string) return res remove_bigrams(doc2)
TfidfVectorizer实例化代码:
from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS as stop_words from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.feature_extraction import text custom_stop_words = [i for i in stop_words] vec = text.TfidfVectorizer(stop_words=custom_stop_words, analyzer='word', ngram_range=(2, 2), preprocessor=remove_bigrams, ) features = vec.fit_transform(doc2)
错误栈:
--------------------------------------------------------------------------- TypeError Traceback (most recent call last) Input In [49], in <cell line: 5>() 3 #t3_cv = CountVectorizer(t2, stop_words = stop_words) 4 vec = text.TfidfVectorizer(stop_words=custom_stop_words, analyzer='word', ngram_range = (2,2), preprocessor = remove_bigrams) ----> 5 features = vec.fit_transform(doc2) File c:\Development_Solutions\Sandbox\SBVE\lib\site-packages\sklearn\feature_extraction\text.py:2079, in TfidfVectorizer.fit_transform(self, raw_documents, y) 2072 self._check_params() 2073 self._tfidf = TfidfTransformer( 2074 norm=self.norm, 2075 use_idf=self.use_idf, 2076 smooth_idf=self.smooth_idf, 2077 sublinear_tf=self.sublinear_tf, 2078 ) -> 2079 X = super().fit_transform(raw_documents) 2080 self._tfidf.fit(X) 2081 # X is already a transformed view of raw_documents so 2082 # we set copy to False File c:\Development_Solutions\Sandbox\SBVE\lib\site-packages\sklearn\feature_extraction\text.py:1338, in CountVectorizer.fit_transform(self, raw_documents, y) 1330 warnings.warn( 1331 "Upper case characters found in" 1332 " vocabulary while 'lowercase'" 1333 " is True. These entries will not" 1334 " be matched with any documents" 1335 ) 1336 break -> 1338 vocabulary, X = self._count_vocab(raw_documents, self.fixed_vocabulary_) 1340 if self.binary: 1341 X.data.fill(1) File c:\Development_Solutions\Sandbox\SBVE\lib\site-packages\sklearn\feature_extraction\text.py:1209, in CountVectorizer._count_vocab(self, raw_documents, fixed_vocab) 1207 for doc in raw_documents: 1208 feature_counter = {} -> 1209 for feature in analyze(doc): 1210 try: 1211 feature_idx = vocabulary[feature] File c:\Development_Solutions\Sandbox\SBVE\lib\site-packages\sklearn\feature_extraction\text.py:113, in _analyze(doc, analyzer, tokenizer, ngrams, preprocessor, decoder, stop_words) 111 doc = preprocessor(doc) 112 if tokenizer is not None: -> 113 doc = tokenizer(doc) 114 if ngrams is not None: 115 if stop_words is not None: TypeError: expected string or bytes-like object
错误原因
TfidfVectorizer的preprocessor参数要求传入的是处理单个字符串的函数:当Vectorizer遍历输入的文档列表时,会把每个单独的文档字符串传入preprocessor,而你写的remove_bigrams函数是接收整个文档列表(list)并返回列表的,导致函数内部遍历单个字符串的每个字符(把字符串当成可迭代对象),最终返回一个字符组成的列表,后续的tokenizer收到的是列表而非字符串,触发类型错误。
解决方案
有两种可行的修正方式:
方式一:修改预处理函数为处理单个字符串
把remove_bigrams改成接收单个字符串、返回处理后的字符串的函数,直接作为preprocessor传入:
def remove_bigrams(doc): gram_2 = ['past performance', 'start date', 'aa aa'] the_string = doc for phrase in gram_2: the_string = the_string.replace(phrase, "") return the_string # 后续Vectorizer代码不变 vec = text.TfidfVectorizer(stop_words=custom_stop_words, analyzer='word', ngram_range=(2, 2), preprocessor=remove_bigrams, ) features = vec.fit_transform(doc2)
方式二:先预处理整个文档列表,再传入Vectorizer
先调用remove_bigrams处理整个doc2,得到预处理后的文档列表,再传入fit_transform,此时不需要设置preprocessor参数:
# 先预处理文档列表 processed_docs = remove_bigrams(doc2) # 实例化Vectorizer时去掉preprocessor参数 vec = text.TfidfVectorizer(stop_words=custom_stop_words, analyzer='word', ngram_range=(2, 2), ) features = vec.fit_transform(processed_docs)
额外说明
如果你的需求是在生成二元组后排除指定的二元组,而不是先从文本中删除这些短语,更合适的做法是:
- 先正常生成所有二元组的TF-IDF特征
- 获取特征名列表,过滤掉不需要的二元组
- 提取对应特征列
示例代码:
vec = text.TfidfVectorizer(stop_words=custom_stop_words, analyzer='word', ngram_range=(2, 2), ) features = vec.fit_transform(doc2) # 获取所有特征名 feature_names = vec.get_feature_names_out() # 过滤掉需要排除的二元组 exclude_grams = {'past performance', 'start date', 'aa aa'} keep_indices = [i for i, name in enumerate(feature_names) if name not in exclude_grams] # 提取保留的特征 filtered_features = features[:, keep_indices]
内容的提问来源于stack exchange,提问作者TER
相关产品推荐
相关产品推荐

