You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用TfidfVectorizer自定义预处理移除二元组时遇TypeError的解决方法

问题描述

尝试移除TfidfVectorizer生成的二元组(bi-grams),自定义了预处理函数remove_bigrams,但执行fit_transform时抛出TypeError: expected string or bytes-like object错误。

测试代码如下:

doc2 = ['this is a test past performance here is another that has aa aa adding builing cat dog horse hurricane', 
        'another that has aa aa and start date and hurricane hitting south carolina']

def remove_bigrams(doc):
    gram_2 = ['past performance', 'start date', 'aa aa']
    res = []
    for record in doc:
        the_string = record
        for phrase in gram_2:
            the_string = the_string.replace(phrase, "")
        res.append(the_string)
    return res

remove_bigrams(doc2)

TfidfVectorizer实例化代码:

from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS as stop_words
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.feature_extraction import text

custom_stop_words = [i for i in stop_words]

vec = text.TfidfVectorizer(stop_words=custom_stop_words,
                           analyzer='word',
                           ngram_range=(2, 2),
                           preprocessor=remove_bigrams,
                          )

features = vec.fit_transform(doc2)

错误栈:

---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
Input In [49], in <cell line: 5>()
      3 #t3_cv = CountVectorizer(t2, stop_words = stop_words)
      4 vec = text.TfidfVectorizer(stop_words=custom_stop_words, analyzer='word', ngram_range = (2,2), preprocessor = remove_bigrams)
----> 5 features = vec.fit_transform(doc2)

File c:\Development_Solutions\Sandbox\SBVE\lib\site-packages\sklearn\feature_extraction\text.py:2079, in TfidfVectorizer.fit_transform(self, raw_documents, y)
   2072 self._check_params()
   2073 self._tfidf = TfidfTransformer(
   2074     norm=self.norm,
   2075     use_idf=self.use_idf,
   2076     smooth_idf=self.smooth_idf,
   2077     sublinear_tf=self.sublinear_tf,
   2078 )
-> 2079 X = super().fit_transform(raw_documents)
   2080 self._tfidf.fit(X)
   2081 # X is already a transformed view of raw_documents so
   2082 # we set copy to False

File c:\Development_Solutions\Sandbox\SBVE\lib\site-packages\sklearn\feature_extraction\text.py:1338, in CountVectorizer.fit_transform(self, raw_documents, y)
   1330             warnings.warn(
   1331                 "Upper case characters found in"
   1332                 " vocabulary while 'lowercase'"
   1333                 " is True. These entries will not"
   1334                 " be matched with any documents"
   1335             )
   1336             break
-> 1338 vocabulary, X = self._count_vocab(raw_documents, self.fixed_vocabulary_)
   1340 if self.binary:
   1341     X.data.fill(1)

File c:\Development_Solutions\Sandbox\SBVE\lib\site-packages\sklearn\feature_extraction\text.py:1209, in CountVectorizer._count_vocab(self, raw_documents, fixed_vocab)
   1207 for doc in raw_documents:
   1208     feature_counter = {}
-> 1209     for feature in analyze(doc):
   1210         try:
   1211             feature_idx = vocabulary[feature]

File c:\Development_Solutions\Sandbox\SBVE\lib\site-packages\sklearn\feature_extraction\text.py:113, in _analyze(doc, analyzer, tokenizer, ngrams, preprocessor, decoder, stop_words)
    111     doc = preprocessor(doc)
    112 if tokenizer is not None:
-> 113     doc = tokenizer(doc)
    114 if ngrams is not None:
    115     if stop_words is not None:

TypeError: expected string or bytes-like object
错误原因

TfidfVectorizer的preprocessor参数要求传入的是处理单个字符串的函数:当Vectorizer遍历输入的文档列表时,会把每个单独的文档字符串传入preprocessor,而你写的remove_bigrams函数是接收整个文档列表(list)并返回列表的,导致函数内部遍历单个字符串的每个字符(把字符串当成可迭代对象),最终返回一个字符组成的列表,后续的tokenizer收到的是列表而非字符串,触发类型错误。

解决方案

有两种可行的修正方式:

方式一:修改预处理函数为处理单个字符串

把remove_bigrams改成接收单个字符串、返回处理后的字符串的函数,直接作为preprocessor传入:

def remove_bigrams(doc):
    gram_2 = ['past performance', 'start date', 'aa aa']
    the_string = doc
    for phrase in gram_2:
        the_string = the_string.replace(phrase, "")
    return the_string

# 后续Vectorizer代码不变
vec = text.TfidfVectorizer(stop_words=custom_stop_words,
                           analyzer='word',
                           ngram_range=(2, 2),
                           preprocessor=remove_bigrams,
                          )
features = vec.fit_transform(doc2)

方式二:先预处理整个文档列表,再传入Vectorizer

先调用remove_bigrams处理整个doc2,得到预处理后的文档列表,再传入fit_transform,此时不需要设置preprocessor参数:

# 先预处理文档列表
processed_docs = remove_bigrams(doc2)

# 实例化Vectorizer时去掉preprocessor参数
vec = text.TfidfVectorizer(stop_words=custom_stop_words,
                           analyzer='word',
                           ngram_range=(2, 2),
                          )
features = vec.fit_transform(processed_docs)

额外说明

如果你的需求是在生成二元组后排除指定的二元组,而不是先从文本中删除这些短语,更合适的做法是:

  1. 先正常生成所有二元组的TF-IDF特征
  2. 获取特征名列表,过滤掉不需要的二元组
  3. 提取对应特征列

示例代码:

vec = text.TfidfVectorizer(stop_words=custom_stop_words,
                           analyzer='word',
                           ngram_range=(2, 2),
                          )
features = vec.fit_transform(doc2)

# 获取所有特征名
feature_names = vec.get_feature_names_out()
# 过滤掉需要排除的二元组
exclude_grams = {'past performance', 'start date', 'aa aa'}
keep_indices = [i for i, name in enumerate(feature_names) if name not in exclude_grams]
# 提取保留的特征
filtered_features = features[:, keep_indices]

内容的提问来源于stack exchange,提问作者TER

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 19:58:10