You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Sklearn文本分类管道时遇空词汇表ValueError求助

解决Sklearn文本分类管道的空词汇表错误

错误本质

ValueError: "empty vocabulary; perhaps the documents only contain stop words" 是因为CountVectorizer在构建词汇表时,没有找到任何符合要求的有效词汇——要么输入文本全是停用词/空内容,要么参数配置过滤掉了所有词汇。

排查步骤

  • 检查输入文本内容:将生成器转换为列表,查看实际文本输出,确认是否存在有效词汇
  • 核对CountVectorizer参数:默认配置会过滤英文停用词、仅保留长度≥2的词,若文本不符合该规则会导致词汇表为空
  • 验证tokenizer输出:确认sequences_to_texts_generator是否将序列正确转换为可读文本,而非空字符串或仅停用词

解决方案

1. 验证并清洗输入文本

先将生成器转为列表,检查文本内容:

train_texts = list(tokenizer.sequences_to_texts_generator(train_text_vec))
print(train_texts[:5])  # 打印前5条文本,确认是否有有效内容

若存在空文本或全停用词的条目,过滤后再训练:

# 过滤空文本
valid_indices = [i for i, text in enumerate(train_texts) if text.strip()]
train_texts_valid = [train_texts[i] for i in valid_indices]
y_train_valid = y_train.argmax(axis=1)[valid_indices]

2. 调整CountVectorizer参数

根据文本特性修改参数,避免过滤有效词汇:

# 示例1:关闭默认停用词过滤(适用于非英文文本)
CountVectorizer(stop_words=None)
# 示例2:允许单字/短词(修改token_pattern)
CountVectorizer(token_pattern=r'\b\w+\b', min_df=1)
# 示例3:自定义停用词列表
CountVectorizer(stop_words=["自定义", "停用词"])

3. 替换生成器为列表输入

Sklearn对生成器的兼容性不如列表,转成列表后训练更稳定:

text_clf = Pipeline([
    ('vect', CountVectorizer(stop_words=None, token_pattern=r'\b\w+\b')),
    ('tfidf', TfidfTransformer()),
    ('clf', RandomForestClassifier(class_weight='balanced', n_estimators=100))
])
text_clf.fit(train_texts_valid, y_train_valid)

4. 检查预处理流程

确认train_text_vec是否为有效序列数据:

  • 排查是否存在全零序列(对应空文本),这类数据需提前过滤
  • 检查tokenizer的配置,确保没有误删除所有非停用词

内容的提问来源于stack exchange,提问作者Rajat Das

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 12:10:48