You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何查看scikit-learn CountVectorizer拟合进度及剩余时间?

解决方案

scikit-learn官方的CountVectorizer和TfidfVectorizer没有内置进度展示功能,你可以通过以下两种方法实现tqdm进度条效果:

方法1:自定义带进度条的Vectorizer子类(最无缝)

重写CountVectorizer的_count_vocab方法,在遍历文档时加入tqdm进度条,TfidfVectorizer可以直接复用该逻辑:

import numpy as np
import scipy.sparse as sp
from collections import defaultdict
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer, _make_int_array
from tqdm import tqdm

class CountVectorizerWithProgress(CountVectorizer):
    def _count_vocab(self, raw_documents, fixed_vocab):
        if fixed_vocab:
            vocabulary = self.vocabulary_
        else:
            vocabulary = defaultdict()
            vocabulary.default_factory = vocabulary.__len__

        analyze = self.build_analyzer()
        j_indices = []
        indptr = [0]
        values = _make_int_array()

        # 仅新增tqdm包装逻辑,其余和官方源码一致
        for doc in tqdm(raw_documents, desc="CountVectorizer拟合进度"):
            feature_counter = {}
            for feature in analyze(doc):
                try:
                    feature_idx = vocabulary[feature]
                    if feature_idx not in feature_counter:
                        feature_counter[feature_idx] = 1
                    else:
                        feature_counter[feature_idx] += 1
                except KeyError:
                    continue

            j_indices.extend(feature_counter.keys())
            values.extend(feature_counter.values())
            indptr.append(len(j_indices))

        j_indices = np.asarray(j_indices, dtype=np.int64)
        indptr = np.asarray(indptr, dtype=np.int64)
        values = np.frombuffer(values, dtype=np.intc)

        X = sp.csr_matrix((values, j_indices, indptr), shape=(len(indptr) - 1, len(vocabulary)), dtype=self.dtype)
        X.sort_indices()
        return vocabulary, X

# 带进度的TfidfVectorizer直接多继承即可
class TfidfVectorizerWithProgress(TfidfVectorizer, CountVectorizerWithProgress):
    pass

使用方式和官方类完全一致:

cvectorizer = CountVectorizerWithProgress(min_df=10, ngram_range=(1,4), max_features=5000)
tvectorizer = TfidfVectorizerWithProgress(min_df=10, ngram_range=(1,4), max_features=5000)
    
cvectorizer.fit(X_train['essay'].values) 
tvectorizer.fit(X_train['essay'].values)

方法2:手动统计词汇表(兼容性更高,无需修改类逻辑)

提前遍历所有文本统计符合要求的词汇,再把生成的词汇表传给Vectorizer,跳过自动统计步骤,遍历过程中可直接加tqdm进度条:

from collections import defaultdict
from tqdm import tqdm

# 构建和你参数匹配的分析器
analyzer = CountVectorizer(min_df=10, ngram_range=(1,4)).build_analyzer()
vocab_count = defaultdict(int)

# 带进度条统计所有词汇的出现频次
for doc in tqdm(X_train['essay'].values, desc="统计词汇进度"):
    for feature in analyzer(doc):
        vocab_count[feature] += 1

# 按照你的参数过滤词汇
min_df = 10
max_features = 5000
filtered_vocab = [k for k, v in sorted(vocab_count.items(), key=lambda x: -v) if v >= min_df][:max_features]
vocab_dict = {k:i for i,k in enumerate(filtered_vocab)}

# 直接传入预生成的词汇表,fit时不会再重复统计
cvectorizer = CountVectorizer(vocabulary=vocab_dict)
tvectorizer = TfidfVectorizer(vocabulary=vocab_dict)

cvectorizer.fit(X_train['essay'].values)
tvectorizer.fit(X_train['essay'].values)

卡顿优化建议

  • 你当前设置的ngram_range=(1,4)会生成大量1-4词的组合,计算量随ngram上限指数增长,若非必要可以调低到(1,2)或(1,3),速度会提升数倍
  • 可以先拿10%的样本跑一遍,预估完整数据集的运行时间,避免无意义的等待
  • 确保你的文本已经做了基础预处理(比如去特殊符号、统一大小写等),减少无效的ngram统计

内容的提问来源于stack exchange,提问作者Anirudh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 17:24:04