如何查看scikit-learn CountVectorizer拟合进度及剩余时间?
解决方案
scikit-learn官方的CountVectorizer和TfidfVectorizer没有内置进度展示功能,你可以通过以下两种方法实现tqdm进度条效果:
方法1:自定义带进度条的Vectorizer子类(最无缝)
重写CountVectorizer的_count_vocab方法,在遍历文档时加入tqdm进度条,TfidfVectorizer可以直接复用该逻辑:
import numpy as np import scipy.sparse as sp from collections import defaultdict from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer, _make_int_array from tqdm import tqdm class CountVectorizerWithProgress(CountVectorizer): def _count_vocab(self, raw_documents, fixed_vocab): if fixed_vocab: vocabulary = self.vocabulary_ else: vocabulary = defaultdict() vocabulary.default_factory = vocabulary.__len__ analyze = self.build_analyzer() j_indices = [] indptr = [0] values = _make_int_array() # 仅新增tqdm包装逻辑,其余和官方源码一致 for doc in tqdm(raw_documents, desc="CountVectorizer拟合进度"): feature_counter = {} for feature in analyze(doc): try: feature_idx = vocabulary[feature] if feature_idx not in feature_counter: feature_counter[feature_idx] = 1 else: feature_counter[feature_idx] += 1 except KeyError: continue j_indices.extend(feature_counter.keys()) values.extend(feature_counter.values()) indptr.append(len(j_indices)) j_indices = np.asarray(j_indices, dtype=np.int64) indptr = np.asarray(indptr, dtype=np.int64) values = np.frombuffer(values, dtype=np.intc) X = sp.csr_matrix((values, j_indices, indptr), shape=(len(indptr) - 1, len(vocabulary)), dtype=self.dtype) X.sort_indices() return vocabulary, X # 带进度的TfidfVectorizer直接多继承即可 class TfidfVectorizerWithProgress(TfidfVectorizer, CountVectorizerWithProgress): pass
使用方式和官方类完全一致:
cvectorizer = CountVectorizerWithProgress(min_df=10, ngram_range=(1,4), max_features=5000) tvectorizer = TfidfVectorizerWithProgress(min_df=10, ngram_range=(1,4), max_features=5000) cvectorizer.fit(X_train['essay'].values) tvectorizer.fit(X_train['essay'].values)
方法2:手动统计词汇表(兼容性更高,无需修改类逻辑)
提前遍历所有文本统计符合要求的词汇,再把生成的词汇表传给Vectorizer,跳过自动统计步骤,遍历过程中可直接加tqdm进度条:
from collections import defaultdict from tqdm import tqdm # 构建和你参数匹配的分析器 analyzer = CountVectorizer(min_df=10, ngram_range=(1,4)).build_analyzer() vocab_count = defaultdict(int) # 带进度条统计所有词汇的出现频次 for doc in tqdm(X_train['essay'].values, desc="统计词汇进度"): for feature in analyzer(doc): vocab_count[feature] += 1 # 按照你的参数过滤词汇 min_df = 10 max_features = 5000 filtered_vocab = [k for k, v in sorted(vocab_count.items(), key=lambda x: -v) if v >= min_df][:max_features] vocab_dict = {k:i for i,k in enumerate(filtered_vocab)} # 直接传入预生成的词汇表,fit时不会再重复统计 cvectorizer = CountVectorizer(vocabulary=vocab_dict) tvectorizer = TfidfVectorizer(vocabulary=vocab_dict) cvectorizer.fit(X_train['essay'].values) tvectorizer.fit(X_train['essay'].values)
卡顿优化建议
- 你当前设置的
ngram_range=(1,4)会生成大量1-4词的组合,计算量随ngram上限指数增长,若非必要可以调低到(1,2)或(1,3),速度会提升数倍 - 可以先拿10%的样本跑一遍,预估完整数据集的运行时间,避免无意义的等待
- 确保你的文本已经做了基础预处理(比如去特殊符号、统一大小写等),减少无效的ngram统计
内容的提问来源于stack exchange,提问作者Anirudh
相关产品推荐
相关产品推荐

