能否通过循环分批训练SVM模型?文本情感分析训练内核崩溃
大样本SVM文本情感分析的问题解决
直接说核心结论:scikit-learn的SVC不支持增量训练,你现在循环里每次调用fit(),模型都会基于当前这20个样本重新初始化参数,之前批次的所有数据都会被完全忽略,最后训练出来的模型只学了最后一批样本,根本达不到用全量数据训练的效果。
原代码的两个关键问题
- 内核崩溃原因:你用BERT的preprocess直接输出
input_word_ids作为特征,每个样本是一串固定长度的token ID(比如128位),样本量一大,高维度特征直接把内存撑爆了,内核自然崩溃。 - 分批fit的错误逻辑:SVC的
fit()是全量训练模式,没有“累积学习”的逻辑,每次调用都会重置模型,相当于覆盖之前的训练结果。
三个可行的解决办法
办法1:用支持增量训练的SVM替代
scikit-learn的SGDClassifier可以设置loss='hinge'来模拟线性SVM的效果,它支持partial_fit()方法,能分批喂数据训练,不会覆盖之前的结果:
tfhub_handle_preprocess="https://tfhub.dev/tensorflow/bert_zh_preprocess/3" bert_preprocess_model = hub.KerasLayer(tfhub_handle_preprocess) def encoder(strlist): return bert_preprocess_model(strlist)["input_word_ids"] from sklearn.preprocessing import FunctionTransformer from sklearn.linear_model import SGDClassifier from sklearn.metrics import accuracy_score from IPython.display import clear_output import numpy as np # 构建流水线 SVM_Model = Pipeline([ ('Encoder', FunctionTransformer(encoder)), ('Svm', SGDClassifier(loss='hinge', random_state=42)) ]) # 注意:首次调用partial_fit必须先传入所有类别,不然模型不知道总共有多少类 all_classes = np.unique(labels) # 用第一个样本初始化模型的类别信息 SVM_Model.named_steps['Svm'].partial_fit(encoder(text[:1]), labels[:1], classes=all_classes) # 分批增量训练 batch_size = 20 for times in range(0, len(text), batch_size): batch_text = text[times:times+batch_size] batch_labels = labels[times:times+batch_size] # 用partial_fit代替fit,实现增量学习 SVM_Model.partial_fit(batch_text, batch_labels) # 验证当前效果 pre = SVM_Model.predict(pre_text[:100]) accuracy = accuracy_score(pre_labels[:100], pre) clear_output(wait=True) print(f'times:{times}, Accuracy:{accuracy}')
办法2:降低特征维度,减少内存占用
直接用input_word_ids当特征太浪费内存,换成BERT的输出embedding(比如[CLS]位置的向量),维度通常是768,比token ID序列紧凑得多,内存压力会大幅降低:
# 加载BERT编码器和预处理模型 tfhub_handle_encoder="https://tfhub.dev/tensorflow/bert_zh_L-12_H-768_A-12/4" tfhub_handle_preprocess="https://tfhub.dev/tensorflow/bert_zh_preprocess/3" bert_preprocess_model = hub.KerasLayer(tfhub_handle_preprocess) bert_model = hub.KerasLayer(tfhub_handle_encoder) def get_bert_embedding(strlist): preprocessed_inputs = bert_preprocess_model(strlist) outputs = bert_model(preprocessed_inputs) # 取[CLS]位置的向量作为整个文本的特征表示 return outputs['pooled_output'] from sklearn.preprocessing import FunctionTransformer from sklearn.svm import SVC SVM_Model = Pipeline([ ('Encoder', FunctionTransformer(get_bert_embedding)), ('Svm', SVC(kernel='linear')) ]) # 此时特征维度大幅降低,很多情况下可以直接全量训练 # 如果还是内存不够,就结合办法1的partial_fit方式分批训练 SVM_Model.fit(text, labels)
办法3:用核外SVM处理超大规模数据
如果数据集大到离谱,还可以用sklearn.svm.LinearSVC配合partial_fit,或者用支持分布式训练的工具,但LinearSVC同样需要先指定所有类别才能增量训练。
内容的提问来源于stack exchange,提问作者Kai
相关产品推荐
相关产品推荐

