重训练多分类模型时遇单一类别报错,如何解决?
问题:单样本重训练触发分类器类数不足错误
需求与现有代码
你要实现的功能:
- 加载已训练模型与未见过的数据集
- 用BERT编码器(SentenceTransformer('all-mpnet-base-v2'))+ 训练好的模型构建流水线
- 随机选数据预测,若结果错误则用该样本重训练模型
现有代码:
random_row = full_df.sample(1) X = random_row[['mmgid', 'body','preprocessedBody']] # select specific columns article = np.random.choice(X['preprocessedBody']) encoded_text = bert_model.encode([article]) prediction = multi_model.predict(encoded_text)[0] print(X['body']) print(f"Prediction: {prediction}") # 询问用户预测是否正确 response = input("Is this prediction correct? (yes/no)") # 若错误则获取正确分类 if response.lower() == 'no': correct_classification = input("What is the correct classification for this article?") # 用该样本重训练模型 multi_model.fit([encoded_text2], [correct_classification]) # 测试更新后的模型 updated_prediction = multi_model.predict([article]) # 显示新预测结果 print(f"The updated model predicts that this article belongs to the '{updated_prediction}' category.")
触发的错误
-> 1183 raise ValueError( 1184 "This solver needs samples of at least 2 classes" 1185 " in the data, but the data contains only one" 1186 " class: %r" 1187 % classes_[0] 1188 ) 1190 if len(self.classes_) == 2: 1191 n_classes = 1 ValueError: This solver needs samples of at least 2 classes in the data, but the data contains only one class: '2'
解决方法
1. 核心原因
调用fit()方法时只传入了单个样本,对应的类别只有一种,而大部分分类器(比如SVM、逻辑回归等)要求训练数据必须包含至少2个不同类别才能完成参数更新。
2. 具体修正步骤
(1)用partial_fit替代fit做增量更新
如果你的模型是scikit-learn的分类器(比如LogisticRegression、SGDClassifier等),可以用partial_fit实现单样本增量训练,需提前指定所有可能的类别:
# 先定义任务中所有可能的分类标签,比如['0','1','2','3'],根据实际情况修改 all_classes = ['0', '1', '2', '3'] # 重训练部分替换fit为partial_fit multi_model.partial_fit(encoded_text, [correct_classification], classes=all_classes)
partial_fit是为增量学习设计的,不需要每次训练都提供多类别数据,只要初始化时指定所有可能类别即可。
(2)修复代码中的变量错误
- 代码里的
encoded_text2未定义,应使用之前生成的encoded_text - 预测更新后的模型时,不能直接传原始文本
article,要先编码成向量再预测:
# 修正后的重训练与预测代码 if response.lower() == 'no': correct_classification = input("What is the correct classification for this article?") # 用partial_fit增量更新模型 multi_model.partial_fit(encoded_text, [correct_classification], classes=all_classes) # 重新编码后预测 updated_prediction = multi_model.predict(encoded_text)[0] print(f"The updated model predicts that this article belongs to the '{updated_prediction}' category.")
(3)备选方案:保留原有训练数据的少量样本
如果你的模型不支持partial_fit,可以把新样本和原有训练数据中的少量多类别样本混合后再调用fit(),但这种方法会增加计算量,且可能覆盖原有模型的部分特征。
内容的提问来源于stack exchange,提问作者Python-data
相关产品推荐
相关产品推荐

