You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

重训练多分类模型时遇单一类别报错,如何解决?

问题:单样本重训练触发分类器类数不足错误

需求与现有代码

你要实现的功能:

  • 加载已训练模型与未见过的数据集
  • 用BERT编码器(SentenceTransformer('all-mpnet-base-v2'))+ 训练好的模型构建流水线
  • 随机选数据预测,若结果错误则用该样本重训练模型

现有代码:

random_row = full_df.sample(1)
X = random_row[['mmgid', 'body','preprocessedBody']]  # select specific columns
article = np.random.choice(X['preprocessedBody'])
encoded_text = bert_model.encode([article])

prediction = multi_model.predict(encoded_text)[0]
print(X['body'])
print(f"Prediction: {prediction}")

# 询问用户预测是否正确
response = input("Is this prediction correct? (yes/no)")
# 若错误则获取正确分类
if response.lower() == 'no':
    correct_classification = input("What is the correct classification for this article?")
    
    # 用该样本重训练模型
    multi_model.fit([encoded_text2], [correct_classification])

    # 测试更新后的模型
    updated_prediction = multi_model.predict([article])

    # 显示新预测结果
    print(f"The updated model predicts that this article belongs to the '{updated_prediction}' category.")

触发的错误

-> 1183     raise ValueError(
   1184         "This solver needs samples of at least 2 classes"
   1185         " in the data, but the data contains only one"
   1186         " class: %r"
   1187         % classes_[0]
   1188     )
   1190 if len(self.classes_) == 2:
   1191     n_classes = 1

ValueError: This solver needs samples of at least 2 classes in the data, but the data contains only one class: '2'

解决方法

1. 核心原因

调用fit()方法时只传入了单个样本,对应的类别只有一种,而大部分分类器(比如SVM、逻辑回归等)要求训练数据必须包含至少2个不同类别才能完成参数更新。

2. 具体修正步骤

(1)用partial_fit替代fit做增量更新

如果你的模型是scikit-learn的分类器(比如LogisticRegression、SGDClassifier等),可以用partial_fit实现单样本增量训练,需提前指定所有可能的类别:

# 先定义任务中所有可能的分类标签,比如['0','1','2','3'],根据实际情况修改
all_classes = ['0', '1', '2', '3']

# 重训练部分替换fit为partial_fit
multi_model.partial_fit(encoded_text, [correct_classification], classes=all_classes)

partial_fit是为增量学习设计的,不需要每次训练都提供多类别数据,只要初始化时指定所有可能类别即可。

(2)修复代码中的变量错误

  • 代码里的encoded_text2未定义,应使用之前生成的encoded_text
  • 预测更新后的模型时,不能直接传原始文本article,要先编码成向量再预测:
# 修正后的重训练与预测代码
if response.lower() == 'no':
    correct_classification = input("What is the correct classification for this article?")
    
    # 用partial_fit增量更新模型
    multi_model.partial_fit(encoded_text, [correct_classification], classes=all_classes)

    # 重新编码后预测
    updated_prediction = multi_model.predict(encoded_text)[0]

    print(f"The updated model predicts that this article belongs to the '{updated_prediction}' category.")

(3)备选方案:保留原有训练数据的少量样本

如果你的模型不支持partial_fit,可以把新样本和原有训练数据中的少量多类别样本混合后再调用fit(),但这种方法会增加计算量,且可能覆盖原有模型的部分特征。

内容的提问来源于stack exchange,提问作者Python-data

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 16:35:17