Sklearn中Hard Voting Classifier平局处理逻辑及自定义规则咨询
sklearn VotingClassifier硬投票平局处理与自定义规则说明
平局默认处理逻辑
硬投票时,VotingClassifier遇到平局(多个类别得票相同且为最高),并非随机选择,而是按照类别标签的升序顺序选取第一个出现最高得票的类别。
比如你提到的场景:1票class 0,4票class 1,4票class 2,此时class 1和class 2得票并列最高,会优先选择标签数值更小的class 1。你之前觉得结果随机,大概率是模型训练过程中的随机性(比如不同训练轮次的模型预测结果有差异)导致的,投票本身的逻辑是确定的。
能否自定义平局时指定class 4?
不行。因为你的任务是三分类,训练数据的类别标签只有0、1、2,VotingClassifier的预测输出只能是训练时见过的类别,无法输出不存在的class 4。
如果需要自定义平局时的选择规则(比如优先选class 2),可以自己实现投票逻辑,或者重写VotingClassifier的预测逻辑,示例如下:
自定义硬投票函数示例
import numpy as np from sklearn.base import BaseEstimator, ClassifierMixin class CustomVotingClassifier(BaseEstimator, ClassifierMixin): def __init__(self, estimators, tie_break_class=2): self.estimators = estimators self.tie_break_class = tie_break_class def fit(self, X, y): # 训练所有子模型 for name, clf in self.estimators: clf.fit(X, y) self.classes_ = np.unique(y) return self def predict(self, X): # 收集所有模型的预测结果 predictions = np.array([clf.predict(X) for name, clf in self.estimators]).T # 统计每个样本的得票 votes = np.array([np.bincount(pred, minlength=len(self.classes_)) for pred in predictions]) # 找到最高得票数 max_votes = votes.max(axis=1) # 处理平局:如果有多个类别得票等于max_votes,指定tie_break_class result = [] for idx in range(len(X)): top_classes = np.where(votes[idx] == max_votes[idx])[0] if len(top_classes) > 1: result.append(self.tie_break_class) else: result.append(top_classes[0]) return np.array(result)
使用时直接实例化这个自定义类,传入子模型和平局时要选的类别即可:
# 假设你已经有9个训练好的模型estimators列表 custom_voter = CustomVotingClassifier(estimators, tie_break_class=2) custom_voter.fit(X_train, y_train) predictions = custom_voter.predict(X_test)
内容的提问来源于stack exchange,提问作者Ivo
相关产品推荐
相关产品推荐

