You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自定义Scikit-learn分类器使用GridSearchCV报样本数不一致如何解决

错误原因

你自定义分类器的score方法实现错误,直接触发了样本数不匹配报错:

  • GridSearchCV默认采用K折交叉验证,每次迭代会拆分训练集(404样本)和验证集(102样本)
  • 你在fit方法中将训练集的离散化标签保存在了实例变量self.yt中,而score方法错误地直接用self.yt(训练集404个标签)和验证集预测结果(102个)计算准确率,样本数完全不匹配。

修复后的代码

只需要调整score方法逻辑,同时移除无用的self.yt缓存即可:

  1. 必须使用score方法传入的验证集真实标签y,而非训练时缓存的标签
  2. 对传入的y用已经拟合好的self.discr做离散化转换(不能重新fit,避免数据泄露)
  3. 移除y的默认值None,保证计算准确率时必须传入真实标签
from sklearn.base import BaseEstimator, ClassifierMixin
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import KBinsDiscretizer
from sklearn.metrics import accuracy_score

class OwnClassifier(BaseEstimator, ClassifierMixin):
    def __init__(self, estimator=None):
        if estimator is None:
            estimator = LogisticRegression(solver='liblinear')
        self.estimator = estimator
        self.discr = KBinsDiscretizer(n_bins=4, encode='ordinal')
        
    def fit(self, X, y):
        # 训练阶段对训练集标签做离散化拟合+转换
        yt = self.discr.fit_transform(y.reshape(-1, 1)).astype(int)
        self.estimator.fit(X, yt.ravel())
        return self
    
    def predict(self, X):
        return self.estimator.predict(X)
    
    def predict_proba(self, X):
        return self.estimator.predict_proba(X)
    
    def score(self, X, y):
        # 验证阶段用已经拟合好的离散器转换验证集标签,再计算准确率
        yt = self.discr.transform(y.reshape(-1, 1)).astype(int)
        y_pred = self.predict(X)
        return accuracy_score(yt.ravel(), y_pred)

内容的提问来源于stack exchange,提问作者user15568232

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 04:18:00