You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多输出预测场景下GridSearchCV未遍历全部交叉验证折的异常问题求助

解决GridSearchCV多输出回归中交叉验证评分异常问题

看起来你的问题出在多进程下自定义评分函数的序列化问题,以及评分函数中不必要的DataFrame转换导致的潜在问题。让我们一步步分析并解决:

问题根源分析

  1. 多进程(n_jobs=-1)导致的自定义异常类序列化失败:你定义了LengthNotEqual自定义异常,但在多进程环境中,这类自定义类如果没有正确序列化,子进程中抛出异常时会被静默处理,导致评分函数返回0(GridSearchCV在捕获异常时可能默认返回0分)。
  2. 不必要的DataFrame转换:将y_true和y_predicted转为DataFrame不仅降低效率,还可能在多进程中引发序列化问题,尤其是当交叉验证拆分后的标签是numpy数组时,转换过程可能隐含未知问题。
  3. 潜在的除以0风险:你的smape函数中没有处理分母为0的情况,虽然你的数据中暂时没有,但极端情况下会导致计算错误。

解决方案

1. 修改评分函数,移除自定义异常并使用numpy数组处理

将评分函数改为纯numpy操作,避免DataFrame转换,同时替换自定义异常为标准的ValueError,并处理分母为0的情况:

from sklearn.metrics import make_scorer
import numpy as np

def smape(observations, predictions):
    observations = np.asarray(observations)
    predictions = np.asarray(predictions)
    
    if len(observations) != len(predictions):
        raise ValueError("The number of observations and predictions don't match!")
    
    N = len(observations)
    denominator = (np.abs(observations) + np.abs(predictions)) / 2
    # 处理分母为0的情况,避免除以0错误
    denominator[denominator == 0] = 1
    
    return 100 / N * np.sum(np.abs(observations - predictions) / denominator)

def final_smape(y_true, y_predicted):
    y_true = np.asarray(y_true)
    y_predicted = np.asarray(y_predicted)
    
    smape_a = smape(y_true[:, 0], y_predicted[:, 0])
    smape_b = smape(y_true[:, 1], y_predicted[:, 1])
    
    return 0.25 * smape_a + 0.75 * smape_b

# 保持scorer定义不变
scorer = make_scorer(final_smape, greater_is_better=False)

2. 先禁用多进程验证问题

暂时将GridSearchCV的n_jobs设为1,验证是否是多进程导致的问题:

model = GridSearchCV(estimator=m, param_grid={}, scoring=scorer, n_jobs=1)
model.fit(X_train, y_train)

如果此时所有交叉验证折的评分都正常,说明之前的多进程序列化问题是核心原因。之后如果需要多进程,可以确保评分函数中没有自定义类,仅使用numpy和sklearn的内置组件(这些都支持序列化)。

3. 验证交叉验证结果

运行修改后的代码后,查看model.cv_results_中的各个split*_test_score,应该所有折都会有有效评分,mean-test-score也会是真实的5折均值,此时best_score_和测试集分数的差距会回归正常(因为DummyRegressor的性能本来就接近训练集和测试集的smape)。

验证效果

修改后,你应该会看到类似这样的输出:

Pipeline(steps=[('std', StandardScaler()), ('dummy', DummyRegressor())])
best score: 12.3  # 接近测试集分数
test score: 12.293068464422076

这样就解决了只有第一个折有有效评分的问题。

内容的提问来源于stack exchange,提问作者Yosef Cohen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 21:12:40