You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

sklearn Pipeline执行cross_val_score时值解包不足错误排查

问题背景

参考在线文章修改编写了两段用于对比不同模型性能的代码,运行交叉验证逻辑时抛出值错误,排查是否为代码逻辑问题或输入数据结构问题。

第一段代码:预处理Pipeline构造函数

功能为定义create_pipe工具函数,通过ColumnTransformer配置分类列常量填充、数值列KNN缺失值填充、独热编码、特征归一化、方差阈值过滤等预处理逻辑,组装为包含预处理模块与分类器的Pipeline对象返回,原代码如下:

def create_pipe(clf):
    
        column_trans = ColumnTransformer(
            [
                ('SimpleImputer', SimpleImputer(strategy ='constant', fill_value = 'unanswered'), categorical_cols.columns),
                ('KNNImputer', KNNImputer(missing_values=np.nan, n_neighbors=17), numeric_cols.columns),
                ('enc', OneHotEncoder(sparse = False, drop=None), list(range(len(categorical_cols.columns)))),
                ('standard_scaler', MinMaxScaler()),
                ('variance', VarianceThreshold(threshold=(.95 * (1 - .95))))
            ],
            remainder='passthrough')
    
        pipeline = Pipeline([('prep',column_trans),
                             ('clf', clf)])
    
        return pipeline

第二段代码:模型交叉验证逻辑

首先定义包含随机森林、逻辑回归的模型字典,随后遍历字典为每个模型创建对应Pipeline实例,调用cross_val_score计算3折交叉验证的f1_macro指标得分,最终输出各模型得分的均值与标准差,原代码如下:

models = {'RFG' : RandomForestRegressor(random_state=42),
          'LogReg' : LogisticRegression(random_state=42)
          }

for name, model, in models.items():
    clf = model
    pipeline = create_pipe(clf)
    scores = cross_val_score(pipeline, 
                             X, 
                             y, 
                             scoring='f1_macro', 
                             cv=3, 
                             n_jobs=1, 
                             error_score='raise')
    print(name, ': Mean f1 Macro: %.3f and Standard Deviation: (%.3f)' % (np.mean(scores), np.std(scores)))

运行报错信息

ValueError: not enough values to unpack (expected 3, got 2)
问题排查结论

这个报错和输入数据结构完全无关,是两个非常明显的代码编写错误:

  • ColumnTransformer配置格式错误:ColumnTransformer要求传入的每个转换器配置必须是长度为3的元组,格式为(转换器名称, 转换器实例, 该转换器作用的列范围)。你在配置列表里写的('standard_scaler', MinMaxScaler())和('variance', VarianceThreshold(threshold=(.95 * (1 - .95))))两个元组只有2个元素,没有指定作用列,程序解析配置时尝试把每个元组解包为3个变量,直接触发了"expected 3, got 2"的错误。
  • 预处理步骤位置错误:ColumnTransformer的作用是对不同列并行执行不同的列转换操作,而MinMaxScaler归一化、VarianceThreshold方差过滤是需要作用在所有拼接完成的特征上的全局处理步骤,不应该放在ColumnTransformer的转换器列表里,应该放在Pipeline中,放在列转换模块之后、分类器之前。

另外还有两个隐藏问题:一是原代码中独热编码指定的作用列是list(range(len(categorical_cols.columns))),这个索引是相对于原始输入数据的,不会因为前面做了分类列填充就自动把前N列识别为分类列,会导致编码作用到错误的列上;二是模型字典里传入了RandomForestRegressor随机森林回归器,但使用的f1_macro是分类任务专属评分指标,回归模型无法输出符合要求的分类预测结果,修完前面的配置错误后这里也会报错,需要替换为RandomForestClassifier。

修正后的参考代码

def create_pipe(clf):
    # 列转换器只处理分列的差异化预处理,同列的多步操作封装为子Pipeline
    column_trans = ColumnTransformer(
        [
            ('cat_process', Pipeline([
                ('imputer', SimpleImputer(strategy ='constant', fill_value = 'unanswered')),
                ('enc', OneHotEncoder(sparse = False, drop=None))
            ]), categorical_cols.columns),
            ('num_process', KNNImputer(missing_values=np.nan, n_neighbors=17), numeric_cols.columns)
        ],
        remainder='passthrough')
    
    # 全局预处理步骤放在列转换之后,和分类器串成完整Pipeline
    pipeline = Pipeline([
        ('prep',column_trans),
        ('standard_scaler', MinMaxScaler()),
        ('variance', VarianceThreshold(threshold=(.95 * (1 - .95)))),
        ('clf', clf)
    ])

    return pipeline

# 回归器替换为分类器,去掉循环语句里多余的逗号
models = {'RFC' : RandomForestClassifier(random_state=42),
          'LogReg' : LogisticRegression(random_state=42)
          }

for name, model in models.items():
    pipeline = create_pipe(model)
    scores = cross_val_score(pipeline, 
                             X, 
                             y, 
                             scoring='f1_macro', 
                             cv=3, 
                             n_jobs=1, 
                             error_score='raise')
    print(f'{name}: Mean f1 Macro: {np.mean(scores):.3f} and Standard Deviation: ({np.std(scores):.3f})')

内容的提问来源于stack exchange,提问作者MarcinKamil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 09:00:52