You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scikit-learn逻辑回归训练报错:y应为一维数组,获(56000,10)形状数组

问题描述

代码片段

import numpy as np 
import pandas as pd 
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.linear_model import LogisticRegression
logistic_regression = LogisticRegression(random_state = 10)
logistic_regression.fit(X_train, y_train)
y_pred_logistic_regression = logistic_regression.predict(X_test)
print(y_pred_logistic_regression.shape)

报错信息

ValueError: y should be a 1d array, got an array of shape (56000, 10) instead

使用Scikit-learn的LogisticRegression训练模型时触发上述错误,传入的y_train形状为(56000,10),需解决该问题。

解决方法

你的y_train是one-hot编码格式的标签(10列对应10个类别),但Scikit-learn的LogisticRegression要求目标变量y是一维的类别索引数组(比如每个样本对应0-9的整数标签),可通过以下方式处理:

  • 方法一:将one-hot标签转换为类别索引
    用numpy的argmax函数提取每一行最大值的索引,该索引即为对应类别:

    # 转换训练集标签
    y_train = np.argmax(y_train, axis=1)
    # 若测试集标签也是one-hot格式,同步转换
    y_test = np.argmax(y_test, axis=1)
    
    # 重新训练模型
    logistic_regression = LogisticRegression(random_state=10)
    logistic_regression.fit(X_train, y_train)
    
  • 方法二:通过标签逆转换适配模型
    若坚持使用one-hot标签,可借助LabelBinarizer逆转换为类别索引,同时指定多分类模式:

    from sklearn.preprocessing import LabelBinarizer
    
    lb = LabelBinarizer()
    y_train = lb.inverse_transform(y_train)
    y_test = lb.inverse_transform(y_test)
    
    logistic_regression = LogisticRegression(random_state=10, multi_class="multinomial")
    logistic_regression.fit(X_train, y_train)
    

补充说明

  • LogisticRegression默认multi_class="auto"会根据标签自动判断,但仅支持一维标签,one-hot格式必须手动转换为类别索引,否则会触发维度错误。
  • 若为多分类任务,转换为类别索引后,模型默认采用OvR(一对多)策略,设置multi_class="multinomial"则会使用Softmax回归。

内容的提问来源于stack exchange,提问作者Shyrush Shrestha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 13:01:41