You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用linear_model.LogisticRegression报错,是否与X、y数据类型有关?

LogisticRegression训练报错排查(数据类型相关)

使用linear_model.LogisticRegression训练模型时出现错误,怀疑问题出在X和y的数据类型上。当前环境为Python 3.8 64位,Scikit-learn版本1.2.0,相关代码如下:

import numpy as np
import matplotlib.pyplot as plt
from sklearn import linear_model

# datat generation

m = 100
w0 = -6
w = np.array([[2], [1]])
X = np.hstack([4*np.random.rand(m,1), 4*np.random.rand(m,1)])

w = np.asmatrix(w)
X = np.asmatrix(X)

y = 1/(1 + np.exp(-w0-X*w)) > 0.5 

C1 = np.where(y == True)[0]
C0 = np.where(y == False)[0]

y = np.empty([m,1])
y[C1] = 1
y[C0] = 0

print(X.shape)
print(y.shape)


from sklearn import linear_model

clf = linear_model.LogisticRegression(solver = 'lbfgs')
clf.fit(X, np.ravel(y))

问题分析

你的代码核心问题在于使用了np.matrix类型,而Scikit-learn模型对numpy数组(np.ndarray)的兼容性更好,矩阵类型的特殊行为可能引发拟合错误:

  • np.matrix的乘法、维度处理逻辑和数组存在差异,模型内部可能无法正确解析该类型输入
  • 虽然np.ravel(y)把标签转换成了一维数组,但输入特征X的矩阵类型仍可能干扰模型拟合流程

修复方案

移除np.asmatrix的转换,全程使用numpy数组,并改用现代numpy的矩阵乘法运算符@:

import numpy as np
import matplotlib.pyplot as plt
from sklearn import linear_model

# 数据生成
m = 100
w0 = -6
w = np.array([[2], [1]])
X = np.hstack([4*np.random.rand(m,1), 4*np.random.rand(m,1)])

# 直接使用数组运算,无需转换为矩阵
y_prob = 1/(1 + np.exp(-w0 - X @ w))
y = y_prob > 0.5 

C1 = np.where(y == True)[0]
C0 = np.where(y == False)[0]

y = np.empty([m,1])
y[C1] = 1
y[C0] = 0

print(X.shape)
print(y.shape)

clf = linear_model.LogisticRegression(solver='lbfgs')
clf.fit(X, np.ravel(y))

补充说明

  • Scikit-learn官方文档优先推荐使用numpy数组作为输入,矩阵类型仅为兼容旧代码保留,日常使用中数组更不易出错
  • X @ w是numpy中统一的矩阵乘法写法,对数组和矩阵都适用,但在数组上的行为更直观可控

内容的提问来源于stack exchange,提问作者chinchin97

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 20:40:55