You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

sklearn线性回归训练报错:输入变量样本数不一致求助

解决sklearn train_test_split导致的样本数不匹配ValueError

问题场景

编写预测logS值的线性回归代码时,运行触发ValueError: Found input variables with inconsistent numbers of samples,调整test_size参数后报错样本数同步变化,但问题始终存在。

原代码

import pandas as pd
from sklearn.metrics import mean_squared_error, r2_score
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression

df=pd.read_csv('https://raw.githubusercontent.com/dataprofessor/data/master/delaney_solubility_with_descriptors.csv')

X = df.drop('logS', axis=1)
y = df['logS']

X_train, y_train, X_test, y_test = train_test_split(X, y, test_size=0.2, random_state=1)

lr = LinearRegression()
lr.fit(X_train, y_train)

报错信息

Traceback (most recent call last):
  File "D:\Python\AI\test\main.py", line 14, in <module>
    lr.fit(X_train, y_train)
  File "D:\Python\AI\test\venv\lib\site-packages\sklearn\base.py", line 1152, in wrapper
    return fit_method(estimator, *args, **kwargs)
  File "D:\Python\AI\test\venv\lib\site-packages\sklearn\linear_model\_base.py", line 678, in fit
    X, y = self._validate_data(
  File "D:\Python\AI\test\venv\lib\site-packages\sklearn\base.py", line 622, in _validate_data
    X, y = check_X_y(X, y, **check_params)
  File "D:\Python\AI\test\venv\lib\site-packages\sklearn\utils\validation.py", line 1164, in check_X_y
    check_consistent_length(X, y)
  File "D:\Python\AI\test\venv\lib\site-packages\sklearn\utils\validation.py", line 407, in check_consistent_length
    raise ValueError(
ValueError: Found input variables with inconsistent numbers of samples: [915, 229]

问题原因

sklearn.model_selection.train_test_split函数的返回值顺序是固定的:训练集特征(X_train)、测试集特征(X_test)、训练集标签(y_train)、测试集标签(y_test)。

你错误地将返回值顺序写成了X_train, y_train, X_test, y_test,导致:

  • X_train是训练集特征(样本数为总样本的80%,即915条)
  • y_train被错误赋值为测试集特征(样本数为总样本的20%,即229条)
    两者样本数不匹配,因此调用lr.fit()时触发报错。

解决方法

修正train_test_split的变量接收顺序,改为正确的顺序即可:

修正后代码

import pandas as pd
from sklearn.metrics import mean_squared_error, r2_score
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression

df=pd.read_csv('https://raw.githubusercontent.com/dataprofessor/data/master/delaney_solubility_with_descriptors.csv')

X = df.drop('logS', axis=1)
y = df['logS']

# 修正变量顺序:X_train, X_test, y_train, y_test
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=1)

lr = LinearRegression()
lr.fit(X_train, y_train)

# 可选:验证模型效果
y_pred = lr.predict(X_test)
print(f"R² Score: {r2_score(y_test, y_pred)}")
print(f"MSE: {mean_squared_error(y_test, y_pred)}")

内容的提问来源于stack exchange,提问作者ketchup

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 20:10:23