You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

拟合模型时样本数不一致报错:已确认数据形状匹配仍存疑问

问题分析与解决方案

兄弟,你这问题的根源太典型了——你把train_test_split的返回值顺序搞反了!

先看你写的这行代码:

X_train, y_train, X_test, y_test = train_test_split(X_sliced, y_sliced)

sklearn里的train_test_split默认返回顺序是:X_train, X_test, y_train, y_test。你把y_train和X_test的位置弄混了,直接导致:

  • X_train确实是拆分后的训练集特征(108797个样本)
  • 但你赋值给y_train的其实是拆分后的测试集特征(36266个样本)
  • 后面的X_test和y_test也全乱了套

这就难怪拟合模型时会报样本数不一致的错误——你的X_train有108797个样本,而你当成y_train的变量只有36266个样本,两者完全不匹配。

修正后的代码

把train_test_split的返回顺序调整正确就行:

X_train, X_test, y_train, y_test = train_test_split(X_sliced, y_sliced)
linreg.fit(X_train, y_train)

额外提个小建议:以后用这类拆分函数时,要么记准返回顺序,要么拆分后打印每个变量的形状校验一下,能避免很多这类低级失误~

内容的提问来源于stack exchange,提问作者HTTP 418

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 08:12:37