You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas构建ML预测模型:X、y列划分两种方式的差异及报错问题

两种特征/标签划分代码的差异及KeyError原因解析

一、两组代码的核心差异

第一组代码

X = np.array(df.drop([predict], 1))
y = np.array(df[predict])
  • 通过np.array()将pandas的DataFrame/Series强制转换为NumPy数组
  • df.drop([predict],1)中的1是旧版pandas的列删除参数(新版推荐用axis=1替代)

第二组代码

X= data.drop(columns=[predict])
y=data[predict]
  • 直接保留pandas原生数据类型:X是DataFrame,y是Series,未做数组转换

二、KeyError=0的原因

循环代码print(predictions[x], x_test[x], y_test[x])的报错完全由数据类型差异导致:

  • 用第一组代码时,x_test是NumPy数组,数组支持整数位置索引,x_test[x]会直接取第x行的特征数据,运行正常。
  • 用第二组代码时,x_test是pandas DataFrame,此时x_test[x]会被pandas解析为按列名索引,而你的DataFrame中不存在名为0的列,因此抛出KeyError=0。

解决方法

方法1:统一转换为NumPy数组(和第一组代码对齐)

X= np.array(data.drop(columns=[predict]))
y=np.array(data[predict])

方法2:用pandas的位置索引方法访问行

修改循环中的索引方式,使用.iloc指定位置索引:

for x in range(len(predictions)):
    print(predictions[x], x_test.iloc[x], y_test.iloc[x])

内容的提问来源于stack exchange,提问作者KKG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 03:42:38