You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

手动实现线性分类器遇KeyError:0报错,求技术解决方案

手动实现线性分类器时出现KeyError:0的问题排查与解决

问题背景

我是Python和机器学习新手,为理解底层原理,尝试不调用API手动编写线性分类器代码,但运行时触发了KeyError:0错误。

实现代码

import numpy as np 
import matplotlib.pyplot as plt 
import pandas as pd 
from sklearn.model_selection import train_test_split

data = pd.read_csv('files/weather.csv', parse_dates= True, index_col=0)
data.head()

data.isnull().sum()

dataset = data[['Humidity3pm','Pressure3pm','RainTomorrow']].dropna()

X = dataset[['Humidity3pm', 'Pressure3pm']]
y = dataset['RainTomorrow']
y = np.array([0 if value == 'No' else 1 for value in y])

X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)


def linear_classifier(X, y, learning_rate=0.01, num_epochs=100):
    num_features = X.shape[1] 
    weights = np.zeros(num_features)
    bias = 0
    
    for epoch in range(num_epochs):
        for i in range(X.shape[0]):
            linear_output = np.dot(X[i], weights) + bias
           
            y_pred = np.sign(linear_output)
            
           
            error = y[i] - y_pred
            # print("The value of error=", error)
            weights = weights + learning_rate * error * X[i]

            bias += learning_rate * error
            
    return weights, bias


weights, bias = linear_classifier(X_train, y_train)   ## This lines gives the error

报错栈

Output exceeds the size limit. Open the full output data in a text editor
---------------------------------------------------------------------------
KeyError                                  Traceback (most recent call last)
File c:\Users\manuc\anaconda3\envs\learning_python\lib\site-packages\pandas\core\indexes\base.py:3802, in Index.get_loc(self, key, method, tolerance)
   3801 try:
-> 3802     return self._engine.get_loc(casted_key)
   3803 except KeyError as err:

File c:\Users\manuc\anaconda3\envs\learning_python\lib\site-packages\pandas\_libs\index.pyx:138, in pandas._libs.index.IndexEngine.get_loc()

File c:\Users\manuc\anaconda3\envs\learning_python\lib\site-packages\pandas\_libs\index.pyx:165, in pandas._libs.index.IndexEngine.get_loc()

File pandas\_libs\hashtable_class_helper.pxi:5745, in pandas._libs.hashtable.PyObjectHashTable.get_item()

File pandas\_libs\hashtable_class_helper.pxi:5753, in pandas._libs.hashtable.PyObjectHashTable.get_item()

KeyError: 0

The above exception was the direct cause of the following exception:

KeyError                                  Traceback (most recent call last)
Cell In[170], line 1
----> 1 weights, bias = linear_classifier(X_train, y_train) 

Cell In[169], line 10, in linear_classifier(X, y, learning_rate, num_epochs)
      8 for epoch in range(num_epochs):
...
   3807     #  InvalidIndexError. Otherwise we fall through and re-raise
   3808     #  the TypeError.
   3809     self._check_indexing_error(key)

KeyError: 0

问题分析与解决

错误原因

X_train是Pandas DataFrame对象,使用X[i]索引时,Pandas会将i当作列名查找,而非行索引。循环到i=0时,DataFrame中没有名为0的列,因此抛出KeyError。而y_train是NumPy数组,用y[i]索引行是正常的。

另外存在潜在问题:标签y取值为0/1,但np.sign(linear_output)返回-1/0/1,会导致误差计算逻辑不一致(比如y[i]为0,y_pred为-1时,error=1),需要统一取值范围。

解决步骤

  1. 将DataFrame转换为NumPy数组:确保可用整数索引访问行数据,有两种方式:
    • 方式一:定义X时直接转换:X = dataset[['Humidity3pm', 'Pressure3pm']].values
    • 方式二:分割数据集后转换:X_train = X_train.values、X_test = X_test.values
  2. 统一预测值取值范围:将np.sign(linear_output)替换为1 if linear_output >=0 else 0,让预测值和标签的0/1取值一致。

修改后的完整代码

import numpy as np 
import matplotlib.pyplot as plt 
import pandas as pd 
from sklearn.model_selection import train_test_split

data = pd.read_csv('files/weather.csv', parse_dates= True, index_col=0)
data.head()

data.isnull().sum()

dataset = data[['Humidity3pm','Pressure3pm','RainTomorrow']].dropna()

# 将X转换为NumPy数组
X = dataset[['Humidity3pm', 'Pressure3pm']].values
y = dataset['RainTomorrow']
y = np.array([0 if value == 'No' else 1 for value in y])

X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)


def linear_classifier(X, y, learning_rate=0.01, num_epochs=100):
    num_features = X.shape[1] 
    weights = np.zeros(num_features)
    bias = 0
    
    for epoch in range(num_epochs):
        for i in range(X.shape[0]):
            linear_output = np.dot(X[i], weights) + bias
           
            # 统一预测值为0/1
            y_pred = 1 if linear_output >= 0 else 0
            
            error = y[i] - y_pred
            # print("The value of error=", error)
            weights = weights + learning_rate * error * X[i]

            bias += learning_rate * error
            
    return weights, bias


weights, bias = linear_classifier(X_train, y_train)

内容的提问来源于stack exchange,提问作者sarika

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 16:44:59