手动实现线性分类器遇KeyError:0报错,求技术解决方案
手动实现线性分类器时出现KeyError:0的问题排查与解决
问题背景
我是Python和机器学习新手,为理解底层原理,尝试不调用API手动编写线性分类器代码,但运行时触发了KeyError:0错误。
实现代码
import numpy as np import matplotlib.pyplot as plt import pandas as pd from sklearn.model_selection import train_test_split data = pd.read_csv('files/weather.csv', parse_dates= True, index_col=0) data.head() data.isnull().sum() dataset = data[['Humidity3pm','Pressure3pm','RainTomorrow']].dropna() X = dataset[['Humidity3pm', 'Pressure3pm']] y = dataset['RainTomorrow'] y = np.array([0 if value == 'No' else 1 for value in y]) X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0) def linear_classifier(X, y, learning_rate=0.01, num_epochs=100): num_features = X.shape[1] weights = np.zeros(num_features) bias = 0 for epoch in range(num_epochs): for i in range(X.shape[0]): linear_output = np.dot(X[i], weights) + bias y_pred = np.sign(linear_output) error = y[i] - y_pred # print("The value of error=", error) weights = weights + learning_rate * error * X[i] bias += learning_rate * error return weights, bias weights, bias = linear_classifier(X_train, y_train) ## This lines gives the error
报错栈
Output exceeds the size limit. Open the full output data in a text editor --------------------------------------------------------------------------- KeyError Traceback (most recent call last) File c:\Users\manuc\anaconda3\envs\learning_python\lib\site-packages\pandas\core\indexes\base.py:3802, in Index.get_loc(self, key, method, tolerance) 3801 try: -> 3802 return self._engine.get_loc(casted_key) 3803 except KeyError as err: File c:\Users\manuc\anaconda3\envs\learning_python\lib\site-packages\pandas\_libs\index.pyx:138, in pandas._libs.index.IndexEngine.get_loc() File c:\Users\manuc\anaconda3\envs\learning_python\lib\site-packages\pandas\_libs\index.pyx:165, in pandas._libs.index.IndexEngine.get_loc() File pandas\_libs\hashtable_class_helper.pxi:5745, in pandas._libs.hashtable.PyObjectHashTable.get_item() File pandas\_libs\hashtable_class_helper.pxi:5753, in pandas._libs.hashtable.PyObjectHashTable.get_item() KeyError: 0 The above exception was the direct cause of the following exception: KeyError Traceback (most recent call last) Cell In[170], line 1 ----> 1 weights, bias = linear_classifier(X_train, y_train) Cell In[169], line 10, in linear_classifier(X, y, learning_rate, num_epochs) 8 for epoch in range(num_epochs): ... 3807 # InvalidIndexError. Otherwise we fall through and re-raise 3808 # the TypeError. 3809 self._check_indexing_error(key) KeyError: 0
问题分析与解决
错误原因
X_train是Pandas DataFrame对象,使用X[i]索引时,Pandas会将i当作列名查找,而非行索引。循环到i=0时,DataFrame中没有名为0的列,因此抛出KeyError。而y_train是NumPy数组,用y[i]索引行是正常的。
另外存在潜在问题:标签y取值为0/1,但np.sign(linear_output)返回-1/0/1,会导致误差计算逻辑不一致(比如y[i]为0,y_pred为-1时,error=1),需要统一取值范围。
解决步骤
- 将DataFrame转换为NumPy数组:确保可用整数索引访问行数据,有两种方式:
- 方式一:定义X时直接转换:
X = dataset[['Humidity3pm', 'Pressure3pm']].values - 方式二:分割数据集后转换:
X_train = X_train.values、X_test = X_test.values
- 方式一:定义X时直接转换:
- 统一预测值取值范围:将
np.sign(linear_output)替换为1 if linear_output >=0 else 0,让预测值和标签的0/1取值一致。
修改后的完整代码
import numpy as np import matplotlib.pyplot as plt import pandas as pd from sklearn.model_selection import train_test_split data = pd.read_csv('files/weather.csv', parse_dates= True, index_col=0) data.head() data.isnull().sum() dataset = data[['Humidity3pm','Pressure3pm','RainTomorrow']].dropna() # 将X转换为NumPy数组 X = dataset[['Humidity3pm', 'Pressure3pm']].values y = dataset['RainTomorrow'] y = np.array([0 if value == 'No' else 1 for value in y]) X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0) def linear_classifier(X, y, learning_rate=0.01, num_epochs=100): num_features = X.shape[1] weights = np.zeros(num_features) bias = 0 for epoch in range(num_epochs): for i in range(X.shape[0]): linear_output = np.dot(X[i], weights) + bias # 统一预测值为0/1 y_pred = 1 if linear_output >= 0 else 0 error = y[i] - y_pred # print("The value of error=", error) weights = weights + learning_rate * error * X[i] bias += learning_rate * error return weights, bias weights, bias = linear_classifier(X_train, y_train)
内容的提问来源于stack exchange,提问作者sarika
相关产品推荐
相关产品推荐

