练习K近邻算法时触发pandas.errors.ParserError,请求排查原因
解决pandas读取breast-cancer-wisconsin.data.txt时的ParserError问题
我在学习《Python机器学习教程》第14部分的K近邻(K Nearest Neighbors)应用时,使用
breast-cancer-wisconsin.data.txt数据集运行以下代码,触发了错误:pandas.errors.ParserError: Error tokenizing data. C error: Expected 7 fields in line 3, saw 11。
我先梳理了你的问题,根源其实很明确:
错误原因
你把数据集内容和Python代码混在了同一个代码块/文件里!当pd.read_csv('breast-cancer-wisconsin.data.txt')执行时,它读取的文件里不仅有数据集的表头和数据行,还包含了import numpy as np这类代码行。pandas会把所有行都当成CSV数据来解析,代码行的字段数量和数据集的11个字段完全不匹配,自然就抛出了字段数量不匹配的ParserError。
解决方案
- 分离数据集和代码
把纯数据集内容单独保存为breast-cancer-wisconsin.data.txt文件,文件里只保留表头和数据行:id,clump_thickness,unif_cell_size,unif_cell_shape,marg_adhesion,single_epith_cell_size,bare_nuclei,bland_chrom,norm_nucleoli,mitoses,class 1000025,5,1,1,1,2,1,3,1,1,2 1002945,5,4,4,5,7,10,3,2,1,2 1015425,3,1,1,1,2,2,3,1,1,2 1016277,6,8,8,1,3,4,3,7,1,2 1017023,4,1,1,3,2,1,3,1,1,2 1017122,8,10,10,8,7,10,9,7,1,4 - 修正代码文件
把Python代码单独放在一个.py文件中,确保pd.read_csv指向的是正确的数据集文件路径。另外注意:scikit-learn的cross_validation模块已经被弃用,建议替换为model_selection,修正后的代码如下:import numpy as np from sklearn import preprocessing, model_selection, neighbors import pandas as pd df = pd.read_csv('breast-cancer-wisconsin.data.txt') df.replace('?', -99999, inplace=True) # df.drop(['id'], 1, inplace=True) X = np.array(df.drop(['class'], 1)) y = np.array(df['class']) X_train, X_test, y_train, y_test = model_selection.train_test_split(X, y, test_size=0.2) clf = neighbors.KNeighborsClassifier() clf.fit(X_train, y_train) accuracy = clf.score(X_test, y_test) print(accuracy) - 验证文件路径
确保.py文件和breast-cancer-wisconsin.data.txt在同一个文件夹下,或者在pd.read_csv里填写完整的文件路径(比如'./data/breast-cancer-wisconsin.data.txt')。
内容的提问来源于stack exchange,提问作者Briancheung
相关产品推荐
相关产品推荐

