为何我的Python脚本出现IndexError?请求排查问题原因
问题排查:Python脚本触发IndexError及LazyClassifier输出空结果
问题描述
运行Python脚本时触发IndexError: single positional indexer is out-of-bounds,且LazyClassifier输出空的评估结果,以下是相关信息:
数据集格式
#ResidueNoInEachProtein,Residue,TrueLabel,Feature1,Feature2,Feature3,Feature4,Feature5,Feature6,Feature7,Feature8,Feature9 0 GLN C 0.000 0.000 0.000 1 1 1 1 1 0 1 THR E 7.057 10.394 0.000 1 1 1 1 1 0 2 VAL E 6.710 9.449 13.140 0 0 0 0 1 0 3 PRO E 6.552 9.752 12.974 0 0 0 0 0 0 4 SER C 6.544 7.584 11.239 0 0 0 0 0 0 5 SER C 5.407 5.140 5.159 0 0 0 0 0 0 6 ASP C 5.485 7.378 5.152 0 0 0 0 0 0 7 GLY C 5.723 9.048 9.571 0 0 0 1 1 0 8 THR C 6.347 9.102 10.812 0 0 0 2 2 0 9 PRO E 6.219 9.620 12.486 0 1 1 3 4 0 10 ILE E 6.412 9.721 12.781 0 0 0 3 4 0 11 ALA E 6.603 10.294 13.140 0 1 1 2 3 0 12 PHE E 7.219 10.586 13.126 0 0 0 2 2 0 13 GLU E 6.939 10.295 13.972 0 0 0 0 1 0 14 ARG E 6.814 10.472 13.764 0 0 0 0 0 0 15 SER E 7.061 9.189 12.947 0 0 0 0 0 0 16 GLY E 6.872 9.856 11.521 0 0 0 0 0 0 17 SER C 6.988 9.388 11.337 0 0 0 0 0 0 18 GLY C 6.903 7.889 9.055 0 0 0 0 0 0
原Python代码
import pandas as pd from sklearn.model_selection import train_test_split from lazypredict.Supervised import LazyClassifier # Load the data from full.regular.txt data = pd.read_csv('full.regular.txt', delim_whitespace=True) # Verify the data structure print("Data preview:") print(data.head()) print("\nData columns:") print(data.columns) # Assuming columns are correctly read, let’s inspect their index positions print("\nColumn positions and data types:") print(data.dtypes) # Extract the true label and features based on given specifications # Column 3 is index 2 and columns 4 to 12 are index 3 to 11 try: y = data.iloc[:, 2] X = data.iloc[:, 3:11] except IndexError as e: print(f"IndexError: {e}") print("The dataset does not have the expected number of columns.") print("Please check the dataset and ensure it matches the expected format.") # Verify the shapes of X and y if the extraction was successful if 'X' in locals() and 'y' in locals(): print("\nFeatures shape:", X.shape) print("Labels shape:", y.shape) # Split the data into training and testing sets X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) # Initialize and run LazyClassifier clf = LazyClassifier(verbose=0, ignore_warnings=True, custom_metric=None) train, test = clf.fit(X_train, X_test, y_train, y_test) # Display the results print("\nTraining set evaluation:") print(train) print("\nTest set evaluation:") print(test)
运行输出
Data preview: #ResidueNoInEachProtein,Residue,TrueLabel,Feature1,Feature2,Feature3,Feature4,Feature5,Feature6,Feature7,Feature8,Feature9 0 GLN C 0.00 0.00 0.00 1 1 1 1 1 0 1 THR E 7.06 10.39 0.00 1 1 1 1 1 0 2 VAL E 6.71 9.45 13.14 0 0 0 0 1 0 3 PRO E 6.55 9.75 12.97 0 0 0 0 0 0 4 SER C 6.54 7.58 11.24 0 0 0 0 0 0 Data columns: Index(['#ResidueNoInEachProtein,Residue,TrueLabel,Feature1,Feature2,Feature3,Feature4,Feature5,Feature6,Feature7,Feature8,Feature9'], dtype='object') Column positions and data types: #ResidueNoInEachProtein,Residue,TrueLabel,Feature1,Feature2,Feature3,Feature4,Feature5,Feature6,Feature7,Feature8,Feature9 int64 dtype: object IndexError: single positional indexer is out-of-bounds The dataset does not have the expected number of columns. Please check the dataset and ensure it matches the expected format. Features shape: (1079134, 8) Labels shape: (1079134,) 100%|██████████| 29/29 [00:00<00:00, 3085.30it/s] Training set evaluation: Empty DataFrame Columns: [Accuracy, Balanced Accuracy, ROC AUC, F1 Score, Time Taken] Index: [] Test set evaluation: Empty DataFrame Columns: [Accuracy, Balanced Accuracy, ROC AUC, F1 Score, Time Taken] Index: []
问题根源分析
数据读取错误:
- 数据集的表头行使用逗号分隔,但数据行使用空格分隔,而你用
delim_whitespace=True读取时,pandas会把整个表头行当成一个列名,同时把数据行错误拆分为2列(第一列是合并的多字段内容,第二列是最后一个数值)。从输出的Data columns可以看到,实际只有1个列名,加上拆分出的第二列,总共只有2列,所以尝试访问iloc[:,2](第三列)必然触发IndexError。
- 数据集的表头行使用逗号分隔,但数据行使用空格分隔,而你用
LazyClassifier空结果原因:
- 输出中显示的
Features shape: (1079134, 8)和Labels shape: (1079134,)并非来自当前读取的错误数据,而是之前运行脚本时留在内存中的旧变量。LazyClassifier处理的是这些无效/不匹配的数据,导致无法生成有效评估结果,最终输出空DataFrame。
- 输出中显示的
解决方案
需要修正数据读取方式,适配表头逗号分隔、数据行空格分隔的格式:
修正后的代码
import pandas as pd from sklearn.model_selection import train_test_split from lazypredict.Supervised import LazyClassifier # 读取表头行,分割得到列名 with open('full.regular.txt', 'r') as f: header_line = f.readline().strip() columns = header_line.split(',') # 读取剩余数据,用空格分隔,指定列名 data = pd.read_csv('full.regular.txt', delim_whitespace=True, skiprows=[0], names=columns) # 验证数据结构 print("Data preview:") print(data.head()) print("\nData columns:") print(data.columns) print("\nColumn positions and data types:") print(data.dtypes) # 提取标签和特征(确认列索引正确) try: y = data.iloc[:, 2] # TrueLabel列,索引2 X = data.iloc[:, 3:11] # Feature1到Feature8,索引3到10(左闭右开,3:11包含3-10) except IndexError as e: print(f"IndexError: {e}") print(f"当前数据集实际列数:{len(data.columns)},请确认列名和索引对应关系") else: print("\nFeatures shape:", X.shape) print("Labels shape:", y.shape) # 拆分数据集 X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) # 运行LazyClassifier clf = LazyClassifier(verbose=0, ignore_warnings=True, custom_metric=None) train, test = clf.fit(X_train, X_test, y_train, y_test) # 显示结果 print("\nTraining set evaluation:") print(train) print("\nTest set evaluation:") print(test)
关键修正点
- 单独读取表头行,用逗号分割得到正确的列名列表。
- 读取数据时跳过表头行,用
names=columns指定列名,同时保留delim_whitespace=True处理空格分隔的数据行。 - 使用
else块替代之前的locals()判断,确保只有在成功提取X和y后才执行后续流程,避免旧变量干扰。
内容的提问来源于stack exchange,提问作者user366312
相关产品推荐
相关产品推荐

