You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何我的Python脚本出现IndexError?请求排查问题原因

问题排查:Python脚本触发IndexError及LazyClassifier输出空结果

问题描述

运行Python脚本时触发IndexError: single positional indexer is out-of-bounds,且LazyClassifier输出空的评估结果,以下是相关信息:

数据集格式

#ResidueNoInEachProtein,Residue,TrueLabel,Feature1,Feature2,Feature3,Feature4,Feature5,Feature6,Feature7,Feature8,Feature9
   0 GLN C   0.000   0.000   0.000 1 1 1 1  1 0
   1 THR E   7.057  10.394   0.000 1 1 1 1  1 0
   2 VAL E   6.710   9.449  13.140 0 0 0 0  1 0
   3 PRO E   6.552   9.752  12.974 0 0 0 0  0 0
   4 SER C   6.544   7.584  11.239 0 0 0 0  0 0
   5 SER C   5.407   5.140   5.159 0 0 0 0  0 0
   6 ASP C   5.485   7.378   5.152 0 0 0 0  0 0
   7 GLY C   5.723   9.048   9.571 0 0 0 1  1 0
   8 THR C   6.347   9.102  10.812 0 0 0 2  2 0
   9 PRO E   6.219   9.620  12.486 0 1 1 3  4 0
  10 ILE E   6.412   9.721  12.781 0 0 0 3  4 0
  11 ALA E   6.603  10.294  13.140 0 1 1 2  3 0
  12 PHE E   7.219  10.586  13.126 0 0 0 2  2 0
  13 GLU E   6.939  10.295  13.972 0 0 0 0  1 0
  14 ARG E   6.814  10.472  13.764 0 0 0 0  0 0
  15 SER E   7.061   9.189  12.947 0 0 0 0  0 0
  16 GLY E   6.872   9.856  11.521 0 0 0 0  0 0
  17 SER C   6.988   9.388  11.337 0 0 0 0  0 0
  18 GLY C   6.903   7.889   9.055 0 0 0 0  0 0

原Python代码

import pandas as pd
from sklearn.model_selection import train_test_split
from lazypredict.Supervised import LazyClassifier

# Load the data from full.regular.txt
data = pd.read_csv('full.regular.txt', delim_whitespace=True)

# Verify the data structure
print("Data preview:")
print(data.head())
print("\nData columns:")
print(data.columns)

# Assuming columns are correctly read, let’s inspect their index positions
print("\nColumn positions and data types:")
print(data.dtypes)

# Extract the true label and features based on given specifications
# Column 3 is index 2 and columns 4 to 12 are index 3 to 11
try:
    y = data.iloc[:, 2]
    X = data.iloc[:, 3:11]
except IndexError as e:
    print(f"IndexError: {e}")
    print("The dataset does not have the expected number of columns.")
    print("Please check the dataset and ensure it matches the expected format.")

# Verify the shapes of X and y if the extraction was successful
if 'X' in locals() and 'y' in locals():
    print("\nFeatures shape:", X.shape)
    print("Labels shape:", y.shape)

    # Split the data into training and testing sets
    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

    # Initialize and run LazyClassifier
    clf = LazyClassifier(verbose=0, ignore_warnings=True, custom_metric=None)
    train, test = clf.fit(X_train, X_test, y_train, y_test)

    # Display the results
    print("\nTraining set evaluation:")
    print(train)
    print("\nTest set evaluation:")
    print(test)

运行输出

Data preview:
                                    #ResidueNoInEachProtein,Residue,TrueLabel,Feature1,Feature2,Feature3,Feature4,Feature5,Feature6,Feature7,Feature8,Feature9
0 GLN C 0.00 0.00  0.00  1 1 1 1 1                                                  0                                                                         
1 THR E 7.06 10.39 0.00  1 1 1 1 1                                                  0                                                                         
2 VAL E 6.71 9.45  13.14 0 0 0 0 1                                                  0                                                                         
3 PRO E 6.55 9.75  12.97 0 0 0 0 0                                                  0                                                                         
4 SER C 6.54 7.58  11.24 0 0 0 0 0                                                  0                                                                         

Data columns:
Index(['#ResidueNoInEachProtein,Residue,TrueLabel,Feature1,Feature2,Feature3,Feature4,Feature5,Feature6,Feature7,Feature8,Feature9'], dtype='object')

Column positions and data types:
#ResidueNoInEachProtein,Residue,TrueLabel,Feature1,Feature2,Feature3,Feature4,Feature5,Feature6,Feature7,Feature8,Feature9    int64
dtype: object
IndexError: single positional indexer is out-of-bounds
The dataset does not have the expected number of columns.
Please check the dataset and ensure it matches the expected format.

Features shape: (1079134, 8)
Labels shape: (1079134,)
100%|██████████| 29/29 [00:00<00:00, 3085.30it/s]

Training set evaluation:
Empty DataFrame
Columns: [Accuracy, Balanced Accuracy, ROC AUC, F1 Score, Time Taken]
Index: []

Test set evaluation:
Empty DataFrame
Columns: [Accuracy, Balanced Accuracy, ROC AUC, F1 Score, Time Taken]
Index: []

问题根源分析

  1. 数据读取错误:

    • 数据集的表头行使用逗号分隔,但数据行使用空格分隔,而你用delim_whitespace=True读取时,pandas会把整个表头行当成一个列名,同时把数据行错误拆分为2列(第一列是合并的多字段内容,第二列是最后一个数值)。从输出的Data columns可以看到,实际只有1个列名,加上拆分出的第二列,总共只有2列,所以尝试访问iloc[:,2](第三列)必然触发IndexError。
  2. LazyClassifier空结果原因:

    • 输出中显示的Features shape: (1079134, 8)和Labels shape: (1079134,)并非来自当前读取的错误数据,而是之前运行脚本时留在内存中的旧变量。LazyClassifier处理的是这些无效/不匹配的数据,导致无法生成有效评估结果,最终输出空DataFrame。

解决方案

需要修正数据读取方式,适配表头逗号分隔、数据行空格分隔的格式:

修正后的代码

import pandas as pd
from sklearn.model_selection import train_test_split
from lazypredict.Supervised import LazyClassifier

# 读取表头行,分割得到列名
with open('full.regular.txt', 'r') as f:
    header_line = f.readline().strip()
    columns = header_line.split(',')

# 读取剩余数据,用空格分隔,指定列名
data = pd.read_csv('full.regular.txt', delim_whitespace=True, skiprows=[0], names=columns)

# 验证数据结构
print("Data preview:")
print(data.head())
print("\nData columns:")
print(data.columns)
print("\nColumn positions and data types:")
print(data.dtypes)

# 提取标签和特征(确认列索引正确)
try:
    y = data.iloc[:, 2]  # TrueLabel列,索引2
    X = data.iloc[:, 3:11]  # Feature1到Feature8,索引3到10(左闭右开,3:11包含3-10)
except IndexError as e:
    print(f"IndexError: {e}")
    print(f"当前数据集实际列数:{len(data.columns)},请确认列名和索引对应关系")
else:
    print("\nFeatures shape:", X.shape)
    print("Labels shape:", y.shape)

    # 拆分数据集
    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

    # 运行LazyClassifier
    clf = LazyClassifier(verbose=0, ignore_warnings=True, custom_metric=None)
    train, test = clf.fit(X_train, X_test, y_train, y_test)

    # 显示结果
    print("\nTraining set evaluation:")
    print(train)
    print("\nTest set evaluation:")
    print(test)

关键修正点

  • 单独读取表头行,用逗号分割得到正确的列名列表。
  • 读取数据时跳过表头行,用names=columns指定列名,同时保留delim_whitespace=True处理空格分隔的数据行。
  • 使用else块替代之前的locals()判断,确保只有在成功提取X和y后才执行后续流程,避免旧变量干扰。

内容的提问来源于stack exchange,提问作者user366312

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 12:38:10