You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行含train_test_split的Python代码出现KeyError,请求排查问题

Fixing the KeyError in Your Train-Test Split Code

Let's break down what's going on here and fix that KeyError step by step.

The Root of the Problem

First, let's walk through your code's logic:

  1. You use train_test_split to split x1 (features, I assume) and x2 (labels) into training and test sets. By default, this function shuffles your data and preserves the original indexes—so X_train is a random subset of x1, not the first 80% of rows in order.
  2. Then you loop over sequential integers (0, 1, 2, ..., 1999) and try to pull values from x2 using those integers as indexes.

Here's why this breaks:

  • If x2 is a pandas Series or DataFrame, its index might not be a simple 0-to-2499 sequence. Even if it was, the shuffle from train_test_split means the rows in X_train don't line up with the first 2000 rows of x2.
  • You're also mixing up the split data: y_train is already the subset of x2 that pairs with X_train—you don't need to go back to the original x2 to get those values.

Simple Fixes

You've got a few easy ways to fix this, depending on what you need:

Option 1: Print the Split Labels Directly (Recommended)

If you want to see the labels that match X_train, just iterate over y_train directly—no index mess required:

X_train, X_test, y_train, y_test = train_test_split(x1, x2, test_size=0.2)
for label in y_train:
    print(label)

Option 2: Use the Actual Indexes from the Split Data

If you specifically need to reference the original x2 using indexes from X_train, use the .index attribute of X_train (since train_test_split keeps matching indexes between X_train and y_train):

X_train, X_test, y_train, y_test = train_test_split(x1, x2, test_size=0.2)
for idx in X_train.index:
    print(x2[idx])

This will give you the exact same output as iterating over y_train—it's just another way to access the same paired data.

Option 3: Reset Indexes for Sequential Access

If you really want to use sequential integer indexes for your split data, reset the indexes after splitting to start from 0:

X_train, X_test, y_train, y_test = train_test_split(x1, x2, test_size=0.2)
X_train = X_train.reset_index(drop=True)
y_train = y_train.reset_index(drop=True)

d = len(X_train)
for i in range(d):
    print(y_train[i])

This ensures your indexes are sequential, so range(d) will work without hitting invalid keys.

Why You Saw Partial Data First

The partial output you got before the error is because some of the early i values (0, 1, 2, ...) happened to exist as valid indexes in x2. Once the loop hit an i that wasn't a valid index in x2, it threw the KeyError.

内容的提问来源于stack exchange,提问作者Arif Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:54:45