You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用LDA降维时将DataFrame转为变量遇KeyError(1)求助

排查KeyError(1)的问题及解决方案

Hey there, let's break down why you're hitting that frustrating KeyError(1) when trying to extract features from your DataFrames.

The Root Cause

When you use syntax like holds[[0,1,2,...8]], Pandas is looking for column names that exactly match those integers (0,1,2...). If your DataFrame's columns aren't named with integers (e.g., they might be default string labels like '0' or custom names like 'feature_1'), Pandas can't find those columns and throws a KeyError. This is a common mix-up between column names and column positions.

Fixes You Can Apply

Here are two straightforward ways to resolve this:

1. Use iloc to select columns by position

iloc lets you index DataFrames by their numerical position (starting from 0), which is exactly what you're trying to do. Update your code like this:

from sklearn.discriminant_analysis import LinearDiscriminantAnalysis as LDA

lda = LDA(n_components=2)

# For the holds DataFrame (9 features + 1 target column)
X = holds.iloc[:, 0:9].values  # Grab first 9 columns (positions 0-8) as features
Y = holds.iloc[:, 9].values     # Grab the 10th column (position 9) as target

# For the latency DataFrame (18 features + 1 target column)
X2 = latency.iloc[:, 0:18].values  # Grab first 18 columns (positions 0-17) as features
Y2 = latency.iloc[:, 18].values     # Grab the 19th column (position 18) as target

Pro tip: If your target column is always the last one in the DataFrame, you can use iloc[:, -1].values for Y/Y2—this works even if you add/remove features later, no need to adjust the index number.

2. Rename columns to integers (if you prefer)

If you want to keep using direct integer indexing, rename your DataFrame columns to match their position indices:

# Rename columns to 0,1,2...n-1
holds.columns = range(holds.shape[1])
latency.columns = range(latency.shape[1])

# Now your original code should work (with a fix for Y2)
X = holds[[0,1,2,3,4,5,6,7,8]].values
Y = holds[9].values
X2 = latency[[0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17]].values
Y2 = latency[18].values  # Fixed this to 18, since latency has 18 features + 1 target

(Side note: In your original code, you had Y2 = latency[9].values—that would grab the 10th column, which is a feature, not your target. Make sure to point to the correct target column position!)

Quick Check to Verify

Before running the code, you can print your column names to confirm what you're working with:

print("Holds columns:", holds.columns)
print("Latency columns:", latency.columns)

This will show you exactly what labels Pandas is using for columns, so you can adjust your indexing accordingly.

内容的提问来源于stack exchange,提问作者Larry Johnson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 07:57:38