如何解决CatBoost代码报错:None of [Index([2,3,4,6,7,9], dtype='int32')] are in the [columns]
问题排查与解决
1. 核心报错修复
报错None of [Index([2, 3, 4, 6, 7, 9], dtype='int32')] are in the [columns]的原因是:在KFold循环中,直接用X[train_index]尝试按行位置索引DataFrame,但pandas会把train_index当成列名匹配,导致找不到对应列。
修复方法:使用.iloc按位置索引行:
# 替换循环内的行索引代码 X_train, X_val = X.iloc[train_index], X.iloc[val_index]
2. 额外问题修复(避免后续报错)
(1)目标变量索引问题
你把y转成了list,而list不支持数组式索引(y[train_index]会报错),应该保留为pandas Series或转成numpy数组:
# 替换y的定义 y = df1['N'].values # 转成numpy数组,支持索引 # 或者直接保留Series:y = df1['N']
(2)模型类型错误
目标变量N是连续数值(如1326、1219等),属于回归任务,但你用了CatBoostClassifier(分类模型),后续用accuracy(分类指标)评估完全不合理。应该改用CatBoostRegressor,并使用回归评估指标(如R²、MAE):
# 替换模型导入与定义 from catboost import CatBoostRegressor model = CatBoostRegressor(max_depth=max_depth, learning_rate=learning_rate, iterations=100, random_seed=42, verbose=False) # Regressor的score方法默认返回R²,也可以手动计算其他回归指标
(3)数据读取问题
你的data.csv是空格分隔的,直接用pd.read_csv可能读取错误,建议指定分隔符:
df1 = pd.read_csv('data.csv', sep='\s+')
完整修正代码
import numpy as np import pandas as pd import seaborn as sns import matplotlib.pyplot as plt from pandas import read_csv df1 = pd.read_csv('data.csv', sep='\s+') # CATBOOST 回归模型 from catboost import CatBoostRegressor from sklearn.model_selection import KFold X = df1.drop('N', axis=1) y = df1['N'].values # 转成numpy数组支持索引 # 设置k折交叉验证 k = 3 kf = KFold(n_splits=k, shuffle=True, random_state=42) # 参数网格 max_depth_values = [4, 6, 8, 10, 12] learning_rate_values = [0.01, 0.05, 0.1, 0.2, 0.3] # 存储结果的DataFrame results_df = pd.DataFrame(columns=['max_depth', 'learning_rate', 'mean_r2_score']) # 交叉验证循环 for max_depth in max_depth_values: for learning_rate in learning_rate_values: print(f"训练模型:max_depth={max_depth}, learning_rate={learning_rate}") r2_scores = [] for train_index, val_index in kf.split(X, y): X_train, X_val = X.iloc[train_index], X.iloc[val_index] y_train, y_val = y[train_index], y[val_index] model = CatBoostRegressor(max_depth=max_depth, learning_rate=learning_rate, iterations=100, random_seed=42, verbose=False) model.fit(X_train, y_train, eval_set=(X_val, y_val)) r2 = model.score(X_val, y_val) r2_scores.append(r2) mean_r2 = np.mean(r2_scores) new_row = {'max_depth': max_depth, 'learning_rate': learning_rate, 'mean_r2_score': mean_r2} results_df = pd.concat([results_df, pd.DataFrame([new_row])], ignore_index=True) # 可视化结果 results_pivot = results_df.pivot(index='max_depth', columns='learning_rate', values='mean_r2_score') plt.figure(figsize=(10, 6)) sns.heatmap(results_pivot, annot=True, fmt=".3f", cmap="YlGnBu") plt.title('CatBoost回归模型平均R²得分热力图') plt.show()
内容的提问来源于stack exchange,提问作者HP N
相关产品推荐
相关产品推荐

