You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决CatBoost代码报错:None of [Index([2,3,4,6,7,9], dtype='int32')] are in the [columns]

问题排查与解决

1. 核心报错修复

报错None of [Index([2, 3, 4, 6, 7, 9], dtype='int32')] are in the [columns]的原因是:在KFold循环中,直接用X[train_index]尝试按行位置索引DataFrame,但pandas会把train_index当成列名匹配,导致找不到对应列。

修复方法:使用.iloc按位置索引行:

# 替换循环内的行索引代码
X_train, X_val = X.iloc[train_index], X.iloc[val_index]

2. 额外问题修复(避免后续报错)

(1)目标变量索引问题

你把y转成了list,而list不支持数组式索引(y[train_index]会报错),应该保留为pandas Series或转成numpy数组:

# 替换y的定义
y = df1['N'].values  # 转成numpy数组,支持索引
# 或者直接保留Series:y = df1['N']

(2)模型类型错误

目标变量N是连续数值(如1326、1219等),属于回归任务,但你用了CatBoostClassifier(分类模型),后续用accuracy(分类指标)评估完全不合理。应该改用CatBoostRegressor,并使用回归评估指标(如R²、MAE):

# 替换模型导入与定义
from catboost import CatBoostRegressor

model = CatBoostRegressor(max_depth=max_depth, learning_rate=learning_rate, iterations=100, random_seed=42, verbose=False)
# Regressor的score方法默认返回R²,也可以手动计算其他回归指标

(3)数据读取问题

你的data.csv是空格分隔的,直接用pd.read_csv可能读取错误,建议指定分隔符:

df1 = pd.read_csv('data.csv', sep='\s+')

完整修正代码

import numpy as np
import pandas as pd
import seaborn as sns 
import matplotlib.pyplot as plt

from pandas import read_csv
df1 = pd.read_csv('data.csv', sep='\s+')

# CATBOOST 回归模型
from catboost import CatBoostRegressor
from sklearn.model_selection import KFold

X = df1.drop('N', axis=1)
y = df1['N'].values  # 转成numpy数组支持索引

# 设置k折交叉验证
k = 3
kf = KFold(n_splits=k, shuffle=True, random_state=42)

# 参数网格
max_depth_values = [4, 6, 8, 10, 12]
learning_rate_values = [0.01, 0.05, 0.1, 0.2, 0.3]

# 存储结果的DataFrame
results_df = pd.DataFrame(columns=['max_depth', 'learning_rate', 'mean_r2_score'])

# 交叉验证循环
for max_depth in max_depth_values:
    for learning_rate in learning_rate_values:
        print(f"训练模型:max_depth={max_depth}, learning_rate={learning_rate}")
        r2_scores = []
        for train_index, val_index in kf.split(X, y):
            X_train, X_val = X.iloc[train_index], X.iloc[val_index]
            y_train, y_val = y[train_index], y[val_index]

            model = CatBoostRegressor(max_depth=max_depth, learning_rate=learning_rate, iterations=100, random_seed=42, verbose=False)
            model.fit(X_train, y_train, eval_set=(X_val, y_val))
            r2 = model.score(X_val, y_val)
            r2_scores.append(r2)

        mean_r2 = np.mean(r2_scores)
        new_row = {'max_depth': max_depth, 'learning_rate': learning_rate, 'mean_r2_score': mean_r2}
        results_df = pd.concat([results_df, pd.DataFrame([new_row])], ignore_index=True)
       
# 可视化结果
results_pivot = results_df.pivot(index='max_depth', columns='learning_rate', values='mean_r2_score')
plt.figure(figsize=(10, 6))
sns.heatmap(results_pivot, annot=True, fmt=".3f", cmap="YlGnBu")
plt.title('CatBoost回归模型平均R²得分热力图')
plt.show()

内容的提问来源于stack exchange,提问作者HP N

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 10:45:07