You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Colab中逻辑回归10折交叉验证报错求助

问题解决:逻辑回归10折交叉验证KeyError报错

报错原因

你的代码里x3_data是Pandas DataFrame类型,直接使用x3_data[train_index]时,Pandas会默认把train_index当作列名去匹配,但KFold.split()返回的是数据的行位置索引,自然找不到对应列,因此触发KeyError。而y3_data如果是Series类型,直接用y3_data[train_index]可行,因为Series的方括号默认按行索引取值。

修复方案

将DataFrame的行提取方式改为.iloc[](按位置索引取值),这是最稳妥的方式,因为KFold返回的索引就是数据的位置序号。

完整可运行代码示例

import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import KFold
from sklearn.metrics import accuracy_score

# 初始化模型与10折交叉验证器
logreg = LogisticRegression(max_iter=1000)  # 增加迭代次数避免收敛问题
kfold = KFold(n_splits=10, shuffle=True, random_state=42)  # 打乱数据保证分布均匀

train_accs, test_accs = [], []

for train_index, test_index in kfold.split(x3_data, y3_data):
    # 用.iloc按位置提取训练/测试集
    x3_train = x3_data.iloc[train_index]
    x3_test = x3_data.iloc[test_index]
    y3_train = y3_data.iloc[train_index]
    y3_test = y3_data.iloc[test_index]
    
    # 训练模型并计算准确率(转成百分比格式)
    logreg.fit(x3_train, y3_train)
    train_acc = accuracy_score(y3_train, logreg.predict(x3_train)) * 100
    test_acc = accuracy_score(y3_test, logreg.predict(x3_test)) * 100
    
    train_accs.append(train_acc)
    test_accs.append(test_acc)

# 整理成指定格式的结果表格
results = pd.DataFrame({
    "Training_Acc": train_accs,
    "Testing_Acc": test_accs
})
results.loc["Av"] = results.mean()

print(results)

额外说明

  • 设置shuffle=True和random_state是为了保证交叉验证划分的一致性,结果可复现。
  • 给逻辑回归模型设置max_iter=1000,避免因数据规模大导致的模型不收敛警告。
  • 最终输出的表格完全符合你期望的格式,包含每折的训练/测试准确率,以及平均值行。

内容的提问来源于stack exchange,提问作者kjnk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 07:05:41