Keras模型超参数寻优前交叉验证及KeyError问题解决
问题:Sequential模型交叉验证时KeyError的修正方法
我希望在为Sequential模型挑选最优超参数集之前执行交叉验证。已知KeyError是由x_train_scaled的列名与cv_train的值不匹配导致的,但不知道该如何修正。相关代码及报错信息如下:
导入依赖库
import math import pandas as pd import numpy as np import tensorflow as tf from tensorflow.keras import Model, Sequential, layers from tensorflow.keras.optimizers import Adam from sklearn.preprocessing import StandardScaler, LabelEncoder from tensorflow.keras.layers import Dense, Dropout, Flatten from sklearn.model_selection import train_test_split, StratifiedKFold
训练集-测试集拆分
# 划分特征与目标变量 x_train, x_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=1)
数据标准化
def scale_datasets(x_train, x_test): """ 标准化训练集和测试集 Z-Score归一化 """ standard_scaler = StandardScaler() x_train_scaled = pd.DataFrame( standard_scaler.fit_transform(x_train), columns=x_train.columns ) x_test_scaled = pd.DataFrame( standard_scaler.transform(x_test), columns = x_test.columns ) return x_train_scaled, x_test_scaled # 执行标准化 x_train_scaled, x_test_scaled = scale_datasets(x_train, x_test) for cv_train, cv_test in kfold.split(x_train_scaled, y_train):
超参数调优(Keras Tuner)
# 使用最优超参数构建模型,并在数据上训练50轮 model = tuner.hypermodel.build(best_hps) history = model.fit(x_train_scaled.iloc[cv_train], y_train.iloc[cv_train], epochs=50, validation_split=0.2)
报错信息
--------------------------------------------------------------------------- KeyError Traceback (most recent call last) /tmp/ipykernel_17/4191267479.py in <module> 4 # Build the model with the optimal hyperparameters and train it on the data for 50 epochs 5 model = tuner.hypermodel.build(best_hps) ----> 6 history = model.fit(x_train_scaled.iloc[cv_train], y_train.iloc[cv_train], epochs=50, validation_split=0.2) 7 8 val_acc_per_epoch = history.history['val_accuracy'] /opt/conda/lib/python3.7/site-packages/pandas/core/frame.py in __getitem__(self, key) 3462 if is_iterator(key): 3463 key = list(key) -> 3464 indexer = self.loc._get_listlike_indexer(key, axis=1)[1] 3465 3466 # take() does not accept boolean indexers /opt/conda/lib/python3.7/site-packages/pandas/core/indexing.py in _get_listlike_indexer(self, key, axis) 1312 keyarr, indexer, new_indexer = ax._reindex_non_unique(keyarr) 1313 -> 1314 self._validate_read_indexer(keyarr, indexer, axis) 1315 1316 if needs_i8_conversion(ax.dtype) or isinstance( /opt/conda/lib/python3.7/site-packages/pandas/core/indexing.py in _validate_read_indexer(self, key, indexer, axis) 1372 if use_interval_msg: 1373 key = list(key) -> 1374 raise KeyError(f"None of [{key}] are in the [{axis_name}]") 1375 1376 not_found = list(ensure_index(key)[missing_mask.nonzero()[0]].unique()) KeyError: "None of [Int64Index([ 0, 1, 2, 3, 4, 5, 6, 7, 8, 10, ... 682, 683, 685, 686, 687, 688, 689, 690, 691, 692], dtype='int64', length=623)] are in the [columns]"
修正方案
1. 修复代码缩进问题
报错核心原因是模型构建与训练代码不在for循环的缩进块内:cv_train是kfold.split()返回的迭代器对象,而非单轮交叉验证的训练集行索引。当代码不在循环内时,iloc[cv_train]会被错误解析为列索引查询,从而触发KeyError。
将超参数调优代码缩进,放入for循环内部:
# 初始化交叉验证器(缺失的关键步骤) kfold = StratifiedKFold(n_splits=5, shuffle=True, random_state=1) # 执行标准化 x_train_scaled, x_test_scaled = scale_datasets(x_train, x_test) # 遍历交叉验证折 for cv_train, cv_test in kfold.split(x_train_scaled, y_train): # 使用最优超参数构建模型,并在当前折的训练集上训练 model = tuner.hypermodel.build(best_hps) history = model.fit(x_train_scaled.iloc[cv_train], y_train.iloc[cv_train], epochs=50, validation_split=0.2)
2. 补充交叉验证器初始化代码
原代码缺少StratifiedKFold的实例化步骤,需添加上述代码中的kfold = StratifiedKFold(...),确保交叉验证器正确配置。
3. 可选:转为Numpy数组避免索引问题
TensorFlow模型可直接处理Numpy数组,无需依赖Pandas索引系统。可将标准化后的DataFrame转为数组,彻底规避索引相关错误:
# 修改标准化函数,返回Numpy数组 def scale_datasets(x_train, x_test): standard_scaler = StandardScaler() x_train_scaled = standard_scaler.fit_transform(x_train) x_test_scaled = standard_scaler.transform(x_test) return x_train_scaled, x_test_scaled # 后续训练直接使用数组,无需iloc history = model.fit(x_train_scaled[cv_train], y_train.iloc[cv_train].values, epochs=50, validation_split=0.2)
内容的提问来源于stack exchange,提问作者melolilili
相关产品推荐
相关产品推荐

