改用StratifiedGroupKFold后GridSearchCV的best_score_返回nan问题
切换到StratifiedGroupKFold后GridSearchCV的best_score_返回nan值
问题场景
将交叉验证策略从StratifiedKFold替换为StratifiedGroupKFold后,GridSearchCV的grid.best_score_返回nan,但原StratifiedKFold方案运行正常。
相关代码
kfold = StratifiedGroupKFold(n_splits=5, shuffle=True, random_state=0) train_indx, test_indx = next(GroupShuffleSplit(random_state=i, test_size=0.2).split(X, Y, groups)) X_train, X_test, Y_train, Y_test = X[train_indx], X[test_indx], Y[train_indx], Y[test_indx] grid = GridSearchCV(clf, par, cv=kfold, scoring='accuracy', n_jobs=-1) grid.fit(X_train, Y_train) print(str(clf).split('(', 1)[0], "// Best Validation Score: {:.5f}".format(grid.best_score_), "// Test Score: {:.5f}".format(accuracy_score(Y_test, grid.best_estimator_.predict(X_test))))
执行结果示例
LogisticRegression // Best Validation Score: nan // Test Score: 0.56164
问题原因
StratifiedGroupKFold是带分组约束的分层交叉验证器,它要求划分fold时同一组的样本不能同时出现在训练集和验证集中,因此必须传入分组信息才能正常工作。你在调用grid.fit()时仅传入X_train和Y_train,未传递训练集对应的分组数据,导致交叉验证无法正确划分样本,最终所有验证分数计算失败,best_score_返回nan。
解决方法
在grid.fit()中添加groups参数,传入训练集对应的分组数据(即groups[train_indx]):
修改后的代码
kfold = StratifiedGroupKFold(n_splits=5, shuffle=True, random_state=0) train_indx, test_indx = next(GroupShuffleSplit(random_state=i, test_size=0.2).split(X, Y, groups)) X_train, X_test, Y_train, Y_test = X[train_indx], X[test_indx], Y[train_indx], Y[test_indx] # 获取训练集对应的分组数据 train_groups = groups[train_indx] grid = GridSearchCV(clf, par, cv=kfold, scoring='accuracy', n_jobs=-1) # 传入groups参数 grid.fit(X_train, Y_train, groups=train_groups) print(str(clf).split('(', 1)[0], "// Best Validation Score: {:.5f}".format(grid.best_score_), "// Test Score: {:.5f}".format(accuracy_score(Y_test, grid.best_estimator_.predict(X_test))))
额外排查点
- 检查训练集的分组数据,避免出现单组样本量过少的情况(比如某组仅1个样本),否则
StratifiedGroupKFold可能无法完成有效分层划分,仍会出现nan - 确认验证集内包含至少两个类别,避免因单一类别导致准确率计算无意义(不过
StratifiedGroupKFold会保证分层,此情况概率较低)
内容的提问来源于stack exchange,提问作者adsasf
相关产品推荐
相关产品推荐

