You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

KernelPCA超参数调优时train_test_split遇Singleton数组错误求助

问题解决:KernelPCA超参数调优时train_test_split报错Singleton array cannot be considered a valid collection

错误原因

从报错回溯和代码逻辑来看,核心问题有三点:

  1. train_test_split参数传入错误:你传入的X_kpca是一个KernelPCA模型实例,而非样本特征矩阵。train_test_split需要处理的是原始特征数据(如你定义的X),模型实例无法被识别为有效数据集,直接触发"Singleton array"错误。
  2. GridSearchCV使用逻辑错误:你将X_kpca作为待调参的模型传入,但正确的做法是传入初始化后的kpca实例;同时KernelPCA是无监督模型,直接用GridSearchCV需要指定无监督评分指标,或结合监督模型完成端到端调参。
  3. 参数网格语法错误:fit_inverse_transform的取值写成(bool, False)不符合要求,应改为布尔值列表[True, False]。

修正后的代码

方案1:仅调优KernelPCA无监督参数(以解释方差为评分依据)

import pandas as pd 
import numpy as np
from sklearn.decomposition import KernelPCA
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.metrics import accuracy_score

# 加载特征与标签
X = meth_clin_sub_nt_2_kipan.iloc[:,7:-1]
y_type = meth_clin_sub_nt_2_kipan["type"]

# 拆分原始特征数据(而非模型实例)
X_train, X_test, y_train, y_test = train_test_split(X, y_type, test_size=0.3, random_state=30)

# 初始化KernelPCA模型
kpca = KernelPCA()

# 修正参数网格的错误取值
param_grid = {
    'n_components': list(range(1,9)),
    'kernel': ('linear', 'poly', 'rbf', 'sigmoid', 'cosine'),  # 移除precomputed,需提前计算核矩阵才可使用
    'degree': list(range(1,9)),
    'tol': [0.0, 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0],
    'fit_inverse_transform': [True, False],  # 修正为合法布尔值列表
    'eigen_solver': ('auto', 'dense', 'arpack', 'randomized'),
    'alpha': [0.0, 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0]
}

# 用无监督评分指标调参
gs = GridSearchCV(kpca, param_grid, cv=10, scoring='explained_variance')
gs.fit(X_train)

# 获取最优模型并完成降维
best_kpca = gs.best_estimator_
X_train_kpca = best_kpca.transform(X_train)
X_test_kpca = best_kpca.transform(X_test)

# 后续可基于降维数据训练分类器
# 示例:
# from sklearn.linear_model import LogisticRegression
# clf = LogisticRegression(max_iter=1000)
# clf.fit(X_train_kpca, y_train)
# print(accuracy_score(y_test, clf.predict(X_test_kpca)))

方案2:结合分类器端到端调优(更适合监督场景)

若最终目标是用降维数据做分类,推荐用Pipeline串联KernelPCA与分类器,同时调优两者参数:

import pandas as pd 
import numpy as np
from sklearn.decomposition import KernelPCA
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.metrics import accuracy_score
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline

X = meth_clin_sub_nt_2_kipan.iloc[:,7:-1]
y_type = meth_clin_sub_nt_2_kipan["type"]

X_train, X_test, y_train, y_test = train_test_split(X, y_type, test_size=0.3, random_state=30)

# 构建Pipeline:先降维,再分类
pipe = Pipeline([
    ('kpca', KernelPCA()),
    ('clf', LogisticRegression(max_iter=1000))
])

# 联合参数网格
param_grid = {
    'kpca__n_components': list(range(1,9)),
    'kpca__kernel': ('linear', 'poly', 'rbf', 'sigmoid', 'cosine'),
    'kpca__degree': list(range(1,9)),
    'kpca__tol': [0.0, 1.0, 2.0],
    'kpca__fit_inverse_transform': [True, False],
    'clf__C': [0.1, 1, 10]
}

# 以分类准确率为评分指标调参
gs = GridSearchCV(pipe, param_grid, cv=10, scoring='accuracy')
gs.fit(X_train, y_train)

# 输出最优模型性能
print(f"交叉验证最优准确率:{gs.best_score_:.4f}")
print(f"测试集准确率:{gs.score(X_test, y_test):.4f}")

额外说明

  • 移除了参数网格中的precomputed核选项,使用该选项需提前计算好核矩阵,否则会触发错误;若确有需求,需先预处理核矩阵再传入。
  • 无监督调参时,explained_variance是适配KernelPCA的评分指标,用于衡量降维后保留的方差比例。

内容的提问来源于stack exchange,提问作者melolilili

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 07:07:04