训练SVM模型应使用训练集还是主数据集?我的流程是否正确?
你的操作存在错误,正确做法如下
核心问题分析
你当前的错误在于直接用验证集(X_val, y_val)拟合GridSearchCV,这会导致参数选择过度适配验证集,同时浪费了训练集的数据用于调参,最终模型的泛化能力会受到严重影响。
正确的调参逻辑是:参数筛选必须基于训练数据(或训练+验证的合并数据),验证集的作用是评估最终模型的性能,而非参与调参过程。
方案一:合并训练与验证集,用GridSearchCV自带交叉验证调参(推荐)
这种方案无需单独划分验证集,GridSearchCV会自动在训练数据内部执行交叉验证,更高效且能避免单次划分验证集的随机性:
# 1. 划分训练集和测试集 X_train, X_test, y_train, y_test = train_test_split(dataTrain, y, test_size=0.2, random_state=10) # 2. 定义参数范围并执行GridSearchCV(基于训练集做5折交叉验证) svm_param_grid = {'C': [0.0001, 0.001, 0.01, 0.1, 10, 100], 'kernel': ['linear']} grid_search = GridSearchCV(svm.SVC(), param_grid=svm_param_grid, cv=5, verbose=50) grid_search.fit(X_train, y_train) # 查看最优参数 print('最优参数:', grid_search.best_params_) # 3. 获取最优模型(GridSearchCV已自动训练好最优模型) best_svm_model = grid_search.best_estimator_ # 4. 用测试集评估模型泛化能力 test_accuracy = best_svm_model.score(X_test, y_test) print('测试集准确率:', test_accuracy)
方案二:保留独立验证集,基于训练集手动调参
如果一定要保留单独的验证集用于最终评估,需在训练集上训练不同参数的模型,用验证集的性能筛选最优参数:
# 1. 划分训练集、验证集、测试集 X_main, X_test, y_main, y_test = train_test_split(dataTrain, y, test_size=0.2, random_state=10) X_train, X_val, y_train, y_val = train_test_split(X_main, y_main, test_size=0.2, random_state=10) # 2. 遍历参数,训练并评估 best_C = None best_val_score = 0 c_candidates = [0.0001, 0.001, 0.01, 0.1, 10, 100] for c in c_candidates: temp_model = svm.SVC(kernel='linear', C=c) temp_model.fit(X_train, y_train) val_score = temp_model.score(X_val, y_val) print(f'C={c} 时,验证集准确率:{val_score}') if val_score > best_val_score: best_val_score = val_score best_C = c print('筛选出的最优C值:', best_C) # 3. 用最优C在训练+验证集上训练(充分利用数据) final_model = svm.SVC(kernel='linear', C=best_C) final_model.fit(X_main, y_main) # 4. 测试集评估 test_accuracy = final_model.score(X_test, y_test) print('测试集准确率:', test_accuracy)
补充说明
方案一的交叉验证方式更稳定,能减少单次划分验证集带来的偶然性,是工业界和学术研究中更常用的调参方法。
内容的提问来源于stack exchange,提问作者Bn.
相关产品推荐
相关产品推荐

