You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何手动for循环调参优于GridSearchCV?如何优化KNN回归效果?

KNN回归调参问题解答

问题场景

我在KNN回归任务中遇到了调参矛盾:手动设置n_neighbors=3时,测试集均方根误差为7057;但用GridSearchCV得到的最优n_neighbors=4,对应误差却高达11473。想请教两个问题:

  1. 为什么手动for循环调参的结果比GridSearchCV更好?
  2. 怎么提升最优邻居数下的预测效果(降低误差)?

代码实现

import pandas as pd
from sklearn.model_selection import train_test_split
import numpy as np
import matplotlib.pyplot as plt
from sklearn.neighbors import KNeighborsRegressor
from sklearn.model_selection import GridSearchCV
from sklearn.metrics import mean_squared_error
from math import sqrt

file_data = pd.read_csv(r"C:\Users\Ashil-Rayan\Desktop\Data\RankHome\housePricelimit2.csv", header=None)
data = file_data.drop(6, axis=1)
target = file_data[6]

z = []
X_train, X_test, y_train, y_test = train_test_split(data, target, test_size=0.2, random_state=42)

# 手动循环寻找最优邻居数
for i in range(1,10):
    knn = KNeighborsRegressor(n_neighbors=i,  weights="distance",)
    knn.fit(X_train, y_train)
    predicts = knn.predict(X_test)
    result = mean_squared_error(y_test, predicts)
    result = sqrt(result)
    score = knn.score(X_test, y_test)
    z.append((i, result, score))
x = []
y = []
for i in z:
    x.append(i[0])
    y.append(i[1])
plt.scatter(x, y)
plt.xlabel("n_neighbors")
plt.ylabel("mean_squared_error")
plt.show()
print(z[1])
print(z[3])

# 使用GridSearchCV调参
parametrs = {"n_neighbors" : range(1, 50),
             "weights" : ["uniform", "distance"]}
Grid = GridSearchCV(KNeighborsRegressor(), parametrs)
Grid.fit(X_train, y_train)
print(Grid.best_params_)
# print(knn.predict([new_data]))
print(knn.score(X_test, y_test))

相关图表

图表1:手动循环调参中n_neighbors与均方根误差的关系散点图
图表2:GridSearchCV参数搜索的相关结果图


问题1:手动调参优于GridSearchCV的原因

  • 评价逻辑差异:手动循环直接用独立测试集计算误差选最优参数,属于“数据泄露”式调参;而GridSearchCV默认用**交叉验证(CV)**在训练集内部划分验证集评估参数,选的是CV得分最高的参数,这个参数是为训练集泛化能力优化的,和直接用测试集选的参数自然可能不同。
  • 参数组合不匹配:手动循环固定了weights="distance",但GridSearchCV同时搜索uniform和distance两种权重,它选出的最优参数(比如n_neighbors=4)大概率搭配的是uniform权重,而你手动测试n_neighbors=4时用的是distance权重,两者不是同一参数组合,误差没有可比性。
  • 代码细节错误:你最后打印的knn.score(X_test, y_test)调用的是手动循环最后一次训练的模型(n_neighbors=9,distance权重),并非GridSearchCV选出的最优模型,这导致你错误判断了最优参数的误差表现。正确做法是用Grid.best_estimator_.predict(X_test)计算测试集误差。

问题2:提升预测效果的方法

  • 强制特征缩放:KNN基于距离计算,对特征尺度极度敏感。用StandardScaler或MinMaxScaler对所有特征做标准化/归一化,避免数值量级大的特征主导距离计算。
  • 优化交叉验证:GridSearchCV默认5折CV,可增加折数(如cv=10)提升参数评估稳定性;针对回归任务,也可尝试时间序列CV或分组CV(如果数据有分组属性)。
  • 扩展参数搜索:除了n_neighbors和weights,还可以搜索metric参数(如"manhattan"、"chebyshev"),不同距离度量可能更适配你的房价数据。
  • 特征工程优化:
    • 用相关性分析、递归特征消除筛选重要特征,移除冗余/噪声特征;
    • 对房价这类右偏分布的目标变量做对数变换,让模型更易拟合;
    • 尝试特征组合(如房间数*面积)生成更具预测性的特征。
  • 尝试KNN变种:比如使用局部加权回归(LWR),或者自定义邻居权重函数,让近邻的权重分配更贴合数据分布。

内容的提问来源于stack exchange,提问作者Kousha Zhiyani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 23:02:04