You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

KNN模型曲线平坦、精度不足求助:股票涨跌预测项目

问题:KNN模型预测股票涨跌时误差率曲线平坦、精度不达预期

我正在完成一个学校项目,需要基于自选变量预测股票涨跌。课堂上学过多种算法,现在想让KNN模型生效,但遇到了两个核心问题:

  • KNN误差率曲线呈直线,没法用肘部法寻找最低误差率
  • 模型精度远低于预期(预期精度45%-75%)

我的数据集包含93条月度数据(2015年至2022年9月),相关截图:

  • 数据集:数据集截图
  • KNN误差率曲线:KNN误差率曲线
  • 精度截图:精度截图

我怀疑存在一些问题,比如混淆矩阵的数据量和数据集相比过少(不太清楚怎么在Excel和Google Colab间正确传递数据)。我直接复制了课堂上的代码:

from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
scaler.fit(df.drop('direction',axis=1))
stand_features = scaler.transform(df.drop('direction',axis=1))
df_stand = pd.DataFrame(stand_features,columns=df.columns[:-1])
df_stand.head()
# import library
from sklearn.model_selection import train_test_split
x_train, x_test, y_train, y_test = train_test_split(stand_features,df['direction'],test_size=0.25)
from sklearn.neighbors import KNeighborsClassifier

knn = KNeighborsClassifier(n_neighbors=1)
knn.fit(x_train,y_train)
y_pred = knn.predict(x_test)
error_rate = []

for i in range(1,40):
  knn = KNeighborsClassifier(n_neighbors=1)
  knn.fit(x_train,y_train)
  pred_i = knn.predict(x_test)
  error_rate.append(np.mean(pred_i != y_test))
plt.figure(figsize=(10,6))
plt.plot(range(1,40),error_rate,color='blue', linestyle='dashed', marker='o', markerfacecolor='red', markersize=10)
plt.title('Error Rate vs K Value')
plt.xlabel('K')
plt.ylabel('Error rate')
from sklearn.metrics import confusion_matrix
from sklearn.metrics import classification_report

knn = KNeighborsClassifier(n_neighbors=2)
knn.fit(x_train,y_train)
y_pred = knn.predict(x_test)

print('WITH K=2')
print('\n')
print(confusion_matrix(y_test,y_pred))
print('\n')
print(classification_report(y_test,y_pred))

我原本期望得到类似期望的KNN曲线的KNN曲线,但实际曲线完全平坦,精度极差。请问有什么解决建议?


解决建议

1. 修复K值循环的核心错误

你误差率曲线平坦的直接原因是循环里一直固定用K=1,根本没遍历不同K值!看这段代码:

error_rate = []

for i in range(1,40):
  knn = KNeighborsClassifier(n_neighbors=1)  # 这里应该用i,不是固定1!
  knn.fit(x_train,y_train)
  pred_i = knn.predict(x_test)
  error_rate.append(np.mean(pred_i != y_test))

把n_neighbors=1改成n_neighbors=i,这样循环才会测试K=1到K=39的不同情况,误差率曲线才会出现正常的波动。

2. 优化数据集与样本划分

  • 93条数据按0.25的测试比例划分,测试集只有约23条样本,混淆矩阵样本量小是正常的,但整体数据量太少会导致模型稳定性差。可以尝试:
    • 调整划分比例,比如把test_size设为0.2,让训练集样本更多
    • 检查direction列的涨跌类别分布,如果存在严重不平衡(比如涨的样本占80%,跌的占20%),需要用重采样(过采样少数类/欠采样多数类)或者在模型里设置class_weight='balanced'参数

3. 特征工程优化

KNN对特征质量非常敏感,从数据集截图看,可能存在特征冗余或区分度不足的问题:

  • 计算特征间的相关性,剔除高度相关的特征(比如皮尔逊相关系数大于0.8的特征)
  • 用特征选择方法(比如sklearn.feature_selection.SelectKBest)筛选对涨跌预测更有价值的特征
  • 检查特征是否有缺失值、异常值,先做预处理(填充缺失值、剔除/修正异常值)

4. KNN模型调优

  • 尝试调整距离度量方式:比如把metric='manhattan'(曼哈顿距离)代替默认的欧氏距离,对异常值更鲁棒
  • 调整权重参数:设置weights='distance',让距离近的样本权重更高,可能提升精度
  • 用交叉验证评估模型:使用sklearn.model_selection.cross_val_score做5折或10折交叉验证,避免单次训练测试划分的随机性导致的结果偏差

5. Excel与Colab数据传递方案

别直接复制粘贴,容易丢数据或格式出错:

  • 在Colab左侧面板点击「文件」→「上传到会话存储」,直接上传Excel文件,然后用pd.read_excel('文件名.xlsx')读取
  • 或者把Excel另存为CSV格式,上传后用pd.read_csv('文件名.csv')读取,兼容性更好

内容的提问来源于stack exchange,提问作者Wild Cat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 03:45:32