You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在sklearn验证曲线上标记指定数据点以查看X、Y值?

解决验证曲线上标记最优数据点的问题

首先,plot_validation_curve的返回值包含了绘图所需的所有原始数据,你可以通过提取这些数据来定位并标记目标点,具体步骤如下:

  1. 拆分返回值,提取核心数据
    plot_validation_curve返回一个元组(ax, param_name, train_scores, test_scores),其中:
  • ax:绘图的Matplotlib轴对象
  • train_scores/test_scores:每个参数对应的交叉验证得分(形状为(len(param_range), cv折数))
  • param_range:你传入的参数取值范围(即np.arange(2,100,2))

先拆分返回值并计算得分的均值(因为每个参数对应多折CV的得分,我们通常看均值):

import numpy as np
from sklearn.model_selection import plot_validation_curve, StratifiedKFold

# 你的原始绘图代码
valid1 = plot_validation_curve(rand_search.best_estimator_, X_train, y_train,
                               cv=StratifiedKFold(n_splits=5), param_range=np.arange(2,100,2),
                               param_name='max_depth', scoring='f1')

# 拆分返回值
ax, param_name, train_scores, test_scores = valid1
param_range = np.arange(2,100,2)  # 和你传入的参数范围一致

# 计算训练集和测试集的得分均值
train_mean = np.mean(train_scores, axis=1)
test_mean = np.mean(test_scores, axis=1)
  1. 定位最优点的索引
    两种方式定位目标点:
  • 如果你已经确定了具体的max_depth值(比如你认为20是最优):
    target_depth = 20
    idx = np.where(param_range == target_depth)[0][0]
    
  • 如果要找测试集F1得分最高的点:
    idx = np.argmax(test_mean)
    target_depth = param_range[idx]
    
  1. 标记并标注数据点
    使用Matplotlib的scatter标记点,annotate显示具体数值:
# 获取目标点的得分
best_score = test_mean[idx]

# 用红色星号标记最优点
ax.scatter(target_depth, best_score, color='red', s=120, marker='*', label='最优参数点')

# 添加数值标注,可根据实际情况调整xytext的位置避免遮挡
ax.annotate(f'max_depth={target_depth}\nF1得分={best_score:.3f}',
            xy=(target_depth, best_score),
            xytext=(target_depth + 6, best_score + 0.03),
            arrowprops=dict(arrowstyle='->', color='darkred'),
            fontsize=10)

# 显示图例
ax.legend()

这样就能在你的验证曲线上清晰标记出选定的最优点,同时显示对应的max_depth和F1得分。

内容的提问来源于stack exchange,提问作者fan-yang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 07:52:42