如何在sklearn验证曲线上标记指定数据点以查看X、Y值?
解决验证曲线上标记最优数据点的问题
首先,plot_validation_curve的返回值包含了绘图所需的所有原始数据,你可以通过提取这些数据来定位并标记目标点,具体步骤如下:
- 拆分返回值,提取核心数据
plot_validation_curve返回一个元组(ax, param_name, train_scores, test_scores),其中:
ax:绘图的Matplotlib轴对象train_scores/test_scores:每个参数对应的交叉验证得分(形状为(len(param_range), cv折数))param_range:你传入的参数取值范围(即np.arange(2,100,2))
先拆分返回值并计算得分的均值(因为每个参数对应多折CV的得分,我们通常看均值):
import numpy as np from sklearn.model_selection import plot_validation_curve, StratifiedKFold # 你的原始绘图代码 valid1 = plot_validation_curve(rand_search.best_estimator_, X_train, y_train, cv=StratifiedKFold(n_splits=5), param_range=np.arange(2,100,2), param_name='max_depth', scoring='f1') # 拆分返回值 ax, param_name, train_scores, test_scores = valid1 param_range = np.arange(2,100,2) # 和你传入的参数范围一致 # 计算训练集和测试集的得分均值 train_mean = np.mean(train_scores, axis=1) test_mean = np.mean(test_scores, axis=1)
- 定位最优点的索引
两种方式定位目标点:
- 如果你已经确定了具体的
max_depth值(比如你认为20是最优):target_depth = 20 idx = np.where(param_range == target_depth)[0][0] - 如果要找测试集F1得分最高的点:
idx = np.argmax(test_mean) target_depth = param_range[idx]
- 标记并标注数据点
使用Matplotlib的scatter标记点,annotate显示具体数值:
# 获取目标点的得分 best_score = test_mean[idx] # 用红色星号标记最优点 ax.scatter(target_depth, best_score, color='red', s=120, marker='*', label='最优参数点') # 添加数值标注,可根据实际情况调整xytext的位置避免遮挡 ax.annotate(f'max_depth={target_depth}\nF1得分={best_score:.3f}', xy=(target_depth, best_score), xytext=(target_depth + 6, best_score + 0.03), arrowprops=dict(arrowstyle='->', color='darkred'), fontsize=10) # 显示图例 ax.legend()
这样就能在你的验证曲线上清晰标记出选定的最优点,同时显示对应的max_depth和F1得分。
内容的提问来源于stack exchange,提问作者fan-yang
相关产品推荐
相关产品推荐

