You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在无调度器的Ray Tune场景下确定模型最优迭代次数

Ray Tune无调度器场景下定位验证集最优迭代节点方法

前置要求

训练函数每完成一轮迭代,必须调用tune.report()上报当前的验证集指标(如验证精度、验证损失等),Ray Tune会自动存储所有迭代的指标数据。

方法1:训练结束后从结果中提取

直接读取tune.run()返回的结果对象,获取最优试验的全量迭代记录,直接定位最优得分对应的迭代:

from ray import tune

def your_train_func(config):
    # 你的训练逻辑
    for epoch in range(config["max_epochs"]):
        # 训练+验证步骤
        train_acc = ... 
        valid_acc = ...
        # 上报指标,可同时上报训练、验证多类指标
        tune.report(train_acc=train_acc, valid_acc=valid_acc, current_epoch=epoch)

# 启动试验,未配置scheduler即无调度器模式
result_grid = tune.run(
    your_train_func,
    config={"max_epochs": 50},
    metric="valid_acc", # 配置你要优化的核心验证指标
    mode="max" # 指标越大越好填max(如精度),越小越好填min(如损失)
)

# 获取效果最优的试验
best_trial = result_grid.get_best_trial()
# 读取最优指标对应的迭代索引
best_iter = best_trial.metric_analysis["valid_acc"]["best_idx"]

print(f"验证集最优得分对应的迭代数:{best_iter}")

该best_iter就是你需要的节点:从该迭代之后如果验证集指标持续走跌/损失持续上升,即代表模型开始过拟合。

方法2:训练过程中实时判断

不需要依赖Ray Tune内置组件,直接在训练函数中自己维护最优迭代记录,还可叠加自定义早停逻辑:

def train_func(config):
    best_valid_score = float("-inf") # 损失类指标初始化为float("inf")
    best_iter = 0
    no_improve_round = 0
    early_stop_patience = 5 # 连续5轮无提升则终止训练

    for epoch in range(config["max_epochs"]):
        # 训练+验证步骤
        valid_acc = validate()

        # 上报指标
        tune.report(valid_acc=valid_acc, best_iter=best_iter, current_epoch=epoch)

        # 自行判断最优迭代
        if valid_acc > best_valid_score: # 损失类指标改为 <
            best_valid_score = valid_acc
            best_iter = epoch
            no_improve_round = 0
        else:
            no_improve_round += 1
        
        if no_improve_round >= early_stop_patience:
            break

补充判断方式

你可以把所有迭代的训练指标、验证指标导出后绘制折线图,当训练指标仍在持续优化,但验证指标开始停滞甚至反向变化时,对应的拐点就是过拟合的起始迭代节点。


内容的提问来源于stack exchange,提问作者Cat-with-a-pipe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 22:57:00