You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用DeepSpeed训练模型时执行数步后自动中断的问题求助

DeepSpeed训练仅执行几步就停止(Autotuning完成后终止)

问题重现

采用标准PyTorch训练循环初始化DeepSpeed引擎:

model_engine, optimizer, _, _ = deepspeed.initialize(
    model=model, 
    model_parameters=model.parameters(),
    config_params=train_params['DEEPSPEED_CONFIG'],
)

调用model_engine.backward(loss)和model_engine.step()执行训练,但最多运行4步就自动停止,日志输出如下:

[2023-05-13 00:01:23,669] [INFO] [logging.py:96:log_dist] [Rank 0] Wrote metrics to /(local_directory)/autotuning_metric.json, /(local_directory)/autotuning_metric.json
Autotuning: done with running current ds config.

已通过nvidia-smi确认无内存耗尽情况,终止通常发生在第5-6次model_engine.step()调用时。

解决方案

  • 修改DeepSpeed自动调优配置
    DeepSpeed的autotuning模块默认会在完成少量调优步骤后终止训练。解决核心是在配置中明确设置调优完成后继续训练:
    在你的DEEPSPEED_CONFIG中添加或更新autotuning字段:

    "autotuning": {
        "enabled": true,
        "stop_after_autotune": false,  // 关键参数:设为false避免调优后终止
        "train_batch_size": 你的训练批次大小,
        "max_trials": 3  // 可根据需求调整调优尝试次数
    }
    
  • 排查训练循环的终止逻辑
    检查训练代码中是否存在误触发的终止条件,比如是否在autotuning阶段结束后,代码错误地跳出了训练循环。例如:

    for epoch in range(num_epochs):
        for step, batch in enumerate(dataloader):
            # 训练步骤
            outputs = model_engine(batch)
            loss = loss_fn(outputs, labels)
            model_engine.backward(loss)
            model_engine.step()
            # 确认此处没有错误的break/return逻辑
    
  • 临时禁用自动调优验证问题
    如果不需要自动调优功能,直接在配置中关闭它,验证训练是否能正常持续运行:

    "autotuning": {
        "enabled": false
    }
    

    若关闭后训练恢复正常,说明问题完全由autotuning配置导致。

  • 升级DeepSpeed版本
    旧版本DeepSpeed的autotuning模块可能存在逻辑bug,导致调优后误终止训练。执行以下命令升级到最新稳定版:

    pip install --upgrade deepspeed
    

内容的提问来源于stack exchange,提问作者nlp4892

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 00:05:23