使用DeepSpeed训练模型时执行数步后自动中断的问题求助
DeepSpeed训练仅执行几步就停止(Autotuning完成后终止)
问题重现
采用标准PyTorch训练循环初始化DeepSpeed引擎:
model_engine, optimizer, _, _ = deepspeed.initialize( model=model, model_parameters=model.parameters(), config_params=train_params['DEEPSPEED_CONFIG'], )调用
model_engine.backward(loss)和model_engine.step()执行训练,但最多运行4步就自动停止,日志输出如下:[2023-05-13 00:01:23,669] [INFO] [logging.py:96:log_dist] [Rank 0] Wrote metrics to /(local_directory)/autotuning_metric.json, /(local_directory)/autotuning_metric.json Autotuning: done with running current ds config.已通过
nvidia-smi确认无内存耗尽情况,终止通常发生在第5-6次model_engine.step()调用时。
解决方案
修改DeepSpeed自动调优配置
DeepSpeed的autotuning模块默认会在完成少量调优步骤后终止训练。解决核心是在配置中明确设置调优完成后继续训练:
在你的DEEPSPEED_CONFIG中添加或更新autotuning字段:"autotuning": { "enabled": true, "stop_after_autotune": false, // 关键参数:设为false避免调优后终止 "train_batch_size": 你的训练批次大小, "max_trials": 3 // 可根据需求调整调优尝试次数 }排查训练循环的终止逻辑
检查训练代码中是否存在误触发的终止条件,比如是否在autotuning阶段结束后,代码错误地跳出了训练循环。例如:for epoch in range(num_epochs): for step, batch in enumerate(dataloader): # 训练步骤 outputs = model_engine(batch) loss = loss_fn(outputs, labels) model_engine.backward(loss) model_engine.step() # 确认此处没有错误的break/return逻辑临时禁用自动调优验证问题
如果不需要自动调优功能,直接在配置中关闭它,验证训练是否能正常持续运行:"autotuning": { "enabled": false }若关闭后训练恢复正常,说明问题完全由autotuning配置导致。
升级DeepSpeed版本
旧版本DeepSpeed的autotuning模块可能存在逻辑bug,导致调优后误终止训练。执行以下命令升级到最新稳定版:pip install --upgrade deepspeed
内容的提问来源于stack exchange,提问作者nlp4892
相关产品推荐
相关产品推荐

