You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch分布式训练初始化阶段出现SIGSEGV错误的原因咨询

PyTorch分布式训练初始化阶段SIGSEGV段错误排查

错误日志

ERROR:torch.distributed.elastic.multiprocessing.api:failed (exitcode: -11) local_rank: 0 (pid: 3680358) of binary: /home/lifesci/ekeys/anaconda3/envs/minigpt4-4/bin/python
Traceback (most recent call last):
  File "/home/lifesci/ekeys/anaconda3/envs/minigpt4-4/bin/torchrun", line 33, in <module>
    sys.exit(load_entry_point('torch==2.0.1', 'console_scripts', 'torchrun')())
  File "/home/lifesci/ekeys/anaconda3/envs/minigpt4-4/lib/python3.10/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 346, in wrapper
    return f(*args, **kwargs)
  File "/home/lifesci/ekeys/anaconda3/envs/minigpt4-4/lib/python3.10/site-packages/torch/distributed/run.py", line 794, in main
    run(args)
  File "/home/lifesci/ekeys/anaconda3/envs/minigpt4-4/lib/python3.10/site-packages/torch/distributed/run.py", line 785, in run
    elastic_launch(
  File "/home/lifesci/ekeys/anaconda3/envs/minigpt4-4/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 134, in __call__
    return launch_agent(self._config, self._entrypoint, list(args))
  File "/home/lifesci/ekeys/anaconda3/envs/minigpt4-4/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 250, in launch_agent
    raise ChildFailedError(
torch.distributed.elastic.multiprocessing.errors.ChildFailedError: 
=========================================================
train.py FAILED
---------------------------------------------------------
Failures:
[1]:
  time      : 2023-10-21_17:36:57
  host      : gnode10.hanhai22.scc.ustc.edu.cn
  rank      : 1 (local_rank: 1)
  exitcode  : -11 (pid: 3680359)
  error_file: <N/A>
  traceback : Signal 11 (SIGSEGV) received by PID 3680359
[2]:
  time      : 2023-10-21_17:36:57
  host      : gnode10.hanhai22.scc.ustc.edu.cn
  rank      : 2 (local_rank: 2)
  exitcode  : -11 (pid: 3680360)
  error_file: <N/A>
  traceback : Signal 11 (SIGSEGV) received by PID 3680360
[3]:
  time      : 2023-10-21_17:36:57
  host      : gnode10.hanhai22.scc.ustc.edu.cn
  rank      : 3 (local_rank: 3)
  exitcode  : -11 (pid: 3680361)
  error_file: <N/A>
  traceback : Signal 11 (SIGSEGV) received by PID 3680361
---------------------------------------------------------
Root Cause (first observed failure):
[0]:
  time      : 2023-10-21_17:36:57
  host      : gnode10.hanhai22.scc.ustc.edu.cn
  rank      : 0 (local_rank: 0)
  exitcode  : -11 (pid: 3680358)
  error_file: <N/A>
  traceback : Signal 11 (SIGSEGV) received by PID 3680358
=========================================================

错误本质与常见原因

Exitcode -11对应SIGSEGV段错误,即进程尝试访问非法内存地址。在分布式训练初始化阶段触发这类错误,常见原因包括:

  • 显存过载:多卡训练时,单卡分配的显存超过剩余容量,比如模型体积过大、batch size未按卡数缩放,初始化时瞬间占满显存导致崩溃。
  • 初始化顺序错误:代码中在调用torch.distributed.init_process_group之前就执行了GPU操作(如模型移至GPU),多进程争抢GPU资源触发内存冲突。
  • 依赖版本不兼容:即使切换了PyTorch和CUDA版本,若cuDNN、NCCL等配套库或GPU驱动与CUDA版本不匹配,底层CUDA操作会触发段错误。
  • 系统资源冲突:GPU被其他进程占用,或服务器共享内存、PCIe链路存在资源竞争,导致分布式进程无法正常初始化CUDA上下文。
  • 自定义算子问题:代码中使用的自定义CUDA算子或第三方库未做分布式兼容处理,多进程下出现内存访问错误。

排查步骤

  • 单卡验证:先以单卡模式运行脚本,确认模型本身可正常初始化,排除基础代码问题。
  • 降低显存负载:缩小batch size,或启用半精度训练(torch.float16),减少单卡显存占用后再尝试分布式。
  • 修正初始化顺序:确保先执行分布式初始化,再绑定对应GPU,示例正确流程:
    import torch.distributed as dist
    dist.init_process_group(backend='nccl')
    local_rank = dist.get_rank()
    torch.cuda.set_device(local_rank)
    model = model.to(local_rank)
    
  • 核对依赖版本:确认GPU驱动、CUDA、cuDNN、NCCL版本与当前PyTorch版本完全匹配。
  • 清理GPU资源:用nvidia-smi查看GPU占用,终止其他占用进程后重启训练。
  • 逐步扩容验证:先以torchrun --nproc_per_node=1启动单进程分布式训练,确认正常后再逐步增加进程数。

内容的提问来源于stack exchange,提问作者ekeys

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 05:34:51