You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行GitHub上的MOTRv2时遇subprocess.CalledProcessError(SIGSEGV:11)错误求助

MOTRv2训练触发SIGSEGV错误求助

执行命令

运行训练脚本:

(motrv2) dcy@ubuntu:~/MOTRv2$ ./tools/train.sh configs/motrv2.args

实际执行的训练命令:

python -m torch.distributed.launch --nproc_per_node=8 --use_env main.py --meta_arch motr --dataset_file e2e_dance --epoch 5 --with_box_refine --lr_drop 4 --lr 2e-4 --lr_backbone 2e-5 --pretrained mot/r50_deformable_detr_plus_iterative_bbox_refinement-checkpoint.pth --batch_size 1 --sample_mode random_interval --sample_interval 10 --sampler_lengths 5 --merger_dropout 0 --dropout 0 --random_drop 0.1 --fp_ratio 0.3 --query_interaction_layer QIMv2 --query_denoise 0.05 --num_queries 10 --append_crowd --det_db det_db_motrv2.json --use_checkpoint --output_dir .

错误信息

分布式初始化完成后抛出:

subprocess.CalledProcessError: Command '['/data1/dcy/anaconda3/envs/motrv2/bin/python', '-u', 'main.py', '--meta_arch', 'motr', '--dataset_file', 'e2e_dance', '--epoch', '5', '--with_box_refine', '--lr_drop', '4', '--lr', '2e-4', '--lr_backbone', '2e-5', '--pretrained', 'mot/r50_deformable_detr_plus_iterative_bbox_refinement-checkpoint.pth', '--batch_size', '1', '--sample_mode', 'random_interval', '--sample_interval', '10', '--sampler_lengths', '5', '--merger_dropout', '0', '--dropout', '0', '--random_drop', '0.1', '--fp_ratio', '0.3', '--query_interaction_layer', 'QIMv2', '--query_denoise', '0.05', '--num_queries', '10', '--append_crowd', '--det_db', 'det_db_motrv2.json', '--use_checkpoint', '--output_dir', '.']' died with <Signals.SIGSEGV: 11>.

运行环境

  • Ubuntu 18.04
  • Python 3.7
  • PyTorch 1.7.1
  • CUDA 11.1

排查建议

  • 降低进程数:当前使用--nproc_per_node=8,若GPU显存不足或硬件存在兼容性问题,尝试改为4或2,减少单节点进程数后测试。
  • 校验预训练文件:确认mot/r50_deformable_detr_plus_iterative_bbox_refinement-checkpoint.pth文件完整无损坏,若下载过程中有中断,重新下载后重试。
  • 匹配CUDA与PyTorch版本:PyTorch 1.7.1官方推荐适配CUDA 11.0,当前用的CUDA 11.1存在版本不匹配风险,可尝试降级CUDA到11.0,或升级PyTorch到1.8+版本(兼容CUDA 11.1)。
  • 禁用梯度检查点:暂时去掉--use_checkpoint参数,该功能可能在部分环境下引发内存访问错误,测试是否能正常运行。
  • 单进程调试:直接运行单进程训练命令(去掉torch.distributed.launch相关参数),获取更详细的错误日志,定位具体出错代码行。
  • 检查数据集路径:确认det_db_motrv2.json和e2e_dance数据集路径正确,无文件缺失或权限问题。

内容的提问来源于stack exchange,提问作者wh1sper13

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 12:15:32