如何高效重新提交失败的Condor作业?无需逐个手动操作
批量重新提交失败Condor作业的高效方法
方法1:利用condor_history筛选历史失败作业并批量提交
Condor的condor_history命令可直接查询已完成(含失败)的作业,通过约束条件精准筛选目标作业,再结合批量命令快速处理:
- 先单独执行筛选命令,确认目标作业ID:
(注:Condor中condor_history -constraint "JobStatus == 5" -format "%d\n" ClusterIdJobStatus == 5对应失败状态,若需筛选被移除的作业可改用JobStatus == 3) - 直接批量重新提交:
condor_history -constraint "JobStatus == 5" -format "%d\n" ClusterId | xargs -I {} condor_submit -cluster {}
方法2:针对队列中未清理的失败作业
若失败作业仍在condor_q队列中未被清理,可改用condor_q筛选提交:
condor_q -constraint "JobStatus == 5" -format "%d\n" ClusterId | xargs -I {} condor_submit -cluster {}
方法3:精细化筛选(如最近N小时失败的作业)
结合时间约束,只重新提交指定时间段内的失败作业,避免重复提交历史过久的作业:
# 筛选过去24小时内失败的作业并提交 condor_history -constraint "(JobStatus == 5) && (JobCurrentStartDate > $(date -d '24 hours ago' +%s))" -format "%d\n" ClusterId | xargs -I {} condor_submit -cluster {}
额外优化
- 可在约束条件中加入作业属性缩小范围,比如指定提交者、作业名:
condor_history -constraint "(Owner == 'your_username') && (JobStatus == 5) && (JobName == 'your_job_name')" -format "%d\n" ClusterId | xargs -I {} condor_submit -cluster {} - 操作前先单独执行筛选部分,确认输出的作业ID无误后,再执行完整提交命令,避免误操作。
内容的提问来源于stack exchange,提问作者StackExchanger
相关产品推荐
相关产品推荐

