You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效重新提交失败的Condor作业?无需逐个手动操作

批量重新提交失败Condor作业的高效方法

方法1:利用condor_history筛选历史失败作业并批量提交

Condor的condor_history命令可直接查询已完成(含失败)的作业,通过约束条件精准筛选目标作业,再结合批量命令快速处理:

  • 先单独执行筛选命令,确认目标作业ID:
    condor_history -constraint "JobStatus == 5" -format "%d\n" ClusterId
    
    (注:Condor中JobStatus == 5对应失败状态,若需筛选被移除的作业可改用JobStatus == 3)
  • 直接批量重新提交:
    condor_history -constraint "JobStatus == 5" -format "%d\n" ClusterId | xargs -I {} condor_submit -cluster {}
    

方法2:针对队列中未清理的失败作业

若失败作业仍在condor_q队列中未被清理,可改用condor_q筛选提交:

condor_q -constraint "JobStatus == 5" -format "%d\n" ClusterId | xargs -I {} condor_submit -cluster {}

方法3:精细化筛选(如最近N小时失败的作业)

结合时间约束,只重新提交指定时间段内的失败作业,避免重复提交历史过久的作业:

# 筛选过去24小时内失败的作业并提交
condor_history -constraint "(JobStatus == 5) && (JobCurrentStartDate > $(date -d '24 hours ago' +%s))" -format "%d\n" ClusterId | xargs -I {} condor_submit -cluster {}

额外优化

  • 可在约束条件中加入作业属性缩小范围,比如指定提交者、作业名:
    condor_history -constraint "(Owner == 'your_username') && (JobStatus == 5) && (JobName == 'your_job_name')" -format "%d\n" ClusterId | xargs -I {} condor_submit -cluster {}
    
  • 操作前先单独执行筛选部分,确认输出的作业ID无误后,再执行完整提交命令,避免误操作。

内容的提问来源于stack exchange,提问作者StackExchanger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 04:42:42