如何恢复/重新提交卡住的SGE作业?qsub与qconf配置求助
Hey there, let's work through your SGE problems step by step—these are really common pain points when dealing with AWS Spot instances or unstable nodes, so I've got practical fixes for you.
There are a few built-in SGE features and custom tools you can use to handle this:
Use SGE's Built-in Rerun Flag
The simplest fix is to explicitly tell SGE your job can be rerun if the node fails. Add the -r y flag when submitting your job:
qsub -r y your_job_script.sh
This flag lets SGE automatically resubmit the job if the underlying node crashes or gets terminated by AWS. You can also set this as a default for all jobs in your queue (see below).
Configure Queue-Level Auto-Rerun
If you want every job in a queue to allow reruns by default, edit the queue's configuration with your root access:
- First, list your active queues to get the target queue name:
qconf -sql - Open the queue's settings for editing:
qconf -mq <your_queue_name> - Find the
rerunparameter in the editor, set it toTrue, then save and exit. Now any job submitted to this queue will automatically be eligible for rerun on node failure.
Custom Monitoring Script for Advanced Control
If you need more flexibility (like limiting rerun counts, checking specific failure reasons), write a simple bash script to monitor your jobs:
#!/bin/bash MAX_RETRIES=3 # Loop through your active jobs (adjust filters as needed) qstat -u $USER | grep -v 'job-ID' | while read line; do JOB_ID=$(echo $line | awk '{print $1}') JOB_STATE=$(echo $line | awk '{print $5}') # Check if job is in error state if [ "$JOB_STATE" = "E" ]; then # Get failure reason FAILURE_REASON=$(qstat -j $JOB_ID | grep 'failed because') if echo "$FAILURE_REASON" | grep -q 'node lost'; then # Track retry count via job comment RETRY_COUNT=$(qstat -j $JOB_ID | grep 'comment' | awk -F'retry=' '{print $2}' | awk '{print $1}') RETRY_COUNT=${RETRY_COUNT:-0} if [ $RETRY_COUNT -lt $MAX_RETRIES ]; then echo "Resubmitting job $JOB_ID (retry $((RETRY_COUNT+1)) of $MAX_RETRIES)" # Resubmit with updated retry count in comment qsub -r y -N $(qstat -j $JOB_ID | grep 'job_name' | awk '{print $3}') -cwd -v RETRY=$((RETRY_COUNT+1)) your_job_script.sh # Clean up the failed job qdel $JOB_ID fi fi fi done
Run this script periodically via cron to automatically check and resubmit eligible failed jobs.
Even if you've set SGE_ROOT, there are a few common reasons this error pops up when using qconf as root:
Ensure SGE_ROOT is Exported in Root's Shell
Environment variables often don't carry over when switching to root (e.g., using su instead of su -). Log in as root directly, or run these commands to set the variables temporarily:
export SGE_ROOT=/path/to/your/sge_installation export PATH=$SGE_ROOT/bin:$SGE_ROOT/bin/$(uname -s | tr '[:upper:]' '[:lower:]')-$(uname -m | sed 's/x86_64/amd64/'):$PATH
To make this permanent, add those lines to /root/.bashrc or /etc/profile so they load every time root logs in.
Use sudo -E to Preserve Environment Variables
If you're running qconf via sudo from a regular user, use the -E flag to keep your existing environment variables (including SGE_ROOT):
sudo -E qconf -mq <your_queue_name>
Verify SGE Configuration File Paths
Check that the core SGE config points to the correct path. Open $SGE_ROOT/default/common/settings.sh and confirm the SGE_ROOT line matches your actual installation directory. If it's wrong, edit it as root.
Check Permissions on SGE_ROOT Directory
Ensure root has full read/write access to the entire SGE installation:
chown -R root:root $SGE_ROOT chmod -R 755 $SGE_ROOT
内容的提问来源于stack exchange,提问作者Shubham Meshram

