Snakemake结合SGE集群(qsub)执行简化规则时返回非零退出状态的问题排查
Let's break down why your simplified raven_assembly rule is failing when submitted via Snakemake, even though the shell command works perfectly manually. Here are the key issues to investigate and fix:
1. The Cluster Job Runs a Nested Snakemake Process (Not Your Shell Command Directly)
Looking at the job script Snakemake generates, it doesn't execute your shell command directly—it launches a second Snakemake process to handle the rule. The non-zero exit code is almost certainly coming from this nested process, not your zcat | head command (which you confirmed produces correct output and runs without errors manually).
Fix/Check: Capture Full Cluster Job Logs
Your current out.txt and errors.txt only log output from your shell command, not the nested Snakemake process. Modify your --cluster flag to capture complete job logs from qsub:
snakemake -r -p --jobs 40 --cluster "qsub -o cluster_out_{jobid}.log -e cluster_err_{jobid}.log -cwd -V"
-cwdensures qsub runs the job in your working directory (avoids path resolution issues)-Vpasses your local environment variables to the cluster node (critical for accessing your conda environment)- The new log files will show errors from the nested Snakemake process, which is likely where the failure occurs.
2. Conda Environment Access on Cluster Nodes
The nested Snakemake process relies on your conda environment at /scratch2/software/anaconda/envs/proj-ergo. If cluster compute nodes can't access this path (e.g., scratch storage isn't mounted on all nodes, or permissions are restricted), the Python/Snakemake call will fail silently or throw an error.
Fix/Check:
- SSH into a compute node and verify the environment path is accessible:
ls /scratch2/software/anaconda/envs/proj-ergo - If the path is unavailable, move the conda environment to a shared filesystem accessible to all nodes, or use
--conda-create-envs-onlyto pre-create the environment on the cluster.
3. Test the Job Script Manually
Save the generated job script (the one starting with #!/bin/sh) to a file like test_job.sh, then submit it directly via qsub:
qsub test_job.sh
This replicates exactly what Snakemake is doing and will reveal the root error in the cluster's native job logs.
4. Refine SGE Cluster Parameters
Your minimal --cluster "qsub" command misses key SGE requirements:
- Add
-Vand-cwdas noted earlier - Match the memory resources specified in your job script: add
-l h_vmem=12G(to align with themem_mb=10903in the generated script)
Final Recommended Command
Putting this all together, run Snakemake with:
snakemake -r -p --jobs 40 --cluster "qsub -o cluster_out_{jobid}.log -e cluster_err_{jobid}.log -cwd -V -l h_vmem=12G"
Check the cluster_err_1.log file for the exact error from the nested Snakemake process—this will point you directly to the root cause.
内容的提问来源于stack exchange,提问作者Joran Martijn

