You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Snakemake结合SGE集群(qsub)执行简化规则时返回非零退出状态的问题排查

Troubleshooting Non-Zero Exit Code in Snakemake + SGE Cluster Execution

Let's break down why your simplified raven_assembly rule is failing when submitted via Snakemake, even though the shell command works perfectly manually. Here are the key issues to investigate and fix:

1. The Cluster Job Runs a Nested Snakemake Process (Not Your Shell Command Directly)

Looking at the job script Snakemake generates, it doesn't execute your shell command directly—it launches a second Snakemake process to handle the rule. The non-zero exit code is almost certainly coming from this nested process, not your zcat | head command (which you confirmed produces correct output and runs without errors manually).

Fix/Check: Capture Full Cluster Job Logs

Your current out.txt and errors.txt only log output from your shell command, not the nested Snakemake process. Modify your --cluster flag to capture complete job logs from qsub:

snakemake -r -p --jobs 40 --cluster "qsub -o cluster_out_{jobid}.log -e cluster_err_{jobid}.log -cwd -V"
  • -cwd ensures qsub runs the job in your working directory (avoids path resolution issues)
  • -V passes your local environment variables to the cluster node (critical for accessing your conda environment)
  • The new log files will show errors from the nested Snakemake process, which is likely where the failure occurs.

2. Conda Environment Access on Cluster Nodes

The nested Snakemake process relies on your conda environment at /scratch2/software/anaconda/envs/proj-ergo. If cluster compute nodes can't access this path (e.g., scratch storage isn't mounted on all nodes, or permissions are restricted), the Python/Snakemake call will fail silently or throw an error.

Fix/Check:

  • SSH into a compute node and verify the environment path is accessible: ls /scratch2/software/anaconda/envs/proj-ergo
  • If the path is unavailable, move the conda environment to a shared filesystem accessible to all nodes, or use --conda-create-envs-only to pre-create the environment on the cluster.

3. Test the Job Script Manually

Save the generated job script (the one starting with #!/bin/sh) to a file like test_job.sh, then submit it directly via qsub:

qsub test_job.sh

This replicates exactly what Snakemake is doing and will reveal the root error in the cluster's native job logs.

4. Refine SGE Cluster Parameters

Your minimal --cluster "qsub" command misses key SGE requirements:

  • Add -V and -cwd as noted earlier
  • Match the memory resources specified in your job script: add -l h_vmem=12G (to align with the mem_mb=10903 in the generated script)

Putting this all together, run Snakemake with:

snakemake -r -p --jobs 40 --cluster "qsub -o cluster_out_{jobid}.log -e cluster_err_{jobid}.log -cwd -V -l h_vmem=12G"

Check the cluster_err_1.log file for the exact error from the nested Snakemake process—this will point you directly to the root cause.

内容的提问来源于stack exchange,提问作者Joran Martijn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 22:17:27