Snakemake中使用name指令时如何通过rules访问规则名称?
Snakemake规则引用问题解析
Snakemake版本
❯ snakemake --version 8.20.6
异常行为
以下Snakefile尝试通过规则的name参数值引用输出:
rule all: input: "test2" rule: name: "test", output: "test", shell: "touch {output} " rule: name: "test2", input: rules.test.output, # input: rules._rules["2"].output, output: "test2", shell: "touch {output} "
运行snakemake -n触发错误:
❯ snakemake -n WorkflowError in file /path/to/Snakefile, line 14: Rule test is not defined in this workflow. Available rules: all, 2
虽然可以通过rules._rules["2"].output访问输出,但索引"2"是Snakemake自动生成的,会随规则顺序变化,完全不可靠。
问题
能否通过rules.test.output或类似方式,按规则的业务名称可靠引用其输出?
更新尝试
我发现以下写法可以正常运行,但不确定是否存在风险:
rule all: input: "test2" rule a: name: "test", output: "test", shell: "touch {output} " rule b: name: "test2", input: rules.a.output, output: "test2", shell: "touch {output} "
运行日志如下:
Building DAG of jobs... Using shell: /usr/bin/bash Provided cores: 16 Rules claiming more threads will be scaled down. Job stats: job count ----- ------- all 1 test 1 test2 1 total 3 Select jobs to execute... Execute 1 jobs... [Sat Oct 12 14:33:20 2024] localrule test: output: test jobid: 2 reason: Missing output files: test resources: tmpdir=/tmp touch test [Sat Oct 12 14:33:20 2024] Finished job 2. 1 of 3 steps (33%) done Select jobs to execute... Execute 1 jobs... [Sat Oct 12 14:33:20 2024] localrule test2: input: test output: test2 jobid: 1 reason: Missing output files: test2; Input files updated by another job: test resources: tmpdir=/tmp touch test2 [Sat Oct 12 14:33:20 2024] Finished job 1. 2 of 3 steps (67%) done Select jobs to execute... Execute 1 jobs... [Sat Oct 12 14:33:20 2024] localrule all: input: test2 jobid: 0 reason: Input files updated by another job: test2 resources: tmpdir=/tmp [Sat Oct 12 14:33:20 2024] Finished job 0. 3 of 3 steps (100%) done Complete log: .snakemake/log/2024-10-12T143320.208718.snakemake.log
解答
核心逻辑
Snakemake的rules对象仅能通过规则定义时的标识符(即rule关键字后紧跟的名称,比如rule a:中的a)来引用规则,name参数只是规则的显示别名,不参与规则的内部引用逻辑。
对两种写法的分析
- 最初错误写法:未指定规则标识符,Snakemake自动用数字作为规则的内部标识,导致无法通过
name参数值(test)引用规则,只能依赖不可靠的自动生成索引。 - 更新后的写法:完全安全且符合规范:
- 每个规则都有明确的自定义标识符(
a、b),通过rules.a、rules.b引用稳定可靠,不受规则顺序影响 name参数仅用于日志和任务的友好显示(比如日志中显示的localrule test),不影响核心引用逻辑
- 每个规则都有明确的自定义标识符(
更优写法建议
如果不需要自定义显示名称,直接将规则标识符设为业务名称会更直观:
rule all: input: "test2" rule test: output: "test", shell: "touch {output}" rule test2: input: rules.test.output, output: "test2", shell: "touch {output}"
这种写法直接用rules.test引用,逻辑清晰,无需额外维护name参数。
内容的提问来源于stack exchange,提问作者ManJi
相关产品推荐
相关产品推荐

