如何让Snakemake为所有匹配输出仅运行一次规则?
在Snakemake中实现单规则批量处理所有样本的方法
问题背景
出于效率考虑,需要让一个Snakemake规则一次性处理所有匹配通配符的样本,而非为每个样本单独运行一次规则。比如你提供的示例代码中,produce_all_csv会被调用3次(对应3个输出文件),但实际需要它仅运行一次,调用类似如下格式的批量工具命令:
tool --inputs input_1.csv,input_2.csv --outputs output_1.csv,output_2.csv
原无法满足需求的代码:
rule all: input: "out_1.csv", "out_2.csv", "out_3.csv", rule produce_all_csv: """ This rule should be called _once_ for _all_ samples Not once per sample """ input: "in_{sample}.csv", output: "out_{sample}.csv", shell: """ # Placeholder for a real command # that takes a list of input files # and produces a list of output file """
惯用解决方法
方法1:直接定义聚合规则(无通配符)
最常用的方式是创建一个不带通配符的规则,把所有输入、输出文件作为完整列表传入,这样Snakemake只会触发该规则一次:
# 先定义样本列表,后续维护更方便 SAMPLES = ["1", "2", "3"] rule all: input: expand("out_{sample}.csv", sample=SAMPLES), rule produce_all_csv: """一次性处理所有样本的规则""" input: infiles = expand("in_{sample}.csv", sample=SAMPLES), output: outfiles = expand("out_{sample}.csv", sample=SAMPLES), shell: """ # 如果工具需要逗号分隔的路径,用','.join格式化 tool --inputs {','.join(input.infiles)} --outputs {','.join(output.outfiles)} """
方法2:使用分组通配符(适用于需要分组的场景)
如果后续需要按批次/分组处理样本,还可以用Snakemake的group通配符特性,把所有样本归为一个组,让规则按组运行:
SAMPLES = ["1", "2", "3"] # 定义一个包含所有样本的分组 ALL_GROUP = "all_samples" rule all: input: expand("out_{sample}.csv", sample=SAMPLES), rule produce_all_csv: input: infiles = expand("in_{sample}.csv", sample=SAMPLES), output: outfiles = expand("out_{sample}.csv", sample=SAMPLES), # 标记规则按分组运行 group: "{group}", params: group = ALL_GROUP, shell: """ tool --inputs {','.join(input.infiles)} --outputs {','.join(output.outfiles)} """
关键要点
- 核心逻辑是让规则的输入输出是完整的文件列表,而非单个带通配符的文件,这样Snakemake会识别这是一个独立的批量任务,而非多个分散的单样本任务。
- 用
expand()函数可以快速生成所有样本对应的文件路径,避免手动重复编写。 - 如果你的工具要求输入输出用特定分隔符(比如逗号),直接用Python的字符串拼接方法格式化即可。
内容的提问来源于stack exchange,提问作者Cornelius Roemer
相关产品推荐
相关产品推荐

