使用Snakemake Checkpoint恢复多通配符未知输出时遇报错求助
Snakemake Checkpoint多通配符场景下的输出恢复问题解决办法
问题背景
使用Snakemake Checkpoint功能恢复预执行未知的输出文件时,当checkpoint输出路径包含{sample}这类多通配符,聚合规则的输入函数会报Missing wildcard values for sample错误,原测试代码如下:
SAMPLE = ["A", "B", "C", "D", "E"] # a target rule to define the desired final output rule all: input: "aggregated.txt" # the checkpoint that shall trigger re-evaluation of the DAG # an number of file is created in a defined directory checkpoint somestep: output: directory("results/{sample}") shell: ''' mkdir -p results mkdir results/{wildcards.sample} # 修正原代码拼写错误mkidr→mkdir cd results/{wildcards.sample} for i in 1 2 3; do touch $i.txt; done ''' # input function for rule aggregate, return paths to all files produced by the checkpoint 'somestep' def aggregate_input(wildcards): checkpoint_output = checkpoints.somestep.get(**wildcards).output[0] return expand("results/{sample}/{i}.txt", sample=wildcards.sample, i=glob_wildcards(os.path.join(checkpoint_output, "{i}.txt")).i) rule aggregate: input: aggregate_input output: "aggregated.txt" shell: "cat {input} > {output}"
错误原因
规则aggregate本身没有定义任何通配符,因此调用aggregate_input时,wildcards对象是空的,无法获取sample值;而checkpointsomestep是针对每个sample实例化运行的,需要先收集所有checkpoint的输出结果再处理。
解决方案
修改aggregate_input函数,不依赖单个sample通配符,而是遍历所有checkpoint运行产生的输出目录,逐个解析并收集文件路径:
import os from snakemake.io import expand, glob_wildcards SAMPLE = ["A", "B", "C", "D", "E"] rule all: input: "aggregated.txt" checkpoint somestep: output: directory("results/{sample}") shell: ''' mkdir -p results/{wildcards.sample} cd results/{wildcards.sample} for i in 1 2 3; do touch $i.txt; done ''' def aggregate_input(wildcards): # 获取所有checkpoint somestep运行产生的输出目录 all_checkpoint_dirs = checkpoints.somestep.get_output() all_files = [] for dir_path in all_checkpoint_dirs: # 从目录路径中解析出sample通配符值 sample = os.path.basename(dir_path) # 获取当前sample目录下的所有i.txt文件 i_vals = glob_wildcards(os.path.join(dir_path, "{i}.txt")).i # 生成完整文件路径并添加到列表 all_files.extend(expand("results/{sample}/{i}.txt", sample=sample, i=i_vals)) return all_files rule aggregate: input: aggregate_input output: "aggregated.txt" shell: "cat {input} > {output}"
修改说明
- 用
checkpoints.somestep.get_output()替代get(**wildcards),获取所有checkpoint实例的输出目录 - 通过
os.path.basename(dir_path)从输出目录路径中提取sample值(因为目录名就是sample值) - 遍历每个sample的目录,收集对应
i.txt文件的路径,最终返回所有文件的列表给聚合规则
验证运行
执行snakemake -j 5后,所有sample下的1.txt/2.txt/3.txt会被正确聚合到aggregated.txt中,无通配符缺失错误。
内容的提问来源于stack exchange,提问作者Ollie White
相关产品推荐
相关产品推荐

