Snakemake中无法正常使用星号作为glob模式的问题排查
问题:Snakemake规则input中使用glob星号模式报错
你遇到的问题不是规则文件错误,也不是Snakemake新版本禁用功能,核心原因是:Snakemake的input字段不会自动解析shell的glob通配符(*),它会把*当作文件名的普通字符处理,所以才会提示找不到per_barcode/sample_*_rep_GGG.txt这个字面意义的文件。
为什么shell里用*能正常运行?
当把*写在shell命令中时,这条命令是交给系统shell执行的,shell会自动展开glob通配符匹配对应文件,所以能正常工作。但这种方式的缺点是Snakemake无法提前检查输入文件是否存在,依赖解析不严谨。
正确的解决方案:在input中用glob函数展开
要让Snakemake正确识别glob匹配的输入文件,需要用Python的glob函数主动展开,修改test_glob规则如下:
from glob import glob barcodes = ["AAA", "GGG", "TTT"] rule end: input: "output/merged.txt" rule test_glob: input: lambda wildcards: glob(f"per_barcode/sample_*_rep_{wildcards.barcode}.txt") output: "output/{barcode}.txt" shell: "head {input} > {output}" rule merge: input: expand("output/{barcode}.txt", barcode = barcodes) output: "output/merged.txt" shell: "cat {input} > {output}"
这里通过lambda函数结合glob,针对每个barcode通配符动态匹配对应的输入文件,Snakemake就能正确解析依赖关系,同时也能提前检查文件是否存在。
内容的提问来源于stack exchange,提问作者Xiaokang
相关产品推荐
相关产品推荐

