Snakemake中MissingInputException报错:批量处理文件失败排查
Snakemake glob_wildcards后出现MissingInputException的问题
问题场景
编写Snakefile读取resources/bin目录下的所有.fasta文件,在results/cctyper目录生成对应同名子目录,代码如下:
strain = glob_wildcards("resources/bin/{isolate}.fasta") rule all: input: expand("results/cctyper/{isolate}/", isolate=strain) rule cctyper: input: "resources/bin/{isolate}.fasta" output: directory("results/cctyper/{isolate}/") conda: "cctyper" shell: """ mkdir {output} touch {output}/cas_operons.tab """
(shell部分为替代代码,因平台问题无法执行原命令)
运行时持续报错:
MissingInputException in rule cctyper in file /Users/cloudy/PRACStest/workflow/Snakefile, line 7: Missing input files for rule cctyper: output: results/cctyper/['SAMEA5061681', 'SAMEA5061682'] wildcards: isolate=['SAMEA5061681', 'SAMEA5061682'] affected files: resources/bin/['SAMEA5061681', 'SAMEA5061682'].fasta
原因分析
glob_wildcards返回的是命名元组,而非直接的字符串列表。直接将strain传给expand时,会把整个元组作为单个参数传递,导致expand生成的路径中,{isolate}被替换为整个样本名列表的字符串形式(如['SAMEA5061681', 'SAMEA5061682']),进而触发规则时,通配符isolate被错误赋值为列表,找不到对应的resources/bin/['SAMEA5061681', 'SAMEA5061682'].fasta文件。
解决方案
需要提取命名元组中对应通配符的列表,即调用strain.isolate(通配符名为isolate),修改后的代码如下:
strain = glob_wildcards("resources/bin/{isolate}.fasta") rule all: input: expand("results/cctyper/{isolate}/", isolate=strain.isolate) rule cctyper: input: "resources/bin/{isolate}.fasta" output: directory("results/cctyper/{isolate}/") conda: "cctyper" shell: """ mkdir -p {output} # 加-p参数避免目录已存在时报错 touch {output}/cas_operons.tab """
另外,建议给mkdir加上-p参数,避免因目录已存在导致的报错。
内容的提问来源于stack exchange,提问作者cloudy
相关产品推荐
相关产品推荐

