Snakemake流水线未生成输出:rule all提示输入文件缺失
问题分析与解决:Snakemake流水线"rule all输入全部缺失"报错
问题描述
编写了一个Snakemake流水线,目标是根据unique.barcodes.txt中的barcodes列表,生成经fastq-grep处理的文件。运行时提示"rule all的输入文件全部缺失",终端列出的目标输出文件格式正确,但shell命令未生成任何内容。
核心错误原因
- Wildcard匹配失效:
fastq_grep规则的输出中,{barcodes}与{sample}直接拼接(无分隔符),Snakemake无法正确区分这两个通配符,导致无法将rule all的目标文件与该规则的输出关联,进而判定没有规则能生成目标文件。 - 重复样本名:
getsamples()函数未对提取的样本名去重,若存在同样本的多个fastq文件,会生成重复的目标文件,干扰流水线逻辑。 - 正则表达式错误:
wildcard_constraints中的[a-z-A-Z]+$写法有误,字符集中的-未放在首尾会被解析为范围,无法正确匹配带连字符的barcode。
修正方案
修正后的代码
refseq = 'refseq.fasta' reads = ['_R1_001', '_R2_001'] def getsamples(): import glob test = glob.glob("*.fastq") samples = [] for i in test: samples.append(i.rsplit('_', 2)[0]) # 对样本名去重,避免重复目标 return list(set(samples)) def getbarcodes(): with open('unique.barcodes.txt') as file: # 过滤空行,避免无效barcode lines = [line.rstrip() for line in file if line.strip()] return lines rule all: input: # 在barcodes与sample之间添加下划线分隔,确保通配符可被正确解析 expand("grepped/{barcodes}_{sample}_R1_001.plate.fastq", barcodes=getbarcodes(), sample=getsamples()), expand("grepped/{barcodes}_{sample}_R2_001.plate.fastq", barcodes=getbarcodes(), sample=getsamples()) rule fastq_grep: input: R1 = "{sample}_R1_001.fastq", R2 = "{sample}_R2_001.fastq" output: out1 = "grepped/{barcodes}_{sample}_R1_001.plate.fastq", out2 = "grepped/{barcodes}_{sample}_R2_001.plate.fastq" wildcard_constraints: # 修正正则表达式,允许字母和连字符 barcodes="[a-zA-Z-]+$", # 限制sample不含下划线,避免匹配混淆 sample="[^_]+" shell: "fastq-grep -i '{wildcards.barcodes}' {input.R1} > {output.out1} && fastq-grep -i '{wildcards.barcodes}' {input.R2} > {output.out2}"
验证步骤
- 确保
unique.barcodes.txt中每行一个有效barcode,无空行。 - 执行
snakemake -n进行预运行检查,确认流水线能正确匹配规则与目标文件。 - 预运行无问题后,执行
snakemake启动流水线。
内容的提问来源于stack exchange,提问作者LucasCortes
相关产品推荐
相关产品推荐

