Snakemake运行SRR数据FastQC时出现MissingRuleException如何解决
报错原因
- 执行命令语法错误
你运行的snakemake -np Snakefile中,直接跟在参数后的Snakefile会被Snakemake识别为需要生成的目标文件,而非要调用的工作流文件。Snakemake默认会自动读取当前目录下名为Snakefile的文件,如需指定自定义路径的Snakefile需要加-s参数,你未加该参数时Snakemake会寻找生成Snakefile的规则,找不到就抛出MissingRuleException。 - Snakefile本身存在多处语法/逻辑错误,即使命令修正也无法正常运行:
- 通配符写法错误:rule fastq的output中不需要加
wildcards.前缀,直接写通配符名即可,你的写法{wildcards.sample_name}属于语法错误。 - 通配符命名不匹配:rule all的输入用的通配符名为
sample,值为SRR编号,但rule fastq的output中写的是sample_name,通配符名不统一无法匹配。 - 函数缩进错误:
get_fastq函数中fq=units["fq1"]一行缩进不符合Python语法,函数执行会报错。 - 输入函数逻辑错误:你将samples的索引设为了sample_name(A/B/C等标识),但
get_fastq尝试用SRR编号作为索引取值,会触发KeyError导致输入函数执行失败。
修复方案
1. 修正执行命令
如果Snakefile在当前工作目录,直接运行:
snakemake -np
如果Snakefile不在当前工作目录,使用-s参数指定路径:
snakemake -np -s /path/to/your/Snakefile
2. 修正Snakefile代码
修正后的完整代码如下:
import os import pandas as pd configfile: "config.yaml" samples = pd.read_csv(config["samples"], sep=",").set_index("sample_name", drop=False) def get_fastq(wildcards): # wildcards.sample对应输出中的{sample}通配符,本身就是SRR编号,可直接拼接路径 return os.path.join(config["path"], f"{wildcards.sample}.fastq.gz") rule all: input: expand(os.path.join(config["path"], "fastq_output/{sample}_fastqc.html"), sample=samples["fq1"].to_list()), expand(os.path.join(config["path"], "fastq_output/{sample}_fastqc.zip"), sample=samples["fq1"].to_list()) rule fastq: input: get_fastq, output: zip = os.path.join(config["path"], "fastq_output/{sample}_fastqc.zip"), html = os.path.join(config["path"], "fastq_output/{sample}_fastqc.html") wrapper: "0.78.0/bio/fastqc"
核心修改点:
- 修正
get_fastq的缩进和逻辑错误,不需要多余的查表和expand操作 - 统一所有位置的通配符名为
sample,和rule all中传入的SRR编号匹配 - 移除output中多余的
wildcards.前缀,符合Snakemake语法规范
内容的提问来源于stack exchange,提问作者olahs
相关产品推荐
相关产品推荐

