Snakemake通配符问题:filt_100k提前执行,touch()修复遇错
问题背景
SPOT-BGC生物信息学流程中,某数据集出现filt_100k聚合规则在前置bbnorm_pe/bbnorm_se步骤未完成时提前执行的异常。尝试用touch()标记任务完成时触发两类错误:
RuleException in file ".../Snakefile", line 325:
Not all output, log and benchmark files of rule bbnorm_pe contain the same wildcards. This is crucial though, in order to avoid that two or more jobs write to the same file.
InputFunctionException in rule bbnorm_pe in file ".../Snakefile", line 325:
Error:
AttributeError: 'Wildcards' object has no attribute 'cohort_mapped_with_sample'
核心原因
- 通配符不匹配:
bbnorm_pe使用{cohort_mapped_with_sample}通配符,而filt_100k的expand()调用用的是{cohort_with_sample},导致Snakemake无法正确解析依赖关系,误判前置任务完成。 touch()使用不当:添加touch()输出时,未确保规则的所有输出、日志文件使用相同通配符集合,触发Snakemake的安全校验。
解决方案
步骤1:统一通配符命名
将bbnorm_pe的通配符与bbnorm_se统一为{cohort_with_sample},消除依赖解析歧义:
rule bbnorm_pe: input: mapped_fq1 = lambda wildcards: fq_R1_names_df.loc[wildcards.cohort_with_sample, 'Location_NonHuman'], mapped_fq2 = lambda wildcards: fq_R2_names_df.loc[wildcards.cohort_with_sample, 'Location_NonHuman'] output: input_kmers = "results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_NON-human_map_input_kmers.png", output_kmers = "results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_NON-human_map_output_kmers.png", normalized_fq1 = "results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_norm.1.fq", normalized_fq2 = "results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_norm.2.fq" log: "logs/BBnorm/{cohort_with_sample}_read_mapping_log.txt" # 线程、资源、参数、shell命令保持不变
步骤2:修复filt_100k的依赖关系
无需额外touch()文件,直接让filt_100k的输入严格对应bbnorm规则的输出,确保Snakemake能追踪所有前置任务:
rule filt_100k: input: pe_fq1 = expand("results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_norm.1.fq", cohort_with_sample=fq_R1_names_df.CohortSample), pe_fq2 = expand("results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_norm.2.fq", cohort_with_sample=fq_R1_names_df.CohortSample), se_fq = expand("results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_norm.SE.fq", cohort_with_sample=fq_SE_names_df.CohortSample), filtering_script = 'workflow/scripts/snakemake_100k_filt.sh' log: filt_100k_log = 'logs/100k_filt.txt' output: completion_100k = 'logs/completion/100k_filt__COMPLETE.txt' shell: 'bash {input.filtering_script}'
步骤3:验证依赖逻辑
运行前执行以下命令生成依赖图,确认filt_100k与所有bbnorm任务的依赖关系正确:
snakemake --dag filt_100k | dot -Tpng > dag.png
备选方案:用touch()实现细粒度任务标记(可选)
若需更精确的任务完成标记,可给每个bbnorm规则添加touch()输出,并让filt_100k依赖这些标记文件:
- 修改
bbnorm_pe添加标记输出:
rule bbnorm_pe: # 输入、原有输出、日志等保持不变 output: input_kmers = "results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_NON-human_map_input_kmers.png", output_kmers = "results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_NON-human_map_output_kmers.png", normalized_fq1 = "results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_norm.1.fq", normalized_fq2 = "results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_norm.2.fq", done = touch("results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_bbnorm_done.txt")
- 同理给
bbnorm_se添加对应done输出 - 更新
filt_100k的输入:
rule filt_100k: input: pe_done = expand("results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_bbnorm_done.txt", cohort_with_sample=fq_R1_names_df.CohortSample), se_done = expand("results/DataNonHuman/BBNorm_Reads/{cohort_with_sample}_bbnorm_done.txt", cohort_with_sample=fq_SE_names_df.CohortSample), filtering_script = 'workflow/scripts/snakemake_100k_filt.sh' # 其余部分不变
关键注意事项
- 所有规则的通配符必须保持一致,Snakemake依赖通配符匹配构建任务依赖链
- 避免在
expand()中使用与前置规则不一致的通配符名称 - 用
snakemake --dry-run提前验证流程逻辑,排查依赖问题
内容的提问来源于stack exchange,提问作者Vi_Varga

