Snakemake工作流中设置Shell输出目录遇语法错误求助
问题分析
报错的核心原因是规则bismark_cov的所有输出文件没有共享相同的通配符集合:
- 第一个输出
results/{sample}/{sample}.{suffix}包含{sample}和{suffix}两个通配符 - 第二个输出
results/{sample}/{sample}_splitting_report.txt仅包含{sample} - 第三个输出
results/{sample}/{C_context}_context_{sample}.txt包含{sample}和{C_context}
Snakemake要求同一个规则的所有输出必须具备完全一致的通配符,否则无法保证每个作业的输出文件唯一,会引发文件写入冲突风险。
修复步骤
1. 修正输入匹配逻辑
原规则的input部分用wildcards.bam_paths是错误的,因为wildcards中不存在bam_paths这个变量。应根据{sample}通配符匹配对应的BAM文件:
input: bam_path=lambda wildcards: next(f for f in bam_paths if os.path.splitext(os.path.basename(f))[0] == wildcards.sample)
2. 统一输出通配符
将规则的通配符限定为{sample},通过expand生成所有该样本对应的输出文件,确保所有输出都仅依赖{sample}:
output: expand('results/{sample}/{sample}.{suffix}', suffix=suffixes), 'results/{sample}/{sample}_splitting_report.txt', expand('results/{sample}/{C_context}_context_{sample}.txt', C_context=contexts)
3. 确保输出目录存在
在Shell命令前添加目录创建命令,避免因目录不存在导致写入失败:
mkdir -p {params.out_dir}
完整修复后的规则代码
rule bismark_cov: input: bam_path=lambda wildcards: next(f for f in bam_paths if os.path.splitext(os.path.basename(f))[0] == wildcards.sample) output: expand('results/{sample}/{sample}.{suffix}', suffix=suffixes), 'results/{sample}/{sample}_splitting_report.txt', expand('results/{sample}/{C_context}_context_{sample}.txt', C_context=contexts) params: out_dir='results/{sample}' shell: """ mkdir -p {params.out_dir} bismark_methylation_extractor {input.bam_path} --parallel 4 \ --paired-end --comprehensive \ --bedGraph --zero_based --output_dir {params.out_dir} """
额外说明
- 原代码中
SAMPLES的生成逻辑是正确的,确保每个BAM文件对应唯一样本名 rule all的输入无需修改,因为它已经通过expand覆盖了所有样本的输出文件
内容的提问来源于stack exchange,提问作者sahuno
相关产品推荐
相关产品推荐

