Snakemake多输入文件样本处理管道报错及优化求助
问题描述
作为Snakemake新手,完成官方教程后尝试搭建处理多输入文件样本的管道,需求如下:
- 读取样本目录(如
data/experiment/sample/)下的所有输入文件,对每个文件运行Python脚本,输出保存至结果目录(如results/experiment/sample/); - 将结果目录中的所有结果文件合并为单个文件(如
results/experiment/sample/merged.txt)。
目录结构
data experiment Sample_1 Sample_1_0.txt Sample_1_1.txt Sample_1_2.txt
当前配置文件 (config.yaml)
exp_dir: experiment sample: Sample_1 distance: 450 max_wavelength: 710 min_wavelength: 575
编写的Snakefile
configfile: "config.yaml" EXP_DIR=config["exp_dir"] SAMPLE=config["sample"] FINAL_FILE=f"results/{EXP_DIR}/{SAMPLE}/{SAMPLE}_complete.txt" wcs = glob_wildcards("data" + "/" + EXP_DIR + "/" + SAMPLE + "/" + SAMPLE + "_{rep}.txt") rule all: input: FINAL_FILE rule detect_peaks: input: "data" + "/" + EXP_DIR + "/" + SAMPLE + "/" + SAMPLE + "_{rep}.txt" output: "results/" + EXP_DIR + "/" + SAMPLE + "/" + SAMPLE + "_{rep}.txt" params: distance=config["distance"], max_wavelength=config["max_wavelength"], min_wavelength=config["min_wavelength"] shell: "python scripts/peak_detection.py -i {input} -o {output} -d {params.distance} " "-m {params.min_wavelength} -x {params.max_wavelength}" rule agg_files: input: expand("data/results" + "/" + EXP_DIR + "/" + SAMPLE + "/" + SAMPLE + "_{rep}.txt", ts=wcs["rep"]) output: FINAL_FILE output: "python scripts/combine_peaks.py -i {input} -o {output}"
报错信息
TypeError in file XXXX, line 29: tuple indices must be integers or slices, not str
错误修复与优化建议
1. 核心报错修复
报错根源在agg_files规则,存在三个问题:
- 路径错误:
expand中写了data/results,正确路径应为results(无需data/前缀); - 通配符变量不匹配:
expand里用了ts=wcs["rep"],但通配符定义的是rep,需改为rep=wcs.rep(或rep=wcs["rep"]); - 重复字段:规则里写了两次
output:,第二个output:应改为shell:。
修正后的agg_files规则:
rule agg_files: input: expand("results/{exp_dir}/{sample}/{sample}_{rep}.txt", exp_dir=EXP_DIR, sample=SAMPLE, rep=wcs.rep) output: FINAL_FILE shell: "python scripts/combine_peaks.py -i {input} -o {output}"
2. 代码优化建议
(1)用f-string简化路径拼接
原代码的路径拼接冗余,可改用f-string简化,注意通配符的大括号需要转义({{rep}}):
# 替换glob_wildcards行 wcs = glob_wildcards(f"data/{EXP_DIR}/{SAMPLE}/{SAMPLE}_{{rep}}.txt") # 替换detect_peaks的input/output rule detect_peaks: input: f"data/{EXP_DIR}/{SAMPLE}/{SAMPLE}_{{rep}}.txt" output: f"results/{EXP_DIR}/{SAMPLE}/{SAMPLE}_{{rep}}.txt"
(2)扩展为多样本支持(可选)
如果后续需要处理多个样本,修改config.yaml为列表形式:
exp_dir: experiment samples: - Sample_1 - Sample_2 distance: 450 max_wavelength: 710 min_wavelength: 575
然后调整Snakefile实现批量处理:
configfile: "config.yaml" EXP_DIR=config["exp_dir"] SAMPLES=config["samples"] # 生成所有样本的最终输出文件列表 FINAL_FILES = [f"results/{EXP_DIR}/{sample}/{sample}_complete.txt" for sample in SAMPLES] rule all: input: FINAL_FILES # 动态获取单个样本的所有重复文件 def get_sample_reps(wildcards): return glob_wildcards(f"data/{EXP_DIR}/{wildcards.sample}/{wildcards.sample}_{{rep}}.txt").rep rule detect_peaks: input: f"data/{EXP_DIR}/{wildcards.sample}/{wildcards.sample}_{{rep}}.txt" output: f"results/{EXP_DIR}/{wildcards.sample}/{wildcards.sample}_{{rep}}.txt" params: distance=config["distance"], max_wavelength=config["max_wavelength"], min_wavelength=config["min_wavelength"] shell: "mkdir -p $(dirname {output}) && python scripts/peak_detection.py -i {input} -o {output} -d {params.distance} " "-m {params.min_wavelength} -x {params.max_wavelength}" rule agg_files: input: expand(f"results/{EXP_DIR}/{wildcards.sample}/{wildcards.sample}_{{rep}}.txt", rep=get_sample_reps(wildcards)) output: f"results/{EXP_DIR}/{wildcards.sample}/{wildcards.sample}_complete.txt" shell: "python scripts/combine_peaks.py -i {input} -o {output}"
(3)自动创建结果目录
在detect_peaks的shell命令前添加mkdir -p $(dirname {output}),确保结果目录不存在时自动创建,避免报错。
内容的提问来源于stack exchange,提问作者Brian M.
相关产品推荐
相关产品推荐

