Snakemake流水线如何基于TSV文件批量生成所有行对应输出
Snakemake 遍历TSV批量运行流水线实现方案
核心逻辑:提前解析TSV中所有记录对应的最终输出文件列表,将其作为流水线总入口规则的输入,Snakemake会自动通过通配符匹配推导依赖链,逐行触发对应任务执行。
第一步:解析TSV生成目标文件列表
在Snakefile最开头添加解析代码,读取TSV内容并生成所有需要产出的最终文件路径。
方案1:使用pandas读取(代码更简洁)
import pandas as pd # 读取TSV文件,指定制表符为分隔符,第一行为表头,所有列按字符串读取避免前导零丢失 annot_table = pd.read_csv("your_annotation.tsv", sep="\t", header=0, dtype=str) # 遍历每一行生成最终输出文件名 target_outputs = [] for _, row in annot_table.iterrows(): target_outputs.append( f"ak_{row['genome']}_{row['locus_tag']}_{row['gene']}_{row['substrate']}" )
注:读取时指定dtype=str可以避免locus_tag这类带前导零的数字型内容被自动转成整数,导致格式和TSV记录不匹配
方案2:使用Python原生csv模块读取(无第三方依赖)
如果不想安装pandas,可以直接用Python标准库实现相同功能:
import csv target_outputs = [] with open("your_annotation.tsv", "r", encoding="utf-8") as f: tsv_reader = csv.DictReader(f, delimiter="\t") for row in tsv_reader: target_outputs.append( f"ak_{row['genome']}_{row['locus_tag']}_{row['gene']}_{row['substrate']}" )
第二步:添加流水线总入口规则
Snakemake会优先执行rule all,将其输入作为所有需要完成的目标,自动推导上游需要运行的任务:
rule all: input: target_outputs
第三步:配置业务处理规则
你之前编写的两个处理规则逻辑基本正确,可直接沿用,仅需注意路径和引号格式即可:
rule extractfeat: input: "/pathtofile/{genome}.gbk" output: "{genome}_{locus_tag}_{gene}_{substrate}.fasta" shell: """ extractfeat {input} {output} -type CDS -value {wildcards.genome}_{wildcards.locus_tag} -describe product,locus_tag """ rule transeq: input: "{genome}_{locus_tag}_{gene}_{substrate}.fasta" output: "ak_{genome}_{locus_tag}_{gene}_{substrate}" shell: """ transeq {input} {output} """
运行注意事项
- 将代码中的
your_annotation.tsv替换为你本地实际的TSV文件路径,同时确认/pathtofile/路径下所有对应基因组名的gbk文件存在 - 执行流水线时,在Snakefile所在目录运行命令
snakemake -j <使用的CPU线程数>即可自动处理TSV中所有行对应的任务 - 正式运行前可以加
-n参数执行空跑,检查生成的任务列表是否和TSV记录数量匹配,避免通配符匹配错误
内容的提问来源于stack exchange,提问作者aetobatus-narinari
相关产品推荐
相关产品推荐

