Snakemake工作流依赖异常及重跑触发问题排查求助
Snakemake工作流异常排查求助
我在使用Snakemake工作流时遇到无法解释的异常:工作流先通过resfinder规则完成宏基因组reads分类,再用filter_and_normalize_resfinder规则计算每百万转录本(TPM)。具体异常如下:
- 测试样本运行正常,但全样本分析时出现两个问题:
- 完全无关的规则触发循环依赖错误(cyclic dependency error)
- 通配符(wildcards)混乱,导致工作流期望的输入文件名异常且缺失(将自身输出文件名的一部分添加到全流程输入文件名中)
- 单独运行到归一化步骤之前的全流程一切正常,但一旦加入归一化步骤,就会触发全流程重跑,此次未出现文件名相关错误
以下是涉及的两个规则代码:
rule resfinder: input: readF = output_dir + "rqcfilter2/deinterleaved/{sample}/{sample}_1.anqdpht.fastq.gz", readR = output_dir + "rqcfilter2/deinterleaved/{sample}/{sample}_2.anqdpht.fastq.gz", output: res_temp = temp(output_dir + "resfinder/{sample}/resfinder_kma/{sample}_all.unsorted.res"), res = output_dir + "resfinder/{sample}/resfinder_kma/{sample}_all.res", log: output_dir + "logs/resfinder/{sample}.log", params: out_dir = output_dir + "resfinder/{sample}/", db_res = db_dir + config["resfinder"]["db_res"], conda: "envs/resfinder.yaml", shell: """ # Checks if databases have been built and builds them otherwise if [ ! -e {params.db_res}/all.comp.b ] then {params.db_res}/INSTALL.py fi # Runs resfinder run_resfinder.py -ifq {input.readF} {input.readR} -o {params.out_dir} --acquired -db_res {params.db_res} # Create a summary results file and sorts it array=( {params.out_dir}resfinder_kma/*.res ) {{ cat ${{array[@]:0:1}}; grep -vh "^#" ${{array[@]:1}}; }} > {output.res_temp} (head -n 1 {output.res_temp} && tail -n +2 {output.res_temp} | sort -nr -k2 -t$'\t') > {output.res} """ # Filters Resfinder results based on 90% identity, 60% coverage, calculates Transcripts per Million, and sorts based on that rule filter_and_normalize_resfinder: input: res_sorted = output_dir + "resfinder/{sample}/resfinder_kma/{sample}_all.res", ref = output_dir + "multiqc/clean/multiqc_report_data/multiqc_fastqc.txt", output: res = output_dir + "resfinder/{sample}_filtered_normalized.res", run: # Reads files ref = pd.read_csv(input.ref, sep="\t") res = pd.read_csv(input.res_sorted, sep="\t") # Filter results using Strict criteria of 60% coverage and 90% identity res = res.query('Template_Coverage >= 60 & Query_Identity >= 90') # Calculates TPM by first normalizing for gene length, and then by total sequencing depth (not just ARG reads) new_res = res new_res['TPM'] = new_res['Depth'] / (new_res['Template_length']/1000) * ((ref.query('Sample == "'+wildcards.sample+'_1.anqdpht"').iat[0, 4])/1000000) # Saves new results new_res.to_csv(output.res, sep="\t", index=False)
若需其他排查信息请告知,感谢帮助。
内容的提问来源于stack exchange,提问作者sgkion
相关产品推荐
相关产品推荐

