You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Snakemake工作流依赖异常及重跑触发问题排查求助

Snakemake工作流异常排查求助

我在使用Snakemake工作流时遇到无法解释的异常:工作流先通过resfinder规则完成宏基因组reads分类,再用filter_and_normalize_resfinder规则计算每百万转录本(TPM)。具体异常如下:

  • 测试样本运行正常,但全样本分析时出现两个问题:
    1. 完全无关的规则触发循环依赖错误(cyclic dependency error)
    2. 通配符(wildcards)混乱,导致工作流期望的输入文件名异常且缺失(将自身输出文件名的一部分添加到全流程输入文件名中)
  • 单独运行到归一化步骤之前的全流程一切正常,但一旦加入归一化步骤,就会触发全流程重跑,此次未出现文件名相关错误

以下是涉及的两个规则代码:

rule resfinder:
    input:
        readF = output_dir + "rqcfilter2/deinterleaved/{sample}/{sample}_1.anqdpht.fastq.gz",
        readR = output_dir + "rqcfilter2/deinterleaved/{sample}/{sample}_2.anqdpht.fastq.gz",
    output:
        res_temp = temp(output_dir + "resfinder/{sample}/resfinder_kma/{sample}_all.unsorted.res"),
        res = output_dir + "resfinder/{sample}/resfinder_kma/{sample}_all.res",
    log:
        output_dir + "logs/resfinder/{sample}.log",
    params:
        out_dir = output_dir + "resfinder/{sample}/",
        db_res = db_dir + config["resfinder"]["db_res"],
    conda:
        "envs/resfinder.yaml",
    shell:
        """
        # Checks if databases have been built and builds them otherwise
        if [ ! -e {params.db_res}/all.comp.b ]
        then
            {params.db_res}/INSTALL.py
        fi

        # Runs resfinder
        run_resfinder.py -ifq {input.readF} {input.readR} -o {params.out_dir} --acquired -db_res {params.db_res}

        # Create a summary results file and sorts it
        array=( {params.out_dir}resfinder_kma/*.res )
        {{ cat ${{array[@]:0:1}}; grep -vh "^#" ${{array[@]:1}}; }} > {output.res_temp}
        (head -n 1 {output.res_temp} && tail -n +2 {output.res_temp}  | sort -nr -k2 -t$'\t') > {output.res}
        """

# Filters Resfinder results based on 90% identity, 60% coverage, calculates Transcripts per Million, and sorts based on that
rule filter_and_normalize_resfinder:
    input:
        res_sorted = output_dir + "resfinder/{sample}/resfinder_kma/{sample}_all.res",
        ref = output_dir + "multiqc/clean/multiqc_report_data/multiqc_fastqc.txt",
    output:
        res = output_dir + "resfinder/{sample}_filtered_normalized.res",
    run:
        # Reads files
        ref = pd.read_csv(input.ref, sep="\t")
        res = pd.read_csv(input.res_sorted, sep="\t")
        
        # Filter results using Strict criteria of 60% coverage and 90% identity
        res = res.query('Template_Coverage >= 60 & Query_Identity >= 90')

        # Calculates TPM by first normalizing for gene length, and then by total sequencing depth (not just ARG reads)
        new_res = res
        new_res['TPM'] = new_res['Depth'] / (new_res['Template_length']/1000) * ((ref.query('Sample == "'+wildcards.sample+'_1.anqdpht"').iat[0, 4])/1000000)

        # Saves new results
        new_res.to_csv(output.res, sep="\t", index=False)

若需其他排查信息请告知,感谢帮助。

内容的提问来源于stack exchange,提问作者sgkion

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 11:57:47