You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Nextflow添加本地MSA搜索步骤后全流程跳过问题求助

Nextflow BOLTZ流程新增步骤后全部被跳过问题排查

问题描述

原本仅将样本表CSV转换为Boltz所需YAML格式的Nextflow BOLTZ流程可正常运行,为替换Boltz默认的云端MSA搜索,新增MMSEQS_COLABFOLDSEARCH步骤及对应CREATE_SAMPLESHEET_YAML_MSA、RUN_BOLTZ_MSA流程后,所有步骤均被跳过,请求协助排查原因。

相关代码

原正常流程代码

workflow BOLTZ {
    take:
    ch_samplesheet  // channel: samplesheet read from --input
    ch_versions     // channel: [ path(versions.yml) ]
    ch_boltz_ccd    // channel: [ path(boltz_ccd) ]
    ch_boltz_model  // channel: [ path(model) ]

    ch_multiqc_files = Channel.empty()

    // CREATE_SAMPLESHEET_YAML
    CREATE_SAMPLESHEET_YAML(
        ch_samplesheet,
    )

    // RUN_BOLTZ
    RUN_BOLTZ(
        CREATE_SAMPLESHEET_YAML.out.samplesheet,
        ch_boltz_model,
        ch_boltz_ccd
    )

    emit:
    versions   = ch_versions
    msa        = RUN_BOLTZ.out.msa
    structures = RUN_BOLTZ.out.structures
    confidence = RUN_BOLTZ.out.confidence
    plddt      = RUN_BOLTZ.out.plddt
}

新增的进程代码

process CREATE_SAMPLESHEET_YAML_MSA {
    cache false
    tag "$meta.id"
    label 'process_single'

    container 'docker://nbtmsh/samplesheet-utils:1.1'

    input:
    tuple val(meta), path(samplesheet)
    path ('*.a3m')

    output:
    tuple val(meta), file('*.yaml'), emit: samplesheet

    when:
    task.ext.when == null || task.ext.when

    script:
    """
    create-samplesheet \
        --directory ./ \
        --msa-dir ./ \
        --yaml \
        --output-file $(sample-name --sanitise --index 0 ./*.fasta).yaml
    """

    stub:
    """
    echo "" > samplesheet.yaml
    """
}

process RUN_BOLTZ_MSA {
    tag "$meta.id"
    label 'process_medium'

    container "/srv/scratch/sbf-pipelines/proteinfold/singularity/boltz.sif"

    input:
    tuple val(meta), path(fasta)
    path ('boltz1_conf.ckpt')
    path ('ccd.pkl')
    path ('**.a3m')

    output:
    path ("boltz_results_${fasta.baseName}/processed/msa/*.npz"), emit: msa
    path ("boltz_results_${fasta.baseName}/processed/structures/*.npz"), emit: structures
    path ("boltz_results_${fasta.baseName}/predictions/${fasta.baseName}/confidence*.json"), emit: confidence
    path ("boltz_results_${fasta.baseName}/predictions/${fasta.baseName}/plddt_*.npz"), emit: plddt

    script:
    """
    boltz predict --use_msa_server "./${fasta.name}" --cache ./
    """
}

修改后的流程代码

workflow BOLTZ {
    take:
    ch_samplesheet  // channel: samplesheet read from --input
    ch_versions     // channel: [ path(versions.yml) ]
    ch_boltz_ccd    // channel: [ path(boltz_ccd) ]
    ch_boltz_model  // channel: [ path(model) ]
    ch_colabfold_params // channel: [ path(colabfold_params) ]
    ch_colabfold_db // channel: [ path(colabfold_db) ]
    ch_uniref30     // channel: [ path(uniref30) ]

    main:
    ch_multiqc_files = Channel.empty()

    // MMSEQS_COLABFOLDSEARCH
    MMSEQS_COLABFOLDSEARCH (
        ch_samplesheet,
        ch_colabfold_params,
        ch_colabfold_db,
        ch_uniref30
    )

    ch_versions = ch_versions.mix(MMSEQS_COLABFOLDSEARCH.out.versions)

    // CREATE_SAMPLESHEET_YAML
    CREATE_SAMPLESHEET_YAML_MSA(
        ch_samplesheet,
        MMSEQS_COLABFOLDSEARCH.out.a3m
    )

    // RUN_BOLTZ
    RUN_BOLTZ_MSA(
        CREATE_SAMPLESHEET_YAML_MSA.out.samplesheet,
        ch_boltz_model,
        ch_boltz_ccd,
        MMSEQS_COLABFOLDSEARCH.out.a3m
    )

    emit:
    versions   = ch_versions
    msa        = RUN_BOLTZ_MSA.out.msa
    structures = RUN_BOLTZ_MSA.out.structures
    confidence = RUN_BOLTZ_MSA.out.confidence
    plddt      = RUN_BOLTZ_MSA.out.plddt
}

MMSEQS_COLABFOLDSEARCH进程代码

process MMSEQS_COLABFOLDSEARCH {
    tag "$meta.id"
    label 'process_high_memory'

    // Exit if running this module with -profile conda / -profile mamba
    if (workflow.profile.tokenize(',').intersect(['conda', 'mamba']).size() >= 1) {
        error("Local MMSEQS_COLABFOLDSEARCH module does not support Conda. Please use Docker / Singularity / Podman instead.")
    }

    container "nf-core/proteinfold_colabfold:dev"

    input:
    tuple val(meta), path(fasta)
    path ('db/params')
    path colabfold_db
    path uniref30

    output:
    tuple val(meta), path("**.a3m"), emit: a3m
    path "versions.yml", emit: versions
    path fasta, emit: fasta

    when:
    task.ext.when == null || task.ext.when

    script:
    def args = task.ext.args ?: ''
    def VERSION = '1.5.2' // WARN: Version information not provided by tool on CLI. Please update this string when bumping container versions.

    """
    ln -r -s $uniref30/uniref30_* ./db
    ln -r -s $colabfold_db/colabfold_envdb* ./db

    /localcolabfold/colabfold-conda/bin/colabfold_search \
        $args \
        --threads $task.cpus ${fasta} \
        ./db \
        "result/"

    cat <<-END_VERSIONS > versions.yml
    "${task.process}":
        colabfold_search: $VERSION
    END_VERSIONS
    """

    stub:
    def VERSION = '1.5.2' // WARN: Version information not provided by tool on CLI. Please update this string when bumping container versions.
    """
    mkdir results
    touch results/${meta.id}.a3m

    cat <<-END_VERSIONS > versions.yml
    "${task.process}":
        colabfold_search: $VERSION
    END_VERSIONS
    """
}

上游调用BOLTZ流程的代码

workflow NFCORE_PROTEINFOLD {
    take:
    // read from the --samplesheet input (this is a csv file)
    samplesheet

    main:
    ch_samplesheet = samplesheet

    [...]

    BOLTZ(
        ch_samplesheet,
        ch_versions,
        PREPARE_BOLTZ_DBS.out.boltz_ccd,
        PREPARE_BOLTZ_DBS.out.boltz_model,
        PREPARE_COLABFOLD_DBS.out.colabfold_db,
        PREPARE_COLABFOLD_DBS.out.uniref30,
        params.num_recycles_colabfold

    )
    ch_versions = ch_versions.mix(BOLTZ.out.versions)
    ch_report_input = ch_report_input.mix(
        BOLTZ
            .out
            .msa
            .join(BOLTZ.out.structures)
            .join(BOLTZ.out.confidence)
            .join(BOLTZ.out.plddt)
            .map { it[0]["model"] = "boltz"; it }
    )
}

排查方向

  1. 参数传递顺序不匹配
    修改后的BOLTZ流程take部分的参数顺序为:ch_samplesheet, ch_versions, ch_boltz_ccd, ch_boltz_model, ch_colabfold_params, ch_colabfold_db, ch_uniref30,但上游调用时传递的顺序是:ch_samplesheet, ch_versions, boltz_ccd, boltz_model, colabfold_db, uniref30, num_recycles_colabfold。这里把ch_colabfold_params的位置传给了colabfold_db,后续参数全部错位,导致MMSEQS_COLABFOLDSEARCH无法获取正确的输入通道,直接跳过。

  2. 通道格式不兼容
    CREATE_SAMPLESHEET_YAML_MSA的输入要求是tuple(val(meta), path(samplesheet)),但上游传入的ch_samplesheet是原始CSV文件通道(非tuple格式),无法和MMSEQS_COLABFOLDSEARCH输出的tuple格式a3m通道正确配对,导致进程没有有效输入而跳过。

  3. 进程输入的通配符匹配问题
    CREATE_SAMPLESHEET_YAML_MSA和RUN_BOLTZ_MSA的输入使用path ('*.a3m')、path ('**.a3m')通配符,但MMSEQS_COLABFOLDSEARCH输出的a3m文件在results/目录下,通道传递时可能没有保留完整路径,导致进程无法匹配到文件,触发跳过。

  4. when条件限制
    所有进程都设置了when: task.ext.when == null || task.ext.when,如果全局或任务级别的task.ext.when被设置为false,会直接跳过所有进程,可检查配置文件或命令行参数是否有相关设置。

内容的提问来源于stack exchange,提问作者NathanBitTheMoon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 13:37:01