Nextflow添加本地MSA搜索步骤后全流程跳过问题求助
问题描述
原本仅将样本表CSV转换为Boltz所需YAML格式的Nextflow BOLTZ流程可正常运行,为替换Boltz默认的云端MSA搜索,新增MMSEQS_COLABFOLDSEARCH步骤及对应CREATE_SAMPLESHEET_YAML_MSA、RUN_BOLTZ_MSA流程后,所有步骤均被跳过,请求协助排查原因。
相关代码
原正常流程代码
workflow BOLTZ { take: ch_samplesheet // channel: samplesheet read from --input ch_versions // channel: [ path(versions.yml) ] ch_boltz_ccd // channel: [ path(boltz_ccd) ] ch_boltz_model // channel: [ path(model) ] ch_multiqc_files = Channel.empty() // CREATE_SAMPLESHEET_YAML CREATE_SAMPLESHEET_YAML( ch_samplesheet, ) // RUN_BOLTZ RUN_BOLTZ( CREATE_SAMPLESHEET_YAML.out.samplesheet, ch_boltz_model, ch_boltz_ccd ) emit: versions = ch_versions msa = RUN_BOLTZ.out.msa structures = RUN_BOLTZ.out.structures confidence = RUN_BOLTZ.out.confidence plddt = RUN_BOLTZ.out.plddt }
新增的进程代码
process CREATE_SAMPLESHEET_YAML_MSA { cache false tag "$meta.id" label 'process_single' container 'docker://nbtmsh/samplesheet-utils:1.1' input: tuple val(meta), path(samplesheet) path ('*.a3m') output: tuple val(meta), file('*.yaml'), emit: samplesheet when: task.ext.when == null || task.ext.when script: """ create-samplesheet \ --directory ./ \ --msa-dir ./ \ --yaml \ --output-file $(sample-name --sanitise --index 0 ./*.fasta).yaml """ stub: """ echo "" > samplesheet.yaml """ } process RUN_BOLTZ_MSA { tag "$meta.id" label 'process_medium' container "/srv/scratch/sbf-pipelines/proteinfold/singularity/boltz.sif" input: tuple val(meta), path(fasta) path ('boltz1_conf.ckpt') path ('ccd.pkl') path ('**.a3m') output: path ("boltz_results_${fasta.baseName}/processed/msa/*.npz"), emit: msa path ("boltz_results_${fasta.baseName}/processed/structures/*.npz"), emit: structures path ("boltz_results_${fasta.baseName}/predictions/${fasta.baseName}/confidence*.json"), emit: confidence path ("boltz_results_${fasta.baseName}/predictions/${fasta.baseName}/plddt_*.npz"), emit: plddt script: """ boltz predict --use_msa_server "./${fasta.name}" --cache ./ """ }
修改后的流程代码
workflow BOLTZ { take: ch_samplesheet // channel: samplesheet read from --input ch_versions // channel: [ path(versions.yml) ] ch_boltz_ccd // channel: [ path(boltz_ccd) ] ch_boltz_model // channel: [ path(model) ] ch_colabfold_params // channel: [ path(colabfold_params) ] ch_colabfold_db // channel: [ path(colabfold_db) ] ch_uniref30 // channel: [ path(uniref30) ] main: ch_multiqc_files = Channel.empty() // MMSEQS_COLABFOLDSEARCH MMSEQS_COLABFOLDSEARCH ( ch_samplesheet, ch_colabfold_params, ch_colabfold_db, ch_uniref30 ) ch_versions = ch_versions.mix(MMSEQS_COLABFOLDSEARCH.out.versions) // CREATE_SAMPLESHEET_YAML CREATE_SAMPLESHEET_YAML_MSA( ch_samplesheet, MMSEQS_COLABFOLDSEARCH.out.a3m ) // RUN_BOLTZ RUN_BOLTZ_MSA( CREATE_SAMPLESHEET_YAML_MSA.out.samplesheet, ch_boltz_model, ch_boltz_ccd, MMSEQS_COLABFOLDSEARCH.out.a3m ) emit: versions = ch_versions msa = RUN_BOLTZ_MSA.out.msa structures = RUN_BOLTZ_MSA.out.structures confidence = RUN_BOLTZ_MSA.out.confidence plddt = RUN_BOLTZ_MSA.out.plddt }
MMSEQS_COLABFOLDSEARCH进程代码
process MMSEQS_COLABFOLDSEARCH { tag "$meta.id" label 'process_high_memory' // Exit if running this module with -profile conda / -profile mamba if (workflow.profile.tokenize(',').intersect(['conda', 'mamba']).size() >= 1) { error("Local MMSEQS_COLABFOLDSEARCH module does not support Conda. Please use Docker / Singularity / Podman instead.") } container "nf-core/proteinfold_colabfold:dev" input: tuple val(meta), path(fasta) path ('db/params') path colabfold_db path uniref30 output: tuple val(meta), path("**.a3m"), emit: a3m path "versions.yml", emit: versions path fasta, emit: fasta when: task.ext.when == null || task.ext.when script: def args = task.ext.args ?: '' def VERSION = '1.5.2' // WARN: Version information not provided by tool on CLI. Please update this string when bumping container versions. """ ln -r -s $uniref30/uniref30_* ./db ln -r -s $colabfold_db/colabfold_envdb* ./db /localcolabfold/colabfold-conda/bin/colabfold_search \ $args \ --threads $task.cpus ${fasta} \ ./db \ "result/" cat <<-END_VERSIONS > versions.yml "${task.process}": colabfold_search: $VERSION END_VERSIONS """ stub: def VERSION = '1.5.2' // WARN: Version information not provided by tool on CLI. Please update this string when bumping container versions. """ mkdir results touch results/${meta.id}.a3m cat <<-END_VERSIONS > versions.yml "${task.process}": colabfold_search: $VERSION END_VERSIONS """ }
上游调用BOLTZ流程的代码
workflow NFCORE_PROTEINFOLD { take: // read from the --samplesheet input (this is a csv file) samplesheet main: ch_samplesheet = samplesheet [...] BOLTZ( ch_samplesheet, ch_versions, PREPARE_BOLTZ_DBS.out.boltz_ccd, PREPARE_BOLTZ_DBS.out.boltz_model, PREPARE_COLABFOLD_DBS.out.colabfold_db, PREPARE_COLABFOLD_DBS.out.uniref30, params.num_recycles_colabfold ) ch_versions = ch_versions.mix(BOLTZ.out.versions) ch_report_input = ch_report_input.mix( BOLTZ .out .msa .join(BOLTZ.out.structures) .join(BOLTZ.out.confidence) .join(BOLTZ.out.plddt) .map { it[0]["model"] = "boltz"; it } ) }
排查方向
参数传递顺序不匹配
修改后的BOLTZ流程take部分的参数顺序为:ch_samplesheet, ch_versions, ch_boltz_ccd, ch_boltz_model, ch_colabfold_params, ch_colabfold_db, ch_uniref30,但上游调用时传递的顺序是:ch_samplesheet, ch_versions, boltz_ccd, boltz_model, colabfold_db, uniref30, num_recycles_colabfold。这里把ch_colabfold_params的位置传给了colabfold_db,后续参数全部错位,导致MMSEQS_COLABFOLDSEARCH无法获取正确的输入通道,直接跳过。通道格式不兼容
CREATE_SAMPLESHEET_YAML_MSA的输入要求是tuple(val(meta), path(samplesheet)),但上游传入的ch_samplesheet是原始CSV文件通道(非tuple格式),无法和MMSEQS_COLABFOLDSEARCH输出的tuple格式a3m通道正确配对,导致进程没有有效输入而跳过。进程输入的通配符匹配问题
CREATE_SAMPLESHEET_YAML_MSA和RUN_BOLTZ_MSA的输入使用path ('*.a3m')、path ('**.a3m')通配符,但MMSEQS_COLABFOLDSEARCH输出的a3m文件在results/目录下,通道传递时可能没有保留完整路径,导致进程无法匹配到文件,触发跳过。when条件限制
所有进程都设置了when: task.ext.when == null || task.ext.when,如果全局或任务级别的task.ext.when被设置为false,会直接跳过所有进程,可检查配置文件或命令行参数是否有相关设置。
内容的提问来源于stack exchange,提问作者NathanBitTheMoon

