Nextflow-Groovy:如何从Manta生成的嵌套目录中输出文件
问题解决:Manta流程输出文件缺失错误
核心问题
你遇到的Missing output file(s) candidateSmallIndels.vcf.gz错误,是因为Nextflow在流程的当前工作目录找不到指定的输出文件——Manta实际将变异文件生成在${params.outdir}/${sample_name}/manta/results/variants/目录下,但你的输出定义里写的是当前目录的文件名,导致路径不匹配。
解决方案
1. 修正输出路径指向
将输出定义中的路径改为Manta实际生成文件的位置,同时通过stageAs简化后续流程的文件引用。
2. 优化输入与脚本逻辑
- 合并重复的输入tuple,减少冗余代码
- 定义统一的
run_dir变量,避免重复书写路径 - 修正publishDir的变量引用问题(原定义中
sample_name在process初始化阶段未赋值)
修改后的完整代码
process manta { errorStrategy 'retry' maxRetries 3 // 用动态闭包指定publishDir,避免提前引用未定义的sample_name publishDir path: { "${params.outdir}/${sample_name}/manta/" }, mode: 'copy' input: // 合并肿瘤/正常样本的bam和bai输入,简化结构 tuple val(sample_id_tumor), path(bqsrbam_tumor_files, stageAs: 'manta_tumorbqsrbam/*') tuple val(sample_id_normal), path(bqsrbam_normal_files, stageAs: 'manta_normalbqsrbam/*') output: // 指向Manta实际生成文件的路径,同时stageAs到当前目录方便后续流程引用 tuple val(sample_name), path("${run_dir}/results/variants/candidateSmallIndels.vcf.gz"), stageAs: "candidateSmallIndels.vcf.gz", emit: manta_small_indels_vcf tuple val(sample_name), path("${run_dir}/results/variants/candidateSmallIndels.vcf.gz.tbi"), stageAs: "candidateSmallIndels.vcf.gz.tbi", emit: manta_small_indels_vcf_tbi tuple val(sample_name), path("${run_dir}/results/variants/candidateSV.vcf.gz"), stageAs: "candidateSV.vcf.gz", emit: manta_candidateSV_vcf tuple val(sample_name), path("${run_dir}/results/variants/candidateSV.vcf.gz.tbi"), stageAs: "candidateSV.vcf.gz.tbi", emit: manta_candidateSV_vcf_tbi tuple val(sample_name), path("${run_dir}/results/variants/diploidSV.vcf.gz"), stageAs: "diploidSV.vcf.gz", emit: manta_diploidSV_vcf tuple val(sample_name), path("${run_dir}/results/variants/diploidSV.vcf.gz.tbi"), stageAs: "diploidSV.vcf.gz.tbi", emit: manta_diploidSV_vcf_tbi tuple val(sample_name), path("${run_dir}/results/variants/somaticSV.vcf.gz"), stageAs: "somaticSV.vcf.gz", emit: manta_somaticSV_vcf tuple val(sample_name), path("${run_dir}/results/variants/somaticSV.vcf.gz.tbi"), stageAs: "somaticSV.vcf.gz.tbi", emit: manta_somaticSV_vcf_tbi script: sample_name = sample_id_normal.split('_N')[0] run_dir = "${params.outdir}/${sample_name}/manta/" """ mkdir -p ${run_dir} # 检查并删除已存在的runWorkflow.py if [ -f "${run_dir}/runWorkflow.py" ]; then rm "${run_dir}/runWorkflow.py" fi # 运行Manta配置脚本 /hpc/packages/configManta.py \\ --normalBam ${bqsrbam_normal_files[0]} \\ --tumorBam ${bqsrbam_tumor_files[0]} \\ --referenceFasta ${params.hg38genome} \\ --runDir ${run_dir} # 执行Manta工作流 ${run_dir}/runWorkflow.py -m local -j 12 --quiet """ }
补充说明
- 输入合并后,
bqsrbam_tumor_files和bqsrbam_normal_files为数组,[0]对应bam文件,[1]对应bai文件(需保证输入tuple的顺序为bam在前、bai在后) stageAs将深层目录的文件映射到流程当前工作目录,后续Strelka2流程可直接引用简化后的文件名- 统一的
run_dir变量让脚本更易维护,避免路径拼写错误
内容的提问来源于stack exchange,提问作者Death Metal
相关产品推荐
相关产品推荐

