Nextflow SNP Calling流程问题:进程间文件传递与索引缺失
解决Nextflow SNP calling流程中bwaMap进程的两个核心问题
问题分析
你的流程存在两个明确问题:
- bwaMap未等待bwaIndex完成:直接使用
params.ref_genome作为bwaMap的参考输入,未建立与bwaIndex进程的依赖关系,导致索引文件还未生成就启动映射,触发fail to locate the index files错误。 - 修剪后reads接收错误:bwaMap的输入定义与trimSequences的输出结构不匹配,导致脚本错误地将样本ID当作reads路径,出现命令中的
R A无效参数。
修改后的完整代码
params.reads = "/home/*_R{1,2}*.gz" params.outdir = "./" params.ref_genome = "/home/GCF_000001735.4_TAIR10.1_genomic.fasta" log.info """\ S N P C A L L I N G P I P E L I N E =================================== reference : ${params.ref_genome} reads : ${params.reads} outdir : ${params.outdir} """ .stripIndent(true) /* * Processes */ process trimSequences { tag "Fastp on $sample_id" publishDir params.outdir, mode: 'copy' cpus 4 debug true input: tuple val(sample_id), path(reads) output: tuple val(sample_id), path("${sample_id}_{1,2}_trimmed.fastq.gz") script: """ fastp -w $task.cpus -i ${reads[0]} -I ${reads[1]} -o ${sample_id}_1_trimmed.fastq.gz -O ${sample_id}_2_trimmed.fastq.gz """ } process bwaIndex { publishDir params.outdir, mode: 'copy' input: path reference output: // 输出参考基因组及其所有索引文件 path "${reference.baseName}.*" script: """ bwa index -a bwtsw $reference """ } process bwaMap { publishDir params.outdir, mode: 'copy' cpus 4 tag "BWA mapping for $sample_id" input: // 接收bwaIndex生成的参考基因组及索引文件 path reference_files // 接收trimSequences输出的样本ID和修剪后reads对 tuple val(sample_id), path(trimmed_reads) output: // 带样本ID的输出文件,避免多样本覆盖 file "${sample_id}_aligned.sam" script: // 从参考文件集合中筛选出fasta文件 def ref_fasta = reference_files.find { it.name.endsWith('.fasta') } """ bwa mem -t $task.cpus $ref_fasta ${trimmed_reads[0]} ${trimmed_reads[1]} > ${sample_id}_aligned.sam """ } /* * Define the workflow */ workflow { Channel .fromFilePairs(params.reads, checkIfExists: true) .set { read_pairs_ch } // 处理reads修剪 trimmed_ch = trimSequences(read_pairs_ch) // 生成参考基因组索引,得到包含fasta和所有索引的通道 index_ch = bwaIndex(params.ref_genome) // 组合索引通道与修剪后reads通道,建立依赖关系 bwaMap(index_ch.combine(trimmed_ch)) }
关键修改说明
- 建立进程依赖:通过
index_ch = bwaIndex(params.ref_genome)获取索引生成的输出通道,再用index_ch.combine(trimmed_ch)将其与修剪后reads通道组合,确保bwaMap必须等待bwaIndex完成后才会启动,同时保证每个样本的映射进程都能获取到完整的索引文件。 - 修正输入结构匹配:bwaMap的输入改为
tuple val(sample_id), path(trimmed_reads),完全匹配trimSequences的输出结构(样本ID + 成对reads路径),解决脚本中错误引用reads的问题。 - 避免文件覆盖:bwaMap的输出文件添加
${sample_id}_前缀,防止多个样本的SAM文件互相覆盖。 - 完整传递参考文件:bwaIndex输出改为
path "${reference.baseName}.*",一次性传递参考基因组和所有索引文件,确保bwa在映射时能找到完整的索引。
内容的提问来源于stack exchange,提问作者Diesel
相关产品推荐
相关产品推荐

