You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Nextflow SNP Calling流程问题:进程间文件传递与索引缺失

解决Nextflow SNP calling流程中bwaMap进程的两个核心问题

问题分析

你的流程存在两个明确问题:

  1. bwaMap未等待bwaIndex完成:直接使用params.ref_genome作为bwaMap的参考输入,未建立与bwaIndex进程的依赖关系,导致索引文件还未生成就启动映射,触发fail to locate the index files错误。
  2. 修剪后reads接收错误:bwaMap的输入定义与trimSequences的输出结构不匹配,导致脚本错误地将样本ID当作reads路径,出现命令中的R A无效参数。

修改后的完整代码

params.reads = "/home/*_R{1,2}*.gz"
params.outdir = "./"
params.ref_genome = "/home/GCF_000001735.4_TAIR10.1_genomic.fasta"

log.info """\
    S N P   C A L L I N G   P I P E L I N E
    ===================================
    reference    : ${params.ref_genome}
    reads        : ${params.reads}
    outdir       : ${params.outdir}
    """
    .stripIndent(true)


/*
 * Processes
 */
process trimSequences {
    tag "Fastp on $sample_id"
    publishDir params.outdir, mode: 'copy'
    cpus 4
    debug true

    input:
    tuple val(sample_id), path(reads)

    output:
    tuple val(sample_id), path("${sample_id}_{1,2}_trimmed.fastq.gz")

    script:
    """
    fastp -w $task.cpus -i ${reads[0]} -I ${reads[1]} -o ${sample_id}_1_trimmed.fastq.gz -O ${sample_id}_2_trimmed.fastq.gz
    """
}

process bwaIndex {
    publishDir params.outdir, mode: 'copy'

    input:
    path reference

    output:
    // 输出参考基因组及其所有索引文件
    path "${reference.baseName}.*"

    script:
    """
    bwa index -a bwtsw $reference
    """
}

process bwaMap {
    publishDir params.outdir, mode: 'copy'
    cpus 4
    tag "BWA mapping for $sample_id"

    input:
    // 接收bwaIndex生成的参考基因组及索引文件
    path reference_files
    // 接收trimSequences输出的样本ID和修剪后reads对
    tuple val(sample_id), path(trimmed_reads)

    output:
    // 带样本ID的输出文件,避免多样本覆盖
    file "${sample_id}_aligned.sam"

    script:
    // 从参考文件集合中筛选出fasta文件
    def ref_fasta = reference_files.find { it.name.endsWith('.fasta') }
    """
    bwa mem -t $task.cpus $ref_fasta ${trimmed_reads[0]} ${trimmed_reads[1]} > ${sample_id}_aligned.sam
    """
}


/*
 * Define the workflow
 */
workflow {

    Channel
        .fromFilePairs(params.reads, checkIfExists: true)
        .set { read_pairs_ch }

    // 处理reads修剪
    trimmed_ch = trimSequences(read_pairs_ch)
    // 生成参考基因组索引,得到包含fasta和所有索引的通道
    index_ch = bwaIndex(params.ref_genome)
    // 组合索引通道与修剪后reads通道,建立依赖关系
    bwaMap(index_ch.combine(trimmed_ch))
}

关键修改说明

  • 建立进程依赖:通过index_ch = bwaIndex(params.ref_genome)获取索引生成的输出通道,再用index_ch.combine(trimmed_ch)将其与修剪后reads通道组合,确保bwaMap必须等待bwaIndex完成后才会启动,同时保证每个样本的映射进程都能获取到完整的索引文件。
  • 修正输入结构匹配:bwaMap的输入改为tuple val(sample_id), path(trimmed_reads),完全匹配trimSequences的输出结构(样本ID + 成对reads路径),解决脚本中错误引用reads的问题。
  • 避免文件覆盖:bwaMap的输出文件添加${sample_id}_前缀,防止多个样本的SAM文件互相覆盖。
  • 完整传递参考文件:bwaIndex输出改为path "${reference.baseName}.*",一次性传递参考基因组和所有索引文件,确保bwa在映射时能找到完整的索引。

内容的提问来源于stack exchange,提问作者Diesel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 03:57:48