Nextflow脚本问题:如何修复all_stats中文件名显示为FIFO序号而非barcode?
问题描述
在以下Nextflow脚本中,cat进程会生成barcode01.fastq.gz和barcode02.fastq.gz文件,所有cat进程的输出被汇总后送入all_stats进程统一处理。但最终生成的all_stats.txt文件中,文件列显示的是1.fastq.gz而非barcode01.fastq.gz(该数字为FIFO序列号而非barcode编号)。需要修改代码使barcode编号能正确显示在统计结果中。
原脚本
def barcodes = (1..2).collect { String.format("barcode%02d", it) } params.orifq = barcodes.collect { "fastq_pass/$it/*.fastq.gz" } Channel .fromPath(params.orifq) .map { it -> [it.name.split("_")[2], it] } .groupTuple() .set{orifq_ch} process cat { debug true publishDir = [ path: "Run/orifq", mode: 'copy' ] input: tuple val(bc), path(fq) output: path("*.fastq.gz") """ cat ${fq} > ${bc}.fastq.gz """ } process all_stats { debug true publishDir = [ path: "Run/stats", mode: 'copy' ] input: path ("*.fastq.gz") output: path ("all_stats.txt"), emit: all_stats """ seqkit stat *.fastq.gz > all_stats.txt """ } workflow { cat(orifq_ch)|collect|all_stats|view }
修改方案
问题根源在于collect操作会让Nextflow将文件重命名为临时序号名,且all_stats进程输入未保留原始文件名。修改步骤如下:
1. 修正cat进程的输出定义
明确指定输出文件名,避免模糊匹配:
process cat { debug true publishDir = [ path: "Run/orifq", mode: 'copy' ] input: tuple val(bc), path(fq) output: path("${bc}.fastq.gz"), emit: cat_out // 明确绑定barcode作为输出文件名 }
2. 修改all_stats进程的输入配置
添加keepName: true参数,确保文件以原始barcode名称复制到进程工作目录:
process all_stats { debug true publishDir = [ path: "Run/stats", mode: 'copy' ] input: path(fq_files, keepName: true) // 保留原始文件名 output: path ("all_stats.txt"), emit: all_stats """ seqkit stat ${fq_files} > all_stats.txt """ }
3. 调整工作流逻辑
移除collect操作,直接将cat进程的输出传递给all_stats:
workflow { cat(orifq_ch) | all_stats | view }
修改后的完整脚本
def barcodes = (1..2).collect { String.format("barcode%02d", it) } params.orifq = barcodes.collect { "fastq_pass/$it/*.fastq.gz" } Channel .fromPath(params.orifq) .map { it -> [it.name.split("_")[2], it] } .groupTuple() .set{orifq_ch} process cat { debug true publishDir = [ path: "Run/orifq", mode: 'copy' ] input: tuple val(bc), path(fq) output: path("${bc}.fastq.gz"), emit: cat_out """ cat ${fq} > ${bc}.fastq.gz """ } process all_stats { debug true publishDir = [ path: "Run/stats", mode: 'copy' ] input: path(fq_files, keepName: true) output: path ("all_stats.txt"), emit: all_stats """ seqkit stat ${fq_files} > all_stats.txt """ } workflow { cat(orifq_ch) | all_stats | view }
修改后,all_stats进程会收到原始barcode命名的文件,seqkit stat生成的统计结果中,文件列将正确显示barcode01.fastq.gz和barcode02.fastq.gz。
内容的提问来源于stack exchange,提问作者Hanjié
相关产品推荐
相关产品推荐

