You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HPC登录节点运行Nextflow调用Gatk去重时出现内存不足错误

问题:Nextflow调用Gatk MarkDuplicates报Java内存/GC线程错误,命令行直接运行正常

在HPC服务器登录节点运行Nextflow流程,未设置任何Java堆内存(XmXX类)参数,调用Gatk的MarkDuplicates模块时出现Java内存不足、无法创建GC线程的错误,但直接在命令行执行完全相同的Gatk命令却能正常运行。


相关流程脚本

main.nf

params.outdir_fastp="/sc/arion/projects/name/user/output_pipeline_nextflow/test_tiny_datasets/trimmed"
params.outdir_index="/sc/arion/projects/path/name/output_pipeline_nextflow/test_tiny_datasets/bwa_index"
params.rawFiles = "/sc/arion/projects/user/name/tiny/tumor/*_R{1,2}_xxx.fastq.gz"
params.outdir_bwa_mem="/sc/arion/projects/user/name/output_pipeline_nextflow/test_tiny_datasets/bwamem"
params.hg38genome="/sc/arion/projects/user/name/reference_genome/neisseria_meningitidis/NM.fasta"
params.gatk_mark_duplicates="/sc/arion/projects/username/path/output_pipeline_nextflow/test_tiny_datasets/gatk_mark_duplicates"

include { FASTP} from './fastp_process.nf'
include {bwa_index} from './index_process.nf'
include { align_bwa_mem} from './bwamem_process.nf'
include { gatk_markduplicates} from './gatk_markduplicates_process.nf'

workflow {
        read_pairs_ch = Channel.fromFilePairs( params.rawFiles )
        FASTP(read_pairs_ch) 
        bwa_index(params.hg38genome)
        align_bwa_mem(FASTP.out.reads,bwa_index.out) 
        gatk_markduplicates(align_bwa_mem.out.sorted_bams)
}

gatk_markduplicates_process.nf

process gatk_markduplicates {

debug true

    publishDir params.gatk_mark_duplicates , mode:"copy"

    input:
        tuple val(sample_id), path(sorted_bam) 
        
    output:
        tuple val(sample_id),path("${sample_id}.dedup.sorted.bam")
        tuple val(sample_id),path("${sample_id}.markdup.metrics.txt")

script:

"""
        echo "$sample_id ${params.outdir_bwa_mem}/${sorted_bam}\n"      
        ml gatk/4.1.3.0

        gatk MarkDuplicates -I ${params.outdir_bwa_mem}/${sorted_bam} \\
        -O ${sample_id}.dedup.sorted.bam \\
        -M ${sample_id}.markdup.metrics.txt

"""
}

错误信息

Error executing process > 'gatk_markduplicates (2)'

Caused by:
  Process `gatk_markduplicates (2)` terminated with an error exit status (1)

Command executed:

  echo "tiny_t_L007 /sc/arion/projects/user/name/output_pipeline_nextflow/test_tiny_datasets/bwamem/tiny_t_L007.sorted.bam
  " An e
  # hs_eml gatk/4.1.3.0
  
          gatk MarkDuplicates -I /sc/arion/projects/user/name/output_pipeline_nextflow/test_tiny_datasets/bwamem/tiny_t_L007.sorted.bam \
          -O tiny_t_L007.dedup.sorted.bam \
          -M tiny_t_L007.markdup.metrics.txt

Command exit status:
  1

Command output:
  tiny_t_L007 /sc/arion/projects/user/name/output_pipeline_nextflow/test_tiny_datasets/bwamem/tiny_t_L007.sorted.bam
  
  #
  # There is insufficient memory for the Java Runtime Environment to continue.
  # Cannot create GC thread. Out of system resources.
  # An error report file with more information is saved as:
  # hs_err_pid148553.log

Command error:
  tiny_t_L007 /sc/arion/projects/user/name/output_pipeline_nextflow/test_tiny_datasets/bwamem/tiny_t_L007.sorted.bam
  
  
  The following have been reloaded with a version change:
    1) java/11.0.2 => java/1.8.0_211
  
  Using GATK jar /hpc/packages/minerva-centos7/gatk/4.1.3.0/gatk-4.1.3.0/gatk-package-4.1.3.0-local.jar
  Running:
      java -Dsamjdk.use_async_io_read_samtools=false -Dsamjdk.use_async_io_write_samtools=true -Dsamjdk.use_async_io_write_tribble=false -Dsamjdk.compression_level=2 -jar /hpc/packages/minerva-centos7/gatk/4.1.3.0/gatk-4.1.3.0/gatk-package-4.1.3.0-local.jar MarkDuplicates -I /sc/arion/projects/user/name/output_pipeline_nextflow/test_tiny_datasets/bwamem/tiny_t_L007.sorted.bam -O tiny_t_L007.dedup.sorted.bam -M tiny_t_L007.markdup.metrics.txt
  #
  # There is insufficient memory for the Java Runtime Environment to continue.
  # Cannot create GC thread. Out of system resources.
  # An error report file with more information is saved as:
  # hs_err_pid148553.log

Work dir:
  /sc/arion/projects/user/name/nextflow_pipeline/scripts_pipeline/work/e1/0b2f8095e9b0caf074691786ed4c1e

Tip: when you have fixed the problem you can continue the execution adding the option `-resume` to the run command line

命令行可正常运行的示例

gatk MarkDuplicates -I /sc/arion/projects/name/name/output_pipeline_nextflow/test_tiny_datasets/bwamem/tiny_t_L003.sorted.bam         -O tiny_t_L003.dedup.sorted.bam         -M tiny_t_L003.markdup.metrics.txt

或

java -Dsamjdk.use_async_io_read_samtools=false -Dsamjdk.use_async_io_write_samtools=true -Dsamjdk.use_async_io_write_tribble=false -Dsamjdk.compression_level=2 -jar /hpc/packages/minerva-centos7/gatk/4.1.3.0/gatk-4.1.3.0/gatk-package-4.1.3.0-local.jar MarkDuplicates -I /sc/arion/projects/username/path/output_pipeline_nextflow/test_tiny_datasets/bwamem/tiny_t_L001.sorted.bam -O tiny_t_L001.dedup.sorted.bam -M tiny_t_L001.markdup.metrics.txt

解决方案

1. 给Nextflow进程分配明确的内存/CPU资源

Nextflow默认不会为进程分配固定资源,而登录节点通常有严格的资源限制。修改gatk_markduplicates_process.nf,添加内存和CPU参数:

process gatk_markduplicates {

    debug true
    memory '8 GB'  # 根据实际需求调整,比如16GB
    cpus 2         # GC线程数量和CPU核心相关,适当分配核心数

    publishDir params.gatk_mark_duplicates , mode:"copy"

    input:
        tuple val(sample_id), path(sorted_bam) 
        
    output:
        tuple val(sample_id),path("${sample_id}.dedup.sorted.bam")
        tuple val(sample_id),path("${sample_id}.markdup.metrics.txt")

script:
"""
        echo "$sample_id ${params.outdir_bwa_mem}/${sorted_bam}\n"      
        ml gatk/4.1.3.0

        gatk MarkDuplicates -I ${params.outdir_bwa_mem}/${sorted_bam} \\
        -O ${sample_id}.dedup.sorted.bam \\
        -M ${sample_id}.markdup.metrics.txt
"""
}

2. 显式设置Java堆内存参数

直接在Gatk命令中指定Java内存选项,避免Nextflow运行环境的默认值不足:

gatk --java-options "-Xmx8G -Xms4G" MarkDuplicates -I ${params.outdir_bwa_mem}/${sorted_bam} \\
-O ${sample_id}.dedup.sorted.bam \\
-M ${sample_id}.markdup.metrics.txt

或者在script块开头设置环境变量:

export JAVA_OPTS="-Xmx8G -Xms4G"
ml gatk/4.1.3.0
gatk MarkDuplicates ...

3. 避免在登录节点运行计算任务

HPC登录节点仅用于提交任务,不适合运行计算密集型流程。配置Nextflow使用集群调度器(如Slurm),让任务自动提交到计算节点:

  • 在nextflow.config中添加Slurm配置:
process {
    executor = 'slurm'
    queue = 'normal'  # 根据集群队列名称调整
}

4. 检查登录节点资源限制

查看当前Shell和Nextflow运行环境的资源限制差异:

  • 命令行执行ulimit -a查看当前限制
  • 在Nextflow的process脚本中加入ulimit -a输出,对比两者的内存、线程限制,必要时联系管理员调整登录节点的资源配额

内容的提问来源于stack exchange,提问作者Death Metal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 08:29:56