You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在20核节点上让两个作业各占用10核运行?

How to Run Two 10-Core Jobs on a 20-Core Node

Got it, let's break this down for you. You have a 20-core node and want to run two jobs each using 10 cores—your existing script is already on the right track, here's how to make it work smoothly:

1. Validate Your Node's Core Count

First, double-check that your target node actually has 20 available cores. Run this command to confirm:

sinfo -o "%N %C"

Look for your node in the output—you should see something like node001 20/20 (total/used cores). If the used count is 0, you're good to go.

2. Keep Your Job Script's Core Configuration

Your existing job.sh already has the correct core setting:

#SBATCH --ntasks-per-node=10 # number of cores

This tells Slurm to allocate 10 cores per job, which is exactly what you need for two jobs to fill the 20-core node.

3. Submit the Two Jobs

There are two reliable ways to get both jobs on the same node:

Option 1: Let Slurm Handle It (Simplest)

Most clusters allow shared node allocation by default. Just submit your jobs sequentially after updating the main.R parameters:

# First job submission
sbatch job.sh

# Modify main.R parameters here...

# Second job submission
sbatch job.sh

Slurm will automatically schedule the second job on the same node as the first, as long as 10 cores are still available.

Option 2: Force Both Jobs to the Same Node (Manual Control)

If you want explicit control over which node the jobs run on:

  1. Submit the first job, then check which node it's using:
    squeue -j <first-job-id>
    
    Look for the NODELIST(REASON) column to get the node name (e.g., node001).
  2. Add a --nodelist line to your job.sh for the second submission:
    #SBATCH --nodelist=node001 # Force job to run on this specific node
    
  3. Submit the second job with this modified script.

4. Critical: Restrict R to Use Only Allocated Cores

A common pitfall: R scripts sometimes default to using all available cores on the node, which would cause your two jobs to fight for resources. Add this to the top of your main.R to limit it to 10 cores:

# Limit R to use exactly 10 cores
library(parallel)
options(mc.cores = 10)

# If using future/promises packages:
library(future)
plan(multisession, workers = 10)

5. Quick Pre-Submission Check

Make sure the sortie directory exists before submitting jobs, otherwise the script will fail to write output/error files:

mkdir -p sortie

Your Final Job Script (With Notes)

Here's your original script with key annotations:

#!/bin/bash
#SBATCH --job-name=main
#SBATCH --mail-use=abc@gmail.com
#SBATCH --mail-type=ALL
#SBATCH --nodes=1 # Correct: we only need one node
#SBATCH --ntasks-per-node=10 # Correct: 10 cores per job
#SBATCH --time=24:00:00 # Adjust as needed
#SBATCH --output=sortie/job.out # Ensure sortie dir exists!
#SBATCH --error=sortie/error.err
module load python
module load R
# module load openmpi # Only needed if your R script uses MPI—skip otherwise
R CMD BATCH '--args ../../' main.R sortie/main.Rout

内容的提问来源于stack exchange,提问作者problème0123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:34:47