如何在20核节点上让两个作业各占用10核运行?
Got it, let's break this down for you. You have a 20-core node and want to run two jobs each using 10 cores—your existing script is already on the right track, here's how to make it work smoothly:
1. Validate Your Node's Core Count
First, double-check that your target node actually has 20 available cores. Run this command to confirm:
sinfo -o "%N %C"
Look for your node in the output—you should see something like node001 20/20 (total/used cores). If the used count is 0, you're good to go.
2. Keep Your Job Script's Core Configuration
Your existing job.sh already has the correct core setting:
#SBATCH --ntasks-per-node=10 # number of cores
This tells Slurm to allocate 10 cores per job, which is exactly what you need for two jobs to fill the 20-core node.
3. Submit the Two Jobs
There are two reliable ways to get both jobs on the same node:
Option 1: Let Slurm Handle It (Simplest)
Most clusters allow shared node allocation by default. Just submit your jobs sequentially after updating the main.R parameters:
# First job submission sbatch job.sh # Modify main.R parameters here... # Second job submission sbatch job.sh
Slurm will automatically schedule the second job on the same node as the first, as long as 10 cores are still available.
Option 2: Force Both Jobs to the Same Node (Manual Control)
If you want explicit control over which node the jobs run on:
- Submit the first job, then check which node it's using:
Look for thesqueue -j <first-job-id>NODELIST(REASON)column to get the node name (e.g.,node001). - Add a
--nodelistline to yourjob.shfor the second submission:#SBATCH --nodelist=node001 # Force job to run on this specific node - Submit the second job with this modified script.
4. Critical: Restrict R to Use Only Allocated Cores
A common pitfall: R scripts sometimes default to using all available cores on the node, which would cause your two jobs to fight for resources. Add this to the top of your main.R to limit it to 10 cores:
# Limit R to use exactly 10 cores library(parallel) options(mc.cores = 10) # If using future/promises packages: library(future) plan(multisession, workers = 10)
5. Quick Pre-Submission Check
Make sure the sortie directory exists before submitting jobs, otherwise the script will fail to write output/error files:
mkdir -p sortie
Your Final Job Script (With Notes)
Here's your original script with key annotations:
#!/bin/bash #SBATCH --job-name=main #SBATCH --mail-use=abc@gmail.com #SBATCH --mail-type=ALL #SBATCH --nodes=1 # Correct: we only need one node #SBATCH --ntasks-per-node=10 # Correct: 10 cores per job #SBATCH --time=24:00:00 # Adjust as needed #SBATCH --output=sortie/job.out # Ensure sortie dir exists! #SBATCH --error=sortie/error.err module load python module load R # module load openmpi # Only needed if your R script uses MPI—skip otherwise R CMD BATCH '--args ../../' main.R sortie/main.Rout
内容的提问来源于stack exchange,提问作者problème0123

