如何在Slurm中设置单作业最大CPU申请量?GPU集群资源管控
Great question—this is a super common pain point when managing GPU clusters, where you need to keep CPU resources reserved for GPU-bound jobs but stop non-GPU workloads from hogging all available cores. Let’s break down the most effective ways to set per-job CPU limits in Slurm:
1. Cluster-Wide Hard Limit (slurm.conf)
The simplest approach is to set a global cap on CPU usage per job using the MaxCPUsPerJob parameter in your slurm.conf file. This applies to all jobs across the cluster, unless you override it at the partition or user level.
For example, if you want every job to max out at 32 CPUs:
MaxCPUsPerJob=32
After updating the config, refresh Slurm to apply changes:
scontrol reconfigure
2. Partition-Specific Limits (Granular Control)
Since you have separate GPU and non-GPU workloads, partition-level limits make perfect sense. You can leave GPU partitions unrestricted (your gres.conf mapping already handles CPU-GPU alignment) while capping CPU usage for non-GPU jobs.
When defining partitions in slurm.conf, add MaxCPUsPerJob to your non-GPU partition:
# GPU partition (no CPU limit, relies on gres.conf mapping) PartitionName=gpu Nodes=gpu-node[1-10] Default=YES State=UP # Non-GPU partition (capped at 16 CPUs per job) PartitionName=general Nodes=cpu-node[1-20] Default=NO MaxCPUsPerJob=16 State=UP
Run scontrol reconfigure to activate these partition rules.
3. User/Account-Level Limits (Fine-Tuned Control)
If you need even more precision—like restricting specific users or teams to lower CPU counts—use sacctmgr to set per-entity limits. This is handy if some users run heavy non-GPU jobs while others only submit GPU workloads.
For example, to limit a user named milad to 8 CPUs per job:
sacctmgr modify user name=milad set MaxCPUsPerJob=8
Or apply the limit to an entire account:
sacctmgr modify account name=research set MaxCPUsPerJob=16
4. Optional: Set Default CPU Requests
To complement hard limits, you can set default CPU values for job submissions so users don’t accidentally request more cores than needed. Add these to slurm.conf:
DefCpusPerTask=1 DefMemPerCPU=4096
This sets a default of 1 CPU and 4GB of memory per task, but users can still request more (up to your MaxCPUsPerJob cap).
Key Tips for Your GPU Cluster
Combining these limits with your existing gres.conf GPU-CPU mapping ensures:
- GPU jobs get the exact CPU cores mapped to their requested GPUs (no resource contention)
- Non-GPU jobs can’t consume all remaining CPU resources, leaving headroom for GPU workloads
Test with a few trial jobs to verify limits are enforced—use squeue -o "%.18i %.9P %.8u %.8T %.10M %.6D %.C %R" to check CPU counts for running jobs.
内容的提问来源于stack exchange,提问作者Milad

