同一节点上Slurm分区CPU过度分配问题排查求助
Slurm共享节点分区资源冲突问题排查
我有一台配备32个CPU的计算节点,已定义两个均使用该节点的Slurm分区A和B。若在分区A提交两个分别请求20个CPU和25个CPU的作业,第二个作业会等待资源;但分别在分区A提交20个CPU、分区B提交25个CPU的作业时,两者均会启动。执行scontrol show node命令显示仅分配了25个CPU。显然两个共享节点的分区未互相遵守资源分配规则,我已设置OverSubscribe=NO但无效果。
以下是我的slurm.conf配置:
MpiDefault=none ProctrackType=proctrack/linuxproc TaskPlugin=task/affinity,task/cgroup ReturnToService=1 # ##Debug SlurmUser=slurm RebootProgram="/usr/sbin/reboot" SlurmctldPidFile=/var/run/slurmctld.pid SlurmdPidFile=/var/run/slurmd.pid SlurmdSpoolDir=/var/spool/slurmd StateSaveLocation=/var/spool/slurmSave # # SCHEDULING SchedulerType=sched/builtin SchedulerParameters=preempt_youngest_first,preempt_strict_order PriorityType=priority/basic SelectType=select/cons_res SelectTypeParameters=CR_Core_Memory,CR_ONE_TASK_PER_CORE PreemptType=preempt/partition_prio PreemptMode=suspend,gang SlurmctldParameters=preempt_send_user_signal AccountingStorageType=accounting_storage/slurmdbd AccountingStorageTRES=gres/license ClusterName=simulation #JobAcctGatherFrequency=30 JobAcctGatherType=jobacct_gather/cgroup #SlurmctldDebug=debug5 #SlurmdDebug=debug5 SlurmctldLogFile=/var/log/slurmLog/slurmctld.log SlurmSchedLogFile=/var/log/slurmLog/slurmsched.log SlurmdLogFile=/var/log/slurmLog/slurmd.log JobCompType=jobcomp/filetxt # # Licenses as generic resources GresTypes=license # # # # COMPUTE NODES NodeName=cl014 Sockets=2 CoresPerSocket=16 ThreadsPerCore=1 State=UNKNOWN RealMemory=237852 PartitionName=A MaxTime=INFINITE PriorityTier=2 PreemptMode=suspend Default=NO Nodes=se-got-cl014 OverSubScribe=NO PartitionName=B MaxTime=INFINITE PriorityTier=2 PreemptMode=suspend Default=NO Nodes=se-got-cl014 OverSubscribe=NO
问题根源与修复方案
1. 节点名称配置不匹配(核心错误)
你定义的节点名称是cl014,但两个分区的Nodes字段填写的是se-got-cl014。Slurm会将这视为两个完全独立的节点,因此调度时不会共享该节点的资源配额,直接导致跨分区提交作业时绕过了资源限制。
修复方法:将两个分区的Nodes值修改为与NodeName一致的cl014:
PartitionName=A MaxTime=INFINITE PriorityTier=2 PreemptMode=suspend Default=NO Nodes=cl014 OverSubScribe=NO PartitionName=B MaxTime=INFINITE PriorityTier=2 PreemptMode=suspend Default=NO Nodes=cl014 OverSubscribe=NO
2. 确认OverSubscribe参数有效性
OverSubscribe=NO的配置本身是正确的,它用于禁止单分区内的资源超配,但前提是Slurm能正确识别节点归属。修复节点名称后,该参数会正常生效,跨分区作业也会受到节点总CPU资源的限制。
3. 验证资源追踪配置
当前的SelectType=select/cons_res和SelectTypeParameters=CR_Core_Memory,CR_ONE_TASK_PER_CORE配置符合需求,它会按核心和内存维度全局追踪节点资源,确保总资源不被超配。
修复后操作
修改配置后,需要重启Slurm服务使设置生效:
systemctl restart slurmctld systemctl restart slurmd
内容的提问来源于stack exchange,提问作者Daniel
相关产品推荐
相关产品推荐

