You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SLURM配置:按GPU请求数自动限CPU/内存或虚拟拆分节点

Absolutely, SLURM can handle both scenarios you're asking about—let's break this down step by step.

1. Auto-Limiting CPU/Memory Based on GPU Count

Yes, you can configure SLURM to automatically scale CPU and memory allocations proportionally to the number of GPUs a user requests. Here's how to set it up:

Core Configuration: cpus-per-gpu and mem-per-gpu

First, define the base resources per GPU for your node. Let's say you have a node with:

  • Total CPUs: 32
  • Total Memory: 128GB
  • Total GPUs: 4
    This gives BaseCPU = 32/4 = 8 and BaseMEM = 128/4 = 32GB per GPU.

You can set these ratios directly on the node using sacctmgr:

sacctmgr modify node <your-node-name> set cpus_per_gpu=8,mem_per_gpu=32G

Or add these parameters to your slurm.conf for the node:

NodeName=<your-node-name> CPUs=32 RealMemory=128000 Gres=gpu:4 CPUsPerGPU=8 MemPerGPU=32000

How It Works for Users

When a user submits a job requesting 2 GPUs with:

sbatch --gres=gpu:2 my_job.sh

SLURM will automatically allocate 2*8=16 CPUs and 2*32=64GB of memory—no need for the user to manually specify --cpus-per-task or --mem.

Enforcing Compliance (Optional)

If you want to prevent users from overriding these ratios (e.g., requesting 2 GPUs but 32 CPUs), use a job_submit.lua script. This script runs when a job is submitted and can adjust or reject jobs that don't follow your resource rules.

A simplified example script might look like this:

function job_submit(job_desc, part_list, submit_uid)
    local gpu_count = job_desc.gres and job_desc.gres:match('gpu:(%d+)') or 0
    if gpu_count > 0 then
        local base_cpu = 8  -- Match your BaseCPU value
        local base_mem = 32000  -- Match your BaseMEM (in MB)
        -- Set CPU count if not specified
        if job_desc.cpus_per_task == nil then
            job_desc.cpus_per_task = base_cpu * gpu_count
        end
        -- Set memory if not specified
        if job_desc.mem == nil then
            job_desc.mem = base_mem * gpu_count .. 'M'
        end
        -- Reject job if requested CPU/memory exceeds the ratio
        if job_desc.cpus_per_task > base_cpu * gpu_count or job_desc.mem:match('%d+')+0 > base_mem * gpu_count then
            slurm.log_info("Job rejected: CPU/memory exceeds ratio for %d GPUs", gpu_count)
            return slurm.ERROR
        end
    end
    return slurm.SUCCESS
end

Place this script in /etc/slurm/job_submit.lua and update slurm.conf to enable it:

JobSubmitPlugins=lua
2. Virtualizing the Node into Smaller Logical Nodes

If you prefer a more rigid split (treating each GPU+BaseCPU+BaseMEM as a separate "virtual node"), you can split your physical node into multiple logical nodes in SLURM.

Example Configuration

Using the same node specs (32 CPUs, 128GB, 4 GPUs), add these entries to slurm.conf:

# Define 4 logical nodes, each with 1 GPU, 8 CPUs, 32GB memory
NodeName=gpu-node-1 CPUs=8 RealMemory=32000 Gres=gpu:1 NodeAddr=<physical-node-ip>
NodeName=gpu-node-2 CPUs=8 RealMemory=32000 Gres=gpu:1 NodeAddr=<physical-node-ip>
NodeName=gpu-node-3 CPUs=8 RealMemory=32000 Gres=gpu:1 NodeAddr=<physical-node-ip>
NodeName=gpu-node-4 CPUs=8 RealMemory=32000 Gres=gpu:1 NodeAddr=<physical-node-ip>

# Create a partition that includes all logical nodes
PartitionName=gpu Nodes=gpu-node-[1-4] Default=YES MaxTime=INFINITE State=UP

How It Works

  • When a user requests 1 GPU, SLURM will allocate one of the logical nodes (8 CPUs + 32GB).
  • For a 2-GPU job, SLURM will allocate two logical nodes (16 CPUs + 64GB)—all on the same physical hardware since they share the same NodeAddr.

Key Notes

  • Ensure the sum of logical node resources matches the physical node's total (no over-provisioning).
  • If you want jobs to stay on the same physical node when requesting multiple GPUs, you can use node features or constraints to group the logical nodes.
  • This approach simplifies resource management for users, as each "node" maps directly to a single GPU and its associated CPU/memory.

内容的提问来源于stack exchange,提问作者Hyperplane

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:56:29