You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SLURM控制器与Worker节点共用同一机器是否可行?如何部署及是否建议?

Is this deployment feasible?

Yes, 100% — SLURM explicitly supports having a single node act as both the controller (slurmctld) and a worker (slurmd). This is a common setup for small-scale clusters, test environments, or situations where hardware resources are limited. Lots of folks run SLURM this way when they don’t have extra hardware for a dedicated head node.

How to implement this setup

Here’s a practical, step-by-step guide to get this up and running:

1. Install SLURM packages

  • On the node that will serve as both controller and worker: Install the full SLURM stack, including the controller daemon, worker daemon, and command-line tools.
    • For Debian/Ubuntu: sudo apt install slurm-wlm slurm-wlm-basic-plugins
    • For RHEL/CentOS/Rocky: sudo dnf install slurm-slurmctld slurm-slurmd slurm-client
  • On your other 3 worker-only nodes: Install just the worker daemon and client tools (skip the controller package).

2. Configure SLURM and authentication

a. Set up slurm.conf

You’ll need a shared slurm.conf file across all nodes. Here’s a minimal example tailored to your 4-node setup:

ControlMachine=node01  # Replace with your controller/worker node's hostname
ControlAddr=192.168.1.100  # Optional: IP of the control node for reliability
SlurmUser=slurm
MpiDefault=none
ProctrackType=proctrack/cgroup
ReturnToService=1
SlurmctldPidFile=/var/run/slurmctld.pid
SlurmdPidFile=/var/run/slurmd.pid
SlurmdSpoolDir=/var/spool/slurmd
SlurmctldSpoolDir=/var/spool/slurmctld
StateSaveLocation=/var/spool/slurmctld
SlurmctldTimeout=300
SlurmdTimeout=300
InactiveLimit=0
MinJobAge=300
KillWait=30
Waittime=0

# Define all nodes, including your control/worker node
NodeName=node01 CPUs=8 State=UNKNOWN  # Adjust CPUs to match your node's core count
NodeName=node02 CPUs=8 State=UNKNOWN
NodeName=node03 CPUs=8 State=UNKNOWN
NodeName=node04 CPUs=8 State=UNKNOWN

# Define a default partition that includes all nodes
PartitionName=normal Nodes=ALL Default=YES MaxTime=INFINITE State=UP
  • Replace node01, the IP, and CPU counts with your actual node details.
  • Copy this file to /etc/slurm/slurm.conf on all nodes, then set permissions: sudo chmod 644 /etc/slurm/slurm.conf && sudo chown root:root /etc/slurm/slurm.conf

b. Set up MUNGE (authentication)

SLURM uses MUNGE for secure node-to-node communication:

  1. Install MUNGE on all nodes:
    • Debian/Ubuntu: sudo apt install munge
    • RHEL/CentOS/Rocky: sudo dnf install munge
  2. On your control/worker node, generate a MUNGE key:
    sudo dd if=/dev/urandom bs=1 count=1024 > /etc/munge/munge.key
    
  3. Copy this key to all other worker nodes (use scp or a shared filesystem):
    sudo scp /etc/munge/munge.key node02:/etc/munge/
    
  4. On all nodes, set the correct permissions for the key:
    sudo chmod 400 /etc/munge/munge.key
    sudo chown munge:munge /etc/munge/munge.key
    
  5. Start and enable MUNGE on all nodes:
    sudo systemctl enable --now munge
    

3. Start SLURM services

  • On your control/worker node: Start both the controller and worker daemons:
    sudo systemctl enable --now slurmctld slurmd
    
  • On your worker-only nodes: Start just the worker daemon:
    sudo systemctl enable --now slurmd
    

4. Verify the setup

  • Run sinfo to check if all nodes are in the UP state:
    sinfo
    
  • Submit test jobs to confirm everything works:
    # Run a hostname command on 1 node (will run on any available worker, including node01)
    srun hostname
    # Run hostname on all 4 nodes
    srun -N4 hostname
    
Should you use this setup?

It depends entirely on your use case:

  • Recommended for: Test environments, small clusters (≤5 nodes), or situations where you have no extra hardware for a dedicated head node. It’s a cost-effective way to get SLURM up and running without extra equipment.
  • Not recommended for: Production clusters with heavy workloads. The slurmctld daemon handles job scheduling, state tracking, and cluster management—if it’s sharing resources with compute jobs, you might face:
    • Delayed job scheduling or slow sinfo/squeue responses
    • Reduced performance for compute tasks (since the controller is competing for CPU/memory)
    • Higher risk of cluster downtime if the node crashes (it takes down both the controller and a worker)
  • If you must use this in production, mitigate risks by:
    • Reserving CPU/memory resources for slurmctld (using cgroups or taskset)
    • Setting up automatic restarts for slurmctld to handle crashes
    • Monitoring the node’s resource usage closely to avoid overload

内容的提问来源于stack exchange,提问作者ajthealchemist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:54:59