SLURM控制器与Worker节点共用同一机器是否可行?如何部署及是否建议?
Is this deployment feasible?
Yes, 100% — SLURM explicitly supports having a single node act as both the controller (slurmctld) and a worker (slurmd). This is a common setup for small-scale clusters, test environments, or situations where hardware resources are limited. Lots of folks run SLURM this way when they don’t have extra hardware for a dedicated head node.
How to implement this setup
Here’s a practical, step-by-step guide to get this up and running:
1. Install SLURM packages
- On the node that will serve as both controller and worker: Install the full SLURM stack, including the controller daemon, worker daemon, and command-line tools.
- For Debian/Ubuntu:
sudo apt install slurm-wlm slurm-wlm-basic-plugins - For RHEL/CentOS/Rocky:
sudo dnf install slurm-slurmctld slurm-slurmd slurm-client
- For Debian/Ubuntu:
- On your other 3 worker-only nodes: Install just the worker daemon and client tools (skip the controller package).
2. Configure SLURM and authentication
a. Set up slurm.conf
You’ll need a shared slurm.conf file across all nodes. Here’s a minimal example tailored to your 4-node setup:
ControlMachine=node01 # Replace with your controller/worker node's hostname ControlAddr=192.168.1.100 # Optional: IP of the control node for reliability SlurmUser=slurm MpiDefault=none ProctrackType=proctrack/cgroup ReturnToService=1 SlurmctldPidFile=/var/run/slurmctld.pid SlurmdPidFile=/var/run/slurmd.pid SlurmdSpoolDir=/var/spool/slurmd SlurmctldSpoolDir=/var/spool/slurmctld StateSaveLocation=/var/spool/slurmctld SlurmctldTimeout=300 SlurmdTimeout=300 InactiveLimit=0 MinJobAge=300 KillWait=30 Waittime=0 # Define all nodes, including your control/worker node NodeName=node01 CPUs=8 State=UNKNOWN # Adjust CPUs to match your node's core count NodeName=node02 CPUs=8 State=UNKNOWN NodeName=node03 CPUs=8 State=UNKNOWN NodeName=node04 CPUs=8 State=UNKNOWN # Define a default partition that includes all nodes PartitionName=normal Nodes=ALL Default=YES MaxTime=INFINITE State=UP
- Replace
node01, the IP, and CPU counts with your actual node details. - Copy this file to
/etc/slurm/slurm.confon all nodes, then set permissions:sudo chmod 644 /etc/slurm/slurm.conf && sudo chown root:root /etc/slurm/slurm.conf
b. Set up MUNGE (authentication)
SLURM uses MUNGE for secure node-to-node communication:
- Install MUNGE on all nodes:
- Debian/Ubuntu:
sudo apt install munge - RHEL/CentOS/Rocky:
sudo dnf install munge
- Debian/Ubuntu:
- On your control/worker node, generate a MUNGE key:
sudo dd if=/dev/urandom bs=1 count=1024 > /etc/munge/munge.key - Copy this key to all other worker nodes (use
scpor a shared filesystem):sudo scp /etc/munge/munge.key node02:/etc/munge/ - On all nodes, set the correct permissions for the key:
sudo chmod 400 /etc/munge/munge.key sudo chown munge:munge /etc/munge/munge.key - Start and enable MUNGE on all nodes:
sudo systemctl enable --now munge
3. Start SLURM services
- On your control/worker node: Start both the controller and worker daemons:
sudo systemctl enable --now slurmctld slurmd - On your worker-only nodes: Start just the worker daemon:
sudo systemctl enable --now slurmd
4. Verify the setup
- Run
sinfoto check if all nodes are in theUPstate:sinfo - Submit test jobs to confirm everything works:
# Run a hostname command on 1 node (will run on any available worker, including node01) srun hostname # Run hostname on all 4 nodes srun -N4 hostname
Should you use this setup?
It depends entirely on your use case:
- Recommended for: Test environments, small clusters (≤5 nodes), or situations where you have no extra hardware for a dedicated head node. It’s a cost-effective way to get SLURM up and running without extra equipment.
- Not recommended for: Production clusters with heavy workloads. The
slurmctlddaemon handles job scheduling, state tracking, and cluster management—if it’s sharing resources with compute jobs, you might face:- Delayed job scheduling or slow
sinfo/squeueresponses - Reduced performance for compute tasks (since the controller is competing for CPU/memory)
- Higher risk of cluster downtime if the node crashes (it takes down both the controller and a worker)
- Delayed job scheduling or slow
- If you must use this in production, mitigate risks by:
- Reserving CPU/memory resources for
slurmctld(using cgroups ortaskset) - Setting up automatic restarts for
slurmctldto handle crashes - Monitoring the node’s resource usage closely to avoid overload
- Reserving CPU/memory resources for
内容的提问来源于stack exchange,提问作者ajthealchemist
相关产品推荐
相关产品推荐

