Slurm忽略sbatch中--distribution=cyclic参数,采用block分配问题求助
Let's break down what's happening here and walk through actionable steps to fix the task distribution issue you're seeing.
First, Understand the Context
You're using Bright Cluster Manager, which manages Slurm services automatically—so manually stopping slurmctld won't work because the cluster monitor will immediately restart it. This is likely part of why your configuration changes aren't taking effect. Let's start with the most probable fixes:
1. Use Bright Cluster's Tools to Apply Slurm Configuration Changes
Bright Cluster Manager maintains its own configuration database, so editing slurm.conf directly might not propagate correctly (or might be overwritten). Instead, use the cmsh tool to update Slurm settings properly:
# Enter the Bright Cluster shell cmsh # Navigate to Slurm configuration slurm configuration # Verify or set the SelectTypeParameters (match your desired setting) set SelectTypeParameters CR_Core # Set the default distribution policy if needed set DefaultTaskDistribution cyclic:cyclic # Commit the changes—Bright will sync configs and restart services automatically commit exit
This ensures your configuration is applied cluster-wide and avoids conflicts with Bright's service monitoring.
2. Verify Active Slurm Configuration
After applying changes, confirm that the running Slurm configuration matches what you expect. Run this command to check key parameters:
scontrol show config | grep -E "SelectType|SelectTypeParameters|DefaultTaskDistribution|SchedulerType"
Look for:
DefaultTaskDistribution = cyclic:cyclic(this should be your default if you didn't set it explicitly)SelectType = select/cons_resandSelectTypeParameters = CR_Core(matches your slurm.conf)
If these don't match, Bright might have overridden your manual slurm.conf edits—using cmsh as above will fix this.
3. Diagnose Job Allocation Details
Let's get concrete data about how your job is being allocated. Submit your test job, then run:
# Replace <JOB_ID> with your actual job ID scontrol show job <JOB_ID>
Look for these fields:
Distribution: Should showcyclic:cyclicif your sbatch parameter took effectNumTasksPerNode: Should list4,4,4for your 12-task, 3-node jobNodeList: Should include your 3 target nodes
If NumTasksPerNode shows 10,1,1, this tells us Slurm is seeing some constraint that's forcing uneven allocation—likely a mismatch between node resources and Slurm's configured resources.
4. Validate Node Resource Configuration
Check that each node's reported CPU resources match your slurm.conf settings (40 cores total per node):
sinfo -N -o "%N %c %C"
This will output each node's total cores (%c) and used/idle/total cores (%C). If any node shows a different total core count, your node image update might have altered the CPU configuration (e.g., disabled cores, changed hyperthreading). You'll need to fix the node image or update slurm.conf (via cmsh) to match the actual hardware.
5. Test with a Minimal Job
Rule out any hidden parameters in your existing sbatch script by testing a minimal job:
#!/bin/bash #SBATCH --ntasks=12 #SBATCH --nodes=3 #SBATCH --distribution=cyclic:cyclic # Print hostname for each task to see distribution srun hostname | sort
If this job still shows uneven distribution (e.g., 10 instances of compute001), the issue is likely with Slurm's node selection logic. Check if nodes have priority weights set:
scontrol show node compute001 | grep Priority
If compute001 has a higher priority, Slurm's backfill scheduler might be favoring it. You can reset node priorities via cmsh or adjust the scheduler's behavior.
Key Takeaway
The most likely culprit is that your manual slurm.conf edits weren't being applied because Bright Cluster Manager was overriding them or restarting slurmctld before changes took effect. Using cmsh to manage Slurm configuration ensures your settings are propagated correctly across the cluster.
内容的提问来源于stack exchange,提问作者Aaron Caba

