You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Slurm忽略sbatch中--distribution=cyclic参数,采用block分配问题求助

Troubleshooting Slurm Cyclic Distribution Ignored After Node Image Update

Let's break down what's happening here and walk through actionable steps to fix the task distribution issue you're seeing.

First, Understand the Context

You're using Bright Cluster Manager, which manages Slurm services automatically—so manually stopping slurmctld won't work because the cluster monitor will immediately restart it. This is likely part of why your configuration changes aren't taking effect. Let's start with the most probable fixes:


1. Use Bright Cluster's Tools to Apply Slurm Configuration Changes

Bright Cluster Manager maintains its own configuration database, so editing slurm.conf directly might not propagate correctly (or might be overwritten). Instead, use the cmsh tool to update Slurm settings properly:

# Enter the Bright Cluster shell
cmsh
# Navigate to Slurm configuration
slurm configuration
# Verify or set the SelectTypeParameters (match your desired setting)
set SelectTypeParameters CR_Core
# Set the default distribution policy if needed
set DefaultTaskDistribution cyclic:cyclic
# Commit the changes—Bright will sync configs and restart services automatically
commit
exit

This ensures your configuration is applied cluster-wide and avoids conflicts with Bright's service monitoring.


2. Verify Active Slurm Configuration

After applying changes, confirm that the running Slurm configuration matches what you expect. Run this command to check key parameters:

scontrol show config | grep -E "SelectType|SelectTypeParameters|DefaultTaskDistribution|SchedulerType"

Look for:

  • DefaultTaskDistribution = cyclic:cyclic (this should be your default if you didn't set it explicitly)
  • SelectType = select/cons_res and SelectTypeParameters = CR_Core (matches your slurm.conf)

If these don't match, Bright might have overridden your manual slurm.conf edits—using cmsh as above will fix this.


3. Diagnose Job Allocation Details

Let's get concrete data about how your job is being allocated. Submit your test job, then run:

# Replace <JOB_ID> with your actual job ID
scontrol show job <JOB_ID>

Look for these fields:

  • Distribution: Should show cyclic:cyclic if your sbatch parameter took effect
  • NumTasksPerNode: Should list 4,4,4 for your 12-task, 3-node job
  • NodeList: Should include your 3 target nodes

If NumTasksPerNode shows 10,1,1, this tells us Slurm is seeing some constraint that's forcing uneven allocation—likely a mismatch between node resources and Slurm's configured resources.


4. Validate Node Resource Configuration

Check that each node's reported CPU resources match your slurm.conf settings (40 cores total per node):

sinfo -N -o "%N %c %C"

This will output each node's total cores (%c) and used/idle/total cores (%C). If any node shows a different total core count, your node image update might have altered the CPU configuration (e.g., disabled cores, changed hyperthreading). You'll need to fix the node image or update slurm.conf (via cmsh) to match the actual hardware.


5. Test with a Minimal Job

Rule out any hidden parameters in your existing sbatch script by testing a minimal job:

#!/bin/bash
#SBATCH --ntasks=12
#SBATCH --nodes=3
#SBATCH --distribution=cyclic:cyclic
# Print hostname for each task to see distribution
srun hostname | sort

If this job still shows uneven distribution (e.g., 10 instances of compute001), the issue is likely with Slurm's node selection logic. Check if nodes have priority weights set:

scontrol show node compute001 | grep Priority

If compute001 has a higher priority, Slurm's backfill scheduler might be favoring it. You can reset node priorities via cmsh or adjust the scheduler's behavior.


Key Takeaway

The most likely culprit is that your manual slurm.conf edits weren't being applied because Bright Cluster Manager was overriding them or restarting slurmctld before changes took effect. Using cmsh to manage Slurm configuration ensures your settings are propagated correctly across the cluster.

内容的提问来源于stack exchange,提问作者Aaron Caba

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:54:44