SLURM未遵循请求资源:提交任务脚本后状态查询异常求助
Let's walk through how to figure out why SLURM might not be honoring your requested resources, even though your job shows as RUNNING.
Step 1: Verify Actual Allocated Resources
The scontrol show job snippet you shared only shows the job state—we need to see the full resource allocation details to compare against your request. Run this command with your job ID to pull key resource info:
scontrol show job <your-job-id> | grep -E 'NodeList|NumNodes|NumTasks|CPUs|Partition'
Look for values like NumNodes=1, NumTasks=1, and CPUs=1 (if you expect a single CPU). If these don't match your #SBATCH directives, that confirms SLURM is deviating from your request.
Step 2: Check Partition Defaults
Your job runs in the debug partition, which might have built-in default resource settings that override your explicit requests. To view the partition's configuration:
sinfo -p debug -o "%P %N %t %C %O"
Pay close attention to the last column (%O), which lists partition-specific defaults (like DefaultCPUsPerTask or DefaultTasksPerNode). If the partition has a default number of CPUs/tasks higher than your request, SLURM might be using that value instead.
Step 3: Explicitly Define All Critical Resources
SLURM often falls back to cluster or partition defaults when resource requests are incomplete. Update your test.sub script to explicitly specify every resource you need, leaving no room for default overrides:
#!/bin/bash #SBATCH --workdir=./ #SBATCH -o test.out #SBATCH --partition=debug #SBATCH --nodes=1 #SBATCH --ntasks=1 #SBATCH --cpus-per-task=1 # Critical: Explicitly set 1 CPU per task #SBATCH --mem-per-cpu=100M # Optional: Add memory limits if needed #SBATCH --requeue #SBATCH --job-name=test x=0 while [ $x -le 100 ]; do echo "Test $x" >> test.out sleep 100 x=$(($x+1)) done
Adding --cpus-per-task=1 forces SLURM to allocate exactly one CPU to your task, aligning perfectly with your --ntasks=1 request.
Step 4: Validate the Fix
Resubmit your updated script with sbatch test.sub, then re-run the scontrol show job command to confirm the allocated resources now match your request. You can also check the job's actual resource usage with:
top -u $USER -p <job-pid>
This will show you if the job is only using the CPU and memory you requested.
Bonus: Check Cluster Policies
If the above steps don't resolve the issue, reach out to your cluster administrator to verify if there are cluster-wide policies (like minimum resource allocations for the debug partition) that are overriding your requests. Administrators can also check SLURM logs for deeper context on resource allocation decisions.
内容的提问来源于stack exchange,提问作者ajthealchemist

