You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SLURM未遵循请求资源:提交任务脚本后状态查询异常求助

Troubleshooting SLURM Resource Request Discrepancy

Let's walk through how to figure out why SLURM might not be honoring your requested resources, even though your job shows as RUNNING.

Step 1: Verify Actual Allocated Resources

The scontrol show job snippet you shared only shows the job state—we need to see the full resource allocation details to compare against your request. Run this command with your job ID to pull key resource info:

scontrol show job <your-job-id> | grep -E 'NodeList|NumNodes|NumTasks|CPUs|Partition'

Look for values like NumNodes=1, NumTasks=1, and CPUs=1 (if you expect a single CPU). If these don't match your #SBATCH directives, that confirms SLURM is deviating from your request.

Step 2: Check Partition Defaults

Your job runs in the debug partition, which might have built-in default resource settings that override your explicit requests. To view the partition's configuration:

sinfo -p debug -o "%P %N %t %C %O"

Pay close attention to the last column (%O), which lists partition-specific defaults (like DefaultCPUsPerTask or DefaultTasksPerNode). If the partition has a default number of CPUs/tasks higher than your request, SLURM might be using that value instead.

Step 3: Explicitly Define All Critical Resources

SLURM often falls back to cluster or partition defaults when resource requests are incomplete. Update your test.sub script to explicitly specify every resource you need, leaving no room for default overrides:

#!/bin/bash
#SBATCH --workdir=./
#SBATCH -o test.out
#SBATCH --partition=debug
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=1  # Critical: Explicitly set 1 CPU per task
#SBATCH --mem-per-cpu=100M  # Optional: Add memory limits if needed
#SBATCH --requeue
#SBATCH --job-name=test
x=0
while [ $x -le 100 ]; do
    echo "Test $x" >> test.out
    sleep 100
    x=$(($x+1))
done

Adding --cpus-per-task=1 forces SLURM to allocate exactly one CPU to your task, aligning perfectly with your --ntasks=1 request.

Step 4: Validate the Fix

Resubmit your updated script with sbatch test.sub, then re-run the scontrol show job command to confirm the allocated resources now match your request. You can also check the job's actual resource usage with:

top -u $USER -p <job-pid>

This will show you if the job is only using the CPU and memory you requested.

Bonus: Check Cluster Policies

If the above steps don't resolve the issue, reach out to your cluster administrator to verify if there are cluster-wide policies (like minimum resource allocations for the debug partition) that are overriding your requests. Administrators can also check SLURM logs for deeper context on resource allocation decisions.

内容的提问来源于stack exchange,提问作者ajthealchemist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:27:31