如何通过Slurm运行NVSHMEM?测试代码PE识别异常求助
NVSHMEM PE识别异常排查
问题描述
刚上手NVSHMEM,编写了简单测试代码,通过Slurm批处理脚本编译运行。预期输出8个PE的启动信息(各PE编号从0到7,总PE数为8),但实际所有进程均显示自身为第0个PE,总PE数为1。
测试代码
#include <nvshmem.h> #include <stdio.h> int main(int argc, char *argv[]) { // 初始化NVSHMEM库 nvshmem_init(); int mype = nvshmem_my_pe(); int npes = nvshmem_n_pes(); fprintf(stdout, "PE %d of %d has started ...\n", mype, npes); // 结束NVSHMEM nvshmem_finalize(); return 0; }
Slurm批处理脚本
#!/bin/bash -l #SBATCH --nodes=2 # 节点数 #SBATCH --ntasks=8 # 任务数 #SBATCH --ntasks-per-node=4 # 每节点任务数 #SBATCH --gpus-per-task=1 # 每个任务对应GPU数 #SBATCH --cpus-per-task=1 # 每个任务对应CPU核心数 #SBATCH --time=00:15:00 # 运行时间(HH:MM:SS) #SBATCH --partition=gpu # 分区 #SBATCH --account=p200301 # 项目账号 #SBATCH --qos=default # QoS级别 module load NCCL OpenMPI CUDA NVSHMEM && nvcc -rdc=true -ccbin g++ -I $NVSHMEM_HOME/include test.cu -o test -L $NVSHMEM_HOME/lib -lnvshmem_host -lnvshmem_device -lucs -lucp && srun -n 8 ./test
输出对比
预期输出
PE 0 of 8 has started ... PE 1 of 8 has started ... PE 2 of 8 has started ... ... PE 7 of 8 has started ...
实际输出
PE 0 of 1 has started ... PE 0 of 1 has started ... ...(共8行相同内容)
排查及解决方法
1. 调整编译方式
NVSHMEM依赖MPI实现进程间通信,直接用nvcc编译可能未正确关联MPI环境。建议使用NVSHMEM提供的nvshmem_compile脚本编译,该脚本会自动处理依赖库和编译选项:
nvshmem_compile -o test test.cu
或者用mpicc配合NVSHMEM的编译参数:
mpicc -rdc=true -ccbin g++ -I $NVSHMEM_HOME/include test.cu -o test -L $NVSHMEM_HOME/lib -lnvshmem -lcudart
2. 设置运行时环境变量
启动程序前设置NVSHMEM_USE_MPI=1,强制NVSHMEM使用MPI作为后端启动方式,确保进程能形成统一的PE集群:
在Slurm脚本的运行步骤前添加:
export NVSHMEM_USE_MPI=1
3. 分离编译与运行步骤
将脚本中的编译和运行分开执行,避免环境变量传递异常,修改后的脚本片段:
module load NCCL OpenMPI CUDA NVSHMEM # 编译代码 nvshmem_compile -o test test.cu # 设置环境变量并运行 export NVSHMEM_USE_MPI=1 srun -n 8 ./test
4. 验证MPI环境可用性
先单独运行MPI测试程序(如mpiexec -n 8 hostname),确认Slurm能正确启动多进程并分配资源,排除MPI环境本身的问题。
内容的提问来源于stack exchange,提问作者LukeTheWalker
相关产品推荐
相关产品推荐

