You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Slurm集群运行Python代码超时停滞问题求助

问题:Slurm集群运行Hugging Face代码超时,本地正常执行

我基于Hugging Face的datasets、transformers库写了Python代码,用来处理GLUE MRPC数据集并加载Roberta模型。本地Linux环境跑完全程只要2分钟,但提交到Slurm集群后,跑了4小时还没完成,最后因为超时被取消。日志只输出了tokenizer相关的警告,后面的逻辑根本没执行,求帮忙排查解决。

我的代码:

from datasets import load_dataset
MAX_LEN = 512
dataset = load_dataset("glue","mrpc")
from transformers import AutoTokenizer
from transformers import RobertaTokenizerFast

#tokenizer =AutoTokenizer.from_pretrained("bert-base-uncased")
tokenizer = RobertaTokenizerFast.from_pretrained("/data/home//raw_roberta/Roberta_Tokenizer", max_length=MAX_LEN, padding='max_length', return_tensors='pt')

print("mapped_dataset")

mapped_dataset = dataset.map(lambda x: tokenizer(x["sentence1"], x["sentence2"], max_length = MAX_LEN, truncation=True, padding='max_length', return_tensors='pt'), batched=True)

print("completeed mapped_dataset")

from transformers import DataCollatorWithPadding
data_collator= DataCollatorWithPadding(tokenizer=tokenizer)

from transformers import AutoModelForSequenceClassification
from transformers import RobertaForMaskedLM

#model = AutoModelForSequenceClassification.from_pretrained("bert-base-uncased",num_labels = 2)
#model = AutoModelForSequenceClassification.from_pretrained("data/home//raw_roberta/Roberta_Model/checkpoint-90000", num_labels = 2)
#model = AutoModelForSequenceClassification.from_pretrained("data/home//raw_roberta/Roberta_Model/checkpoint-90000")
base_model = RobertaForMaskedLM.from_pretrained('/data/home//raw_roberta/Roberta_Model/checkpoint-90000').roberta

from transformers import TrainingArguments
print(base_model.config)

本地运行日志(耗时约2分钟):

The tokenizer class you load from this checkpoint is not the same type as the class this function is called from. It may result in unexpected tokenization.
The tokenizer class you load from this checkpoint is 'BertTokenizer'.
The class this function is called from is 'RobertaTokenizer'.
Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.
The tokenizer class you load from this checkpoint is not the same type as the class this function is called from. It may result in unexpected tokenization.
The tokenizer class you load from this checkpoint is 'BertTokenizer'.
The class this function is called from is 'RobertaTokenizerFast'.
Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.
mapped_dataset
completeed mapped_dataset
RobertaConfig {
  "_name_or_path": "/data/home//raw_roberta/Roberta_Model/checkpoint-90000",
  "architectures": [
    "RobertaForMaskedLM"
  ],
  "attention_probs_dropout_prob": 0.1,
  "bos_token_id": 0,
  "classifier_dropout": null,
  "eos_token_id": 2,
  "hidden_act": "gelu",
  "hidden_dropout_prob": 0.1,
  "hidden_size": 768,
  "initializer_range": 0.02,
  "intermediate_size": 3072,
  "layer_norm_eps": 1e-12,
  "max_position_embeddings": 514,
  "model_type": "roberta",
  "num_attention_heads": 12,
  "num_hidden_layers": 12,
  "pad_token_id": 1,
  "position_embedding_type": "absolute",
  "torch_dtype": "float32",
  "transformers_version": "4.33.2",
  "type_vocab_size": 1,
  "use_cache": true,
  "vocab_size": 52000
}

Slurm集群运行日志(运行4小时后超时取消):

/data/home//anaconda3/envs/py38v1/lib/python3.8/site-packages/scipy/__init__.py:138: UserWarning: A NumPy version >=1.16.5 and <1.23.0 is required for this version of SciPy (detected version 1.24.4)
  warnings.warn(f"A NumPy version >={np_minversion} and <{np_maxversion} is required for this version of "
The tokenizer class you load from this checkpoint is not the same type as the class this function is called from. It may result in unexpected tokenization. 
The tokenizer class you load from this checkpoint is 'BertTokenizer'. 
The class this function is called from is 'RobertaTokenizer'.
Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.
The tokenizer class you load from this checkpoint is not the same type as the class this function is called from. It may result in unexpected tokenization. 
The tokenizer class you load from this checkpoint is 'BertTokenizer'. 
The class this function is called from is 'RobertaTokenizerFast'.
Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.
slurmstepd: error: *** JOB xxxxx ON compute-9-0 CANCELLED AT 2024-02-05T08:35:30 DUE TO TIME LIMIT ***

排查与解决方法

1. 数据集加载的网络阻塞

本地大概率已经缓存了GLUE MRPC数据集,但Slurm集群节点可能无法访问Hugging Face的数据集服务器,导致load_dataset("glue","mrpc")一直卡在下载环节。

  • 解决:把本地缓存的数据集复制到集群的共享存储路径,然后指定本地路径加载:
    # 替换成集群上的数据集路径
    dataset = load_dataset("glue", "mrpc", data_dir="/cluster/shared/datasets/glue/mrpc")
    
    或者提前在集群节点手动下载好数据集,确保路径有读取权限。

2. Tokenizer映射的性能瓶颈

代码里dataset.map用了batched=True但没指定batch_size,默认批次可能过大,加上集群节点内存/CPU资源不足,导致处理速度极慢。另外return_tensors='pt'会让map返回PyTorch张量,占用大量内存,进一步拖慢速度。

  • 解决:给map指定合理的批次大小,同时去掉return_tensors='pt'(后续用DataCollatorWithPadding处理更高效):
    mapped_dataset = dataset.map(
        lambda x: tokenizer(x["sentence1"], x["sentence2"], max_length=MAX_LEN, truncation=True, padding='max_length'),
        batched=True,
        batch_size=64  # 根据节点资源调整,比如32、64、128
    )
    

3. 模型路径解析异常

代码里的模型路径有双斜杠/data/home//raw_roberta/...,部分环境下可能导致路径解析错误,集群节点无法正确读取模型文件,进而卡在加载环节。

  • 解决:把路径改成单斜杠/data/home/raw_roberta/Roberta_Model/checkpoint-90000,同时确认集群节点对该路径有读取权限。

4. 环境依赖版本不兼容

集群日志显示NumPy版本(1.24.4)不符合SciPy的要求(>=1.16.5且<1.23.0),这种依赖不兼容可能导致底层库出现隐性阻塞或错误。

  • 解决:在集群上创建和本地一致的conda环境,安装相同版本的依赖:
    conda create -n py38v1 python=3.8
    conda activate py38v1
    # 替换成本地使用的具体版本号
    pip install datasets==2.14.5 transformers==4.33.2 torch==2.0.1 numpy==1.22.4 scipy==1.10.1
    

5. Slurm作业资源配置不足

提交Slurm作业时如果没指定足够的CPU核心、内存,代码会因为资源不够运行缓慢,最终超时。

  • 解决:在Slurm脚本里增加资源配置,比如:
    #SBATCH --nodes=1
    #SBATCH --ntasks-per-node=4  # 核心数,根据需求调整
    #SBATCH --mem=16G  # 内存,根据模型大小调整
    #SBATCH --time=01:00:00  # 超时时间,设置足够完成的时长
    

内容的提问来源于stack exchange,提问作者user23352539

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 17:51:01