Ubuntu服务器下载Llama2-70b-chat模型报错:设备空间不足求解
Llama2-70b-chat模型下载时出现"No space left on device"错误的解决方法
问题背景
在Ubuntu服务器的Jupyter Notebook中运行PyTorch代码,尝试从Hugging Face下载Llama2-70b-chat模型并保存到本地,用于GPU部署,但运行时报错OSError: [Errno 28] No space left on device。已确认服务器主磁盘空间充足,疑惑是否为GPU内存耗尽导致。
运行代码
from torch import cuda, bfloat16 import transformers model_id = 'meta-llama/Llama-2-70b-chat-hf' device = f'cuda:{cuda.current_device()}' if cuda.is_available() else 'cpu' # set quantization configuration to load large model with less GPU memory # this requires the `bitsandbytes` library bnb_config = transformers.BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type='nf4', bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=bfloat16 ) # begin initializing HF items, need auth token for these # hf_auth = '<YOUR_API_KEY>' hf_auth = apikey model_config = transformers.AutoConfig.from_pretrained( model_id, use_auth_token=hf_auth ) model = transformers.AutoModelForCausalLM.from_pretrained( model_id, trust_remote_code=True, config=model_config, quantization_config=bnb_config, device_map='auto', use_auth_token=hf_auth ) model.save_model('/save_path/') model.eval() print(f"Model loaded on {device}")
错误信息
File ~/anaconda3/envs/LLMenv/lib/python3.10/site-packages/huggingface_hub/file_download.py:544, in http_get(url, temp_file, proxies, resume_size, headers, timeout, max_retries, expected_size) 542 if chunk: # filter out keep-alive new chunks 543 progress.update(len(chunk)) --> 544 temp_file.write(chunk) 546 if expected_size is not None and expected_size != temp_file.tell(): 547 raise EnvironmentError( 548 f"Consistency check failed: file should be of size {expected_size} but has size" 549 f" {temp_file.tell()} ({displayed_name}).\nWe are sorry for the inconvenience. Please retry download and" 550 " pass `force_download=True, resume_download=False` as argument.\nIf the issue persists, please let us" 551 " know by opening an issue on https://github.com/huggingface/huggingface_hub." 552 ) File ~/anaconda3/envs/LLMenv/lib/python3.10/tempfile.py:483, in _TemporaryFileWrapper.__getattr__.<locals>.func_wrapper(*args, **kwargs) 481 @_functools.wraps(func) 482 def func_wrapper(*args, **kwargs): --> 483 return func(*args, **kwargs) OSError: [Errno 28] No space left on device
问题分析与解决方法
错误原因
该错误并非GPU内存耗尽,而是Hugging Face下载模型时使用的临时目录磁盘空间不足。默认情况下,huggingface_hub会将临时文件存放在系统临时目录(如/tmp),若该分区空间较小,就会触发此错误。
解决步骤
指定临时目录到空间充足的磁盘
在代码开头添加环境变量配置,修改临时文件与模型缓存的存储路径:import os # 替换为服务器上空间充足的路径 os.environ['HF_HOME'] = '/path/to/large/disk/huggingface_cache' os.environ['TMPDIR'] = '/path/to/large/disk/tmp'这样模型下载的缓存和临时文件都会存储到指定的大空间目录。
直接指定模型缓存目录
在from_pretrained方法中添加cache_dir参数,明确缓存路径:model_config = transformers.AutoConfig.from_pretrained( model_id, use_auth_token=hf_auth, cache_dir='/path/to/large/disk/huggingface_cache' ) model = transformers.AutoModelForCausalLM.from_pretrained( model_id, trust_remote_code=True, config=model_config, quantization_config=bnb_config, device_map='auto', use_auth_token=hf_auth, cache_dir='/path/to/large/disk/huggingface_cache' )临时清理系统临时目录
若不想修改路径,可手动清理系统临时目录的旧文件:sudo rm -rf /tmp/*但此方法仅能临时解决,后续大模型下载仍可能触发空间不足问题。
确认GPU内存适配
本次错误虽与GPU内存无关,但需注意:Llama2-70b-chat经4bit量化后,所需GPU内存约为15-20GB,需确保服务器GPU显存足够。若显存不足,会出现CUDA out of memory错误,而非磁盘空间错误。
内容的提问来源于stack exchange,提问作者user3476463
相关产品推荐
相关产品推荐

