You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ubuntu服务器下载Llama2-70b-chat模型报错:设备空间不足求解

Llama2-70b-chat模型下载时出现"No space left on device"错误的解决方法

问题背景

在Ubuntu服务器的Jupyter Notebook中运行PyTorch代码,尝试从Hugging Face下载Llama2-70b-chat模型并保存到本地,用于GPU部署,但运行时报错OSError: [Errno 28] No space left on device。已确认服务器主磁盘空间充足,疑惑是否为GPU内存耗尽导致。

运行代码

from torch import cuda, bfloat16
import transformers

model_id = 'meta-llama/Llama-2-70b-chat-hf'

device = f'cuda:{cuda.current_device()}' if cuda.is_available() else 'cpu'

# set quantization configuration to load large model with less GPU memory
# this requires the `bitsandbytes` library
bnb_config = transformers.BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type='nf4',
    bnb_4bit_use_double_quant=True,
    bnb_4bit_compute_dtype=bfloat16
)

# begin initializing HF items, need auth token for these
# hf_auth = '<YOUR_API_KEY>'

hf_auth = apikey
model_config = transformers.AutoConfig.from_pretrained(
    model_id,
    use_auth_token=hf_auth
)

model = transformers.AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    config=model_config,
    quantization_config=bnb_config,
    device_map='auto',
    use_auth_token=hf_auth
)
model.save_model('/save_path/')

model.eval()
print(f"Model loaded on {device}")

错误信息

File ~/anaconda3/envs/LLMenv/lib/python3.10/site-packages/huggingface_hub/file_download.py:544, in http_get(url, temp_file, proxies, resume_size, headers, timeout, max_retries, expected_size)
    542     if chunk:  # filter out keep-alive new chunks
    543         progress.update(len(chunk))
--> 544         temp_file.write(chunk)
    546 if expected_size is not None and expected_size != temp_file.tell():
    547     raise EnvironmentError(
    548         f"Consistency check failed: file should be of size {expected_size} but has size"
    549         f" {temp_file.tell()} ({displayed_name}).\nWe are sorry for the inconvenience. Please retry download and"
    550         " pass `force_download=True, resume_download=False` as argument.\nIf the issue persists, please let us"
    551         " know by opening an issue on https://github.com/huggingface/huggingface_hub."
    552     )

File ~/anaconda3/envs/LLMenv/lib/python3.10/tempfile.py:483, in _TemporaryFileWrapper.__getattr__.<locals>.func_wrapper(*args, **kwargs)
    481 @_functools.wraps(func)
    482 def func_wrapper(*args, **kwargs):
--> 483     return func(*args, **kwargs)

OSError: [Errno 28] No space left on device

问题分析与解决方法

错误原因

该错误并非GPU内存耗尽,而是Hugging Face下载模型时使用的临时目录磁盘空间不足。默认情况下,huggingface_hub会将临时文件存放在系统临时目录(如/tmp),若该分区空间较小,就会触发此错误。

解决步骤

  • 指定临时目录到空间充足的磁盘
    在代码开头添加环境变量配置,修改临时文件与模型缓存的存储路径:

    import os
    # 替换为服务器上空间充足的路径
    os.environ['HF_HOME'] = '/path/to/large/disk/huggingface_cache'
    os.environ['TMPDIR'] = '/path/to/large/disk/tmp'
    

    这样模型下载的缓存和临时文件都会存储到指定的大空间目录。

  • 直接指定模型缓存目录
    在from_pretrained方法中添加cache_dir参数,明确缓存路径:

    model_config = transformers.AutoConfig.from_pretrained(
        model_id,
        use_auth_token=hf_auth,
        cache_dir='/path/to/large/disk/huggingface_cache'
    )
    
    model = transformers.AutoModelForCausalLM.from_pretrained(
        model_id,
        trust_remote_code=True,
        config=model_config,
        quantization_config=bnb_config,
        device_map='auto',
        use_auth_token=hf_auth,
        cache_dir='/path/to/large/disk/huggingface_cache'
    )
    
  • 临时清理系统临时目录
    若不想修改路径,可手动清理系统临时目录的旧文件:

    sudo rm -rf /tmp/*
    

    但此方法仅能临时解决,后续大模型下载仍可能触发空间不足问题。

  • 确认GPU内存适配
    本次错误虽与GPU内存无关,但需注意:Llama2-70b-chat经4bit量化后,所需GPU内存约为15-20GB,需确保服务器GPU显存足够。若显存不足,会出现CUDA out of memory错误,而非磁盘空间错误。

内容的提问来源于stack exchange,提问作者user3476463

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 09:02:09