使用Llama 2时遇RuntimeError:无GPU无法量化求助
Llama 2项目量化报错:No GPU found. A GPU is needed for quantization.
问题详情
在PyCharm开发基于Llama 2的项目时,运行代码触发量化相关错误,核心报错信息:
RuntimeError: No GPU found. A GPU is needed for quantization.
完整错误栈:
Traceback (most recent call last): File "/Users/satarupadeb/miniconda3/lib/python3.11/site-packages/streamlit/runtime/scriptrunner/script_runner.py", line 534, in _run_script exec(code, module.__dict__) File "/Users/satarupadeb/Desktop/investment advisor/app.py", line 39, in <module> tokenizer, model = get_tokenizer_model() ^^^^^^^^^^^^^^^^^^^^^ File "/Users/satarupadeb/miniconda3/lib/python3.11/site-packages/streamlit/runtime/caching/cache_utils.py", line 212, in wrapper return cached_func(*args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/satarupadeb/miniconda3/lib/python3.11/site-packages/streamlit/runtime/caching/cache_utils.py", line 241, in __call__ return self._get_or_create_cached_value(args, kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/satarupadeb/miniconda3/lib/python3.11/site-packages/streamlit/runtime/caching/cache_utils.py", line 267, in _get_or_create_cached_value return self._handle_cache_miss(cache, value_key, func_args, func_kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/satarupadeb/miniconda3/lib/python3.11/site-packages/streamlit/runtime/caching/cache_utils.py", line 321, in _handle_cache_miss computed_value = self._info.func(*func_args, **func_kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/satarupadeb/Desktop/investment advisor/app.py", line 34, in get_tokenizer_model model = AutoModelForCausalLM.from_pretrained(name, cache_dir='./model/' ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/satarupadeb/miniconda3/lib/python3.11/site-packages/transformers/models/auto/auto_factory.py", line 566, in from_pretrained return model_class.from_pretrained( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/satarupadeb/miniconda3/lib/python3.11/site-packages/transformers/modeling_utils.py", line 2897, in from_pretrained raise RuntimeError("No GPU found. A GPU is needed for quantization.") RuntimeError: No GPU found. A GPU is needed for quantization.
触发报错的代码片段:
@st.cache_resource def get_tokenizer_model(): # Create tokenizer tokenizer = AutoTokenizer.from_pretrained(name, cache_dir='./model/', use_auth_token=auth_token) # Create model model = AutoModelForCausalLM.from_pretrained(name, cache_dir='./model/' , use_auth_token=auth_token, torch_dtype=torch.float16, rope_scaling={"type": "dynamic", "factor": 2}, load_in_8bit=True) return tokenizer, model tokenizer, model = get_tokenizer_model()
解决方法
方法1:使用GPU运行
如果设备配备NVIDIA GPU,确保已安装对应版本的CUDA驱动和支持CUDA的PyTorch,load_in_8bit=True的量化操作会自动在GPU上执行,解决报错。
方法2:关闭8-bit量化(适配CPU环境)
无GPU时,直接移除load_in_8bit=True参数,并将torch_dtype改为CPU兼容的torch.float32:
@st.cache_resource def get_tokenizer_model(): tokenizer = AutoTokenizer.from_pretrained(name, cache_dir='./model/', use_auth_token=auth_token) model = AutoModelForCausalLM.from_pretrained(name, cache_dir='./model/' , use_auth_token=auth_token, torch_dtype=torch.float32, rope_scaling={"type": "dynamic", "factor": 2}) return tokenizer, model
注意:CPU运行大模型速度极慢,建议优先使用GPU,或切换到参数更小的Llama 2模型(如7B版本)。
方法3:CPU环境下启用8-bit量化(bitsandbytes支持)
若想在CPU上实现量化,需安装最新版bitsandbytes,并通过BitsAndBytesConfig配置量化参数:
from transformers import BitsAndBytesConfig, AutoTokenizer, AutoModelForCausalLM @st.cache_resource def get_tokenizer_model(): bnb_config = BitsAndBytesConfig( load_in_8bit=True, llm_int8_enable_fp32_cpu_offload=True ) tokenizer = AutoTokenizer.from_pretrained(name, cache_dir='./model/', use_auth_token=auth_token) model = AutoModelForCausalLM.from_pretrained(name, cache_dir='./model/' , use_auth_token=auth_token, torch_dtype=torch.float16, rope_scaling={"type": "dynamic", "factor": 2}, quantization_config=bnb_config, device_map='cpu') return tokenizer, model
该方案性能仍远不及GPU,仅作为临时替代方案。
内容的提问来源于stack exchange,提问作者MIMI
相关产品推荐
相关产品推荐

