部署BitsAndBytes bf16量化Gemma模型遇双错误求助
问题描述
在Hugging Face Spaces使用Gradio运行BitsAndBytes bf16量化的Gemma-2-2b模型时,遇到两个关键错误:
- RuntimeError:未使用的kwargs
['_load_in_4bit', '_load_in_8bit', 'quant_method'],这些参数未被BitsAndBytesConfig支持 - AttributeError:
'frozenset' object has no attribute 'discard',来自bitsandbytes库中对available_devices的操作
错误信息片段
Runtime Error
退出码:1。原因:未使用的kwargs: ['_load_in_4bit', '_load_in_8bit', 'quant_method']。这些kwargs未在<class 'transformers.utils.quantization_config.BitsAndBytesConfig'>中使用。
...
available_devices.discard("cpu") ** # Only Intel CPU is supported by BNB at the moment
AttributeError: 'frozenset' object has no attribute 'discard'
现有代码(app.py)
import gradio as gr from huggingface_hub import InferenceClient from transformers import AutoTokenizer, AutoModelForCausalLM css = """ html, body { margin: 0; padding: 0; height: 100%; overflow: hidden; } body::before { content: ''; position: fixed; top: 0; left: 0; width: 100vw; height: 100vh; background-image: url('https://png.pngtree.com/background/20230413/original/pngtree-medical-color-cartoon-blank-background-picture-image_2422159.jpg'); background-size: cover; background-repeat: no-repeat; opacity: 0.60; background-position: center; z-index: -1; } .gradio-container { display: flex; flex-direction: column; justify-content: center; align-items: center; height: 100vh; } """ model_id = "harishnair04/Gemma-medtr-2b-sft" tokenizer = AutoTokenizer.from_pretrained(model_id) gemma_model = AutoModelForCausalLM.from_pretrained(model_id) tokenizer.pad_token_id = tokenizer.eos_token_id def respond(input1): template = "Instruction:\n{instruction}\n\nResponse:\n{response}" inputs = tokenizer(input1, return_tensors="pt") out = gemma_model.generate(**inputs, temperature=0.4, do_sample=True, max_new_tokens=200) return tokenizer.decode(out[0], skip_special_tokens=True) chat_interface = gr.Interface( respond, inputs="text", outputs="text", title="MT CHAT", description="Gemma 2b finetuned on medical transcripts", css=css ) chat_interface.launch()
当前依赖(requirements.txt)
huggingface_hub==0.25.2 gradio transformers>=4.45.1 torch gguf>=0.10.0 sentencepiece numpy==1.21.0 bitsandbytes accelerate>=0.26.0 https://github.com/bitsandbytes-foundation/bitsandbytes/releases/download/continuous-release_multi-backend-refactor/bitsandbytes-0.44.1.dev0-py3-none-manylinux_2_24_x86_64.whl
已尝试操作
- 检查依赖版本兼容性
- 排查未使用的量化参数警告
- 移除fp16设置尝试避免精度冲突
解决方案
1. 修正依赖版本,解决frozenset属性错误
使用bitsandbytes稳定版替换开发版,同时匹配兼容的transformers和accelerate版本,避免库内部的属性错误。更新后的requirements.txt如下:
huggingface_hub==0.25.2 gradio==4.44.0 transformers==4.44.2 torch==2.1.2 gguf>=0.10.0 sentencepiece==0.1.99 numpy==1.21.0 bitsandbytes==0.43.1 accelerate==0.24.1
说明:bitsandbytes 0.44.1.dev0开发版存在available_devices被设为frozenset的bug,稳定版0.43.1可解决该问题;同时指定transformers和accelerate的兼容版本,避免参数不匹配问题。
2. 显式配置BitsAndBytesConfig,消除未使用kwargs错误
在加载模型时,明确声明量化配置,覆盖模型自带的冲突参数。修改app.py中模型加载部分:
import gradio as gr from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig import torch # ... 保留原有css代码 ... model_id = "harishnair04/Gemma-medtr-2b-sft" tokenizer = AutoTokenizer.from_pretrained(model_id) # 显式配置bf16量化参数 bnb_config = BitsAndBytesConfig( load_in_4bit=False, load_in_8bit=False, bnb_4bit_use_double_quant=False, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16 ) gemma_model = AutoModelForCausalLM.from_pretrained( model_id, quantization_config=bnb_config, device_map="auto", trust_remote_code=True ) # 确保pad_token正确设置 tokenizer.pad_token = tokenizer.eos_token tokenizer.pad_token_id = tokenizer.eos_token_id # ... 保留原有respond函数和Gradio启动代码 ...
说明:显式指定量化配置,覆盖模型默认的_load_in_4bit等未被当前BitsAndBytesConfig支持的参数;device_map="auto"自动分配模型到可用设备,trust_remote_code=True确保加载自定义模型配置。
3. 额外验证
- 确保Hugging Face Spaces的硬件配置支持bf16(如使用NVIDIA T4/A10G等GPU)
- 启动前清理环境缓存,确保依赖完全重新安装
内容的提问来源于stack exchange,提问作者doniker99

