如何通过配置将Hugging Face小模型(如小型Llama2)转为bfloat16?
解决配置设置bfloat16无效的问题
你的问题核心是修改config的时机错误:你先调用AutoModelForCausalLM.from_config(config)创建了模型,之后才去设置config.torch_dtype,此时模型已经完成初始化,修改配置根本不会影响已创建的模型参数 dtype。
正确处理方式
在调用from_config创建模型之前,就把torch_dtype配置好,这样模型初始化时会直接用指定的 dtype 生成参数。
修改后的代码示例
def get_smaller_llama2(hidden_size : int = 2048, num_hidden_layers : int = 12, return_tokenizer: bool = False, verbose : bool = False, ): config = AutoConfig.from_pretrained("meta-llama/Llama-2-7b-hf") # 先调整模型结构参数 config.hidden_size = hidden_size config.num_hidden_layers = num_hidden_layers # 关键步骤:在创建模型前设置torch_dtype config.torch_dtype = torch.bfloat16 if (torch.cuda.is_available() and torch.cuda.get_device_capability(torch.cuda.current_device())[0] >= 8) else torch.float32 # 可选:在from_config时额外指定torch_dtype,双重保险 smaller_model = AutoModelForCausalLM.from_config(config, torch_dtype=config.torch_dtype) device = torch.device(f"cuda:{0}" if torch.cuda.is_available() else "cpu") # 模型已为目标dtype,直接移动到设备即可 smaller_model = smaller_model.to(device) print(f'Model dtype: {next(iter(smaller_model.parameters())).dtype}') if verbose: print(f'config: {config}') print("Smaller number of parameters:", sum(p.numel() for p in smaller_model.parameters())) print(f'Model device: {next(iter(smaller_model.parameters())).device}') print(f'Model dtype: {next(iter(smaller_model.parameters())).dtype}') print() if return_tokenizer: tokenizer = AutoTokenizer.from_pretrained('meta-llama/Llama-2-7b-hf', padding_side="right", use_fast=False, trust_remote_code=True, use_auth_token=True) return smaller_model, tokenizer return smaller_model
额外注意事项
- 如果需要在模型创建后转换 dtype,不要只调用
model.to(device),要明确指定 dtype:smaller_model = smaller_model.to(device, dtype=torch.bfloat16) - bfloat16 需要显卡算力 >=8.0(如A100、RTX30/40系列)支持,设备不支持时会自动 fallback 到 float32,属于正常行为
内容的提问来源于stack exchange,提问作者Charlie Parker
相关产品推荐
相关产品推荐

