You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过配置将Hugging Face小模型(如小型Llama2)转为bfloat16?

解决配置设置bfloat16无效的问题

你的问题核心是修改config的时机错误:你先调用AutoModelForCausalLM.from_config(config)创建了模型,之后才去设置config.torch_dtype,此时模型已经完成初始化,修改配置根本不会影响已创建的模型参数 dtype。

正确处理方式

在调用from_config创建模型之前,就把torch_dtype配置好,这样模型初始化时会直接用指定的 dtype 生成参数。

修改后的代码示例

def get_smaller_llama2(hidden_size : int = 2048, 
                       num_hidden_layers : int = 12, 
                       return_tokenizer: bool = False, 
                       verbose : bool = False,
                       ):
    config = AutoConfig.from_pretrained("meta-llama/Llama-2-7b-hf")
    # 先调整模型结构参数
    config.hidden_size = hidden_size
    config.num_hidden_layers = num_hidden_layers
    # 关键步骤:在创建模型前设置torch_dtype
    config.torch_dtype = torch.bfloat16 if (torch.cuda.is_available() and torch.cuda.get_device_capability(torch.cuda.current_device())[0] >= 8) else torch.float32
    
    # 可选:在from_config时额外指定torch_dtype,双重保险
    smaller_model = AutoModelForCausalLM.from_config(config, torch_dtype=config.torch_dtype)
    
    device = torch.device(f"cuda:{0}" if torch.cuda.is_available() else "cpu")
    # 模型已为目标dtype,直接移动到设备即可
    smaller_model = smaller_model.to(device)
    
    print(f'Model dtype: {next(iter(smaller_model.parameters())).dtype}')
    if verbose:
        print(f'config: {config}')
        print("Smaller number of parameters:", sum(p.numel() for p in smaller_model.parameters()))
        print(f'Model device: {next(iter(smaller_model.parameters())).device}')
        print(f'Model dtype: {next(iter(smaller_model.parameters())).dtype}')
        print()
    if return_tokenizer:
        tokenizer = AutoTokenizer.from_pretrained('meta-llama/Llama-2-7b-hf', padding_side="right", use_fast=False, trust_remote_code=True, use_auth_token=True)
        return smaller_model, tokenizer
    return smaller_model

额外注意事项

  • 如果需要在模型创建后转换 dtype,不要只调用model.to(device),要明确指定 dtype:smaller_model = smaller_model.to(device, dtype=torch.bfloat16)
  • bfloat16 需要显卡算力 >=8.0(如A100、RTX30/40系列)支持,设备不支持时会自动 fallback 到 float32,属于正常行为

内容的提问来源于stack exchange,提问作者Charlie Parker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 22:05:10