You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求协助将google/flan-t5-large PyTorch模型转换为GGUF格式

解决Flan-T5-Large转GGUF格式的KeyError问题

错误原因

Flan-T5属于Encoder-Decoder架构的T5系列模型,其配置文件用d_model表示嵌入维度,而llama.cpp的convert.py默认只读取Decoder-only模型(如Llama)的hidden_size字段,因此触发KeyError: 'hidden_size'。

解决方案

方案1:使用llama.cpp最新版的T5专用转换参数

  1. 更新llama.cpp到最新版本:
cd llama.cpp
git pull
make
  1. 执行转换时添加--t5参数,明确指定模型类型:
python convert.py models/google/flan-t5-large --t5 --outfile flan-t5-large-f16.gguf --outtype f16

--outtype可指定量化类型,f16为半精度,也可选择q4_0、q4_1等量化格式。

方案2:手动修改convert.py适配T5配置(旧版llama.cpp适用)

如果无法更新到最新版,可修改convert.py的字段映射逻辑:

  1. 打开llama.cpp/convert.py,找到loadHFTransformerJson函数(约203行位置)。
  2. 修改嵌入维度和注意力头数的读取逻辑:
# 原代码
n_embd           = config["hidden_size"]
n_head           = config["num_attention_heads"]
# 修改为
n_embd           = config.get("hidden_size") or config.get("d_model")
n_head           = config.get("num_attention_heads") or config.get("num_heads")
  1. 保存修改后执行转换命令:
python convert.py models/ --outfile flan-t5-large-f16.gguf --outtype f16

测试模型

转换完成后,可使用llama.cpp的main工具验证模型加载:

./main -m flan-t5-large-f16.gguf -p "Translate to French: Hello world" -n 50

内容的提问来源于stack exchange,提问作者Akella Niranjan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 16:07:44