请求协助将google/flan-t5-large PyTorch模型转换为GGUF格式
解决Flan-T5-Large转GGUF格式的KeyError问题
错误原因
Flan-T5属于Encoder-Decoder架构的T5系列模型,其配置文件用d_model表示嵌入维度,而llama.cpp的convert.py默认只读取Decoder-only模型(如Llama)的hidden_size字段,因此触发KeyError: 'hidden_size'。
解决方案
方案1:使用llama.cpp最新版的T5专用转换参数
- 更新llama.cpp到最新版本:
cd llama.cpp git pull make
- 执行转换时添加
--t5参数,明确指定模型类型:
python convert.py models/google/flan-t5-large --t5 --outfile flan-t5-large-f16.gguf --outtype f16
--outtype可指定量化类型,f16为半精度,也可选择q4_0、q4_1等量化格式。
方案2:手动修改convert.py适配T5配置(旧版llama.cpp适用)
如果无法更新到最新版,可修改convert.py的字段映射逻辑:
- 打开
llama.cpp/convert.py,找到loadHFTransformerJson函数(约203行位置)。 - 修改嵌入维度和注意力头数的读取逻辑:
# 原代码 n_embd = config["hidden_size"] n_head = config["num_attention_heads"] # 修改为 n_embd = config.get("hidden_size") or config.get("d_model") n_head = config.get("num_attention_heads") or config.get("num_heads")
- 保存修改后执行转换命令:
python convert.py models/ --outfile flan-t5-large-f16.gguf --outtype f16
测试模型
转换完成后,可使用llama.cpp的main工具验证模型加载:
./main -m flan-t5-large-f16.gguf -p "Translate to French: Hello world" -n 50
内容的提问来源于stack exchange,提问作者Akella Niranjan
相关产品推荐
相关产品推荐

