加载微调后LlamaForCausalLM状态字典时遇尺寸不匹配错误
解决LlamaForCausalLM加载微调模型时的参数尺寸不匹配问题
问题复现
RuntimeError: Error(s) in loading state_dict for LlamaForCausalLM:
size mismatch for model.embed_tokens.weight: copying a param with shape [128258, 4096] from checkpoint, the shape in current model is [128256, 4096].
size mismatch for lm_head.weight: copying a param with shape [128258, 4096] from checkpoint, the shape in current model is [128256, 4096].
问题核心是微调后的checkpoint词表大小(128258)与原Llama-3-8B-Instruct模型的词表大小(128256)不一致,导致参数形状不匹配。
解决方案
方法一:加载时指定匹配的词表大小
初始化模型时显式设置vocab_size为checkpoint的尺寸,确保模型结构与权重匹配:
from transformers import LlamaForCausalLM, LlamaConfig # 基于原模型配置,修改词表大小为checkpoint的128258 config = LlamaConfig.from_pretrained("meta-llama/Meta-Llama-3-8B-Instruct", vocab_size=128258) # 加载私有库中的微调模型 model = LlamaForCausalLM.from_pretrained( "你的私有库模型路径", config=config, device_map="auto" )
方法二:截取权重匹配原模型词表(仅适用于新增token未训练的场景)
如果微调时新增的2个token未在训练数据中出现、权重无有效更新,可以手动截取checkpoint中的对应参数:
import torch from transformers import LlamaForCausalLM # 加载微调后的checkpoint权重 checkpoint = torch.load("你的模型权重文件路径") # 截取前128256行权重,匹配原模型词表尺寸 checkpoint["model.embed_tokens.weight"] = checkpoint["model.embed_tokens.weight"][:128256, :] checkpoint["lm_head.weight"] = checkpoint["lm_head.weight"][:128256, :] # 加载原模型并替换权重 model = LlamaForCausalLM.from_pretrained( "meta-llama/Meta-Llama-3-8B-Instruct", state_dict=checkpoint, device_map="auto" )
方法三:修正微调流程(从根源避免问题)
检查微调阶段是否存在添加特殊token但未同步更新模型的操作:
- 若使用
tokenizer.add_tokens()扩展词表,必须同步执行model.resize_token_embeddings(len(tokenizer))更新嵌入层和输出层尺寸 - 合并模型时,确保使用扩展词表后的tokenizer和模型结构,而非原模型的默认配置
内容的提问来源于stack exchange,提问作者DigiSpocDeera
相关产品推荐
相关产品推荐

