You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

加载微调后LlamaForCausalLM状态字典时遇尺寸不匹配错误

解决LlamaForCausalLM加载微调模型时的参数尺寸不匹配问题

问题复现

RuntimeError: Error(s) in loading state_dict for LlamaForCausalLM:
size mismatch for model.embed_tokens.weight: copying a param with shape [128258, 4096] from checkpoint, the shape in current model is [128256, 4096].
size mismatch for lm_head.weight: copying a param with shape [128258, 4096] from checkpoint, the shape in current model is [128256, 4096].

问题核心是微调后的checkpoint词表大小(128258)与原Llama-3-8B-Instruct模型的词表大小(128256)不一致,导致参数形状不匹配。


解决方案

方法一:加载时指定匹配的词表大小

初始化模型时显式设置vocab_size为checkpoint的尺寸,确保模型结构与权重匹配:

from transformers import LlamaForCausalLM, LlamaConfig

# 基于原模型配置,修改词表大小为checkpoint的128258
config = LlamaConfig.from_pretrained("meta-llama/Meta-Llama-3-8B-Instruct", vocab_size=128258)
# 加载私有库中的微调模型
model = LlamaForCausalLM.from_pretrained(
    "你的私有库模型路径",
    config=config,
    device_map="auto"
)

方法二:截取权重匹配原模型词表(仅适用于新增token未训练的场景)

如果微调时新增的2个token未在训练数据中出现、权重无有效更新,可以手动截取checkpoint中的对应参数:

import torch
from transformers import LlamaForCausalLM

# 加载微调后的checkpoint权重
checkpoint = torch.load("你的模型权重文件路径")

# 截取前128256行权重,匹配原模型词表尺寸
checkpoint["model.embed_tokens.weight"] = checkpoint["model.embed_tokens.weight"][:128256, :]
checkpoint["lm_head.weight"] = checkpoint["lm_head.weight"][:128256, :]

# 加载原模型并替换权重
model = LlamaForCausalLM.from_pretrained(
    "meta-llama/Meta-Llama-3-8B-Instruct",
    state_dict=checkpoint,
    device_map="auto"
)

方法三:修正微调流程(从根源避免问题)

检查微调阶段是否存在添加特殊token但未同步更新模型的操作:

  • 若使用tokenizer.add_tokens()扩展词表,必须同步执行model.resize_token_embeddings(len(tokenizer))更新嵌入层和输出层尺寸
  • 合并模型时,确保使用扩展词表后的tokenizer和模型结构,而非原模型的默认配置

内容的提问来源于stack exchange,提问作者DigiSpocDeera

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 02:14:54