加载HuggingFace模型执行掩码语言建模任务的正确方式及警告处理
问题场景
我在测试不同模型在掩码语言建模任务中的表现,使用的提示语是:
prompt = "The Milky Way is a [MASK] galaxy"
需要获取模型对掩码token的预测结果。
模型加载代码
我用以下代码加载掩码语言建模专用模型:
from transformers import AutoModelForMaskedLM, AutoTokenizer model = AutoModelForMaskedLM.from_pretrained('bert-base-cased') model.eval() tokenizer = AutoTokenizer.from_pretrained('bert-base-cased', truncation=True)
出现的警告
运行后出现如下警告:
Some weights of the model checkpoint at bert-base-uncased were not used when initializing BertForMaskedLM: ['cls.seq_relationship.bias', 'bert.pooler.dense.bias', 'cls.seq_relationship.weight', 'bert.pooler.dense.weight']
已有参考解答
在HuggingFace社区的讨论中,有类似问题但只涉及['cls.seq_relationship.weight', 'cls.seq_relationship.bias']两组权重,对应的解答是:
It tells you that by loading the bert-base-uncased checkpoint in the BertForMaskedLM architecture, you're dropping two weights: ['cls.seq_relationship.weight', 'cls.seq_relationship.bias'].
These are the weights used for next-sentence prediction, which aren't necessary for Masked Language Modeling.
If you're only interested in doing masked language modeling, then you can safely disregard this warning.
我的疑问
但我加载模型时还会丢弃['bert.pooler.dense.bias', 'bert.pooler.dense.weight']这两组权重,不确定在未微调模型的情况下,丢弃这些权重会不会影响掩码语言建模的性能。另外,如果用model = AutoModel.from_pretrained('bert-base-cased')加载模型不会有警告,但这类模型无法用于预测掩码token。
请问正确的做法是继续用AutoModelForMaskedLM加载模型并忽略警告(假设这些权重对掩码任务无用),还是有其他解决方法?
直接忽略警告即可,不会影响掩码语言建模性能
cls.seq_relationship相关权重是用来做下一句预测任务的,完全和掩码语言建模无关,丢弃不影响。bert.pooler相关权重是用来生成整个句子的向量表示(比如用于分类任务),掩码语言建模只依赖Transformer编码器输出的token级表示,再经过cls.predictions层(掩码预测专用头)来预测[MASK],所以pooler的权重根本不会被调用,丢弃后完全不影响掩码预测的结果。
为什么用AutoModel加载没有警告但不能用?
AutoModel只加载了BERT的基础编码器结构,没有附带掩码语言建模专用的cls.predictions头,所以确实无法完成[MASK]的预测任务,必须用AutoModelForMaskedLM才能加载带预测头的模型结构。额外优化:消除警告的小技巧
如果实在不想看到警告,可以在加载模型时添加ignore_mismatched_sizes=True参数,明确告诉Transformers库忽略权重不匹配的情况:model = AutoModelForMaskedLM.from_pretrained('bert-base-cased', ignore_mismatched_sizes=True)
内容的提问来源于stack exchange,提问作者Penguin

