如何让BERT预测未收录新Token?以Scholz预测问题为例
问题:让BERT正确预测自定义Token(Scholz)
问题重现
正常工作示例
使用bert-base-uncased进行Mask预测时,已知词汇可正常被预测:
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased') model = BertForMaskedLM.from_pretrained('bert-base-uncased') fill_mask_pipeline_pre = pipeline("fill-mask", model=model, tokenizer=tokenizer) sentence_test = "Olaf is the chancellor of germany. [MASK] is the chancellor of germany." prediction = fill_mask_pipeline_pre(sentence_test)[:3]
→ 第一个预测结果为“Olaf”,符合预期。
异常情况
当测试语句替换为未被BERT识别的专有名词时:
sentence_test = "Scholz is the chancellor of germany. [MASK] is the chancellor of germany."
期望预测结果为“Scholz”,但从未出现。原因是BertTokenizer.from_pretrained('bert-base-uncased')会将“Scholz”拆分为sc和##holz两个子词,模型无法输出完整的“Scholz”作为预测结果。
已尝试的无效操作:
- 直接将“Scholz”添加至分词器词表
- 添加Token后,用大量包含“Scholz”的文本微调模型
解决方案
要让模型能预测新增的完整Token,需确保分词器与模型的词表、嵌入层同步更新,步骤如下:
1. 正确添加自定义Token到分词器
添加Token时需指定add_prefix_space=True(适配uncased模型的分词逻辑),确保“Scholz”被识别为单个Token:
from transformers import BertTokenizer, BertForMaskedLM, pipeline # 加载原始分词器 tokenizer = BertTokenizer.from_pretrained('bert-base-uncased') # 添加自定义Token,设置add_prefix_space保证分词一致性 tokenizer.add_tokens(["scholz"], add_prefix_space=True) # 验证:tokenizer.tokenize("scholz") 应输出 ["scholz"] # 保存更新后的分词器(可选,方便复用) tokenizer.save_pretrained("./updated_bert_tokenizer")
2. 扩展模型的嵌入层维度
模型嵌入层的维度必须与分词器词表大小匹配,否则无法识别新增Token:
# 加载原始模型 model = BertForMaskedLM.from_pretrained('bert-base-uncased') # 扩展嵌入层,自动初始化新增Token的权重 model.resize_token_embeddings(len(tokenizer)) # 保存更新后的模型(可选) model.save_pretrained("./updated_bert_model")
3. 针对性微调模型
微调需保证训练数据中“Scholz”以完整Token形式被处理,且Mask策略贴合测试场景:
- 训练数据需包含大量类似
"Scholz is the chancellor of germany. [MASK] is the chancellor of germany."的句子,提高[MASK]替换为“Scholz”的比例 - 使用
DataCollatorForLanguageModeling处理数据,确保Mask操作覆盖目标专有名词位置 - 调整微调参数:学习率设为5e-51e-4,训练轮次35轮(根据数据量调整),避免过拟合或欠拟合
4. 验证预测效果
加载更新后的分词器和模型测试:
tokenizer = BertTokenizer.from_pretrained("./updated_bert_tokenizer") model = BertForMaskedLM.from_pretrained("./updated_bert_model") fill_mask_pipeline = pipeline("fill-mask", model=model, tokenizer=tokenizer) sentence_test = "Scholz is the chancellor of germany. [MASK] is the chancellor of germany." prediction = fill_mask_pipeline(sentence_test)[:3] print(prediction)
此时“scholz”应出现在预测结果中。
关键注意点
- 必须验证分词器对“Scholz”的处理结果,确保其被识别为单个Token
- 微调数据的上下文要与测试场景一致,数据质量优先于数量
- 不可跳过
resize_token_embeddings步骤,否则模型无法映射新增Token到对应的嵌入向量
内容的提问来源于stack exchange,提问作者Maximilian Huber
相关产品推荐
相关产品推荐

