You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让BERT预测未收录新Token?以Scholz预测问题为例

问题:让BERT正确预测自定义Token(Scholz)

问题重现

正常工作示例

使用bert-base-uncased进行Mask预测时,已知词汇可正常被预测:

tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = BertForMaskedLM.from_pretrained('bert-base-uncased')
fill_mask_pipeline_pre = pipeline("fill-mask", model=model, tokenizer=tokenizer)
sentence_test = "Olaf is the chancellor of germany. [MASK] is the chancellor of germany."
prediction = fill_mask_pipeline_pre(sentence_test)[:3]

→ 第一个预测结果为“Olaf”,符合预期。

异常情况

当测试语句替换为未被BERT识别的专有名词时:

sentence_test = "Scholz is the chancellor of germany. [MASK] is the chancellor of germany."

期望预测结果为“Scholz”,但从未出现。原因是BertTokenizer.from_pretrained('bert-base-uncased')会将“Scholz”拆分为sc和##holz两个子词,模型无法输出完整的“Scholz”作为预测结果。

已尝试的无效操作:

  • 直接将“Scholz”添加至分词器词表
  • 添加Token后,用大量包含“Scholz”的文本微调模型

解决方案

要让模型能预测新增的完整Token,需确保分词器与模型的词表、嵌入层同步更新,步骤如下:

1. 正确添加自定义Token到分词器

添加Token时需指定add_prefix_space=True(适配uncased模型的分词逻辑),确保“Scholz”被识别为单个Token:

from transformers import BertTokenizer, BertForMaskedLM, pipeline

# 加载原始分词器
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
# 添加自定义Token,设置add_prefix_space保证分词一致性
tokenizer.add_tokens(["scholz"], add_prefix_space=True)
# 验证:tokenizer.tokenize("scholz") 应输出 ["scholz"]
# 保存更新后的分词器(可选,方便复用)
tokenizer.save_pretrained("./updated_bert_tokenizer")

2. 扩展模型的嵌入层维度

模型嵌入层的维度必须与分词器词表大小匹配,否则无法识别新增Token:

# 加载原始模型
model = BertForMaskedLM.from_pretrained('bert-base-uncased')
# 扩展嵌入层,自动初始化新增Token的权重
model.resize_token_embeddings(len(tokenizer))
# 保存更新后的模型(可选)
model.save_pretrained("./updated_bert_model")

3. 针对性微调模型

微调需保证训练数据中“Scholz”以完整Token形式被处理,且Mask策略贴合测试场景:

  • 训练数据需包含大量类似"Scholz is the chancellor of germany. [MASK] is the chancellor of germany."的句子,提高[MASK]替换为“Scholz”的比例
  • 使用DataCollatorForLanguageModeling处理数据,确保Mask操作覆盖目标专有名词位置
  • 调整微调参数:学习率设为5e-51e-4,训练轮次35轮(根据数据量调整),避免过拟合或欠拟合

4. 验证预测效果

加载更新后的分词器和模型测试:

tokenizer = BertTokenizer.from_pretrained("./updated_bert_tokenizer")
model = BertForMaskedLM.from_pretrained("./updated_bert_model")
fill_mask_pipeline = pipeline("fill-mask", model=model, tokenizer=tokenizer)

sentence_test = "Scholz is the chancellor of germany. [MASK] is the chancellor of germany."
prediction = fill_mask_pipeline(sentence_test)[:3]
print(prediction)

此时“scholz”应出现在预测结果中。

关键注意点

  • 必须验证分词器对“Scholz”的处理结果,确保其被识别为单个Token
  • 微调数据的上下文要与测试场景一致,数据质量优先于数量
  • 不可跳过resize_token_embeddings步骤,否则模型无法映射新增Token到对应的嵌入向量

内容的提问来源于stack exchange,提问作者Maximilian Huber

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 00:15:29