创建仅保留部分词汇的HF分词器后测试失败,求排查原因
问题原因及解决方法
核心原因:GPT2Tokenizer的BPE机制不触发unk_token
GPT2Tokenizer基于字节对编码(BPE)设计,默认没有unk_token逻辑——它会将任何未知文本拆分为字节级别的子词(哪怕是单个字符的字节),因此无论你怎么缩小词汇表,它都不会主动输出unk_token,自然不会把未知词汇替换成"the"。
另外你可能还忽略了两个细节:
- 仅设置
unk_token = "the"但未同步更新unk_token_id,导致分词器无法正确映射unk_token到对应的ID; - 深拷贝原分词器后,内部的BPE合并规则还是原大词汇表的规则,和新的小词汇表不匹配,分词逻辑依然沿用旧规则。
解决步骤
1. 重写分词逻辑,强制替换未知token
在得到分词后的token列表后,手动检查每个token是否在新词汇表中,不在的就替换为"the",再进行编码:
def get_tokenizer_with_subset_of_vocab(original_tokenizer, fraction=0.1): # 深拷贝原分词器 tokenizer = deepcopy(original_tokenizer) # 筛选非特殊词汇并随机采样 non_special_tokens = [tok for tok in tokenizer.vocab if tok not in tokenizer.all_special_tokens] sampled_tokens = random.sample(non_special_tokens, int(len(non_special_tokens)*fraction)) # 构建新词汇表:采样词汇+原特殊词汇 new_vocab = {tok: idx for idx, tok in enumerate(sampled_tokens + tokenizer.all_special_tokens)} # 更新双向映射 tokenizer.ids_to_tokens = {idx: tok for tok, idx in new_vocab.items()} tokenizer.vocab = new_vocab # 确保unk_token"the"在词汇表中并设置对应ID tokenizer.unk_token = "the" if "the" not in tokenizer.vocab: tokenizer.vocab["the"] = len(tokenizer.vocab) tokenizer.ids_to_tokens[len(tokenizer.ids_to_tokens)] = "the" tokenizer.unk_token_id = tokenizer.vocab["the"] # 重写分词方法,替换未知token original_tokenize = tokenizer._tokenize def custom_tokenize(text): tokens = original_tokenize(text) return [tok if tok in tokenizer.vocab else tokenizer.unk_token for tok in tokens] tokenizer._tokenize = custom_tokenize return tokenizer
2. 调整单元测试逻辑
确保测试时验证稀有词汇被替换为"the":
def _test0_does_hacky_fraction_tokenizer_work(): from transformers import GPT2Tokenizer original_tokenizer = GPT2Tokenizer.from_pretrained("gpt2") small_tokenizer = get_tokenizer_with_subset_of_vocab(original_tokenizer, fraction=0.01) test_text = "This is a test sentence with rare words like xyzzypqrst" encoded = small_tokenizer.encode(test_text, return_tensors="pt") decoded = small_tokenizer.decode(encoded[0], skip_special_tokens=True) # 断言稀有词汇已被替换 assert "xyzzypqrst" not in decoded assert "the" in decoded
3. 额外注意事项
- 必须确保"the"存在于新词汇表中,否则替换时会出现新的未知token;
- 重写
_tokenize方法会改变分词器默认行为,需确保不影响其他功能; - 如果不需要子词拆分,可以考虑使用基于词表的分词器(如BertWordPieceTokenizer),这类分词器更容易触发unk_token。
内容的提问来源于stack exchange,提问作者Charlie Parker
相关产品推荐
相关产品推荐

