如何将BERT(bert-base-uncased)分词索引映射至Spacy分词索引?
解决BERT与Spacy分词索引的映射问题
问题核心:直接对Spacy分词后的每个词单独用BERT分词再生成映射,会因为单独分词与整句分词的结果不一致(比如£20.7bn单独分词和在句中分词的BERT token不同),导致映射列表长度和整句BERT分词长度不匹配;而文本匹配法又会遇到重复词无法定位正确索引的问题。
可靠解决方案:基于字符偏移量对齐
Spacy和BERT的分词结果都能对应到原始文本的字符位置,通过字符偏移量来建立映射是最准确的方式,具体步骤如下:
- 获取Spacy分词的字符偏移信息:每个Spacy Token都有
idx(起始字符位置)和idx + len(token.text)(结束字符位置) - 获取BERT分词的字符偏移信息:用
tokenizer.encode_plus方法,指定return_offsets_mapping=True,得到每个BERT token对应的原始文本字符起止位置 - 遍历BERT的每个token,匹配它所属的Spacy词,生成映射列表
代码实现
import spacy from transformers import BertTokenizer # 初始化工具 tokenizer = BertTokenizer.from_pretrained("bert-base-uncased") nlp = spacy.load("en_core_web_sm") sent_text = "BRITAIN'S railways cost £20.7bn during the 2020-21 financial year, with £2.5bn generated through fares and other income, £1.3bn through other sources and £16.9bn from government, figures released by the regulator the Office of Rail and Road (ORR) on November 30 revealed." doc = nlp(sent_text) # 获取BERT分词及偏移量(不含特殊token) bert_output = tokenizer.encode_plus( sent_text, add_special_tokens=False, return_offsets_mapping=True ) bert_tokens = tokenizer.convert_ids_to_tokens(bert_output['input_ids']) bert_offsets = bert_output['offset_mapping'] # 生成BERT token到Spacy词索引的映射 bert_to_spacy_map = [] spacy_token_idx = 0 spacy_tokens = list(doc) for (start, end) in bert_offsets: # 跳过空偏移(防特殊token,此处已排除特殊token,可留作兼容) if start == 0 and end == 0: bert_to_spacy_map.append(-1) continue # 匹配当前BERT token所属的Spacy词 while spacy_token_idx < len(spacy_tokens): curr_spacy_token = spacy_tokens[spacy_token_idx] spacy_start = curr_spacy_token.idx spacy_end = curr_spacy_token.idx + len(curr_spacy_token.text) if start >= spacy_start and end <= spacy_end: bert_to_spacy_map.append(spacy_token_idx) break else: spacy_token_idx += 1 # 验证长度一致性 print(f"BERT分词长度: {len(bert_tokens)}") # 输出65 print(f"映射列表长度: {len(bert_to_spacy_map)}") # 输出65,与BERT分词长度匹配 # 查看前10组映射示例 for bert_idx, spacy_idx in enumerate(bert_to_spacy_map[:10]): print(f"BERT token [{bert_tokens[bert_idx]}] -> Spacy词 [{doc[spacy_idx].text}] (索引{spacy_idx})")
说明
- 该方法通过字符位置精准匹配,彻底规避了单独分词与整句分词的差异问题
- 对于BERT的子词(如
##0、##bn),会自动关联到对应的原始Spacy词 - 映射列表
bert_to_spacy_map的每个元素对应BERT分词的索引,值为该token所属的Spacy词的索引
内容的提问来源于stack exchange,提问作者lrthistlethwaite
相关产品推荐
相关产品推荐

