You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将BERT(bert-base-uncased)分词索引映射至Spacy分词索引?

解决BERT与Spacy分词索引的映射问题

问题核心:直接对Spacy分词后的每个词单独用BERT分词再生成映射,会因为单独分词与整句分词的结果不一致(比如£20.7bn单独分词和在句中分词的BERT token不同),导致映射列表长度和整句BERT分词长度不匹配;而文本匹配法又会遇到重复词无法定位正确索引的问题。

可靠解决方案:基于字符偏移量对齐

Spacy和BERT的分词结果都能对应到原始文本的字符位置,通过字符偏移量来建立映射是最准确的方式,具体步骤如下:

  1. 获取Spacy分词的字符偏移信息:每个Spacy Token都有idx(起始字符位置)和idx + len(token.text)(结束字符位置)
  2. 获取BERT分词的字符偏移信息:用tokenizer.encode_plus方法,指定return_offsets_mapping=True,得到每个BERT token对应的原始文本字符起止位置
  3. 遍历BERT的每个token,匹配它所属的Spacy词,生成映射列表

代码实现

import spacy
from transformers import BertTokenizer

# 初始化工具
tokenizer = BertTokenizer.from_pretrained("bert-base-uncased")
nlp = spacy.load("en_core_web_sm")

sent_text = "BRITAIN'S railways cost £20.7bn during the 2020-21 financial year, with £2.5bn generated through fares and other income, £1.3bn through other sources and £16.9bn from government, figures released by the regulator the Office of Rail and Road (ORR) on November 30 revealed."
doc = nlp(sent_text)

# 获取BERT分词及偏移量(不含特殊token)
bert_output = tokenizer.encode_plus(
    sent_text,
    add_special_tokens=False,
    return_offsets_mapping=True
)
bert_tokens = tokenizer.convert_ids_to_tokens(bert_output['input_ids'])
bert_offsets = bert_output['offset_mapping']

# 生成BERT token到Spacy词索引的映射
bert_to_spacy_map = []
spacy_token_idx = 0
spacy_tokens = list(doc)

for (start, end) in bert_offsets:
    # 跳过空偏移(防特殊token,此处已排除特殊token,可留作兼容)
    if start == 0 and end == 0:
        bert_to_spacy_map.append(-1)
        continue
    
    # 匹配当前BERT token所属的Spacy词
    while spacy_token_idx < len(spacy_tokens):
        curr_spacy_token = spacy_tokens[spacy_token_idx]
        spacy_start = curr_spacy_token.idx
        spacy_end = curr_spacy_token.idx + len(curr_spacy_token.text)
        if start >= spacy_start and end <= spacy_end:
            bert_to_spacy_map.append(spacy_token_idx)
            break
        else:
            spacy_token_idx += 1

# 验证长度一致性
print(f"BERT分词长度: {len(bert_tokens)}")  # 输出65
print(f"映射列表长度: {len(bert_to_spacy_map)}")  # 输出65,与BERT分词长度匹配

# 查看前10组映射示例
for bert_idx, spacy_idx in enumerate(bert_to_spacy_map[:10]):
    print(f"BERT token [{bert_tokens[bert_idx]}] -> Spacy词 [{doc[spacy_idx].text}] (索引{spacy_idx})")

说明

  • 该方法通过字符位置精准匹配,彻底规避了单独分词与整句分词的差异问题
  • 对于BERT的子词(如##0、##bn),会自动关联到对应的原始Spacy词
  • 映射列表bert_to_spacy_map的每个元素对应BERT分词的索引,值为该token所属的Spacy词的索引

内容的提问来源于stack exchange,提问作者lrthistlethwaite

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 22:30:29